Balasubramanian Raman

dblp:57/6641 · also Raman Balasubramanian · DBLP profile ↗
← Back
127ranked-venue papers
0as first author
56since 2021 · last 2026
0000-0001-6277-6267ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 70 · 26 since 2021Artificial intelligence and machine learning · 41 · 27 since 2021Computer networks · 10 · 6 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Systems, architecture and hardware · 3 · 1 since 2021Security and privacy · 2Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 DTMIR-Pro: Domain Translation with Prompt-based Latent-Space Generalization for Multi-Weather Image Restoration
abstract
Multi-weather image restoration seeks to recover scene visibility under rainy, snowy, and hazy conditions, thereby enhancing high-level vision tasks. Existing methods typically train on combined datasets with single-type weather degradations, limiting their generalization to real-world scenarios involving mixed degradations. Domain translation has emerged as a viable solution by generating diverse weather-degraded variants of the same scene. However, current approaches require separate models for each degradation type, resulting in increased system complexity. To address this, we propose DTMIR-Pro, a prompt-based domain translation framework with latent space generalization for multi-weather image restoration. A single trainable network performs multi-domain translation using domain-adaptive prompts and dynamic kernel selection via a proposed Dynamic Multi-Head Attention block to learn diverse degradation patterns. The restoration network takes translated outputs and employs a Multi-Weather Fusion Block with global-local feature streams to capture complex degradations. Furthermore, we introduce a Similarity-Based Encoder Routing mechanism to transfer domain-specific features from the translation encoder to the restoration stage. Extensive experiments on both synthetic and real-world weather-degraded datasets demonstrate the effectiveness and generalizability of the proposed method. The code is made available at https://github.com/AshutoshKulkarni4998/DTMIR-Pro.
Ashutosh Kulkarni, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Balasubramanian Raman
WACV5
2026 Hierarchical fractal reservoirs: Unsupervised multiscale symbolic graphs for long-horizon prediction of chaotic flows
Shradhdha Agrawal, Shreyas Rajesh Dahale, Sutirtha Ghosh, Balasubramanian Raman
Neurocomputing5
2026 Adaptive group-weighted convolutions: Enhancing equivariant networks via learnable symmetry importance maps
Kanishk Sharma, Balasubramanian Raman
Neurocomputing3
2026 Editorial to special issue on selected extended works from 9th international conference on computer vision & image processing (CVIP) 2024
Mohan Kankanhalli, Balasubramanian Raman, M. Subrahmanyam 0001, Jagadeesh Kakarla, Sambit Bakshi
Image Vis. Comput.2
2026 DAIRNet: Degradation-aware All-in-one Image Restoration Network with cross-channel feature interaction
Amit Monga, Hemkant Nehete, Tharun Kumar Reddy, Balasubramanian Raman
J. Vis. Commun. Image Represent.4
2026 Semantic spherical mixup: geometry-aware data augmentation in latent manifolds for parameter-efficient language model adaptation
Balasubramanian Raman
Knowl. Inf. Syst.2
2026 Thalamocortical burst gating enables event-locked memory
Balasubramanian Raman
Knowl. Based Syst.2
2026 Topologically informed echo state networks via poincaré return maps for chaotic time-series
Sutirtha Ghosh, Balasubramanian Raman
Neural Networks4
2026 ASM-DiffConvNet: Physics-Guided Difference Convolution Network for Single-Image Restoration
abstract
This work proposes a physics-guided unified deep learning architecture for single image restoration targeting dehazing, deraining, and low-light enhancement. The architecture first estimates the transmission map and airlight under an atmospheric scattering model, and then refines the result with a grayscale prior. A DiffConv feature extractor is proposed which combines vanilla and difference convolutions with a Laplacian branch (to capture high-frequency features). During inference, its branches are re-parameterized into a single kernel for reducing computational complexity. The grayscale prior replaces the Y channel in the YCbCr space to suppress noise and color artifacts, while a refinement stage uses Spatial Feature Transform (SFT) to inject structural features from this grayscale prior into the RGB domain. Experiments on standard benchmarks show consistent improvements in PSNR and SSIM at lower computational cost.
Hemkant Nehete, Amit Monga, Tharun Kumar Reddy, Balasubramanian Raman
IEEE Signal Process. Lett.4
2026 PHtNN: Prediction of Heart Disease Risk Using Twin Neural Network
abstract
Cardiovascular disease is one of the primary causes of increasing mortality rates globally, spanning various types of ailments. Aside from a healthy lifestyle, prediction, prognostication, and early diagnosis can all contribute to lower mortality rates. The massive variation in the economic growth and development of countries worldwide, accompanied by the irregular availability of medical experts and radiologists, is a major impediment to early diagnosis. Researchers are working on prediction systems that will aid doctors and radiologists in prognostication and assessment by providing diagnostics to the human race without regard to geographical, economic, or financial inequalities. The use of a computational intelligence-based medical imaging prediction system to either prognosticate or detect and further diagnose the disease is becoming more popular. In this work, a computational intelligence-based prediction system, PHtNN , for heart disease diagnosis has been proposed. PHtNN uses the Multiple Factor Analysis (MFA) to extract features from the heart disease multi-datasets, VA Long Beach, Switzerland, Hungarian, Cleveland, and Z-Alizadeh Sani, and train the model by using twin neural network. The system, PHtNN , is validated using the hold-out validation scheme with a ratio of 3:1. Experimental results reveal that PHtNN outperforms several previous baseline approaches in terms of accuracy and improves the system’s efficiency; as a result, it can assist medical experts in diagnosing cardiac patients.
Ankur Gupta 0002, Rahul Kumar 0003, Balasubramanian Raman, Harkirat Singh Arora
ACM Trans. Intell. Syst. Technol.3
2025 Frozen in Time: Parameter-Efficient Time Series Transformers via Reservoir-Induced Feature Expansion and Fixed Random Dynamics
abstract
Transformers are the de-facto choice for sequence modelling, yet their quadratic self-attention and weak temporal bias can make long-range forecasting both expensive and brittle. We introduce FreezeTST, a lightweight hybrid that interleaves frozen random-feature (reservoir) blocks with standard trainable Transformer layers. The frozen blocks endow the network with rich nonlinear memory at no optimisation cost; the trainable layers learn to query this memory through self-attention. The design cuts trainable parameters and also lowers wall-clock training time, while leaving inference complexity unchanged. On seven standard long-term forecasting benchmarks, FreezeTST consistently matches or surpasses specialised variants such as Informer, Autoformer, and PatchTST; with substantially lower compute. Our results show that embedding reservoir principles within Transformers offers a simple, principled route to efficient long-term time-series prediction. In the interest of reproducibility, we release our implementation at github.com/deepdyn/Frozen-Transformers and provide a full technical appendix (proofs, ablations, hyperparameters) in the arXiv preprint at DOI: 10.48550/arXiv:2508.18130.
Anupriya Dey, Balasubramanian Raman
ECAI4
2025 SFCola-Net: Spatial-Frequency Collaborative Attention Network for Image Restoration
abstract
Ambient factors such as fog, rain, and haze significantly degrade images, adversely impacting the performance of computer vision systems in autonomous vehicles. Numerous image restoration architectures have been developed to address this challenge, with local and non-local attention-based methods gaining significant traction for their promising results. However, existing approaches predominantly focus on either local or non-local attention mechanisms, which limits their ability to comprehensively restore the image quality. Additionally, nonlocal attention methods, while leveraging the self-similarity of natural images, often struggle to accurately model long-range dependencies due to excessive degradation and limited information in spatial domain. To overcome these limitations, this work proposes SFCola-Net, a novel spatial-frequency collaborative attention multiscale network that combines local and non-local features for effective image restoration. The proposed SFColaNet architecture restores the image features in both spatial and frequency domains, effectively handling areas with complex textures. Furthermore, patch-wise non-local attention model is proposed, enabling the network to capture long-range features. The proposed network demonstrates significant improvements across various image restoration tasks, including image dehazing, deraining, and low light enhancement, thereby enhancing the robustness and reliability of computer vision systems in autonomous vehicles.
Hemkant Nehete, Amit Monga, Tharun Kumar Reddy, Balasubramanian Raman
IJCNN4
2025 Multimodal Interpretable Depression Analysis Using Visual, Physiological, Audio and Textual Data
abstract
Motivated by depression's significant impact on global health, this work proposes MultiDepNet, a novel multi-modal interpretable depression detection system integrating visual, physiological, audio, and textual data. Through ded-icated feature extraction methods (MTCNN for video, TS-CAN for physiological, ResNet-18 for audio, and RoBERTa for text modalities) and a strategic fusion of modality-specific networks including CNN-RNN, Transformer, MLP, and ResNet-18, it achieves significant advancements in depression detection. Its performance, evaluated across four benchmark datasets (AVEC 2013, AVEC 2014, DAIC, and E-DAIC), demonstrates average MAE of 5.64, RMSE of 7.15, accuracy of 74.19%, precision of 0.7373, re-call of 0.7378, and F1 of 0.7376. It also implements a Multiviz-based interpretability mechanism that computes each modality's contribution to the model's performance. The results reveal the visual modality to be the most signifi-cant, contributing 37.88% towards depression detection.
Puneet Kumar 0003, Shreshtha Misra, Zhuhong Shao, Balasubramanian Raman
WACV5
2025 DFTQuake: Tripartite Fourier attention and dendrite network for real-time early prediction of earthquake magnitude and peak ground acceleration
Anushka Joshi, Nithya Reddy Vedium, Balasubramanian Raman
Eng. Appl. Artif. Intell.3
2025 Brain-lateralized reservoirs: Neuro-anatomically constrained echo state networks for chaotic and physiological signals
Shreyas Rajesh Dahale, Yash, Balasubramanian Raman
Knowl. Based Syst.5
2025 Interpretable image emotion recognition: A domain adaptation approach using facial expressions
Puneet Kumar 0003, Balasubramanian Raman
Multim. Tools Appl.2
2025 VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition
abstract
This paper proposes a multimodal emotion recognition system, VIsual Spoken Textual Additive Net (VISTANet), to classify emotions reflected by input containing image, speech, and text into discrete classes. A new interpretability technique, K-Average Additive exPlanation (KAAP), has been developed that identifies important visual, spoken, and textual features leading to predicting a particular emotion class. The VISTANet fuses information from image, speech, and text modalities using a hybrid of intermediate and late fusion. It automatically adjusts the weights of their intermediate outputs while computing the weighted average. The KAAP technique computes the contribution of each modality and corresponding features toward predicting a particular emotion class. To mitigate the insufficiency of multimodal emotion datasets labelled with discrete emotion classes, we have constructed the IIT-R MMEmoRec dataset consisting of images, corresponding speech and text, and emotion labels (‘angry,’ ‘happy,’ ‘hate,’ and ‘sad’). The VISTANet has resulted in an overall emotion recognition accuracy of 80.11% on the IIT-R MMEmoRec dataset using visual, spoken, and textual modalities, outperforming single or dual-modality configurations. The code and data can be accessed at github.com/MIntelligence-Group/MMEmoRec.
Puneet Kumar 0003, Sarthak Malik, Balasubramanian Raman
IEEE Trans. Affect. Comput.3
2024 RefMOS: A Robust Referred Moving Object Segmentation framework based on text query
abstract
Referred Moving object segmentation is a very challenging task in automated video surveillance applications as it requires additional information to learn about object representation referred by natural language expression. In segmenting specific moving objects targeted by a text, suppressing other moving as well as stationary objects is a crucial task. A better context needs to be learned where linguistic, spatial, and temporal features need to be taken into account. In this work, we have proposed a robust referred moving object segmentation (RefMOS) framework to capture moving objects referred by text query. Most of the earlier state-of-the-art methods exploit a different type of supervision by treating video frames as images but lack temporal information during processing. In this work, we have proposed an inter-frame movement detector (IFCD) module, which extracts the movement information between the consecutive frames and helps integrate temporal information with spatial visual features. Language embedding is utilized to capture the information of referred moving objects in the text by extracting linguistic features from a pre-trained language model, i.e., BERT. Furthermore, the cross-entropy loss and SGD optimizer are used to train the network. Our RefMOS framework competes with the state-of-the-art approaches and achieves 48.6 mean IOU on the ref-DAVIS 17 dataset.
Prafulla Saxena, Susim Mukul Roy, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Balasubramanian Raman
AVSS6
2024 Neuro-Emotional Mapping of Human Emotions via EEG Signals
abstract
This study presents a novel approach to emotion recognition using Electroencephalogram (EEG) signals and machine learning, with a focus on deciphering the neural signatures of human emotions. We utilized EmoNeuroDB, a dataset of 720 EEG recordings from 40 participants to explore the efficacy of Random Forest classifiers, optimized through grid search and evaluated via cross-validation, to classify six basic emotions. Our methodology integrates comprehensive signal processing techniques to extract meaningful features from the EEG signals, facilitating the accurate classification of emotional states. The experiment design, involving a sequence of emotion elicitation and relaxation phases, coupled with advanced data analysis, underscores the potential of combining EEG and machine learning for real-time emotion recognition. The findings contribute to the understanding of emotional processing in the brain, offering implications for enhancing human-computer interaction and psychological research.
Sutirtha Ghosh, Balasubramanian Raman
FG5
2024 Fusing Multimodal Streams for Improved Group Emotion Recognition in Videos
Piyush Dhamdhere, Balasubramanian Raman
ICPR (21)3
2024 Differentially Private Spiking Variational Autoencoder
Srishti Yadav, Anshul Pundhir, Tanish Goyal, Balasubramanian Raman, Sanjeev Kumar 0001
ICPR (15)4
2024 Integrating Physiological Signals with Dynamical Attention Networks for Personality Trait Analysis
abstract
In the expanding field of personality psychology, the assessment and analysis of personality traits are of paramount importance. This paper introduces a novel deep learning network designed to predict the OCEAN personality traits (openness, conscientiousness, extraversion, agreeableness, and neuroticism), leveraging physiological signals such as electroencephalogram (EEG), electrocardiogram (ECG), and galvanic skin response (GSR). The model incorporates convolutional layers and is augmented with specifically engineered attention modules that meticulously process EEG, ECG, and GSR signals. By integrating a novel blend of component-wise attention alongside gating mechanisms, the model significantly enhances feature selection and representation. This is further complemented by the inclusion of a temporal attention module that emphasizes significant timesteps in the data, ensuring a more dynamic and responsive analysis. The model’s architecture culminates in a robust and integrated feature representation, encapsulating the nuanced interplay of the processed physiological signals. This composite representation is subsequently flattened and fed into a dense module, designed for the accurate prediction of personality traits. Our approach not only marks a significant stride in understanding the intricate relationship between physiological signals and personality traits but also lays the groundwork for diverse applications. These range from advancements in psychological research and personalized health monitoring to refining human-computer interaction paradigms. By bridging the gap between physiological data and psychological traits, this research contributes to the burgeoning field of affective computing and offers a new vista in the personalized assessment of individual traits.
Richa, Kishore Babu Nampalle, Balasubramanian Raman
IJCNN5
2024 Towards Ethical Dermatology: Mitigating Bias in Skin Condition Classification
abstract
Deep neural networks have proven to be highly effective in efficiently handling various medical tasks, including classification and segmentation in the healthcare domain. Specifically, convolutional neural networks (CNNs) have gained significant popularity in aiding dermatologists with the diagnosis of skin lesions. The CNNs are found to surpass the performance of dermatologists and reduce the need for manual intervention. However, a potential drawback of deep neural networks is their susceptibility to biased predictions towards minority subgroups within the dataset they are trained on, as they effectively learn from the underlying data distribution. In dermatology datasets, researchers have highlighted the negative impact caused by the under-representation of individuals with dark skin tones, which leads to social and medical discrimination. This data imbalance poses challenges when deploying deep neural networks on a larger scale, as it can result in unfair outcomes and exacerbate existing biases. To address this issue, we utilize a debiasing approach based on variational autoencoders that effectively mitigate the data bias present in deep neural networks. To evaluate the effectiveness of our proposed approach, we conducted experiments using a recent benchmark dataset known as Fitzpatrick-17k. This dataset consists of clinical images representing various skin conditions, each labeled for the disease class and Fitzpatrick skin tone. By leveraging the inherent latent data distribution across different classes, we demonstrated significant improvements in mitigating the bias related to skin tone within the dermatology dataset. As a result, we observed an enhanced classification rate for the under-represented minority classes. To validate our proposed approach, we employed both quantitative and qualitative measures. Our model not only outperformed existing benchmark approaches for skin condition classification but also effectively addressed the detrimental effects caused by inherent bias in the classifier. Furthermore, we employed the Uniform Manifold Approximation and Projection (UMAP) technique to visually showcase the robust learning of the data distribution by our model, providing evidence of its generalization capabilities. We will share our source code publicly for reproducibility to facilitate further research and validation of our approach.
Anshul Pundhir, Sanchit Verma, Balasubramanian Raman
IJCNN3
2024 SOFIM: Stochastic Optimization Using Regularized Fisher Information Matrix
abstract
This paper introduces a new stochastic optimization method based on the regularized Fisher information matrix (FIM), named SOFIM, which can efficiently utilize the FIM to approximate the Hessian matrix for finding Newton’s gradient update in large-scale stochastic optimization of machine learning models. It can be viewed as a variant of natural gradient descent, where the challenge of storing and calculating the full FIM is addressed through making use of the regularized FIM and directly finding the gradient update direction via Sherman-Morrison matrix inversion. Additionally, like the popular Adam method, SOFIM uses the first moment of the gradient to address the issue of non-stationary objectives across mini-batches due to heterogeneous data. The utilization of the regularized FIM and Sherman-Morrison matrix inversion leads to the improved convergence rate with the same space and time complexities as stochastic gradient descent (SGD) with momentum. The extensive experiments on training deep learning models using several benchmark image classification datasets demonstrate that the proposed SOFIM outperforms SGD with momentum and several state-of-the-art Newton optimization methods in term of the convergence speed for achieving the pre-specified objectives of training and test losses as well as test accuracy.
Mrinmay Sen, A. K. Qin 0001, Gayathri C, Raghu Kishore N, Yen-Wei Chen 0001, Balasubramanian Raman
IJCNN6
2024 Unveiling Robustness of Spiking Neural Networks against Data Poisoning Attacks
abstract
Spiking Neural Networks (SNNs) are gaining attention as a potential evolution of Artificial Neural Networks, mimicking neural computing like the human brain. Known for their energy efficiency and sparse trigger event-driven operation in neuromorphic computing, SNNs’ resilience against adversarial attacks still needs to be explored. This study assesses large-scale SNNs in medical and non-medical datasets and reveals their inefficiency in medical image classification. We introduced three adversarial attacks on SNNs and observed a significant performance drop with increasing attack severity. We mainly proposed a lightweight SNN-based model that outperforms large-scale SNNs in medical image classification and also found robust against different adversarial attacks. We also introduced a novel metric, Attack Diversion Score, to quantify the performance divergence of SNNs during attacks. Our model, employing spatial learning through time, is memory and power-efficient, hence, suitable for computer-aided diagnosis. Using three datasets, our approach is validated against Spiking ResNet-18 and Spiking VGG-11 and found robust against different data poisoning attacks. We affirm the utility of our model through several quantitative and qualitative measures, which have also proven its effectiveness. The source code of our implementation will be publicly available to foster reproducibility and support future research.
Srishti Yadav, Anshul Pundhir, Balasubramanian Raman, Sanjeev Kumar 0001
IJCNN3
2024 Towards Engagement Prediction: A Cross-Modality Dual-Pipeline Approach using Visual and Audio Features
abstract
Engagement estimation is crucial for advancing natural human-computer interaction, allowing artificial agents to dynamically adjust their responses based on user engagement levels and creating more intuitive and immersive experiences. Despite advancements in automating real-time engagement estimation, challenges persist in real-world scenarios due to the complex nature of multi-modal human social signals. This paper proposes a novel cross-modality fusion-based methodology to address these challenges by leveraging multi-modal data. Our approach integrates visual and audio features, such as facial motion, acoustic characteristics, Contrastive Language-Image Pretraining (CLIP), and semantic embeddings. These features first pass through a transformer encoder, are then combined and processed through a cross-modal fusion mechanism, ensuring robust integration. The final integrated features are then used to predict engagement scores. This hierarchical and self-normalizing approach enhances the accuracy of engagement estimation by effectively capturing dependencies within and between modalities. The experiments are conducted on multimediate's NoXI and MPIIGroupInteraction datasets and the results demonstrates competitive performance in estimating engagement levels, addressing the complex, context-dependent nature of human engagement. Specifically, our approach achieves a Global Concordance Correlation Coefficient (CCC) score approximately (56.1%) higher than the baseline. This work contributes to developing more intelligent and responsive artificial systems, enhancing user experiences across various interactive applications.
Surbhi Madan, Abhinav Dhall, Balasubramanian Raman
ACM Multimedia5
2024 Interpretable multimodal emotion recognition using hybrid fusion of speech and image data
Puneet Kumar 0003, Sarthak Malik, Balasubramanian Raman
Multim. Tools Appl.3
2024 Towards improved U-Net for efficient skin lesion segmentation
Kishore Babu Nampalle, Anshul Pundhir, Pushpamanjari Ramesh Jupudi, Balasubramanian Raman
Multim. Tools Appl.4
2024 An integrated approach for prediction of magnitude using deep learning techniques
Anushka Joshi, Balasubramanian Raman, C. Krishna Mohan
Neural Comput. Appl.2
2024 A Deep Attention Model for Onsite Estimation of Earthquake Epicenter Distance and Magnitude
abstract
The onsite early warning techniques that issue earthquake alerts based on the seismic response of single stations have proven to be quite successful in detecting damage. The magnitude of an earthquake and the epicentral distance are vital parameters for accessing the intensity of destruction at an observation point during an earthquake. This article presents a novel earthquake early warning (EEW) system model designed to predict epicentral distance. The proposed model synergizes the strengths of deep learning and shallow machine learning (ML) techniques, offering a new perspective on seismic event prediction. Specifically, the study introduces an architecture that uses a long short-term memory (LSTM) network with two attention mechanisms to extract high-level features from the initial 3 s of the primary waveform. These attention mechanisms are built for time- and feature-based dependencies. Global site-related features such as shear-wave velocities at various depths and station coordinates are also incorporated, enhancing the model’s predictive capacity. Following this, the scalable ML algorithm XGBoost is applied. The fusion of deep and shallow learning methods applied in the onsite prediction of epicentral distance makes this a significant contribution to the early detection of epicenter distance, azimuth, and focal depth. The next step is the single-station magnitude detection from the predicted epicenter distance, azimuth, focal depth, and other magnitude-dependent parameters. The study shows remarkable results for magnitude detection. The findings indicate the potential for improved prediction of earthquake parameters, contributing toward the ongoing goal of enhancing EEW systems and reducing the destructive impacts of seismic events.
Anushka Joshi, Balasubramanian Raman
IEEE Trans. Geosci. Remote. Sens.3
2023 Classification of Hard and Soft Wheat Species Using Hyperspectral Imaging and Machine Learning Models
Nitin Tyagi, Balasubramanian Raman, Neerja Mittal Garg
ICONIP (14)2
2023 Zero-shot learning based cross-lingual sentiment analysis for sanskrit text with insufficient labeled data
Puneet Kumar 0003, Kshitij Pathania, Balasubramanian Raman
Appl. Intell.3
2023 Medical image security and authenticity via dual encryption
Kishore Babu Nampalle, Shriansh Manhas, Balasubramanian Raman
Appl. Intell.3
2023 Brain strokes classification by extracting quantum information from CT scans
Anjali Gautam, Balasubramanian Raman
Multim. Tools Appl.2
2023 Affective Feedback Synthesis Towards Multimodal Text and Image Data
abstract
In this article, we have defined a novel task of affective feedback synthesis that generates feedback for input text and corresponding images in a way similar to humans responding to multimodal data. A feedback synthesis system has been proposed and trained using ground-truth human comments along with image–text input. We have also constructed a large-scale dataset consisting of images, text, Twitter user comments, and the number of likes for the comments by crawling news articles through Twitter feeds. The proposed system extracts textual features using a transformer-based textual encoder. The visual features have been extracted using a Faster region-based convolutional neural networks model. The textual and visual features have been concatenated to construct multimodal features that the decoder uses to synthesize the feedback. We have compared the results of the proposed system with baseline models using quantitative and qualitative measures. The synthesized feedbacks have been analyzed using automatic and human evaluation. They have been found to be semantically similar to the ground-truth comments and relevant to the given text–image input.
Puneet Kumar 0003, Gaurav Bhatt, Omkar Ingle, Daksh Goyal, Balasubramanian Raman
ACM Trans. Multim. Comput. Commun. Appl.5
2022 Towards Improved Skin Lesion Classification using Metadata Supervision
abstract
Skin cancer is one of the most common types of cancer worldwide. Nowadays, Computer-Aided Diagnosis (CAD) systems are being adopted to diagnose skin cancer as they help reduce the manual burden on doctors and provide high reliability. Most of the contributions made towards developing algorithms for skin lesion classification have mainly considered the imaging modality only. However, in practical scenarios, it has been found that dermatologists also consider the patient’s demographics to refine their outcome. Based on this fact, this paper proposes a multimodal fusion-based approach toward improved skin lesion classification using patient metadata. We have presented a novel algorithm that combines the patient’s clinical information to guide the image features to improve the skin lesion classification. Also, the proposed work solved the issue of missing and unknown values in the patients’ metadata to assist in further improvements in skin cancer classification. The proposed approach is validated through extensive experimentation and surpasses available state-of-the-art methods on the benchmark dataset (PAD-UFES-20). We have evaluated the proposed method quantitatively and qualitatively and found it robust enough to classify skin lesion categories effectively. The proposed approach outperformed the benchmark results significantly over five Convolutional Neural Network (CNN) architectures during the evaluation. Our results can be reproduced§.
Anshul Pundhir, Saurabh Dadhich, Ananya Agarwal, Balasubramanian Raman
ICPR4
2022 C-CADZ: computational intelligence system for coronary artery disease detection using Z-Alizadeh Sani dataset
Ankur Gupta 0002, Rahul Kumar 0003, Harkirat Singh Arora, Balasubramanian Raman
Appl. Intell.4
2022 Parkinson's disease diagnosis using neural networks: Survey and comprehensive evaluation
Muhammad Tanveer 0001, Ashraf Haroon Rashid, Rahul Kumar 0003, Balasubramanian Raman
Inf. Process. Manag.4
2022 Crypt-OR: A privacy-preserving distributed cloud computing framework for object-removal in the encrypted images
Vishesh Kumar Tanwar, Balasubramanian Raman, Rama Bhargava
J. Netw. Comput. Appl.2
2022 SecureDL: A privacy preserving deep learning model for image recognition over cloud
Vishesh Kumar Tanwar, Balasubramanian Raman, Amitesh Singh Rajput, Rama Bhargava
J. Vis. Commun. Image Represent.2
2022 Classification of COVID-19 from chest x-ray images using deep features and correlation coefficient
abstract
COVID-19 is a viral disease that in the form of a pandemic has spread in the entire world, causing a severe impact on people's well being. In fighting against this deadly disease, a pivotal step can prove to be an effective screening and diagnosing step to treat infected patients. This can be made possible through the use of chest X-ray images. Early detection using the chest X-ray images can prove to be a key solution in fighting COVID-19. Many computer-aided diagnostic (CAD) techniques have sprung up to aid radiologists and provide them a secondary suggestion for the same. In this study, we have proposed the notion of Pearson Correlation Coefficient (PCC) along with variance thresholding to optimally reduce the feature space of extracted features from the conventional deep learning architectures, ResNet152 and GoogLeNet. Further, these features are classified using machine learning (ML) predictive classifiers for multi-class classification among COVID-19, Pneumonia and Normal. The proposed model is validated and tested on publicly available COVID-19 and Pneumonia and Normal dataset containing an extensive set of 768 images of COVID-19 with 5216 training images of Pneumonia and Normal patients. Experimental results reveal that the proposed model outperforms other previous related works. While the achieved results are encouraging, further analysis on the COVID-19 images can prove to be more reliable for effective classification.
Rahul Kumar 0003, Ridhi Arora, Vipul Bansal, Vinodh J. Sahayasheela, Himanshu Buckchash, Javed Imran, N. Narayanan 0001, Ganesh Namasivayam Pandian, Balasubramanian Raman
Multim. Tools Appl.9
2022 CBSN: Comparative measures of normalization techniques for brain tumor segmentation using SRCNet
Rahul Kumar 0003, Ankur Gupta 0002, Harkirat Singh Arora, Balasubramanian Raman
Multim. Tools Appl.4
2022 A BERT based dual-channel explainable text emotion recognition system
Puneet Kumar 0003, Balasubramanian Raman
Neural Networks2
2022 Privacy-Preserving Distribution and Access Control of Personalized Healthcare Data
abstract
The popularity of wearable smart healthcare devices has led to the emergence of a new service paradigm. However, in order to improve service quality, the manufacturers and online service providers collect massive data. This is a big concern as medical data are extremely sensitive. A few schemes have been proposed to overcome this problem. But, they suffer from security risks and overall increased complexity. Also, there is no implicit entity authentication and data integrity involved. We address these problems by allowing rectified data access through a directing authority, known as the transcryptor, using polymorphic encryption. Entity authentication and data integrity are achieved by smartly organizing data access and key information packets. The performance of the proposed approach is tested over different modalities data with varying sizes, whereas the security analysis is demonstrated using a challenge-response game model. The comparison with the state-of-the-art schemes illustrates superiority of the proposed approach.
Amitesh Singh Rajput, Balasubramanian Raman
IEEE Trans. Ind. Informatics2
2022 IBRDM: An Intelligent Framework for Brain Tumor Classification Using Radiomics- and DWT-based Fusion of MRI Sequences
abstract
Brain tumors are one of the critical malignant neurological cancers with the highest number of deaths and injuries worldwide. They are categorized into two major classes, high-grade glioma (HGG) and low-grade glioma (LGG), with HGG being more aggressive and malignant, whereas LGG tumors are less aggressive, but if left untreated, they get converted to HGG. Thus, the classification of brain tumors into the corresponding grade is a crucial task, especially for making decisions related to treatment. Motivated by the importance of such critical threats to humans, we propose a novel framework for brain tumor classification using discrete wavelet transform-based fusion of MRI sequences and Radiomics feature extraction. We utilized the Brain Tumor Segmentation 2018 challenge training dataset for the performance evaluation of our approach, and we extract features from three regions of interest derived using a combination of several tumor regions. We used wrapper method-based feature selection techniques for selecting a significant set of features and utilize various machine learning classifiers, Random Forest, Decision Tree, and Extra Randomized Tree for training the model. For proper validation of our approach, we adopt the five-fold cross-validation technique. We achieved state-of-the-art performance considering several performance metrics, 〈Acc,Sens,Spec,F1-score,MCC,AUC〉 ≡ 〈 98.60%, 99.05%, 97.33%, 99.05%, 96.42%, 98.19% 〉, whereAcc,Sens,Spec,F1-score,MCC, andAUCrepresents the accuracy, sensitivity, specificity, F1-score, Matthews correlation coefficient, and area-under-the-curve, respectively. We believe our proposed approach will play a crucial role in the planning of clinical treatment and guidelines before surgery.
Rahul Kumar 0003, Ankur Gupta 0002, Harkirat Singh Arora, Balasubramanian Raman
ACM Trans. Internet Techn.4
2022 GraSP: Local Grassmannian Spatio-Temporal Patterns for Unsupervised Pose Sequence Recognition
abstract
Many applications of action recognition, especially broad domains like surveillance or anomaly-detection, favor unsupervised methods considering that exhaustive labeling of actions is not possible. However, very limited work has happened in this domain. Moreover, the existing self-supervised approaches suffer from their dependency upon labeled data for finetuning. To this end, this paper puts forward a manifold based unsupervised pose-sequence recognition approach that leverages only the natural biases present in the data. It works by clustering the projections of temporal derivatives of the fragmented data on the Grassmann manifold. Temporal derivatives are formed by the inter-frame gradients with local and global metrics. To commensurate with this, a dynamic view-invariant pose representation is proposed. Additionally, a variable aggregation step is introduced for better feature vector quantization. Extensive empirical evaluation and ablations on several challenging datasets under three categories confirm the superiority of the proposed approach in contrast to current methods.
Himanshu Buckchash, Balasubramanian Raman
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Pansharpening Scheme Using Bi-dimensional Empirical Mode Decomposition and Neural Network
abstract
The pansharpening is a combination of multispectral (MS) and panchromatic (PAN) images that produce a high-spatial-spectral-resolution MS images. In multiresolution analysis–based pansharpening schemes, some spatial and spectral distortions are found. It can be reduced by adding spatial detail images of the PAN image into MS images. In the convolution neural network– (CNN) based method, the lowpass filter image extracted by the CNN model when MS and PAN images are directly applied into the input. The feature values are very high and reduce the conversion efficiency. In the proposed scheme, bi-dimensional empirical mode decomposition is used to extract the spatial detail information of the PAN image to reduce the feature values of the input. This extracted PAN image information is applied to the CNN to produce the non-linear changes in the image pixels and transformed into the perfect spatial detail image. It identifies the spatial and spectral detail quantity for the proposed scheme and it also varies with the different datasets automatically of the same satellite images. Simulation results in the context of qualitative and quantitative analysis demonstrate the effectiveness of proposed scheme applied on datasets collected by different satellites.
Nidhi Saxena, Balasubramanian Raman
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Hybrid Fusion Based Approach for Multimodal Emotion Recognition with Insufficient Labeled Data
abstract
In this paper, a deep learning based fusion approach has been proposed to classify the emotions portrayed by image and corresponding text into discrete emotion classes. The proposed method first implements intermediate fusion on image and text inputs and then applies late fusion on image, text, and intermediate fusion’s output. We have also come up with a way to handle the unavailability of labeled multimodal emotional data. We have prepared a new dataset built on Balanced Twitter for Sentiment Analysis dataset (B-T4SA) dataset containing an image, text, and emotion labels, i.e., ‘happy,’ ‘sad,’ ‘hate’ and ‘anger.’ The emotion recognition accuracy of 90.20% has been achieved by the proposed method. Along with multi-class emotion recognition, we’ve also compared the sentiment classification results and found the proposed method to perform better than the benchmark approaches.
Puneet Kumar 0003, Vedanti Khokher, Yukti Gupta, Balasubramanian Raman
ICIP4
2021 Towards the Explainability of Multimodal Speech Emotion Recognition
Puneet Kumar 0003, Vishesh Kaushik, Balasubramanian Raman
Interspeech3
2021 Towards zero shot learning of geometry of motion streams and its application to anomaly recognition
Himanshu Buckchash, Balasubramanian Raman
Expert Syst. Appl.2
2021 Towards accurate classification of skin cancer from dermatology images
abstract
Abstract Skin cancer is the most well‐known disease found in the individuals who are exposed to the Sun's ultra‐violet (UV) radiations. It is identified when skin tissues on the epidermis grow in an uncontrolled manner and appears to be of different colour than the normal skin tissues. This paper focuses on predicting the class of dermascopic images as benign and malignant. A new feature extraction method has been proposed to carry out this work which can extract relevant features from image texture. Local and gradient information from and directions of images has been utilized for feature extraction. After that images are classified using machine learning algorithms by using those extracted features. The efficacy of the proposed feature extraction method has been proved by conducting several experiments on the publicly available image dataset 2016 International Skin Imaging Collaboration (ISIC 2016). The classification results obtained by the method are also compared with state‐of‐the‐art feature extraction methods which show that it performs better than others. The evaluation criteria used to obtain the results are accuracy, true positive rate (TPR) and false positive rate (FPR) where TPR and FPR are used for generating receiver operating characteristic curves.
Anjali Gautam, Balasubramanian Raman
IET Image Process.2
2021 Deep neural network hyper-parameter tuning through twofold genetic approach
Puneet Kumar 0003, Shalini Batra, Balasubramanian Raman
Soft Comput.3
2021 Discriminative Auto-Encoding for Classification and Representation Learning Problems
abstract
Auto-Encoders (AE) play an important role in feature extraction, fusion and representation learning. Particularly, in case of medical datasets, where labeled data is often scarce, they perform better than transfer-learning approaches. Many auto-encoding methods have been proposed, however, present methods either preserve the latent space (with lower predictive power) or provide good predictive performance (with distorted latent space). This work presents a Discriminative Auto-Encoding (DiscAE) approach which provides better representations with decent reconstructions. Clustering constraint is imposed through a discriminator on the latent space of a simple auto-encoder to enforce Gaussian distribution of the latent space. The network is trained by alternating between decoder and the discriminator. Furthermore, the role of noise is explored in improving the reconstructions and regularization of the latent space. Unlike Variational Auto-Encoders (VAE), DiscAE provides better performance by utilizing the labeled data. The working of DiscAE under a semi-supervised setting is also discussed. To support our claims, extensive experiments and ablations were carried out on three benchmark datasets. It was found that DiscAE outperforms the existing auto-encoding approaches on predictive tasks while maintaining the quality of reconstructions.
Vipul Bansal, Himanshu Buckchash, Balasubramanian Raman
IEEE Signal Process. Lett.3
2021 Efficient Method and Architecture for Real-Time Video Defogging
abstract
Real-time video defogging has a huge demand in intelligent transportation, advanced driver assistance systems (ADAS), long-range surveillance, autonomous aerial vehicles and endoscopic surgery. Most of these applications are constrained by stringent frame rate, power and memory budget. Till date, the methods devised for such requirements are very rare. Therefore, this paper proposes an efficient method and very-large-scale integration (VLSI) architecture for resource-constrained embedded system targeting real-time video defogging. The method and architecture are co-designed to achieve high throughput while consuming less resources and power. The architecture is divided into four parts namely, atmospheric light estimation unit, airlight adjustment unit, transmission estimation unit, and pixel restoration unit. The atmospheric light estimation unit employs a$3\times 3$tile based approach that eliminates the requirement of large buffer memory for$15\times 15$support size. The transmission is estimated using dark channel prior and gradient threshold based Gaussian filtering approach. In order to overcome flickering artifacts, an adaptive airlight updating scheme is employed. From the quantitative and qualitative evaluations, it is observed that the proposed method outperforms the existing hardware approaches. Furthermore, the field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC) implementations of the architecture achieve a high throughput of 200 MPixels/s and 600 MPixels/s, respectively. The ASIC implementation dissipates 9.03 mW power at 200MHz. Moreover, the proposed design does not require any external memory such as dynamic random-access memory (DRAM), thus making it suitable for on-chip processing that can be closely integrated with an image sensor.
Rahul Kumar 0007, Balasubramanian Raman, Brajesh Kumar Kaushik
IEEE Trans. Intell. Transp. Syst.2
2021 -Score-Based Secure Biomedical Model for Effective Skin Lesion Segmentation Over eHealth Cloud
abstract
This study aims to process the private medical data over eHealth cloud platform. The current pandemic situation, caused by Covid19 has made us to realize the importance of automatic remotely operated independent services, such as cloud. However, the cloud servers are developed and maintained by third parties, and may access user’s data for certain benefits. Considering these problems, we propose a specialized method such that the patient’s rights and changes in medical treatment can be preserved. The problem arising due to Melanoma skin cancer is carefully considered and a privacy-preserving cloud-based approach is proposed to achieve effective skin lesion segmentation. The work is accomplished by the development of a Z -score-based local color correction method to differentiate image pixels from ambiguity, resulting the segmentation quality to be highly improved. On the other hand, the privacy is assured by partially order homomorphic Permutation Ordered Binary (POB) number system and image permutation. Experiments are performed over publicly available images from the ISIC 2016 and 2017 challenges, as well as PH dataset, where the proposed approach is found to achieve significant results over the encrypted images (known as encrypted domain), as compared to the existing schemes in the plain domain (unencrypted images). We also compare the results with the winners of the ISBI 2016 and 2017 challenges, and show that the proposed approach achieves a very close result with them, even after processing test images in the encrypted domain. Security of the proposed approach is analyzed using a challenge-response game model.
Amitesh Singh Rajput, Vishesh Kumar Tanwar, Balasubramanian Raman
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Novel Architecture for Lifting Discrete Wavelet Packet Transform With Arbitrary Tree Structure
abstract
This brief presents a novel pipelined VLSI architecture for computing discrete wavelet packet transform (DWPT) with an arbitrary wavelet tree. Coefficients for different levels are computed in a series of stages. Each stage consists of a bypassed wavelet filter and circuit for reordering intermediate coefficients. The proposed lifting-based wavelet filter computes high- and low-pass coefficients in series. In order to accommodate the arbitrary tree structure, the filter either computes the coefficients or bypass the samples. The reordering of intermediate coefficients forms a subband required for next-level computation. The coefficients are computed in a serial manner and reordering of intermediate coefficients reduce not only the memory elements but also the circuit complexity. The proposed pipelined architecture reduces the requirement of memory elements by 50%. Furthermore, the hardware implementation results show that the area and power requirement are reduced by 33% and 20%, respectively.
Gyanendra Singh, Samba Raju Chiluveru, Balasubramanian Raman, Manoj Tripathy, Brajesh Kumar Kaushik
IEEE Trans. Very Large Scale Integr. Syst.3
2020 Human Motion Generation by Stochastic Conditioning of Deep Recurrent Networks On Pose Manifolds
abstract
Human motion generation is a stochastic process. The 3D motion generation task requires efficient regulation of stochasticity and a controlled approach for error-accumulation. Current generation approaches either fail to check error-amplitude or to preserve the signal. In this paper, we present a stochastic approach for 3D human motion generation. To this end, we design a fully differentiable, end-to-end, block-based autoregressive recurrent neural network (RNN) architecture. The proposed model incorporates variable auto-conditioning length along with probabilistic variational inference on the RNN hidden-state, to regulate stochasticity. We separately train an auto-encoder to bound skeletons on a known manifold of valid-poses. We extensively test the proposed approach on publicly available Motion-Capture benchmarks. The quantitative and qualitative evaluations indicate the superiority of the proposed approach in comparison to state-of-the-art on long-term motion generation while achieving comparable performance on short-term prediction task.
Himanshu Buckchash, Balasubramanian Raman
ICIP2
2020 End-to-end Triplet Loss based Emotion Embedding System for Speech Emotion Recognition
abstract
In this paper, an end-to-end neural embedding system based on triplet loss and residual learning has been proposed for speech emotion recognition. The proposed system learns the embeddings from the emotional information of the speech utterances. The learned embeddings are used to recognize the emotions portrayed by given speech samples of various lengths. The proposed system implements Residual Neural Network architecture. It is trained using softmax pretraining and triplet loss function. The weights between the fully connected and embedding layers of the trained network are used to calculate the embedding values. The embedding representations of various emotions are mapped onto a hyperplane, and the angles among them are computed using the cosine similarity. These angles are utilized to classify a new speech sample into its appropriate emotion class. The proposed system has demonstrated 91.67% and 64.44% accuracy while recognizing emotions for RAVDESS and IEMOCAP dataset, respectively.
Puneet Kumar 0003, Sidharth Jain, Balasubramanian Raman, Partha Pratim Roy 0001, Masakazu Iwamura
ICPR3
2020 DuTriNet: Dual-Stream Triplet Siamese Network for Self-Supervised Action Recognition by Modeling Temporal Correlations
abstract
Self Supervised Learning (SSL) is the task of training a model independent of human annotations. A recent path-breaking SSL work is OpenAI's GPT-3. Very limited work has happened on SSL for Action Recognition (AR). Present SSL models either leverage only spatial data or use suboptimal frame sampling algorithms. To this end, we present a comprehensive study and propose DuTriNet for SSL in AR. We introduce the idea of temporal protraction by fusing optical-flow information with weber-maps inside parallel data-streams which share weights under a triplet Siamese architecture. We also propose flow-intensity based non-parametric frame sampling algorithm. Extensive experiments and ablations have been performed on two publicly available benchmark datasets for AR. Our findings suggest the suitability of DuTriNet for SSL.
Himanshu Buckchash, Balasubramanian Raman
ICTAI2
2020 Skin Cancer Classification from Dermoscopic Images using Feature Extraction Methods
abstract
Melanoma is a type of skin cancer that is mainly caused by intense UV exposure. If melanoma is identified at an early stage, then it is generally remediable. However, if it is not diagnosed properly, cancer can grow to rest of the body which then makes it difficult to cure and can be lethal. Conventionally, melanoma is diagnosed through visual methods and biopsies but their accuracy may not be reliable for all the cases. Hence, the risks involved for such a diagnosis have emerged identification and classification of melanoma as benign or malignant a very important research problem in medical imaging. This paper employs various feature descriptors like local binary pattern (LBP), complete LBP (CLBP) and their variants, which are based on histogram mapping such as uniform, rotation invariant and rotation invariant uniform patterns. The extracted features are then used to train different classifiers such as decision tree, random forest (RF), support vector machine (SVM) and k nearest neighbour (kNN). A comparative study of the various feature descriptors and classifiers are analyzed for accurate identification and classification of melanoma as benign or malignant. An image dataset which has been used in our work has been downloaded from ISIC-Archive, which consists of 947 dermoscopic images and the dataset is made freely available online by realizing the importance of the research. The best accuracy has been obtained by using RF in CLBP with an accuracy of 80.3%.
Anjali Gautam, Balasubramanian Raman
TENCON2
2020 A privacy-preserving protocol for efficient nighttime haze removal using cloud based automatic reference image selection and color transfer as a service
Amitesh Singh Rajput, Balasubramanian Raman
Comput. Commun.2
2020 Privacy-preserving human action recognition as a remote cloud service using RGB-D sensors and deep CNN
abstract
Cloud-based expert systems are highly emerging nowadays. However, the data owners and cloud service providers are not in the same trusted domain in practice. For the sake of data privacy, sensitive data usually has to be encrypted before outsourcing which makes effective cloud utilization a challenging task. Taking this concern into account, we propose a novel cloud-based approach to securely recognize human activities. A few schemes exist in the literature for secure recognition. However, they suffer from the problem of constrained data and are vulnerable to re-identification attack, where advanced deep learning models are used to predict an object’s identity. We address these problems by considering color and depth data, and securing them using position based superpixel transformation. The proposed transformation is designed by actively involving additional noise while resizing the underlying image. Due to this, a higher degree of obfuscation is achieved. Further, in spite of securing the complete video, we secure only four images, that is, one motion history image and three depth motion maps which are highly saving the data overhead. The recognition is performed using a four stream deep Convolutional Neural Network (CNN), where each stream is based on pre-trained MobileNet architecture. Experimental results show that the proposed approach is the best suitable candidate in “security-recognition accuracy (%)” trade-off relation among other image obfuscation as well as state-of-the-art schemes. Moreover, a number of security tests and analyses demonstrate robustness of the proposed approach.
Amitesh Singh Rajput, Balasubramanian Raman, Javed Imran
Expert Syst. Appl.2
2020 Efficient extraction of consistent bit locations from binarized iris features
Debanjan Sadhya, Kanjar De, Balasubramanian Raman, Partha Pratim Roy 0001
Expert Syst. Appl.3
2020 Vector ordering and regression learning-based ranking for dynamic summarisation of user videos
abstract
Dynamic video summarisation (video skimming) is a process of generating a shorter video (video skim) as a summary of a given video, which helps in its easier and quicker comprehension. In this study, an efficient dynamic summarisation approach for user videos is proposed using vector ordering for ranking video units (frames/shots). User videos are casually shot unscripted videos, where skimming involves the selection of its interesting part(s) ignoring many uninteresting ones. The concept of R‐ordering of vectors is employed to find a representative frame, which is used to perform relative ranking of the video frames. It is theoretically shown that significance is given to each element of a frame's feature vector while computing the importance scores that lead to the frame ranks used for skimming. Furthermore, the allocation of different weights to the features involved is also achieved using linear and Gaussian process regressions. Through extensive experiments considering several standard datasets with human‐labelled ground truth, the proposed approach is demonstrated to be efficient and to perform better than the relevant state‐of‐the‐art.
Vivekraj V. K, Debashis Sen, Balasubramanian Raman
IET Image Process.3
2020 Guest Editorial Special Issue on Edge-Cloud Interplay Based on SDN and NFV for Next-Generation IoT Applications
abstract
With significant and continuing advances in information and communication technologies, the Internet of Things (IoT) will play an increasingly important role in domains, such as healthcare, transportation, finance, and energy. In an IoT system, billions of devices (e.g., sensors, wearables, and smart appliances) are connected to the global network infrastructure, and one associated phenomenon is the generation of a large volume of data. Apart from data volume, the velocity, variety, and veracity of these data will pose a significant burden on conventional networking infrastructures. However, as sensor and fifth-generation (5G) cellular technologies advance, so will the pervasiveness of IoT deployment. Parallel to this trend, cloud computing has been integrated with IoT in order to address limitations in existing IoT networks (e.g., storage and computing resources), and examples include Google cloud dataflow and Amazon IoT. However, cloud-centric IoT solutions may not be suited for delay-sensitive and computationally intensive applications, for example, due to resource availability, end-to-end latency, bandwidth, etc. Increasingly, large-scale IoT deployments demand high connectivity, interoperability, and orchestration which are necessary for minimizing latency and maximizing throughput. This highlights the importance of a distributed computing platform that can support the interactions between IoT and cloud computing systems.
Sahil Garg, Song Guo 0001, Vincenzo Piuri, Kim-Kwang Raymond Choo, Balasubramanian Raman
IEEE Internet Things J.5
2020 Fast Griffin Lim based waveform generation strategy for text-to-speech synthesis
Puneet Kumar 0003, Vikas Maddukuri, Nagasai Madamshettib, Kishore KG, Sahit Sai Sriram Kavurub, Balasubramanian Raman, Partha Pratim Roy 0001
Multim. Tools Appl.7
2020 Local gradient of gradient pattern: a robust image descriptor for the classification of brain strokes from computed tomography images
Anjali Gautam, Balasubramanian Raman
Pattern Anal. Appl.2
2020 2DInpaint: A novel privacy-preserving scheme for image inpainting in an encrypted domain over the cloud
Vishesh Kumar Tanwar, Balasubramanian Raman, Amitesh Singh Rajput, Rama Bhargava
Signal Process. Image Commun.2
2020 CryptoLesion: A Privacy-preserving Model for Lesion Segmentation Using Whale Optimization over Cloud
abstract
The low-cost, accessing flexibility, agility, and mobility of cloud infrastructures have attracted medical organizations to store their high-resolution data in encrypted form. Besides storage, these infrastructures provide various image processing services for plain (non-encrypted) images. Meanwhile, the privacy and security of uploaded data depend upon the reliability of the service provider(s). The enforcement of laws towards privacy policies in health-care organizations, for not disclosing their patient’s sensitive and private medical information, restrict them to utilize these services. To address these privacy concerns for melanoma detection, we propose CryptoLesion , a privacy-preserving model for segmenting lesion region using whale optimization algorithm (WOA) over the cloud in the encrypted domain (ED). The user’s image is encrypted using a permutation ordered binary number system and a random stumble matrix. The task of segmentation is accomplished by dividing an encrypted image into a pre-defined number of clusters whose optimal centroids are obtained by WOA in ED, followed by the assignment of each pixel of an encrypted image to the unique centroid. The qualitative and quantitative analysis of CryptoLesion is evaluated over publicly available datasets provided in The International Skin Imaging Collaboration Challenges in 2016, 2017, 2018, and PH 2 dataset. The segmented results obtained by CryptoLesion are found to be comparable with the winners of respective challenges. CryptoLesion is proved to be secure from a probabilistic viewpoint and various cryptographic attacks. To the best of our knowledge, CryptoLesion is first moving towards the direction of lesion segmentation in ED.
Vishesh Kumar Tanwar, Balasubramanian Raman, Amitesh Singh Rajput, Rama Bhargava
ACM Trans. Multim. Comput. Commun. Appl.2
2020 Dynamic texture recognition using local tetra pattern - three orthogonal planes (LTrP-TOP)
Amit, Balasubramanian Raman, Debanjan Sadhya
Vis. Comput.2
2020 Deep motion templates and extreme learning machine for sign language recognition
Javed Imran, Balasubramanian Raman
Vis. Comput.2
2019 Privacy-Preserving Smart Surveillance Using Local Color Correction and Optimized ElGamal Cryptosystem over Cloud
abstract
The emergence of cloud computing in integration with smart multimedia devices has created an attractive business model today. However, due to the involvement of third party servers, there is a risk of privacy for highly confidential data like surveillance images/videos. Moreover, due to inconsistent lightning conditions, there is a usual requirement of post-processing the captured multimedia for better appearance. Addressing these problems, we propose a novel cloud based privacy-preserving approach for image color enhancement in this paper. Unlike existing color correction schemes, where colors of the test image are processed in plain domain with visible image contents, we propose to perform color correction operations in the encrypted domain over cloud. As a consequence, superior results are achieved along with complete privacy assurance. In addition, we propose a block-based image encryption method using logistic-tent system and ElGamal cryptosystem. As a result, size of the encrypted image is significantly reduced as compared to the naive approach. Experimental results are performed under various tests and the proposed approach is found to be highly effective as compared to state-of-the-art schemes. Moving ahead, security strength of the proposed approach is demonstrated through a challenge response game model.
Amitesh Singh Rajput, Balasubramanian Raman
CLOUD2
2019 Development of a clustering based fusion framework for locating the most consistent IrisCodes bits
Debanjan Sadhya, Kanjar De, Balasubramanian Raman, Partha Pratim Roy 0001
Inf. Sci.3
2019 Segmentation of ischemic stroke lesion from 3d mr images using random forest
Anjali Gautam, Balasubramanian Raman
Multim. Tools Appl.2
2019 Facial emotion classification using concatenated geometric and textural features
Debashis Sen, Samyak Datta, Balasubramanian Raman
Multim. Tools Appl.3
2019 Dense motion analysis of German finger spellings
Vishesh Kumar Tanwar, Himanshu Buckchash, Balasubramanian Raman, Rama Bhargava
Multim. Tools Appl.3
2019 Representation learning using step-based deep multi-modal autoencoders
Gaurav Bhatt, Piyush Jha, Balasubramanian Raman
Pattern Recognit.3
2019 Generation of Cancelable Iris Templates via Randomized Bit Sampling
abstract
Iris-based biometric models are widely recognized to be one of the most accurate forms for authenticating individual identities. Features extracted from the captured iris images (known as IrisCodes) conventionally get stored in their native format over a data repository. However, from a security aspect, the stored templates are highly vulnerable to a wide spectrum of adversarial attack forms. The study in this paper addresses this issue by introducing a privacy-preserving and secure biometric scheme based on the notion of locality sensitive hashing (LSH). In this paper, we have generated cancelable IrisCode features, coined as locality sampled code (LSC), which simultaneously provides strong security guarantees and satisfactory system performance. The functionality of our proposed framework pivots around the fact that intra-class IrisCode samples are “close” to each other, due to which they hash to the same location. Alternatively, the inter-class IrisCodes features are comparatively dissimilar and consequently hash to different locations. We have rigorously examined the intrinsic properties of the LSCs by estimating the intra-class and inter-class collision probabilities for two distinct IrisCodes. Furthermore, we have formally analyzed the security guarantees of non-invertibility, revocability, and unlinkability in our model by establishing various bounds on the adversarial success probability. Extensive empirical tests on the CASIAv3 and IITD benchmark iris databases demonstrate the superior performance of our proposed model, for which we have obtained the best EERs of 0.105% and 1.4%, respectively.
Debanjan Sadhya, Balasubramanian Raman
IEEE Trans. Inf. Forensics Secur.2
2018 Privacy Preserving Image Scaling Using 2D Bicubic Interpolation Over the Cloud
abstract
This paper presents an efficient scheme for privacy preserving image scaling over the cloud. The proposed scheme advances the emerging trend of privacy preserving cloud computing and supports desired operations required for image scaling in the encrypted domain. We use scaling and randomization followed by modulo operation for generation of image shares and employ 2D bicubic interpolation for scaling user images in the encrypted domain. The previous scheme used bilinear interpolation utilizing four neighboring pixel intensity values to approximate intermediate ones while scaling the image. On the contrary, we use sixteen neighboring pixel intensity values utilizing bicubic interpolation by leveraging considerable efforts to support the processing of user images over cloud servers. The proposed scheme is validated through various experiments followed by a comparison with the existing scheme. Additionally, security analysis demonstrates the robustness of the proposed scheme under various attack scenarios.
Vishesh Kumar Tanwar, Amitesh Singh Rajput, Balasubramanian Raman, Rama Bhargava
SMC3
2018 Secure data deduplication using secret sharing schemes over cloud
Priyanka Singh 0001, Nishant Agarwal, Balasubramanian Raman
Future Gener. Comput. Syst.3
2018 Don't just sign use brain too: A novel multimodal approach for user identification and verification
Rajkumar Saini, Barjinder Kaur, Priyanka Singh 0001, Pradeep Kumar 0002, Partha Pratim Roy 0001, Balasubramanian Raman
Inf. Sci.6
2018 Reversible data hiding based on Shamir's secret sharing for color images over cloud
Priyanka Singh 0001, Balasubramanian Raman
Inf. Sci.2
2018 Text recognition in scene image and video frame using Color Channel selection
Ayan Kumar Bhunia, Partha Pratim Roy 0001, Balasubramanian Raman, Umapada Pal 0001
Multim. Tools Appl.4
2018 Cloud based image color transfer and storage in encrypted domain
Amitesh Singh Rajput, Balasubramanian Raman
Multim. Tools Appl.2
2018 CryptoCT: towards privacy preserving color transfer and storage over cloud
Amitesh Singh Rajput, Balasubramanian Raman
Multim. Tools Appl.2
2018 Just process me, without knowing me: a secure encrypted domain processing based on Shamir secret sharing and POB number system
Priyanka Singh 0001, Balasubramanian Raman, Manoj Misra
Multim. Tools Appl.2
2018 Local neighborhood difference pattern: A new feature descriptor for natural and texture image retrieval
Manisha Verma, Balasubramanian Raman
Multim. Tools Appl.2
2018 A (n, n) threshold non-expansible XOR based visual cryptography with unique meaningful shares
Priyanka Singh 0001, Balasubramanian Raman, Manoj Misra
Signal Process.2
2018 Toward Encrypted Video Tampering Detection and Localization Based on POB Number System Over Cloud
abstract
The unlimited growth in the amount of multimedia content has shifted the global infrastructure to the cloud-based multimedia hosting. However, the high probability of security breaches of the content is increasing demand for secure solutions toward this end. One such feasible solution is to encrypt the content to unreadable form before outsourcing to the cloud-based servers. In this paper, the media information, specifically the video content is distributed into multiple random shares based on the permutation ordered binary number system. The information remains fully concealed without any leakage at the cloud servers. Even if the attacks endanger the integrity of the information, the proposed scheme is enriched with the capability of chalking out accurately the altered pixels. These tampered regions are subsequently reflected in the reconstructed video frames obtained at authentic entity end possessing the secret keys required to build back the original content. Moreover, the proposed scheme is capable of detecting temporal attacks on the video frames by employing sequence-based authentication bits. The robustness of the proposed scheme has been validated under different attack scenarios and the scheme is found to be performing satisfactorily well.
Priyanka Singh 0001, Balasubramanian Raman, Nishant Agarwal
IEEE Trans. Circuits Syst. Video Technol.2
2017 A Deep Learning Frame-Work for Recognizing Developmental Disorders
abstract
Developmental Disorders are chronic disabilities that have a severe impact on the day to day functioning of a large section of the human population. Recognizing developmental disorders from facial images is an important but a relatively unexplored challenge in the field of computer vision. This paper proposes a novel framework to detect developmental disorders from facial images. A spectrum of disorders constituting of Autism Spectrum Disorder, Cerebral Palsy, Fetal Alcohol Syndrome, Down syndrome, Intellectual disability and Progeria have been considered for recognition. The framework relies on Deep Convolutional Neural Networks (DCNN) for feature extraction. A new data-set comprising of images of subjects with these disabilities was built for testing the performance of the frame work. This model has been tested on different age groups, individual disabilities and has also been compared to a similar model that uses human intelligence to identify different developmental disorders. The results indicate that the model performs better than average human intelligence in terms of differentiating amongst different disabilities and is able to recognize subjects with these developmental disorders with an accuracy of 98.80%.
Pushkar Shukla, Tanu Gupta, Aradhya Saini, Priyanka Singh 0001, Balasubramanian Raman
WACV5
2017 Evaluation of periocular features for kinship verification in the wild
Bhavik Patel, R. P. Maheshwari 0001, Balasubramanian Raman
Comput. Vis. Image Underst.3
2017 Rotation and script independent text detection from video frames using sub pixel mapping
Anshul Mittal, Partha Pratim Roy 0001, Priyanka Singh 0001, Balasubramanian Raman
J. Vis. Commun. Image Represent.4
2017 An image copyright protection system using chaotic maps
Asha Rani 0005, Balasubramanian Raman
Multim. Tools Appl.2
2017 A secured robust watermarking scheme based on majority voting concept for rightful ownership assertion
Priyanka Singh 0001, Balasubramanian Raman
Multim. Tools Appl.2
2017 A multimodal biometric watermarking system for digital images in redundant discrete wavelet transform
Priyanka Singh 0001, Balasubramanian Raman, Partha Pratim Roy 0001
Multim. Tools Appl.2
2017 Prediction of advertisement preference by fusing EEG response and sentiment analysis
Himaanshu Gauba, Pradeep Kumar 0002, Partha Pratim Roy 0001, Priyanka Singh 0001, Debi Prosad Dogra, Balasubramanian Raman
Neural Networks6
2017 A secure image sharing scheme based on SVD and Fractional Fourier Transform
Priyanka Singh 0001, Balasubramanian Raman, Manoj Misra
Signal Process. Image Commun.2
2017 Secure Cloud-Based Image Tampering Detection and Localization Using POB Number System
abstract
The benefits of high-end computation infrastructure facilities provided by cloud-based multimedia systems are attracting people all around the globe. However, such cloud-based systems possess security issues as third party servers become involved in them. Rendering data in an unreadable form so that no information is revealed to the cloud data centers will serve as the best solution to these security issues. One such image encryption scheme based on a Permutation Ordered Binary Number System has been proposed in this work. It distributes the image information in totally random shares, which can be stored at the cloud data centers. Further, the proposed scheme authenticates the shares at the pixel level. If any tampering is done at the cloud servers, the scheme can accurately identify the altered pixels via authentication bits and localizes the tampered area. The tampered portion is also reflected back in the reconstructed image that is obtained at the authentic user end. The experimental results validate the efficacy of the proposed scheme against various kinds of possible attacks, tested with a variety of images. The tamper detection accuracy has been computed on a pixel basis and found to be satisfactorily high for most of the tampering scenarios.
Priyanka Singh 0001, Balasubramanian Raman, Nishant Agarwal, Pradeep K. Atrey
ACM Trans. Multim. Comput. Commun. Appl.2
2016 Vector R-ordering based selection of segments for video skimming
abstract
Video skimming is a process of generating a shorter yet fully comprehensible version of a given video as its dynamic summary. A generic skimming system involves division of the video into segments and selecting the segments based on their suitability. The suitability is often obtained considering various features of the video and combining their individual contributions. Suggesting that the combination causes loss of information, we propose collective representation of the individual contributions in the form of a vector and use vector reduced (R)-ordering to judge the suitability. R-ordering based tree-structured organization and similarity levels of the video segments are employed to determine the suitability. Comparing with user generated summaries, we show that a video summary generated by a general skimming approach using R-ordering will be more effective in covering the important parts of a given video than when a feature combination is used.
Vivekraj V. K, Balasubramanian Raman, Debashis Sen
ICPR2
2016 Compass local binary patterns for gender recognition of facial photographs and sketches
Bhavik Patel, R. P. Maheshwari 0001, Balasubramanian Raman
Neurocomputing3
2016 Visible watermarking based on importance and just noticeable distortion of image regions
Himanshu Agarwal, Debashis Sen, Balasubramanian Raman, Mohan Kankanhalli
Multim. Tools Appl.3
2016 An image copyright protection scheme by encrypting secret data with the host image
Asha Rani 0005, Balasubramanian Raman
Multim. Tools Appl.2
2016 A novel colour- and texture-based image retrieval technique using multi-resolution local extrema peak valley pattern and RGB colour histogram
Madhumanti Dey, Balasubramanian Raman, Manisha Verma
Pattern Anal. Appl.2
2015 Fusion of submanifold and local texture features for palmprint authentication
abstract
In this paper, a palmprint based authentication system is proposed. The proposed system makes use of locality sensitive discriminant analysis and directional local extrema patterns. A locality sensitive discriminant analysis has the capability to project the palmprints into a new space, where the palmprints belonging to same class become closer, whereas those of different classes move far apart. Locality sensitive discriminant analysis gathers the discriminating information based on the local texture of the data and available ground truth values. A directional local extrema pattern extracts the directional edge information based on local extrema in 0°, 45°, 90°, and 135° directions in an image. A palmprint contains several lines in various directions. A directional local extrema pattern may prove beneficial to extract edge information from palmprints. The two kinds of features are fused at score level for authentication purpose. The PolyU pamprint dataset has been used for evaluation, and the results are found promising. To establish the improvements in palmprint recognition, the results of proposed system has been compared with several well established state of the art algorithms.
Asha Rani 0005, Manisha Verma, Balasubramanian Raman
VCIP3
2015 Local extrema co-occurrence pattern for color and texture image retrieval
Manisha Verma, Balasubramanian Raman, M. Subrahmanyam 0001
Neurocomputing2
2015 Center symmetric local binary co-occurrence pattern for texture, face and bio-medical image retrieval
Manisha Verma, Balasubramanian Raman
J. Vis. Commun. Image Represent.2
2015 Image watermarking in real oriented wavelet transform domain
Himanshu Agarwal, Pradeep K. Atrey, Balasubramanian Raman
Multim. Tools Appl.3
2015 Blind reliable invisible watermarking method in wavelet domain for face image watermark
Himanshu Agarwal, Balasubramanian Raman, Ibrahim Venkat
Multim. Tools Appl.2
2014 A robust watermarking scheme exploiting balanced neural tree for rightful ownership protection
Asha Rani 0005, Balasubramanian Raman, Sanjeev Kumar 0001
Multim. Tools Appl.2
2013 Discrete fractional wavelet transform and its application to multiple encryption
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
Inf. Sci.3
2013 A new aspect in robust digital watermarking
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
Multim. Tools Appl.3
2013 Joint watermarking and encryption for still visual data
Nidhi Taneja, Gaurav Bhatnagar, Balasubramanian Raman, Indra Gupta
Multim. Tools Appl.3
2012 Image and Video Encryption based on Dual Space-Filling Curves
abstract
In this paper, a new encryption scheme with three different modes of operations is proposed based on dual space-filling curves (SFSs) and a fractional wavelet transform (FrWT). This scheme is initially proposed for images and then extended to videos. The core idea behind the proposed schemes is to decompose an image/video first by the FrWT followed by the shuffling of each sub-band coefficients by means of a dual SFC. At last, an inverse FrWT is performed to get the encrypted image/video. A reliable decryption process is also proposed to construct the original image from the encrypted image. The experimental results, security and comparative analysis demonstrate the efficiency and robustness of the proposed scheme. Further, this paper also proposes an efficient implementation of an FrWT based on chaotic maps.
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
Comput. J.3
2012 A new robust adjustable logo watermarking scheme
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
Comput. Secur.3
2012 Expert system design using wavelet and color vocabulary trees for image retrieval
M. Subrahmanyam 0001, R. P. Maheshwari 0001, Balasubramanian Raman
Expert Syst. Appl.3
2012 Fractional dual tree complex wavelet transform and its application to biometric security during communication and transmission
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
Future Gener. Comput. Syst.3
2012 Combinational domain encryption for still visual data
Nidhi Taneja, Balasubramanian Raman, Indra Gupta
Multim. Tools Appl.2
2012 Chaos based cryptosystem for still visual data
Nidhi Taneja, Balasubramanian Raman, Indra Gupta
Multim. Tools Appl.2
2012 A blind watermarking algorithm based on fractional Fourier transform and visual cryptography
Sanjay Rawat 0002, Balasubramanian Raman
Signal Process.2
2012 Local maximum edge binary patterns: A new descriptor for image retrieval and object tracking
M. Subrahmanyam 0001, R. P. Maheshwari 0001, Balasubramanian Raman
Signal Process.3
2012 Local Tetra Patterns: A New Feature Descriptor for Content-Based Image Retrieval
abstract
In this paper, we propose a novel image indexing and retrieval algorithm using local tetra patterns (LTrPs) for content-based image retrieval (CBIR). The standard local binary pattern (LBP) and local ternary pattern (LTP) encode the relationship between the referenced pixel and its surrounding neighbors by computing gray-level difference. The proposed method encodes the relationship between the referenced pixel and its neighbors, based on the directions that are calculated using the first-order derivatives in vertical and horizontal directions. In addition, we propose a generic strategy to compute nth-order LTrP using (n - 1)th-order horizontal and vertical derivatives for efficient CBIR and analyze the effectiveness of our proposed algorithm by combining it with the Gabor transform. The performance of the proposed method is compared with the LBP, the local derivative patterns, and the LTP based on the results obtained using benchmark image databases viz., Corel 1000 database (DB1), Brodatz texture database (DB2), and MIT VisTex database (DB3). Performance analysis shows that the proposed method improves the retrieval result from 70.34%/44.9% to 75.9%/48.7% in terms of average precision/average recall on database DB1, and from 79.97% to 85.30% and 82.23% to 90.02% in terms of average retrieval rate on databases DB2 and DB3, respectively, as compared with the standard LBP.
M. Subrahmanyam 0001, R. P. Maheshwari 0001, Balasubramanian Raman
IEEE Trans. Image Process.3
2012 A New Fractional Random Wavelet Transform for Fingerprint Security
abstract
In this correspondence paper, the wavelet transform, which is an important tool in signal and image processing, has been generalized by coalescing wavelet transform and fractional random transform. The new transform, i.e., fractional random wavelet transform (FrRnWT) inherits the excellent mathematical properties of wavelet transform and fractional random transform. Possible applications of the proposed transform are in biometrics, image compression, image transmission, transient signal processing, etc. In this correspondence paper, biometrics is chosen as the primary application; and hence, a new technique is proposed for securing fingerprints during communication and transmission over insecure channel.
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
IEEE Trans. Syst. Man Cybern. Part A3
2011 A new robust reference logo watermarking scheme
Gaurav Bhatnagar, Balasubramanian Raman
Multim. Tools Appl.2
2010 Real Time Human Visual System Based Framework for Image Fusion
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
ICISP3
2006 Binary Tree Based Linear Time Fingerprint Matching
abstract
Fingerprint matching algorithm is a key step in fingerprint recognition system. Though there are many existing matching algorithms, there has been inability to match fingerprints in linear time. In this paper we present a novel biometric approach to match fingerprints that run in linear time. We match the minutiae in the fingerprint by constructing a Nearest Neighbor Vector (NNV) considering its k-nearest neighbors. The consolidation of these matched minutiae points is done by incorporating them in binary tree that propagates simultaneously in both fingerprints. This helps our algorithm to run in O(n) time in contrast to many existing algorithms when reference core point is available. We analyze the resulting improvement in computational complexity and present experimental evaluation over FVC2002 database.
Mayur D. Jain, S. Nalin Pradeep, C. Prakash, Balasubramanian Raman
ICIP4
2006 Palmprint Recognition: Two level Structure Matching
abstract
We introduce palmprint recognition, one of the most reliable personal identification methods in the biometric technology. In this paper, a new approach to the palmprint matching by constructing local and global line feature structures is presented. The datum point of palmprint acts as an important registration due to its remarkable advantage about its spatial location. Initially we define all possible local line feature structures constructed with adjacent lines, around datum point. Through a first level match using these local line feature structures, we get best matched line features. Using these best matched line features, we construct a global line feature structure of palmprint with datum point as reference. This global line feature structure is spread across in four quadrants, datum point being the origin. We then carry out a second level match using this global structure of palmprint to reliably determine its uniqueness. The two levels of matching using local and global line feature structures helps in effective palmprint recognition. With several palmprint images, we tested out proposed verification system and the experimental result shows that the performance of our algorithm is good.
S. Nalin Pradeep, Mayur D. Jain, C. Prakash, Balasubramanian Raman
IJCNN4
2005 Computing hierarchical curve-skeletons of 3D objects
Nicu D. Cornea, Deborah Silver, Xiaosong Yuan, Balasubramanian Raman
Vis. Comput.4