EDBT 2026 Demo / reviewers in the wild / expert
Santosh Kumar Vipparthi
dblp:148/8541
· DBLP profile ↗
53ranked-venue papers
3as first author
38since 2021 · last 2027
0000-0002-5672-3537ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 24 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | A fractional hierarchical attention model for enhanced context-guided visual storytelling
Prem Shanker Yadav, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi |
Inf. Process. Manag. | 3 |
| 2026 | QuEENet: Quantum-Enhanced Expressive Network for Image ClassificationabstractThis paper presents QuEENet, a hybrid quantum-classical architecture for image classification that incorporates parameterized quantum circuits within a convolution neural network. This study investigates how quantum circuit expressivity and entanglement strategies influence classification performance, with a focus on configurations involving a CNOT gate followed by a rotational gate on the target qubit. Non-Clifford gates, such as Rx/Ry/Rzsupports larger state-space coverage and expressivity in quantum models. The proposed QuEENet explored the aspect of non-Clifford gates in parameterized quantum circuits. While non-Clifford gates are theoretically critical for universal quantum computation, but their role in image classification task is unexplored. Experimental results across multiple benchmark datasets suggest that while increased expressivity via non-Clifford gates can be beneficial, it should be carefully balanced with circuit interpretability and trainability. QuEENet demonstrates that hybrid models can leverage quantum circuits not merely as architectural novelties, but as controllable modules for enhancing learning in classical pipelines. An extensive ablation study was conducted across multiple datasets to highlight the effects of Clifford and non-Clifford gate combinations and entanglement configurations. Shashank Bayal, Rushikesh Govind Dawane, Komal, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 4 |
| 2026 | DTMIR-Pro: Domain Translation with Prompt-based Latent-Space Generalization for Multi-Weather Image RestorationabstractMulti-weather image restoration seeks to recover scene visibility under rainy, snowy, and hazy conditions, thereby enhancing high-level vision tasks. Existing methods typically train on combined datasets with single-type weather degradations, limiting their generalization to real-world scenarios involving mixed degradations. Domain translation has emerged as a viable solution by generating diverse weather-degraded variants of the same scene. However, current approaches require separate models for each degradation type, resulting in increased system complexity. To address this, we propose DTMIR-Pro, a prompt-based domain translation framework with latent space generalization for multi-weather image restoration. A single trainable network performs multi-domain translation using domain-adaptive prompts and dynamic kernel selection via a proposed Dynamic Multi-Head Attention block to learn diverse degradation patterns. The restoration network takes translated outputs and employs a Multi-Weather Fusion Block with global-local feature streams to capture complex degradations. Furthermore, we introduce a Similarity-Based Encoder Routing mechanism to transfer domain-specific features from the translation encoder to the restoration stage. Extensive experiments on both synthetic and real-world weather-degraded datasets demonstrate the effectiveness and generalizability of the proposed method. The code is made available at https://github.com/AshutoshKulkarni4998/DTMIR-Pro. Ashutosh Kulkarni, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Balasubramanian Raman |
WACV | 3 |
| 2026 | A motion flow guided MicroNet framework for micro expression recognition
Monu Verma, Santosh Kumar Vipparthi, Mohamed Abdel-Mottaleb |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | FedHC: Enhanced federated learning with Hessian and cosine correlation for proximal correlation
Kushall Singh, Monu Verma, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi, G. Shankara Raju Kosuru, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb |
Knowl. Based Syst. | 4 |
| 2026 | FedHMed: Adaptive progressive loss and KL-divergence regularization for federated heterogeneous medical image classification tasks
Kushall Singh, Monu Verma, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb |
Knowl. Based Syst. | 4 |
| 2026 | ME-NAS: A Micro Expression Feature Adaptive Neural Architecture SearchabstractConvolution neural networks (CNN) have emerged as a prevailing paradigm for micro-expression recognition (MER) yet, it is inefficient and time-intensive to design optimal CNN-based MER models manually. In recent times, the neural architecture search (NAS) has garnered attention due to its automatic CNN architecture searching ability. However, the performance of NAS in MER is limited by challenges such as rapid duration, subtle intensity, and a mismatch between architecture and cell-level search. The existing search space, which stacks 12 cells with 3 transition paths (downsample, upsample, and same resolution), creates deep networks that may diminish minute spatiotemporal features due to progressive convolution and pooling. Therefore, motivated by these factors, in this article, we introduce a novel approach, the Micro-Expression Feature Adaptive NAS (ME-NAS), to analyze true human emotions through MER. While NAS has gained attention for its automatic CNN architecture search ability, its application in MER faces challenges due to ingrained challenges (rapid duration, subtle and low intensity) and the discrepancy between architecture and cell-level search. The existing NAS architecture search space is designed by stacking 12 cells with 3 transition paths (downsample, upsample, and same resolution), resulting in a deep network. Such deep networks may diminish minute spatiotemporal features due to the progressive convolution and pooling operations. Motivated by these factors, we designed a new NAS algorithm: ME-NAS. The ME-NAS comprises f (EXPERT) in architecture search, along with refined and complementary feature derivative (ReCODE) operations in cell-level search. The EXPERT aims to trace the optimal paths instead of covering all possible paths between cells. The ReCODE operations capture micro-level variations from spatial and temporal domains by introducing 24 3D convolution operations. The proposed ReCODE and EXPERT search space jointly lead to the search for a robust and shallow CNN architecture for micro-expressions (MEs). The proposed ME-NAS is evaluated on six datasets: CASME-I, CASME-II, CAS(ME) \({}^{2}\) , SAMM, SMIC, and MEGC-19 composite, with two evaluation strategies: LOSO and cross-domain, respectively. The experimental results manifest that the proposed ME-NAS outperformed the state-of-the-art approaches on both evaluation strategies. Monu Verma, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2025 | TRUST: Time-Domain Residual Unsupervised Stability Technique for Improved Heart Rate EstimationabstractCamera-based estimation of vital signs is a promising method for non-contact health monitoring, which analyzes minute changes in video data. However, the creation of accurate models for this task is challenging due to the scarcity of datasets that possess synchronized vital sign recordings. Our research enhances an existing non-contrastive unsupervised learning technique for extracting rPPG signals, which does not necessitate ground-truth signals during the training process. We have incorporated new time-domain loss functions and added a feature stabilization block to improve the model's stability and accuracy in detecting low-level features. Additionally, we have devised a metric to evaluate the feature instability in the model's final layer. Our experiments on four public datasets demonstrate that our method surpasses the performance of current state-of-the-art methods. These advancements make our approach a significant breakthrough in the development of scalable deep-learning models for camera-based heart-rate estimation. Shahzad Ahmad 0002, Sania Bano, Sukalpa Chanda, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 4 |
| 2025 | PULSE: Physiological Understanding with Liquid Signal ExtractionabstractThe non-contact estimation of vital signs, particularly heart rate, from video data is a promising method for remote health monitoring. 3D convolutional layers are widely used for this task due to their ability to capture both spatial and temporal features. However, traditional 3D convolutions, while effective in many cases, lack the capacity to adjust dy-namically to the temporal variability inherent in physiological signals such as remote photoplethysmography (rPPG), which are characterized by subtle frequency changes over time. To address this, we propose PULSE (Physiological Understanding with Liquid Signal Extraction), a frame-work that employs Liquid Time-Constant (LTC) models with 3D convolutional layers to enhance temporal sensitivity and improve the extraction of these fine-grained rPPG signals. In PULSE, traditional 3D-conv layers are deployed for ini-tial feature extraction, while LTC-based 3D-conv layers dy-namically adapt and guide the temporal processing, allowing the model to better track and interpret the subtle variations in heart rate signals under different conditions, such as motion artifacts and lighting changes. We evaluated the effectiveness of PULSE in an unsupervised training setting, demonstrating that our solution performs well even in the absence of labeled datasets a common challenge in rPPG signal extraction. Experimental evaluations on three public datasets confirm that PULSE achieves comparable or supe-rior results to existing methods, proving its robustness and efficacy for real-world, non-contact health monitoring applications. Shahzad Ahmad 0002, Sania Bano, Sachin Verma, Yogesh S. Rawat, Sukalpa Chanda, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 6 |
| 2025 | Phaseformer: Phase-Based Attention Mechanism for Underwater Image Restoration and BeyondabstractQuality degradation is observed in underwater images due to the effects of light refraction and absorption by water, leading to issues like color cast, haziness, and limited visibility. This degradation negatively affects the performance of autonomous underwater vehicles used in marine applications. To address these challenges, we propose a lightweight phase-based transformer network with 1.77M parameters for underwater image restoration (UIR). Our approachfocuses on effectively extracting non-contaminated features using a phase-based self-attention mechanism. We also introduce an optimized phase attention block to restore structural information by propagating prominent attentive features from the input. We evaluate our method on both synthetic (UIEB, UFO-120) and real-world (UIEB, U45, UCCS, SQUID) underwater image datasets. Additionally, we demonstrate its effectiveness for low-light image enhancement using the LOL dataset. Through extensive ab-lation studies and comparative analysis, it is clear that the proposed approach outperforms existing state-of-the-art (SOTA) methods. Code is available at Phaseformer. Md Raqib Khan, Anshul Negi, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 5 |
| 2025 | USWformer: Efficient Sparse Wavelet Transformer for Underwater Image EnhancementabstractTransformer-based methods have shown great promise in underwater image enhancement (UIE) tasks due to their capability to model long-range dependencies, which are vital for reconstructing clear images. While numerous effective attention mechanisms have been devised to handle the computational requirements of transformers, they frequently incorporate redundant information and noisy interactions from irrelevant regions. Additionally, the current methods focusing solely on the raw pixel space constrains the exploration of the underwater image frequency dynamics, thus hindering the models from fully leveraging their potential for producing high-quality images. To address these challenges, we propose USWformer, an efficient UIE Sparse Wavelet Transformer Network (1.19 M parameters) to eliminate the redundant features in both the spatial and frequency domains. The USWformer consists of two fundamental components: a Sparse Wavelet Self-Attention (SWSA) block and a Multi-scale Wavelet Feed-Forward Network (MWFN). The SWSA block selectively preserves essential attention scores from the keys corresponding to each query, adjusting the feature details. MWFN further diminishes the feature redundancy in the aggregated features thereby improving the enhancement of the underwater images. We assess the efficacy of our approach across benchmark datasets comprising synthetic and real-world under-water images, showcasing its superiority via thorough ablation studies and comparative analyses. Nancy Mehta, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 3 |
| 2025 | Hierarchical motion magnification
Jasdeep Singh, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, G. Sankara Raju Kosuru, Hasan Al-Marzouqi |
Neurocomputing | 2 |
| 2025 | Learnable directional scale space filters for video motion magnification
Jasdeep Singh, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, G. Sankara Raju Kosuru, Hasan Al-Marzouqi |
Knowl. Based Syst. | 2 |
| 2025 | Cross-centroid ripple pattern for facial expression recognitionabstractAbstract In this paper, we propose a new feature descriptor Cross-Centroid Ripple Pattern (CRIP) for facial expression recognition. CRIP encodes the transitional pattern of a facial expression by incorporating a cross-centroid relationship between two ripples located at radius r 1 and r 2 respectively. These ripples are generated by dividing the local neighborhood region into subregions. Thus, CRIP has the ability to preserve macro and microstructural variations in an extensive region, which enables it to deal with side views and spontaneous expressions. Furthermore, gradient information between cross centroid ripples provides strength to capture prominent edge features in active patches: eyes, nose, and mouth, that define the disparities between different facial expressions. Cross-centroid information also provides robustness to irregular illumination. Moreover, CRIP utilizes the averaging behavior of pixels at subregions that yields robustness to deal with noisy conditions. The performance of the proposed descriptor is evaluated on seven comprehensive expression datasets consisting of challenging conditions such as age, pose, ethnicity, and illumination variations. The experimental results show that our descriptor consistently achieved a better accuracy rate as compared to existing state-of-the-art approaches. Monu Verma, Santosh Kumar Vipparthi |
Multim. Tools Appl. | 2 |
| 2025 | A novel approach for image retrieval in remote sensing using vision-language-based image caption generation
Prem Shanker Yadav, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi |
Multim. Tools Appl. | 3 |
| 2024 | Zero Reference based Low-light Enhancement with Wavelet OptimizationabstractImages captured in low light conditions usually suffer from poor visibility, a high amount of noise, and little information stored in the dark image, which has a negative impact on subsequent processing for outdoor computer vision applications. Presently, numerous deep learning based methods achieved superior performance with multi-exposure paired training data or additional information. However, obtaining multi-exposure data samples is a tedious task in real-time scenarios. To mitigate this challenge, we propose a zero reference based learnable wavelet approach without multi-exposure paired training data requirement for low-light image enhancement. Our proposed approach generates the low light image and learns to project an image into noise free similar looking image, then we enhance the image using retinex theory. Further, we have proposed learnable wavelet block to remove the hidden noise amplified while enhancement. We introduce Gaussian-based supervision to improve the smoothness of the image. Extensive experimental analysis on synthetic as well as real-world images, along with thorough ablation study demonstrate the effectiveness of our proposed method over the existing state-of-the-art methods for low-light image enhancement. The code is provided at https://github.com/vision-lab-sggsiet/Zero-Reference-based-Low-light-Enhancement-with-Wavelet-Optimization. Vivek Deshmukh, Adinath Madhavrao Dukre, Ashutosh Kulkarni, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Anil Balaji Gonde |
AVSS | 5 |
| 2024 | AeroDehazeNet: Exploiting Selective Multi-Scale Transformers for Aerial Image DehazingabstractRemote sensing is the task of analyzing and acquiring useful information from satellite images captured at a far distance from the earth’s surface. These images are vulnerable to degradation due to the presence of mist or haze. Existing methods either make use of prior information to estimate haze free images, or use CNN architectures based on generative adversarial networks (GANs) or Transformers. Though the state-of-the-art transformer-based architectures helped to dehaze the aerial images, they lacked the ability to capture multi-scale dependencies of the image. Identifying this shortcoming, we propose AeroDehazeNet based on a transformer that captures multi-scale dependencies along with global dependencies of the image. Our network comprises of three key components: (1) a multi-scale selective attention (MScA) network to attentively process the multi-scale information in an image, (2) residual attention network (RAN in feed forward network responsible for distilling non-degraded features passed from MScA, and (3) high frequency dominant skip connection (HFDS) block for passing diverse features (low frequency and high frequency) prominent with multi-scale edge features from encoder levels to adjacent decoder levels. The extensive quantitative and qualitative comparisons with existing methods on synthetic and realworld data plus exhaustive ablation study demonstrate the efficacy of our proposed network over transformer based state-of-the-art architectures with comparatively less number of parameters and FLOPs. Testing code is available at https://github.com/KartikGonde/AeroDehazeNet. Kartik Gonde, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Pramod Patil, Vinod V. Kimbahune |
AVSS | 3 |
| 2024 | RefMOS: A Robust Referred Moving Object Segmentation framework based on text queryabstractReferred Moving object segmentation is a very challenging task in automated video surveillance applications as it requires additional information to learn about object representation referred by natural language expression. In segmenting specific moving objects targeted by a text, suppressing other moving as well as stationary objects is a crucial task. A better context needs to be learned where linguistic, spatial, and temporal features need to be taken into account. In this work, we have proposed a robust referred moving object segmentation (RefMOS) framework to capture moving objects referred by text query. Most of the earlier state-of-the-art methods exploit a different type of supervision by treating video frames as images but lack temporal information during processing. In this work, we have proposed an inter-frame movement detector (IFCD) module, which extracts the movement information between the consecutive frames and helps integrate temporal information with spatial visual features. Language embedding is utilized to capture the information of referred moving objects in the text by extracting linguistic features from a pre-trained language model, i.e., BERT. Furthermore, the cross-entropy loss and SGD optimizer are used to train the network. Our RefMOS framework competes with the state-of-the-art approaches and achieves 48.6 mean IOU on the ref-DAVIS 17 dataset. Prafulla Saxena, Susim Mukul Roy, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Balasubramanian Raman |
AVSS | 4 |
| 2024 | Luminate: Linguistic Understanding and Multi-Granularity Interaction for Video Object SegmentationabstractReferring Video Object Segmentation (R-VOS) is a challenging task that involves segmenting objects in a video based on linguistic descriptions. In this paper, we introduce a novel multi-granularity referring video Object segmentation framework, termed as LUMINATE. The LUMINATE framework introduces a streamlined approach to cross-modal fusion. The proposed LUMINATE enhanced interaction between visual and textual modalities begins with cross-attention between the vision encoder’s query and the text encoder’s key-value pairs, and vice versa. The results are then concatenated with the respective queries of the vision and text encoders, fostering a comprehensive understanding of semantic relationships. The combined features are fed into the Transformer Encoder for further refinement and integration into the segmentation pipeline. Extensive experiments on benchmark datasets, including Ref-DAVIS, demonstrate that our proposed LUMINATE approach achieves better results than state-of-the-art methods in terms of Jaccard and F-measure evaluation metrics. Furthermore, the efficiency of our multi-object R-VOS variant is highlighted, achieving a threefold speed improvement while maintaining satisfactory segmentation performance. The proposed approach contributes to advancing the capabilities of R-VOS models, paving the way for improved multimodal reasoning and real-world applications. Rahul Tekchandani, Ritik Maheshwari, Praful Hambarde, Satya Narayan Tazi, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
ICIP | 5 |
| 2024 | Frequency Modulated Deformable Transformer for Underwater Image Enhancement
Adinath Madhavrao Dukre, Vivek Deshmukh, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, Anil Balaji Gonde, M. Subrahmanyam 0001 |
ICPR (32) | 5 |
| 2024 | Probing Attention-Driven Normalizing Flow Network for Low-Light Image Enhancement
Nancy Mehta, K. N. Prakash, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
ICPR (32) | 4 |
| 2024 | Attentive Color Fusion Transformer Network (ACFTNet) for Underwater Image Enhancement
Mohd Ubaid Wani, Md Raqib Khan, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
ICPR (21) | 5 |
| 2024 | Fusing Image and Text Features for Scene Sentiment Analysis Using Whale-Honey Badger Optimization Algorithm (WHBOA)
Prem Shanker Yadav, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi |
ICPR (2) | 3 |
| 2024 | Spectroformer: Multi-Domain Query Cascaded Transformer Network For Underwater Image EnhancementabstractUnderwater images often suffer from color distortion, haze, and limited visibility due to light refraction and absorption in water. These challenges significantly impact autonomous underwater vehicle applications, necessitating efficient image enhancement techniques. To address these challenges, we propose a Multi-Domain Query Cascaded Transformer Network for underwater image enhancement. Our approach includes a novel Multi-Domain Query Cascaded Attention mechanism that integrates localized transmission features and global illumination features. To improve feature propagation from the encoder to the decoder, we propose a Spatio-Spectro Fusion-Based Attention Block. Additionally, we introduce a Hybrid Fourier-Spatial Up-sampling Block, which uniquely combines Fourier and spatial upsampling techniques to enhance feature resolution effectively. We evaluate our method on benchmark synthetic and real-world underwater image datasets, demonstrating its superiority through extensive ablation studies and comparative analysis. The testing code is available at: https://github.com/Mdraqibkhan/Spectroformer. Md Raqib Khan, Nancy Mehta, Shruti S. Phutke, Santosh Kumar Vipparthi, Sukumar Nandi, M. Subrahmanyam 0001 |
WACV | 5 |
| 2024 | C2AIR: Consolidated Compact Aerial Image Haze RemovalabstractAerial image haze removal deals with improving the visibility and quality of images captured from aerial platforms, such as drones and satellites. Aerial images are commonly used in various applications such as environmental monitoring, and disaster response. These applications usually require cleaner data for accurate functioning. However, atmospheric conditions such as haze or fog can significantly degrade the quality of these images, reducing their contrast, color saturation, and sharpness, making it difficult to extract meaningful information from them. Existing methods rely on computationally heavy and haze density (light, moderate, dense) specific architectures for aerial image dehazing. In light of these limitations, we propose a novel lightweight and consolidated approach for aerial image dehazing. In this approach, we propose Density Aware Query Modulated Block for learning weather degradations in input features and guiding the restoration process. Further, we propose Cross Collaborative Feed-Forward Block for learning to restore varying sizes of the structures in the input images. Finally, we propose Gated Adaptive Feature Fusion block to achieve inter-scale and intra-feature attentive fusion, effective for aerial image restoration. Extensive analysis on benchmark aerial image dehazing datasets and real-world images, along with detailed ablation studies validate the effectiveness of the proposed approach. Further, we have analysed our method for other restoration task such as underwater image enhancement to experiment its wide applicability. The code is available at https://github.com/AshutoshKulkarni4998/C2AIR. Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 3 |
| 2024 | Triplet-set feature proximity learning for video anomaly detection
Kuldeep Marotirao Biradar, Murari Mandal, Sachin Dube, Santosh Kumar Vipparthi, Dinesh Kumar Tyagi |
Image Vis. Comput. | 4 |
| 2023 | RNAS-MER: A Refined Neural Architecture Search with Hybrid Spatiotemporal Operations for Micro-Expression RecognitionabstractExisting neural architecture search (NAS) methods comprise linear connected convolution operations and use ample search space to search task-driven convolution neural networks (CNN). These CNN models are computationally expensive and diminish the quality of receptive fields for tasks like micro-expression recognition (MER) with limited training samples. Therefore, we propose a refined neural architecture search strategy to search for a tiny CNN architecture for MER. In addition, we introduced a refined hybrid module (RHM) for inner-level search space and an optimal path explore network (OPEN) for outer-level search space. The RHM focuses on discovering optimal cell structures by incorporating a multilateral hybrid spatiotemporal operation space. Also, spatiotemporal attention blocks are embedded to refine the aggregated cell features. The OPEN search space aims to trace an optimal path between the cells to generate a tiny spatiotemporal CNN architecture instead of covering all possible tracks. The aggregate mix of RHM and OPEN search space availed the NAS method to robustly search and design an effective and efficient framework for MER. Compared with contemporary works, experiments reveal that the RNAS-MER is capable of bridging the gap between NAS algorithms and MER tasks. Furthermore, RNAS-MER achieves new state-of-the-art performances on challenging MER benchmarks, including 0.8511%, 0.7620%, 0.9078% and 0.8235% UAR on COMPOSITE, SMIC, CASME-II and SAMM datasets respectively. Monu Verma, Priyanka Lubal, Santosh Kumar Vipparthi, Mohamed Abdel-Mottaleb |
WACV | 3 |
| 2023 | SBI-DHGR: Skeleton-based intelligent dynamic hand gestures recognition
Satya Narayan, Arka Prokash Mazumdar, Santosh Kumar Vipparthi |
Expert Syst. Appl. | 3 |
| 2023 | Efficient neural architecture search for emotion recognition
Monu Verma, Murari Mandal, M. Satish Kumar Reddy, Yashwanth Reddy Meedimale, Santosh Kumar Vipparthi |
Expert Syst. Appl. | 5 |
| 2023 | HyFiNet: Hybrid feature attention network for hand gesture recognition
Gopa Bhaumik, Monu Verma, Mahesh Chandra Govil, Santosh Kumar Vipparthi |
Multim. Tools Appl. | 4 |
| 2022 | HYPE: CNN Based HYbrid PrEcoding Framework for 5G and Beyond
Deepti Sharma 0003, Kuldeep Marotirao Biradar, Santosh Kumar Vipparthi, Ramesh Babu Battula |
AINA (2) | 3 |
| 2022 | RIChEx: A Robust Inter-Frame Change Exposure for Segmenting Moving ObjectsabstractMoving object segmentation in video plays an important role in many computer vision applications such as outdoor video surveillance, robotics, self-driving cars, etc. Existing work comprises complicated training procedures considering more number of history frames, fine-tuning on test sequences, and pre-trained models adopted from other applications. Therefore, this paper proposes an end-to-end robust inter-frame change exposure (RIChEx) framework for segmenting moving objects using three consecutive frames. The RIChEx consists inter-frame change gradient (IFCG) aggregation approach responsible for extracting the relevant foreground features to capture substantial changes and spatio-temporal structural dependencies between the consecutive frames. This module consists of multi-scale addition (MSA) blocks for extracting multi-scale features in early layers. Furthermore, to enhance the temporal learning capability of the module, previous recurrent responses and optical flow information are integrated to estimate probable foreground availability and salient object motions. We have evaluated the proposed framework on the benchmark DAVIS-16 dataset to validate the model’s performance. Prafulla Saxena, Kuldeep Marotirao Biradar, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi |
ICIP | 4 |
| 2022 | Scene Independency Matters: An Empirical Study of Scene Dependent and Scene Independent Evaluation for CNN-Based Change DetectionabstractVisual change detection in video is one of the essential tasks in computer vision applications. Recently, a number of supervised deep learning methods have achieved top performance over the benchmark datasets for change detection. However, inconsistent training-testing data division schemes adopted by these methods have led to documentation of incomparable results. We address this crucial issue through our own propositions for benchmark comparative analysis. The existing works have evaluated the model in scene dependent evaluation setup which makes it difficult to assess the generalization capability of the model in completely unseen videos. It also leads to inflated results. Therefore, in this paper, we present a completely scene independent evaluation strategy for a comprehensive analysis of the model design for change detection. We propose well-defined scene independent and scene dependent experimental frameworks for training and evaluation over the benchmark CDnet 2014, LASIESTA and SBMI2015 datasets. A cross-data evaluation is performed with PTIS dataset to further measure the robustness of the models. We designed a fast and lightweight online end-to-end convolutional network called ChangeDet (speed-58.8 fps and model size-1.59 MB) in order to achieve robust performance in completely unseen videos. The ChangeDet estimates the background through a sequence of maximum multi-spatial receptive feature (MMSR) blocks using past temporal history. The contrasting features are produced through the assimilation of temporal median and contemporary features from the current frame. Further, these features are processed through an encoder-decoder to detect pixel-wise changes. The proposed ChangeDet outperforms the existing state-of-the-art methods in all four benchmark datasets. Murari Mandal, Santosh Kumar Vipparthi |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | An Empirical Review of Deep Learning Frameworks for Change Detection: Model Design, Experimental Frameworks, Challenges and Research NeedsabstractVisual change detection, aiming at segmentation of video frames into foreground and background regions, is one of the elementary tasks in computer vision and video analytics. The applications of change detection include anomaly detection, object tracking, traffic monitoring, human machine interaction, behavior analysis, action recognition, and visual surveillance. Some of the challenges in change detection include background fluctuations, illumination variation, weather changes, intermittent object motion, shadow, fast/slow object motion, camera motion, heterogeneous object shapes and real-time processing. Traditionally, this problem has been solved using hand-crafted features and background modelling techniques. In recent years, deep learning frameworks have been successfully adopted for robust change detection. This article aims to provide an empirical review of the state-of-the-art deep learning methods for change detection. More specifically, we present a detailed analysis of the technical characteristics of different model designs and experimental frameworks. We provide model design based categorization of the existing approaches, including the 2D-CNN, 3D-CNN, ConvLSTM, multi-scale features, residual connections, autoencoders and GAN based methods. Moreover, an empirical analysis of the evaluation settings adopted by the existing deep learning methods is presented. To the best of our knowledge, this is a first attempt to comparatively analyze the different evaluation frameworks used in the existing deep change detection methods. Finally, we point out the research needs, future directions and draw our own conclusions. Murari Mandal, Santosh Kumar Vipparthi |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | AutoMER: Spatiotemporal Neural Architecture Search for Microexpression RecognitionabstractFacial microexpressions offer useful insights into subtle human emotions. This unpremeditated emotional leakage exhibits the true emotions of a person. However, the minute temporal changes in the video sequences are very difficult to model for accurate classification. In this article, we propose a novel spatiotemporal architecture search algorithm, AutoMER for microexpression recognition (MER). Our main contribution is a new parallelogram design-based search space for efficient architecture search. We introduce a spatiotemporal feature module named 3-D singleton convolution for cell-level analysis. Furthermore, we present four such candidate operators and two 3-D dilated convolution operators to encode the raw video sequences in an end-to-end manner. To the best of our knowledge, this is the first attempt to discover 3-D convolutional neural network (CNN) architectures with a network-level search for MER. The searched models using the proposed AutoMER algorithm are evaluated over five microexpression data sets: CASME-I, SMIC, CASME-II, CAS(ME) ∧2 , and SAMM. The proposed generated models quantitatively outperform the existing state-of-the-art approaches. The AutoMER is further validated with different configurations, such as downsampling rate factor, multiscale singleton 3-D convolution, parallelogram, and multiscale kernels. Overall, five ablation experiments were conducted to analyze the operational insights of the proposed AutoMER. Monu Verma, M. Satish Kumar Reddy, Yashwanth Reddy Meedimale, Murari Mandal, Santosh Kumar Vipparthi |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | ExtriDeNet: an intensive feature extrication deep network for hand gesture recognition
Gopa Bhaumik, Monu Verma, Mahesh Chandra Govil, Santosh Kumar Vipparthi |
Vis. Comput. | 4 |
| 2021 | One for All: An End-to-End Compact Solution for Hand Gesture RecognitionabstractThe HGR is a quite challenging task as its performance is influenced by various aspects such as illumination variations, cluttered backgrounds, spontaneous capture, etc. The conventional CNN networks for HGR are following two stage pipeline to deal with the various challenges: complex signs, illumination variations, complex and cluttered backgrounds. The existing approaches needs expert expertise as well as auxiliary computation at stage 1 to remove the complexities from the input images. Therefore, in this paper, we proposes an novel end-to-end compact CNN framework: fine grained feature attentive network for hand gesture recognition (Fit-Hand) to solve the challenges as discussed above. The pipeline of the proposed architecture consists of two main units: FineFeat module and dilated convolutional (Conv) layer. The FineFeat module extracts fine grained feature maps by employing attention mechanism over multiscale receptive fields. The attention mechanism is introduced to capture effective features by enlarging the average behaviour of multi-scale responses. Moreover, dilated convolution provides global features of hand gestures through a larger receptive field. In addition, integrated layer is also utilized to combine the features of FineFeat module and dilated layer which enhances the discriminability of the network by capturing complementary context information of hand postures. The effectiveness of Fit-Hand is evaluated by using subject dependent (SD) and subject independent (SI) validation setup over seven benchmark datasets: MUGD-I, MUGD-II, MUGD-III, MUGD-IV, MUGD-V, Finger Spelling and OUHANDS, respectively. Furthermore, to investigate the deep insights of the proposed Fit-Hand framework, we performed ten ablation study Monu Verma, Santosh Kumar Vipparthi |
IJCNN | 3 |
| 2021 | 3DCD: Scene Independent End-to-End Spatiotemporal Feature Learning Framework for Change Detection in Unseen VideosabstractChange detection is an elementary task in computer vision and video processing applications. Recently, a number of supervised methods based on convolutional neural networks have reported high performance over the benchmark dataset. However, their success depends upon the availability of certain proportions of annotated frames from test video during training. Thus, their performance on completely unseen videos or scene independent setup is undocumented in the literature. In this work, we present a scene independent evaluation (SIE) framework to test the supervised methods in completely unseen videos to obtain generalized models for change detection. In addition, a scene dependent evaluation (SDE) is also performed to document the comparative analysis with the existing approaches. We propose a fast (speed-25 fps) and lightweight (0.13 million parameters, model size-1.16 MB) end-to-end 3D-CNN based change detection network (3DCD) with multiple spatiotemporal learning blocks. The proposed 3DCD consists of a gradual reductionist block for background estimation from past temporal history. It also enables motion saliency estimation, multi-schematic feature encoding-decoding, and finally foreground segmentation through several modular blocks. The proposed 3DCD outperforms the existing state-of-the-art approaches evaluated in both SIE and SDE setup over the benchmark CDnet 2014, LASIESTA and SBMI2015 datasets. To the best of our knowledge, this is a first attempt to present results in clearly defined SDE and SIE setups in three change detection datasets. Murari Mandal, Vansh Dhar, Santosh Kumar Vipparthi, Mohamed Abdel-Mottaleb |
IEEE Trans. Image Process. | 4 |
| 2020 | Non-Linearities Improve OrigiNet based on Active Imaging for Micro Expression RecognitionabstractMicro expression recognition (MER)is a very challenging task as the expression lives very short in nature and demands feature modeling with the involvement of both spatial and temporal dynamics. Existing MER systems exploit CNN networks to spot the significant features of minor muscle movements and subtle changes. However, existing networks fail to establish a relationship between spatial features of facial appearance and temporal variations of facial dynamics. Thus, these networks were not able to effectively capture minute variations and subtle changes in expressive regions. To address these issues, we introduce an active imaging concept to segregate active changes in expressive regions of a video into a single frame while preserving facial appearance information. Moreover, we propose a shallow CNN network: hybrid local receptive field based augmented learning network (OrigiNet) that efficiently learns significant features of the micro-expressions in a video. In this paper, we propose a new refined rectified linear unit (RReLU), which overcome the problem of vanishing gradient and dying ReLU. RReLU extends the range of derivatives as compared to existing activation functions. The RReLU not only injects a nonlinearity but also captures the true edges by imposing additive and multiplicative property. Furthermore, we present an augmented feature learning block to improve the learning capabilities of the network by embedding two parallel fully connected layers. The performance of proposed OrigiNet is evaluated by conducting leave one subject out experiments on four comprehensive ME datasets. The experimental results demonstrate that OrigiNet outperformed state-of-the-art techniques with less computational complexity. Monu Verma, Santosh Kumar Vipparthi, Girdhari Singh |
IJCNN | 2 |
| 2020 | MOR-UAV: A Benchmark Dataset and Baselines for Moving Object Recognition in UAV VideosabstractVisual data collected from Unmanned Aerial Vehicles (UAVs) has opened a new frontier of computer vision that requires automated analysis of aerial images/videos. However, the existing UAV datasets primarily focus on object detection. An object detector does not differentiate between the moving and non-moving objects. Given a real-time UAV video stream, how can we both localize and classify the moving objects, i.e. perform moving object recognition (MOR) The MOR is one of the essential tasks to support various UAV vision-based applications including aerial surveillance, search and rescue, event recognition, urban and rural scene understanding.To the best of our knowledge, no labeled dataset is available for MOR evaluation in UAV videos. Therefore, in this paper, we introduce MOR-UAV, a large-scale video dataset for MOR in aerial videos. We achieve this by labeling axis-aligned bounding boxes for moving objects which requires less computational resources than producing pixel-level estimates. We annotate 89,783 moving object instances collected from 30 UAV videos, consisting of 10,948 frames in various scenarios such as weather conditions, occlusion, changing flying altitude and multiple camera views. We assigned the labels for two categories of vehicles (car and heavy vehicle). Furthermore, we propose a deep unified framework MOR-UAVNet for MOR in UAV videos. Since, this is a first attempt for MOR in UAV videos, we present 16 baseline results based on the proposed framework over the MOR-UAV dataset through quantitative and qualitative experiments. We also analyze the motion-salient regions in the network through multiple layer visualizations. The MOR-UAVNet works online at inference as it requires only few past frames. Moreover, it doesn't require predefined target initialization from user. Experiments also demonstrate that the MOR-UAV dataset is quite challenging. Murari Mandal, Lav Kush Kumar, Santosh Kumar Vipparthi |
ACM Multimedia | 3 |
| 2020 | MotionRec: A Unified Deep Framework for Moving Object RecognitionabstractIn this paper we present a novel deep learning framework to perform online moving object recognition (MOR) in streaming videos. The existing methods for moving object detection (MOD) only computes class-agnostic pixel-wise binary segmentation of video frames. On the other hand, the object detection techniques do not differentiate between static and moving objects. To the best of our knowledge, this is a first attempt for simultaneous localization and classification of moving objects in a video, i.e. MOR in a single-stage deep learning framework. We achieve this by labelling axis-aligned bounding boxes for moving objects which requires less computational resources than producing pixel-level estimates. In the proposed MotionRec, both temporal and spatial features are learned using past history and current frames respectively. First, the background is estimated with a temporal depth reductionist (TDR) block. Then the estimated background, current frame and temporal median of recent observations are assimilated to encode spatiotemporal motion saliency. Moreover, feature pyramids are generated from these motion saliency maps to perform regression and classification at multiple levels of feature abstractions. MotionRec works online at inference as it requires only few past frames for MOR. Moreover, it doesn't require predefined target initialization from user. We also annotated axis-aligned bounding boxes (42,614 objects (14,814 cars and 27,800 person) in 24,923 video frames of CDnet 2014 dataset) due to lack of available benchmark datasets for MOR. The performance is observed qualitatively and quantitatively in terms of mAP over a defined unseen test set. Experiments show that the proposed MotionRec significantly improves over strong baselines with RetinaNet architectures for MOR. Murari Mandal, Lav Kush Kumar, Mahipal Singh Saran, Santosh Kumar Vipparthi |
WACV | 4 |
| 2020 | AVDNet: A Small-Sized Vehicle Detection Network for Aerial Visual DataabstractDetection of small-sized targets in aerial views is a challenging task due to the smallness of vehicle size, complex background, and monotonic object appearances. In this letter, we propose a one-stage vehicle detection network (AVDNet) to robustly detect small-sized vehicles in aerial scenes. In AVDNet, we introduced ConvRes residual blocks at multiple scales to alleviate the problem of vanishing features for smaller objects caused because of the inclusion of deeper convolutional layers. These residual blocks, along with enlarged output feature map, ensure the robust representation of the salient features for small-sized objects. Furthermore, we proposed a recurrent-feature aware visualization (RFAV) technique to analyze the network behavior. We also created a new airborne image data set (ABD) by annotating 1396 new objects in 79 aerial images for our experiments. The effectiveness of AVDNet is validated on VEDAI, DLR-3K, DOTA, and the combined (VEDAI, DLR-3K, DOTA, and ABD) data set. Experimental results demonstrate the significant performance improvement of the proposed method over state-of-the-art detection techniques in terms of mAP, computation, and space complexity. Murari Mandal, Manal Shah, Prashant Meena, Sanhita Devi, Santosh Kumar Vipparthi |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2020 | LEARNet: Dynamic Imaging Network for Micro Expression RecognitionabstractUnlike prevalent facial expressions, micro expressions have subtle, involuntary muscle movements which are short-lived in nature. These minute muscle movements reflect true emotions of a person. Due to the short duration and low intensity, these micro-expressions are very difficult to perceive and interpret correctly. In this paper, we propose the dynamic representation of micro-expressions to preserve facial movement information of a video in a single frame. We also propose a Lateral Accretive Hybrid Network (LEARNet) to capture micro-level features of an expression in the facial region. The LEARNet refines the salient expression features in accretive manner by incorporating accretion layers (AL) in the network. The response of the AL holds the hybrid feature maps generated by prior laterally connected convolution layers. Moreover, LEARNet architecture incorporates the cross decoupled relationship between convolution layers which helps in preserving the tiny but influential facial muscle change information. The visual responses of the proposed LEARNet depict the effectiveness of the system by preserving both high- and micro-level edge features of facial expression. The effectiveness of the proposed LEARNet is evaluated on four benchmark datasets: CASME-I, CASME-II, CAS(ME)'2 and SMIC. The experimental results after investigation show a significant improvement of 4.03%, 1.90%, 1.79% and 2.82% as compared with ResNet on CASME-I, CASME-II, CAS(ME)'2 and SMIC datasets respectively. Monu Verma, Santosh Kumar Vipparthi, Girdhari Singh, M. Subrahmanyam 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | SSSDET: Simple Short and Shallow Network for Resource Efficient Vehicle Detection in Aerial ScenesabstractDetection of small-sized targets is of paramount importance in many aerial vision-based applications. The commonly deployed low cost unmanned aerial vehicles (UAVs) for aerial scene analysis are highly resource constrained in nature. In this paper we propose a simple short and shallow network (SSSDet) to robustly detect and classify small-sized vehicles in aerial scenes. The proposed SSSDet is up to 4× faster, requires 4.4× less FLOPs, has 30× less parameters, requires 31× less memory space and provides better accuracy in comparison to existing state-of-the-art detectors. Thus, it is more suitable for hardware implementation in real-time applications. We also created a new airborne image dataset (ABD) by annotating 1396 new objects in 79 aerial images for our experiments. The effectiveness of the proposed method is validated on the existing VEDAI, DLR-3K, DOTA and Combined dataset. The SSSDet outperforms state-of-the-art detectors in term of accuracy, speed, compute and memory efficiency. Murari Mandal, Manal Shah, Prashant Meena, Santosh Kumar Vipparthi |
ICIP | 4 |
| 2019 | ANTIC: antithetic isomeric cluster patterns for medical image retrieval and change detectionabstractIn this study, new feature descriptors are designed for medical image retrieval and change detection applications, respectively. Inspired by isomerism, the authors propose a novel feature descriptor named antithetic isomeric cluster pattern (ANTIC). The ANTIC is defined by the two properties: cluster patterns and antithetic isomerism (ANTI). The cluster pattern corresponds to successive pixel intensity differences at antithetical orientations. Furthermore, the ANTI is characterised by two aspects: first, the clusters are oppositely oriented (antithetical) to each other and second, both adhere to a defined isomeric property. The ANTIC identifies the lines and corner point information in the local neighbourhood across various directions. To attain enhanced robustness, they further proposed multiresolution ANTIC by integrating the multiresolution Gaussian filter. Moreover, to reduce the feature dimensionality, they extended their work to rotation invariant features. The proposed method outperforms other widely used feature descriptors in biomedical and retinopathy image retrieval applications. In addition, they extracted spatiotemporal features by designing intra‐ANTIC and inter‐ANTIC to detect motion changes in video sequences. They validated the effectiveness of these features by conducting experiments on CDNet 2014 dataset. The proposed descriptor achieves better performance in various challenging conditions for change detection as compared to other state‐of‐the‐art techniques. Murari Mandal, Mallika Chaudhary, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Anil Balaji Gonde, Shyam Krishna Nagar |
IET Comput. Vis. | 3 |
| 2019 | SOD-CED: salient object detection for noisy images using convolution encoder-decoderabstractDuring the last decade, there has been profound progress in the field of visual saliency. However, there still exist various major challenges that hinder the detection performance for scenes with complex composition, presence of additive noise, objects of diverse scale and rotations etc. Generally, images with additive noise have low spatial resolution and blurred edges, which affects the learning capability of the network and causes inaccurate detection. In order to address these issues, in this study, the authors propose a fully convolutional neural network which jointly denoise the input maps by learning edges and contrast details, followed by learning of residing salient details via colour spatial maps in an end‐to‐end fashion. Their framework employs convolutional layers that use gradient and contrast details of images to denoise the areas with high edge density. After denoising, the denoised images are subjected to salient object detection (SOD) using convolutional layers. The effectiveness of the proposed network is evaluated on benchmark datasets. The experimental results demonstrate the significant performance improvement of the proposed method over state‐of‐the‐art detection techniques. Maheep Singh, Mahesh Chandra Govil, Emmanuel S. Pilli, Santosh Kumar Vipparthi |
IET Comput. Vis. | 4 |
| 2019 | Regional adaptive affinitive patterns (RADAP) with logical operators for facial expression recognitionabstractAutomated facial expression recognition plays a significant role in the study of human behaviour analysis. In this study, the authors propose a robust feature descriptor named regional adaptive affinitive patterns (RADAP) for facial expression recognition. The RADAP computes positional adaptive thresholds in the local neighbourhood and encodes multi‐distance magnitude features which are robust to intra‐class variations and irregular illumination variation in an image. Furthermore, they established cross‐distance co‐occurrence relations in RADAP by using logical operators. They proposed XRADAP, ARADAP, and DRADAP using xor, adder and decoder, respectively. The XRADAP engrains the quality of robustness to intra‐class variations in RADAP features using pairwise co‐occurrence. Similarly, ARADAP and DRADAP extract more stable and illumination invariant features and capture the minute expression features which are usually missed by regular descriptors. The performance of the proposed methods is evaluated by conducting experiments on nine benchmark datasets Cohn–Kanade+ (CK+), Japanese female facial expression (JAFFE), Multimedia Understanding Group (MUG), MMI, OULU‐CASIA, Indian spontaneous expression database, DISFA, AFEW and Combined (CK+, JAFFE, MUG, MMI & GEMEP‐FERA) database in both person dependent and person independent setup. The experimental results demonstrate the effectiveness of the proposed method over state‐of‐the‐art approaches. Murari Mandal, Monu Verma, Sonakshi Mathur, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Kranthi Kumar Deveerasetty |
IET Image Process. | 4 |
| 2019 | 3DFR: A Swift 3D Feature Reductionist Framework for Scene Independent Change DetectionabstractIn this paper we propose an end-to-end swift 3D feature reductionist framework (3DFR) for scene independent change detection. The 3DFR framework consists of three feature streams: a swift 3D feature reductionist stream (AvFeat), a contemporary feature stream (ConFeat) and a temporal median feature map. These multilateral foreground/background features are further refined through an encoder-decoder network. As a result, the proposed framework not only detects temporal changes but also learns high-level appearance features. Thus, it incorporates the object semantics for effective change detection. Furthermore, the proposed framework is validated through a scene independent evaluation scheme in order to demonstrate the robustness and generalization capability of the network. The performance of the proposed method is evaluated on the benchmark CDnet 2014 dataset. The experimental results show that the proposed 3DFR network outperforms the state-of-the-art approaches. Murari Mandal, Vansh Dhar, Santosh Kumar Vipparthi |
IEEE Signal Process. Lett. | 4 |
| 2018 | CANDID: Robust Change Dynamics and Deterministic Update Policy for Dynamic Background SubtractionabstractBackground subtraction in video provides the preliminary information essential for many computer vision applications. In this paper, we propose a sequence of approaches named CANDID to solve the change detection problem in challenging video scenarios. The CANDID adaptively initializes the pixel-level distance threshold and update rate. These parameters are updated by computing the change dynamics at a location. Further, the background model is maintained by formulating a deterministic update policy. The performance of the proposed method is evaluated over various challenging scenarios such as dynamic background and extreme weather conditions. The qualitative and quantitative measures of the proposed method outperform the existing state-of-the-art approaches. Murari Mandal, Prafulla Saxena, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
ICPR | 3 |
| 2018 | QUEST: Quadriletral Senary Bit Pattern for Facial Expression RecognitionabstractFacial expression has significant role to analyzing human cognitive state. Deriving an accurate facial appearance representation is critical task for an automatic facial expression recognition application. This paper provides a new feature descriptor named as Quadrilateral Senary bit Pattern for facial expression recognition. The QUEST pattern encoded the intensity changes by emphasizing relationship between neighboring and reference pixels by dividing them into two quadrilaterals in local neighborhood. Thus, the resultant gradient edges reveal the transitional variation information, that improves the classification rate by discriminating expression classes. Moreover, it also enhances the capability of the descriptor to deal with view point variations and illumination changes. The trine relationship in quadrilateral structure helps to extract the expressive edges and suppressing noise elements to enhance the robustness to noisy conditions. The QUEST pattern generates a six-bit compact code, which improve the efficiency of the FER system with more discriminability. The effectiveness of proposed method is evaluated by conducting several experiments on four benchmark datasets: MMI, GEMEP-FERA, OULU-CASIA and ISED. The experimental results show better performance of the proposed method as compared to existing state-art-the approaches. Monu Verma, Prafulla Saxena, Santosh Kumar Vipparthi, Girdhari Singh |
SMC | 3 |
| 2016 | Local directional mask maximum edge patterns for image retrieval and face recognitionabstractThis study proposes a new feature descriptor, local directional mask maximum edge pattern, for image retrieval and face recognition applications. Local binary pattern (LBP) and LBP variants collect the relationship between the centre pixel and its surrounding neighbours in an image. Thus, LBP based features are very sensitive to the noise variations in an image. Whereas the proposed method collects the maximum edge patterns (MEP) and maximum edge position patterns (MEPP) from the magnitude directional edges of face/image. These directional edges are computed with the aid of directional masks. Once the directional edges (DE) are computed, the MEP and MEPP are coded based on the magnitude of DE and position of maximum DE. Further, the robustness of the proposed method is increased by integrating it with the multiresolution Gaussian filters. The performance of the proposed method is tested by conducting four experiments onopen access series of imaging studies‐magnetic resonance imaging, Brodatz, MIT VisTex and Extended Yale B databases for biomedical image retrieval, texture retrieval and face recognition applications. The results after being investigated the proposed method shows a significant improvement as compared with LBP and LBP variant features in terms of their evaluation measures on respective databases. Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Anil Balaji Gonde, Q. M. Jonathan Wu |
IET Comput. Vis. | 1 |
| 2015 | Local Gabor maximum edge position octal patterns for image retrieval
Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Shyam Krishna Nagar, Anil Balaji Gonde |
Neurocomputing | 1 |
| 2014 | Expert image retrieval system using directional local motif XoR patterns
Santosh Kumar Vipparthi, Shyam Krishna Nagar |
Expert Syst. Appl. | 1 |