VLDB 2026 Research / reviewers in the wild / expert
Monu Verma
dblp:223/9848
· DBLP profile ↗
16ranked-venue papers
11as first author
12since 2021 · last 2026
0000-0003-4962-882XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Level Volumetric Transformer for Early Prediction of Response to Neoadjuvant Chemotherapy in Locally Advanced Breast CancerabstractNeoadjuvant Chemotherapy (NAC) is a standard treatment for locally advanced breast cancer, where achieving a Pathological Complete Response (pCR) is the primary goal. Early identification of patient response is vital for personalizing treatment and avoiding unnecessary toxicity. Recent advances in deep learning have produced models predicting NST effectiveness using breast MRI scans. However, many of these models rely on time-consuming tumor segmentation, requiring expert involvement. To address these challenges and enhance clinical applicability, we propose a Multi-Level Volumetric Transformer (MVT-Former), a novel model that predicts NAC response directly from non-segmented, full-field breast MRI data combined with relevant clinical information. The primary novelty of this work lies in its specialized dual-transformer design: (1) the Multi-Level Convolutional Spatial Transformer (MLCS-Former), which utilizes multi-scale convolutions and a Global Convolutional Attention (GCA) mechanism to extract fine-grained textural and morphological features from 2D MRI slices without manual annotations; and (2) the Volume Feature Learning Transformer (VFL-Former), which captures 3D structural changes and long-range dependencies across the entire MRI volume. We evaluated the MVT-Former on the I-SPY-1 TRIAL dataset and results demonstrate that the proposed model outperforms state-of-the-art methods, achieving superior performance across key metrics, including area under the curve, accuracy, sensitivity, and specificity. Monu Verma, Fernando Collado-Mesa, Mohamed Abdel-Mottaleb |
ACM Trans. Comput. Heal. | 1 |
| 2026 | A motion flow guided MicroNet framework for micro expression recognition
Monu Verma, Santosh Kumar Vipparthi, Mohamed Abdel-Mottaleb |
J. Vis. Commun. Image Represent. | 1 |
| 2026 | FedHC: Enhanced federated learning with Hessian and cosine correlation for proximal correlation
Kushall Singh, Monu Verma, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi, G. Shankara Raju Kosuru, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb |
Knowl. Based Syst. | 2 |
| 2026 | FedHMed: Adaptive progressive loss and KL-divergence regularization for federated heterogeneous medical image classification tasks
Kushall Singh, Monu Verma, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb |
Knowl. Based Syst. | 2 |
| 2026 | ME-NAS: A Micro Expression Feature Adaptive Neural Architecture SearchabstractConvolution neural networks (CNN) have emerged as a prevailing paradigm for micro-expression recognition (MER) yet, it is inefficient and time-intensive to design optimal CNN-based MER models manually. In recent times, the neural architecture search (NAS) has garnered attention due to its automatic CNN architecture searching ability. However, the performance of NAS in MER is limited by challenges such as rapid duration, subtle intensity, and a mismatch between architecture and cell-level search. The existing search space, which stacks 12 cells with 3 transition paths (downsample, upsample, and same resolution), creates deep networks that may diminish minute spatiotemporal features due to progressive convolution and pooling. Therefore, motivated by these factors, in this article, we introduce a novel approach, the Micro-Expression Feature Adaptive NAS (ME-NAS), to analyze true human emotions through MER. While NAS has gained attention for its automatic CNN architecture search ability, its application in MER faces challenges due to ingrained challenges (rapid duration, subtle and low intensity) and the discrepancy between architecture and cell-level search. The existing NAS architecture search space is designed by stacking 12 cells with 3 transition paths (downsample, upsample, and same resolution), resulting in a deep network. Such deep networks may diminish minute spatiotemporal features due to the progressive convolution and pooling operations. Motivated by these factors, we designed a new NAS algorithm: ME-NAS. The ME-NAS comprises f (EXPERT) in architecture search, along with refined and complementary feature derivative (ReCODE) operations in cell-level search. The EXPERT aims to trace the optimal paths instead of covering all possible paths between cells. The ReCODE operations capture micro-level variations from spatial and temporal domains by introducing 24 3D convolution operations. The proposed ReCODE and EXPERT search space jointly lead to the search for a robust and shallow CNN architecture for micro-expressions (MEs). The proposed ME-NAS is evaluated on six datasets: CASME-I, CASME-II, CAS(ME) \({}^{2}\) , SAMM, SMIC, and MEGC-19 composite, with two evaluation strategies: LOSO and cross-domain, respectively. The experimental results manifest that the proposed ME-NAS outperformed the state-of-the-art approaches on both evaluation strategies. Monu Verma, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2025 | Cross-centroid ripple pattern for facial expression recognitionabstractAbstract In this paper, we propose a new feature descriptor Cross-Centroid Ripple Pattern (CRIP) for facial expression recognition. CRIP encodes the transitional pattern of a facial expression by incorporating a cross-centroid relationship between two ripples located at radius r 1 and r 2 respectively. These ripples are generated by dividing the local neighborhood region into subregions. Thus, CRIP has the ability to preserve macro and microstructural variations in an extensive region, which enables it to deal with side views and spontaneous expressions. Furthermore, gradient information between cross centroid ripples provides strength to capture prominent edge features in active patches: eyes, nose, and mouth, that define the disparities between different facial expressions. Cross-centroid information also provides robustness to irregular illumination. Moreover, CRIP utilizes the averaging behavior of pixels at subregions that yields robustness to deal with noisy conditions. The performance of the proposed descriptor is evaluated on seven comprehensive expression datasets consisting of challenging conditions such as age, pose, ethnicity, and illumination variations. The experimental results show that our descriptor consistently achieved a better accuracy rate as compared to existing state-of-the-art approaches. Monu Verma, Santosh Kumar Vipparthi |
Multim. Tools Appl. | 1 |
| 2023 | RNAS-MER: A Refined Neural Architecture Search with Hybrid Spatiotemporal Operations for Micro-Expression RecognitionabstractExisting neural architecture search (NAS) methods comprise linear connected convolution operations and use ample search space to search task-driven convolution neural networks (CNN). These CNN models are computationally expensive and diminish the quality of receptive fields for tasks like micro-expression recognition (MER) with limited training samples. Therefore, we propose a refined neural architecture search strategy to search for a tiny CNN architecture for MER. In addition, we introduced a refined hybrid module (RHM) for inner-level search space and an optimal path explore network (OPEN) for outer-level search space. The RHM focuses on discovering optimal cell structures by incorporating a multilateral hybrid spatiotemporal operation space. Also, spatiotemporal attention blocks are embedded to refine the aggregated cell features. The OPEN search space aims to trace an optimal path between the cells to generate a tiny spatiotemporal CNN architecture instead of covering all possible tracks. The aggregate mix of RHM and OPEN search space availed the NAS method to robustly search and design an effective and efficient framework for MER. Compared with contemporary works, experiments reveal that the RNAS-MER is capable of bridging the gap between NAS algorithms and MER tasks. Furthermore, RNAS-MER achieves new state-of-the-art performances on challenging MER benchmarks, including 0.8511%, 0.7620%, 0.9078% and 0.8235% UAR on COMPOSITE, SMIC, CASME-II and SAMM datasets respectively. Monu Verma, Priyanka Lubal, Santosh Kumar Vipparthi, Mohamed Abdel-Mottaleb |
WACV | 1 |
| 2023 | Efficient neural architecture search for emotion recognition
Monu Verma, Murari Mandal, M. Satish Kumar Reddy, Yashwanth Reddy Meedimale, Santosh Kumar Vipparthi |
Expert Syst. Appl. | 1 |
| 2023 | HyFiNet: Hybrid feature attention network for hand gesture recognition
Gopa Bhaumik, Monu Verma, Mahesh Chandra Govil, Santosh Kumar Vipparthi |
Multim. Tools Appl. | 2 |
| 2022 | AutoMER: Spatiotemporal Neural Architecture Search for Microexpression RecognitionabstractFacial microexpressions offer useful insights into subtle human emotions. This unpremeditated emotional leakage exhibits the true emotions of a person. However, the minute temporal changes in the video sequences are very difficult to model for accurate classification. In this article, we propose a novel spatiotemporal architecture search algorithm, AutoMER for microexpression recognition (MER). Our main contribution is a new parallelogram design-based search space for efficient architecture search. We introduce a spatiotemporal feature module named 3-D singleton convolution for cell-level analysis. Furthermore, we present four such candidate operators and two 3-D dilated convolution operators to encode the raw video sequences in an end-to-end manner. To the best of our knowledge, this is the first attempt to discover 3-D convolutional neural network (CNN) architectures with a network-level search for MER. The searched models using the proposed AutoMER algorithm are evaluated over five microexpression data sets: CASME-I, SMIC, CASME-II, CAS(ME) ∧2 , and SAMM. The proposed generated models quantitatively outperform the existing state-of-the-art approaches. The AutoMER is further validated with different configurations, such as downsampling rate factor, multiscale singleton 3-D convolution, parallelogram, and multiscale kernels. Overall, five ablation experiments were conducted to analyze the operational insights of the proposed AutoMER. Monu Verma, M. Satish Kumar Reddy, Yashwanth Reddy Meedimale, Murari Mandal, Santosh Kumar Vipparthi |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | ExtriDeNet: an intensive feature extrication deep network for hand gesture recognition
Gopa Bhaumik, Monu Verma, Mahesh Chandra Govil, Santosh Kumar Vipparthi |
Vis. Comput. | 2 |
| 2021 | One for All: An End-to-End Compact Solution for Hand Gesture RecognitionabstractThe HGR is a quite challenging task as its performance is influenced by various aspects such as illumination variations, cluttered backgrounds, spontaneous capture, etc. The conventional CNN networks for HGR are following two stage pipeline to deal with the various challenges: complex signs, illumination variations, complex and cluttered backgrounds. The existing approaches needs expert expertise as well as auxiliary computation at stage 1 to remove the complexities from the input images. Therefore, in this paper, we proposes an novel end-to-end compact CNN framework: fine grained feature attentive network for hand gesture recognition (Fit-Hand) to solve the challenges as discussed above. The pipeline of the proposed architecture consists of two main units: FineFeat module and dilated convolutional (Conv) layer. The FineFeat module extracts fine grained feature maps by employing attention mechanism over multiscale receptive fields. The attention mechanism is introduced to capture effective features by enlarging the average behaviour of multi-scale responses. Moreover, dilated convolution provides global features of hand gestures through a larger receptive field. In addition, integrated layer is also utilized to combine the features of FineFeat module and dilated layer which enhances the discriminability of the network by capturing complementary context information of hand postures. The effectiveness of Fit-Hand is evaluated by using subject dependent (SD) and subject independent (SI) validation setup over seven benchmark datasets: MUGD-I, MUGD-II, MUGD-III, MUGD-IV, MUGD-V, Finger Spelling and OUHANDS, respectively. Furthermore, to investigate the deep insights of the proposed Fit-Hand framework, we performed ten ablation study Monu Verma, Santosh Kumar Vipparthi |
IJCNN | 1 |
| 2020 | Non-Linearities Improve OrigiNet based on Active Imaging for Micro Expression RecognitionabstractMicro expression recognition (MER)is a very challenging task as the expression lives very short in nature and demands feature modeling with the involvement of both spatial and temporal dynamics. Existing MER systems exploit CNN networks to spot the significant features of minor muscle movements and subtle changes. However, existing networks fail to establish a relationship between spatial features of facial appearance and temporal variations of facial dynamics. Thus, these networks were not able to effectively capture minute variations and subtle changes in expressive regions. To address these issues, we introduce an active imaging concept to segregate active changes in expressive regions of a video into a single frame while preserving facial appearance information. Moreover, we propose a shallow CNN network: hybrid local receptive field based augmented learning network (OrigiNet) that efficiently learns significant features of the micro-expressions in a video. In this paper, we propose a new refined rectified linear unit (RReLU), which overcome the problem of vanishing gradient and dying ReLU. RReLU extends the range of derivatives as compared to existing activation functions. The RReLU not only injects a nonlinearity but also captures the true edges by imposing additive and multiplicative property. Furthermore, we present an augmented feature learning block to improve the learning capabilities of the network by embedding two parallel fully connected layers. The performance of proposed OrigiNet is evaluated by conducting leave one subject out experiments on four comprehensive ME datasets. The experimental results demonstrate that OrigiNet outperformed state-of-the-art techniques with less computational complexity. Monu Verma, Santosh Kumar Vipparthi, Girdhari Singh |
IJCNN | 1 |
| 2020 | LEARNet: Dynamic Imaging Network for Micro Expression RecognitionabstractUnlike prevalent facial expressions, micro expressions have subtle, involuntary muscle movements which are short-lived in nature. These minute muscle movements reflect true emotions of a person. Due to the short duration and low intensity, these micro-expressions are very difficult to perceive and interpret correctly. In this paper, we propose the dynamic representation of micro-expressions to preserve facial movement information of a video in a single frame. We also propose a Lateral Accretive Hybrid Network (LEARNet) to capture micro-level features of an expression in the facial region. The LEARNet refines the salient expression features in accretive manner by incorporating accretion layers (AL) in the network. The response of the AL holds the hybrid feature maps generated by prior laterally connected convolution layers. Moreover, LEARNet architecture incorporates the cross decoupled relationship between convolution layers which helps in preserving the tiny but influential facial muscle change information. The visual responses of the proposed LEARNet depict the effectiveness of the system by preserving both high- and micro-level edge features of facial expression. The effectiveness of the proposed LEARNet is evaluated on four benchmark datasets: CASME-I, CASME-II, CAS(ME)'2 and SMIC. The experimental results after investigation show a significant improvement of 4.03%, 1.90%, 1.79% and 2.82% as compared with ResNet on CASME-I, CASME-II, CAS(ME)'2 and SMIC datasets respectively. Monu Verma, Santosh Kumar Vipparthi, Girdhari Singh, M. Subrahmanyam 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Regional adaptive affinitive patterns (RADAP) with logical operators for facial expression recognitionabstractAutomated facial expression recognition plays a significant role in the study of human behaviour analysis. In this study, the authors propose a robust feature descriptor named regional adaptive affinitive patterns (RADAP) for facial expression recognition. The RADAP computes positional adaptive thresholds in the local neighbourhood and encodes multi‐distance magnitude features which are robust to intra‐class variations and irregular illumination variation in an image. Furthermore, they established cross‐distance co‐occurrence relations in RADAP by using logical operators. They proposed XRADAP, ARADAP, and DRADAP using xor, adder and decoder, respectively. The XRADAP engrains the quality of robustness to intra‐class variations in RADAP features using pairwise co‐occurrence. Similarly, ARADAP and DRADAP extract more stable and illumination invariant features and capture the minute expression features which are usually missed by regular descriptors. The performance of the proposed methods is evaluated by conducting experiments on nine benchmark datasets Cohn–Kanade+ (CK+), Japanese female facial expression (JAFFE), Multimedia Understanding Group (MUG), MMI, OULU‐CASIA, Indian spontaneous expression database, DISFA, AFEW and Combined (CK+, JAFFE, MUG, MMI & GEMEP‐FERA) database in both person dependent and person independent setup. The experimental results demonstrate the effectiveness of the proposed method over state‐of‐the‐art approaches. Murari Mandal, Monu Verma, Sonakshi Mathur, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Kranthi Kumar Deveerasetty |
IET Image Process. | 2 |
| 2018 | QUEST: Quadriletral Senary Bit Pattern for Facial Expression RecognitionabstractFacial expression has significant role to analyzing human cognitive state. Deriving an accurate facial appearance representation is critical task for an automatic facial expression recognition application. This paper provides a new feature descriptor named as Quadrilateral Senary bit Pattern for facial expression recognition. The QUEST pattern encoded the intensity changes by emphasizing relationship between neighboring and reference pixels by dividing them into two quadrilaterals in local neighborhood. Thus, the resultant gradient edges reveal the transitional variation information, that improves the classification rate by discriminating expression classes. Moreover, it also enhances the capability of the descriptor to deal with view point variations and illumination changes. The trine relationship in quadrilateral structure helps to extract the expressive edges and suppressing noise elements to enhance the robustness to noisy conditions. The QUEST pattern generates a six-bit compact code, which improve the efficiency of the FER system with more discriminability. The effectiveness of proposed method is evaluated by conducting several experiments on four benchmark datasets: MMI, GEMEP-FERA, OULU-CASIA and ISED. The experimental results show better performance of the proposed method as compared to existing state-art-the approaches. Monu Verma, Prafulla Saxena, Santosh Kumar Vipparthi, Girdhari Singh |
SMC | 1 |