VLDB 2026 Research / reviewers in the wild / expert
Niladri B. Puhan
dblp:80/4084 · also Niladri Bihari Puhan
· DBLP profile ↗
25ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-5932-1579ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 5 since 2021Security and privacy · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing the cloud burst prediction over mountainous terrains of northwestern Himalaya of India: a deep learning approach
Dhananjay Trivedi, Omveer Sharma, Sandeep Pattnaik, Niladri B. Puhan |
Neural Comput. Appl. | 5 |
| 2025 | Enhancing Intensification Forecast of Tropical Cyclones in the North Indian Ocean Basin Using Fused Vision Transformer and Convolutional Block Attention With ResNet ModelabstractThe accurate forecasting of tropical cyclone (TC) intensity is challenging because of the large uncertainties in the physical processes and lack of observations over the north Indian Ocean basins (NIO), i.e., the Bay of Bengal (BoB) and the Arabian Sea (AS). These limitations lead to more inaccuracy in the intensity forecast of physics-based dynamical models. Here, we design a fused deep learning (DL) network comprised of a Vision Transformer (ViT) and Convolutional Block Attention Model (CBAM) with ResNet, known as HASTVi, for cyclone intensity forecast with a lead time of up to 24 hours. The model has been trained and tested using infrared satellite images obtained from the Indian Space Research Organization (ISRO) and the India Meteorological Department (IMD) best estimates of TC intensity. A total of 38 TCs were used to train and test the model, out of which 35 TCs (7274 images) were used for training and 3 TCs (694 images) were used for testing purposes. The three most devastating TCs, viz. Amphan (2020), Fani (2019), and Tauktae (2021) were used for testing purposes due to their immense socioeconomic impact. The proposed vision HASTVi model shows the Mean absolute error (MAE) of 2.90 kts, 3.14 kts, and 5.38 kts, for 3hr forecast, which is better than the state-of-the-art Convolutional Long Short Term Memory (ConvLSTM) model, having an MAE of 15.35 kts, 18.99 kts, and 29.90 kts for Taukate, Fani, and Amphan, respectively. In addition, the HASTVi model has been tested for the rapid intensification phase of the TC over the NIO basin. These findings have direct implications for improving the TC early warning systems over the NIO basins. Omveer Sharma, Dhananjay Trivedi, Sandeep Pattnaik, Niladri B. Puhan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Interpretative Attention Networks for Structural Component Recognition
Abhishek Uniyal, Bappaditya Mandal, Niladri B. Puhan, Padmalochan Bera |
ICPR (16) | 3 |
| 2024 | Transformer based composite network for autonomous driving trajectory prediction on multi-lane highways
Omveer Sharma, Nirod C. Sahoo, Niladri B. Puhan |
Appl. Intell. | 3 |
| 2023 | Visual Attention Assisted GamesabstractIn this work, we propose a committee of attention models developed for improving the deep reinforcement learning frequently used for games. The game environment is manifested with spatial and temporal attention mechanisms so as to focus on important regions while playing the games. We propose that when an agent’s visual attention space is streamlined and strategic temporally coherent representations are used, generalisation will be faster than other traditional architectures. The proposed spatial attention mechanism’s output enables direct analysis of the information recognised by the agent to choose its actions, allowing for a more straightforward interpretation of the provided state. We evaluate several techniques across a variety of games to reinforce our argument. Extensive experimental results on five Atari 2600 games demonstrate that an agent that makes use of this framework is capable of outperforming state-of-the-art models on ATARI tasks while being interpretable. Bappaditya Mandal, Niladri B. Puhan, Varma Homi Anil |
CoG | 2 |
| 2023 | Improvement in District Scale Heavy Rainfall Prediction Over Complex Terrain of North East India Using Deep LearningabstractPredicting heavy rainfall events (HREs) in real time poses a significant challenge in India, particularly in complex terrain regions like Assam, where these hydro-meteorological events frequently associated with flash floods with severe consequences over region. The devastating HREs in June 2022 led to numerous casualties, extensive damage, and economic losses exceeding 200 crore, necessitating the evacuation of over 4 million individuals. As we write this paper Assam again going through immense flooding situation in now i.e. June2023. Due to the limitations of deterministic numerical weather models in accurately forecasting these events, the study explores the incorporation of deep learning (DL) models, specifically U-Nets, using simulated daily accumulated rainfall outputs from various parametrization schemes. Over a four-day period in June 2022, the U-Net based model demonstrated superior skills in predicting rainfall at the district scale, achieving a Mean Absolute Error (MAE) of less than 12mm, outperforming individual and ensemble model outputs. Comparing the DL model’s performance to the Weather Research and Forecasting (WRF) forecasts, it exhibited a remarkable 64.78% reduction in MAE across Assam. Notably, the proposed model accurately predicted HREs in specific districts such as Barpeta, Kamrup, Kokrajhar, and Nalbari, showcasing improved spatial variation compared to the WRF model. The DL model’s predictions aligned with actual rainfall (> 150 mm) observations from the India Meteorological Department (IMD), while the WRF forecasts consistently underestimated rainfall intensity (< 100 mm). Furthermore, the proposed model achieved a high prediction accuracy of 77.9% in categorical rainfall prediction, significantly outperforming the WRF schemes by 38.1%. Omveer Sharma, Dhananjay Trivedi, Sandeep Pattnaik, Vivekananda Hazra, Niladri B. Puhan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Holistic Feature Reconstruction-Based 3-D Attention Mechanism for Cross-Spectral Periocular RecognitionabstractRecognition of individuals in cross-spectral environments involves situations where probe and gallery images are captured in distinct wavelength ranges and hence differ significantly in terms of appearance and illumination. In case of heterogeneous periocular images, alleviating this wide appearance gap and learning to extract illumination-invariant features from such local regions become cumbersome, giving rise to a challenging research problem. In this work, we design a novel holistic feature reconstruction-based attention module (H-FRAM) to refine and generate discriminative convolutional features. In contrast to existing spatial and channel attention mechanisms that compute 2-D or 1-D attention weights respectively, H-FRAM calculates the importance of each feature location by performing multi-linear principal component analysis on full 3-D tensor space. In H-FRAM, we claim that the feature reconstruction error, computed in a holistic manner, plays a crucial role in determining the relevance of feature locations. This reconstruction error is found by projecting the input feature map onto a multi-dimensional eigenspace that captures most of the important variations. To our knowledge, this is the first work which explores subspace learning approach in the context of 3-D attention mechanism to capture discriminative feature information. We have created an in-house cross-spectral periocular dataset containing visible and near-infrared images from 200 classes. The images are captured in unconstrained acquisition setup involving unsupervised eye and head movements as well as accessory variations (face masks and eyeglasses). Extensive experiments and ablation studies show that the proposed network achieves state-of-the-art recognition performances on the existing and in-house datasets for heterogeneous and homogeneous periocular recognition as well as heterogeneous face recognition. Niladri B. Puhan, Sushree Sangeeta Behera |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Kernelized dynamic convolution routing in spatial and channel interaction for attentive concrete defect recognition
Gaurab Bhattacharya, Niladri B. Puhan, Bappaditya Mandal |
Signal Process. Image Commun. | 2 |
| 2021 | Recent advances in motion and behavior planning techniques for software architecture of autonomous vehicles: A state-of-the-art survey
Omveer Sharma, Nirod C. Sahoo, Niladri B. Puhan |
Eng. Appl. Artif. Intell. | 3 |
| 2021 | Multi-Deformation Aware Attention Learning for Concrete Structural Defect ClassificationabstractIn this work, we propose a deep multi-deformation aware attention learning (MDAL) architecture comprising of multi-scale committee of attention (MSCA) and fine-grained feature induced attention (FGIA) modules to classify multi-target multi-class defects in concrete structures found in civil infrastructures. The MDAL network is composed of interleaved MSCA and FGIA modules to encode crucial fine-grained deformation-aware information from concrete images. The novel attention mechanism is able to localize specific defect regions within an image and extracts crucial discriminative information in multi-scale fashion ranging from coarser to finer features without using any preprocessing step, such as region-of-interest selection or denoising. Our proposed attention mechanism enables the MDAL architecture to automatically classify multiple overlapping defect classes present in the concrete images and leads to an end-to-end trainable deep network. Experimental results on three large concrete defect datasets and ablation studies show that our MDAL network outperforms the current state-of-the-art methodologies significantly. Gaurab Bhattacharya, Bappaditya Mandal, Niladri B. Puhan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Interleaved Deep Artifacts-Aware Attention Mechanism for Concrete Structural Defect ClassificationabstractAutomatic machine classification of concrete structural defects in images poses significant challenges because of multitude of problems arising from the surface texture, such as presence of stains, holes, colors, poster remains, graffiti, marking and painting, along with uncontrolled weather conditions and illuminations. In this paper, we propose an interleaved deep artifacts-aware attention mechanism (iDAAM) to classify multi-target multi-class and single-class defects from structural defect images. Our novel architecture is composed of interleaved fine-grained dense modules (FGDM) and concurrent dual attention modules (CDAM) to extract local discriminative features from concrete defect images. FGDM helps to aggregate multi-layer robust information with wide range of scales to describe visually-similar overlapping defects. On the other hand, CDAM selects multiple representations of highly localized overlapping defect features and encodes the crucial spatial regions from discriminative channels to address variations in texture, viewing angle, shape and size of overlapping defect classes. Within iDAAM, FGDM and CDAM are interleaved to extract salient discriminative features from multiple scales by constructing an end-to-end trainable network without any preprocessing steps, making the process fully automatic. Experimental results and extensive ablation studies on three publicly available large concrete defect datasets show that our proposed approach outperforms the current state-of-the-art methodologies. Gaurab Bhattacharya, Bappaditya Mandal, Niladri B. Puhan |
IEEE Trans. Image Process. | 3 |
| 2020 | Variance-guided attention-based twin deep network for cross-spectral periocular recognition
Sushree Sangeeta Behera, Sapna S. Mishra, Bappaditya Mandal, Niladri B. Puhan |
Image Vis. Comput. | 4 |
| 2019 | A Survey on Smooth Path Generation Techniques for Nonholonomic Autonomous Vehicle SystemsabstractAutonomous vehicles (AVs) are at the heart of academic and industry research because of various advantages such as safety improvement, lower energy and fuel consumption, exploitation of the road network, reduced congestion and greater mobility. Obstacle avoidance, searching the safest path to follow, generation of suitable manoeuver, optimization based most comfortable trajectory along with road boundaries and traffic rules are important concerns in motion planning. The purpose of this paper is to review the existing approaches for trajectory planning in terms of their feasibility, dynamic constraints, handling obstacles, and optimality for comfort. Omveer Sharma, Nirod C. Sahoo, Niladri B. Puhan |
IECON | 3 |
| 2019 | Multi-Level Dual-Attention Based CNN for Macular Optical Coherence Tomography ClassificationabstractIn this letter, we propose a multi-level dual-attention model to classify two common macular diseases, age-related macular degeneration (AMD) and diabetic macular edema (DME) from normal macular eye conditions using optical coherence tomography (OCT) imaging technique. Our approach unifies the dual-attention mechanism at multi-levels of the pre-trained deep convolutional neural network (CNN). It provides a focused learning mechanism by taking into account both multi-level features based attention focusing on the salient coarser features and self-attention mechanism attending higher entropy regions of the finer features. Our proposed method enables the network to automatically focus on the relevant parts of the input images at different levels of feature subspaces. This leads to a more locally deformation-aware feature generation and classification. The proposed approach does not require pre-processing steps such as extraction of region of interest, denoising, and retinal flattening, making the network more robust and fully automatic. Experimental results on two macular OCT databases show the superior performance of our proposed approach as compared to the current state-of-the-art methodologies. Sapna S. Mishra, Bappaditya Mandal, Niladri B. Puhan |
IEEE Signal Process. Lett. | 3 |
| 2018 | Deep residual network with regularised fisher framework for detection of melanomaabstractOf all the skin cancer that is prevalent, melanoma has the highest mortality rates. Melanoma becomes life threatening when it penetrates deep into the dermis layer unless detected at an early stage, it becomes fatal since it has a tendency to migrate to other parts of our body. This study presents an automated non‐invasive methodology to assist the clinicians and dermatologists for detection of melanoma. Unlike conventional computational methods which require (expensive) domain expertise for segmentation and hand crafted feature computation and/or selection, a deep convolutional neural network‐based regularised discriminant learning framework which extracts low‐dimensional discriminative features for melanoma detection is proposed. Their approach minimises the whole of within‐class variance information and maximises the total class variance information. The importance of various subspaces arising in the within‐class scatter matrix followed by dimensionality reduction using total class variance information is analysed for melanoma detection. Experimental results on ISBI 2016, MED‐NODE, PH2 and the recent ISBI 2017 databases show the efficacy of their proposed approach as compared to other state‐of‐the‐art methodologies. Nazneen N. Sultana, Bappaditya Mandal, Niladri B. Puhan |
IET Comput. Vis. | 3 |
| 2018 | Unconstrained handwritten digit recognition using perceptual shape primitives
Kalyan S. Dash, Niladri B. Puhan, Ganapati Panda |
Pattern Anal. Appl. | 2 |
| 2017 | Periocular recognition in cross-spectral scenarioabstractPeriocular recognition has been an active area of research in the past few years. In spite of the advancements made in this area, the cross-spectral matching of visible (VIS) and near-infrared (NIR) periocular images remains a challenge. In this paper, we propose a method based on illumination normalization of VIS and NIR periocular images. Specifically, the approach involves normalizing the images using the difference of Gaussian (DoG) filtering, followed by the computation of a descriptor that captures structural details in the illumination normalized images using histogram of oriented gradients (HOG). Finally, the feature vectors corresponding to the query and the enrolled image are compared using the cosine similarity metric to generate a matching score. Performance of our algorithm has been evaluated on three publicly available benchmark databases of cross-spectral periocular images. Our approach yields significant improvement in performance over the existing approach. Sushree Sangeeta Behera, Mahesh Gour, Vivek Kanhangad, Niladri B. Puhan |
IJCB | 4 |
| 2017 | FoodNet: Recognizing Foods Using Ensemble of Deep NetworksabstractIn this letter, we propose a protocol for an automatic food recognition system that identifies the contents of the meal from the images of the food. We developed a multilayered convolutional neural network (CNN) pipeline that takes advantages of the features from other deep networks and improves the efficiency. Numerous traditional handcrafted features and methods are explored, among which CNNs are chosen as the best performing features. Networks are trained and fine-tuned using preprocessed images and the filter outputs are fused to achieve higher accuracy. Experimental results on the largest real-world food recognition database ETH Food-101 and newly contributed Indian food image database demonstrate the effectiveness of the proposed methodology as compared to many other benchmark deep learned CNN frameworks. Paritosh Pandey, Akella Deepthi, Bappaditya Mandal, Niladri B. Puhan |
IEEE Signal Process. Lett. | 4 |
| 2016 | BESAC: Binary External Symmetry Axis Constellation for unconstrained handwritten character recognition
Kalyan S. Dash, Niladri B. Puhan, Ganapati Panda |
Pattern Recognit. Lett. | 2 |
| 2015 | Handwritten numeral recognition using non-redundant Stockwell transform and bio-inspired optimal zoningabstractHandwritten digit recognition is one of the challenging problems of character recognition because of the large variation in writing styles of individuals and the presence of similar looking shapes of different numerals. Most of the feature extraction techniques are based on statistical or topological attributes of the image in its spatial domain, barring few works attempting feature extraction in a transformed domain. Another challenge is the optimal selection of zones while extracting features from localised zones of the unknown (test) image. In most of the cases, the recognition phase, being isolated from the training phase makes it impossible to adaptively improve the feature selection using the knowledge obtained from error analysis. In this study, the authors propose a feature extraction technique, new to the character recognition problem, using non‐redundant Stockwell transform. Another transformed domain feature extraction using Slantlet coefficients is proposed. They also propose to use bio‐inspired and evolutionary computing‐based optimisation techniques to adaptively select the optimal zone arrangement in the feature selection stage from the knowledge of classification accuracy. The proposed methods are experimentally validated on handwritten digit database of Odia language which proves to outperform any recognition accuracy reported before. Kalyan S. Dash, Niladri B. Puhan, Ganapati Panda |
IET Image Process. | 2 |
| 2014 | Chord oriented gap feature for offline signature verificationabstractIn this paper, we address offline signature verification by proposing a new Partial Invariant Chord Oriented Gap (PICOG) feature. The new heuristically developed feature is conceptualized after observing the directional variation of the gaps (sequence of white pixels) between signature strokes. A set of unique and partial invariant chords is identified using genuine and forgery training signatures in the writer dependent system. Different sets of PICOG chords are selected for each writer by defining a threshold (djnv). The similarity threshold (dth) is computed by performing another training step using the PICOG chords. A majority score based approach is selected to determine if the testing signature is genuine or forgery. A maximum accuracy of 82.27% is obtained on the widely used and publicly available, noisy signature database (CEDAR). M. Manoj Kumar, Niladri B. Puhan |
ICARCV | 2 |
| 2009 | High Capacity Data Hiding in Binary Document Images
Niladri B. Puhan, Anthony Tung Shuen Ho, Farook Sattar |
IWDW | 1 |
| 2007 | Secure authentication watermarking for localization against the Holliman-Memon attack
Niladri B. Puhan, Anthony Tung Shuen Ho |
Multim. Syst. | 1 |
| 2005 | Secure Tamper Localization in Binary Document Image Authentication
Niladri B. Puhan, Anthony Tung Shuen Ho |
KES (4) | 1 |
| 2004 | Imperceptible data embedding in sharply-contrasted binary imagesabstractData embedding in sharply-contrasted binary images like text, drawing, signature and cartoon is a challenging issue due to simple pixel statistics in such images. Arbitrary modification to the pixels can be visually perceptible in the process of data embedding. The use of a valid perceptual model is important to minimize the effect of such visual distortion in binary images. In this paper, a novel perceptual model is used to embed significant amount of information such that the original and the marked images before and after data embedding process are perceptually similar. In our model, the distortion that occurs after flipping a pixel is estimated on the curvature-weighted distance difference (CWDD) measure between two contour segments. Anthony Tung Shuen Ho, Niladri B. Puhan, Anamitra Makur, Pina Marziliano, Yong Liang Guan 0001 |
ICARCV | 2 |