EDBT 2026 Demo / reviewers in the wild / expert
Ardhendu Behera
dblp:25/3860
· DBLP profile ↗
40ranked-venue papers
14as first author
21since 2021 · last 2026
0000-0003-0276-9000ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 9 first-author · 11 since 2021Artificial intelligence and machine learning · 21 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorSystems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-Grained Action Recognition
Imtiaz Ul Hassan, Nik Bessis, Ardhendu Behera |
ICPR (5) | 3 |
| 2026 | Towards Label-Free Single-Cell Phenotyping Using Multi-task Learning
Saqib Nazir, Ardhendu Behera |
ICPR (5) | 2 |
| 2026 | HistDiT: A Structure-Aware Latent Conditional Diffusion Model for High-Fidelity Virtual Staining in Histopathology
Aasim Bin Saleem, Amr Ahmed 0002, Ardhendu Behera, Hafeezullah Amin, Iman Yi Liao, Mahmoud Khattab, Jia-Wern Pan, Haslina Makmur |
ICPR (13) | 3 |
| 2026 | VIM-Net: A voxel-interaction multimodal network for 3D object detection
Minghan Wang, Xijiong Wang, Yonghuai Liu, Ardhendu Behera, Baowen Zhang |
Pattern Recognit. | 6 |
| 2026 | From Virtual Environments to Real-World Trials: Emerging Trends in Autonomous DrivingabstractAutonomous driving technologies have achieved significant advances in recent years, yet their real-world deployment remains constrained by data scarcity, safety requirements, and the need for generalization across diverse environments. In response, synthetic data and virtual environments have emerged as powerful enablers, offering scalable, controllable, and richly annotated scenarios for training and evaluation. This survey presents a comprehensive review of recent developments at the intersection of autonomous driving, simulation technologies, and synthetic datasets. We organize the landscape across three core dimensions: 1) the use of synthetic data for perception and planning, 2) digital twin-based simulation for system validation, and 3) domain adaptation strategies bridging synthetic and real-world data. We also highlight the role of vision-language models and simulation realism in enhancing scene understanding and generalization. A detailed taxonomy of datasets, tools, and simulation platforms is provided, alongside an analysis of trends in benchmark design. Finally, we discuss critical challenges and open research directions, including Sim2Real transfer, scalable safety validation, cooperative autonomy, and simulation-driven policy learning, that must be addressed to accelerate the path toward safe, generalizable, and globally deployable autonomous driving systems. Aditya Humnabadkar, Arindam Sikdar, Benjamin Cave, Huaizhong Zhang, Nik Bessis, Ardhendu Behera |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | LANet: Enhancing Plant Disease Recognition Through Transfer Learning and Layer AttentionabstractAccurate recognition of plant diseases plays a critical role in maintaining agricultural productivity and food security. This study proposes LANet, an innovative network architecture aimed at improving the precision of identifying and categorizing plant diseases from images. LANet consists of two primary components. Firstly, Layer Attention applies self-attention across feature layers at different scales to capture the weights between cross-scale features, addressing the data loss caused by downsampling and enhancing plant species recognition. Secondly, we introduce a transfer learning method, where a plant species classification network is initially trained, its parameters are then frozen, and several convolutional layers are added to train the disease classification network. This allows the model to leverage the learned classification information for more accurate disease recognition. Our model demonstrates an accuracy of 85.62%, outperforming nine other models significantly. Comparative analyses using the public PlantVillage dataset highlight the superior performance of our method in disease recognition. Ardhendu Behera, Fuzhong Li, Wuping Zhang, Yonghuai Liu |
IPAS | 2 |
| 2025 | Interweaving Insights: High-Order Feature Interaction for Fine-Grained Visual RecognitionabstractThis paper presents a novel approach for Fine-Grained Visual Classification (FGVC) by exploring Graph Neural Networks (GNNs) to facilitate high-order feature interactions, with a specific focus on constructing both inter- and intra-region graphs. Unlike previous FGVC techniques that often isolate global and local features, our method combines both features seamlessly during learning via graphs. Inter-region graphs capture long-range dependencies to recognize global patterns, while intra-region graphs delve into finer details within specific regions of an object by exploring high-dimensional convolutional features. A key innovation is the use of shared GNNs with an attention mechanism coupled with the Approximate Personalized Propagation of Neural Predictions (APPNP) message-passing algorithm, enhancing information propagation efficiency for better discriminability and simplifying the model architecture for computational efficiency. Additionally, the introduction of residual connections improves performance and training stability. Comprehensive experiments showcase state-of-the-art results on benchmark FGVC datasets, affirming the efficacy of our approach. This work underscores the potential of GNN in modeling high-level feature interactions, distinguishing it from previous FGVC methods that typically focus on singular aspects of feature representation. Our source code is available at https://github.com/Arindam-1991/I2-HOFI. Arindam Sikdar, Yonghuai Liu, Siddhardha Kedarisetty, Yitian Zhao, Amr Ahmed 0002, Ardhendu Behera |
Int. J. Comput. Vis. | 6 |
| 2025 | SLPDR: A Benchmark for Ship License Plate Detection and RecognitionabstractShip identification is a prerequisite for the intelligent management of maritime transportation, yet existing research is confined to broad ship detection and categorization, which only provides the ship’s location or type instead of its identification. Inspired by the research on the Car License Plate (CLP), we make the first attempt to propose the concept of the Ship License Plate (SLP). In addition, the limited data hinders research on ship identification. To overcome this obstacle, we construct the first large-scale Ship License Plate Detection and Recognition (SLPDR) dataset, which contains 1,472 ship identities and 88,862 images. In addition, this paper proposes an SLP detection model named YOLO-SSA and evaluates this model as well as typical detection methods on the SLPDR dataset. The experimental results demonstrate that the proposed YOLO-SSA achieves better SLP detection performance by enhancing the features where ships and SLPs are located. Furthermore, we explore the prospective applications of SLPs in intelligent maritime transportation, including ship monitoring and berth management. Project web page: https://vsislab.github.io/SLPDR/ Youmei Zhang, Ran Song 0001, Yonghuai Liu, Ardhendu Behera, Mingxin Zhang 0006, Wei Zhang 0021 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Driving Through Graphs: a Bipartite Graph for Traffic Scene AnalysisabstractWe introduce a novel approach for traffic scene analysis in driving videos by exploring spatio-temporal relationships captured by a temporal frame-to-frame (f2f) bipartite graph, eliminating the need for complex image-level high-dimensional feature extraction. Instead, we rely on object detectors that provide bounding box information. The proposed graph approach efficiently connects objects across frames where nodes represent essential object attributes, and edges signify interactions based on simple spatial metrics such as distance and angles between objects. A key innovation is the integration of dynamic edge attributes, computed using Multilayer Perceptrons (MLP) by exploring this spatial metric. These attributes enhance our Interaction-aware Graph Neural Networks (IA-GNNs) framework by adapting the PageRank-driven approximate personalized propagation of neural predictions (APPNP) scheme and graph attention mechanism in a novel way. This has significantly improved our model’s ability to understand spatio-temporal interactions of multiple objects in traffic scenarios. We have rigorously evaluated our approach on two benchmark datasets, METEOR and INTERACTION, demonstrating its accuracy in analyzing traffic scenarios. This streamlined, graph-based strategy marks a significant shift towards more efficient and insightful traffic scene analysis using video data. Our source code is available at: https://github.com/Addy-1998/Bip_DTG. Aditya Humnabadkar, Arindam Sikdar, Huaizhong Zhang, Ardhendu Behera |
ICIP | 5 |
| 2024 | Exploiting Label Uncertainty for Enhanced 3D Object Detection From Point CloudsabstractAccurate detection of objects from LiDAR point clouds is crucial for autonomous driving and environment modeling. However, uncertainties in ground truth labels due to occlusions, sparsity, and truncation can hinder model training and performance. This paper introduces two strategies to address these issues: 1) Soft Regression Loss (SoRL) and 2) Discrete Quantization Sampling (DQS). SoRL utilizes Gaussian distributions for object predictions, measuring uncertainty based on the probability of ground truth labels within these distributions. This method effectively accounts for deviations in object location and orientation. Meanwhile, DQS introduces uncertainty scores for dynamic sample selection, aiming to refine the quality of positive samples for regression. Based on the proposed modules, we design a lightweight multi-stage object detection framework. Notably, these modules can enhance existing 3D object detection methods without affecting significantly inference speeds. Experiments over benchmark datasets show the effectiveness of our method, especially for cars in sparse point clouds. Yonghuai Liu, Ardhendu Behera, Ran Song 0001, Hejin Yuan |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Regional Attention Network (RAN) for Head Pose and Fine-Grained Gesture RecognitionabstractAffect is often expressed via non-verbal body language such as actions/gestures, which are vital indicators for human behaviors. Recent studies on recognition of fine-grained actions/gestures in monocular images have mainly focused on modeling spatial configuration of body parts representing body pose, human-objects interactions and variations in local appearance. The results show that this is a brittle approach since it relies on accurate body parts/objects detection. In this work, we argue that there exist local discriminative semantic regions, whose “informativeness” can be evaluated by the attention mechanism for inferring fine-grained gestures/actions. To this end, we propose a novel end-to-endregional attention network (RAN), which is a fully convolutional neural network (CNN) to combine multiple contextual regions through attention mechanism, focusing on parts of the images that are most relevant to a given task. Our regions consist of one or more consecutive cells and are adapted from the strategies used in computing HOG (Histogram of Oriented Gradient) descriptor. The model is extensively evaluated on ten datasets belonging to 3 different scenarios: 1) head pose recognition, 2) drivers state recognition, and 3) human action and facial expression recognition. The proposed approach outperforms the state-of-the-art by a considerable margin in different metrics. Ardhendu Behera, Zachary Wharton, Yonghuai Liu, Morteza Ghahremani, Swagat Kumar, Nik Bessis |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | A review of computer vision-based approaches for physical rehabilitation and assessmentabstractAbstract The computer vision community has extensively researched the area of human motion analysis, which primarily focuses on pose estimation, activity recognition, pose or gesture recognition and so on. However for many applications, like monitoring of functional rehabilitation of patients with musculo skeletal or physical impairments, the requirement is to comparatively evaluate human motion. In this survey, we capture important literature on vision-based monitoring and physical rehabilitation that focuses on comparative evaluation of human motion during the past two decades and discuss the state of current research in this area. Unlike other reviews in this area, which are written from a clinical objective, this article presents research in this area from a computer vision application perspective. We propose our own taxonomy of computer vision-based rehabilitation and assessment research which are further divided into sub-categories to capture novelties of each research. The review discusses the challenges of this domain due to the wide ranging human motion abnormalities and difficulty in automatically assessing those abnormalities. Finally, suggestions on the future direction of research are offered. Bappaditya Debnath, Mary O'Brien, Motonori Yamaguchi, Ardhendu Behera |
Multim. Syst. | 4 |
| 2022 | SR-GNN: Spatial Relation-Aware Graph Neural Network for Fine-Grained Image CategorizationabstractOver the past few years, a significant progress has been made in deep convolutional neural networks (CNNs)-based image recognition. This is mainly due to the strong ability of such networks in mining discriminative object pose and parts information from texture and shape. This is often inappropriate for fine-grained visual classification (FGVC) since it exhibits high intra-class and low inter-class variances due to occlusions, deformation, illuminations, etc. Thus, an expressive feature representation describing global structural information is a key to characterize an object/ scene. To this end, we propose a method that effectively captures subtle changes by aggregating context-aware features from most relevant image-regions and their importance in discriminating fine-grained categories avoiding the bounding-box and/or distinguishable part annotations. Our approach is inspired by the recent advancement in self-attention and graph neural networks (GNNs) approaches to include a simple yet effective relation-aware feature transformation and its refinement using a context-aware attention mechanism to boost the discriminability of the transformed feature in an end-to-end learning process. Our model is evaluated on eight benchmark datasets consisting of fine-grained objects and human-object interactions. It outperforms the state-of-the-art approaches by a significant margin in recognition accuracy. Asish Bera, Zachary Wharton, Yonghuai Liu, Nik Bessis, Ardhendu Behera |
IEEE Trans. Image Process. | 5 |
| 2022 | Deep CNN, Body Pose, and Body-Object Interaction Features for Drivers' Activity MonitoringabstractAutomatic recognition and prediction of in-vehicle human activities has a significant impact on the next generation of driver assistance and intelligent autonomous vehicles. In this article, we present a novel single image driver action recognition algorithm inspired by human perception that often focuses selectively on parts of the images to acquire information at specific places which are distinct to a given task. Unlike existing approaches, we argue that human activity is a combination of pose and semantic contextual cues. In detail, we model this by considering the configuration of body joints, their interaction with objects being represented as a pairwise relation to capture the structural information. Our body-pose and body-object interaction representation is built to be semantically rich and meaningful, which is highly discriminative even though it is coupled with a basic linear SVM classifier. We also propose a Multi-stream Deep Fusion Network (MDFN) for combining high-level semantics with CNN features. Our experimental results demonstrate that the proposed approach significantly improves the drivers’ action recognition accuracy on two exacting datasets. Ardhendu Behera, Zachary Wharton, Alexander Keidel, Bappaditya Debnath |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Retinal Structure Detection in OCTA Image via Voting-Based Multitask LearningabstractAutomated detection of retinal structures, such as retinal vessels (RV), the foveal avascular zone (FAZ), and retinal vascular junctions (RVJ), are of great importance for understanding diseases of the eye and clinical decision-making. In this paper, we propose a novel Voting-based Adaptive Feature Fusion multi-task network (VAFF-Net) for joint segmentation, detection, and classification of RV, FAZ, and RVJ in optical coherence tomography angiography (OCTA). A task-specific voting gate module is proposed to adaptively extract and fuse different features for specific tasks at two levels: features at different spatial positions from a single encoder, and features from multiple encoders. In particular, since the complexity of the microvasculature in OCTA images makes simultaneous precise localization and classification of retinal vascular junctions into bifurcation/crossing a challenging task, we specifically design a task head by combining the heatmap regression and grid classification. We take advantage of three different en face angiograms from various retinal layers, rather than following existing methods that use only a single en face. We carry out extensive experiments on three OCTA datasets acquired using different imaging devices, and the results demonstrate that the proposed method performs on the whole better than either the state-of-the-art single-purpose methods or existing multi-task learning solutions. We also demonstrate that our multi-task learning method generalizes across other imaging modalities, such as color fundus photography, and may potentially be used as a general multi-task learning tool. We also construct three datasets for multiple structure detection, and part of these datasets with the source code and evaluation benchmark have been released for public access. Jinkui Hao, Ting Shen, Xueli Zhu 0002, Yonghuai Liu, Ardhendu Behera, Dan Zhang 0026, Bang Chen, Jiang Liu 0001, Jiong Zhang 0004, Yitian Zhao |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Context-aware Attentional Pooling (CAP) for Fine-grained Visual ClassificationabstractDeep convolutional neural networks (CNNs) have shown a strong ability in mining discriminative object pose and parts information for image recognition. For fine-grained recognition, context-aware rich feature representation of object/scene plays a key role since it exhibits a significant variance in the same subcategory and subtle variance among different subcategories. Finding the subtle variance that fully characterizes the object/scene is not straightforward. To address this, we propose a novel context-aware attentional pooling (CAP) that effectively captures subtle changes via sub-pixel gradients, and learns to attend informative integral regions and their importance in discriminating different subcategories without requiring the bounding-box and/or distinguishable part annotations. We also introduce a novel feature encoding by considering the intrinsic consistency between the informativeness of the integral regions and their spatial structures to capture the semantic correlation among them. Our approach is simple yet extremely effective and can be easily applied on top of a standard classification backbone network. We evaluate our approach using six state-of-the-art (SotA) backbone networks and eight benchmark datasets. Our method significantly outperforms the SotA approaches on six datasets and is very competitive with the remaining two. Ardhendu Behera, Zachary Wharton, Pradeep Hewage, Asish Bera |
AAAI | 1 |
| 2021 | An attention-driven hierarchical multi-scale representation for visual recognition
Zachary Wharton, Ardhendu Behera, Asish Bera |
BMVC | 2 |
| 2021 | Attentional Learn-able Pooling for Human Activity RecognitionabstractHuman activity/behaviour monitoring and recognition is a key for facilitating humans robot interaction, and allows robots for a better scheduling of future operations. It is challenging and often addressed at different levels, such as human activity classification, future activity prediction and monitoring of the on-going activities. The paper proposes a novel attention-based learn-able pooling mechanism for human activity classification from RGB videos. Recently, most of the best performing human activity recognition approaches are based on 3D skeleton positions. The 3D skeleton positions are not always available in videos captured using RGB cameras, which are widely used in robotics applications. RGB videos contain rich spatio-temporal information and processing them semantically is a difficult task. Moreover, accurately capturing spatial information and long-term temporal dependencies is the key to achieving high recognition accuracy. We use an existing Convolutional Neural Network for image recognition to extract video features which are then processed using our innovative application of attention mechanism to focus the network on features that are more important for discrimination. Afterwards, we use a novel learn-able pooling mechanism to extract activity-aware spatio-temporal cues for efficient activity recognition. The proposed pooling mechanism learns the structural information from hidden states of a bidirectional Long Short-Term Memory network via Fisher Vectors. Bappaditya Debnath, Mary O'Brien, Swagat Kumar, Ardhendu Behera |
ICRA | 4 |
| 2021 | Coarse Temporal Attention Network (CTA-Net) for Driver's Activity RecognitionabstractThere is significant progress in recognizing traditional human activities from videos focusing on highly distinctive actions involving discriminative body movements, body-object and/or human-human interactions. Driver's activities are different since they are executed by the same subject with similar body parts movements, resulting in subtle changes. To address this, we propose a novel framework by exploiting the spatiotemporal attention to model the subtle changes. Our model is named Coarse Temporal Attention Network (CTA-Net), in which coarse temporal branches are introduced in a trainable glimpse network. The goal is to allow the glimpse to capture high-level temporal relationships, such as `during', `before' and `after' by focusing on a specific part of a video. These branches also respect the topology of the temporal dynamics in the video, ensuring that different branches learn meaningful spatial and temporal changes. The model then uses an innovative attention mechanism to generate high-level action specific contextual information for activity recognition by exploring the hidden states of an LSTM. The attention mechanism helps in learning to decide the importance of each hidden state for the recognition task by weighing them when constructing the representation of the video. Our approach is evaluated on four publicly accessible datasets and significantly outperforms the state-of-the-art by a considerable margin with only RGB video as input. Zachary Wharton, Ardhendu Behera, Yonghuai Liu, Nik Bessis |
WACV | 2 |
| 2021 | Deep learning-based effective fine-grained weather forecasting modelabstractAbstract It is well-known that numerical weather prediction (NWP) models require considerable computer power to solve complex mathematical equations to obtain a forecast based on current weather conditions. In this article, we propose a novel lightweight data-driven weather forecasting model by exploring temporal modelling approaches of long short-term memory (LSTM) and temporal convolutional networks (TCN) and compare its performance with the existing classical machine learning approaches, statistical forecasting approaches, and a dynamic ensemble method, as well as the well-established weather research and forecasting (WRF) NWP model. More specifically Standard Regression (SR), Support Vector Regression (SVR), and Random Forest (RF) are implemented as the classical machine learning approaches, and Autoregressive Integrated Moving Average (ARIMA), Vector Auto Regression (VAR), and Vector Error Correction Model (VECM) are implemented as the statistical forecasting approaches. Furthermore, Arbitrage of Forecasting Expert (AFE) is implemented as the dynamic ensemble method in this article. Weather information is captured by time-series data and thus, we explore the state-of-art LSTM and TCN models, which is a specialised form of neural network for weather prediction. The proposed deep model consists of a number of layers that use surface weather parameters over a given period of time for weather forecasting. The proposed deep learning networks with LSTM and TCN layers are assessed in two different regressions, namely multi-input multi-output and multi-input single-output. Our experiment shows that the proposed lightweight model produces better results compared to the well-known and complex WRF model, demonstrating its potential for efficient and accurate weather forecasting up to 12 h. Pradeep Hewage, Marcello Trovati, Ella Grishikashvili Pereira, Ardhendu Behera |
Pattern Anal. Appl. | 4 |
| 2021 | Attend and Guide (AG-Net): A Keypoints-Driven Attention-Based Deep Network for Image RecognitionabstractThis article presents a novel keypoints-based attention mechanism for visual recognition in still images. Deep Convolutional Neural Networks (CNNs) for recognizing images with distinctive classes have shown great success, but their performance in discriminating fine-grained changes is not at the same level. We address this by proposing an end-to-end CNN model, which learns meaningful features linking fine-grained changes using our novel attention mechanism. It captures the spatial structures in images by identifying semantic regions (SRs) and their spatial distributions, and is proved to be the key to modeling subtle changes in images. We automatically identify these SRs by grouping the detected keypoints in a given image. The "usefulness" of these SRs for image recognition is measured using our innovative attentional mechanism focusing on parts of the image that are most relevant to a given task. This framework applies to traditional and fine-grained image recognition tasks and does not require manually annotated regions (e.g. bounding-box of body parts, objects, etc.) for learning and prediction. Moreover, the proposed keypoints-driven attention mechanism can be easily integrated into the existing CNN models. The framework is evaluated on six diverse benchmark datasets. The model outperforms the state-of-the-art approaches by a considerable margin using Distracted Driver V1 (Acc: 3.39%), Distracted Driver V2 (Acc: 6.58%), Stanford-40 Actions (mAP: 2.15%), People Playing Musical Instruments (mAP: 16.05%), Food-101 (Acc: 6.30%) and Caltech-256 (Acc: 2.59%) datasets. Asish Bera, Zachary Wharton, Yonghuai Liu, Nik Bessis, Ardhendu Behera |
IEEE Trans. Image Process. | 5 |
| 2020 | Rotation Axis Focused Attention Network (RAFA-Net) for Estimating Head Pose
Ardhendu Behera, Zachary Wharton, Pradeep Hewage, Swagat Kumar |
ACCV (5) | 1 |
| 2020 | Orderly Disorder in Point Cloud Domain
Morteza Ghahremani, Bernard Tiddeman, Yonghuai Liu, Ardhendu Behera |
ECCV (28) | 4 |
| 2020 | Unsupervised Monocular Depth Estimation for Night-Time Images Using Adversarial Domain Feature Adaptation
V. Madhu Babu, Sourav Garg, Anima Majumder, Swagat Kumar, Ardhendu Behera |
ECCV (28) | 5 |
| 2020 | Attention-Driven Body Pose Encoding for Human Activity RecognitionabstractThis article proposes a novel attention-based body pose encoding for human activity recognition. Most of the existing human activity recognition approaches based on 3D pose data often enrich the input data using additional handcrafted representations such as velocity, super-normal vectors, pairwise relations, and so on. The enriched data complements the 3D body joint position data and improves model performance. In this paper, we propose a novel approach that learns enhanced feature representations from a given sequence of 3D body joints. To achieve this encoding, the approach exploits two body pose streams: 1) a spatial stream which encodes the spatial relationship between various body joints at each time point to learn spatial structure involving the spatial distribution of different body joints 2) a temporal stream that learns the temporal variation of individual body joints over the entire sequence duration to present a temporally enhanced representation. Afterwards, these two pose streams are fused with a multi-head attention mechanism. We also capture the contextual information from the RGB video stream using a deep Convolutional Neural Network (CNN) model combined with a multi-head attention and a bidirectional Long Short-Term Memory (LSTM) network. Finally, the RGB video stream is combined with the fused body pose stream to give a novel end-to-end deep model for effective human activity recognition. The proposed model is evaluated on three datasets including the challenging NTU-RGBD dataset and achieves state-of-the-art results. Bappaditya Debnath, Mary O'Brien, Swagat Kumar, Ardhendu Behera |
ICPR | 4 |
| 2020 | Temporal convolutional neural (TCN) network for an effective weather forecasting using time-series data from the local weather stationabstractAbstract Non-predictive or inaccurate weather forecasting can severely impact the community of users such as farmers. Numerical weather prediction models run in major weather forecasting centers with several supercomputers to solve simultaneous complex nonlinear mathematical equations. Such models provide the medium-range weather forecasts, i.e., every 6 h up to 18 h with grid length of 10–20 km. However, farmers often depend on more detailed short-to medium-range forecasts with higher-resolution regional forecasting models. Therefore, this research aims to address this by developing and evaluating a lightweight and novel weather forecasting system, which consists of one or more local weather stations and state-of-the-art machine learning techniques for weather forecasting using time-series data from these weather stations. To this end, the system explores the state-of-the-art temporal convolutional network (TCN) and long short-term memory (LSTM) networks. Our experimental results show that the proposed model using TCN produces better forecasting compared to the LSTM and other classic machine learning approaches. The proposed model can be used as an efficient localized weather forecasting tool for the community of users, and it could be run on a stand-alone personal computer. Pradeep Hewage, Ardhendu Behera, Marcello Trovati, Ella Grishikashvili Pereira, Morteza Ghahremani, Francesco Palmieri 0002, Yonghuai Liu |
Soft Comput. | 2 |
| 2019 | A CNN Model for Head Pose Recognition using Wholes and RegionsabstractHead pose recognition and monitoring is key to many real-world applications, since it is a vital indicator for human attention and behavior. Currently, head pose is often computed by localizing landmarks on a targeted face and solving 2D to 3D correspondence problem with a mean head model. Recent research has shown that this is a brittle approach since it relies entirely on the accuracy of landmark detection, the extraneous head model and an ad-hoc alignment step. Recent work has also shown that the best-performing methods often combine multiple low-level image features with high-level contextual cues. In this paper, we present a novel end-to-end deep network, which is inspired by these ideas and explores regions within an image to capture topological changes due to changes in viewpoint. We adapt the existing state-of-the-art deep CNNs to use more than one region for accurate head pose recognition. Our regions consist of one or more consecutive cells and is adapted from the strategies used in computing HOG descriptor. Extensive experimental results on head pose recognition using four different large-scale datasets, demonstrate that the proposed approach outperforms many state-of-the-art deep CNN models. We also compare our pose recognition performance with the latest OpenFace 2.0 facial behavior analysis toolkit. In addition, we contribute head pose annotation to a large-scale dataset (VGGFace2). Ardhendu Behera, Andrew G. Gidney, Zachary Wharton, Daniel Robinson, Keiron Quinn |
FG | 1 |
| 2018 | Latent Body-Pose guided DenseNet for Recognizing Driver's Fine-grained Secondary ActivitiesabstractThe following topics are dealt with: object detection; learning (artificial intelligence); feature extraction; video surveillance; video signal processing; image classification; convolutional neural nets; image motion analysis; computer vision; object tracking. Ardhendu Behera, Alexander Keidel |
AVSS | 1 |
| 2018 | Adapting MobileNets for mobile based upper body pose estimationabstractHuman pose estimation through deep learning has achieved very high accuracy over various difficult poses. However, these are computationally expensive and are often not suitable for mobile based systems. In this paper, we investigate the use of MobileNets, which is well-known to be a light-weight and efficient CNN architecture for mobile and embedded vision applications. We adapt MobileNets for pose estimation inspired by the hourglass network. We introduce a novel split stream architecture at the final two layers of the MobileNets. This approach reduces over-fitting, resulting in improvement in accuracy and reduction in parameter size. We also show that by maintaining part of the original network we are able to improve accuracy by transferring the learned features from ImageNet pre-trained MobileNets. The adapted model is evaluated on the FLIC dataset. Our network out-performed the default MobileNets for pose estimation, as well as achieved performance comparable to the state of the art results while reducing inference time significantly. Bappaditya Debnath, Mary O'Brien, Motonori Yamaguchi, Ardhendu Behera |
AVSS | 4 |
| 2018 | A Vision-based Transfer Learning Approach for Recognizing Behavioral Symptoms in People with DementiaabstractWith an aging population that continues to grow, dementia is a major global health concern. It is a syndrome in which there is a deterioration in memory, thinking, be-havior and the ability to perform activities of daily living. Depression and aggressive behavior are the most upsetting and challenging symptoms of dementia. Automatic recognition of these behaviors would not only be useful to alert family members and caregivers, but also helpful in planning and managing daily activities of people with dementia (PwD). In this work, we propose a vision-based approach that unifies transfer learning and deep convolutional neural network (CNN) for the effective recognition of behavioral symptoms. We also compare the performance of state-of-the-art CNN features with the hand-crafted HOG-feature, as well as their combination using a basic linear SVM. The proposed method is evaluated on a newly created dataset, which is based on the dementia storyline in ITVs Emmerdale episodes. The Alzheimer's Society has described it as a "realistic portrayal"1of the condition to raise awareness of the issues surrounding dementia. Zachary Wharton, Erik Thomas, Bappaditya Debnath, Ardhendu Behera |
AVSS | 4 |
| 2014 | Real-time Activity Recognition by Discerning Qualitative Relationships Between Randomly Chosen Visual Features
Ardhendu Behera, Anthony G. Cohn 0001, David C. Hogg |
BMVC | 1 |
| 2013 | Robust abandoned object detection integrating wide area visual surveillance and social context
James M. Ferryman, David C. Hogg, Jan Sochman, Ardhendu Behera, José A. Rodríguez-Serrano, Simon F. Worgan, Longzhen Li, Valerie Leung, Murray Evans, Philippe Cornic, Stéphane Herbin, Stefan Schlenger, Michael Dose |
Pattern Recognit. Lett. | 4 |
| 2012 | Egocentric Activity Monitoring and Recovery
Ardhendu Behera, David C. Hogg, Anthony G. Cohn 0001 |
ACCV (3) | 1 |
| 2012 | Workflow Activity Monitoring Using Dynamics of Pair-Wise Qualitative Spatial Relations
Ardhendu Behera, Anthony G. Cohn 0001, David C. Hogg |
MMM | 1 |
| 2011 | Exploiting petri-net structure for activity classification and user instruction within an industrial settingabstractLive workflow monitoring and the resulting user interaction in industrial settings faces a number of challenges. A formal workflow may be unknown or implicit, data may be sparse and certain isolated actions may be undetectable given current visual feature extraction technology. This paper attempts to address these problems by inducing a structural workflow model from multiple expert demonstrations. When interacting with a naive user, this workflow is combined with spatial and temporal information, under a Bayesian framework, to give appropriate feedback and instruction. Structural information is captured by translating a Markov chain of actions into a simple place/transition petri-net. This novel petri-net structure maintains a continuous record of the current workbench configuration and allows multiple sub-sequences to be monitored without resorting to second order processes. This allows the user to switch between multiple sub-tasks, while still receiving informative feedback from the system. As this model captures the complete workflow, human inspection of safety critical processes and expert annotation of user instructions can be made. Activity classification and user instruction results show a significant on-line performance improvement when compared to the existing Hidden Markov Model or pLSA based state of the art. Further analysis reveals that the majority of our model's classification errors are caused by small de-synchronisation events rather than significant workflow deviations. We conclude with a discussion of the generalisability of the induced place/transition petri-net to other activity recognition tasks and summarise the developments of this model. Simon F. Worgan, Ardhendu Behera, Anthony G. Cohn 0001, David C. Hogg |
ICMI | 2 |
| 2008 | DocMIR: An automatic document-based indexing system for meeting retrieval
Ardhendu Behera, Denis Lalanne, Rolf Ingold |
Multim. Tools Appl. | 1 |
| 2005 | Influence of fusion strategies on feature-based identification of low-resolution documentsabstractThe paper describes a method by which one could use the documents captured from low-resolution handheld devices to retrieve the originals of those documents from a document store. The method considers conjunctively two complementary feature sets. First, the geometrical distribution of the color in the document's 2D image plane is preferred. Secondly, the shallow layout features is considered due to the poor resolution of the captured documents. We propose in this article to fuse those two complementary feature sets in order to improve document identification performance. Finally, in order to test the influence of merging strategies on document identification performance, a synergic method is proposed and evaluated relative to a similar method in which feature sets are simply considered sequentially. Ardhendu Behera, Denis Lalanne, Rolf Ingold |
ACM Symposium on Document Engineering | 1 |
| 2005 | Enhancement of Layout-based Identification of Low-resolution Documents using Geometrical Color DistributionabstractThis paper proposes a multi-signature document identification method that works robustly with low-resolution documents captured from handheld devices. The proposed method is based on the extraction of a visual signature containing both (a) the color content distribution in the image plane of the document, i.e. the color signature, and (b) the shallow layout structure of the document, i.e. the layout signature. The color distribution is first considered, in order to filter documents with very dissimilar colors, and the identification is finally done on the remaining set using the layout signature. An evaluation, that compares our color and layout-based method with the layout signature alone, is finally presented. Ardhendu Behera, Denis Lalanne, Rolf Ingold |
ICDAR | 1 |
| 2004 | Visual signature based identification of Low-resolution document imagesabstractIn this paper, we present (a) a method for identifying documents captured from low-resolution devices such as web-cams, digital cameras or mobile phones and (b) a technique for extracting their textual content without performing OCR. The first method associates a hierarchically structured visual signature to the low-resolution document image and further matches it with the visual signatures of the original high-resolution document images, stored in PDF form in a repository. The matching algorithm follows the signature hierarchy, which speeds-up the search by guiding it towards fruitful solution spaces. In a second step, the content of the original PDF document is extracted, structured, and matched with its corresponding high-resolution visual signature. Finally, the matched content is attached to the low-resolution document image's visual signature, which greatly enriches the document's content and indexing. We present in this article both these identification and extraction methods and evaluate them on various documents, resolutions and lighting conditions, using different capture devices. Ardhendu Behera, Denis Lalanne, Rolf Ingold |
ACM Symposium on Document Engineering | 1 |
| 2004 | Looking at projected documents: event detection & document identificationabstractIn the context of a multimodal application, the article proposes an image-based method for bridging the gap between document excerpts and video extracts. The approach, called document image alignment, takes advantage of the observable events related to documents that are visible during meetings. In particular, the article presents a new method for detecting slide changes in slideshows, its evaluation, and a preliminary work on document identification. Ardhendu Behera, Denis Lalanne, Rolf Ingold |
ICME | 1 |