Vijay John

dblp:47/5067 · DBLP profile ↗
← Back
33ranked-venue papers
21as first author
17since 2021 · last 2026
0000-0002-9553-0906ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 12 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 14 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Frame-Level Driver Drowsiness Detection with Deep Embedded Clustering-based Pseudo Label Refinement
Vijay John, Yasutomo Kawanishi
IV1
2026 View-aware Cross-modal Distillation for Multi-view Action Recognition
abstract
The widespread use of multi-sensor systems has increased research in multi-view action recognition. While existing approaches in multi-view setups with fully overlapping sensors benefit from consistent view coverage, partially overlapping settings where actions are visible in only a subset of views remain underexplored. This challenge becomes more severe in real-world scenarios, as many systems provide only limited input modalities and rely on sequence-level annotations instead of dense frame-level labels. In this study, we propose View-aware Cross-modal Knowledge Distillation (ViCoKD), a framework that distills knowledge from a fully supervised multi-modal teacher to a modality- and annotation-limited student. ViCoKD employs a cross-modal adapter with cross-modal attention, allowing the student to exploit multi-modal correlations while operating with incomplete modalities. Moreover, we propose a View-aware Consistency module to address view misalignment, where the same action may appear differently or only partially across viewpoints. It enforces prediction alignment when the action is co-visible across views, guided by human-detection masks and confidence-weighted Jensen–Shannon divergence between their predicted class distributions. Experiments on the real-world MultiSensor-Home dataset show that ViCoKD consistently outperforms competitive distillation methods across multiple backbones and environments, delivering significant gains and surpassing the teacher model under limited conditions.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide
WACV3
2026 Hierarchical graph attention networks with spatio-temporal class tokens for distributed audio-visual event classification
abstract
Abstract This paper presents a novel multi-view multimodal graph learning framework for distributed audio-visual event classification using synchronized sequences from multi-microphone and multi-camera sensors. Existing approaches often rely on simple aggregation strategies for multi-view multimodal inputs, which fail to adequately capture the complex spatio-temporal relationships both within and across modalities. To address this limitation, we propose a graph attention network architecture with individual frame-level sensor nodes for each microphone and camera, and three types of frame-level spatio-temporal nodes. In this framework, within each temporal frame, audio spatio-temporal nodes connect to microphones, video spatio-temporal nodes to cameras, and audio-video spatio-temporal nodes to all sensor nodes. Temporal edges further interconnect each spatio-temporal node with its corresponding node in preceding frames. This architecture enables dynamic aggregation of sensor node features through attention-weighted mechanisms, generating updated spatio-temporal nodes that capture intra-modal, inter-modal, intra-frame, and inter-frame relational dependencies. Experimental results on the MM-Office and MM-OR datasets demonstrate that the proposed framework significantly outperforms existing baseline methods for audio-visual event classification, highlighting its superior capability in modeling complex spatio-temporal dependencies across distributed sensor networks.
Vijay John, Yasutomo Kawanishi
Multim. Tools Appl.1
2026 MultiSensor-Home: Multi-modal multi-view dataset and benchmarks for action recognition in home environments
abstract
Multi-modal multi-view action recognition is a rapidly growing area in computer vision, with important applications in surveillance, smart homes, and assistive robotics. However, existing datasets often fail to capture real-world challenges such as distributed sensor layouts, asynchronous data streams, and limited frame-level annotations. To address these limitations, we introduce MultiSensor-Home, a novel multi-modal multi-view dataset specifically designed for realistic residential environments. It comprises 5,250 untrimmed videos recorded in two distinct residential environments, Home-1 and Home-2, using five distributed RGB and audio sensor units, where each video contains multiple sequential actions along with background segments. Each frame in the recorded sequences is manually annotated with fine-grained frame-level action labels, making it, to the best of our knowledge, the first densely annotated multi-view dataset for home activity recognition. To benchmark this dataset, we further propose Act ion Selection Learning-guided Transformer-based Sensor Fusion (ActFusion), a unified method that jointly models temporal dynamics and cross-view correspondence. It dynamically models cross-view relationships and selects informative frames, enabling robust training under both frame-level supervision, where the start and end timings of each action are labeled, and sequence-level supervision, where only action labels are provided. To support reproducible evaluation, we establish a comprehensive benchmark with standardized training and testing protocols. Extensive experiments on MultiSensor-Home and the existing MM-Office datasets show that ActFusion consistently outperforms baseline methods across diverse scenarios. By capturing the challenges in a realistic setting, MultiSensor-Home sets a new benchmark and encourages future research on robust and generalizable action recognition methods.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide
Pattern Recognit.3
2025 Modelling Spatio-Temporal Dynamics by Graph Attention Network for Distributed Multi-Microphone Sound Event Classification
abstract
This paper introduces a novel distributed multi-microphone sound event classification framework that uses graph attention networks to model spatial and temporal relationships between distributed multi-microphones. Existing methods utilize naive aggregation approaches like concatenation, averaging, or maximum operations for multiple sensor inputs resulting in information loss and limitations in capturing the complex spatial and temporal relationships in the multi-microphone sequence. To address this, we propose a framework based on the graph attention network to model the complex multi-microphone relationship. Our framework is based on two graph structures: a spatio-temporal graph (STG), which is algorithmically modeled to capture inter-microphone and inter-frame relationships, and a learnable fully-connected spatial graph (FCSG), which is designed to capture complementary details. Utilizing them, a graph attention network-based aggregation module effectively updates the graph nodes resulting in improved event classification accuracy. Experimental results on the MM-Office dataset demonstrate that our proposed framework significantly outperforms baseline methods for the event classification task.
Vijay John, Yasutomo Kawanishi
AVSS1
2025 Cross-modal Emotion-specific Attention model for Multimodal Emotion Recognition
abstract
Emotion recognition plays a crucial role in humanrobot interaction, where accurately interpreting human emotions through multiple modalities is essential for heartfelt communication. Although previous multimodal emotion recognition models have shown reasonable performance, there are two difficulties: (1) they often struggle to capture the fine-grained interactions between the modalities. (2) The prominent features are different across different emotions. To address these difficulties, we propose a novel Cross-modal Emotion-specific Attention model (CEA) for multimodal emotion recognition. The proposed model has two key components to address the difficulties above: (1) To capture the fine-grained interactions, we introduce the dense interaction matrix representation. (2) To focus more on emotion-specific features, we also introduce the specific emotion tokens. Combining these two components enhances the model’s ability to capture subtle emotional nuances and improves overall recognition accuracy. We evaluate the performance of the proposed architecture on the CREMA-D public audiovisual datasets through comprehensive ablation studies and comparison with baseline models. The results demonstrate that our Cross-modal Emotion-specific Attention model significantly outperforms the baseline methods, confirming its effectiveness in enhancing emotion recognition accuracy.
Jia-Yi Chen, Vijay John, Yasutomo Kawanishi
FG2
2025 MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
abstract
Multi-modal multi-view action recognition is a rapidly growing field in computer vision, offering significant potential for applications in surveillance. However, current datasets often fail to address real-world challenges such as widearea distributed settings, asynchronous data streams, and the lack of frame-level annotations. Furthermore, existing methods face difficulties in effectively modeling inter-view relationships and enhancing spatial feature learning. In this paper, we introduce the MultiSensor-Home dataset, a novel benchmark designed for comprehensive action recognition in home environments, and also propose the Multi-modal Multi-view Transformer-based Sensor Fusion (MultiTSF) method. The proposed MultiSensor-Home dataset features untrimmed videos captured by distributed sensors, providing high-resolution RGB and audio data along with detailed multi-view frame-level action labels. The proposed MultiTSF method leverages a Transformer-based fusion mechanism to dynamically model inter-view relationships. Furthermore, the proposed method integrates a human detection module to enhance spatial feature learning, guiding the model to prioritize frames with human activity to enhance action the recognition accuracy. Experiments on the proposed MultiSensor-Home and the existing MM-Office datasets demonstrate the superiority of MultiTSF over the state-of-the-art methods. Quantitative and qualitative results highlight the effectiveness of the proposed method in advancing real-world multi-modal multi-view action recognition.
Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide
FG3
2025 Reproducibility Companion Paper: Enhancing Model Interpretability with Local Attribution over Global Exploration
abstract
Reproducibility is indispensable for transferring explainable-AI algorithms from academic prototypes to production systems. This companion paper documents the artefacts, procedures, and outcomes that reproduce the empirical claims of ''Enhancing Model Interpretability with Local Attribution over Global Exploration'' (ACM MM 2024). We release a containerised archive containing source code, data-serialisation scripts, one-click executables, and a detailed README, all conforming to the ACM Multimedia reproducibility guidelines. The regenerated Insertion and Deletion scores deviate by only 2.2% on average. In addition, an exhaustive 10, 20, 30 3 grid-search over key hyper-parameters reveals a new configuration, (30, 20, 30), that improves the Insertion score of three convolutional backbones by 7.51% without additional code changes. These artefacts provide a rigorous, extensible foundation for future research on local attribution methods. Our code is available at: https://github.com/LMBTough/LA/
Zhibo Jin, Jiayu Zhang 0001, Fang Chen 0001, Jianlong Zhou, Vijay John, Florian Spiess 0001
ACM Multimedia6
2025 Multimodal Cascaded Framework with Multimodal Latent Loss Functions Robust to Missing Modalities
abstract
Despite interest in multimodal classification, few studies have addressed the missing modality problem in which an incomplete multimodal input with one or more missing modalities is classified as the target class. The missing modality problem is shown to reduce the classification accuracy as the discriminative power of the obtained feature space is reduced. In this study, we address the missing modality problem in multimodal classification using a novel cascaded framework. The proposed framework is formulated in the feature space to address the missing modality problem by generating complete multimodal data from incomplete multimodal data. Subsequently, an optimal multimodal data is obtained by feature selection of the generated and original data. The proposed cascaded framework consists of three steps: feature extraction, feature generation, and classification. The framework is formulated to handle both complete and incomplete multimodal data simultaneously. The cascaded framework is trained using novel latent loss functions: missing modality joint loss, centroid joint loss, and latent prior loss. These loss functions, based on metric learning, are designed to ensure that data from the same class remain proximate in the latent space irrespective of the presence or absence of modality data. The cascaded framework is validated on bimodal audio-visible RAVDESS and trimodal audio-visible-thermal Speaking Faces datasets. The experimental results show that the cascaded framework improves classification accuracy even with incomplete multimodal data.
Vijay John, Yasutomo Kawanishi
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Generating Pseudo-Strong Labels from Weak Labels for Distributed Multi-Microphone Sound Event Detection
Vijay John, Yasutomo Kawanishi
ICPR (10)1
2023 Audio-Visual Sensor Fusion Framework Using Person Attributes Robust to Missing Visual Modality for Person Recognition
Vijay John, Yasutomo Kawanishi
MMM (2)1
2023 Multimodal Cascaded Framework with Metric Learning Robust to Missing Modalities for Person Classification
abstract
This paper addresses the missing modality problem in multimodal person classification, where an incomplete multimodal input with one modality missing is classified into predefined person classes. A multimodal cascaded framework with three deep learning models is proposed, where model parameters, outputs, and latent space learnt at a given step are transferred to the model in a subsequent step. The cascaded framework addresses the missing modality problem by, firstly, generating the complete multimodal data from the incomplete multimodal data in the feature space via a latent space. Subsequently, the generated and original multimodal features are effectively merged and embedded into a final latent space to estimate the person label. During the learning phase, the cascaded framework uses two novel latent loss functions, the missing modality joint loss, and latent prior loss to learn the different latent spaces. The missing modality joint loss ensures that the similar class latent data are close to each other, even if a modality is missing. In the cascaded framework, the latent prior loss learns the final latent space using a previously learnt latent space as a prior. The proposed framework is validated on the audio-visible RAVDESS and the visible-thermal Speaking Faces datasets. A detailed comparative analysis and an ablation analysis are performed, which demonstrate that the proposed framework enhances the robustness of person classification even under conditions of missing modalities, reporting an average of 21.75% increase and 25.73% increase over the baseline algorithms on the RAVDESS and Speaking Faces datasets.
Vijay John, Yasutomo Kawanishi
MMSys1
2022 Audio and Video-based Emotion Recognition using Multimodal Transformers
abstract
Emotion recognition, an important research problem in human-robot interactions, is primarily achieved by extracting human emotions from audio and visual data. State-of-the-art performance is reported by audio-visual sensor fusion algorithms using deep learning models such as CNN, RNN, and LSTM. However, the RNN and LSTM are shown to be limited in handling the long-term dependencies over the entire input sequence. In this work, we propose to improve the performance of audio-visual emotion recognition using a novel transformer-based model, containing three transformer branches, named multimodal transformers. The three transformer branches, in our work, compute the audio self-attention, the video self-attention, and the audio-video cross attention. The self-attention branches identify the most relevant information in the audio and video input, while the cross-attention branch identifies the most relevant audio-video interactive information. The relevant information from these three branches report the best performance in our ablation study. We also propose a novel temporal embedding scheme, termed block embedding, to add the temporal information to the visual feature, derived from the multiple frames in the video. The proposed architecture is validated using the RAVDESS, CREMA-D, and SAVEE audio-visual public datasets. A detailed ablation study and comparative analysis with baseline models is performed. The results show that the proposed multi-modal transformer framework is better than the baseline methods.
Vijay John, Yasutomo Kawanishi
ICPR1
2022 A Multimodal Sensor Fusion Framework Robust to Missing Modalities for Person Recognition
abstract
Utilizing the sensor characteristics of the audio, visible camera, and thermal camera, the robustness of person recognition can be enhanced. Existing multimodal person recognition frameworks are primarily formulated assuming that multimodal data is always available. In this paper, we propose a novel trimodal sensor fusion framework using the audio, visible, and thermal camera, which addresses the missing modality problem. In the framework, a novel deep latent embedding framework, termed the AVTNet, is proposed to learn multiple latent embeddings. Also, a novel loss function, termed missing modality loss, accounts for possible missing modalities based on the triplet loss calculation while learning the individual latent embeddings. Additionally, a joint latent embedding utilizing the trimodal data is learnt using the multi-head attention transformer, which assigns attention weights to the different modalities. The different latent embeddings are subsequently used to train a deep neural network. The proposed framework is validated on the Speaking Faces dataset. A comparative analysis with baseline algorithms shows that the proposed framework significantly increases the person recognition accuracy while accounting for missing modalities.
Vijay John, Yasutomo Kawanishi
MMAsia1
2022 Distribution-dependent feature selection for deep neural networks
Xuebin Zhao, Weifu Li, Hong Chen 0004, Yingjie Wang 0007, Vijay John
Appl. Intell.6
2021 Deep Fusion-based Visible and Thermal Camera Forecasting using Seq2Seq GAN
Vijay John, Annamalai Lakshmanan, Ali Boyali, Simon Thompson 0002, Seiichi Mita
IV1
2021 Deep Learning Thermal Image Translation for Night Vision Perception
abstract
Context enhancement is critical for the environmental perception in night vision applications, especially for the dark night situation without sufficient illumination. In this article, we propose a thermal image translation method, which can translate thermal/infrared (IR) images into color visible (VI) images, called IR2VI. The IR2VI consists of two cascaded steps: translation from nighttime thermal IR images to gray-scale visible images (GVI), which is called IR-GVI; and the translation from GVI to color visible images (CVI), which is known as GVI-CVI in this article. For the first step, we develop the Texture-Net, a novel unsupervised image translation neural network based on generative adversarial networks. Texture-Net can learn the intrinsic characteristics from the GVI and integrate them into the IR image. In comparison with the state-of-the-art unsupervised image translation methods, the proposed Texture-Net is able to address some common challenges, e.g., incorrect mapping and lack of fine details, with a structure connection module and a region-of-interest focal loss. For the second step, we investigated the state-of-the-art gray-scale image colorization methods and integrate the deep convolutional neural network into the IR2VI framework. The results of the comprehensive evaluation experiments demonstrate the effectiveness of the proposed IR2VI image translation method. This solution will contribute to the environmental perception and understanding in varied night vision applications.
Shuo Liu 0009, Mingliang Gao 0001, Vijay John, Zheng Liu 0002, Erik Blasch
ACM Trans. Intell. Syst. Technol.3
2020 Enhancing Depth Quality of Stereo Vision using Deep Learning-based Prior Information of the Driving Environment
abstract
Generation of high density depth values of the driving environment is indispensable for autonomous driving. Stereo vision is one of the practical and effective methods to generate these depth values. However, the accuracy of the stereo vision is limited by texture-less regions, such as sky and road areas, and repeated patterns in the image. To overcome these problems, we propose to enhance the stereo generated depth by incorporating prior information of the driving environment. Prior information, generated by deep learning-based U-Net model, is utilized in a novel post-processing mathematical framework to refine the stereo generated depth. The proposed mathematical framework is formulated as an optimization problem, which refines the errors due to texture-less regions and repeated patterns. Owing to its mathematical formulation, the post-processing framework is not a black-box and is explainable, and can be readily utilized for depth maps generated by any stereo vision algorithm. The proposed framework is qualitatively validated on the acquired dataset and KITTI dataset. The results obtained show that the proposed framework improves the stereo depth generation accuracy.
Weifu Li, Vijay John, Seiichi Mita
ICPR2
2019 Multi-Agent Reinforcement Learning for Autonomous On Demand Vehicles
abstract
In this study, we elaborate the procedure of designing a supervisory controller for the Autonomous Transit on Demand Vehicle (ATODV) system. Reinforcement learning is implemented to reduce the mean waiting time of the passengers, and a cost function is introduced to penalize the energy consumption of the electric vehicles. A stochastic simulation environment for an ATODV pilot project is coded in the Python environment to train the autonomous cart decision process as agents with artificial intelligence. Passenger group behavior, get-on and get-off times, destinations are modeled as random variables. A single Deep Q-Learning Network is trained subject to multi-agent settings. The ATODV system's independent decision making for the carts to reduce the passenger's waiting time while constraining the energy consumption and empty vehicle motion is evaluated.
Ali Boyali, Naohisa Hashimoto, Vijay John, Tankut Acarman
IV3
2019 RVNet: Deep Sensor Fusion of Monocular Camera and Radar for Image-Based Obstacle Detection in Challenging Environments
Vijay John, Seiichi Mita
PSIVT1
2018 Free Space, Visible and Missing Lane Marker Estimation using the PsiNet and Extra Trees Regression
abstract
In this paper, a vision-based multilabel deep learning framework is combined with an extra trees regression framework to estimate the free space, visible ego-lane markers and missing ego-lane markers. The multilabel deep learning framework, the PsiNet, with two semantic segmentation layers and one multiclass classifer layer, estimates the free space and visible ego-lane markers, while the deep learning-based extra trees regression framework estimates the missing ego-lane markers. The missing ego-lane markers are predicted using image-based deep features extracted from the multilabel framework. To account for spatial variation in the missing ego lane markers, multiple extra trees regression models are trained. During testing, the multiclass label estimated by the multilabel framework is used to retrieve the corresponding extra trees regression model. The proposed framework combining the deep learning-based semantic segmentation and regression frameworks is termed the PsiNet-ET framework. We validate our proposed framework using the multiple acquired datasets. A comparative analysis with baseline algorithms, along with a parametric analysis is performed. The experimental results show that the proposed framework robustly estimates the free space, visible and missing lane markers even for challenging road scenes.
Vijay John, Nithilan Meenakshi Karunakaran, Chunzhao Guo, Kiyosumi Kidono, Seiichi Mita
ICPR1
2018 Sensor Fusion of Intensity and Depth Cues using the ChiNet for Semantic Segmentation of Road Scenes
abstract
Vision-based environment perception is an important research topic for autonomous driving and advanced driver assistance systems. Vision sensors, such as the monocular camera and stereo camera, are widely used for environment perception. The monocular camera provides the appearance information like intensity, and the stereo camera provides the depth information. The appearance and depth information are complementary, and their effective fusion would result in robust environment perception. Consequently, in this paper, we propose a novel deep learning framework, termed as the ChiNet, for the effective sensor fusion of the appearance and depth information for free space and road object estimation. The ChiNet has two input branches and two output branches. The ChiNet input branches contains separate branches for the intensity and depth information. For the output branches, the ChiNet contains separate branches for the free space and road object semantic segmentation. A comparative of the proposed framework with state-of-the-art baseline algorithms is performed using an acquired dataset. Moreover, a detailed parameter analysis is performed to validate the ChiNet architecture as well as the advantages of sensor fusion. The experimental results show that the ChiNet is better than baseline algorithms. We also show that the proposed ChiNet architecture is better than other variations of the ChiNet architecture.
Vijay John, Nithilan Meenakshi Karunakaran, Seiichi Mita, Hossein Tehrani Niknejad, M. Konishi, Kazuhisa Ishimaru, Tomoyuki Oishi
Intelligent Vehicles Symposium1
2018 Vision and Dead Reckoning-based End-to-End Parking for Autonomous Vehicles
abstract
In this paper, a combined vision and dead reckoning-based parking system for end-to-end driving are proposed. Standard autonomous parking frameworks contain multiple modules with each module having its own limitation. On the other hand, the proposed parking framework consists of a single end-to-end module, which reduces these inherent limitations. In the proposed deep learning-based parking system, a novel iterative two-stage learning framework is utilized to predict the steering angles and gear status using a front and back mounted monocular camera. In the first stage of the proposed framework, the encoder-decoder architecture is used to predict an initial estimate of the steering angle trajectory from multiple frames of the front or the back monocular camera. The camera used for steering estimated is selected using the gear status estimate. The gear status is predefined during initialization and estimated subsequently in the second stage of the proposed framework. In the second stage of the proposed framework, the initial estimate of the steering angle trajectory along with the vehicles heading angle, and absolute position is given as an input to the long short-term memory network to estimate the optimal steering angle and gear status. The proposed framework is validated on an acquired dataset. A comparative analysis of baseline algorithms and detailed parametric analysis are performed. The experimental results show that the proposed framework is better than the baseline end-to-end algorithms.
Rathour Swarn, Vijay John, Nithilan Meenakshi Karunakaran, Seiichi Mita
Intelligent Vehicles Symposium2
2017 Automated driving by monocular camera using deep mixture of experts
abstract
In this paper, we propose a real-time vision-based filtering algorithm for steering angle estimation in autonomous driving. A novel scene-based particle filtering algorithm is used to estimate and track the steering angle using images obtained from a monocular camera. Highly accurate proposal distributions and likelihood are modeled for the second order particle filter, at the scene-level, using deep learning. For every road scene, an individual proposal distribution and likelihood model is learnt for the corresponding particle filter. The proposal distribution is modeled using a novel long short term memory network-mixture-of-expert-based regression framework. To facilitate the learning of highly accurate proposal distributions, each road scene is partitioned into straight driving, left turning and right turning sub-partitions. Subsequently, each expert in the regression framework accurately model the expert driver's behavior within a specific partition of the given road scene. Owing to the accuracy of the modelled proposal distributions, the steering angle is robustly tracked, even with a limited number of sampled particles. The sampled particles are assigned importance weights using a deep learning-based likelihood. The likelihood is modeled with a convolutional neural network and extra trees-based regression framework, which predicts the steering angle for a given image. We validate our proposed algorithm using multiple sequences. We perform a detailed parameter analysis and a comparative analysis of our proposed algorithm with different baseline algorithms. Experimental results show that the proposed algorithm can robustly track the steering angles with few particles in real-time even for challenging scenes.
Vijay John, Seiichi Mita, Hossein Tehrani Niknejad, Kazuhisa Ishimaru
Intelligent Vehicles Symposium1
2017 Real-time hand posture and gesture-based touchless automotive user interface using deep learning
abstract
In this study, a vision based in-car entertainment user interface is presented. The user interface is designed using a hand posture and gesture recognition algorithm in deep learning framework. The hand posture recognition algorithm is formulated using the convolutional neural network to perform the fundamental tasks in the user interface. The hand gesture recognition algorithm is formulated using the long-term recurrent convolutional neural network to intuitively interact with the touchless automotive user interface in a detailed manner. In the recurrent deep learning framework, typically, the gesture frames are taken from a uniformly sampled image sequence. In this work, the recurrent structure is enhanced using a reduced number of input frames captured from the image sequence. The reduced input frames or key frames represent the action present in the video sequence. Sparse dictionary learning provide reliable key frame extraction from video sequences. However, sparse dictionary learning is computationally expensive, and are individually optimized for every video sequence. In this paper, we propose to approximate sparse dictionary learning using a non-linear regression framework. The multilayer perceptron is utilized to model the non-linear regression framework. The optimal neural network architecture is identified after a detailed evaluation. We evaluate the proposed recognition methods on public datasets. The proposed methods yield a recognition accuracy of 92% and 90% for pose and gestures, respectively. The combined hand posture and gesture recognition takes 82ms which is a reasonable for real time implementation.
Vijay John, Makoto Umetsu, Ali Boyali, Seiichi Mita, Masayuki Imanishi, Norio Sanma, Syunsuke Shibata
Intelligent Vehicles Symposium1
2017 3D point cloud map based vehicle localization using stereo camera
abstract
Nowadays, the driverless automobiles have become a near reality and are going to become widely available. For autonomous navigation, the vehicles need to precise localize itself within a pre-defined map. In this paper, we propose a novel algorithm for the problem of three-dimensional (3D) point cloud map (PCL) based localization using a stereo camera. This 3D point cloud map consists of dense 3D geometric information and intensity measures of surface reflectivity value generated by the 3D light detection and ranging (LIDAR) scanner based mapping system. Although some LIDAR based localization algorithms have been proposed, in this paper we present a comparable centimeter-level accuracy localization algorithm using much cheaper and commodity stereo camera. Specifically, at each candidate position we transform the 3D data points from the real-world coordinate system to the camera coordinate system and synthetic the virtual depth and intensity images from the 3D PCL map. We localize the ego vehicle by estimating the transformation between the real-world and vehicle coordinates in each frame by matching these virtual images with the stereo depth and intensity images. In the experiment part, we show that although the 3D map was generated 3 years ago, the proposed algorithm still can produce reliable localization results even in many difficult cases, such as shadow, dynamic objects, new lane marker and night.
Yuquan Xu, Vijay John, Seiichi Mita, Hossein Tehrani Niknejad, Kazuhisa Ishimaru, Sakiko Nishino
Intelligent Vehicles Symposium2
2016 Fast road scene segmentation using deep learning and scene-based models
abstract
Pixel-labeling approaches using semantic segmentation play an important role in road scene understanding. In recent years, deep learning approaches such as the deconvolutional neural network have been used for semantic segmentation, obtaining state-of-the-art results. However, the segmentation results have limited object delineation. In this paper, we adopt the de-convolutional neural network to perform the semantic segmentation of the road scene using colour and depth information. Moreover, we improve the network's limited object delineation within a computationally efficient framework using novel features that are learnt at the pixel-level and patch-level for different road scenes. The patch-level features represent the road scene geometry. On the other hand, the learnt pixel-level features represent the appearance and depth information. The features learnt for the different road scenes are indexed with the scene's pre-defined label. Following the indexing, the random forest classifier is trained to retrieve the relevant geometric and appearance-depth features for a given road scene. The retrieved features are then used to refine identified error regions in the initial semantic segmentation estimate. Our proposed algorithm is evaluated on an acquired dataset and compared with state-of-the-art baseline algorithms. We also perform a detailed parametric evaluation of our proposed framework. The experimental results show that our proposed algorithm reports better accuracy.
Vijay John, Kiyosumi Kidono, Chunzhao Guo, Hossein Tehrani Niknejad, Seiichi Mita, Kasuhiza Ishimaru
ICPR1
2015 Real-Time Lane Estimation Using Deep Features and Extra Trees Regression
Vijay John, Zheng Liu 0002, Chunzhao Guo, Seiichi Mita, Kiyosumi Kidono
PSIVT1
2014 Charting-based subspace learning for video-based human action classification
Vijay John, Emanuele Trucco
Mach. Vis. Appl.1
2013 Solving Person Re-identification in Non-overlapping Camera using Efficient Gibbs Sampling
abstract
This paper proposes a novel probabilistic approach for appearance-based person reidentification in non-overlapping camera networks.It accounts for varying illumination, varying camera gain and has low computational complexity.More specifically, we present a graphical model where we model the person's appearance in addition to camera illumination and gain.We analytically derive the solutions for the person's appearance and camera properties, and use a novel constant time Gibbs sampling scheme to estimate the identification labels.We validate our algorithm on two indoor datasets and perform a comparative analysis with existing algorithms.We demonstrate significantly increased re-identification accuracy in addition to significantly reducing the computational complexity on our datasets.
Vijay John, Gwenn Englebienne, Ben J. A. Kröse
BMVC1
2013 Person re-identification using height-based gait in colour depth camera
abstract
We address the problem of person re-identification in colour-depth camera using the height temporal information of people. Our proposed gait-based feature corresponds to the frequency response of the height temporal information. We demonstrate that the discriminative periodic motion associated with human gait is encoded within the height temporal information. Additionally, we also investigate the discriminative ability of a novel feature vector obtained by the integration of the height temporal information with a color and height-based appearance model. Given the proposed features we adopt a feature selection scheme for each person, based on KL-divergence, to identify the discriminative subset of frequency bins for the height-based gait feature that enhances the overall identification accuracy. To identify the test person, we formulate a probabilistic matching framework incorporating the selected frequency bins. We validate our algorithm on the publicly available TUM-GAID dataset and our studio datasets and report over 80% accuracy for 75 people with our combined feature, significantly better than standard colour-based features. Additionally, we also observe that height-based gait features reports over 90% for smaller population and are rotation invariant, robust to appearance noise and occlusion.
Vijay John, Gwenn Englebienne, Ben J. A. Kröse
ICIP1
2010 Markerless Multi-view Articulated Pose Estimation Using Adaptive Hierarchical Particle Swarm Optimisation
Spela Ivekovic, Vijay John, Emanuele Trucco
EvoApplications (1)2
2010 Markerless human articulated tracking using hierarchical particle swarm optimisation
Vijay John, Emanuele Trucco, Spela Ivekovic
Image Vis. Comput.1