Yasushi Yagi

dblp:67/6296 · DBLP profile ↗
← Back
184ranked-venue papers
24as first author
20since 2021 · last 2026
0000-0002-3546-8071ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 148 · 22 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 108 · 3 first-author · 11 since 2021Systems, architecture and hardware · 35 · 15 first-authorHuman-computer interaction and ubiquitous computing · 21 · 8 since 2021Security and privacy · 20 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Short-Term Temporal Behavioral Drift in Smartwatch User Authentication: A Case Study Using Apple Watch Sensor Logs
Maharage Nisansala Sevwandi Perera, Takeshi Kawamoto, Allam Shehata, Franziska Zimmer, Ryosuke Kobayashi, Mhd Irvan, Rie Shigetomi Yamaguchi, Yasushi Yagi
SECRYPT (1)8
2025 Behavioral Signature Decoding: Facial Landmark-based Graph Learning for Cybernetic Avatar Authentication
abstract
With the rapid advancement of AI-generated videos, distinguishing synthetic content from genuine human-driven content has become increasingly difficult, threatening the integrity of human authenticity and creative expression. In this context, Cybernetic Avatar (CA) is introduced as a digital entity that mirrors a remote operator’s facial expressions, gestures, and speech in virtual environments, posing new challenges for secure identity verification. A critical threat emerges when unauthorized users manipulate a CA, potentially deceiving both systems and human observers. This paper addresses the CA Authentication problem, which seeks to verify the true teleoperator behind a CA video despite the CA’s mutable appearance and expressive behaviors. More specifically, we propose a robust CA authentication framework that leverages spatio-temporal facial behavior captured from the CA video to authenticate the legitimate teleoperator. To effectively learn identity-sensitive motion patterns (signature), we develop a Behavior Signature Decoder Graph Convolutional Network (BSDec-GCN) that constructs a constrained spatio-temporal graph to amplify identity-specific dynamics and suppress inter-user ambiguity. Furthermore, we introduce a dual landmark and graph-level losses that boost discrimination. The comprehensive experiments and the thorough ablation studies demonstrate the reliability of the proposed framework with a competitive performance against the existing baseline methods. To the best of our knowledge, this work presents the first graph-based learning approach tailored for Cybernetic Avatar authentication, opening a new direction for securing virtual identity in the era of AI-mediated communication.
Ammar Alsherfawi, Jianhang Zhou, Allam Shehata, Yasushi Yagi
IJCB4
2025 Towards Robust and Efficient Continuous Face Authentication via Disentangled Representation Learning and Adaptive Identity Completion
abstract
In this paper, we focus on the problem of continuous face authentication, which aims to verify a user’s identity persistently over time from streaming or sequential facial inputs. Unlike conventional face authentication methods that operate in a one-time or set-to-set manner, our approach produces segment-wise identity decisions, enabling real-time identity monitoring in dynamic scenarios. To tackle challenges such as temporal variation and identity-irrelevant noise (e.g., pose, blur, occlusion), we propose a novel three-stage framework comprising: (1) disentangled representation learning to suppress identity-irrelevant components in pre-trained embeddings, (2) attention-based intra-segment aggregation to extract robust identity cues within each segment, and (3) inter-segment adaptive identity feature completion to incrementally integrate new segments with prior identity representations for stable and efficient prediction. Extensive experiments on IJB-B, IJB-C benchmark datasets, and our own collected dataset under real-world video conferencing conditions demonstrate that our method achieves competitive accuracy with notable improvements in inference efficiency (up to 2× speed-up) over state-of-the-art baselines. These results validate the effectiveness and practicality of our framework for real-world continuous authentication systems.
Xiang Li 0028, Chi Xu 0003, Yasushi Yagi
IJCB3
2025 Reconstruct and De-identify (RaD): A Joint Task Framework for Face Reconstruction and De-identification Leveraging the 3D Morphable Model Explainability
abstract
Current face de-identification methods often struggle to balance robust privacy protection with preserving image utility. While existing approaches effectively obscure identity, they frequently degrade visual quality, limiting their practical applicability. To address this challenge, we propose a robust reconstruction and de-identification framework (RaD) that leverages the 3D Morphable Model (3DMM) explicit representation. We first utilize a CNN encoder to predict the disentangled 3DMM coefficients, enabling a coarse reconstruction via a differentiable renderer. To enrich facial detail beyond the 3DMM’s topology and statistical priors, we introduce a personalized albedo generator (PAG) that adds fine texture details. For de-identification, we then apply a transformer-based identity protector (IP) to manipulate the 3DMM shape and texture parameters in order to conceal identity-sensitive features while preserving the photorealism of the reconstructed images. Finally, an Image Enhancement Module (IEM) refines the outputs, removing any artifacts and further enhancing visual quality. Extensive experiments on multiple benchmarks demonstrate that our framework outperforms state-of-the-art methods both quantitatively and qualitatively, making it well-suited for applications that require realistic reconstructions along with privacy preservation without compromising image quality.
Allam Shehata, Mohamad Ammar Alsherfawi Aljazaerly, Yasushi Yagi
IJCB3
2025 EEG-based User Authentication in Realistic Scenarios: From Solo Reading to Dialog Games
abstract
While electroencephalography (EEG)-based user authentication has demonstrated strong potential in controlled laboratory conditions, its reliability in more realistic settings remains underexplored. In this work, we investigate EEG-based user authentication in realistic interactive scenarios involving cognitively and behaviorally rich tasks. We collected EEG data from subjects under two conditions: solo reading aloud and two-person dialog-based games. These scenarios represent practical, everyday activities where users produce speech and engage in turn-taking interactions, introducing non-stationarity, muscle artifacts, and attention shifts that challenge conventional EEG-based models. To evaluate the user authentication performance under these challenges, we systematically apply various deep learning backbone models and loss functions under a fair and consistent experimental protocol. Specifically, we first divide the whole EEG signal into short windows with pre-defined length and stride. For each window, a time-frequency representation is computed using a continuous wavelet transform, resulting in a sequence of 2D time-frequency maps that capture the non-stationary characteristics of EEG signals. These maps are then fed into various backbone models, such as long short term memory (LSTM) and convolutional neural network (CNN), which are trained to extract robust identity-discriminative features. To further enhance inter-subject separability, we employ the Softmax loss or ArcFace loss during training. Experimental results demonstrate that even under dynamic and less controlled conditions, EEG signals retain individual-specific patterns, yielding high authentication accuracy. These findings highlight the feasibility of extending EEG-based biometrics to more natural environments.
Chi Xu 0003, Xiang Li 0028, Shuqiong Wu, Yasushi Yagi
IJCB4
2025 Warp Gait Across Ages: Cross-age Gait Video Translation with Part-aware Flow Warping
abstract
Cross-age gait video translation aims to translate a gait video of a subject captured at a certain age into another age while preserving the individual identity and realism. This has a wide range of applications, including age-invariant gait recognition. In this paper, we propose a method for cross-age gait video translation using a spatially smooth geometric warping field. More specifically, we employ a variant of the spatial transformer network (STN), called flow warping, which achieves accurate coarse-to-fine warping field inference. In addition, we extend the flow warping framework by introducing gait stance-dependent warping fields to better represent the temporal variations in the warping fields. We also develop an approach called Part-STN to infer part-dependent warping fields and then merge them into a unified warping field to retain consistency within each body part. We then train Part-STN with loss functions considering three aspects: aging effect (using an age group classification loss), identity preservation (using a cycle-consistency loss and triplet loss on the representation learned from a pretrained gait recognition model), and realism (using an adversarial loss and smoothness loss on estimated flow warping). Our framework outperforms state-of-the-art cross-age gait video translation methods quantitatively and qualitatively on the largest gait database OULP-Age, with respect to both age group classification, identity recognition and realism.
Yiyi Zhang 0002, Hanchong Yan, Liqing Zhang 0001, Yasushi Yagi
IJCB4
2025 Pedestrian Motion Reconstruction: A Large-scale Benchmark via Mixed Reality Rendering with Multiple Perspectives and Modalities
abstract
Reconstructing pedestrian motion from dynamic sensors, with a focus on pedestrian intention, is crucial for advancing autonomous driving safety. However, this task is challenging due to data limitations arising from technical complexities, safety, and cost concerns. We introduce the Pedestrian Motion Reconstruction (PMR) dataset, which focuses on pedestrian intention to reconstruct behavior using multiple perspectives and modalities. PMR is developed from a mixed reality platform that combines real-world realism with the extensive, accurate labels of simulations, thereby reducing costs and risks. It captures the intricate dynamics of pedestrian interactions with objects and vehicles, using different modalities for a comprehensive understanding of human-vehicle interaction. Analyses show that PMR can naturally exhibit pedestrian intent and simulate extreme cases. PMR features a vast collection of data from 54 subjects interacting across 12 urban settings with 7 objects, encompassing 12,138 sequences with diverse weather conditions and vehicle speeds. This data provides a rich foundation for modeling pedestrian intent through multi-view and multi-modal insights. We also conduct comprehensive benchmark assessments across different modalities to thoroughly evaluate pedestrian motion reconstruction methods.
Yiyi Zhang 0002, Xinhao Hu, Li Niu 0002, Jianfu Zhang 0003, Yasushi Makihara, Yasushi Yagi, Wenlong Liao, Junchi Yan, Liqing Zhang 0001
ICLR7
2025 Two-stage CNN with weakly supervised segmentation for skin lesion classification
Anggasta Aji Azhari, Novanto Yudistira, Agus Wahyu Widodo, Yasushi Yagi
Multim. Tools Appl.4
2025 Predicting Future Cognitive Decline From Long-Term Observations of Dual-Task Performance Data
abstract
Early stage detection of cognitive decline is crucial for effective prevention and treatment of dementia. However, current approaches based on MRI or biomarkers are expensive and impractical, making them unsuitable for early-stage detection from daily measurements. A suitable option is the dual-task paradigm, which involves simultaneously performing two tasks (typically a physical task combined with a cognitive task). This approach has proven effective in assessing daily cognitive status. The underlying principle is that dual-task performance reflects the maximum cognitive load that can be handled by participants, which in turn reflects their current cognitive function. However, a one-time dual-task test cannot predict future changes in cognitive function. In this study, we present the first attempt at leveraging long-term observations of dual-task performance data. Our results show that changes in dual-task performance over time are associated with future cognitive changes. Our approach extracts temporal features from six months of dual-task performance data, and predicts future cognitive decline over the next two years using a machine learning model. Our experimental results yielded an accuracy comparable to that returned by MRI scans, thus demonstrating that the proposed approach can achieve early detection of future cognitive decline from routine dual-task measurements.
Shuqiong Wu, Tomoya Noguchi, Fumio Okura, Yasushi Yagi
IEEE J. Biomed. Health Informatics4
2024 On Cropping for Gait Recognition: Does Constant-velocity Locomotion Assumption Improve Gait Recognition Accuracy?
abstract
Most of gait recognition studies focus on feature extraction and matching steps by using cropped image sequences as inputs. We usually use frame-by-frame tight bounding boxes (TBBs) obtained by pedestrian detection or instance segmentation for cropping. Cropping by the TBB, however, suffers from apparent scale changes within gait period (e.g., apparent height changes between single/double support phases), which may cause a drop in gait recognition accuracy. We therefore propose a method of cropping for gait recognition to better preserve the apparent scale by introducing constant-velocity locomotion (CVL) assumption for a short period (e.g., one second). We derive that a bounding box sequence (BBS) in the 2D image coordinate under CVL assumption in the 3D camera coordinate, is represented by non-linear interpolation between the starting and ending frames without camera calibration parameters. We then estimate BBS parameters (i.e., bounding boxes at the starting and ending frames) by generalized Hough transform with voting from pedestrian region proposals. Experiments with OU-MVLP show that the proposed cropping improves the gait recognition accuracies for both model-based and appearance-based approaches.
Tappei Okimura, Xiang Li 0028, Chi Xu 0003, Yasushi Yagi
IJCB4
2023 Online Model-based Gait Age and Gender Estimation
abstract
This paper presents an online human model-based framework for gait-based age and gender estimation from a sequence of monocular frames. More specifically, we fine-tune a human mesh recovery model (i.e., HMR) to estimate the shape and pose parameters of a predefined 3D human model (i.e., SMPL). We then utilize the estimated parameters to predict the age and gender of the walking subject. To make the age and gender estimation task more favorable for real-time applications, we consider estimating the corresponding probability distributions of age and gender, which preserve the prediction uncertainty. Experiments on the world’s largest multi-view gait age and gender estimation dataset showed the superiority of the proposed method compared to the existing appearance-based baseline. We implement online standalone and client-server systems based on the proposed framework to demonstrate the performance of real-time estimation. We further propose a geometric correction step to the input gait sequence for a more generalization capability of the online system.
Allam Shehata, Mohamad Ammar Alsherfawi Aljazaerly, Levin Gäher, Xiang Li 0028, Yasushi Makihara, Yasushi Yagi
IJCB6
2023 Natural Image Matting with Attended Global Context
Yiyi Zhang 0002, Li Niu 0002, Yasushi Makihara, Jianfu Zhang 0003, Weijie Zhao 0003, Yasushi Yagi, Liqing Zhang 0001
J. Comput. Sci. Technol.6
2023 Action Recognition From a Single Coded Image
abstract
The unprecedented success of deep convolutional neural networks (CNN) on the task of video-based human action recognition assumes the availability of good resolution videos and resources to develop and deploy complex models. Unfortunately, certain budgetary and environmental constraints on the camera system and the recognition model may not be able to accommodate these assumptions and require reducing their complexity. To alleviate these issues, we introduce a deep sensing solution to directly recognize human actions from coded exposure images. Our deep sensing solution consists of a binary CNN-based encoder network that emulates the capturing of a coded exposure image of a dynamic scene using a coded exposure camera, followed by a 2D CNN for recognizing human action in the captured coded exposure image. Furthermore, we propose a novel knowledge distillation framework to jointly train the encoder and the action recognition model and show that the proposed training approach improves the action recognition accuracy by an absolute margin of 6.2%, 2.9%, and 7.9% on Something$^{2}$-v2, Kinetics-400, and UCF-101 datasets, respectively, in comparison to our previous approach. Finally, we built a prototype coded exposure camera using LCoS to validate the feasibility of our deep sensing solution. Our evaluation of the prototype camera show results that are consistent with the simulation results.
Sudhakar Kumawat, Tadashi Okawara, Michitaka Yoshida, Hajime Nagahara, Yasushi Yagi
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Annotator-dependent uncertainty-aware estimation of gait relative attributes
abstract
In this paper, we describe an uncertainty-aware estimation framework for gait relative attributes. We specifically design a two-stream network model that takes a pair of gait videos as input. It then outputs a corresponding pair of Gaussian distributions of gait absolute attribute scores and annotator-dependent gait relative attribute label distributions. Moreover, we propose a differentiable annotator-independent uncertainty layer to estimate the gait relative attribute score distribution from the absolute distributions then map it to a relative attribute label distribution using the computation of cumulative distribution functions. Furthermore, we propose another annotator-dependent uncertainty layer to estimate the uncertainty on the gait relative attribute labels in terms of a set of trainable transition matrices. Finally, we design a joint loss function on the relative attribute label distribution to learn the model parameters. Experiments on two gait relative attribute datasets demonstrated the effectiveness of the proposed method against baselines in quantitative and qualitative evaluations.
Allam Shehata, Yasushi Makihara, Daigo Muramatsu, Md. Atiqur Rahman Ahad, Yasushi Yagi
Pattern Recognit.5
2023 Occlusion-Aware Human Mesh Model-Based Gait Recognition
abstract
Partial occlusion of the human body caused by obstacles or a limited camera field of view often occurs in surveillance videos, which affects the performance of gait recognition in practice. Existing methods for gait recognition against occlusion require a bounding box or the height of a full human body as a prerequisite, which is unobserved in occlusion scenarios. In this paper, we propose an occlusion-aware model-based gait recognition method that works directly on gait videos under occlusion without the above-mentioned prerequisite. Specifically, given a gait sequence that only contains non-occluded body parts in the images, we directly fit a skinned multi-person linear (SMPL)-based human mesh model to the input images without any pre-normalization or registration of the human body. We further use the pose and shape features extracted from the estimated SMPL model for recognition purposes, and use the extracted camera parameters in the occlusion attenuation module to reduce intra-subject variation in human model fitting caused by occlusion pattern differences. Experiments on occlusion samples simulated from the OU-MVLP dataset demonstrated the effectiveness of the proposed method, which outperformed state-of-the-art gait recognition methods by about 15% rank-1 identification rate and 2% equal error rate in the identification and verification scenarios, respectively.
Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi
IEEE Trans. Inf. Forensics Secur.4
2022 Investigating strategies towards adversarially robust time series classification
abstract
Deep neural networks have been shown to be vulnerable against specifically-crafted perturbations designed to affect their predictive performance. Such perturbations, formally termed ‘adversarial attacks’ have been designed for various domains in the literature, most prominently in computer vision and more recently, in time series classification. Therefore there is a need to derive robust strategies to defend deep networks from such attacks. In this work we propose to establish axioms of robustness against adversarial attacks in time series classification. We subsequently design a suitable experimental methodology and empirically validate the hypotheses put forth. Results obtained from our investigations confirm the proposed hypotheses, and provide a strong empirical baseline with a view to mitigating the effects of adversarial attacks in deep time series classification.
Mubarak G. Abdu-Aguye, Walid Gomaa 0001, Yasushi Makihara, Yasushi Yagi
Pattern Recognit. Lett.4
2021 Estimation of Gait Relative Attribute Distributions using a Differentiable Trade-off Model of Optimal and Uniform Transports
abstract
This paper describes a method for estimating gait relative attribute distributions. Existing datasets for gait relative attributes have only three-grade annotations, which cannot be represented in the form of distributions. Thus, we first create a dataset with seven-grade annotations for five gait relative attributes (i.e., beautiful, graceful, cheerful, imposing, and relaxed). Second, we design a deep neural network to handle gait relative attribute distributions. Although the ground-truth (i.e., annotation) is given in a relative (or pairwise) manner with some degree of uncertainty (i.e., inconsistency among multiple annotators), it is desirable for the system to output an absolute attribute distribution for each gait input. Therefore, we develop a model that converts a pair of absolute attribute distributions into a relative attribute distribution. More specifically, we formulate the conversion as a transportation process from one absolute attribute distribution to the other, then derive a differentiable model that determines the trade-off between optimal transport and uniform transport. Finally, we learn the network parameters by minimizing the dissimilarity between the estimated and ground-truth distributions through the Kullback–Leibler divergence and the expectation dissimilarity. Experimental results show that the proposed method successfully estimates both absolute and relative attribute distributions.
Yasushi Makihara, Yuta Hayashi, Allam Shehata, Daigo Muramatsu, Yasushi Yagi
IJCB5
2021 Real-Time Gait-Based Age Estimation and Gender Classification from a Single Image
abstract
In this paper, we propose a unified real-time framework for gait-based age estimation and gender classification that uses just a single image, which reduces the latency in video capturing compared with the existing methods based on a gait cycle. To cope with the problem of lacking motion information in the input single image, we first reconstruct a gait cycle of a silhouette sequence from the input image via a gait cycle reconstruction network. The reconstructed gait cycle is then fed into a state-of-the-art gait recognition network for feature representation learning, which is further used to obtain the class of the gender and the estimated probability distribution of integer age labels. Unlike the existing methods focusing on the gait sequences captured from the side view, the proposed method is applicable to the gait images from an arbitrary view with a single trained model, which is more suitable for real-world application scenarios (e.g., automatic access control). Stand-alone and client-server online systems were implemented based on the proposed method, which validates the real-time/online property in actual scenes. The experiments on the world's largest multi-view gait dataset demonstrate the effectiveness of the proposed method, which achieves performance improvement compared with the benchmark algorithms.
Chi Xu 0003, Yasushi Makihara, Ruochen Liao, Hirotaka Niitsuma, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003
WACV6
2021 Action recognition using kinematics posture feature on 3D skeleton joint locations
abstract
Action recognition is a very widely explored research area in computer vision and related fields. We propose Kinematics Posture Feature (KPF) extraction from 3D joint positions based on skeleton data for improving the performance of action recognition. In this approach, we consider the skeleton 3D joints as kinematics sensors. We propose Linear Joint Position Feature (LJPF) and Angular Joint Position Feature (AJPF) based on 3D linear joint positions and angles between bone segments. We then combine these two kinematics features for each video frame for each action to create the KPF feature sets. These feature sets encode the variation of motion in the temporal domain as if each body joint represents kinematics position and orientation sensors. In the next stage, we process the extracted KPF feature descriptor by using a low pass filter, and segment them by using sliding windows with optimized length. This concept resembles the approach of processing kinematics sensor data. From the segmented windows, we compute the Position-based Statistical Feature (PSF). These features consist of temporal domain statistical features (e.g., mean, standard deviation, variance, etc.). These statistical features encode the variation of postures (i.e., joint positions and angles) across the video frames. For performing classification, we explore Support Vector Machine (Linear), RNN, CNNRNN, and ConvRNN model. The proposed PSF feature sets demonstrate prominent performance in both statistical machine learning- and deep learning-based models. For evaluation, we explore five benchmark datasets namely UTKinect-Action3D, Kinect Activity Recognition Dataset (KARD), MSR 3D Action Pairs, Florence 3D, and Office Activity Dataset (OAD). To prevent overfitting, we consider the leave-one-subject-out framework as the experimental setup and perform 10-fold cross-validation. Our approach outperforms several existing methods in these benchmark datasets and achieves very promising classification performance.
Md. Atiqur Rahman Ahad, Masud Ahmed, Anindya Das Antar, Yasushi Makihara, Yasushi Yagi
Pattern Recognit. Lett.5
2021 Cross-View Gait Recognition Using Pairwise Spatial Transformer Networks
abstract
In this paper, we propose a pairwise spatial transformer network (PSTN) for cross-view gait recognition, which reduces unwanted feature mis-alignment due to view differences before a recognition step for better performance. The proposed PSTN is a unified CNN architecture that consists of a pairwise spatial transformer (PST) and subsequent recognition network (RN). More specifically, given a matching pair of gait features from different source and target views, the PST estimates a non-rigid deformation field to register the features in the matching pair into their intermediate view, which mitigates distortion by registration compared with the case of direct deformation from the source view to target view. The registered matching pair is then fed into the RN to output a dissimilarity score. Although registration may reduce not only intra-subject variations but also inter-subject variations, we can still achieve a good trade-off between them using a loss function designed to optimize recognition accuracy. Experiments on three publicly available gait datasets demonstrate that the proposed method yields superior performance for both verification and identification scenarios by combining any gait recognition network benchmarks with the PST.
Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003
IEEE Trans. Circuits Syst. Video Technol.4
2020 Descriptor-Free Multi-view Region Matching for Instance-Wise 3D Reconstruction
Takuma Doi, Fumio Okura, Toshiki Nagahara, Yasuyuki Matsushita, Yasushi Yagi
ACCV (5)5
2020 End-to-End Model-Based Gait Recognition
Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Yasushi Yagi, Shiqi Yu 0001, Mingwu Ren
ACCV (3)4
2020 Gait Recognition via Semi-supervised Disentangled Representation Learning to Identity and Covariate Features
abstract
Existing gait recognition approaches typically focus on learning identity features that are invariant to covariates (e.g., the carrying status, clothing, walking speed, and viewing angle) and seldom involve learning features from the covariate aspect, which may lead to failure modes when variations due to the covariate overwhelm those due to the identity. We therefore propose a method of gait recognition via disentangled representation learning that considers both identity and covariate features. Specifically, we first encode an input gait template to get the disentangled identity and covariate features, and then decode the features to simultaneously reconstruct the input gait template and the canonical version of the same subject with no covariates in a semi-supervised manner to ensure successful disentanglement. We finally feed the disentangled identity features into a contrastive/triplet loss function for a verification/identification task. Moreover, we find that new gait templates can be synthesized by transferring the covariate feature from one subject to another. Experimental results on three publicly available gait data sets demonstrate the effectiveness of the proposed method compared with other state-of-the-art methods.
Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Yasushi Yagi, Mingwu Ren
CVPR4
2020 Gait Recognition from a Single Image Using a Phase-Aware Gait Cycle Reconstruction Network
Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003
ECCV (19)4
2020 Detecting Adversarial Attacks In Time-Series Data
abstract
In recent times, deep neural networks have seen increased adoption in highly critical tasks. They are also susceptible to adversarial attacks, which are specifically crafted changes made to input samples which lead to erroneous output from such models. Such attacks have been shown to affect different types of data such as images and more recently, time-series data. Such susceptibility could have catastrophic consequences, depending on the domain.We propose a method for detecting Fast Gradient Sign Method (FGSM) and Basic Iterative Method (BIM) adversarial attacks as adapted for time-series data. We frame the problem as an instance of outlier detection and construct a normalcy model based on information and chaos-theoretic measures, which can then be used to determine whether unseen samples are normal or adversarial. Our approach shows promising performance on several datasets from the 2015 UCR Time Series Archive, reaching up to 97% detection accuracy in the best case.
Mubarak G. Abdu-Aguye, Walid Gomaa 0001, Yasushi Makihara, Yasushi Yagi
ICASSP4
2020 How Confident Are You in Your Estimate of a Human Age? Uncertainty-aware Gait-based Age Estimation by Label Distribution Learning
abstract
Gait-based age estimation is one of key techniques for many applications (e.g., finding lost children/aged wanders). It is well known that the age estimation uncertainty is highly dependent on ages (i.e., it is generally small for children while is large for adults/the elderly), and it is important to know the uncertainty for the above-mentioned applications. We therefore propose a method of uncertainty-aware gait-based age estimation by introducing a label distribution learning framework. More specifically, we design a network which takes an appearance-based gait feature as an input and outputs discrete label distributions in the integer age domain. Experiments with the world-largest gait database OULP-Age show that the proposed method can successfully represent the uncertainty of age estimation and also outperforms or is comparable to the state-of-the-art methods.
Atsuya Sakata, Yasushi Makihara, Noriko Takemura, Daigo Muramatsu, Yasushi Yagi
IJCB5
2020 DeformGait: Gait Recognition under Posture Changes using Deformation Patterns between Gait Feature Pairs
abstract
In this paper, we propose a unified convolutional neural network (CNN) framework for robust gait recognition against posture changes (e.g., those induced by walking speed changes). In order to mitigate the posture changes, we first register an input matching pair of gait features with different postures by a deformable registration network, which estimates a deformation field to transform the input pair both into their intermediate posture. The pair of the registered features is then fed into a recognition network. Furthermore, ways of the deformation (i.e., deformation patterns) can differ between the same subject pairs (e.g., only posture deformation) and different subject pairs (e.g., not only posture deformation but also body shape deformation), which implies the deformation pattern can be another cue to distinguish the same subject pairs from the different subject pairs. We therefore introduce another recognition network whose input is the deformation pattern. Finally, the deformable registration network, and the two recognition networks for the registered features and the deformation patterns, constitute the whole framework, named DeformGait, and they are trained in an end-to-end manner by minimizing a loss function which is appropriately designed for each of verification and identification scenario. Experiments on the publicly available dataset containing the largest speed variations demonstrate that the proposed method achieves the state-of-the-art performance in both identification and verification scenarios.
Chi Xu 0003, Daisuke Adachi, Yasushi Makihara, Yasushi Yagi, Jianfeng Lu 0003
IJCB4
2020 Action Recognition from a Single Coded Image
abstract
Cameras are prevalent in society at the present time, for example, surveillance cameras, and smartphones equipped with cameras and smart speakers. There is an increasing demand to analyze human actions from these cameras to detect unusual behavior or within a man-machine interface for Internet of Things (IoT) devices. For a camera, there is a trade-off between spatial resolution and frame rate. A feasible approach to overcome this trade-off is compressive video sensing. Compressive video sensing uses random coded exposure and reconstructs higher than read out of sensor frame rate video from a single coded image. It is possible to recognize an action in a scene from a single coded image because the image contains multiple temporal information for reconstructing a video. In this paper, we propose reconstruction-free action recognition from a single coded exposure image. We also proposed deep sensing framework which models camera sensing and classification models into convolutional neural network (CNN) and jointly optimize the coded exposure and classification model simultaneously. We demonstrated that the proposed method can recognize human actions from only a single coded image. We also compared it with competitive inputs, such as low-resolution video with a high frame rate and high-resolution video with a single frame in simulation and real experiments.
Tadashi Okawara, Michitaka Yoshida, Hajime Nagahara, Yasushi Yagi
ICCP4
2020 Deep Gait Relative Attribute using a Signed Quadratic Contrastive Loss
abstract
This paper presents a deep learning-based method to estimate gait attributes (e.g., stately, cool, relax, etc.). Similarly to the existing studies on relative attribute, human perception-based annotations on the gait attributes are given to pairs of gait videos (i.e., the first one is better, tie, and the second one is better), and the relative annotations are utilized to train a ranking model of the gait attribute. More specifically, we design a Siamese (i.e., two-stream) network which takes a pair of gait inputs and output gait attribute score for each. We then introduce a suitable loss function called a signed contrastive loss to train the network parameters with the relative annotation. Unlike the existing loss functions for learning to rank does not inherit a nice property of a quadratic contrastive loss, the proposed signed quadratic contrastive loss function inherits the nice property. The quantitative evaluation results reveal that the proposed method shows better or comparable accuracies of relative attribute prediction against the baseline methods.
Yuta Hayashi, Allam Shehata, Yasushi Makihara, Daigo Muramatsu, Yasushi Yagi
ICPR5
2020 Adaptive Pooling Is All You Need: An Empirical Study on Hyperparameter-insensitive Human Action Recognition Using Wearable Sensors
abstract
A plethora of techniques have been proposed in human action recognition fields, and particularly deep learning-based methods such as convolutional neural networks (CNNs) have achieved impressive results. Usually, there is need to tune hyper-parameters in the deep neural network (e.g., filter size, stride) to achieve reasonable results. Such hyper-parameter tuning is, however, extremely time and resource-intensive even for small models. In this paper, we posit that the inclusion of an adaptive pooling in CNNs used for human action recognition largely eliminates the need for hyper-parameter tuning. Specifically, we demonstrated our idea for human action recognition using inertial sensor data (i.e., a temporal sequence) with a one-dimensional adaptive pooling. We compared the adaptive pooling to conventional CNNs with randomly chosen hyper-parameters using a publicly available data set for human action recognition. Experimental results showed that the adaptive pooling achieved better accuracy than the conventional CNNs.
Mubarak G. Abdu-Aguye, Walid Gomaa 0001, Yasushi Makihara, Yasushi Yagi
IJCNN4
2020 Identifying motion pathways in highly crowded scenes: A non-parametric tracklet clustering approach
Allam S. Hassanein, Mohamed E. Hussein 0001, Walid Gomaa 0001, Yasushi Makihara, Yasushi Yagi
Comput. Vis. Image Underst.5
2020 Gait recognition invariant to carried objects using alpha blending generative adversarial networks
abstract
Gait recognition invariant to carried objects (COs) is very difficult in a real-life scene because the COs can have various shapes and sizes, in addition to unpredictable carrying locations (e.g., front, back, and side, or multiple locations). Therefore, in this paper, we propose a robust method for gait recognition against various COs by reconstructing a gait template without COs. A straightforward approach is to directly generate a gait template without COs given a gait template with COs as the input using a conventional generative adversarial network. There is, however, a potential risk of unnecessarily altering parts that were originally unaffected by COs (e.g., leg parts for a person carrying a backpack). Because we do not want to touch such unaffected parts in the original template, we first estimate a gait template without COs, and then blend it with the original template by an estimated alpha matte that indicates the blending parameters. We then create an alpha-blended template from the original template and the generated template without COs based on the estimated alpha matte. We use two independent generators to estimate the alpha matte and the generated template without COs. Finally, we feed the alpha-blended gait template into a state-of-the-art discrimination network for gait recognition. The experimental results on three publicly available gait databases with real-life COs demonstrate the state-of-the-art performance of the proposed method.
Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Yasushi Yagi, Mingwu Ren
Pattern Recognit.4
2019 On the Feasibility of On-body Roaming Models in Human Activity Recognition
Mubarak G. Abdu-Aguye, Walid Gomaa 0001, Yasushi Makihara, Yasushi Yagi
ICINCO (1)4
2019 Reflectance and Shape Estimation with a Light Field Camera Under Natural Illumination
Trung Ngo Thanh, Hajime Nagahara, Ko Nishino, Rin-Ichiro Taniguchi, Yasushi Yagi
Int. J. Comput. Vis.5
2019 Speed-Invariant Gait Recognition Using Single-Support Gait Energy Image
abstract
Gait is one of the most popular behavioral biometrics because it can be authenticated at a distance from a camera without subject cooperation. Speed differences between matching pairs, however, cause significant performance drops in gait recognition, and gait mode difference (i.e., walking versus running) makes gait recognition further challenging. We therefore propose a speed-invariant gait representation called single-support GEI (SSGEI), which realizes a good trade-off between speed invariance and stability by aggregating multiple frames around single-support phases. In addition, to mitigate the pose differences between walking and running modes at single-support phases, we morph walking and running SSGEIs into intermediate SSGEIs between walking and running mode, where we exploit a free-form deformation field from the walking or running modes to the intermediate mode obtained by training data. We finally apply Gabor filtering and spatial metric learning as postprocessing for further accuracy improvement. Experiments on two publicly available datasets, the OU-ISIR Treadmill Dataset A and the CASIA-C Dataset demonstrate that the proposed method yields the state-of-the-art accuracies in both identification and verification scenarios with a low computational cost.
Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003
Multim. Tools Appl.4
2019 Gait-based age progression/regression: a baseline and performance evaluation by age group classification and cross-age gait identification
abstract
Gait is believed to be an advanced behavioral biometric that can be perceived at a large distance from a camera without subject cooperation and hence is favorable for many applications in surveillance and forensics. However, appearance differences caused by human aging may significantly reduce the performance of gait recognition. Modeling the aging process on gait features is one of the possible solutions to this problem, and it may inspire more potential applications, such as finding lost children and examining health status. To the best of our knowledge, this topic has not been studied in the literature. Motivated by the fact that aging effects are mainly reflected in the shape and appearance deformations of the gait feature, we propose a baseline algorithm for gait-based age progression and regression using a generic geometric transformation between different age groups, in conjunction with the gait energy image, which is an appearance-based gait feature frequently used in the gait analysis community, to render gait aging and reverse aging effects simultaneously. Various evaluations were conducted through gait-based age group classification and cross-age gait identification to validate the performance of the proposed method, in addition to providing several insights for future research on the subject.
Chi Xu 0003, Yasushi Makihara, Yasushi Yagi, Jianfeng Lu 0003
Mach. Vis. Appl.3
2019 Material Classification from Time-of-Flight Distortions
abstract
This paper presents a material classification method using an off-the-shelf Time-of-Flight (ToF) camera. The proposed method is built upon a key observation that the depth measurement by a ToF camera is distorted for objects with certain materials, especially with translucent materials. We show that this distortion is due to the variation of time domain impulse responses across materials and also due to the measurement mechanism of the ToF cameras. Specifically, we reveal that the amount of distortion varies according to the modulation frequency of the ToF camera, the object material, and the distance between the camera and object. Our method uses the depth distortion of ToF measurements as a feature for classification and achieves material classification of a scene. Effectiveness of the proposed method is demonstrated by numerical evaluations and real-world experiments, showing its capability of material classification, even for visually indistinguishable objects.
Kenichiro Tanaka, Yasuhiro Mukaigawa, Takuya Funatomi, Hiroyuki Kubo, Yasuyuki Matsushita, Yasushi Yagi
IEEE Trans. Pattern Anal. Mach. Intell.6
2019 On Input/Output Architectures for Convolutional Neural Network-Based Cross-View Gait Recognition
abstract
In this paper, we discuss input/output architectures for convolutional neural network (CNN)-based cross-view gait recognition. For this purpose, we consider two aspects: verification versus identification and the tradeoff between spatial displacements caused by subject difference and view difference. More specifically, we use the Siamese network with a pair of inputs and contrastive loss for verification and a triplet network with a triplet of inputs and triplet ranking loss for identification. The aforementioned CNN architectures are insensitive to spatial displacement, because the difference between a matching pair is calculated at the last layer after passing through the convolution and max pooling layers; hence, they are expected to work relatively well under large view differences. By contrast, because it is better to use the spatial displacement to its best advantage because of the subject difference under small view differences, we also use CNN architectures where the difference between a matching pair is calculated at the input level to make them more sensitive to spatial displacement. We conducted experiments for cross-view gait recognition and confirmed that the proposed architectures outperformed the state-of-the-art benchmarks in accordance with their suitable situations of verification/identification tasks and view differences.
Noriko Takemura, Yasushi Makihara, Daigo Muramatsu, Tomio Echigo, Yasushi Yagi
IEEE Trans. Circuits Syst. Video Technol.5
2019 Joint Intensity Transformer Network for Gait Recognition Robust Against Clothing and Carrying Status
abstract
Clothing and carrying status variations are the two key factors that affect the performance of gait recognition because people usually wear various clothes and carry all kinds of objects, while walking in their daily life. These covariates substantially affect the intensities within conventional gait representations such as gait energy images. Hence, to properly compare a pair of input gait features, an appropriate metric for joint intensity is needed in addition to the conventional spatial metric. We therefore propose a unified joint intensity transformer network for gait recognition that is robust against various clothing and carrying statuses. Specifically, the joint intensity transformer network is a unified deep learning-based architecture containing three parts: a joint intensity metric estimation net, a joint intensity transformer, and a discrimination network. First, the joint intensity metric estimation net uses a well-designed encoder-decoder network to estimate a sample-dependent joint intensity metric for a pair of input gait energy images. Subsequently, a joint intensity transformer module outputs the spatial dissimilarity of two gait energy images using the metric learned by the joint intensity metric estimation net. Third, the discrimination network is a generic convolution neural network for gait recognition. In addition, the joint intensity transformer network is designed with different loss functions depending on the gait recognition task (i.e., a contrastive loss function for the verification task and a triplet loss function for the identification task). The experiments on the world's largest datasets containing various clothing and carrying statuses demonstrate the state-of-the-art performance of the proposed method.
Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Yasushi Yagi, Mingwu Ren
IEEE Trans. Inf. Forensics Secur.4
2018 Probabilistic Plant Modeling via Multi-View Image-to-Image Translation
abstract
This paper describes a method for inferring three-dimensional (3D) plant branch structures that are hidden under leaves from multi-view observations. Unlike previous geometric approaches that heavily rely on the visibility of the branches or use parametric branching models, our method makes statistical inferences of branch structures in a probabilistic framework. By inferring the probability of branch existence using a Bayesian extension of image-to-image translation applied to each of multi-view images, our method generates a probabilistic plant 3D model, which represents the 3D branching pattern that cannot be directly observed. Experiments demonstrate the usefulness of the proposed approach in generating convincing branch structures in comparison to prior approaches.
Takahiro Isokane, Fumio Okura, Ayaka Ide, Yasuyuki Matsushita, Yasushi Yagi
CVPR5
2018 Depth error correction for projector-camera based consumer depth cameras
abstract
This paper proposes a depth measurement error model for consumer depth cameras such as the Microsoft Kinect, and a corresponding calibration method. These devices were originally designed as video game interfaces, and their output depth maps usually lack sufficient accuracy for 3D measurement. Models have been proposed to reduce these depth errors, but they only consider camera-related causes. Since the depth sensors are based on projector-camera systems, we should also consider projector-related causes. Also, previous models require disparity observations, which are usually not output by such sensors, so cannot be employed in practice. We give an alternative error model for projector-camera based consumer depth cameras, based on their depth measurement algorithm, and intrinsic parameters of the camera and the projector; it does not need disparity values. We also give a corresponding new parameter estimation method which simply needs observation of a planar board. Our calibrated error model allows use of a consumer depth sensor as a 3D measuring device. Experimental results show the validity and effectiveness of the error model and calibration procedure.
Hirotake Yamazoe, Hitoshi Habe, Ikuhisa Mitsugami, Yasushi Yagi
Comput. Vis. Media4
2018 Gait-based human age estimation using age group-dependent manifold learning and regression
abstract
Human age estimation from gait is expected to be an important technology for a variety of applications such as automatic customer counting for marketing research or automatic age-based access control restriction for a specific area because the gait can be observable at a distance from a camera (e.g., CCTV). Although the aging process of gait significantly differs among age groups (e.g., children, adults, and the elderly), previous studies on gait-based human age estimation employ a single age group-independent estimation model that suffers from large estimation errors when the age variation increases. We therefore propose an age group-dependent gait-based human age estimation method for better accuracy. Specifically, in the training phase, we first compose age groups that are well-separated from each other by clustering gait features along with their age labels. We then learn a classifier that classifies the gait features for multiple age groups using a directed acyclic graph support vector machine. Next, we learn an age regression model for each age group using support vector regression with a Gaussian kernel in conjunction with a manifold learning technique, i.e., orthogonal locality preserving projection, to better characterize the gait feature. In the test phase, given a gait feature, it is first classified into an age group and then its age is estimated with the age regression model of the classified age group. Experimental results on a gait database that has the world’s largest population of participants ranging from 2 to 90 years old demonstrate the state-of-the-art performance of the proposed method.
Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Yasushi Yagi, Mingwu Ren
Multim. Tools Appl.4
2018 3D level set method for blastomere segmentation of preimplantation embryos in fluorescence microscopy images
Andrey Grushnikov, Ritsuya Niwayama, Takeo Kanade, Yasushi Yagi
Mach. Vis. Appl.4
2018 Gait Recognition Based on Normal Distance Maps
abstract
Gait is a commonly used biometric for human recognition. Its main advantage relies on its ability to identify people at distances at which other biometrics fail. In this paper, we develop a new approach for gait recognition that combines the distance transform with curvatures of local contours. We call our gait feature template the normal distance map. Our method encodes both body shapes and boundary curvatures into a novel feature descriptor that is more robust than existing gait representations. We evaluate our approach on the widely used and challenging USF and CASIA-B datasets. Furthermore, we evaluate it on the OU-ISIR gait dataset, the largest one available in the literature, to obtain statistically reliable results. We verify our approach is significantly superior to the current state-of-the-art under most conditions.
Hazem El-Alfy, Ikuhisa Mitsugami, Yasushi Yagi
IEEE Trans. Cybern.3
2017 Reflectance and Shape Estimation with a Light Field Camera under Natural Illumination
Trung Ngo Thanh, Hajime Nagahara, Ko Nishino, Rin-Ichiro Taniguchi, Yasushi Yagi
BMVC5
2017 Joint Intensity and Spatial Metric Learning for Robust Gait Recognition
abstract
This paper describes a joint intensity metric learning method to improve the robustness of gait recognition with silhouette-based descriptors such as gait energy images. Because existing methods often use the difference of image intensities between a matching pair (e.g., the absolute difference of gait energies for the l1-norm) to measure a dissimilarity, large intrasubject differences derived from covariate conditions (e.g., large gait energies caused by carried objects vs. small gait energies caused by the background), may wash out subtle intersubject differences (e.g., the difference of middle-level gait energies derived from motion differences). We therefore introduce a metric on joint intensity to mitigate the large intrasubject differences as well as leverage the subtle intersubject differences. More specifically, we formulate the joint intensity and spatial metric learning in a unified framework and alternately optimize it by linear or ranking support vector machines. Experiments using the OU-ISIR treadmill data set B with the largest clothing variation and large population data set with bag, β version containing carrying status in the wild demonstrate the effectiveness of the proposed method.
Yasushi Makihara, Atsuyuki Suzuki, Daigo Muramatsu, Xiang Li 0028, Yasushi Yagi
CVPR5
2017 Material Classification Using Frequency-and Depth-Dependent Time-of-Flight Distortion
abstract
This paper presents a material classification method using an off-the-shelf Time-of-Flight (ToF) camera. We use a key observation that the depth measurement by a ToF camera is distorted in objects with certain materials, especially with translucent materials. We show that this distortion is caused by the variations of time domain impulse responses across materials and also by the measurement mechanism of the existing ToF cameras. Specifically, we reveal that the amount of distortion varies according to the modulation frequency of the ToF camera, the material of the object, and the distance between the camera and object. Our method uses the depth distortion of ToF measurements as features and achieves material classification of a scene. Effectiveness of the proposed method is demonstrated by numerical evaluation and real-world experiments, showing its capability of even classifying visually similar objects.
Kenichiro Tanaka, Yasuhiro Mukaigawa, Takuya Funatomi, Hiroyuki Kubo, Yasuyuki Matsushita, Yasushi Yagi
CVPR6
2017 Recovering Inner Slices of Layered Translucent Objects by Multi-Frequency Illumination
abstract
This paper describes a method for recovering appearance of inner slices of translucent objects. The appearance of a layered translucent object is the summed appearance of all layers, where each layer is blurred by a depth-dependent point spread function (PSF). By exploiting the difference of low-pass characteristics of depth-dependent PSFs, we develop a multi-frequency illumination method for obtaining the appearance of individual inner slices. Specifically, by observing the target object with varying the spatial frequency of checker-pattern illumination, our method recovers the appearance of inner slices via computation. We study the effect of non-uniform transmission due to inhomogeneity of translucent objects and develop a method for recovering clear inner slices based on the pixel-wise PSF estimates under the assumption of spatial smoothness of inner slice appearances. We quantitatively evaluate the accuracy of the proposed method by simulations and qualitatively show faithful recovery using real-world scenes.
Kenichiro Tanaka, Yasuhiro Mukaigawa, Hiroyuki Kubo, Yasuyuki Matsushita, Yasushi Yagi
IEEE Trans. Pattern Anal. Mach. Intell.5
2016 Gait Energy Response Function for Clothing-Invariant Gait Recognition
Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Daigo Muramatsu, Yasushi Yagi, Mingwu Ren
ACCV (2)5
2016 Speed Invariance vs. Stability: Cross-Speed Gait Recognition Using Single-Support Gait Energy Image
Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003
ACCV (2)4
2016 Recovering Transparent Shape from Time-of-Flight Distortion
abstract
This paper presents a method for recovering shape and normal of a transparent object from a single viewpoint using a Time-of-Flight (ToF) camera. Our method is built upon the fact that the speed of light varies with the refractive index of the medium and therefore the depth measurement of a transparent object with a ToF camera may be distorted. We show that, from this ToF distortion, the refractive light path can be uniquely determined by estimating a single parameter. We estimate this parameter by introducing a surface normal consistency between the one determined by a light path candidate and the other computed from the corresponding shape. The proposed method is evaluated by both simulation and real-world experiments and shows faithful transparent shape recovery.
Kenichiro Tanaka, Yasuhiro Mukaigawa, Hiroyuki Kubo, Yasuyuki Matsushita, Yasushi Yagi
CVPR5
2016 View Transformation Model Incorporating Quality Measures for Cross-View Gait Recognition
abstract
Cross-view gait recognition authenticates a person using a pair of gait image sequences with different observation views. View difference causes degradation of gait recognition accuracy, and so several solutions have been proposed to suppress this degradation. One useful solution is to apply a view transformation model (VTM) that encodes a joint subspace of multiview gait features trained with auxiliary data from multiple training subjects, who are different from test subjects (recognition targets). In the VTM framework, a gait feature with a destination view is generated from that with a source view by estimating a vector on the trained joint subspace, and gait features with the same destination view are compared for recognition. Although this framework improves recognition accuracy as a whole, the fit of the VTM depends on a given gait feature pair, and causes an inhomogeneously biased dissimilarity score. Because it is well known that normalization of such inhomogeneously biased scores improves recognition accuracy in general, we therefore propose a VTM incorporating a score normalization framework with quality measures that encode the degree of the bias. From a pair of gait features, we calculate two quality measures, and use them to calculate the posterior probability that both gait features originate from the same subjects together with the biased dissimilarity score. The proposed method was evaluated against two gait datasets, a large population gait dataset of over-ground walking (course dataset) and a treadmill gait dataset. The experimental results show that incorporating the quality measures contributes to accuracy improvement in many cross-view settings.
Daigo Muramatsu, Yasushi Makihara, Yasushi Yagi
IEEE Trans. Cybern.3
2015 Recovering inner slices of translucent objects by multi-frequency illumination
abstract
This paper describes a method for recovering appearance of inner slices of translucent objects. The outer appearance of translucent objects is a summation of the appearance of slices at all depths, where each slice is blurred by depth-dependent point spread functions (PSFs). By exploiting the difference of low-pass characteristics of depth-dependent PSFs, we develop a multi-frequency illumination method for obtaining the appearance of individual inner slices using a coaxial projector-camera setup. Specifically, by measuring the target object with varying the spatial frequency of checker patterns emitted from a projector, our method recovers inner slices via a simple linear solution method. We quantitatively evaluate accuracy of the proposed method by simulations and show qualitative recovery results using real-world scenes.
Kenichiro Tanaka, Yasuhiro Mukaigawa, Hiroyuki Kubo, Yasuyuki Matsushita, Yasushi Yagi
CVPR5
2015 Effective part-based gait identification using frequency-domain gait entropy features
Md. Rokanujjaman, Md. Altab Hossain, Yasushi Makihara, Yasushi Yagi
Multim. Tools Appl.6
2015 Onboard monocular pedestrian detection by combining spatio-temporal hog with structure from motion algorithm
Chunsheng Hua, Yasushi Makihara, Yasushi Yagi, Shun Iwasaki, Keisuke Miyagawa
Mach. Vis. Appl.3
2015 Similar gait action recognition using an inertial sensor
Trung Ngo Thanh, Yasushi Makihara, Hajime Nagahara, Yasuhiro Mukaigawa, Yasushi Yagi
Pattern Recognit.5
2015 Gait-Based Person Recognition Using Arbitrary View Transformation Model
abstract
Gait recognition is a useful biometric trait for person authentication because it is usable even with low image resolution. One challenge is robustness to a view change (cross-view matching); view transformation models (VTMs) have been proposed to solve this. The VTMs work well if the target views are the same as their discrete training views. However, the gait traits are observed from an arbitrary view in a real situation. Thus, the target views may not coincide with discrete training views, resulting in recognition accuracy degradation. We propose an arbitrary VTM (AVTM) that accurately matches a pair of gait traits from an arbitrary view. To realize an AVTM, we first construct 3D gait volume sequences of training subjects, disjoint from the test subjects in the target scene. We then generate 2D gait silhouette sequences of the training subjects by projecting the 3D gait volume sequences onto the same views as the target views, and train the AVTM with gait features extracted from the 2D sequences. In addition, we extend our AVTM by incorporating a part-dependent view selection scheme (AVTM_PdVS), which divides the gait feature into several parts, and sets part-dependent destination views for transformation. Because appropriate destination views may differ for different body parts, the part-dependent destination view selection can suppress transformation errors, leading to increased recognition accuracy. Experiments using data sets collected in different settings show that the AVTM improves the accuracy of cross-view matching and that the AVTM_PdVS further improves the accuracy in many cases, in particular, verification scenarios.
Daigo Muramatsu, Akira Shiraishi, Yasushi Makihara, Md. Zasim Uddin, Yasushi Yagi
IEEE Trans. Image Process.5
2014 Estimating Depth of Layered Structure Based on Multispectral Speckle Correlation
abstract
Some objects have layered structures in which a dynamic region is covered by a static layer. In this paper, we propose a new experiment for estimating the depth of the dynamic region using speckle analysis. The speckle is caused by the mutual interference of a coherence laser. We use two characteristics of the speckle. One is the temporal stability of the speckle pattern and the other is the wavelength dependency of the transmittance of the laser. We estimate the depth by computing correlations of speckle patterns using multispectral lasers. Experimental results using a simulated skin show that multispectral speckle correlation can be used for analyzing a layered structure.
Takahiro Matsumura, Yasuhiro Mukaigawa, Yasushi Yagi
3DV3
2014 Gait Recognition under Speed Transition
abstract
This paper describes a method of gait recognition from image sequences wherein a subject is accelerating or decelerating. As a speed change occurs due to a change of pitch (the first-order derivative of a phase, namely, a gait stance) and/or stride, we model this speed change using a cylindrical manifold whose azimuth and height corresponds to the phase and the stride, respectively. A radial basis function (RBF) interpolation framework is used to learn subject specific mapping matrices for mapping from manifold to image space. Given an input image sequence of speed transited gait of a test subject, we estimate the mapping matrix of the test subject as well as the phase and stride sequence using an energy minimization framework considering the following three points: (1) fitness of the synthesized images to the input image sequence as well as to an eigenspace constructed by exemplars of training subjects, (2) smoothness of the phase and the stride sequence, and (3) pitch and stride fitness to the pitch-stride preference model. Using the estimated mapping matrix, we synthesize a constant-speed gait image sequence, and extract a conventional period-based gait feature from it for matching. We conducted experiments using real speed transited gait image sequences with 179 subjects and demonstrated the effectiveness of the proposed method.
Al Mansur, Yasushi Makihara, Muhammad Rasyid Aqmar, Yasushi Yagi
CVPR4
2014 Surface Normal Deconvolution: Photometric Stereo for Optically Thick Translucent Objects
Chika Inoshita, Yasuhiro Mukaigawa, Yasuyuki Matsushita, Yasushi Yagi
ECCV (2)4
2014 Score-level fusion by generalized Delaunay triangulation
abstract
This paper describes a method for score-level fusion in multi-cue two-class classification problems. Fusion based on the probability density function (PDF) of multiple scores given for each class is a promising approach because it guarantees optimality as long as the estimated PDFs are correct. Instead of lattice-type control points used in previous non-parametric density-based approaches, floating control points (FCPs) are introduced to improve scalability and the whole posterior distribution is represented by interpolation or extrapolation using generalized Delaunay triangulation. Given a set of FCPs obtained by k-means, posteriors on the FCPs are estimated by an energy minimization framework using training samples. The experiments, using both simulation data as well as several types of real data from three publicly available score databases for multi-cue biometric authentication, demonstrate the effectiveness of the proposed method.
Yasushi Makihara, Daigo Muramatsu, Haruyuki Iwama, Trung Ngo Thanh, Yasushi Yagi, Md. Altab Hossain
IJCB5
2014 Cross-view gait recognition using view-dependent discriminative analysis
abstract
Gait is a unique and promising behavioral biometrics which allows to authenticate a person even at a distance from the camera. Since a matching pair of gait features are often drawn from different views due to differences in camera position/attitude and walking directions in the real world, it is important to cope with cross-view gait recognition. In this paper, we propose a discriminative approach to cross-view gait recognition using view-dependent projection matrices, unlike the existing discriminant approaches which utilize only a single common projection matrix for different views. We demonstrated the effectiveness of the proposed method through cross-view gait recognition experiments with two publicly available gait datasets. In addition, since the success of the discriminant analysis relies on the training sample size, we show the effect of transfer learning across two gait datasets as well as provide the rigorous sensitivity analysis of the proposed method against the number of training subjects ranging from 10 to approximately 1,000 subjects.
Al Mansur, Yasushi Makihara, Daigo Muramatsu, Yasushi Yagi
IJCB4
2014 Light Transport Refocusing for Unknown Scattering Medium
abstract
In this paper we propose a new light transport refocusing method for depth estimation as well as for investigation inside scattering media with unknown scattering properties. Propagated visible light rays through scattering media are utilized in our proposed refocusing method. We use 2D light source to illuminate the scattering media and 2D image sensor for capturing transported rays. The proposed method that uses 4D light transport can clearly visualize shallow depth, as well as deep depth plane of the medium. We apply our light transport refocusing method for depth estimation using conventional depth-from-focus method and for clear visualization by descattering the light rays passing through the medium. To evaluate the effectiveness we have done experiments using acrylic and milk-water type scattering medium in various optical and geometrical conditions. Finally, we show up the results of depth estimation and clear visualization, as well as with numeric evaluation.
Md. Abdul Mannan, Seiichi Tagawa, Toru Tamaki, Hajime Nagahara, Yasuhiro Mukaigawa, Yasushi Yagi
ICPR6
2014 Segmenting Reddish Lesions in Capsule Endoscopy Images Using a Gastrointestinal Color Space
abstract
Segmenting reddish lesions in capsule endoscopy (CE) images is an initial step for further computer-assisted applications such as image enhancement, abnormal measurement/tracking, and so on. In this paper, we propose an automatic segmentation method that is successful even with CE image including unclear reddish lesions. To obtain this, the proposed method seeks good features to discriminate the reddish lesions from normal tissues. For implementations, we first extract only meaningful regions in a CE image through a pre-segmentation step. The proposed features then are extracted for the meaningful regions in stead of the whole image. We approaches segmentation task through considering a statistical operator for the extracted features, that is local mean image. Candidates of the abnormal regions are located in the local mean image with assistants of a diffusion process. Evaluations in the experiments confirm effectiveness of the proposed method with both qualitative and quantitative measurement.
Hai Vu, Tomio Echigo, Yuma Imura, Yukiko Yanagawa, Yasushi Yagi
ICPR5
2014 A Web-Based Education System for Reading Video Capsule Endoscopy
abstract
The interpretive skills of medical doctors and medical technologists who examine video capsule endoscopy (VCE) in clinical practice are usually improved through hands-on courses. Such courses require that a large volume of cases be undertaken as part of the training, and they thus consume a considerable amount of the trainees' time. This paper describes an e-learning system that reduces the training time in addition to enhancing the quality of the educational process with regard to reading VCE. To achieve this goal, we focused on organizing training courses in order to appropriate for the laborious conditions that exist when reading VCE. The designed courses help the trainees acquire knowledge of abnormal regions and become familiar with reading VCE before taking examinations under conditions similar to those actual clinical practice. The proposed training modality was developed as an e-learning application on the World Wide Web. Thus, it can be easily extended to a wide range of trainees. In the experiments, 20 participants completed the self-learning training procedures in approximately 3 hours. The proposed system is much faster than conventional hands-on courses, which require a minimum of 8 hours. Furthermore, the trainees' learning performances in the final examinations confirmed that the proposed system is particularly effective for inexperienced examining doctors.
Hai Vu, Yukiko Yanagawa, Tomio Echigo, Masatsugu Shiba, Hirotoshi Okazaki, Yasuhiro Fujiwara, Tetsuo Arakawa, Yasushi Yagi
ICPR8
2014 Gait recognition by fluctuations
Muhammad Rasyid Aqmar, Yusuke Fujihara, Yasushi Makihara, Yasushi Yagi
Comput. Vis. Image Underst.4
2014 The largest inertial sensor-based gait database and performance evaluation of gait-based personal authentication
Trung Ngo Thanh, Yasushi Makihara, Hajime Nagahara, Yasuhiro Mukaigawa, Yasushi Yagi
Pattern Recognit.5
2014 Many-to-Many Superpixel Matching for Robust Tracking
abstract
We present a robust tracking method based on many-to-many image superpixel matching (MMM). Our MMM tracker represents a target and its background using two sets of superpixels. Multiple hypotheses for superpixel matching are considered for better tracking performance. For each superpixel in an input image, k matching candidates are searched in the representative sets using approximate k -NN searching. The degree of matching is measured using foreground likelihood and matching probability assignment. The superpixel matching results are projected onto a displacement confidence map that depicts the motion probabilities of all the superpixels. During the projection, the displacements confidence of the superpixels are regularized by kernel methods. We estimate the target position by searching for the maximum probability on the displacement confidence map. The experimental results confirm that our superpixel matching achieves better performance than other trackers.
Junqiu Wang, Yasushi Yagi
IEEE Trans. Cybern.2
2013 Descattering of transmissive observation using Parallel High-Frequency Illumination
abstract
The inner structures of an object can be measured by capturing transmissive images. However, the recorded images of a translucent object tend to be unclear due to strong scattering of light inside the object. In this paper, we propose a descattering approach based on Parallel High-frequency Illumination. We show in this paper that the original high-frequency illumination method and the various extended techniques can be uniformly defined as a separation of overlapped and non-overlapped light rays. Also, we show that transmissive light rays do not overlap each other by constructing a parallel projection/measurement system for performing both illumination and observation. We have developed a measurement system that consists of a camera and projector with telecentric lenses and have evaluated descattering effects by extracting transmissive light rays.
Kenichiro Tanaka, Yasuhiro Mukaigawa, Yasuyuki Matsushita, Yasushi Yagi
ICCP4
2013 Two-Point Gait: Decoupling Gait from Body Shape
abstract
Human gait modeling (e.g., for person identification) largely relies on image-based representations that muddle gait with body shape. Silhouettes, for instance, inherently entangle body shape and gait. For gait analysis and recognition, decoupling these two factors is desirable. Most important, once decoupled, they can be combined for the task at hand, but not if left entangled in the first place. In this paper, we introduce Two-Point Gait, a gait representation that encodes the limb motions regardless of the body shape. Two-Point Gait is directly computed on the image sequence based on the two point statistics of optical flow fields. We demonstrate its use for exploring the space of human gait and gait recognition under large clothing variation. The results show that we can achieve state-of-the-art person recognition accuracy on a challenging dataset.
Stephen Lombardi, Ko Nishino, Yasushi Makihara, Yasushi Yagi
ICCV4
2013 Shape priors extraction and application for geodesic distance transforms in images and videos
Junqiu Wang, Yasushi Yagi
Pattern Recognit. Lett.2
2013 Inverse Dynamics for Action Recognition
abstract
Pose-based approaches for human action recognition are attractive owing to their accurate use of human motion information. Traditionally, such approaches used kinematic features for classification. However, in addition to having high dimensions and a small interclass variation, kinematic features do not consider the interaction of the environment on human motion. In this paper, we propose a method for action recognition using dynamic features, derived by applying inverse dynamics to a physics-based representation of the human body. The physics-based model is articulated and actuated with muscles and consists of joints with variable stiffness. Dynamic features under consideration include the torques from the knee and hip joints of both legs and, implicitly, gravity, ground reaction forces, and the pose of the remaining body parts. These features are more discriminative than kinematic features, resulting in a low-dimensional representation for human actions, which preserves much of the information of the original high-dimensional pose. This low-dimensional feature achieves good classification performance even with a relatively small training data set in a simple classification framework such as a hidden Markov model. The effectiveness of the proposed method is demonstrated through experiments on the Carnegie Mellon University motion capture data set and Osaka University Kinect action data set with various actions.
Al Mansur, Yasushi Makihara, Yasushi Yagi
IEEE Trans. Cybern.3
2012 Efficient Background Subtraction under Abrupt Illumination Variations
Junqiu Wang, Yasushi Yagi
ACCV (1)2
2012 Video from nearly still: An application to low frame-rate gait recognition
abstract
In this paper, we propose a temporal super resolution approach for quasi-periodic image sequence such as human gait. The proposed method effectively combines example-based and reconstruction-based temporal super resolution approaches. A periodic image sequence is expressed as a manifold parameterized by a phase and a standard manifold is learned from multiple high frame-rate sequences in the training stage. In the test stage, an initial phase for each frame of an input low frame-rate image sequence is estimated based on the standard manifold at first, and the manifold reconstruction and the phase estimation are then iterated to generate better high frame-rate images in the energy minimization framework that ensures the fitness to both the input images and the standard manifold. The proposed method is applied to low frame-rate gait recognition and experiments with real data of 100 subjects demonstrate a significant improvement by the proposed method, particularly for quite low frame-rate videos (e.g., 1 fps).
Naoki Akae, Al Mansur, Yasushi Makihara, Yasushi Yagi
CVPR4
2012 Shape from Single Scattering for Translucent Objects
Chika Inoshita, Yasuhiro Mukaigawa, Yasuyuki Matsushita, Yasushi Yagi
ECCV (2)4
2012 Person re-identification using view-dependent score-level fusion of gait and color features
Ryo Kawai, Yasushi Makihara, Chunsheng Hua, Haruyuki Iwama, Yasushi Yagi
ICPR5
2012 Can gait fluctuations improve gait recognition?
Yasushi Makihara, Yusuke Fujihara, Yasushi Yagi
ICPR3
2012 View-invariant gait recognition from low frame-rate videos
Al Mansur, Yasushi Makihara, Yasushi Yagi
ICPR3
2012 Point cloud transport
Hozuma Nakajima, Yasushi Makihara, Hsu Hsu, Ikuhisa Mitsugami, Mitsuru Nakazawa, Hirotake Yamazoe, Hitoshi Habe, Yasushi Yagi
ICPR8
2012 Dynamic scene reconstruction using asynchronous multiple Kinects
Mitsuru Nakazawa, Ikuhisa Mitsugami, Yasushi Makihara, Hozuma Nakajima, Hitoshi Habe, Hirotake Yamazoe, Yasushi Yagi
ICPR7
2012 8-D reflectance field for computational photography
Seiichi Tagawa, Yasuhiro Mukaigawa, Yasushi Yagi
ICPR3
2012 Inertial-sensor-based walking action recognition using robust step detection and inter-class relationships
Trung Ngo Thanh, Yasushi Makihara, Hajime Nagahara, Yasuhiro Mukaigawa, Yasushi Yagi
ICPR5
2012 Easy depth sensor calibration
Hirotake Yamazoe, Hitoshi Habe, Ikuhisa Mitsugami, Yasushi Yagi
ICPR4
2012 Gait recognition using images of oriented smooth pseudo motion
abstract
This paper proposes a method of gait recognition using not only shape feature but also motion feature from silhouette image sequences. The inner silhouette motion called pseudo motion is constructed by dividing the silhouette shape into small clusters and by computing many-to-many correspondence via earth mover's morphing framework. The raw pseudo motion, however, tends to be locally fluctuated in the spatio-temporal domain, and hence the spatio-temporal regularization is imposed to provide the smooth pseudo motion. The smooth pseudo motion image sequences are further partitioned into images with eight different orientations, and then averaged over each gait period to produce images of oriented smooth pseudo motion. Both shape and motion cues are integrated in score-level fusion framework based on linear logistic regression and the single-dimensional fused distance is returned by the learned optimal weights. The experiments with the publicly available gait database show the effectiveness of the proposed method compared with the case where the shape information is used alone.
Yasushi Makihara, Betria Silvana Rossa, Yasushi Yagi
SMC3
2012 Attacks using random forgery against DTW-based online signature verification algorithm
abstract
We investigated the false accept rates against randomly forged signatures written by each signer by using a DTW-based online signature verification algorithm. The experimental results show that attacks using some signers' signatures are stronger than those using skilled forged signatures and that some signers' signatures can be a wolf against the algorithm considered.
Daigo Muramatsu, Yasushi Yagi
SMC2
2012 Pedestrian detection based on appearance, motion, and shadow information
abstract
We present a new pedestrian detection algorithm that considers multiple information sources. Appearance-based detection methods face difficulties such as appearance variations and occlusions. Shape-based methods can have false positives on shadows since they usually have similar shapes with foreground objects. To deal with these problems, we use appearance, motion, and shadow information simultaneously in our detection method. We detect pedestrians using shape information of both foreground and shadow regions. Then, we filter the detection results based on motion information if available. The proposed method gives low false positives due to the integration of multiple information sources. Moreover, it alleviates the problem brought by occlusion since casted shadows are observable when foreground objects are occluded. Our experimental results show that the proposed algorithm provides good performance in difficult situations.
Junqiu Wang, Yasushi Yagi
SMC2
2012 The OU-ISIR Gait Database Comprising the Large Population Dataset and Performance Evaluation of Gait Recognition
abstract
This paper describes the world's largest gait database-the “OU-ISIR Gait Database, Large Population Dataset”-and its application to a statistically reliable performance evaluation of vision-based gait recognition. Whereas existing gait databases include at most 185 subjects, we construct a larger gait database that includes 4007 subjects (2135 males and 1872 females) with ages ranging from 1 to 94 years. The dataset allows us to determine statistically significant performance differences between currently proposed gait features. In addition, the dependences of gait-recognition performance on gender and age group are investigated and the results provide several novel insights, such as the gradual change in recognition performance with human growth.
Haruyuki Iwama, Mayu Okumura, Yasushi Makihara, Yasushi Yagi
IEEE Trans. Inf. Forensics Secur.4
2011 The optimal camera arrangement by a performance model for gait recognition
abstract
Recently, many gait recognition algorithms are proposed, and the optimal camera arrangement is necessary to maximize the performance. In this paper, we propose the optimal camera arrangement by using a performance model that considers observation conditions comprehensively. We select silhouette resolution, observation view, and its local and global changes as the observation conditions affecting the performance. Then, training sets composed of pairs of the observation conditions and the performance is obtained by gait recognition experiments under several camera arrangements. A performance model is constructed by applying Gaussian Processes Regression to the training set. The optimal arrangement is determined by estimating the performance for each camera arrangement with the performance model. The effectiveness of the proposed method is demonstrated by experiments of performance estimation with a training set including 17 subjects and the optimal camera arrangement.
Naoki Akae, Yasushi Makihara, Yasushi Yagi
FG3
2011 Gait recognition using periodic temporal super resolution for low frame-rate videos
abstract
This paper describes a method of gait recognition where both a gallery and a probe are based on low frame-rate videos. The sparsity of phases (stances) per gait period makes it much harder to match the gait using existing gait recognition algorithms. Consequently, we introduce a super resolution technique to generate a high frame-rate periodic image sequence as a preprocess to matching. First, the initial phase for each frame is estimated based on an exemplar of a high frame-rate gait image sequence. Images between a pair of adjacent frames sorted by the estimated phases are then filled using a morphing technique to avoid ghosting effects. Next, a manifold of the periodic gait image sequence is reconstructed based on the estimated phase and morphed images. Finally, the phase estimation and manifold reconstruction are iterated to generate better high frame-rate images in the energy minimization framework. Experiments with real data on 100 subjects demonstrate the effectiveness of the proposed method particularly for low frame-rate videos of less than 5 fps.
Naoki Akae, Yasushi Makihara, Yasushi Yagi
IJCB3
2011 Score-level fusion based on the direct estimation of the Bayes error gradient distribution
abstract
This paper describes a method of score-level fusion to optimize a Receiver Operating Characteristic (ROC) curve for multimodal biometrics. When the Probability Density Functions (PDFs) of the multimodal scores for each client and imposter are obtained from the training samples, it is well known that the isolines of a function of probabilistic densities, such as the likelihood ratio, posterior, or Bayes error gradient, give the optimal ROC curve. The success of the probability density-based methods depends on the PDF estimation for each client and imposter, which still remains a challenging problem. Therefore, we introduce a frame work of direct estimation of the Bayes error gradient that bypasses the troublesome PDF estimation for each client and imposter. The lattice-type control points are allocated in a multiple score space, and the Bayes error gradients on the control points are then estimated in a comprehensive manner in the energy minimization framework including not only the data fitness of the training samples but also the boundary conditions and monotonic increase constraints to suppress the over-training. The experimental results for both simulation and real public data show the effectiveness of the proposed method.
Yasushi Makihara, Daigo Muramatsu, Yasushi Yagi, Md. Altab Hossain
IJCB3
2011 Gait-based age estimation using a whole-generation gait database
abstract
This paper addresses gait-based age estimation using a large-scale whole-generation gait database. Previous work on gait-based age estimation evaluated their methods using databases that included only 170 subjects at most with a limited age variation, which was insufficient to statistically demonstrate the possibility of gait-based age estimation. Therefore, we first constructed a much larger whole generation gait database which includes 1,728 subjects with ages ranging from 2 to 94 years. We then provided a base line algorithm for gait-based age estimation implemented by Gaussian process regression, which has achieved successes in the face-based age estimation field, in conjunction with silhouette-based gait features such as an averaged silhouette (or Gait Energy Image) which has been used extensively in many gait recognition algorithms. Finally, experiments using the whole-generation gait database demonstrated the viability of gait-based age estimation.
Yasushi Makihara, Mayu Okumura, Haruyuki Iwama, Yasushi Yagi
IJCB4
2011 Phase registration in a gallery improving gait authentication
abstract
In this paper, we propose a method of inertial sensor-based gait authentication by inter-period phase registration of an owner's gallery. In spite of the importance for gait authentication of constructing a gallery of phase-registered gait patterns, previous implementations just relied on simple methods of period detection based on heuristic knowledge such as local peaks/valleys or local auto-correlation of the gait signals. Consequently, we propose to improve a gait gallery by incorporating a phase registration technique which globally optimizes inter-period phase consistency in an energy minimization framework. However, the previous phase registration technique suffers from a phase distortion problem due to ambiguities in the combination of a periodic signal function and a phase evolution function. We present a linear phase evolution prior to constructing an undistorted gait signal for better matching performance. Experiments using real gait signals from 32 subjects show that the proposed methods outperform the latest methods in the field.
Trung Ngo Thanh, Yasushi Makihara, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Yasushi Yagi
IJCB6
2011 Action recognition using dynamics features
abstract
In this paper, we propose a method of action recognition using dynamics features based on physics model. The dynamics features are composed of torques from knee and hip joints of both legs and implicitly include the gravity, ground reaction forces, and the pose of the remaining body parts. These features are more discriminative than the kinematics features, and they result in a low dimensional representation of a human action which preserves much information of the original high dimensional pose. This low dimensional feature allows us to achieve a good classification performance even with a relatively small training data in a simple classification framework such as HMM. The effectiveness of the proposed method is demonstrated through experiments on the CMU motion capture dataset with various actions.
Al Mansur, Yasushi Makihara, Yasushi Yagi
ICRA3
2010 Foreground and Shadow Segmentation Based on a Homography-Correspondence Pair
Haruyuki Iwama, Yasushi Makihara, Yasushi Yagi
ACCV (4)3
2010 Gait Analysis of Gender and Age Using a Large-Scale Multi-view Gait Database
Yasushi Makihara, Hidetoshi Mannami, Yasushi Yagi
ACCV (2)3
2010 Temporal Super Resolution from a Single Quasi-periodic Image Sequence Based on Phase Registration
Yasushi Makihara, Atsushi Mori, Yasushi Yagi
ACCV (1)3
2010 Phase Registration of a Single Quasi-Periodic Signal Using Self Dynamic Time Warping
Yasushi Makihara, Trung Ngo Thanh, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Yasushi Yagi
ACCV (3)6
2010 Earth Mover's Morphing: Topology-Free Shape Morphing Using Cluster-Based EMD Flows
Yasushi Makihara, Yasushi Yagi
ACCV (4)2
2010 Hemispherical Confocal Imaging Using Turtleback Reflector
Yasuhiro Mukaigawa, Seiichi Tagawa, Ramesh Raskar, Yasuyuki Matsushita, Yasushi Yagi
ACCV (1)6
2010 Analysis of light transport in scattering media
abstract
We propose a new method to analyze light transport in homogeneous scattering media. The incident light undergoes multiple bounces in translucent objects, and produces a complex light field. Our method analyzes the light transport in two steps. First, single and multiple scattering are separated by projecting high-frequency stripe patterns. Then, multiple scattering is decomposed into each bounce component based on the light transport equation. The light field for each bounce is recursively estimated. Experimental results show that light transport in scattering media can be decomposed and visualized for each bounce.
Yasuhiro Mukaigawa, Yasushi Yagi, Ramesh Raskar
CVPR2
2010 Silhouette transformation based on walking speed for gait identification
abstract
We propose a method of gait silhouette transformation from one speed to another to cope with walking speed changes in gait identification. When a person changes his/her walking speed, dynamic features (e.g. stride and joint angle) are changed while static features (e.g. thigh and shin lengths) are unchanged. Based on the fact, firstly, static and dynamic features are separated from gait silhouettes by fitting a human model. Secondly, a factorization-based speed transformation model for the dynamic features is created using a training set for multiple persons on multiple speeds. This model can transform the dynamic features from a reference speed to another arbitrary speed. Finally, silhouettes are restored by combining the unchanged static features and the transformed dynamic features. Evaluation by gait identification using silhouette-based frequency-domain features shows the effectiveness of the proposed method.
Akira Tsuji, Yasushi Makihara, Yasushi Yagi
CVPR3
2010 How to Control Acceptance Threshold for Biometric Signatures with Different Confidence Values?
abstract
In the biometric verification, authentication is given when a distance of biometric signatures between enrollment and test phases is less than an acceptance threshold, and the performance is usually evaluated by a so-called Receiver Operating Characteristics (ROC) curve expressing a trade off between False Rejection Rate (FRR) and False Acceptance Rate (FAR). On the other hand, it is also well known that the performance is significantly affected by the situation differences between enrollment and test phases. This paper describes a method to adaptively control an acceptance threshold with quality measures derived from situation differences so as to optimize the ROC curve. We show that the optimal evolution of the adaptive threshold in the domain of the distance and quality measure is equivalent to a constant evolution in the domain of the error gradient defined as a ratio of a total error rate to a total acceptance rate. An experiment with simulation data demonstrates that the proposed method outperforms the previous methods, particularly under a lower FAR or FRR tolerance condition.
Yasushi Makihara, Md. Altab Hossain, Yasushi Yagi
ICPR3
2010 Cluster-Pairwise Discriminant Analysis
abstract
Pattern recognition problems often suffer from the larger intra-class variation due to situation variations such as pose, walking speed, and clothing variations in gait recognition. This paper describes a method of discriminant subspace analysis focused on situation cluster pair. In training phase, both a situation cluster discriminant subspace and class discriminant subspaces for the situation cluster pair by using training samples of non recognition-target classes. In testing phase, given a matching pair of patterns of recognition-target classes, posterior of situation cluster pairs is estimated at first, and then the distance is calculated in the corresponding cluster-pairwise class discriminant subspace. The experiments both with simulation data and real data show the effectiveness of the proposed method.
Yasushi Makihara, Yasushi Yagi
ICPR2
2010 Gait Recognition Using Period-Based Phase Synchronization for Low Frame-Rate Videos
abstract
This paper proposes a method for period-based gait trajectory matching in the eigenspace using phase synchronization for low frame-rate videos. First, a gait period is detected by maximizing the normalized autocorrelation of the gait silhouette sequence for the temporal axis. Next, a gait silhouette sequence is expressed as a trajectory in the eigenspace and the gait phase is synchronized by time stretching and time shifting of the trajectory based on the detected period. In addition, multiple period-based matching results are integrated via statistical procedures for more robust matching in the presence of fluctuations among gait sequences. Results of experiments conducted with 185 subjects to evaluate the performance of the gait verification with various spatial and temporal resolutions, demonstrate the effectiveness of the proposed method.
Atsushi Mori, Yasushi Makihara, Yasushi Yagi
ICPR3
2010 Color Analysis for Segmenting Digestive Organs in VCE
abstract
This paper presents an efficient method for automatically segmenting the digestive organs in a Video Capsule Endoscopy (VCE) sequence. The method is based on unique characteristics of color tones of the digestive organs. We first introduce a color model of the gastrointestinal (GI) tract containing the color components of GI wall and non-wall regions. Based on the wall regions extracted from images, the distribution along the time dimension for each color component is exploited to learn the dominant colors that are candidates for discriminating digestive organs. The strongest candidates are then combined to construct a representative signal to detect the boundary of two adjacent regions. The results of experiments are comparable with previous works, but computation cost is more efficient.
Hai Vu, Yasushi Yagi, Tomio Echigo, Masatsugu Shiba, Kazuhide Higuchi, Tetsuo Arakawa, Keiko Yagi
ICPR2
2010 Visual tracking and segmentation using appearance and spatial information of patches
abstract
Object tracking and segmentation find a wide range of applications in robotics. Tracking and segmentation are difficult in cluttered and dynamic backgrounds. We propose a tracking and segmentation algorithm in which tracking and segmentation are performed consecutively. We separate input images into disjoint patches using an efficient oversegmentation algorithm. Objects and their background are described by bags of patches. We classify the patches in a new frame by searching k nearest neighbors. K-d trees are constructed using these patches to reduce computational complexity. Target location is estimated coarsely by running the mean-shift algorithm. Based on the estimated locations, we classify the patches again using appearance and spatial information. This strategy out-performs direct segmentation of patches based on appearance information only. Experimental results show that the proposed algorithm provides good performance on difficult sequences with clutter.
Junqiu Wang, Yasushi Yagi
ICRA2
2010 Clothing-invariant gait identification using part-based clothing categorization and adaptive weight control
Md. Altab Hossain, Yasushi Makihara, Junqiu Wang, Yasushi Yagi
Pattern Recognit.4
2009 Adaptive-Scale Robust Estimator Using Distribution Model Fitting
Trung Ngo Thanh, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Masahiko Yachida, Yasushi Yagi
ACCV (3)6
2009 People Tracking and Segmentation Using Efficient Shape Sequences Matching
Junqiu Wang, Yasushi Yagi, Yasushi Makihara
ACCV (2)2
2009 Dense 3D reconstruction method using a single pattern for fast moving object
abstract
Dense 3D reconstruction of extremely fast moving objects could contribute to various applications such as body structure analysis and accident avoidance and so on. The actual cases for scanning we assume are, for example, acquiring sequential shape at the moment when an object explodes, or observing fast rotating turbine's blades. In this paper, we propose such a technique based on a one-shot scanning method that reconstructs 3D shape from a single image where dense and simple pattern are projected onto an object. To realize dense 3D reconstruction from a single image, there are several issues to be solved; e.g. instability derived from using multiple colors, and difficulty on detecting dense pattern because of influence of object color and texture compression. This paper describes the solutions of the issues by combining two methods, that is (1) an efficient line detection technique based on de Bruijn sequence and belief propagation, and (2) an extension of shape from intersections of lines method. As a result, a scanning system that can capture an object in fast motion has been actually developed by using a high-speed camera. In the experiments, the proposed method successfully captured the sequence of dense shapes of an exploding balloon, and a breaking ceramic dish at 300–1000 fps.
Ryusuke Sagawa, Yuichi Ota, Yasushi Yagi, Ryo Furukawa 0001, Naoki Asada, Hiroshi Kawasaki
ICCV3
2009 An adaptive-scale robust estimator for motion estimation
abstract
Although RANSAC is the most widely used robust estimator in computer vision, it has certain limitations making it ineffective in some situations, such as the motion estimation problem, in which uncertainty on the image features changes according to the capturing conditions. The greatest problem is that the threshold used by RANSAC to detect inliers cannot be changed adaptively; instead it is fixed by the user. An adaptive scale algorithm must therefore be applied in such cases. In this paper, we propose a new adaptive scale robust estimator that adaptively finds the best solution with the best scale to fit the inliers, without the need for predefined information. Our new adaptive scale estimator matches the residual probability density from an estimate and the standard Gaussian probability density function to find the best inlier scale. Our algorithm is evaluated in several motion estimation experiments under varying conditions and the results are compared with several of the latest adaptive-scale robust estimators.
Trung Ngo Thanh, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Masahiko Yachida, Yasushi Yagi
ICRA6
2009 Towards an Interpretation of Intestinal Motility Using Capsule Endoscopy Image Sequences
Hai Vu, Tomio Echigo, Ryusuke Sagawa, Keiko Yagi, Masatsugu Shiba, Kazuhide Higuchi, Tetsuo Arakawa, Yasushi Yagi
PSIVT8
2009 Wearable imaging system for capturing omnidirectional movies from a first-person perspective
abstract
We propose a novel wearable imaging system that can capture omnidirectional movies from the viewpoint of the camera wearer. The imaging system solves the problems of resolution uniformity and gaze matching that conventional approaches do not address. We combine cameras with curved mirrors that control the projection of the imaging system to produce uniform resolution. Use of the mirrors also enables the viewpoint to be moved closer to the eyes of the camera wearer, thus reducing gaze mismatching. The optics, including the curved mirror, have been designed to form an objective projection. The capability of the designed optics is evaluated with respect to resolution, aberration, and gaze matching. We have developed a prototype based on the designed optics for practical use. Capability of the prototype and effectiveness of first-person perspective omnidirectional movies were demonstrated through quantitative evaluations and presentation experiments to ordinary people, respectively.
Kazuaki Kondo, Yasuhiro Mukaigawa, Yasushi Yagi
VRST3
2009 Adaptive Mean-Shift Tracking With Auxiliary Particles
abstract
We present a new approach for robust and efficient tracking by incorporating the efficiency of the mean-shift algorithm with the multihypothesis characteristics of particle filtering in an adaptive manner. The aim of the proposed algorithm is to cope with problems that were brought about by sudden motions and distractions. The mean-shift tracking algorithm is robust and effective when the representation of a target is sufficiently discriminative, the target does not jump beyond the bandwidth, and no serious distractions exist. We propose a novel two-stage motion estimation method that is efficient and reliable. If a sudden motion is detected by the motion estimator, some particle-filtering-based trackers can be used to outperform the mean-shift algorithm, at the expense of using a large particle set. In our approach, the mean-shift algorithm is used, as long as it provides reasonable performance. Auxiliary particles are introduced to cope with distractions and sudden motions when such threats are detected. Moreover, discriminative features are selected according to the separation of the foreground and background distributions when threats do not exist. This strategy is important, because it is dangerous to update the target model when the tracking is in an unsteady state. We demonstrate the performance of our approach by comparing it with other trackers in tracking several challenging image sequences.
Junqiu Wang, Yasushi Yagi
IEEE Trans. Syst. Man Cybern. Part B2
2008 Dynamic scene shape reconstruction using a single structured light pattern
abstract
3D acquisition techniques to measure dynamic scenes and deformable objects with little texture are extensively researched for applications like the motion capturing of human facial expression. To allow such measurement, several techniques using structured light have been proposed. These techniques can be largely categorized into two types. The first involves techniques to temporally encode positional information of a projector’s pixels using multiple projected patterns, and the second involves techniques to spatially encode positional information into areas or color spaces. Although the former allows dense reconstruction with a sufficient number of patterns, it has difficulty in scanning objects in rapid motion. The latter technique uses only a single pattern, so this problem can be resolved, however, it often uses complex patterns or color intensities, which are weak to noise, shape distortions, or textures. Thus, it remains an open problem to achieve dense and stable 3D acquisition in real cases. In this paper, we propose a technique to achieve dense shape reconstruction that requires only a single-frame image of a grid pattern. The proposed technique also has the advantage of being robust in terms of image processing.
Hiroshi Kawasaki, Ryo Furukawa 0001, Ryusuke Sagawa, Yasushi Yagi
CVPR4
2008 One-shot range scanner using coplanarity constraints
abstract
Methods for scanning dynamic scenes are important in many applications and many systems using structured light have been proposed. Many of these systems use either multiple patterns projected rapidly or a single pattern. Although the former allows dense reconstruction with a sufficient number of patterns, it has difficulty in capturing objects in rapid motion. The latter technique uses only a single pattern and have no such difficulties, however, they often have stability problems and their result tend to have low resolution. In this paper, we develop a system to achieve dense and accurate 3D measurement from only a single image. The proposed system also has the advantage of being robust in terms of image processing.
Ryo Furukawa 0001, Huynh Quang Huy Viet, Hiroshi Kawasaki, Ryusuke Sagawa, Yasushi Yagi
ICIP5
2008 Patch-based adaptive tracking using spatial and appearance information
abstract
We present a patch-based tracking algorithm in which both appearance and spatial information are taken into account for target localization. We decompose a target into several patches based on appearance similarity and spatial distribution. Each patch has its distinctive appearance and spatial distribution. Appearance information is described by kernels which are non-parametric; while spatial information is represented by spatial Gaussians. The overall motion is estimated by mean shift algorithm. The motion is refined based on the likelihood images computed using pixel classification. The proposed tracker provides better position and likelihood images.
Junqiu Wang, Yasushi Yagi
ICIP2
2008 Clothes-invariant gait identification using part-based adaptive weight control
abstract
This paper describes a method of part-based gait identification under substantial clothes variations. When clothes types between a gallery and a probe are different, silhouettes fairly change for some parts and subject discrimination capability decrease for those parts. Therefore, we exploit the discrimination capability as a matching weight for each part and control the weights adaptively based on a distribution of distances between a probe and all the galleries. As a result of experiments with our clothes-variation gait dataset, the proposedmethod achievedmuch better performance than a whole-based approach.
Md. Altab Hossain, Yasushi Makihara, Wang Junqui, Yasushi Yagi
ICPR4
2008 Scale-invariant density-based clustering initialization algorithm and its application
abstract
In this paper, we bring out a new density-based clustering initialization algorithm which is invariant to the scale factor. Instead of using the scale factor while the cluster initialization, in this research, we determine the number and position of clusters according to the changes of cluster density with the division and agglomeration processes. During the division process, the initial cluster seeds are produced by a self-propagate method according to the density changes. The number of clusters is determined by agglomerating pair of RNN (reciprocal nearest neighbor) cluster seeds, when the density of newly merged cluster is increased. When no more cluster seeds can be merged any more, the remained number of cluster seeds is regarded as the real cluster number. Through various experiments, the effectiveness of the proposed algorithm has been proved.
Chunsheng Hua, Ryusuke Sagawa, Yasushi Yagi
ICPR3
2008 Silhouette extraction based on iterative spatio-temporal local color transformation and graph-cut segmentation
abstract
We propose an iterative scheme of spatio-temporal local color transformation of background and graph-cut segmentation for silhouette extraction. Given an initial background subtraction, spatio-temporal background color transformation is processed for fitting modeled background colors to input background ones under a different illumination condition. After foreground colors are modeled based on the fit background, spatio-temporal graph-cut algorithm is applied to acquire a foreground/background segmentation result. Because these two processes need well-segmented background and well-fit background each other, they are iterated in turn to obtain better silhouette extraction results. Silhouette extraction experiments for a walking human on a treadmill show the effectiveness of the proposed method.
Yasushi Makihara, Yasushi Yagi
ICPR2
2008 Analysis of subsurface scattering under generic illumination
abstract
We present a new method of analyzing subsurface scattering occurring in a translucent object from a single image taken under generic illumination. In our method, diffuse subsurface reflectance in the subsurface scattering model can be linearly solved by quantizing the distances between each pair of surface points. Then, the dipole approximation is fit to the diffuse subsurface reflectance. By applying our method to real images, we confirm that the parameters of subsurface scattering can be computed for different materials.
Yasuhiro Mukaigawa, Kazuya Suzuki, Yasushi Yagi
ICPR3
2008 Switching local and covariance matching for efficient object tracking
abstract
The covariance tracker finds the targets in consecutive frames by global searching. Covariance tracking has achieved impressive successes thanks to its ability of capturing spatial and statistical properties as well as the correlations between them. Nevertheless, the covariance tracker is relatively inefficient due to its heavy computational cost of model updating and comparing the model with the covariance matrices of the candidate regions. Moreover, it is not good at dealing with articulated object tracking since integral histograms are employed to accelerate the searching process. In this work, we aim to alleviate the computational burden by selecting appropriate tracking approaches. We compute foreground probabilities of pixels and localize the target by local searching when the tracking is in steady states. Covariance tracking is performed when distractions, sudden motions or occlusions are detected. Different from the traditional covariance tracker, we use log-Euclidean metrics instead of Riemannian invariant metrics which are more computationally expensive. The proposed tracking algorithm has been verified on many video sequences. It proves more efficient than the covariance tracker. It is also effective in dealing with occlusions, which are an obstacle for local mode-seeking trackers such as the mean-shift tracker.
Junqiu Wang, Yasushi Yagi
ICPR2
2008 Accurate calibration of intrinsic camera parameters by observing parallel light pairs
abstract
This study describes a method of estimating the intrinsic parameters of a perspective camera. In previous calibration methods for perspective cameras, the intrinsic and extrinsic parameters are estimated simultaneously during calibration. Thus, the intrinsic parameters depend on the estimation of the extrinsic parameters, which is inconsistent with the fact that intrinsic parameters are independent of extrinsic ones. Moreover, in a situation where the extrinsic parameters are not used, only the intrinsic parameters need to be estimated. In this case, an intrinsic parameter, such as focal length, is not sufficiently robust to combat the image processing noise, that is absorbed by both parameter types, during calibration. We therefore propose a new method that allows the estimation of intrinsic parameters without estimating the extrinsic parameters. In order to calibrate the intrinsic parameters, the proposed method observes parallel light pairs that are projected on different points. This is accomplished by applying the constraint that the relative angle of two parallel rays is constant irrespective of where the rays are projected. This method focuses only on intrinsic parameters and the calibrations are sufficiently robust as demonstrated in this study. Moreover, our method can visualize the error of the calibrated result and the degeneracy of the input data.
Ryusuke Sagawa, Yasushi Yagi
ICRA2
2008 Robust and real-time egomotion estimation using a compound omnidirectional sensor
abstract
We propose a new egomotion estimation algorithm for a compound omnidirectional camera. Image features are detected by a conventional feature detector and then quickly classified into near and far features by checking infinity on the omnidirectional image of the compound omnidirectional sensor. Egomotion estimation is performed in two steps: first, rotation is recovered using far features; then translation is estimated from near features using the estimated rotation. RANSAC is used for estimations of both rotation and translation. Experiments in various environments show that our approach is robust and provides good accuracy in real-time for large motions.
Trung Ngo Thanh, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Masahiko Yachida, Yasushi Yagi
ICRA6
2008 Human tracking and segmentation supported by silhouette-based gait recognition
abstract
Gait recognition has recently gained attention as an effective approach to identify individuals at a distance from a camera. Most existing gait recognition algorithms assume that people have been tracked and silhouettes have been segmented successfully. Tacking and segmentation are, however, very difficult especially for articulated objects such as human beings. Therefore, we present an integrated algorithm for tracking and segmentation supported by gait recognition. After the tracking module produces initial results consisting of bounding boxes and foreground likelihood images, the gait recognition module searches for the optimal silhouette-based gait models corresponding to the results. Then, the segmentation module tries to segment people out using the provided gait silhouette sequence as shape priors. Experiments on real video sequences show the effectiveness of the proposed approach.
Junqiu Wang, Yasushi Makihara, Yasushi Yagi
ICRA3
2008 Integrating Color and Shape-Texture Features for Adaptive Real-Time Object Tracking
abstract
We extend the standard mean-shift tracking algorithm to an adaptive tracker by selecting reliable features from color and shape-texture cues according to their descriptive ability. The target model is updated according to the similarity between the initial and current models, and this makes the tracker more robust. The proposed algorithm has been compared with other trackers using challenging image sequences, and it provides better performance.
Junqiu Wang, Yasushi Yagi
IEEE Trans. Image Process.2
2007 Synchronized Ego-Motion Recovery of Two Face-to-Face Cameras
Jinshi Cui, Yasushi Yagi, Hongbin Zha, Yasuhiro Mukaigawa, Kazuaki Kondo
ACCV (1)2
2007 Multiplexed Illumination for Measuring BRDF Using an Ellipsoidal Mirror and a Projector
Yasuhiro Mukaigawa, Kohei Sumino, Yasushi Yagi
ACCV (2)3
2007 Mirror Localization for Catadioptric Imaging System by Observing Parallel Light Pairs
Ryusuke Sagawa, Nobuya Aoki, Yasushi Yagi
ACCV (1)3
2007 Gait Identification Based on Multi-view Observations Using Omnidirectional Camera
Kazushige Sugiura, Yasushi Makihara, Yasushi Yagi
ACCV (1)3
2007 Discriminative Mean Shift Tracking with Auxiliary Particles
Junqiu Wang, Yasushi Yagi
ACCV (1)2
2007 High-Speed Measurement of BRDF using an Ellipsoidal Mirror and a Projector
abstract
Measuring BRDF (bi-directional reflectance distribution function) requires huge amounts of time because a target object must be illuminated from all incident angles and the reflected lights must be measured from all reflected angles. In this paper, we present a high-speed method to measure BRDFs using an ellipsoidal mirror and a projector. Our method makes it possible to change incident angles without a mechanical drive. Moreover, the omni-directional reflected lights from the object can be measured by one static camera at once. Our prototype requires only fifty minutes to measure anisotropic BRDFs, even if the lighting interval is one degree.
Yasuhiro Mukaigawa, Kohei Sumino, Yasushi Yagi
CVPR3
2007 High Dynamic Range Camera using Reflective Liquid Crystal
abstract
High dynamic range images (HDRIs) are needed for capturing scenes that include drastic lighting changes. This paper presents a method to improve the dynamic range of a camera by using a reflective liquid crystal. The system consists of a camera and a reflective liquid crystal placed in front of the camera. By controlling the attenuation rate of the liquid crystal, the scene radiance for each pixel is adaptively controlled. After the control, the original scene radiance is derived from the attenuation rate of the liquid crystal and the radiance obtained by the camera. A prototype system has been developed and tested for a scene that includes drastic lighting changes. The radiance of each pixel was independently controlled and the HDRIs were obtained by calculating the original scene radiance from these results.
Hidetoshi Mannami, Ryusuke Sagawa, Yasuhiro Mukaigawa, Tomio Echigo, Yasushi Yagi
ICCV5
2007 Mirror Localization for a Catadioptric Imaging System by Projecting Parallel Lights
abstract
This paper describes a method of mirror localization to calibrate a catadioptric imaging system. Even though the calibration of a catadioptric system includes the estimation of various parameters, in this paper we focus on the localization of the mirror. Since some previously proposed methods assume a single view point system, they have strong restrictions on the position and shape of the mirror. We propose a method that uses parallel lights to simplify the geometry of projection for estimating the position of the mirror, thereby not restricting the position or shape of the mirror. Further, we omit the translation process between the camera and calibration objects from the parameters to be estimated by observing some parallel lights from a different direction. We obtain the constraints on the projection and compute the error between the model of the mirror and the measurements. The position of the mirror is estimated by minimizing the error. We also test our method by simulation and real experiments, and finally we evaluate the accuracy of our method.
Ryusuke Sagawa, Nobuya Aoki, Yasuhiro Mukaigawa, Tomio Echigo, Yasushi Yagi
ICRA5
2007 Robust and Real-time Rotation Estimation of Compound Omnidirectional Sensor
abstract
Camera ego-motion consists of translation and rotation, in which rotation can be described simply by distant features. We present a robust rotation estimation using distant features given by our compound omnidirectional sensor. Features are detected by a conventional feature detector, and then distant features are identified by checking the infinity on the omnidirectional image of the compound sensor. The rotation matrix is estimated between consecutive video frames using RANSAC with only distant features. Experiments with various environments show that our approach is robust and also gives reasonable accuracy in real-time.
Trung Ngo Thanh, Hajime Nagahara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Masahiko Yachida, Yasushi Yagi
ICRA6
2007 Contraction Detection in Small Bowel from an Image Sequence of Wireless Capsule Endoscopy
Hai Vu, Tomio Echigo, Ryusuke Sagawa, Keiko Yagi, Masatsugu Shiba, Kazuhide Higuchi, Tetsuo Arakawa, Yasushi Yagi
MICCAI (1)8
2007 Adaptive dynamic range camera with reflective liquid crystal
Hidetoshi Mannami, Ryusuke Sagawa, Yasuhiro Mukaigawa, Tomio Echigo, Yasushi Yagi
J. Vis. Commun. Image Represent.5
2006 Matching Gait Image Sequences in the Frequency Domain for Tracking People at a Distance
Ryusuke Sagawa, Yasushi Makihara, Tomio Echigo, Yasushi Yagi
ACCV (2)4
2006 Gait Recognition Using a View Transformation Model in the Frequency Domain
Yasushi Makihara, Ryusuke Sagawa, Yasuhiro Mukaigawa, Tomio Echigo, Yasushi Yagi
ECCV (3)5
2006 Evaluation of HBP Mirror System for Remote Surveillance
abstract
The HBP (horizontal fixed viewpoint biconical paraboloidal) mirror is an anisotropic convex mirror that has a property of inhomogeneous angular resolution about azimuth angle. In this paper, we investigate the effectiveness of the HBP mirror system for remote surveillance. We developed a real remote surveillance system that is constructed by the HBP mirror system mounted on an electric cart. Through the surveillance experiments, surveyors usually looked almost front views, and they paid attention to interesting objects only when the cart approaches them. Since the HBP mirror system has high resolution in frontal view, it seems to work well in the surveillance. We also constructed a simulational remote surveillance environment in order to quantitatively compare the HBP mirror system with a conventional omnidirectional mirror system under fair experimental conditions. As a practical task, we assumed object searching in a virtually constructed devastated area. We confirmed that objects can be detected earlier and with certainty by the HBP mirror system
Kazuaki Kondo, Yasuhiro Mukaigawa, Toshiya Suzuki, Yasushi Yagi
IROS4
2005 Real Time 3D Environment Modeling for a Mobile Robot by Aligning Range Image Sequence
Ryusuke Sagawa, Nanaho Osawa, Tomio Echigo, Yasushi Yagi
BMVC4
2005 Non-isotropic Omnidirectional Imaging System for an Autonomous Mobile Robot
abstract
A real-time omnidirectional imaging system that can acquire an omnidirectional field of view at video rate using a convex mirror was applied to a variety of conditions. The imaging system consists of an isotropic convex mirror and a camera pointing vertically toward the mirror with its optical axis aligned with the mirror's optical axis. Because of these optics, angular resolution is independent of the azimuth angle. However, it is important for a mobile robot to find and avoid obstacles in its path. We consider that angular resolution in the direction of the robot's moving needs higher resolution than that of its lateral view. In this paper, we propose a non-isotropic omnidirectional imaging system for navigating a mobile robot.
Kazuaki Kondo, Yasushi Yagi, Masahiko Yachida
ICRA2
2005 Stereovision with a Single Camera and Multiple Mirrors
abstract
You can create catadioptric omnidirectional stereovision using several mirrors with a single camera. These systems have interesting advantages, for instance in the case of mobile robot navigation and environment reconstruction. Our paper aims at estimating the” quality” of such stereovision system. What happens when the number of mirrors increases? Is it better to increase the base-line or to increase the number of mirrors? We propose some criteria and a methodology to compare different significant categories (seven): three already existing systems and four new designs that we propose. We also study and propose a global comparison between the best configurations.
El Mustapha Mouaddib, Ryusuke Sagawa, Tomio Echigo, Yasushi Yagi
ICRA4
2005 Calibration of lens distortion by structured-light scanning
abstract
This paper describes a new method to automatically calibrate lens distortion of wide-angle lenses. We project structured-light patterns using a flat display to generate a map between the display and the image coordinate systems. This approach has two advantages. First, it is easier to take correspondences of image and marker (display) coordinates around the edge of a camera image than using a usual marker, e.g. a checker board. Second, since we can easily construct a dense map, a simple linear interpolation is enough to create an undistorted image. Our method is not restricted by the distortion parameters because it directly generates the map. We have evaluated the accuracy of our method and the error becomes smaller than results by parameter fitting.
Ryusuke Sagawa, Masaya Takatsuji, Tomio Echigo, Yasushi Yagi
IROS4
2005 Iconic Memory-Based Omnidirectional Route Panorama Navigation
abstract
A route navigation method for a mobile robot with an omnidirectional image sensor is described. The route is memorized from a series of consecutive omnidirectional images of the horizon when the robot moves to its goal. While the robot is navigating to the goal point, input is matched against the memorized spatio-temporal route pattern by using dual active contour models and the exact robot position and orientation is estimated from the converged shape of the active contour models.
Yasushi Yagi, Kousuke Imai, Kentaro Tsuji, Masahiko Yachida
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 Evaluation of Iconic Memory-based ORP Navigation
abstract
A route navigation method for a mobile robot with an omni-directional image sensor is described. The route is memorized from a series of consecutive omni-directional images at the horizon when the robot moves to the goal. While the robot is navigating to the goal point, the input is matched against the memorized spatio-temporal route pattern by using dual active contour models and the exact robot position and orientation is estimated from the converged shape of active contour models. In this paper, we evaluated the precision of the localization of the robot under several different moving conditions and environments.
Yasushi Yagi, Kentaro Tsuji, Masahiko Yachida
ICRA1
2004 Compound catadioptric stereo sensor for omnidirectional object detection
abstract
This paper describes a novel system for detecting objects close to our sensor. For real time detection and portability, we have developed a small sensor with compound spherical mirrors. Since an object is projected onto each mirror, our method computes the range by a catadioptric stereo method. Our method creates a lookup table of corresponding points for an infinite range. If an object is close enough to the sensor, the projected points of the object are different from these corresponding points. Thus, our method can detect near objects by taking the differences in intensity of the corresponding points between the images in the mirrors. We show the experimental setup of our sensor and the result for detecting near objects.
Ryusuke Sagawa, Naoki Kurita, Tomio Echigo, Yasushi Yagi
IROS4
2004 Editorial: Research in Japan on Omni-Directional Sensors and Their Applications
Yasushi Yagi, Katsushi Ikeuchi
Int. J. Comput. Vis.1
2004 Real-Time Omnidirectional Image Sensors
Yasushi Yagi, Masahiko Yachida
Int. J. Comput. Vis.1
2004 Super wide viewer using catadioptrical optics
abstract
Many applications have used a head-mounted display (HMD), such as in virtual and mixed realities and telepresence. However, the field of view (FOV) of commercial HMD systems is too narrow for feeling immersion. In this paper, we propose a super-wide field of view head-mounted display consisting of an ellipsoidal mirror and a hyperboloidal curved mirror. The horizontal FOV of the proposed HMD is 180 degrees and includes the peripheral vision of humans. It increases the reality and immersion for users.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
ACM Trans. Graph.2
2003 Iconic memory-based onmidirectional route panorama navigation
abstract
A route navigation method for a mobile robot with an omni-directional image sensor is described. The route is memorized from a series of consecutive omnidirectional images at the horizon when the robot moves to the goal. While the robot is navigating to the goal point, the input is matched against the memorized spatio-temporal route pattern by using dual active contour models and the exact robot position and orientation is estimated from the converged shape of active contour models.
Yasushi Yagi, Kousuke Imai, Masahiko Yachida
ICRA1
2003 Wide field of view catadioptrical head-mounted display
abstract
Many applications have used a Head-Mounted Display (HMD), such as in virtual and mixed realities, and tele-presence. The advantage of HMD systems is the ease of feeling a 3D world in the display of animation or movies. However, the field of view (FOV) of commercial HMD systems is too narrow for feeling immersion. The horizontal FOV of many commercial HMDs is around 60 degrees, significantly narrower than that of humans. In this paper, we propose a super wide field of view catadioptrical head-mounted display consisting of an ellipsoidal and a hyperboloidal curved mirror. The horizontal FOV of the proposed HMD is 180 degrees and includes the peripheral view of humans. It increases reality and immersion of users. As well, the central region (60 degrees) of the FOV can measure 3D distances using stereoscopics.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
IROS2
2003 Super wide viewer using catadioptrical optics
abstract
Many applications have used a Head-Mounted Display (HMD), such as in virtual and mixed realities, and tele-presence. The advantage of HMD systems is the ease of feeling a 3D world in the display of animation or movies. However, the field of view (FOV) of commercial HMD systems is too narrow for feeling immersion. The horizontal FOV of many commercial HMDs is around 60 degrees, significantly narrower than that of humans. In this paper, we propose a super wide field of view head-mounted display consisting of an ellipsoidal and a hyperboloidal curved mirror. The horizontal FOV of the proposed HMD is 180 degrees and includes the peripheral view of humans. It increases reality and immersion of users. As well, the central region (60 degrees) of the FOV can measure 3D distances using stereoscopics.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
VRST2
2003 Rolling and swaying motion estimation for a mobile robot by using omnidirectional optical flows
Yasushi Yagi, Wataru Nishi, Nels E. Benson, Masahiko Yachida
Mach. Vis. Appl.1
2003 Superresolution modeling using an omnidirectional image sensor
abstract
Recently, many virtual reality and robotics applications have been called on to create virtual environments from real scenes. A catadioptric omnidirectional image sensor composed of a convex mirror can simultaneously observe a 360-degree field of view making it useful for modeling man-made environments such as rooms, corridors, and buildings, because any landmarks around the sensor can be taken in and tracked in its large field of view. However, the angular resolution of the omnidirectional image is low because of the large field of view captured. Hence, the resolution of surface texture patterns on the three-dimensional (3-D) scene model generated is not sufficient for monitoring details. To overcome this, we propose a high resolution scene texture generation method that combines an omnidirectional image sequence using image mosaic and superresolution techniques.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
IEEE Trans. Syst. Man Cybern. Part B2
2002 Resolution Improving Method for a 3D Environment Modeling using Omnidirectional Image Sensor
abstract
Recently, many applications in virtual reality and robotics require to create virtual environment from a real scene. A catadioptric omnidirectional image sensor composed of a convex mirror can observe a 360-degree field of view at once. It is useful for modeling a man-made environment such as a room, a corridor and a building, because the landmarks around the sensor can be taken and tracked by its large field of view. However, the angular resolution of the omnidirectional image is low owing to capturing the large field of view. Therefore, the resolution of texture patterns of each surface on the generated 3D scene model is not enough for monitoring details. To solve this problem, we propose a high-resolution scene texture generation method that combines an omnidirectional image sequence using image mosaic and super-resolution techniques.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
ICRA2
2002 Face identification using an omnidirectional image sequence
abstract
Face is one of the most attractive information for personal identification. In this paper, we propose the face identification method from an omnidirectional image sequence. Since the omnidirectional image sensor, HyperOmni Vision, observes a 360-degree view around the robot, it can observe a global azimuth information of the person (face). It tracks the human face while the person walks around the camera. Under an assumption of smooth human motion, we identify the corresponding person from facial database.
Yu Ohara, Yasushi Yagi, Taro Yokoyama, Masahiko Yachida
IROS2
2001 Resolution improving method from multi-focal omnidirectional images
abstract
The omnidirectional image sensor named HyperOmniVision, is composed of a hyperboloidal mirror and conventional video camera. It can observe a 360 degree field of view and can transform an input image to a perspective image. However, it has an intrinsic problem where the image resolution of HyperOmniVision is lower than that of an ordinary video camera, because it has a structure whereby only one CCD captures whole periphery scene. We propose a resolution improvement method using sub-pixel displaced and multi-focused images.
Hajime Nagahara, Yasushi Yagi, Masahiko Yachida
ICIP (1)2
2000 Active Contour Road Model for Smart Vehicle
abstract
We propose a method to solve the general problem of road tracking and 3D-shape reconstruction for a smart vehicle. The method assumes that the road boundaries are parallel and that the width of the road is constant. We then detect and track the road region in the image using active contour models subject to a parallelism constraint. The system then generates a 3D-road model from a single image. We evaluate the effectiveness of the method by applying to real road scenes comprising more than 2000 images.
Yasushi Yagi, Yoshiteru Kawasaki, Masahiko Yachida, J. Michael Brady
ICPR1
2000 3D Line Segment Reconstruction by Using HyperOmni Vision and Omnidirectional Hough Transforming
abstract
We (1995) have proposed an omnidirectional image sensor called HyperOmni Vision. This sensor can acquire omnidirectional and perspective images in real time. In this paper we present a new Hough transform method for an omnidirectional image by using a cubic Hough space. Furthermore, we describe a method for reconstructing 3D line segments from an omnidirectional input image sequence obtained by HyperOmni Vision with known motion.
Kazumasa Yamazawa, Yasushi Yagi, Masahiko Yachida
ICPR2
2000 Environmental Map Generation and Egomotion Estimation in a Dynamic Environment for an Omnidirectional Image Sensor
abstract
Generation of a stationary environmental map is one of the important tasks for vision based robot navigation. Observational errors in the generated environmental map accumulate in long movements of the robot. To generate a large environmental map, it is desirable not to assume known robot motion. In this paper, under the assumption of unknown translational motions of the robot, we propose a method to generate a stationary environmental map, and estimate the egomotion of a robot in a dynamic environment by using an omnidirectional image sensor. Since both robot and objects move in the environment, the stationary map generation and the robot egomotion estimation by using a single camera are difficult because of correspondence ambiguity caused by occlusion. The proposed method can detect a moving object and find occlusion and mismatching by evaluating the estimation error of each object location.
Yasushi Yagi, Kouichi Shouya, Masahiko Yachida
ICRA1
2000 Generation of stationary environmental map under unknown robot motion
abstract
Under the assumption of known motion of a robot, environmental maps of a real scene can be successfully generated by monitoring azimuth changes in an image. Several researchers have used this property for robot navigation. However, it is difficult to observe the exact motion parameters of the robot because of encoder measurement error. Therefore, observational errors in the generated environmental map accumulate in long movements of the robot. To generate a large environmental map, it is desirable not to assume known robot motion. In this paper, under the assumption of unknown motions of the robot, we propose a method to generate an environmental map and estimate the egomotion of a robot, by using an omnidirectional image sensor.
Yasushi Yagi, Hiroaki Hamada, Nels E. Benson, Masahiko Yachida
IROS1
1999 Reactive visual navigation based on omnidirectional sensing-path following and collision avoidance
abstract
Described here is a visual navigation method for navigating a mobile robot along a man-made route such as a corridor or a street. We have proposed an image sensor, named HyperOmniVision, with a hyperboloidal mirror for vision based navigation of the mobile robot. This sensing system can acquire an omnidirectional view around the robot in real time. In the case of the man-made route, road boundaries between the ground plane and wall appears as a close-looped curve in the image. By making use of this optical characteristic the robot can avoid obstacles and move along the corridor by tracking the close-looped curve with an active contour model. Experiments that have been done in a real environment are described.
Yasushi Yagi, Hiroyuki Nagai, Kazumasa Yamazawa, Masahiko Yachida
IROS1
1998 Facial Contour Extraction Model
Taro Yokoyama, Yasushi Yagi, Masahiko Yachida
FG2
1998 Active contour model for extracting human faces
abstract
In this paper, we propose a facial contour extraction model. The influence of environmental changes on the model is small, and it is based on an active contour model (ACM). Our method has three characteristics: global shape constraint of axis-symmetry, dual-scale filtering and iterative initialization. We apply the method to real human faces of more than 60 people with our face recognition system for evaluating effectiveness.
Taro Yokoyama, Yasushi Yagi, Masahiko Yachida
ICPR2
1998 Building Local Floor Map by Use of Ultrasonic and Omni-Directional Vision Sensor
abstract
In this paper, we propose a new fusion approach which uses the ultrasonic sensor aided by an omnidirectional vision sensor to give a grid based free space around the robot. By use of the ultrasonic sensor, the robot can obtain a conservative range information based on our nearby range filtering method. This filtering can give a more reliable result considering the sensor's problem of specular reflection. Also, by use of the special omni-directional vision sensor we developed, the color and edge information can be obtained in a single picture and mapped to the ground plane by the inverse perspective transformation. Thus the range, color and edge information can all be expressed on a metric grid-based representation which forms the basis of our radial fusion processing. Results in an indoor cluttered environment are given which show the usefulness of our proposed sensor fusion approach.
Shih-Chieh Wei, Yasushi Yagi, Masahiko Yachida
ICRA2
1998 Route Presentation for Mobile Robot Navigation by Omnidirectional Route Panorama Fourier Transformation
abstract
Described here is a route navigation method for a mobile robot with an omnidirectional image sensor. The route is memorized by a series of two dimensional Fourier power spectra of consecutive omnidirectional images at the horizon while the robot moves to the goal position. While the robot is navigating to the goal point, it is controlled by comparing the principal axis of inertia of its current position with that of the memorized Fourier power spectrum.
Yasushi Yagi, Syuji Fujimura, Masahiko Yachida
ICRA1
1997 Map generation for multiple image sensing sensor MISS under unknown robot egomotion
abstract
We design a new multiple image sensing sensor (MISS), which consists of a single camera and two different types of optics: an omnidirectional image sensor COPIS (conic projection image sensor) and binocular vision, for navigating the robot and understanding interesting objects in an environment by integrating both sensory data. Since COPIS observes a 360 degree view around the robot, it can observe global information of features (azimuths of vertical edges). On the other hand, binocular vision can obtain a sequence of stereo images although a field of view is limited by the visual angle of lens. In this paper, we integrate merits of omnidirectional vision and binocular vision, and propose an efficient sensing system for generating an environmental map by using MISS. Without an assumption of the exact known motion of the robot, MISS can generate an environmental map and estimate egomotion of the robot from both sensory information sources. In particular, the method is suitable for map generation during long robot movement because deadlocking error of the robot does not accumulate. The system has been evaluated on a prototype of MISS in a real environment.
Yasushi Yagi, Kazuhro Egami, Masahiko Yachida
IROS1
1996 Gesture recognition using colored gloves
abstract
This paper proposes a method to recognize hand gestures from a monocular image. We use colored gloves to detect specific region of hand easily. By using this method, we can treat easily the occlusion problem due to color information. Hand gesture are recognized by the decision tree made from the image features automatically by our proposed method. There are another pattern recognition methods using image features like nearest-neighbor method. But it is better for computation time to use the decision tree because the decision tree uses only few selected important features. In this paper, we show a learning result by the CAD model and recognition results in real images by our proposed method.
Yoshio Iwai, Ken Watanabe, Yasushi Yagi, Masahiko Yachida
ICPR3
1996 Rolling motion estimation for mobile robot by using omnidirectional image sensor HyperOmniVision
abstract
Described here is a method for estimating a rolling motion of a mobile robot from optical flows. We have proposed an image sensor with a hyperboloidal mirror for vision based navigation of the mobile robot. Its name is HyperOmniVision. This sensing system can acquire an omnidirectional view around the robot, in real-time, with use of the hyperboloidal mirror. The radial component of optical flow in HyperOmniVision has a periodic characteristic. The proposed method makes use of this characteristic to estimate robustly the rolling motion of the robot.
Yasushi Yagi, Wataru Nishi, Kazumasa Yamazawa, Masahiko Yachida
ICPR1
1996 The integration of an environmental map observed by multiple mobile robots with omnidirectional image sensor COPIS
abstract
We describe a method for generating the environmental map by cooperative observation among multiple mobile robots with omnidirectional visual sensor. Under the assumption of the known motion of the robot, an environmental map of a scene is generated by monitoring azimuth change in the image. First, each robot communicates and exchanges information of sensory date, then, can estimate a cooperative robot by evaluating relative directions of motion and estimated locations of vertical edges. Finally, a global environmental map is generated by integrating both environmental maps.
Yasushi Yagi, Shinichi Izuhara, Masahiko Yachida
IROS1
1996 Stabilization for mobile robot by using omnidirectional optical flow
abstract
Described here is a method for estimating rolling and swaying motions of a mobile robot from optical flows. We have proposed an image sensor with a hyperboloidal mirror for vision based navigation of the mobile robot. Its name is HyperOmni Vision. This sensing system can acquire an omnidirectional view around the robot, in real-time, with use of the hyperboloidal mirror. The radial component of optical flow in HyperOmniVision has a periodic characteristic. The circumferential component of optical flow has a symmetric characteristic. The proposed method makes use of these characteristics to estimate robustly the rolling and swaying motions of the robot.
Yasushi Yagi, Wataru Nishi, Kazumasa Yamazawa, Masahiko Yachida
IROS1
1995 Evaluating Effectivity of Map Generation by Tracking Vertical Edges in Omnidirectional Image Sequence
abstract
We designed a new omnidirectional image sensor, called COPIS (Conic Projection Image Sensor), to guide the navigation of a mobile robot. The feature of COPIS is passive sensing of the omnidirectional image of the environment, in real-time, using a conic mirror. COPIS is a suitable sensor for visual navigation in a real environment. Under the assumption of the known motion of the robot, an environmental map of an indoor scene is generated by monitoring the azimuth change in the image. We did several experiments in the simple indoor environment. The precision of the obtained environmental maps was sufficient for robot navigation in such environment. In these experiments, to examine the potential of COPIS against the effects of observational errors in real-time navigation, we simplified the image processing method and the experimental environment. In this paper, we improve the image processing method for extracting and tracking vertical edges, taking care of the reliability, and also evaluate the effectiveness of COPIS in a real indoor and outdoor environment.
Yasushi Yagi, Kazuya Sato, Masahiko Yachida
ICRA1
1995 Obstacle Detection with Omnidirectional Image Sensor HyperOmni Vision
abstract
Described here is an image sensor with a hyperboloidal mirror for vision based navigation of a mobile robot. Its name is HyperOmni Vision. This sensing system can acquire an omnidirectional view around the robot, in real-time, with use of a hyperboloidal mirror. The authors show a prototype of a mobile robot system with HyperOmni Vision and a method for estimating the motion of the robot and finding unknown obstacles.
Kazumasa Yamazawa, Yasushi Yagi, Masahiko Yachida
ICRA2
1995 Map-based navigation for a mobile robot with omnidirectional image sensor COPIS
abstract
We designed a new omnidirectional image sensor COPIS (Conic Projection Image Sensor) to guide the navigation of a mobile robot. The feature of COPIS is passive sensing of the omnidirectional image of the environment, in real-time (at the frame rate of a TV camera), using a conic mirror. COPIS is a suitable sensor for visual navigation in a real world environment. We report here a method for navigating a robot by detecting the azimuth of each object in the omnidirectional image. The azimuth is matched with the given environmental map. The robot can precisely estimate its own location and motion (the velocity of the robot) because COPIS observes a 360/spl deg/ view around the robot, even when all edges are not extracted correctly from the omnidirectional image. The robot can avoid colliding against unknown obstacles and estimate locations by detecting azimuth changes, while moving about in the environment. Under the assumption of the known motion of the robot, an environmental map of an indoor scene is generated by monitoring azimuth change in the image.>
Yasushi Yagi, Yoshimitsu Nishizawa, Masahiko Yachida
IEEE Trans. Robotics Autom.1
1994 Multiple Visual Sensing System for Mobile Robot
abstract
In this paper, we propose a new multiple visual sensing sensor (MISS), which combines with an omnidirectional image sensor COPIS (COnic Projection Image Sensor) and binocular vision, for navigating the robot and understanding interesting objects in an environment by integrating both sensory data. Since COPIS observes a 360 degree view around the robot, the robot can always estimate its own location and motion precisely. The location of unknown objects can also be estimated. COPIS can observe a global and precise information of features (vertical edges) in real-time, however, it is difficult to understand their details of shapes. On the other hand, the mobile robot has binocular vision and obtains a sequence of stereo images from the environment. We extend the principle of trinocular vision to establish correspondences between a sequence of binocular images. Although a view field of the binocular vision is limited by a visual angle of lens, the binocular vision is useful for understand spatial configuration of the environment. Therefore we integrate both merits and propose an efficient sensing system. The system has been evaluated on the prototype sensor in actual environment.>
Yasushi Yagi, Hitoshi Okumura, Masahiko Yachida
ICRA1
1994 Detection of unknown moving objects by reciprocation of observed information between mobile robot
abstract
Described here is a method for navigating multiple mobile robots which estimate locations and motions of unknown moving objects. Each robot with conic projection image sensor COPIS can observe an omnidirectional view around the robot in real-time with use of a conic mirror. First, each robot communicates and exchanges information of sensory data. Then, unknown moving objects and a reciprocating robot are discriminated, and motions and locations of unknown moving objects are estimated by cooperative observation of azimuth changes between the two reciprocating robots. >
Yasushi Yagi, Ya Lin, Masahiko Yachida
IROS1
1994 Real-time omnidirectional image sensor (COPIS) for vision-guided navigation
abstract
Describes a conic projection image sensor (COPIS) and its application: navigating a mobile robot in a manner that avoids collisions with objects approaching from any direction. The COPIS system acquires an omnidirectional view around the robot, in real-time, with use of a conic mirror. Based on the assumption of constant linear motion of the robot and objects, the objects moving along collision paths are detected by monitoring azimuth changes. Confronted with such objects, the robot changes velocity to avoid collision and determines locations and velocities.>
Yasushi Yagi, Shinjiro Kawato, Saburo Tsuji
IEEE Trans. Robotics Autom.1
1993 Omnidirectional imaging with hyperboloidal projection
abstract
Described here is an image sensor with a hyperboloidal mirror for vision based navigation of a mobile robot. Its name is HyperOmni Vision. This sensing system can acquire an omnidirectional view around the robot, in real-time, with use of a hyperboloidal mirror.
Kazumasa Yamazawa, Yasushi Yagi, Masahiko Yachida
IROS2
1992 Map based navigation of the mobile robot using omnidirectional image sensor COPIS
abstract
The authors propose a novel omnidirectional image sensor, COPIS (Conic Projection Image Sensor), for guiding the navigation of a mobile robot. It performs passive sensing of the omnidirectional image of the environment in real time (at the frame rate of a TV camera) using a conic mirror. COPIS is a suitable sensor for visual navigation in the real-world environment with moving objects. The authors describe a method for estimating the location and the motion of the robot by detecting the azimuth of each object in the omnidirectional image. In this method, the azimuth is matched with the given environmental map. They also present a method to avoid collision against unknown obstacles and estimate their locations by detecting their azimuth changes while the robot is moving in the environment. Using the COPIS system, several experiments in the real world were performed.>
Yasushi Yagi, Yoshimitsu Nishizawa, Masahiko Yachida
ICRA1
1991 Real-time generation of environmental map and obstacle avoidance using omnidirectional image sensor with conic mirror
abstract
An omnidirectional image sensor COPIS (conic projection image sensor) is proposed for guiding navigation of a mobile robot. It features passive sensing of the omnidirectional environment in real-time using a conic mirror. Because the conic mirror is used, its image is under conic projection; where the azimuth of each point in the scene appears in the image as its direction from the image center. The authors describe COPIS and its application to guide the navigation of a mobile robot. The COPIS system acquires an omnidirectional view around the robot in real-time by using a conic mirror. Under the assumption of constant motion of the robot, locations of objects around the robot can be estimated by detecting their azimuth changes in the omnidirectional image. Using this method, the robot generates an environmental map of an indoor scene while it is moving in the environment. A method to avoid collision against objects by detecting their azimuth changes is presented.>
Yasushi Yagi, Masahiko Yachida
CVPR1
1991 Collision avoidance using omnidirectional image sensor (COPIS)
abstract
A conic projection image sensor (COPIS) and its application to guide the navigation of a mobile robot are described. The COPIS system acquires an omnidirectional view around the robot in real-time by using a conic mirror. Under the assumption of constant linear motion of the robot and moving objects, objects moving along collision paths are found by monitoring their azimuth changes. If such objects are found, the robot changes its velocity to avoid collision. Judgment as to whether each object is static or moving and estimation of its location and velocity are possible by the change of the robot's velocity.>
Yasushi Yagi, Shinjiro Kawato, Saburo Tsuji
ICRA1
1991 Estimating location and avoiding collision against unknown obstacles for the mobile robot using omnidirectional image sensor COPIS
abstract
Proposes a new omnidirectional image sensor, COPIS (conic projection image sensor), for guiding navigation of a mobile robot. It features passive sensing of omnidirectional images of the environment in real-time (at the frame rate of a TV camera) using a conic mirror. COPIS is a suitable sensor for visual navigation in real world environment with moving objects. The paper describes a method for estimating the location and the motion of the robot by detecting the azimuth of each object in the omnidirectional image. It also presents a method to avoid collision against unknown obstacles and estimates their locations by detecting their azimuth changes while the robot is moving in the environment.>
Yasushi Yagi, Yoshimitsu Nishizawa, Masahiko Yachida
IROS1
1985 Dynamic scene analysis for a mobile robot in a man-made environment
abstract
Analysis of scene viewed continuously from a robot moving in a man-made environment, such as a building or a plant, yields useful information for the navigation. The knowledge on the environment, richness of the scene in vertical edges and the flatness of the floor, is arranged in constraints for the dynamic scene analysis. The rotational component of camera motion is estimated first from image points invarient from translation. After compensating for movements by the rotation between the consecutive images, the foci of expansion of translational motion of both the robot and moving objects are determined.
Saburo Tsuji, Yasushi Yagi, Minoru Asada
ICRA2