Takayoshi Yamashita

dblp:64/5510 · DBLP profile ↗
← Back
75ranked-venue papers
8as first author
35since 2021 · last 2026
0000-0003-2631-9856ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 3 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 8 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Systems, architecture and hardware · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Capture-Calibrate-Coach: A Graph-Based Framework for Knowledge Monitoring Estimation and Adaptive Feedback
Li Chen 0032, Cheng Tang 0001, Boxuan Ma, Daisuke Deguchi, Takayoshi Yamashita, Atsushi Shimada 0001
AIED7
2026 Few-Shot Adaptive Open-Set Object Detection with Personalized Scene Generation
Yuzuru Nakamura, Yasunori Ishii, Takayoshi Yamashita
ICPR (5)3
2025 From Reflections to Motifs: A Graph-Based Analysis of Learners' Knowledge Construction
Li Chen 0032, Cheng Tang 0001, Daisuke Deguchi, Takayoshi Yamashita, Atsushi Shimada 0001
AIED (6)5
2025 OD-RASE: Ontology-Driven Risk Assessment and Safety Enhancement for Autonomous Driving
Kota Shimomura, Masaki Nambata, Atsuya Ishikawa, Ryota Mimura, Koki Inoue, Takayoshi Yamashita, Takayuki Kawabuchi
ICCV6
2025 Weight Pruning to Mitigate Class-Specific Accuracy Degradation for LiDAR-Based 3D Object Detection
abstract
The realization of autonomous driving systems requires efficient and accurate 3D object detection to identify objects such as vehicles, pedestrians, and cyclists within the driving environment using point cloud data. For achieving both high speed processing and high accuracy, it is necessary to reduce model size by model compression techniques, such as pruning, while maintaining performance. However, pruning for 3D object detection tasks has not been extensively studied, and the effects of applying existing pruning methods to 3D object detection models remain unclear. In this paper, we clarify the problems of pruning 3D object detection models with existing methods through preliminary experiments, and propose a pruning method suitable for 3D object detection models that solves these problems. Our preliminary experiments reveal that existing pruning methods significantly degrade detection performance for specific object classes. To address this issue, we propose a pruning method that preserves class-specific knowledge, mitigating biased accuracy degradation across different object classes. Experimental results on the KITTI dataset demonstrate that the proposed method can be combined with existing pruning methods without conflicts and achieves higher accuracy than existing methods.
Tenshi Ito, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
IV3
2025 Single-agent vs. Multi-agent LLM Strategies for Automated Student Reflection Assessment
Li Chen 0032, Cheng Tang 0001, Valdemar Svábenský, Daisuke Deguchi, Takayoshi Yamashita, Atsushi Shimada 0001
PAKDD (5)6
2024 DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention
NguyenHuu BaoLong, Yuzhi Shi, Tsubasa Hirakawa, Takayoshi Yamashita, Tohgoroh Matsui, Hironobu Fujiyoshi
ACCV (10)5
2024 Faster Convergence and Uncorrelated Gradients in Self-Supervised Online Continual Learning
Koyo Imai, Naoto Hayashi, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
ACCV (8)4
2024 Active Domain Adaptation with False Negative Prediction for Object Detection
abstract
Domain adaptation adapts models to various scenes with different appearances. In this field, active domain adaptation is crucial in effectively sampling a limited number of data in the target domain. We propose an active domain adaptation method for object detection, focusing on quantifying the undetectability of objects. Existing methods for active sampling encounter challenges in considering undetected objects while estimating the uncertainty of model predictions. Our proposed active sampling strategy addresses this issue using an active learning approach that simultaneously accounts for uncertainty and undetectability. Our newly proposed False Negative Prediction Module evaluates the undetectability of images containing undetected objects, enabling more informed active sampling. This approach considers previously overlooked undetected objects, thereby reducing false negative errors. Moreover, using unlabeled data, our proposed method utilizes uncertainty-guided pseudo-labeling to enhance domain adaptation further. Extensive experiments demonstrate that the performance of our proposed method closely rivals that of fully supervised learning while requiring only a fraction of the labeling efforts needed for the latter.
Yuzuru Nakamura, Yasunori Ishii, Takayoshi Yamashita
CVPR3
2024 Deep Single Image Camera Calibration by Heatmap Regression to Recover Fisheye Images Under Manhattan World Assumption
abstract
A Manhattan world lying along cuboid buildings is useful for camera angle estimation. However, accurate and robust angle estimation from fisheye images in the Manhattan world has remained an open challenge because general scene images tend to lack constraints such as lines, arcs, and vanishing points. To achieve higher accuracy and robustness, we propose a learning-based calibration method that uses heatmap regression, which is similar to pose estimation using keypoints, to detect the directions of labeled image coordinates. Simultaneously, our two estimators recover the rotation and remove fisheye distortion by remapping from a general scene image. Without considering vanishing-point constraints, we find that additional points for learning-based methods can be defined. To compensate for the lack of vanishing points in images, we introduce auxiliary diagonal points that have the optimal 3D arrangement of spatial uniformity. Extensive experiments demonstrated that our method outperforms conventional methods on large-scale datasets and with off-the-shelf cameras.
Nobuhiko Wakai, Satoshi Sato, Yasunori Ishii, Takayoshi Yamashita
CVPR4
2024 Layer-Wise Relevance Propagation with Conservation Property for ResNet
Seitaro Otsuki, Tsumugi Iida, Félix Doublet, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura
ECCV (43)5
2024 Enhancing the Accuracy of Predicting Students Grades in Open-Ended Questions through Adjustments to Attention Weights
Masaki Koike, Hirokazu Kohama, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
EDM4
2024 Binary-Decomposed Vision Transformer: Compressing and Accelerating Vision Transformer by Binary Decomposition
abstract
Vision Transformers (ViTs) have emerged as versatile and high-performance models for various tasks such as image classification, object detection, and semantic segmentation. However, the ViT-L model, which demonstrates high accuracy, has a large number of parameters (307M), leading to increased computational requirements. To deploy ViTs on embedded devices and similar platforms, it is crucial to compress the model size and accelerate the inference process. In this paper, we propose the Binary-decomposed Vision Transformer (BdViT), a method for model compression and accelerated inference for ViTs models. BdViT consists of weight binarization based on vector decomposition and quantization of multiplication and addition operations, which does not require retraining model parameters. Through evaluation experiments using image recognition datasets, we demonstrated that BdViT can significantly reduce the number of parameters while mitigating performance degradation.
Ryota Kondo, Hiroaki Minoura, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
ICIP4
2024 Human-like Guidance by Generating Navigation Using Spatial-Temporal Scene Graph
abstract
Vehicle navigation systems use both GPS and map data, primarily information derived from map data. Conventional navigation systems assume that the user will look directly at the display to check information. Simultaneously provided text and voice often play only a supplementary role, which can lead to driver distraction and misinterpretation. In contrast, human navigation utilizes visual information, potentially reducing the cognitive load on drivers. Human-like Guidance is aimed at realizing a driving assistance system that supports navigation akin to human guidance. Implementing Human-like Guidance, requires the handling of video footage from in-vehicle cameras during vehicle operation, suggesting the need for an approach combining image recognition and language model. However, images captured during operation often contain superfluous information, making the selection of relevant objects for navigation challenging. Moreover, relying solely on image information makes it difficult to consider the relationship with surrounding objects. Therefore, this study proposes a Spatial-Temporal Scene Graph that can represent spatial and temporal information of objects from driving scene videos. Furthermore, we achieve Human-like Guidance through navigation generation using features extracted from the Spatial-Temporal Scene Graph. Our results show that our proposed method improves the accuracy of navigation generation accuracy compared to traditional image-based navigation methods. In addition, the use of a Spatial-Temporal Scene Graph enables the generation of human-like navigation that focuses on the movements of surrounding vehicle objects.
Hayato Suzuki, Kota Shimomura, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Shota Okubo, Nanri Takuya
IV4
2024 High-Precision for Multi-Task Learning from In-Vehicle Camera using BiFPN
abstract
Multi-task learning is effective for object detection and segmentation, which are closely related to each other and necessary for automated driving. However, there is a problem with the learning process in conventional multi-task learning models. In multi-task learning, common features among downstream tasks are first extracted by a backbone network. Then, these features are used for different downstream tasks. Since the required feature is different depending on the downstream task, it is necessary to extract features suitable for each downstream task. In this paper, we propose a multi-tasking model that introduces BiFPN feature fusion method for automated driving tasks and the Next-ViT model utilizing CNN and Transformer to extract features. From the evaluation experiments of automated driving tasks, we confirmed that the proposed method improves the accuracy of multi-task learning.
Hidenori Itaya, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
IV4
2024 LLM-Driven Ontology Learning to Augment Student Performance Analysis in Higher Education
Cheng Tang 0001, Li Chen 0032, Daisuke Deguchi, Takayoshi Yamashita, Atsushi Shimada 0001
KSEM (3)5
2024 Learning Intra-class Multimodal Distributions with Orthonormal Matrices
abstract
In this paper, we address the challenges of representing feature distributions which have multimodality within a class in deep neural networks. Existing online clustering methods employ sub-centroids to capture intra-class variations. However, conducting online clustering faces some limitations, i.e., online clustering assigns only a single sub-centroid to a feature vector extracted from a backbone and ignores the relationship between the other sub-centroids and the feature vector, and updating sub-centroids in an online clustering manner incurs significant storage costs. To address these limitations, we propose a novel method utilizing orthonormal matrices instead of sub-centroids for relaxing discrete assignments into continuous assignments. We update the orthonormal matrices using a gradient-based method, which eliminates the need for online clustering or additional storage. Experimental results on the CIFAR and ImageNet datasets exhibit that the proposed method outperforms current online clustering techniques in classification accuracy, sub-category discovery, and transferability, providing an efficient solution to the challenges posed by complex recognition targets.
Jumpei Goto, Yohei Nakata, Kiyofumi Abe, Yasunori Ishii, Takayoshi Yamashita
WACV5
2023 Embedding Human Knowledge into Spatio-Temproal Attention Branch Network in Video Recognition via Temporal attention
Saki Noguchi, Yuzhi Shi, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
BMVC4
2023 Recommending Learning Actions Using Neural Network
abstract
Many studies applying neural networks to the field of education have focused on student performance prediction and explainability of their decisions. While those studies introduced neural networks into educational settings, such networks cannot directly support student learnings in place of teachers. Therefore, we present a method that uses a general Transformer encoder to recommend appropriate learning actions for improving student performance. By considering the attention weight of a low-performing student to be close to that of a high-performing student, our method recommends the learning materials and actions for learning the materials. To evaluate the effectiveness of our method, we trained a deep neural network (DNN) on a private dataset of student operations (e.g., NEXT, PREV, OPEN) on digital learning materials obtained from a Japanese university. The number of operations divided by each learning material and by type of operation are input to the DNN, and the DNN outputs the student’s grade on 5-point scale. We applied our method with this trained DNN to samples that successfully predicted grades, and the number of operations increased on the basis of the recommended learning materials and actions. By re-inputting modified sample into the DNN, we then observe how the student performance changes. The results of this simple experiment indicate that more students improved their performance with both the material-based and operation-based recommendations than with random recommendations. The percentage of students whose grades improved tended to be larger for those with low grades. Specifically, the improvement ratio for students with the two lowest grades was over 90% by operation-based recommendation. This is consistent with our intuition that low-performing students are more likely to improve.
Hirokazu Kohama, Yuki Ban, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Akitoshi Itai, Hiroyasu Usami
ICCE4
2023 This Looks Like It Rather Than That: ProtoKNN For Similarity-Based Classifiers
Yuki Ukai, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
ICLR3
2023 Visual Explanation for Cooperative Behavior in Multi-Agent Reinforcement Learning
abstract
Multi-agent reinforcement learning (MARL) can acquire cooperative behavior among agents by training multiple agents in the same environment. Therefore, it is expected to be applied to complex tasks in real environments, such as traffic signal control in a traffic environment and cooperative behavior of robots. In this study, using the multi-actor-attention-critic (MAAC) with the actor-critic method as a basis, we introduce an attention head for the actor that calculates the agent's action. In contrast to the critic in MAAC, which shares the attention head among all the agents, the attention head of the actor in our method is constructed independently for each agent. This allows the attention head of the actor to calculate actor-attention (indicating which other agents are gazed at by each agent) and to acquire cooperative behavior. We visualize actor-attention to analyze the basis of agents’ decisions for cooperative behavior. Using single_spread, which is a multi-agent environment for cooperative problems, we show that the basis of decisions for cooperative behavior can be easily analyzed. We also demonstrate that our method efficiently obtains cooperative behavior considering other agents through quantitative evaluation of the cooperative behavior.
Hidenori Itaya, Tom Sagawa, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
IJCNN4
2023 Analyzing the Accuracy, Representations, and Explainability of Various Loss Functions for Deep Learning
abstract
Deep learning utilizes a vast amounts of training data and updates weight parameters so as to minimize the loss between a predicted probability and a ground truth label. Generally, we use cross-entropy as the loss function. Although loss functions for image classification other than cross-entropy exist, their efficacy has not been adequately investigated. In this work, we extensively analyze models trained with different loss functions and clarify the properties of each. Specifically, we analyze the feature space and explainability as well as the classification accuracy on various benchmark datasets and network architectures. For feature space and explainability, we investigate the effectiveness of each loss function by quantitative and qualitative evaluations. We then discuss the properties and improvements of each.
Tenshi Ito, Hiroki Adachi, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
IJCNN4
2022 Visual Explanation Generation Based on Lambda Attention Branch Networks
Tsumugi Iida, Takumi Komatsu, Kanta Kaneda, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura
ACCV (2)5
2022 Few-shot Adaptive Object Detection with Cross-Domain CutMix
Yuzuru Nakamura, Yasunori Ishii, Yuki Maruyama, Takayoshi Yamashita
ACCV (6)4
2022 Deep Ensemble Learning by Diverse Knowledge Distillation for Fine-Grained Object Classification
Naoki Okamoto, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
ECCV (11)3
2022 Rethinking Generic Camera Models for Deep Single Image Camera Calibration to Recover Rotation and Fisheye Distortion
Nobuhiko Wakai, Satoshi Sato, Yasunori Ishii, Takayoshi Yamashita
ECCV (18)4
2022 Class-Wise FM-NMS for Knowledge Distillation of Object Detection
abstract
The trade-off between accuracy and speed for an object detection model is important. When we implement an object detection model in embedded devices, a lightweight model can accelerate the detection speed. Meanwhile, the detection accuracy will be decreased. In this paper, we propose a knowledge distillation method for a lightweight object detection model. The proposed method introduces an improved feature map novel non-maximum suppression (FM-NMS) method. The improved FM-NMS uses different focus size with respect to each object class, which can suppress false positives and improve detection accuracy. In our experiments, we use onestage object detection methods, YOLOv4 as a teacher model and YOLOv4-tiny as a student model, and we apply the proposed method to them. The experimental results demonstrate that the proposed method improves the detection accuracy of the student model while maintaining the lightweight model size.
Lyuzhuang Liu, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
ICIP3
2022 Refining Design Spaces in Knowledge Distillation for Deep Collaborative Learning
abstract
Knowledge distillation is one of the most widely utilized methods to improve the performance of a model. The knowledge transfer graph has been proposed for deep collaborative learning that enables a rich diversity of bidirectional knowledge distillation. However, exploring a knowledge transfer graph is difficult due to the many potential combinations it can have, so it is not clear how accurate the resultant graphs will actually be. To address this issue, we propose a method for designing the search space with step by step and analyze the trends of graphs to design graphs with high accuracy on the basis of the acquired results. Experiments on the CIFAR-100 dataset show that we confirm that the accuracy of the best knowledge transfer graph in the search space is better than that derived using the asynchronous successive halving algorithm. We also demonstrate that the explored knowledge transfer graphs can be transferred to different datasets.
Sachi Iwata, Soma Minami, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
ICPR4
2022 Action Spotting in Soccer Videos Using Multiple Scene Encoders
abstract
Action spotting, which temporally localizes specific actions in a video, is an important task for understanding high-level semantic information. In this paper, we formulate the action spotting task to one of scene sequence recognition and propose a model with multiple scene encoders to capture scene changes around the timestamp where an action occurs. We divide the input into multiple subsets to reduce the influence of scene context that is temporally distant, and feed every subset into a scene encoder to learn scene context in every subset. Because the optimal temporal length for time windows (chunks) is different for each action, we analyze the influence of chunk sizes for action spotting. The experimental results on the public SoccerNet-v2 dataset demonstrate state-of-the-art accuracy. By using embedding features, our method obtains an Average-mAP of 75.3%. In addition, we confirm that the performance can be improved by using optimal chunk sizes for different actions.
Yuzhi Shi, Hiroaki Minoura, Takayoshi Yamashita, Tsubasa Hirakawa, Hironobu Fujiyoshi, Mitsuru Nakazawa, Yeongnam Chae, Björn Stenger
ICPR3
2022 Solving the Deadlock Problem with Deep Reinforcement Learning Using Information from Multiple Vehicles
abstract
Autonomous driving system controls a vehicle using path planning. Path planning for automated vehicles observes a vehicle and the surrounding information and plans a trajectory on the basis of rule-based approach. However, the rule-based path planning cannot generate an appropriate trajectory for complex scenes, such as two vehicles passes each other at an intersection without traffic lights. Such complex scene is called deadlock. For avoiding the deadlock, it is very costly to create rules manually. In this paper, we propose a multi-agent deep reinforcement learning method to generate appropriate trajectories at the deadlock scenes. The proposed method consists of a single feature extractor and actor-critic branches. Moreover, we introduce a mask-attention mechanism for visual explanation. By taking a look at the obtained attention maps, we can confirm the obtained agent and the reason of the behavior. For evaluating our method, we develop a simulator environment of autonomous driving that produces a certain deadlock scene. The experimental results with the developed environment show that the proposed method can generate trajectories avoiding deadlocks.
Tsuyoshi Goto, Hidenori Itaya, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
IV4
2021 Performance Prediction and Importance Analysis Using Transformer
Akiyoshi Satake, Hironobu Fujiyoshi, Takayoshi Yamashita
ICCE3
2021 Semantic Segmentation And Change Detection By Multi-Task U-Net
abstract
Change detection involves extracting the changed regions from images taken of the same place at different times. Potential applications are automatically updating of HD maps or identifying damages caused by natural disasters. However, conventional change detection methods merely detect changed regions without classifying them. In this paper, we propose a change detection method that can estimate the object class of a changed region. Our method extends a U-Net as a multi-task learning framework and estimates changed regions and semantic segmentation simultaneously. We propose using the pixel-wise classification probabilities of semantic segmentation for detecting changed regions rather than the conventional L2 norm-based difference of feature maps. In our experiments, we show that our method can improve change detection performance and estimate the classes of corresponding changed objects.
Shungo Tsutsui, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
ICIP3
2021 Visual Explanation using Attention Mechanism in Actor-Critic-based Deep Reinforcement Learning
abstract
Deep reinforcement learning (DRL) has great potential for acquiring the optimal action in complex environments such as games and robot control. However, it is difficult to analyze the decision-making of the agent, i.e., the reasons it selects the action acquired by learning. In this work, we propose Mask-Attention A3C (Mask A3C), which introduces an attention mechanism into Asynchronous Advantage Actor-Critic (A3C), which is an actor-critic-based DRL method, and can analyze the decision-making of an agent in DRL. A3C consists of a feature extractor that extracts features from an image, a policy branch that outputs the policy, and a value branch that outputs the state value. In this method, we focus on the policy and value branches and introduce an attention mechanism into them. The attention mechanism applies a mask processing to the feature maps of each branch using mask-attention that expresses the judgment reason for the policy and state value with a heat map. We visualized mask-attention maps for games on the Atari 2600 and found we could easily analyze the reasons behind an agent's decision-making in various game tasks. Furthermore, experimental results showed that the agent could achieve a higher performance by introducing the attention mechanism.
Hidenori Itaya, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura
IJCNN3
2021 Iterative Coarse-to-Fine 6D-Pose Estimation Using Back-propagation
abstract
We propose a 6D pose estimation method for an object from a single RGB image for a robotic grasping task. Many approaches estimate pose parameters from images taken from other viewpoints and use deep learning to achieve high accuracy. However, most of these methods are not robust to changes in object texture, and there is a possibility that the correct pose cannot be estimated by only one-time inference. Our aims are to reduce the number of failure cases and improve the accuracy by a novel architecture using the iterative backpropagation of a pose decoder network and pose estimation on intermediate representation. The error between random and target pose parameters are backpropagated to a neural network and the gradient for approaching the target pose is obtained. The pose parameter is updated using the obtained gradient, the error is calculated again, and backpropagation is re-performed. Repeating this process, we estimate a more accurate pose. Experiments using our own dataset show that estimation accuracy is improved and the number of failure cases is reduced. Furthermore, estimation by coarse-to-fine iterative processing is more accurate and faster. We also experiment with grasping using a UR5 robot and show that the robot can grasp objects without depth information when using the pose estimated by the proposed method.
Ryosuke Araki, Kousuke Mano, Tadanori Hirano, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
IROS5
2021 Image Captioning for Near-Future Events from Vehicle Camera Images and Motion Information
abstract
Image captioning is a task to generate a sentence explaining an input image. In autonomous driving, image captioning is expected for providing linguistic explanations of autonomous driving control decision-making because it can reduce the psychological burden on passengers and prevent accidents. Current image-captioning methods are limited to generating a caption for an input image and not generating captions for events in the near future. It is important to generate captions for any event that will happen in the near future to prevent accidents and alert passengers. Therefore, we created a task to generate an explanatory sentence of near-future events using images observed from past to present. For this task, we propose a near-future image-captioning method suitable for in-vehicle camera images. Our experiments using the Berkeley Deep Drive eXplanation Dataset showed that the proposed method can appropriately generate captions for near-future events.
Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
IV3
2020 Knowledge Transfer Graph for Deep Collaborative Learning
Soma Minami, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
ACCV (4)3
2020 Spatial Temporal Attention Graph Convolutional Networks with Mechanics-Stream for Skeleton-Based Action Recognition
Katsutoshi Shiraki, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
ACCV (5)3
2020 Improving reliability of attention branch network by introducing uncertainty
abstract
Convolutional neural networks (CNNs) are being used in various fields related to image recognition and are achieving high recognition accuracy. However, most existing CNNs do not consider uncertainty in their predictions; that is, they do not account for the difficulty of prediction, and the extent to which their predictions are reliable is unclear. This problem is considered to be the cause of erroneous decisions when we use CNNs in practice. By considering the uncertainty of the prediction result, it is thought that recognition accuracy would improve, and erroneous decisions would be suppressed. We propose a Bayesian attention branch network (Bayesian ABN) that incorporates uncertainty into an attention branch network (ABN). The method incorporates a Bayesian neural network (Bayesian NN) into the ABN to account for uncertainty in the prediction result. Also, it outputs prediction results from two branches and chooses the one having the lower uncertainty. In evaluations using standard object recognition datasets, we confirmed that the proposed method improves the accuracy and reliability of CNNs.
Takuya Tsukahara, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
ICPR3
2020 MT-DSSD: Deconvolutional Single Shot Detector Using Multi Task Learning for Object Detection, Segmentation, and Grasping Detection
abstract
This paper presents the multi-task Deconvolutional Single Shot Detector (MT-DSSD), which runs three tasks-object detection, semantic object segmentation, and grasping detection for a suction cup-in a single network based on the DSSD. Simultaneous execution of object detection and segmentation by multi-task learning improves the accuracy of these two tasks. Additionally, the model detects grasping points and performs the three recognition tasks necessary for robot manipulation. The proposed model can perform fast inference, which reduces the time required for grasping operation. Evaluations using the Amazon Robotics Challenge (ARC) dataset showed that our model has better object detection and segmentation performance than comparable methods, and robotic experiments for grasping show that our model can detect the appropriate grasping point.
Ryosuke Araki, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
ICRA4
2020 Video Object Detection and Tracking based on Angle Consistency between Motion and Flow
Toshiki Seo, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
IV3
2019 Attention Branch Network: Learning of Attention Mechanism for Visual Explanation
abstract
Visual explanation enables humans to understand the decision making of deep convolutional neural network (CNN), but it is insufficient to contribute to improving CNN performance. In this paper, we focus on the attention map for visual explanation, which represents a high response value as the attention location in image recognition. This attention region significantly improves the performance of CNN by introducing an attention mechanism that focuses on a specific region in an image. In this work, we propose Attention Branch Network (ABN), which extends a response-based visual explanation model by introducing a branch structure with an attention mechanism. ABN can be applicable to several image recognition tasks by introducing a branch for the attention mechanism and is trainable for visual explanation and image recognition in an end-to-end manner. We evaluate ABN on several image recognition tasks such as image classification, fine-grained recognition, and multiple facial attribute recognition. Experimental results indicate that ABN outperforms the baseline models on these image recognition tasks while generating an attention map for visual explanation. Our code is available.
Hiroshi Fukui, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
CVPR3
2019 Fast and Precise Detection of Object Grasping Positions with Eigenvalue Templates
abstract
Fast Graspability Evaluation (FGE) has been proposed as a method for detecting grasping positions on objects and is now being used for industrial robots. FGE uses convolution of hand templates with regions on the target object to estimate the optimum grasping posture. However, the hand opening width and rotation angles must be set with high resolution to achieve highly accurate results and the computational load is high. To address that issue, we propose a method in which hand templates are represented in compact form for faster processing by using singular value decomposition. Applying singular value decomposition enables hand templates to be represented as linear combinations of a small number of eigenvalue templates and eigenfunctions. Eigenfunctions take discrete values, but response values can be calculated with arbitrary parameters by fitting a continuous function. Experimental results show that the proposed method reduces computation time by two thirds while maintaining the same detection accuracy as conventional FGE for both parallel hands and three-finger hands.
Kousuke Mano, Takahiro Hasegawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Yukiyasu Domae
ICRA3
2019 Detecting layered structures of partially occluded objects for bin picking
abstract
When robots engage in bin picking of multiple objects, a failure in grasping partially occluded objects may occur because other objects may overlap the desired ones. Therefore, the layered structure of objects needs to be detected, and the picking order needs to be established. In this paper, we propose a new dataset that evaluates not only the area of objects but also the layered structures of objects. In this dataset, three tasks are targeted: object detection, semantic segmentation, and segmentation of occluded areas for bin picking of multiple objects. The dataset, called “the Amazon Robotics Challenge (ARC) Multi-task Dataset” contains 1,500 RGB images and depth images, including all scenes containing bounding box labels, semantic segmentation labels, and occluded area labels. This enables representing the layered structure of overlapped objects with a tree structure. A benchmark of the ARC multi-task dataset demonstrated that occluded areas could be segmented using a Mask regional convolutional neural network (R-CNN) and that layered structures of objects could be predicted. Our dataset is available at the following URL:http://mprg.jp/research/arc_dataset_2017_e.
Yusuke Inagaki, Ryosuke Araki, Takayoshi Yamashita, Hironobu Fujiyoshi
IROS3
2019 Visual Explanation by Attention Branch Network for End-to-end Learning-based Self-driving
abstract
Self-driving decides an appropriate control considering the surrounding environment. To this end, self-driving control methods by using a convolutional neural network (CNN) have been studied, which directly input the vehicle-mounted camera image to a network and output a steering directory. However, if we need to control not only steering but also throttle, it is necessary to grasp the state of the car itself in addition to the surrounding environment. Moreover, in order to use CNNs for critical applications such as self-driving, it is important to analyze where the network focuses on the image and to understand the decision making. In this work, we propose a method to solve these problems. First, to control both steering and throttle simultaneously, we propose using the current vehicle speed as the state of the car itself. Second, we introduce an attention branch network (ABN) architecture to a self-driving model, which enables visually analyzing the reason of the self-driving decision making by using an attention map. Experimental results with a driving simulator demonstrate that our method controls a car stably, and we can analyze the decision making by using the attention map.
Keisuke Mori, Hiroshi Fukui, Takuya Murase, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi
IV5
2018 Compactification of Affine Transformation Filter Using Tensor Decomposition
abstract
Keypoint matching is used in a variety of tasks such as specific object recognition and panoramic image generation. Affine-SIFT (ASIFT) enables affine invariant matching by generating many affine transformation images of an input image. It describes the scale-invariant feature transform (SIFT) features of the generated image. However, ASIFT must perform multiple costly online computations for affine transformation. We represent an oriented FAST and rotated BRIEF (ORB) descriptor in a linear filter subjected to many affine transformations. We calculate the affine features by convolving the generated filter with the patch image. However, convolving the 19,200 filters generated by the affine transformation is inefficient. In order to reduce the convolution processing, the affine transformation filter is made compact by a factorization method. We built a 4-order tensor using the affine transformation filter. The 4-order tensor decomposes into the Tucker model. We reduce dimensions appropriately for each mode. In this way, we propose a compact and accurate feature description. Our evaluation experiments confirmed that the proposed method reduces the processing time to 19% while maintaining the same precision as singular value decomposition, which is the conventional method.
Kohei Kawai, Takahiro Hasegawa, Yuji Yamauchi, Takayoshi Yamashita, Hironobu Fujiyoshi
ICIP4
2018 Multiple Skip Connections of Dilated Convolution Network for Semantic Segmentation
abstract
Semantic segmentation is a task to estimate class for each pixel. This task also have received benefit from the deep ConvNet and it has achieves high accuracy. In the semantic segmentation from in-vehicle camera image, the object size such as a pedestrian or a vehicle fluctuates according to the distance from the camera. We propose scale aware semantic segmentation method especially small object. The contributions of the method are 1) to feed the features of small region by multiple skip connections, 2) to extract context from multiple receptive field by multiple dilated convolution blocks. The proposed method has achieved high accuracy in the Cityscapes dataset. The comparison with state-of-the-art methods, it has achieved the comparable performance at category IoU and iIoU metrics.
Takayoshi Yamashita, Hironori Furukawa, Hironobu Fujiyoshi
ICIP1
2017 Accelerating Computation of Exemplar-SVM by Binary Approximation based on Matrix Decomposition
Takato Kurokawa, Yuji Yamauchi, Mitsuru Ambai, Takayoshi Yamashita, Hironobu Fujiyoshi
BMVC4
2017 Students' Performance Prediction Using Data of Multiple Courses by Recurrent Neural Network
Fumiya Okubo, Takayoshi Yamashita, Atsushi Shimada 0001, Shin'ichi Konomi
ICCE2
2017 Word Recognition by Combining Outline Emphasis and Synthesize Background
Yukihiro Achiha, Takayoshi Yamashita, Mitsuru Nakazawa, Soh Masuko, Yuji Yamauchi, Hironobu Fujiyoshi
ICEC2
2017 A neural network approach for students' performance prediction
abstract
In this paper, we propose a method for predicting final grades of students by a Recurrent Neural Network (RNN) from the log data stored in the educational systems. We applied this method to the log data from 108 students and examined the accuracy of prediction. From the experimental results, comparing with multiple regression analysis, it is confirmed that an RNN is effective to early prediction of final grades.
Fumiya Okubo, Takayoshi Yamashita, Atsushi Shimada 0001, Hiroaki Ogata
LAK2
2016 Pedestrian and part position detection using a regression-based multiple task deep convolutional neural network
abstract
In driving support systems, it is not only necessary to detect the position of pedestrians, but also to estimate the distance between a pedestrian and the vehicle. In general approaches using monocular cameras, the upper and lower positions of each pedestrian are detected using a bounding box obtained from a pedestrian detection technique. The distance between the pedestrian and the vehicle is then estimated using these positions and the camera parameters. This conventional framework uses independent pedestrian detection and position detection processes to estimate the distance. In this paper, we propose a method to detect both the pedestrian and their position simultaneously using a regression-based deep convolutional neural network (DCNN). This simultaneous detection method is possible to train efficient features for both tasks, because it is attention to head and leg regions from given labels. In the experiments, our method improves the performance of pedestrian detection compared with the DCNN which detects only pedestrian. The proposed approach also improves the detection accuracy of the head and leg positions compared with the methods that detect only these positions. Using the results of position detection and the camera parameters, our method achieves distance estimation to within 5% error.
Takayoshi Yamashita, Hiroshi Fukui, Yuji Yamauchi, Hironobu Fujiyoshi
ICPR1
2016 Design of a Low-false-positive Gesture for a Wearable Device
abstract
As smartwatches are becoming more widely used in society, gesture recognition, as an important aspect of interaction with smartwatches, is attracting attention. An accelerometer that is incorporated in a device is often used to recognize gestures. However, a gesture is often detected falsely when a similar pattern of action occurs in daily life. In this paper, we present a novel method of designing a new gesture that reduces false detection. We refer to such a gesture as a low-false-positive (LFP) gesture. The proposed method enables a gesture design system to suggest LFP motion gestures automatically. The user of the system can design LFP gestures more easily and quickly than what has been possible in previous work. Our method combines primitive gestures to create an LFP gesture. The combination of primitive gestures is recognized quickly and accurately by a random forest algorithm using our method. We experimentally demonstrate the good recognition performance of our method for a designed gesture with a high recognition rate and without false detection.
Ryo Kawahata, Atsushi Shimada 0001, Takayoshi Yamashita, Hideaki Uchiyama, Rin-Ichiro Taniguchi
ICPRAM3
2016 Robust pedestrian attribute recognition for an unbalanced dataset using mini-batch training with rarity rate
abstract
Pedestrian attributes are significant information for Advanced Driver Assistance System(ADAS). Pedestrian attributes such as body poses, face orientations and open umbrella are meant action or state of pedestrian. In general, this information is recognized using independent classifiers for each task. Performing all of these separate tasks is too time-consuming at the testing stage. In addition, the processing time increases with the number of tasks. To address this problem, multi-task learning or heterogeneous learning is able to train a single classifier to perform multiple tasks. In particular, heterogeneous learning is able to simultaneously train regression and recognition tasks, because reducing both training and testing time. However, heterogeneous learning tends to result in a lower accuracy rate for classes with a few training samples. In this paper, we propose a method to improve the performance of heterogeneous learning for such classes. We introduce a rarity rate based on the importance and class probability of each task. The appropriate rarity rate is assigned to each training sample. Thus, the samples in a mini-batch for training a deep convolutional neural network are augmented by this rarity rate to focus on the class with a few samples. Our heterogeneous learning approach with the rarity rate attains better performance on pedestrian attribute recognition, especially for classes representing open umbrellas.
Hiroshi Fukui, Takayoshi Yamashita, Yuji Yamauchi, Hironobu Fujiyoshi, Hiroshi Murase
Intelligent Vehicles Symposium2
2015 Multiple-Hypothesis Affine Region Estimation with Anisotropic LoG Filters
abstract
We propose a method for estimating multiple-hypothesis affine regions from a keypoint by using an anisotropic Laplacian-of-Gaussian (LoG) filter. Although conventional affine region detectors, such as Hessian/Harris-Affine, iterate to find an affine region that fits a given image patch, such iterative searching is adversely affected by an initial point. To avoid this problem, we allow multiple detections from a single keypoint. We demonstrate that the responses of all possible anisotropic LoG filters can be efficiently computed by factorizing them in a similar manner to spectral SIFT. A large number of LoG filters that are densely sampled in a parameter space are reconstructed by a weighted combination of a limited number of representative filters, called "eigenfilters", by using singular value decomposition. Also, the reconstructed filter responses of the sampled parameters can be interpolated to a continuous representation by using a series of proper functions. This results in efficient multiple extrema searching in a continuous space. Experiments revealed that our method has higher repeatability than the conventional methods.
Takahiro Hasegawa, Mitsuru Ambai, Kohta Ishikawa, Gou Koutaki, Yuji Yamauchi, Takayoshi Yamashita, Hironobu Fujiyoshi
ICCV6
2015 Facial point detection based on a convolutional neural network with optimal mini-batch procedure
abstract
We propose a Convolutional Neural Network (CNN)-based method to ensure both robustness to variations in facial pose and real-time processing. Although the robustness of CNNs has attracted attention in various fields, the training process suffers from difficulties in parameter setting and the manner in which training samples are provided. We demonstrate a manner of providing samples that results in a better network. We consider four methods: 1) subset with augmentation, 2) random selection, 3) fixed-person subset, and 4) the conventional approach. Experimental results indicate that the subset with augmentation technique has sufficient variations and quantity to obtain the best performance. Our CNN-based method is robust under facial pose variations, and achieves better performance. In addition, since our networks structure is simple, processing takes approximately 10ms for one face on a standard CPU.
Masatoshi Kimura, Takayoshi Yamashita, Yuji Yamauchi, Hironobu Fujiyoshi
ICIP2
2015 SWAP-NODE: A regularization approach for deep convolutional neural networks
abstract
The regularization is important for training of a deep network. One of breakthrough approach is dropout. It randomly deletes a certain number of activations in each layer in the feed-forward step of the training process. The dropout significantly reduces an effect of over-fitting and improves test performance. We introduce a new regularization approach for deep learning, called the swap-node. The swap-node, which is applied to a fully connected layer, swaps the activation values of two nodes randomly selected with a certain probability. Empirical evaluation shows that the network using the swap-node performs the best on MNIST, CIFAR-10, and SVHN. We also demonstrate superior performance of a combination of the swap-node and dropout on these datasets.
Takayoshi Yamashita, Masayuki Tanaka 0001, Yuji Yamauchi, Hironobu Fujiyoshi
ICIP1
2015 Facial point detection using convolutional neural network transferred from a heterogeneous task
abstract
We present a novel training approach that uses convolutional neural network for facial part detection. In the proposed training procedure, we use the parameters of a network obtained for a heterogeneous task as the initial parameters of the network for a target task. We employ a convolutional neural network for facial part labeling in the heterogeneous task, and then transfer the trained network so as to provide initial parameters of the network for facial point detection. This transfer of network is advantageous in the training for a target task in that 1) the network obtains representation kernels for extraction of facial part regions and 2) the network reduces detection errors at distant positions. The performance of the proposed method applied to BioID and Labeled Face Parts in the Wild datasets is comparable to that of state-of-the art methods. In addition, since our network structure is simple, processing takes approximately 3ms for one face on a standard CPU.
Takayoshi Yamashita, Taro Watasue, Yuji Yamauchi, Hironobu Fujiyoshi
ICIP1
2015 Fast 3D edge detection by using decision tree from depth image
abstract
T3D edge detection from a depth image is an important technique of 3D object recognition in preprocessing. There are three types of 3D edges in a depth image called jump, convex roof, and concave roof edges. Conventional 3D edge detection based on ring operators has been proposed. The conventional ring operator can detect three types of 3D edges by classifying the response of Fourier transforms. Since the conventional method needs to apply Fourier transforms to all pixels of a depth image, real-time processing cannot be done due to high computational cost. Therefore, this paper presents a fast and reliable method of detecting three types of 3D edges by using a decision tree. The decision tree is trained under supervised learning from numerous synthesized depth images and labels by capturing depth relations between candidate pixels and pixels on a ring operator to classify 3D edges. The experimental results revealed that the proposed method has 25 times faster than the conventional method. This paper also presents some examples of 3D line and 3D convex corner detection based on results obtained with the proposed method.
Masaya Kaneko, Takahiro Hasegawa, Yuji Yamauchi, Takayoshi Yamashita, Hironobu Fujiyoshi, Hiroshi Murase
IROS4
2015 Pedestrian detection based on deep convolutional neural network with ensemble inference network
abstract
Pedestrian detection is an active research topic for driving assistance systems. To install pedestrian detection in a regular vehicle, however, there is a need to reduce its cost and ensure high accuracy. Although many approaches have been developed, vision-based methods of pedestrian detection are best suited to these requirements. In this paper, we propose the methods based on Convolutional Neural Networks (CNN) that achieves high accuracy in various fields. To achieve such generalization, our CNN-based method introduces Random Dropout and Ensemble Inference Network (EIN) to the training and classification processes, respectively. Random Dropout selects units that have a flexible rate, instead of the fixed rate in conventional Dropout. EIN constructs multiple networks that have different structures in fully connected layers. The proposed methods achieves comparable performance to state-of-the-art methods, even though the structure of the proposed methods are considerably simpler.
Hiroshi Fukui, Takayoshi Yamashita, Yuji Yamauchi, Hironobu Fujiyoshi, Hiroshi Murase
Intelligent Vehicles Symposium2
2014 Asymmetric Feature Representation for Object Recognition in Client Server System
Yuji Yamauchi, Mitsuru Ambai, Ikuro Sato, Yuichi Yoshida, Hironobu Fujiyoshi, Takayoshi Yamashita
ACCV (1)6
2014 Hand posture recognition based on bottom-up structured deep convolutional neural network with curriculum learning
abstract
Hand posture recognition has tremendous potential in the field of natural user interactions. There were many advances in research in recent years but there are still limitations regarding its usage in unfavorable live situations where hand posture variation, illumination change or background complexity are an issue. In cases like these, recognizing the hand posture is a difficult task. As such, we considered reducing the difficulty of the task by using curriculum learning with intermediate information. We proceeded to divide the complex architecture of the hand posture recognition task into two easier ones: 1) Extraction of the hand shape under clutter background with illumination change, 2) Recognition of the hand posture from a binary image. In order to do so, we propose here a bottom-up structured deep convolutional neural network incorporating a special layer for binary image extraction. Our proposed method also employs state-of-the art techniques for deep learning to obtain generalization. As a result, we achieved better recognition performances of the hand posture under clutter background compared to the baseline method.
Takayoshi Yamashita, Taro Watasue
ICIP1
2014 To Be Bernoulli or to Be Gaussian, for a Restricted Boltzmann Machine
abstract
We introduce a method that automatically selects appropriate RBM types according to the visible unit distribution. The distribution of a visible unit strongly depends on a dataset. For example, binary data can be considered as pseudo binary distribution with high peaks at 0 and 1. For real-value data, the distribution can be modeled by single Gaussian model or Gaussian mixture model. Our proposed method selects appropriate RBM according to the distribution of each unit. We employ the Gaussian mixture model to determine whether the visible unit distribution is the pseudo binary or the Gaussian mixre. According to this distribution, we can select a Bernoulli-Bernoulli RBM(BBRBM) or a Gaussian-Bernoulli RBM(GBRBM). Furthermore, we employ normalization process to obtain a smoothed Gaussian mixture distribution. This allowed us to reduce variations such as illumination changes in the input data. After experimentation with MNIST, CBCL and our own dataset, our proposed method obtained the best recognition performance and further shortened the convergence time of the learning process.
Takayoshi Yamashita, Masayuki Tanaka 0001, Eiji Yoshida, Yuji Yamauchi, Hironobu Fujiyoshi
ICPR1
2013 Robust object tracking using a context based on the relation of object and background
abstract
Object tracking is complicated by perspective changes to both the object and background caused by object and the camera motion, object-to-object and object-to-background occlusion and illumination changes. Conventional object tracking method focus on distinguishing the target from the background. Adopting a new perspective, we propose a context aware tracking by a collaborative model of both the object and its surrounding background. In this model we introduce the “probability of tracking failure” that determines the feature similarity and the spatial relationship of the target and the surrounding background, as the “target-surrounding context”, which appears effective to predict the likelihood of a tracking failure. This target-surrounding context can be used to prevent tracking failure where the background has similar objects. In the experimental of the scene which occurs occlusion by similar objects, the proposed method outperformed most of the conventional methods.
Takayoshi Yamashita, Hironobu Fujiyoshi
ICASSP1
2012 Adaptation of boosted pedestrian detectors by feature reselection
abstract
Adaptation of pre-trained boosted pedestrian detectors to specific scenes is an important yet difficult task in computer vision. To address this problem, a feature reselection strategy is proposed in this paper. The proposed method identifies weak classifiers which do not well adapt to the specific scene, and replaces them with retrained weak classifiers. This feature reselection strategy has the following advantages: 1) it does not need original offline training data, but only uses a few online samples from the target scene; 2) the adapted detector preserves the generality of the generic detector, resulting in very few false positives; and 3) it can adapt a generic detector to a specific scene with very fast speed due to its parallel nature. Experiments on challenging pedestrian detection datasets demonstrate that our proposed strategy can significantly improve the performance of pre-trained boosted detectors in specific scenes with very low computation cost and very little labeling work.
Genquan Duan, Haizhou Ai, Takayoshi Yamashita
ICIP4
2012 How to Select Useful Hand Shapes for Hand Gesture Recognition System
Atsushi Shimada 0001, Takayoshi Yamashita, Rin-Ichiro Taniguchi
ICPRAM (2)2
2011 Robust Contour Tracking by Combining Region and Boundary Information
abstract
This paper presents a new object tracking model that systematically combines region and boundary features. Besides traditional region features (intensity/color and texture), we design a new boundary-based object detector for accurate and robust tracking in low-contrast and complex scenes, which usually appear in the commonly used monochrome surveillance systems. In our model, region feature-based energy terms are characterized by probability models, and boundary feature terms include edge and frame difference. With a new weighting term, a novel energy functional is proposed to systematically combine the region and boundary-based components, and it is minimized by a level set evolution equation. For an efficient computational cost, motion information is utilized for new frame level set initialization. Compared with region feature-based models, the experimental results show that the proposed model significantly improves the performance under different circumstances, especially for objects in low-contrast and complex environments.
Ling Cai 0003, Takayoshi Yamashita, Yiren Xu, Xin Yang 0007
IEEE Trans. Circuits Syst. Video Technol.3
2010 Human Pose Estimation Using Exemplars and Part Based Refinement
Yanchao Su, Haizhou Ai, Takayoshi Yamashita, Shihong Lao
ACCV (2)3
2010 Combined Top-Down/Bottom-Up Human Articulated Pose Estimation Using AdaBoost Learning
abstract
In this paper, a novel human articulated pose estimation method based on AdaBoost algorithm is presented. The human articulated pose is estimated by locating major human joint positions. We learn the classifiers on a normalized image for classifying each pixel position into a certain category. Two different kinds of classifiers, bottom-up joint position classifier and top-down skeleton classifier, are combined to achieve final results. HOG (Histogram of Oriented Gradient) feature is used for training both classifiers. Our human pose estimation system consists of three models, human detection, view classification, and pose estimation. The implemented system can automatically estimate human pose of different views. Experiment results are reported to show our proposed method can work on relatively small-size human images without using human silhouettes as a prerequisite, which is very efficient, robust and accurate enough for potential applications in visual surveillance.
Haizhou Ai, Takayoshi Yamashita, Shihong Lao
ICPR3
2009 Towards Robust Object Detection: Integrated Background Modeling Based on Spatio-temporal Features
Tatsuya Tanaka, Atsushi Shimada 0001, Rin-Ichiro Taniguchi, Takayoshi Yamashita, Daisaku Arita
ACCV (1)4
2009 SURF Tracking
abstract
Most motion-based tracking algorithms assume that objects undergo rigid motion, which is most likely disobeyed in real world. In this paper, we present a novel motion-based tracking framework which makes no such assumptions. Object is represented by a set of local invariant features, whose motions are observed by a feature correspondence process. A generative model is proposed to depict the relationship between local feature motions and object global motion, whose parameters are learned efficiently by an on-line EM algorithm. And the object global motion is estimated in term of maximum likelihood of observations. Then an updating mechanism is employed to adapt object representation. Experiments show that our framework is flexible and robust in dealing with appearance changes, background clutter, illumination changes and occlusion.
Takayoshi Yamashita, Shihong Lao
ICCV2
2008 Human tracking based on Soft Decision Feature and online real boosting
abstract
Online Boosting is an effective incremental learning method which can update weak classifiers efficiently according to the object being trackedt. It is a promising technique for online object tracking to adapt tothe appearance variations of objects during tracking process. However, proposed online-boosting based tracking methods update and select weak classifiers from fixed the offline learned weak classifiers, which might not be an optimal selection for object appearance variations. In this paper, we propose a new feature adjusting strategy for online boosting called Soft Decision Feature. We combine it with online real AdaBoost to achieve better tracking performance in scenes with human pose and posture variations. Experiment result demonstrates that it can successfully deal with the human posture variation scenes that conventional online boosting tracking methods fails to deal with.
Takayoshi Yamashita, Hironobu Fujiyoshi, Shihong Lao, Masato Kawade
ICPR1
2008 Tracking in Low Frame Rate Video: A Cascade Particle Filter with Discriminative Observers of Different Life Spans
abstract
Tracking objects in low frame rate (LFR) video or with abrupt motion poses two main difficulties which most conventional tracking methods can hardly handle: 1) poor motion continuity and increased search space; 2) fast appearance variation of target and more background clutter due to increased search space. In this paper, we address the problem from a view which integrates conventional tracking and detection, and present a temporal probabilistic combination of discriminative observers of different lifespans. Each observer is learned from different ranges of samples, with different subsets of features, to achieve varying levels of discriminative power at varying cost. An efficient fusion and temporal inference is then done by a cascade particle filter which consists of multiple stages of importance sampling. Experiments show significantly improved accuracy of the proposed approach in comparison with existing tracking methods, under the condition of LFR data and abrupt motion of both target and camera.
Yuan Li 0022, Haizhou Ai, Takayoshi Yamashita, Shihong Lao, Masato Kawade
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 Tracking in Low Frame Rate Video: A Cascade Particle Filter with Discriminative Observers of Different Lifespans
abstract
Tracking object in low frame rate video or with abrupt motion poses two main difficulties which conventional tracking methods can barely handle: 1) poor motion continuity and increased search space; 2) fast appearance variation of target and more background clutter due to increased search space. In this paper, we address the problem from a view which integrates conventional tracking and detection, and present a temporal probabilistic combination of discriminative observers of different lifespans. Each observer is learned from different ranges of samples, with different subsets of features, to achieve varying level of discriminative power at varying cost. An efficient fusion and temporal inference is then done by a cascade particle filter which consists of multiple stages of importance sampling. Experiments show significantly improved accuracy of the proposed approach in comparison with existing tracking methods, under the condition of low frame rate data and abrupt motion of both target and camera.
Yuan Li 0022, Haizhou Ai, Takayoshi Yamashita, Shihong Lao, Masato Kawade
CVPR3
2007 Online Real Boosting for Object Tracking Under Severe Appearance Changes and Occlusion
abstract
Robust visual tracking is always a challenging but yet intriguing problem owing to the appearance variability of target objects. In this paper we propose a novel method to handle large changes in appearance based on online real-value boosting, which is utilized to incrementally learn a strong classifier to distinguish between objects and their background. By incorporating online real boosting into a particle filter framework, our tracking algorithm shows a strong adaptability for different target objects which undergo severe appearance changes during the tracking process.
Takayoshi Yamashita, Shihong Lao, Masato Kawade, Feihu Qi
ICASSP (1)2
2007 Incremental Learning of Boosted Face Detector
abstract
In recent years, boosting has been successfully applied to many practical problems in pattern recognition and computer vision fields such as object detection and tracking. As boosting is an offline training process with beforehand collected data, once learned, it cannot make use of any newly arriving ones. However, an offline boosted detector is to be exploited online and inevitably there must be some special cases that are not covered by those beforehand collected training data. As a result, the inadaptable detector often performs badly in diverse and changeful environments which are ordinary for many real-life applications. To alleviate this problem, this paper proposes an incremental learning algorithm to effectively adjust a boosted strong classifier with domain-partitioning weak hypotheses to online samples, which adopts a novel approach to efficient estimation of training losses received from offline samples. By this means, the offline learned general-purpose detectors can be adapted to special online situations at a low extra cost, and still retains good generalization ability for common environments. The experiments show convincing results of our incremental learning approach on challenging face detection problems with partial occlusions and extreme illuminations.
Chang Huang, Haizhou Ai, Takayoshi Yamashita, Shihong Lao, Masato Kawade
ICCV3