EDBT 2026 Demo / reviewers in the wild / expert
Hironobu Fujiyoshi
dblp:79/2304
· DBLP profile ↗
82ranked-venue papers
3as first author
27since 2021 · last 2025
0000-0001-7391-4725ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 1 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 2 first-author · 14 since 2021Systems, architecture and hardware · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Temporal Cues for Semi-Supervised Multi-View 3D Object DetectionabstractWhile recent advancements in camera-based 3D object detection demonstrate remarkable performance, they require thousands or even millions of human-annotated frames. This requirement significantly inhibits their deployment in various locations and sensor configurations. To address this gap, we propose a performant semi-supervised framework that leverages unlabeled RGB-only driving sequences - data easily collected with cost-effective RGB cameras - to significantly improve temporal, camera-only 3D detectors. We observe that the standard semi-supervised pseudo-labeling paradigm underperforms in this temporal, camera-only setting due to poor 3D localization of pseudo-labels. To address this, we train a single 3D detector to handle RGB sequences both forward and backward in time, then ensemble both its forwards and backwards pseudo-labels for semi-supervised learning. We further improve the pseudo-label quality by leveraging 3D object tracking to infill missing detections and by eschewing simple confidence thresholding in favor of using the auxiliary 2D detection head to filter 3D predictions. Finally, to enable the backbone to learn directly from the unlabeled data itself, we introduce an object-query conditioned masked reconstruction objective. Our framework demonstrates remarkable performance improvement on large-scale autonomous driving datasets nuScenes and nuPlan. Jinhyung Park, Navyata Sanghvi, Hiroki Adachi, Yoshihisa Shibata, Shawn Hunt, Shinya Tanaka, Hironobu Fujiyoshi, Kris Makoto Kitani |
CVPR | 7 |
| 2025 | Weight Pruning to Mitigate Class-Specific Accuracy Degradation for LiDAR-Based 3D Object DetectionabstractThe realization of autonomous driving systems requires efficient and accurate 3D object detection to identify objects such as vehicles, pedestrians, and cyclists within the driving environment using point cloud data. For achieving both high speed processing and high accuracy, it is necessary to reduce model size by model compression techniques, such as pruning, while maintaining performance. However, pruning for 3D object detection tasks has not been extensively studied, and the effects of applying existing pruning methods to 3D object detection models remain unclear. In this paper, we clarify the problems of pruning 3D object detection models with existing methods through preliminary experiments, and propose a pruning method suitable for 3D object detection models that solves these problems. Our preliminary experiments reveal that existing pruning methods significantly degrade detection performance for specific object classes. To address this issue, we propose a pruning method that preserves class-specific knowledge, mitigating biased accuracy degradation across different object classes. Experimental results on the KITTI dataset demonstrate that the proposed method can be combined with existing pruning methods without conflicts and achieves higher accuracy than existing methods. Tenshi Ito, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 4 |
| 2024 | DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention
NguyenHuu BaoLong, Yuzhi Shi, Tsubasa Hirakawa, Takayoshi Yamashita, Tohgoroh Matsui, Hironobu Fujiyoshi |
ACCV (10) | 7 |
| 2024 | Faster Convergence and Uncorrelated Gradients in Self-Supervised Online Continual Learning
Koyo Imai, Naoto Hayashi, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ACCV (8) | 5 |
| 2024 | SAM Helps SSL: Mask-guided Attention Bias for Self-supervised Learning
Kensuke Taguchi, Takehiko Kawai, Wataru Imaeda, Hironobu Fujiyoshi |
BMVC | 4 |
| 2024 | Layer-Wise Relevance Propagation with Conservation Property for ResNet
Seitaro Otsuki, Tsumugi Iida, Félix Doublet, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura |
ECCV (43) | 6 |
| 2024 | Enhancing the Accuracy of Predicting Students Grades in Open-Ended Questions through Adjustments to Attention Weights
Masaki Koike, Hirokazu Kohama, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
EDM | 5 |
| 2024 | Binary-Decomposed Vision Transformer: Compressing and Accelerating Vision Transformer by Binary DecompositionabstractVision Transformers (ViTs) have emerged as versatile and high-performance models for various tasks such as image classification, object detection, and semantic segmentation. However, the ViT-L model, which demonstrates high accuracy, has a large number of parameters (307M), leading to increased computational requirements. To deploy ViTs on embedded devices and similar platforms, it is crucial to compress the model size and accelerate the inference process. In this paper, we propose the Binary-decomposed Vision Transformer (BdViT), a method for model compression and accelerated inference for ViTs models. BdViT consists of weight binarization based on vector decomposition and quantization of multiplication and addition operations, which does not require retraining model parameters. Through evaluation experiments using image recognition datasets, we demonstrated that BdViT can significantly reduce the number of parameters while mitigating performance degradation. Ryota Kondo, Hiroaki Minoura, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICIP | 5 |
| 2024 | Human-like Guidance by Generating Navigation Using Spatial-Temporal Scene GraphabstractVehicle navigation systems use both GPS and map data, primarily information derived from map data. Conventional navigation systems assume that the user will look directly at the display to check information. Simultaneously provided text and voice often play only a supplementary role, which can lead to driver distraction and misinterpretation. In contrast, human navigation utilizes visual information, potentially reducing the cognitive load on drivers. Human-like Guidance is aimed at realizing a driving assistance system that supports navigation akin to human guidance. Implementing Human-like Guidance, requires the handling of video footage from in-vehicle cameras during vehicle operation, suggesting the need for an approach combining image recognition and language model. However, images captured during operation often contain superfluous information, making the selection of relevant objects for navigation challenging. Moreover, relying solely on image information makes it difficult to consider the relationship with surrounding objects. Therefore, this study proposes a Spatial-Temporal Scene Graph that can represent spatial and temporal information of objects from driving scene videos. Furthermore, we achieve Human-like Guidance through navigation generation using features extracted from the Spatial-Temporal Scene Graph. Our results show that our proposed method improves the accuracy of navigation generation accuracy compared to traditional image-based navigation methods. In addition, the use of a Spatial-Temporal Scene Graph enables the generation of human-like navigation that focuses on the movements of surrounding vehicle objects. Hayato Suzuki, Kota Shimomura, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Shota Okubo, Nanri Takuya |
IV | 5 |
| 2024 | High-Precision for Multi-Task Learning from In-Vehicle Camera using BiFPNabstractMulti-task learning is effective for object detection and segmentation, which are closely related to each other and necessary for automated driving. However, there is a problem with the learning process in conventional multi-task learning models. In multi-task learning, common features among downstream tasks are first extracted by a backbone network. Then, these features are used for different downstream tasks. Since the required feature is different depending on the downstream task, it is necessary to extract features suitable for each downstream task. In this paper, we propose a multi-tasking model that introduces BiFPN feature fusion method for automated driving tasks and the Next-ViT model utilizing CNN and Transformer to extract features. From the evaluation experiments of automated driving tasks, we confirmed that the proposed method improves the accuracy of multi-task learning. Hidenori Itaya, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 5 |
| 2023 | Embedding Human Knowledge into Spatio-Temproal Attention Branch Network in Video Recognition via Temporal attention
Saki Noguchi, Yuzhi Shi, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
BMVC | 5 |
| 2023 | Recommending Learning Actions Using Neural NetworkabstractMany studies applying neural networks to the field of education have focused on student performance prediction and explainability of their decisions. While those studies introduced neural networks into educational settings, such networks cannot directly support student learnings in place of teachers. Therefore, we present a method that uses a general Transformer encoder to recommend appropriate learning actions for improving student performance. By considering the attention weight of a low-performing student to be close to that of a high-performing student, our method recommends the learning materials and actions for learning the materials. To evaluate the effectiveness of our method, we trained a deep neural network (DNN) on a private dataset of student operations (e.g., NEXT, PREV, OPEN) on digital learning materials obtained from a Japanese university. The number of operations divided by each learning material and by type of operation are input to the DNN, and the DNN outputs the student’s grade on 5-point scale. We applied our method with this trained DNN to samples that successfully predicted grades, and the number of operations increased on the basis of the recommended learning materials and actions. By re-inputting modified sample into the DNN, we then observe how the student performance changes. The results of this simple experiment indicate that more students improved their performance with both the material-based and operation-based recommendations than with random recommendations. The percentage of students whose grades improved tended to be larger for those with low grades. Specifically, the improvement ratio for students with the two lowest grades was over 90% by operation-based recommendation. This is consistent with our intuition that low-performing students are more likely to improve. Hirokazu Kohama, Yuki Ban, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Akitoshi Itai, Hiroyasu Usami |
ICCE | 5 |
| 2023 | This Looks Like It Rather Than That: ProtoKNN For Similarity-Based Classifiers
Yuki Ukai, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICLR | 4 |
| 2023 | Visual Explanation for Cooperative Behavior in Multi-Agent Reinforcement LearningabstractMulti-agent reinforcement learning (MARL) can acquire cooperative behavior among agents by training multiple agents in the same environment. Therefore, it is expected to be applied to complex tasks in real environments, such as traffic signal control in a traffic environment and cooperative behavior of robots. In this study, using the multi-actor-attention-critic (MAAC) with the actor-critic method as a basis, we introduce an attention head for the actor that calculates the agent's action. In contrast to the critic in MAAC, which shares the attention head among all the agents, the attention head of the actor in our method is constructed independently for each agent. This allows the attention head of the actor to calculate actor-attention (indicating which other agents are gazed at by each agent) and to acquire cooperative behavior. We visualize actor-attention to analyze the basis of agents’ decisions for cooperative behavior. Using single_spread, which is a multi-agent environment for cooperative problems, we show that the basis of decisions for cooperative behavior can be easily analyzed. We also demonstrate that our method efficiently obtains cooperative behavior considering other agents through quantitative evaluation of the cooperative behavior. Hidenori Itaya, Tom Sagawa, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IJCNN | 5 |
| 2023 | Analyzing the Accuracy, Representations, and Explainability of Various Loss Functions for Deep LearningabstractDeep learning utilizes a vast amounts of training data and updates weight parameters so as to minimize the loss between a predicted probability and a ground truth label. Generally, we use cross-entropy as the loss function. Although loss functions for image classification other than cross-entropy exist, their efficacy has not been adequately investigated. In this work, we extensively analyze models trained with different loss functions and clarify the properties of each. Specifically, we analyze the feature space and explainability as well as the classification accuracy on various benchmark datasets and network architectures. For feature space and explainability, we investigate the effectiveness of each loss function by quantitative and qualitative evaluations. We then discuss the properties and improvements of each. Tenshi Ito, Hiroki Adachi, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IJCNN | 5 |
| 2023 | Uncertainty-Aware Interactive LiDAR Sampling for Deep Depth CompletionabstractProgrammable scan LiDAR is able to measure arbitrary areas and is expected to be used in various applications. In this paper, we study a LiDAR sampling strategy for deep depth completion of a programmable scan LiDAR with an RGB camera. General data sampling strategies include adaptive approaches such as active learning, in which candidate data are assessed through a task model for data selection and then the selected data pool is updated sequentially. Although it is an effective approach, the adaptive approach requires many iterations involving the inference process to assess the candidate data, which is not suitable for LiDAR systems. Therefore, we propose a novel interactive LiDAR sampling method without each inference process. Our key insights are that we assess sampling candidates by depth estimation uncertainty and virtually update the uncertainty by an approximation of the candidate assessment. This enables us to add interactivity to the model state without requiring each inference process. We demonstrate the effectiveness of our method on the KITTI dataset and the generalization performance on the NYU-Depth-v2 dataset in comparison with a conventional adaptive LiDAR sampling method, and we find superior results in the depth completion task. We also show ablation studies to analyze our approach. Kensuke Taguchi, Shogo Morita, Yusuke Hayashi, Wataru Imaeda, Hironobu Fujiyoshi |
WACV | 5 |
| 2022 | Visual Explanation Generation Based on Lambda Attention Branch Networks
Tsumugi Iida, Takumi Komatsu, Kanta Kaneda, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura |
ACCV (2) | 6 |
| 2022 | Deep Ensemble Learning by Diverse Knowledge Distillation for Fine-Grained Object Classification
Naoki Okamoto, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ECCV (11) | 4 |
| 2022 | Class-Wise FM-NMS for Knowledge Distillation of Object DetectionabstractThe trade-off between accuracy and speed for an object detection model is important. When we implement an object detection model in embedded devices, a lightweight model can accelerate the detection speed. Meanwhile, the detection accuracy will be decreased. In this paper, we propose a knowledge distillation method for a lightweight object detection model. The proposed method introduces an improved feature map novel non-maximum suppression (FM-NMS) method. The improved FM-NMS uses different focus size with respect to each object class, which can suppress false positives and improve detection accuracy. In our experiments, we use onestage object detection methods, YOLOv4 as a teacher model and YOLOv4-tiny as a student model, and we apply the proposed method to them. The experimental results demonstrate that the proposed method improves the detection accuracy of the student model while maintaining the lightweight model size. Lyuzhuang Liu, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICIP | 4 |
| 2022 | Refining Design Spaces in Knowledge Distillation for Deep Collaborative LearningabstractKnowledge distillation is one of the most widely utilized methods to improve the performance of a model. The knowledge transfer graph has been proposed for deep collaborative learning that enables a rich diversity of bidirectional knowledge distillation. However, exploring a knowledge transfer graph is difficult due to the many potential combinations it can have, so it is not clear how accurate the resultant graphs will actually be. To address this issue, we propose a method for designing the search space with step by step and analyze the trends of graphs to design graphs with high accuracy on the basis of the acquired results. Experiments on the CIFAR-100 dataset show that we confirm that the accuracy of the best knowledge transfer graph in the search space is better than that derived using the asynchronous successive halving algorithm. We also demonstrate that the explored knowledge transfer graphs can be transferred to different datasets. Sachi Iwata, Soma Minami, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICPR | 5 |
| 2022 | Action Spotting in Soccer Videos Using Multiple Scene EncodersabstractAction spotting, which temporally localizes specific actions in a video, is an important task for understanding high-level semantic information. In this paper, we formulate the action spotting task to one of scene sequence recognition and propose a model with multiple scene encoders to capture scene changes around the timestamp where an action occurs. We divide the input into multiple subsets to reduce the influence of scene context that is temporally distant, and feed every subset into a scene encoder to learn scene context in every subset. Because the optimal temporal length for time windows (chunks) is different for each action, we analyze the influence of chunk sizes for action spotting. The experimental results on the public SoccerNet-v2 dataset demonstrate state-of-the-art accuracy. By using embedding features, our method obtains an Average-mAP of 75.3%. In addition, we confirm that the performance can be improved by using optimal chunk sizes for different actions. Yuzhi Shi, Hiroaki Minoura, Takayoshi Yamashita, Tsubasa Hirakawa, Hironobu Fujiyoshi, Mitsuru Nakazawa, Yeongnam Chae, Björn Stenger |
ICPR | 5 |
| 2022 | Solving the Deadlock Problem with Deep Reinforcement Learning Using Information from Multiple VehiclesabstractAutonomous driving system controls a vehicle using path planning. Path planning for automated vehicles observes a vehicle and the surrounding information and plans a trajectory on the basis of rule-based approach. However, the rule-based path planning cannot generate an appropriate trajectory for complex scenes, such as two vehicles passes each other at an intersection without traffic lights. Such complex scene is called deadlock. For avoiding the deadlock, it is very costly to create rules manually. In this paper, we propose a multi-agent deep reinforcement learning method to generate appropriate trajectories at the deadlock scenes. The proposed method consists of a single feature extractor and actor-critic branches. Moreover, we introduce a mask-attention mechanism for visual explanation. By taking a look at the obtained attention maps, we can confirm the obtained agent and the reason of the behavior. For evaluating our method, we develop a simulator environment of autonomous driving that produces a certain deadlock scene. The experimental results with the developed environment show that the proposed method can generate trajectories avoiding deadlocks. Tsuyoshi Goto, Hidenori Itaya, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 5 |
| 2021 | Performance Prediction and Importance Analysis Using Transformer
Akiyoshi Satake, Hironobu Fujiyoshi, Takayoshi Yamashita |
ICCE | 2 |
| 2021 | Semantic Segmentation And Change Detection By Multi-Task U-NetabstractChange detection involves extracting the changed regions from images taken of the same place at different times. Potential applications are automatically updating of HD maps or identifying damages caused by natural disasters. However, conventional change detection methods merely detect changed regions without classifying them. In this paper, we propose a change detection method that can estimate the object class of a changed region. Our method extends a U-Net as a multi-task learning framework and estimates changed regions and semantic segmentation simultaneously. We propose using the pixel-wise classification probabilities of semantic segmentation for detecting changed regions rather than the conventional L2 norm-based difference of feature maps. In our experiments, we show that our method can improve change detection performance and estimate the classes of corresponding changed objects. Shungo Tsutsui, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICIP | 4 |
| 2021 | Visual Explanation using Attention Mechanism in Actor-Critic-based Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) has great potential for acquiring the optimal action in complex environments such as games and robot control. However, it is difficult to analyze the decision-making of the agent, i.e., the reasons it selects the action acquired by learning. In this work, we propose Mask-Attention A3C (Mask A3C), which introduces an attention mechanism into Asynchronous Advantage Actor-Critic (A3C), which is an actor-critic-based DRL method, and can analyze the decision-making of an agent in DRL. A3C consists of a feature extractor that extracts features from an image, a policy branch that outputs the policy, and a value branch that outputs the state value. In this method, we focus on the policy and value branches and introduce an attention mechanism into them. The attention mechanism applies a mask processing to the feature maps of each branch using mask-attention that expresses the judgment reason for the policy and state value with a heat map. We visualized mask-attention maps for games on the Atari 2600 and found we could easily analyze the reasons behind an agent's decision-making in various game tasks. Furthermore, experimental results showed that the agent could achieve a higher performance by introducing the attention mechanism. Hidenori Itaya, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura |
IJCNN | 4 |
| 2021 | Iterative Coarse-to-Fine 6D-Pose Estimation Using Back-propagationabstractWe propose a 6D pose estimation method for an object from a single RGB image for a robotic grasping task. Many approaches estimate pose parameters from images taken from other viewpoints and use deep learning to achieve high accuracy. However, most of these methods are not robust to changes in object texture, and there is a possibility that the correct pose cannot be estimated by only one-time inference. Our aims are to reduce the number of failure cases and improve the accuracy by a novel architecture using the iterative backpropagation of a pose decoder network and pose estimation on intermediate representation. The error between random and target pose parameters are backpropagated to a neural network and the gradient for approaching the target pose is obtained. The pose parameter is updated using the obtained gradient, the error is calculated again, and backpropagation is re-performed. Repeating this process, we estimate a more accurate pose. Experiments using our own dataset show that estimation accuracy is improved and the number of failure cases is reduced. Furthermore, estimation by coarse-to-fine iterative processing is more accurate and faster. We also experiment with grasping using a UR5 robot and show that the robot can grasp objects without depth information when using the pose estimated by the proposed method. Ryosuke Araki, Kousuke Mano, Tadanori Hirano, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IROS | 6 |
| 2021 | Image Captioning for Near-Future Events from Vehicle Camera Images and Motion InformationabstractImage captioning is a task to generate a sentence explaining an input image. In autonomous driving, image captioning is expected for providing linguistic explanations of autonomous driving control decision-making because it can reduce the psychological burden on passengers and prevent accidents. Current image-captioning methods are limited to generating a caption for an input image and not generating captions for events in the near future. It is important to generate captions for any event that will happen in the near future to prevent accidents and alert passengers. Therefore, we created a task to generate an explanatory sentence of near-future events using images observed from past to present. For this task, we propose a near-future image-captioning method suitable for in-vehicle camera images. Our experiments using the Berkeley Deep Drive eXplanation Dataset showed that the proposed method can appropriately generate captions for near-future events. Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 4 |
| 2020 | Knowledge Transfer Graph for Deep Collaborative Learning
Soma Minami, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ACCV (4) | 4 |
| 2020 | Spatial Temporal Attention Graph Convolutional Networks with Mechanics-Stream for Skeleton-Based Action Recognition
Katsutoshi Shiraki, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ACCV (5) | 4 |
| 2020 | Improving reliability of attention branch network by introducing uncertaintyabstractConvolutional neural networks (CNNs) are being used in various fields related to image recognition and are achieving high recognition accuracy. However, most existing CNNs do not consider uncertainty in their predictions; that is, they do not account for the difficulty of prediction, and the extent to which their predictions are reliable is unclear. This problem is considered to be the cause of erroneous decisions when we use CNNs in practice. By considering the uncertainty of the prediction result, it is thought that recognition accuracy would improve, and erroneous decisions would be suppressed. We propose a Bayesian attention branch network (Bayesian ABN) that incorporates uncertainty into an attention branch network (ABN). The method incorporates a Bayesian neural network (Bayesian NN) into the ABN to account for uncertainty in the prediction result. Also, it outputs prediction results from two branches and chooses the one having the lower uncertainty. In evaluations using standard object recognition datasets, we confirmed that the proposed method improves the accuracy and reliability of CNNs. Takuya Tsukahara, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICPR | 4 |
| 2020 | MT-DSSD: Deconvolutional Single Shot Detector Using Multi Task Learning for Object Detection, Segmentation, and Grasping DetectionabstractThis paper presents the multi-task Deconvolutional Single Shot Detector (MT-DSSD), which runs three tasks-object detection, semantic object segmentation, and grasping detection for a suction cup-in a single network based on the DSSD. Simultaneous execution of object detection and segmentation by multi-task learning improves the accuracy of these two tasks. Additionally, the model detects grasping points and performs the three recognition tasks necessary for robot manipulation. The proposed model can perform fast inference, which reduces the time required for grasping operation. Evaluations using the Amazon Robotics Challenge (ARC) dataset showed that our model has better object detection and segmentation performance than comparable methods, and robotic experiments for grasping show that our model can detect the appropriate grasping point. Ryosuke Araki, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICRA | 5 |
| 2020 | Video Object Detection and Tracking based on Angle Consistency between Motion and Flow
Toshiki Seo, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 4 |
| 2019 | Attention Branch Network: Learning of Attention Mechanism for Visual ExplanationabstractVisual explanation enables humans to understand the decision making of deep convolutional neural network (CNN), but it is insufficient to contribute to improving CNN performance. In this paper, we focus on the attention map for visual explanation, which represents a high response value as the attention location in image recognition. This attention region significantly improves the performance of CNN by introducing an attention mechanism that focuses on a specific region in an image. In this work, we propose Attention Branch Network (ABN), which extends a response-based visual explanation model by introducing a branch structure with an attention mechanism. ABN can be applicable to several image recognition tasks by introducing a branch for the attention mechanism and is trainable for visual explanation and image recognition in an end-to-end manner. We evaluate ABN on several image recognition tasks such as image classification, fine-grained recognition, and multiple facial attribute recognition. Experimental results indicate that ABN outperforms the baseline models on these image recognition tasks while generating an attention map for visual explanation. Our code is available. Hiroshi Fukui, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
CVPR | 4 |
| 2019 | Fast and Precise Detection of Object Grasping Positions with Eigenvalue TemplatesabstractFast Graspability Evaluation (FGE) has been proposed as a method for detecting grasping positions on objects and is now being used for industrial robots. FGE uses convolution of hand templates with regions on the target object to estimate the optimum grasping posture. However, the hand opening width and rotation angles must be set with high resolution to achieve highly accurate results and the computational load is high. To address that issue, we propose a method in which hand templates are represented in compact form for faster processing by using singular value decomposition. Applying singular value decomposition enables hand templates to be represented as linear combinations of a small number of eigenvalue templates and eigenfunctions. Eigenfunctions take discrete values, but response values can be calculated with arbitrary parameters by fitting a continuous function. Experimental results show that the proposed method reduces computation time by two thirds while maintaining the same detection accuracy as conventional FGE for both parallel hands and three-finger hands. Kousuke Mano, Takahiro Hasegawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Yukiyasu Domae |
ICRA | 4 |
| 2019 | Detecting layered structures of partially occluded objects for bin pickingabstractWhen robots engage in bin picking of multiple objects, a failure in grasping partially occluded objects may occur because other objects may overlap the desired ones. Therefore, the layered structure of objects needs to be detected, and the picking order needs to be established. In this paper, we propose a new dataset that evaluates not only the area of objects but also the layered structures of objects. In this dataset, three tasks are targeted: object detection, semantic segmentation, and segmentation of occluded areas for bin picking of multiple objects. The dataset, called “the Amazon Robotics Challenge (ARC) Multi-task Dataset” contains 1,500 RGB images and depth images, including all scenes containing bounding box labels, semantic segmentation labels, and occluded area labels. This enables representing the layered structure of overlapped objects with a tree structure. A benchmark of the ARC multi-task dataset demonstrated that occluded areas could be segmented using a Mask regional convolutional neural network (R-CNN) and that layered structures of objects could be predicted. Our dataset is available at the following URL:http://mprg.jp/research/arc_dataset_2017_e. Yusuke Inagaki, Ryosuke Araki, Takayoshi Yamashita, Hironobu Fujiyoshi |
IROS | 4 |
| 2019 | Visual Explanation by Attention Branch Network for End-to-end Learning-based Self-drivingabstractSelf-driving decides an appropriate control considering the surrounding environment. To this end, self-driving control methods by using a convolutional neural network (CNN) have been studied, which directly input the vehicle-mounted camera image to a network and output a steering directory. However, if we need to control not only steering but also throttle, it is necessary to grasp the state of the car itself in addition to the surrounding environment. Moreover, in order to use CNNs for critical applications such as self-driving, it is important to analyze where the network focuses on the image and to understand the decision making. In this work, we propose a method to solve these problems. First, to control both steering and throttle simultaneously, we propose using the current vehicle speed as the state of the car itself. Second, we introduce an attention branch network (ABN) architecture to a self-driving model, which enables visually analyzing the reason of the self-driving decision making by using an attention map. Experimental results with a driving simulator demonstrate that our method controls a car stably, and we can analyze the decision making by using the attention map. Keisuke Mori, Hiroshi Fukui, Takuya Murase, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 6 |
| 2018 | Compactification of Affine Transformation Filter Using Tensor DecompositionabstractKeypoint matching is used in a variety of tasks such as specific object recognition and panoramic image generation. Affine-SIFT (ASIFT) enables affine invariant matching by generating many affine transformation images of an input image. It describes the scale-invariant feature transform (SIFT) features of the generated image. However, ASIFT must perform multiple costly online computations for affine transformation. We represent an oriented FAST and rotated BRIEF (ORB) descriptor in a linear filter subjected to many affine transformations. We calculate the affine features by convolving the generated filter with the patch image. However, convolving the 19,200 filters generated by the affine transformation is inefficient. In order to reduce the convolution processing, the affine transformation filter is made compact by a factorization method. We built a 4-order tensor using the affine transformation filter. The 4-order tensor decomposes into the Tucker model. We reduce dimensions appropriately for each mode. In this way, we propose a compact and accurate feature description. Our evaluation experiments confirmed that the proposed method reduces the processing time to 19% while maintaining the same precision as singular value decomposition, which is the conventional method. Kohei Kawai, Takahiro Hasegawa, Yuji Yamauchi, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICIP | 5 |
| 2018 | Multiple Skip Connections of Dilated Convolution Network for Semantic SegmentationabstractSemantic segmentation is a task to estimate class for each pixel. This task also have received benefit from the deep ConvNet and it has achieves high accuracy. In the semantic segmentation from in-vehicle camera image, the object size such as a pedestrian or a vehicle fluctuates according to the distance from the camera. We propose scale aware semantic segmentation method especially small object. The contributions of the method are 1) to feed the features of small region by multiple skip connections, 2) to extract context from multiple receptive field by multiple dilated convolution blocks. The proposed method has achieved high accuracy in the Cityscapes dataset. The comparison with state-of-the-art methods, it has achieved the comparable performance at category IoU and iIoU metrics. Takayoshi Yamashita, Hironori Furukawa, Hironobu Fujiyoshi |
ICIP | 3 |
| 2017 | Accelerating Computation of Exemplar-SVM by Binary Approximation based on Matrix Decomposition
Takato Kurokawa, Yuji Yamauchi, Mitsuru Ambai, Takayoshi Yamashita, Hironobu Fujiyoshi |
BMVC | 5 |
| 2017 | Word Recognition by Combining Outline Emphasis and Synthesize Background
Yukihiro Achiha, Takayoshi Yamashita, Mitsuru Nakazawa, Soh Masuko, Yuji Yamauchi, Hironobu Fujiyoshi |
ICEC | 6 |
| 2016 | Misclassification tolerable learning for robust pedestrian orientation classificationabstractIn this paper, we propose a multiclass classifier training method which reduces “fatal” misclassifications by cost-relaxation of “tolerable” misclassifications in one-against-all classifiers training, named misclassification tolerable learning. In a binary classifier in the one-against-all classifiers, we introduce a new class group “conceptually similar classes,” whose class labels are similar to the positive class. In the case of pedestrian orientation classification, the conceptually similar classes are defined as neighboring orientations to the positive orientation. We consider the misclassification of the conceptually similar classes to the positive class as tolerable misclassification. By relaxing the cost of the tolerable misclassifications, our proposed classification method reduces fatal misclassifications of non-similar classes. We evaluated the cost-relaxation effectiveness on several public datasets and confirmed that the proposed method outperforms the normal SVM on all of the datasets in the soft criterion by achieving 78.63% recognition rate on PDC Dataset. Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi |
ICPR | 5 |
| 2016 | Pedestrian and part position detection using a regression-based multiple task deep convolutional neural networkabstractIn driving support systems, it is not only necessary to detect the position of pedestrians, but also to estimate the distance between a pedestrian and the vehicle. In general approaches using monocular cameras, the upper and lower positions of each pedestrian are detected using a bounding box obtained from a pedestrian detection technique. The distance between the pedestrian and the vehicle is then estimated using these positions and the camera parameters. This conventional framework uses independent pedestrian detection and position detection processes to estimate the distance. In this paper, we propose a method to detect both the pedestrian and their position simultaneously using a regression-based deep convolutional neural network (DCNN). This simultaneous detection method is possible to train efficient features for both tasks, because it is attention to head and leg regions from given labels. In the experiments, our method improves the performance of pedestrian detection compared with the DCNN which detects only pedestrian. The proposed approach also improves the detection accuracy of the head and leg positions compared with the methods that detect only these positions. Using the results of position detection and the camera parameters, our method achieves distance estimation to within 5% error. Takayoshi Yamashita, Hiroshi Fukui, Yuji Yamauchi, Hironobu Fujiyoshi |
ICPR | 4 |
| 2016 | Robust pedestrian attribute recognition for an unbalanced dataset using mini-batch training with rarity rateabstractPedestrian attributes are significant information for Advanced Driver Assistance System(ADAS). Pedestrian attributes such as body poses, face orientations and open umbrella are meant action or state of pedestrian. In general, this information is recognized using independent classifiers for each task. Performing all of these separate tasks is too time-consuming at the testing stage. In addition, the processing time increases with the number of tasks. To address this problem, multi-task learning or heterogeneous learning is able to train a single classifier to perform multiple tasks. In particular, heterogeneous learning is able to simultaneously train regression and recognition tasks, because reducing both training and testing time. However, heterogeneous learning tends to result in a lower accuracy rate for classes with a few training samples. In this paper, we propose a method to improve the performance of heterogeneous learning for such classes. We introduce a rarity rate based on the importance and class probability of each task. The appropriate rarity rate is assigned to each training sample. Thus, the samples in a mini-batch for training a deep convolutional neural network are augmented by this rarity rate to focus on the class with a few samples. Our heterogeneous learning approach with the rarity rate attains better performance on pedestrian attribute recognition, especially for classes representing open umbrellas. Hiroshi Fukui, Takayoshi Yamashita, Yuji Yamauchi, Hironobu Fujiyoshi, Hiroshi Murase |
Intelligent Vehicles Symposium | 4 |
| 2015 | Multiple-Hypothesis Affine Region Estimation with Anisotropic LoG FiltersabstractWe propose a method for estimating multiple-hypothesis affine regions from a keypoint by using an anisotropic Laplacian-of-Gaussian (LoG) filter. Although conventional affine region detectors, such as Hessian/Harris-Affine, iterate to find an affine region that fits a given image patch, such iterative searching is adversely affected by an initial point. To avoid this problem, we allow multiple detections from a single keypoint. We demonstrate that the responses of all possible anisotropic LoG filters can be efficiently computed by factorizing them in a similar manner to spectral SIFT. A large number of LoG filters that are densely sampled in a parameter space are reconstructed by a weighted combination of a limited number of representative filters, called "eigenfilters", by using singular value decomposition. Also, the reconstructed filter responses of the sampled parameters can be interpolated to a continuous representation by using a series of proper functions. This results in efficient multiple extrema searching in a continuous space. Experiments revealed that our method has higher repeatability than the conventional methods. Takahiro Hasegawa, Mitsuru Ambai, Kohta Ishikawa, Gou Koutaki, Yuji Yamauchi, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICCV | 7 |
| 2015 | Facial point detection based on a convolutional neural network with optimal mini-batch procedureabstractWe propose a Convolutional Neural Network (CNN)-based method to ensure both robustness to variations in facial pose and real-time processing. Although the robustness of CNNs has attracted attention in various fields, the training process suffers from difficulties in parameter setting and the manner in which training samples are provided. We demonstrate a manner of providing samples that results in a better network. We consider four methods: 1) subset with augmentation, 2) random selection, 3) fixed-person subset, and 4) the conventional approach. Experimental results indicate that the subset with augmentation technique has sufficient variations and quantity to obtain the best performance. Our CNN-based method is robust under facial pose variations, and achieves better performance. In addition, since our networks structure is simple, processing takes approximately 10ms for one face on a standard CPU. Masatoshi Kimura, Takayoshi Yamashita, Yuji Yamauchi, Hironobu Fujiyoshi |
ICIP | 4 |
| 2015 | SWAP-NODE: A regularization approach for deep convolutional neural networksabstractThe regularization is important for training of a deep network. One of breakthrough approach is dropout. It randomly deletes a certain number of activations in each layer in the feed-forward step of the training process. The dropout significantly reduces an effect of over-fitting and improves test performance. We introduce a new regularization approach for deep learning, called the swap-node. The swap-node, which is applied to a fully connected layer, swaps the activation values of two nodes randomly selected with a certain probability. Empirical evaluation shows that the network using the swap-node performs the best on MNIST, CIFAR-10, and SVHN. We also demonstrate superior performance of a combination of the swap-node and dropout on these datasets. Takayoshi Yamashita, Masayuki Tanaka 0001, Yuji Yamauchi, Hironobu Fujiyoshi |
ICIP | 4 |
| 2015 | Facial point detection using convolutional neural network transferred from a heterogeneous taskabstractWe present a novel training approach that uses convolutional neural network for facial part detection. In the proposed training procedure, we use the parameters of a network obtained for a heterogeneous task as the initial parameters of the network for a target task. We employ a convolutional neural network for facial part labeling in the heterogeneous task, and then transfer the trained network so as to provide initial parameters of the network for facial point detection. This transfer of network is advantageous in the training for a target task in that 1) the network obtains representation kernels for extraction of facial part regions and 2) the network reduces detection errors at distant positions. The performance of the proposed method applied to BioID and Labeled Face Parts in the Wild datasets is comparable to that of state-of-the art methods. In addition, since our network structure is simple, processing takes approximately 3ms for one face on a standard CPU. Takayoshi Yamashita, Taro Watasue, Yuji Yamauchi, Hironobu Fujiyoshi |
ICIP | 4 |
| 2015 | Fast 3D edge detection by using decision tree from depth imageabstractT3D edge detection from a depth image is an important technique of 3D object recognition in preprocessing. There are three types of 3D edges in a depth image called jump, convex roof, and concave roof edges. Conventional 3D edge detection based on ring operators has been proposed. The conventional ring operator can detect three types of 3D edges by classifying the response of Fourier transforms. Since the conventional method needs to apply Fourier transforms to all pixels of a depth image, real-time processing cannot be done due to high computational cost. Therefore, this paper presents a fast and reliable method of detecting three types of 3D edges by using a decision tree. The decision tree is trained under supervised learning from numerous synthesized depth images and labels by capturing depth relations between candidate pixels and pixels on a ring operator to classify 3D edges. The experimental results revealed that the proposed method has 25 times faster than the conventional method. This paper also presents some examples of 3D line and 3D convex corner detection based on results obtained with the proposed method. Masaya Kaneko, Takahiro Hasegawa, Yuji Yamauchi, Takayoshi Yamashita, Hironobu Fujiyoshi, Hiroshi Murase |
IROS | 5 |
| 2015 | Pedestrian detection based on deep convolutional neural network with ensemble inference networkabstractPedestrian detection is an active research topic for driving assistance systems. To install pedestrian detection in a regular vehicle, however, there is a need to reduce its cost and ensure high accuracy. Although many approaches have been developed, vision-based methods of pedestrian detection are best suited to these requirements. In this paper, we propose the methods based on Convolutional Neural Networks (CNN) that achieves high accuracy in various fields. To achieve such generalization, our CNN-based method introduces Random Dropout and Ensemble Inference Network (EIN) to the training and classification processes, respectively. Random Dropout selects units that have a flexible rate, instead of the fixed rate in conventional Dropout. EIN constructs multiple networks that have different structures in fully connected layers. The proposed methods achieves comparable performance to state-of-the-art methods, even though the structure of the proposed methods are considerably simpler. Hiroshi Fukui, Takayoshi Yamashita, Yuji Yamauchi, Hironobu Fujiyoshi, Hiroshi Murase |
Intelligent Vehicles Symposium | 4 |
| 2015 | Pedestrian orientation classification utilizing single-chip coaxial RGB-ToF cameraabstractThis paper proposes a method for pedestrian orientation classification. In image recognition, the accuracy is often degraded by the influence of background. In addition, it is also difficult to remove the background and extract only the human body from an image. To overcome these problems, we utilize a single-chip RGB-ToF camera. This camera can acquire RGB and depth images along the same optical axis at the same moment, and thus segmentation of the RGB image becomes easier by using the coaxial depth image. Our proposed method segmented a human body from its background accurately, which lead to the improvement of the accuracy of pedestrian orientation classification. Fumito Shinmura, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi |
Intelligent Vehicles Symposium | 6 |
| 2014 | Asymmetric Feature Representation for Object Recognition in Client Server System
Yuji Yamauchi, Mitsuru Ambai, Ikuro Sato, Yuichi Yoshida, Hironobu Fujiyoshi, Takayoshi Yamashita |
ACCV (1) | 5 |
| 2014 | Keypoint detection by cascaded fastabstractWhen the FAST method for detecting corner features at high speed is applied to images that include complex textures (regions that include foliage, shrubbery, etc.), many corners that are not needed for object recognition are detected because FAST defines corner features on the basis of a 16-pixel bounding circle. To overcome that problem, we propose the Cascaded FAST that defines corners on the basis of similarity in terms of intensity, continuity and orientation in a broader range of areas (20, 16, and 12 pixel bounding circles). Also, cascading three decision trees trained by the FAST approach enables high-speed corner detection in which non-corners are eliminated early in the process. Furthermore, Cascaded FAST determines scale by using an image pyramid and determines orientation at high speed by using a framework for referencing surrounding pixels. Takahiro Hasegawa, Yuji Yamauchi, Mitsuru Ambai, Yuichi Yoshida, Hironobu Fujiyoshi |
ICIP | 5 |
| 2014 | To Be Bernoulli or to Be Gaussian, for a Restricted Boltzmann MachineabstractWe introduce a method that automatically selects appropriate RBM types according to the visible unit distribution. The distribution of a visible unit strongly depends on a dataset. For example, binary data can be considered as pseudo binary distribution with high peaks at 0 and 1. For real-value data, the distribution can be modeled by single Gaussian model or Gaussian mixture model. Our proposed method selects appropriate RBM according to the distribution of each unit. We employ the Gaussian mixture model to determine whether the visible unit distribution is the pseudo binary or the Gaussian mixre. According to this distribution, we can select a Bernoulli-Bernoulli RBM(BBRBM) or a Gaussian-Bernoulli RBM(GBRBM). Furthermore, we employ normalization process to obtain a smoothed Gaussian mixture distribution. This allowed us to reduce variations such as illumination changes in the input data. After experimentation with MNIST, CBCL and our own dataset, our proposed method obtained the best recognition performance and further shortened the convergence time of the learning process. Takayoshi Yamashita, Masayuki Tanaka 0001, Eiji Yoshida, Yuji Yamauchi, Hironobu Fujiyoshi |
ICPR | 5 |
| 2013 | Robust object tracking using a context based on the relation of object and backgroundabstractObject tracking is complicated by perspective changes to both the object and background caused by object and the camera motion, object-to-object and object-to-background occlusion and illumination changes. Conventional object tracking method focus on distinguishing the target from the background. Adopting a new perspective, we propose a context aware tracking by a collaborative model of both the object and its surrounding background. In this model we introduce the “probability of tracking failure” that determines the feature similarity and the spatial relationship of the target and the surrounding background, as the “target-surrounding context”, which appears effective to predict the likelihood of a tracking failure. This target-surrounding context can be used to prevent tracking failure where the background has similar objects. In the experimental of the scene which occurs occlusion by similar objects, the proposed method outperformed most of the conventional methods. Takayoshi Yamashita, Hironobu Fujiyoshi |
ICASSP | 2 |
| 2012 | Human detection by Haar-like filtering using depth information
Sho Ikemura, Hironobu Fujiyoshi |
ICPR | 2 |
| 2011 | Moving object detection with background model based on spatio-temporal textureabstractBackground subtraction is a common method for detecting moving objects, but it is yet a difficult problem to distinguish moving objects from backgrounds when these backgrounds change significantly. Hence, we propose a method for detecting moving objects with a background model that covers dynamic changes in backgrounds utilizing a spatio-temporal texture named “Space-Time Patch”, which describes motion and appearance, whereas conventional textures describe appearance only. Our experimental results show the proposed method outperforms one conventional method in three scenes: in an outdoor scene where leaves and branches of a tree are waving in intermittent wind, in an indoor scene where ceiling lights are turned on and off frequently, and in an escalator scene beside a window facing outdoors where some passengers are leaning over the hand-rail. Ryo Yumiba, Masanori Miyoshi, Hironobu Fujiyoshi |
WACV | 3 |
| 2010 | Real-Time Human Detection Using Relational Depth Similarity Features
Sho Ikemura, Hironobu Fujiyoshi |
ACCV (4) | 2 |
| 2010 | Fostering UML Modeling Skills and Social Skills through Programming EducationabstractIn this research, we attempted to support the learning of the UML modeling skills and social skills required in software development scenarios as part of programming education in the Department of Engineering. We conducted a class based on PBL in which the learners formed teams to build a robot using LEGO Mindstorms. The results confirmed that through the classes, the learners showed improvements in both modeling skills and social skills. These results demonstrate the educational effectiveness of class design based on PBL and using the theme of building a robot, and the effectiveness of the modeling template created through this research. Norio Ishii, Yuri Suzuki 0001, Hironobu Fujiyoshi, Takashi Fujii |
CSEE&T | 3 |
| 2010 | A method for estimating cut-edit points in personal videosabstractWe analyze the tendencies in choosing cut-edit points in personal video content edited by users on the Internet, and develop a method to automatically estimate cut-edit points based on the results. When we investigated the relationship among cut-edit points in personal videos using a space-time patch feature(ST-patch feature), we realized that cut-edits were done in frames with a low Continuous Rank-Increase Measure (CRIM) value and a high Motion Correlation (MC) value calculated from the ST-patch feature. Therefore, we propose a method for estimating cut-edit points based on CRIM and MC values. Experimental results indicate that we obtained a recall ratio of 61.3% and a precision ratio of 50.4%. Takuya Furukawa, Hironobu Fujiyoshi, Akiyuki Nomura |
ICME | 2 |
| 2010 | Human-Area Segmentation by Selecting Similar Silhouette Images Based on Weak-Classifier ResponseabstractHuman-area segmentation is a major issue in video surveillance. Many existing methods estimate individual human areas from the foreground area obtained by background subtraction, but the effects of camera movement can make it difficult to obtain a background image. We have achieved human-area segmentation requiring no background image by using chamfer matching to match the results of human detection using Real AdaBoost with silhouette images. Although accuracy in chamfer matching drops as the number of templates increases, the proposed method enables segmentation accuracy to be improved by selecting silhouette images similar to the matching target beforehand based on response values from weak classifiers in Real AdaBoost. Hiroaki Ando, Hironobu Fujiyoshi |
ICPR | 2 |
| 2009 | Video Segmentation Using Iterated Graph Cuts Based on Spatio-temporal Volumes
Tomoyuki Nagahashi, Hironobu Fujiyoshi, Takeo Kanade |
ACCV (2) | 2 |
| 2009 | Designing a Programming Course to Foster Creativity using UML Modeling Template
Norio Ishii, Yuki Nagao, Yuri Suzuki 0001, Hironobu Fujiyoshi, Takashi Fujii |
CSEDU (2) | 4 |
| 2009 | Method for generating videowith virtual camerawork using bi-directional object tracking between keyframesabstractPersonal video sharing services such as YouTube have become popular because videos can easily be recorded in high-definition (HD) using a personal camcorder. However, it is difficult to broadcast an HD video via the Internet because of the large amount of data involved. We present a novel method for generating videos with virtual camerawork that is based on object tracking technology. Once a user specifies the positions of the object on keyframes, our method can be used to generate virtual camerawork between two keyframes in a row on the basis of the results of bi-directional tracking. We evaluated our method with subjective experiments and demonstrated its effectiveness. Yudai Shinoki, Hironobu Fujiyoshi |
ICME | 2 |
| 2009 | A Method for Visualizing Pedestrian Traffic Flow Using SIFT Feature Point Tracking
Yuji Tsuduki, Hironobu Fujiyoshi |
PSIVT | 2 |
| 2008 | Incoherent motion detection using a time-series Gram matrix featureabstractThis paper proposes a new method for incoherent motion recognition from video sequences. We use time-series spatio-temporal intensity gradients within a space-time patch. Using a global space-time patch, we found that the gradient feature allows us to distinguish an incoherent motion from a coherent motion without segmentalion. Furthermore the algorithm can run in real time even on an embedded device. In this paper, we verify motion recognition performance for actions which we consider coherent (walk/run) and incoherent (turn/squat/inverse walk). To identify the multiple motion classes, we use linear discriminant analysis and the KNN method. As a result, Our method can distinguish multiple-class motion patterns with a detection rate of about 80%. Also the detection rute of incoherent motions is 100% with a false positive rate of less than 10 %. Masato Kazui, Masanori Miyoshi, Shoji Muramatsu, Hironobu Fujiyoshi |
ICPR | 4 |
| 2008 | Shot boundary detection using co-occurrence of global motion in video streamabstractWe propose a method of shot boundary detection based on the co-occurrence of global motion in video stream. In addition to the conventional features based on appearance and local motion, we apply ST (Space-Time) patch analysis for detecting global motion in video stream. And then we perform shot boundary detection by constructing AdaBoost classifiers which represent the co-occurrence of global motion and the conventional features. Experimental results show that our method had 3.8% higher F-measure value than that of the conventional method for gradual shot boundary detection. Yosuke Murai, Hironobu Fujiyoshi |
ICPR | 2 |
| 2008 | A method of feature selection using contribution ratio based on boostingabstractAdaBoost and support vector machines (SVM) algorithms are commonly used in the field of object recognition. As classifiers, their classification performance is sensitive to affected by feature sets. To improve this performance, in addition to using the classifiers for accurate selection of feature sets, attention must be given to determining which feature subset to use in the classifier. Evaluating feature sets using a margin of the decision boundary of an SVM classifier proposed by Kugler is a solution for this problem. However, the margin in an SVM is sometimes large due to outliers. This paper presents a feature selection method that uses a contribution ratio based on boosting, which is effective for evaluating features. By comparing our method to the conventional one that uses a confident margin, we found that our method can select better feature sets using the contribution ratio obtained from boosting. Masamitsu Tsuchiya, Hironobu Fujiyoshi |
ICPR | 2 |
| 2008 | Human tracking based on Soft Decision Feature and online real boostingabstractOnline Boosting is an effective incremental learning method which can update weak classifiers efficiently according to the object being trackedt. It is a promising technique for online object tracking to adapt tothe appearance variations of objects during tracking process. However, proposed online-boosting based tracking methods update and select weak classifiers from fixed the offline learned weak classifiers, which might not be an optimal selection for object appearance variations. In this paper, we propose a new feature adjusting strategy for online boosting called Soft Decision Feature. We combine it with online real AdaBoost to achieve better tracking performance in scenes with human pose and posture variations. Experiment result demonstrates that it can successfully deal with the human posture variation scenes that conventional online boosting tracking methods fails to deal with. Takayoshi Yamashita, Hironobu Fujiyoshi, Shihong Lao, Masato Kawade |
ICPR | 2 |
| 2008 | People detection based on co-occurrence of appearance and spatiotemporal featuresabstractThis paper presents a method for detecting people based on the co-occurrence of appearance and spatiotemporal features. Histograms of oriented gradients(HOG) are used as appearance features, and the results of pixel state analysis are used as spatiotemporal features. The pixel state analysis classifies foreground pixels as either stationary or transient. The appearance and spatiotemporal features are projected into subspaces in order to reduce the dimensions of the vectors by principal component analysis(PCA). The cascade AdaBoost classifier is used to represent the co-occurrence of the appearance and spatiotemporal features. The use of feature co-occurrence, which captures the similarity of appearance, motion, and spatial information within the people class, makes it an effective detector. Experimental results show that the performance of our method is about 29% better than that of the conventional method. Yuji Yamauchi, Hironobu Fujiyoshi, Bon-Woo Hwang, Takeo Kanade |
ICPR | 2 |
| 2007 | Combined Object Detection and Segmentation by Using Space-Time Patches
Yasuhiro Murai, Hironobu Fujiyoshi, Takeo Kanade |
ACCV (1) | 2 |
| 2007 | Image Segmentation Using Iterated Graph Cuts Based on Multi-scale Smoothing
Tomoyuki Nagahashi, Hironobu Fujiyoshi, Takeo Kanade |
ACCV (2) | 2 |
| 2007 | Mean-Shift-Based Color Tracking in Illuminance Change
Yuji Hayashi, Hironobu Fujiyoshi |
RoboCup | 2 |
| 2006 | Generating a Time Shrunk Lecture Video by Event DetectionabstractStreaming a lecture video via the Internet is important for e-learning. We have developed a system that generates a lecture video using virtual camerawork based on shooting techniques of broadcast cameramen. However, viewing a full-length video takes time for students. In this paper, we propose a method for generating a time shrunk lecture video using event detection. We detect two kinds of events: a speech period and a chalkboard writing period. A speech period is detected by voice activity detection with LPC cepstrum and classified into speech or non-speech using Mahalanobis distance. To detect chalkboard writing periods, we use a graph cuts technique to segment a precise region of interests such as an instructor. By deleting content-free periods, i.e, period without the events of speech and writing, and fast-forwarding writing periods, our method can generate a time shrunk lecture video automatically. The resulting generated video is about 20%~30% shorter than the original video in time. This is almost the same as the results of manual editing by a human operator Takao Yokoi, Hironobu Fujiyoshi |
ICME | 2 |
| 2006 | Road Observation and Information Providing System for Supporting Mobility of PedestrianabstractWe have been developing the Robotic Communication Terminals (RCT) as a mobility support system for elderly and disabled people, which assists their impaired elements of mobility - recognition, actuation, and information access. The RCT consists of three types of terminals: "Environment Embedded Terminal (EET)", "user-carried mobile terminal", and "user-carrying mobile terminal". The EET system robustly detects moving objects at an outdoor surveillance site all day, and presents walkers with information about their surroundings. In this paper, as a part of the EET, we propose a method for detecting moving objects based on temporal differencing using adaptive thresholding calculated from intensity changes in the past few frames. For 23 cases of video evaluation, a high detection rate was measured under the variations caused by meteorological effects. We also have developed a test bed system that provides real-time road information detected by the EET to the map-based terminal. In the case of three clients, we demonstrate that users can receive the information from the EET with a 0.152[sec] time delay. Hironobu Fujiyoshi, Takeshi Komura, Ikuko Eguchi Yairi, Kentaro Kayama |
ICVS | 1 |
| 2005 | Virtual camerawork for generating lecture video from high resolution imagesabstractWe propose a method for generating a dynamic lecture video from the high resolution images recorded by a HDV camcorder. The lecture images are cropped to track the region of the instructor. Our approach uses bilateral filtering to avoid jittery motion caused by temporal differencing, and pseudo camera motion (panning) based on shooting technique of broadcast cameraman is generated. We evaluated our method, and showed that the effectiveness of our algorithm was verified through subjective experiments. Takao Yokoi, Hironobu Fujiyoshi |
ICME | 2 |
| 2005 | Mosaic-Based Global Vision System for Small Size Robot League
Yuji Hayashi, Seiji Tohyama, Hironobu Fujiyoshi |
RoboCup | 3 |
| 2005 | A New Practice Course for Freshmen Using RoboCup Based Small Robots
Yasunori Nagasaka, Morihiko Saeki, Shoichi Shibata, Hironobu Fujiyoshi, Takashi Fujii, Toshiyuki Sakata |
RoboCup | 4 |
| 2005 | Robust and Accurate Detection of Object Orientation and ID Without Color Segmentation
Shoichi Shimizu, Tomoyuki Nagahashi, Hironobu Fujiyoshi |
RoboCup | 3 |
| 2004 | A Method of Pseudo Stereo Vision from Images of Cameras Shutter Timing Adjusted
Hironobu Fujiyoshi, Shoichi Shimizu, Yasunori Nagasaka, Tomoichi Takahashi |
RoboCup | 1 |
| 2001 | Algorithms for cooperative multisensor surveillanceabstractThe Video Surveillance and Monitoring (VSAM) team at Carnegie Mellon University (CMU) has developed an end-to-end, multicamera surveillance system that allows a single human operator to monitor activities in a cluttered environment using a distributed network of active video sensors. Video understanding algorithms have been developed to automatically detect people and vehicles, seamlessly track them using a network of cooperating active sensors, determine their three-dimensional locations with respect to a geospatial site model, and present this information to a human operator who controls the system through a graphical user interface. The goal is to automatically collect and disseminate real-time information to improve the situational awareness of security providers and decision makers. The feasibility of real-time video surveillance has been demonstrated within a multicamera testbed system developed on the campus of CMU. This paper presents an overview of the issues and algorithms involved in creating this semiautonomous, multicamera surveillance system. Robert T. Collins, Alan J. Lipton, Hironobu Fujiyoshi, Takeo Kanade |
Proc. IEEE | 3 |
| 1998 | Real-time human motion analysis by image skeletonizationabstractIn this paper a process is described for analysing the motion of a human target in a video stream. Moving targets are detected and their boundaries extracted. From these, a "star" skeleton is produced. Two motion cues are determined from this skeletonization: body posture, and cyclic motion of skeleton segments. These cues are used to determine human activities such as walking or running, and even potentially, the target's gait. Unlike other methods, this does not require an a priori human model, or a large number of "pixels on target". Furthermore, it is computationally inexpensive, and thus ideal for real-world video applications such as outdoor video surveillance. Hironobu Fujiyoshi, Alan J. Lipton |
WACV | 1 |
| 1998 | Moving target classification and tracking from real-time videoabstractThis paper describes an end-to-end method for extracting moving targets from a real-time video stream, classifying them into predefined categories according to image-based properties, and then robustly tracking them. Moving targets are detected using the pixel wise difference between consecutive image frames. A classification metric is applied these targets with a temporal consistency constraint to classify them into three categories: human, vehicle or background clutter. Once classified targets are tracked by a combination of temporal differencing and template matching. The resulting system robustly identifies targets of interest, rejects background clutter and continually tracks over large distances and periods of time despite occlusions, appearance changes and cessation of target motion. Alan J. Lipton, Hironobu Fujiyoshi, Raju S. Patil |
WACV | 2 |