EDBT 2026 Demo / reviewers in the wild / expert
Tsubasa Hirakawa
dblp:141/9933
· DBLP profile ↗
35ranked-venue papers
3as first author
24since 2021 · last 2025
0000-0003-3851-5221ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 2 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Weight Pruning to Mitigate Class-Specific Accuracy Degradation for LiDAR-Based 3D Object DetectionabstractThe realization of autonomous driving systems requires efficient and accurate 3D object detection to identify objects such as vehicles, pedestrians, and cyclists within the driving environment using point cloud data. For achieving both high speed processing and high accuracy, it is necessary to reduce model size by model compression techniques, such as pruning, while maintaining performance. However, pruning for 3D object detection tasks has not been extensively studied, and the effects of applying existing pruning methods to 3D object detection models remain unclear. In this paper, we clarify the problems of pruning 3D object detection models with existing methods through preliminary experiments, and propose a pruning method suitable for 3D object detection models that solves these problems. Our preliminary experiments reveal that existing pruning methods significantly degrade detection performance for specific object classes. To address this issue, we propose a pruning method that preserves class-specific knowledge, mitigating biased accuracy degradation across different object classes. Experimental results on the KITTI dataset demonstrate that the proposed method can be combined with existing pruning methods without conflicts and achieves higher accuracy than existing methods. Tenshi Ito, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 2 |
| 2024 | DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention
NguyenHuu BaoLong, Yuzhi Shi, Tsubasa Hirakawa, Takayoshi Yamashita, Tohgoroh Matsui, Hironobu Fujiyoshi |
ACCV (10) | 4 |
| 2024 | Faster Convergence and Uncorrelated Gradients in Self-Supervised Online Continual Learning
Koyo Imai, Naoto Hayashi, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ACCV (8) | 3 |
| 2024 | Layer-Wise Relevance Propagation with Conservation Property for ResNet
Seitaro Otsuki, Tsumugi Iida, Félix Doublet, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura |
ECCV (43) | 4 |
| 2024 | Enhancing the Accuracy of Predicting Students Grades in Open-Ended Questions through Adjustments to Attention Weights
Masaki Koike, Hirokazu Kohama, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
EDM | 3 |
| 2024 | Binary-Decomposed Vision Transformer: Compressing and Accelerating Vision Transformer by Binary DecompositionabstractVision Transformers (ViTs) have emerged as versatile and high-performance models for various tasks such as image classification, object detection, and semantic segmentation. However, the ViT-L model, which demonstrates high accuracy, has a large number of parameters (307M), leading to increased computational requirements. To deploy ViTs on embedded devices and similar platforms, it is crucial to compress the model size and accelerate the inference process. In this paper, we propose the Binary-decomposed Vision Transformer (BdViT), a method for model compression and accelerated inference for ViTs models. BdViT consists of weight binarization based on vector decomposition and quantization of multiplication and addition operations, which does not require retraining model parameters. Through evaluation experiments using image recognition datasets, we demonstrated that BdViT can significantly reduce the number of parameters while mitigating performance degradation. Ryota Kondo, Hiroaki Minoura, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICIP | 3 |
| 2024 | Human-like Guidance by Generating Navigation Using Spatial-Temporal Scene GraphabstractVehicle navigation systems use both GPS and map data, primarily information derived from map data. Conventional navigation systems assume that the user will look directly at the display to check information. Simultaneously provided text and voice often play only a supplementary role, which can lead to driver distraction and misinterpretation. In contrast, human navigation utilizes visual information, potentially reducing the cognitive load on drivers. Human-like Guidance is aimed at realizing a driving assistance system that supports navigation akin to human guidance. Implementing Human-like Guidance, requires the handling of video footage from in-vehicle cameras during vehicle operation, suggesting the need for an approach combining image recognition and language model. However, images captured during operation often contain superfluous information, making the selection of relevant objects for navigation challenging. Moreover, relying solely on image information makes it difficult to consider the relationship with surrounding objects. Therefore, this study proposes a Spatial-Temporal Scene Graph that can represent spatial and temporal information of objects from driving scene videos. Furthermore, we achieve Human-like Guidance through navigation generation using features extracted from the Spatial-Temporal Scene Graph. Our results show that our proposed method improves the accuracy of navigation generation accuracy compared to traditional image-based navigation methods. In addition, the use of a Spatial-Temporal Scene Graph enables the generation of human-like navigation that focuses on the movements of surrounding vehicle objects. Hayato Suzuki, Kota Shimomura, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Shota Okubo, Nanri Takuya |
IV | 3 |
| 2024 | High-Precision for Multi-Task Learning from In-Vehicle Camera using BiFPNabstractMulti-task learning is effective for object detection and segmentation, which are closely related to each other and necessary for automated driving. However, there is a problem with the learning process in conventional multi-task learning models. In multi-task learning, common features among downstream tasks are first extracted by a backbone network. Then, these features are used for different downstream tasks. Since the required feature is different depending on the downstream task, it is necessary to extract features suitable for each downstream task. In this paper, we propose a multi-tasking model that introduces BiFPN feature fusion method for automated driving tasks and the Next-ViT model utilizing CNN and Transformer to extract features. From the evaluation experiments of automated driving tasks, we confirmed that the proposed method improves the accuracy of multi-task learning. Hidenori Itaya, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 3 |
| 2023 | Embedding Human Knowledge into Spatio-Temproal Attention Branch Network in Video Recognition via Temporal attention
Saki Noguchi, Yuzhi Shi, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
BMVC | 3 |
| 2023 | Recommending Learning Actions Using Neural NetworkabstractMany studies applying neural networks to the field of education have focused on student performance prediction and explainability of their decisions. While those studies introduced neural networks into educational settings, such networks cannot directly support student learnings in place of teachers. Therefore, we present a method that uses a general Transformer encoder to recommend appropriate learning actions for improving student performance. By considering the attention weight of a low-performing student to be close to that of a high-performing student, our method recommends the learning materials and actions for learning the materials. To evaluate the effectiveness of our method, we trained a deep neural network (DNN) on a private dataset of student operations (e.g., NEXT, PREV, OPEN) on digital learning materials obtained from a Japanese university. The number of operations divided by each learning material and by type of operation are input to the DNN, and the DNN outputs the student’s grade on 5-point scale. We applied our method with this trained DNN to samples that successfully predicted grades, and the number of operations increased on the basis of the recommended learning materials and actions. By re-inputting modified sample into the DNN, we then observe how the student performance changes. The results of this simple experiment indicate that more students improved their performance with both the material-based and operation-based recommendations than with random recommendations. The percentage of students whose grades improved tended to be larger for those with low grades. Specifically, the improvement ratio for students with the two lowest grades was over 90% by operation-based recommendation. This is consistent with our intuition that low-performing students are more likely to improve. Hirokazu Kohama, Yuki Ban, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Akitoshi Itai, Hiroyasu Usami |
ICCE | 3 |
| 2023 | This Looks Like It Rather Than That: ProtoKNN For Similarity-Based Classifiers
Yuki Ukai, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICLR | 2 |
| 2023 | Visual Explanation for Cooperative Behavior in Multi-Agent Reinforcement LearningabstractMulti-agent reinforcement learning (MARL) can acquire cooperative behavior among agents by training multiple agents in the same environment. Therefore, it is expected to be applied to complex tasks in real environments, such as traffic signal control in a traffic environment and cooperative behavior of robots. In this study, using the multi-actor-attention-critic (MAAC) with the actor-critic method as a basis, we introduce an attention head for the actor that calculates the agent's action. In contrast to the critic in MAAC, which shares the attention head among all the agents, the attention head of the actor in our method is constructed independently for each agent. This allows the attention head of the actor to calculate actor-attention (indicating which other agents are gazed at by each agent) and to acquire cooperative behavior. We visualize actor-attention to analyze the basis of agents’ decisions for cooperative behavior. Using single_spread, which is a multi-agent environment for cooperative problems, we show that the basis of decisions for cooperative behavior can be easily analyzed. We also demonstrate that our method efficiently obtains cooperative behavior considering other agents through quantitative evaluation of the cooperative behavior. Hidenori Itaya, Tom Sagawa, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IJCNN | 3 |
| 2023 | Analyzing the Accuracy, Representations, and Explainability of Various Loss Functions for Deep LearningabstractDeep learning utilizes a vast amounts of training data and updates weight parameters so as to minimize the loss between a predicted probability and a ground truth label. Generally, we use cross-entropy as the loss function. Although loss functions for image classification other than cross-entropy exist, their efficacy has not been adequately investigated. In this work, we extensively analyze models trained with different loss functions and clarify the properties of each. Specifically, we analyze the feature space and explainability as well as the classification accuracy on various benchmark datasets and network architectures. For feature space and explainability, we investigate the effectiveness of each loss function by quantitative and qualitative evaluations. We then discuss the properties and improvements of each. Tenshi Ito, Hiroki Adachi, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IJCNN | 3 |
| 2022 | Visual Explanation Generation Based on Lambda Attention Branch Networks
Tsumugi Iida, Takumi Komatsu, Kanta Kaneda, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura |
ACCV (2) | 4 |
| 2022 | Deep Ensemble Learning by Diverse Knowledge Distillation for Fine-Grained Object Classification
Naoki Okamoto, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ECCV (11) | 2 |
| 2022 | Class-Wise FM-NMS for Knowledge Distillation of Object DetectionabstractThe trade-off between accuracy and speed for an object detection model is important. When we implement an object detection model in embedded devices, a lightweight model can accelerate the detection speed. Meanwhile, the detection accuracy will be decreased. In this paper, we propose a knowledge distillation method for a lightweight object detection model. The proposed method introduces an improved feature map novel non-maximum suppression (FM-NMS) method. The improved FM-NMS uses different focus size with respect to each object class, which can suppress false positives and improve detection accuracy. In our experiments, we use onestage object detection methods, YOLOv4 as a teacher model and YOLOv4-tiny as a student model, and we apply the proposed method to them. The experimental results demonstrate that the proposed method improves the detection accuracy of the student model while maintaining the lightweight model size. Lyuzhuang Liu, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICIP | 2 |
| 2022 | Refining Design Spaces in Knowledge Distillation for Deep Collaborative LearningabstractKnowledge distillation is one of the most widely utilized methods to improve the performance of a model. The knowledge transfer graph has been proposed for deep collaborative learning that enables a rich diversity of bidirectional knowledge distillation. However, exploring a knowledge transfer graph is difficult due to the many potential combinations it can have, so it is not clear how accurate the resultant graphs will actually be. To address this issue, we propose a method for designing the search space with step by step and analyze the trends of graphs to design graphs with high accuracy on the basis of the acquired results. Experiments on the CIFAR-100 dataset show that we confirm that the accuracy of the best knowledge transfer graph in the search space is better than that derived using the asynchronous successive halving algorithm. We also demonstrate that the explored knowledge transfer graphs can be transferred to different datasets. Sachi Iwata, Soma Minami, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICPR | 3 |
| 2022 | Action Spotting in Soccer Videos Using Multiple Scene EncodersabstractAction spotting, which temporally localizes specific actions in a video, is an important task for understanding high-level semantic information. In this paper, we formulate the action spotting task to one of scene sequence recognition and propose a model with multiple scene encoders to capture scene changes around the timestamp where an action occurs. We divide the input into multiple subsets to reduce the influence of scene context that is temporally distant, and feed every subset into a scene encoder to learn scene context in every subset. Because the optimal temporal length for time windows (chunks) is different for each action, we analyze the influence of chunk sizes for action spotting. The experimental results on the public SoccerNet-v2 dataset demonstrate state-of-the-art accuracy. By using embedding features, our method obtains an Average-mAP of 75.3%. In addition, we confirm that the performance can be improved by using optimal chunk sizes for different actions. Yuzhi Shi, Hiroaki Minoura, Takayoshi Yamashita, Tsubasa Hirakawa, Hironobu Fujiyoshi, Mitsuru Nakazawa, Yeongnam Chae, Björn Stenger |
ICPR | 4 |
| 2022 | Forest-Related SDG Issues Monitoring for Data-Scarce Regions Employing Machine Learning and Remote Sensing - A Case Study for Ena City, JapanabstractWe proposed a combined machine learning approach with a deep convolutional neural network (CNN) to monitor forest utilization toward Sustainable Development Goals (SDGs) for data-scarce regions. First, we employed the Random Forest (RF) classifier using Google Earth Engine (GEE) for forest mapping. Then, we designed a deep CNN architecture that works for tree species/age mapping from coarse and polygonal ground-truth data. The proposed network has U-shape and comprises 3D Atrous Convolutions. The model was optimized by a weighted cross-entropy loss function. We trained the model with times-series Sentinel 1, 2, and Digital Elevation Model (DEM) data with sparse annotations. Our proposed models achieved 94.5% overall accuracy (OA) for forest mapping, 77.80% (OA) for tree species, and 81.74% (OA) for tree age classification, respectively in Ena city, Japan. The outcome of our study indicates the potential of remote sensing and machine learning in monitoring forest development, conservation, and utilization toward SDGs from coarse ground-truth data. Our source code for the implementation is available at: https://github.com/anhp95/forest_attr_segment Anh Phan, Kiyoshi Takejima, Tsubasa Hirakawa, Hiromichi Fukui |
IGARSS | 3 |
| 2022 | Solving the Deadlock Problem with Deep Reinforcement Learning Using Information from Multiple VehiclesabstractAutonomous driving system controls a vehicle using path planning. Path planning for automated vehicles observes a vehicle and the surrounding information and plans a trajectory on the basis of rule-based approach. However, the rule-based path planning cannot generate an appropriate trajectory for complex scenes, such as two vehicles passes each other at an intersection without traffic lights. Such complex scene is called deadlock. For avoiding the deadlock, it is very costly to create rules manually. In this paper, we propose a multi-agent deep reinforcement learning method to generate appropriate trajectories at the deadlock scenes. The proposed method consists of a single feature extractor and actor-critic branches. Moreover, we introduce a mask-attention mechanism for visual explanation. By taking a look at the obtained attention maps, we can confirm the obtained agent and the reason of the behavior. For evaluating our method, we develop a simulator environment of autonomous driving that produces a certain deadlock scene. The experimental results with the developed environment show that the proposed method can generate trajectories avoiding deadlocks. Tsuyoshi Goto, Hidenori Itaya, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 3 |
| 2021 | Semantic Segmentation And Change Detection By Multi-Task U-NetabstractChange detection involves extracting the changed regions from images taken of the same place at different times. Potential applications are automatically updating of HD maps or identifying damages caused by natural disasters. However, conventional change detection methods merely detect changed regions without classifying them. In this paper, we propose a change detection method that can estimate the object class of a changed region. Our method extends a U-Net as a multi-task learning framework and estimates changed regions and semantic segmentation simultaneously. We propose using the pixel-wise classification probabilities of semantic segmentation for detecting changed regions rather than the conventional L2 norm-based difference of feature maps. In our experiments, we show that our method can improve change detection performance and estimate the classes of corresponding changed objects. Shungo Tsutsui, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICIP | 2 |
| 2021 | Visual Explanation using Attention Mechanism in Actor-Critic-based Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) has great potential for acquiring the optimal action in complex environments such as games and robot control. However, it is difficult to analyze the decision-making of the agent, i.e., the reasons it selects the action acquired by learning. In this work, we propose Mask-Attention A3C (Mask A3C), which introduces an attention mechanism into Asynchronous Advantage Actor-Critic (A3C), which is an actor-critic-based DRL method, and can analyze the decision-making of an agent in DRL. A3C consists of a feature extractor that extracts features from an image, a policy branch that outputs the policy, and a value branch that outputs the state value. In this method, we focus on the policy and value branches and introduce an attention mechanism into them. The attention mechanism applies a mask processing to the feature maps of each branch using mask-attention that expresses the judgment reason for the policy and state value with a heat map. We visualized mask-attention maps for games on the Atari 2600 and found we could easily analyze the reasons behind an agent's decision-making in various game tasks. Furthermore, experimental results showed that the agent could achieve a higher performance by introducing the attention mechanism. Hidenori Itaya, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Komei Sugiura |
IJCNN | 2 |
| 2021 | Iterative Coarse-to-Fine 6D-Pose Estimation Using Back-propagationabstractWe propose a 6D pose estimation method for an object from a single RGB image for a robotic grasping task. Many approaches estimate pose parameters from images taken from other viewpoints and use deep learning to achieve high accuracy. However, most of these methods are not robust to changes in object texture, and there is a possibility that the correct pose cannot be estimated by only one-time inference. Our aims are to reduce the number of failure cases and improve the accuracy by a novel architecture using the iterative backpropagation of a pose decoder network and pose estimation on intermediate representation. The error between random and target pose parameters are backpropagated to a neural network and the gradient for approaching the target pose is obtained. The pose parameter is updated using the obtained gradient, the error is calculated again, and backpropagation is re-performed. Repeating this process, we estimate a more accurate pose. Experiments using our own dataset show that estimation accuracy is improved and the number of failure cases is reduced. Furthermore, estimation by coarse-to-fine iterative processing is more accurate and faster. We also experiment with grasping using a UR5 robot and show that the robot can grasp objects without depth information when using the pose estimated by the proposed method. Ryosuke Araki, Kousuke Mano, Tadanori Hirano, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IROS | 4 |
| 2021 | Image Captioning for Near-Future Events from Vehicle Camera Images and Motion InformationabstractImage captioning is a task to generate a sentence explaining an input image. In autonomous driving, image captioning is expected for providing linguistic explanations of autonomous driving control decision-making because it can reduce the psychological burden on passengers and prevent accidents. Current image-captioning methods are limited to generating a caption for an input image and not generating captions for events in the near future. It is important to generate captions for any event that will happen in the near future to prevent accidents and alert passengers. Therefore, we created a task to generate an explanatory sentence of near-future events using images observed from past to present. For this task, we propose a near-future image-captioning method suitable for in-vehicle camera images. Our experiments using the Berkeley Deep Drive eXplanation Dataset showed that the proposed method can appropriately generate captions for near-future events. Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 2 |
| 2020 | Knowledge Transfer Graph for Deep Collaborative Learning
Soma Minami, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ACCV (4) | 2 |
| 2020 | Spatial Temporal Attention Graph Convolutional Networks with Mechanics-Stream for Skeleton-Based Action Recognition
Katsutoshi Shiraki, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ACCV (5) | 2 |
| 2020 | Improving reliability of attention branch network by introducing uncertaintyabstractConvolutional neural networks (CNNs) are being used in various fields related to image recognition and are achieving high recognition accuracy. However, most existing CNNs do not consider uncertainty in their predictions; that is, they do not account for the difficulty of prediction, and the extent to which their predictions are reliable is unclear. This problem is considered to be the cause of erroneous decisions when we use CNNs in practice. By considering the uncertainty of the prediction result, it is thought that recognition accuracy would improve, and erroneous decisions would be suppressed. We propose a Bayesian attention branch network (Bayesian ABN) that incorporates uncertainty into an attention branch network (ABN). The method incorporates a Bayesian neural network (Bayesian NN) into the ABN to account for uncertainty in the prediction result. Also, it outputs prediction results from two branches and chooses the one having the lower uncertainty. In evaluations using standard object recognition datasets, we confirmed that the proposed method improves the accuracy and reliability of CNNs. Takuya Tsukahara, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICPR | 2 |
| 2020 | MT-DSSD: Deconvolutional Single Shot Detector Using Multi Task Learning for Object Detection, Segmentation, and Grasping DetectionabstractThis paper presents the multi-task Deconvolutional Single Shot Detector (MT-DSSD), which runs three tasks-object detection, semantic object segmentation, and grasping detection for a suction cup-in a single network based on the DSSD. Simultaneous execution of object detection and segmentation by multi-task learning improves the accuracy of these two tasks. Additionally, the model detects grasping points and performs the three recognition tasks necessary for robot manipulation. The proposed model can perform fast inference, which reduces the time required for grasping operation. Evaluations using the Amazon Robotics Challenge (ARC) dataset showed that our model has better object detection and segmentation performance than comparable methods, and robotic experiments for grasping show that our model can detect the appropriate grasping point. Ryosuke Araki, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
ICRA | 3 |
| 2020 | Correction of Seasonal Effects on VIIRS DNB Monthly Composites by Using Stable Lit Data and Regression Convolutional Neural NetworkabstractIn the past years, satellite-observed nighttime lights have been one of the widely used geospatial data products. The Visible Infrared Imaging Radiometer Suite (VIIRS) Day/Night Band (DNB) is currently one of the highest-quality nighttime lights data. However, nighttime lights radiance can be strongly affected by seasonal factors including variations in snow and vegetation. Therefore, uncorrected nighttime lights data may cause strong bias in quantitative researches such as estimation of population and economic activity. In recent years, there have been several researches aiming at investigation of impacts of seasonal factors on nighttime lights radiance. However, it is still lack of method for correction of seasonal effects in the data. In this paper, we propose a new seasonal-effects correction algorithm for VIIRS DNB monthly composite data. This algorithm is based on utilization of stable lit pixels as reference and a regression convolutional neural network (CNN) to match uncorrected data to the reference data. Experimental results show that the algorithm significantly improved correlation (R2) between total sum of night lights (TOL) and electric power consumption (EPC), from nearly 0.78 to 0.93, with data covering the whole Japanese territory. Visual inspection shows that the brightness of snow-affected regions were strongly reduced after undergoing correction. Chuc Man Duc, Tsubasa Hirakawa, Hiromichi Fukui |
IGARSS | 2 |
| 2020 | Video Object Detection and Tracking based on Angle Consistency between Motion and Flow
Toshiki Seo, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 2 |
| 2019 | Attention Branch Network: Learning of Attention Mechanism for Visual ExplanationabstractVisual explanation enables humans to understand the decision making of deep convolutional neural network (CNN), but it is insufficient to contribute to improving CNN performance. In this paper, we focus on the attention map for visual explanation, which represents a high response value as the attention location in image recognition. This attention region significantly improves the performance of CNN by introducing an attention mechanism that focuses on a specific region in an image. In this work, we propose Attention Branch Network (ABN), which extends a response-based visual explanation model by introducing a branch structure with an attention mechanism. ABN can be applicable to several image recognition tasks by introducing a branch for the attention mechanism and is trainable for visual explanation and image recognition in an end-to-end manner. We evaluate ABN on several image recognition tasks such as image classification, fine-grained recognition, and multiple facial attribute recognition. Experimental results indicate that ABN outperforms the baseline models on these image recognition tasks while generating an attention map for visual explanation. Our code is available. Hiroshi Fukui, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
CVPR | 2 |
| 2019 | Visual Explanation by Attention Branch Network for End-to-end Learning-based Self-drivingabstractSelf-driving decides an appropriate control considering the surrounding environment. To this end, self-driving control methods by using a convolutional neural network (CNN) have been studied, which directly input the vehicle-mounted camera image to a network and output a steering directory. However, if we need to control not only steering but also throttle, it is necessary to grasp the state of the car itself in addition to the surrounding environment. Moreover, in order to use CNNs for critical applications such as self-driving, it is important to analyze where the network focuses on the image and to understand the decision making. In this work, we propose a method to solve these problems. First, to control both steering and throttle simultaneously, we propose using the current vehicle speed as the state of the car itself. Second, we introduce an attention branch network (ABN) architecture to a self-driving model, which enables visually analyzing the reason of the self-driving decision making by using an attention map. Experimental results with a driving simulator demonstrate that our method controls a car stably, and we can analyze the decision making by using the attention map. Keisuke Mori, Hiroshi Fukui, Takuya Murase, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi |
IV | 4 |
| 2016 | Defocus-aware Dirichlet particle filter for stable endoscopic video frame recognitionabstractBACKGROUND AND OBJECTIVE: A computer-aided system for colorectal endoscopy could provide endoscopists with important helpful diagnostic support during examinations. A straightforward means of providing an objective diagnosis in real time might be for using classifiers to identify individual parts of every endoscopic video frame, but the results could be highly unstable due to out-of-focus frames. To address this problem, we propose a defocus-aware Dirichlet particle filter (D-DPF) that combines a particle filter with a Dirichlet distribution and defocus information. METHODS: We develop a particle filter with a Dirichlet distribution that represents the state transition and likelihood of each video frame. We also incorporate additional defocus information by using isolated pixel ratios to sample from a Rayleigh distribution. RESULTS: We tested the performance of the proposed method using synthetic and real endoscopic videos with a frame-wise classifier trained on 1671 images of colorectal endoscopy. Two synthetic videos comprising 600 frames were used for comparisons with a Kalman filter and D-DPF without defocus information, and D-DPF was shown to be more robust against the instability of frame-wise classification results. Computation time was approximately 88ms/frame, which is sufficient for real-time applications. We applied our method to 33 endoscopic videos and showed that the proposed method can effectively smoothen highly unstable probability curves under actual defocus of the endoscopic videos. CONCLUSION: The proposed D-DPF is a useful tool for smoothing unstable results of frame-wise classification of endoscopic videos to support real-time diagnosis during endoscopic examinations. Tsubasa Hirakawa, Toru Tamaki, Bisser Raytchev, Kazufumi Kaneda, Tetsushi Koide, Shigeto Yoshida, Yoko Kominami, Shinji Tanaka |
Artif. Intell. Medicine | 1 |
| 2016 | Corrigendum to "Defocus-aware Dirichlet particle filter for stable endoscopic video frame recognition" [Artif. Intell. Med. 68 (March 2016) 1-16]
Tsubasa Hirakawa, Toru Tamaki, Bisser Raytchev, Kazufumi Kaneda, Tetsushi Koide, Shigeto Yoshida, Yoko Kominami, Shinji Tanaka |
Artif. Intell. Medicine | 1 |
| 2013 | Smoothing posterior probabilities with a particle filter of dirichlet distribution for stabilizing colorectal NBI endoscopy recognitionabstractThis paper proposes a method for smoothing the posterior probabilities obtained from classification results of time series input. We deal with this problem as a filtering problem with Dirichlet distribution and develop a particle filtering for this task. As a practical example of smoothing, we apply the proposed method to stabilizing NBI endoscopy recognition results over time. Experimental results demonstrate that our approach can effectively smooth highly unstable probability curves. Tsubasa Hirakawa, Toru Tamaki, Bisser Raytchev, Kazufumi Kaneda, Tetsushi Koide, Yoko Kominami, Rie Miyaki, Taiji Matsuo, Shigeto Yoshida, Shinji Tanaka |
ICIP | 1 |