EDBT 2026 Demo / reviewers in the wild / expert
Kanji Tanaka 0003
dblp:27/3752-3 · also Tanaka Kanji 0003
· DBLP profile ↗
24ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-1143-5478ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 6 first-author · 5 since 2021Systems, architecture and hardware · 14 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HOI-LCD: Leveraging Humans as Dynamic Landmarks Toward Thermal Loop Closing Even in Complete Darkness
Tatsuro Sakai, Yanshuo Bai, Kanji Tanaka 0003, Wuhao Xie, Jonathan Tay Yu Liang, Daiki Iwata |
ICINCO (2) | 3 |
| 2025 | Continual Multi-Robot Learning from Black-Box Visual Place Recognition Models
Kenta Tsukahara, Kanji Tanaka 0003, Daiki Iwata, Jonathan Tay Yu Liang, Wuhao Xie |
ICINCO (1) | 2 |
| 2024 | Robot Traversability Prediction: Towards Third-Person-View Extension of Walk2Map with Photometric and Physical ConstraintsabstractWalk2Map has emerged as a promising data-driven method to generate indoor traversability maps based solely on pedestrian trajectories, offering great potential for indoor robot navigation. In this study, we investigate a novel approach called Walk2Map++, which involves replacing Walk2Map’s first-person sensor (i.e., IMU) with a human observing third-person view from the robot’s onboard camera. However, human observation from a third-person camera is significantly ill-posed due to visual uncertainties resulting from occlusion, nonlinear perspective, depth ambiguity, and human-to-human interaction. To regularize the ill-posedness, we propose integrating two types of constraints: photometric (i.e., occlusion ordering) and physical (i.e., collision avoidance). We demonstrate that these constraints can be effectively inferred from the interaction between past and present observations, human trackers, and object reconstructions. We depict the seamless integration of asynchronous map optimization events, like loop closure, into the real-time traversability map, facilitating incremental and efficient map refinement. We validate the efficacy of our enhanced methodology through rigorous fusion and comparison with established techniques, demonstrating its capability to advance traversability prediction in complex indoor environments. The code and datasets associated with this study are available for further research and adoption in the field at https://github.com/jonathantyl97/HO3-SLAM. Jonathan Tay Yu Liang, Kanji Tanaka 0003 |
IROS | 2 |
| 2022 | TaylorMade Visual Burr Detection for High-mix Low-volume Production of Non-convex Cylindrical Metal Objects
Kyosuke Tashiro, Koji Takeda, Shogo Aoki, Haoming Ye, Hiroki Tomoe, Kanji Tanaka 0003 |
ICPRAM | 6 |
| 2022 | Domain Invariant Siamese Attention Mask for Small Object Change Detection via Everyday Indoor Robot NavigationabstractThe problem of image change detection via every-day indoor robot navigation is explored from a novel perspective of the self-attention technique. Detecting semantically non-distinctive and visually small changes remains a key challenge in the robotics community. Intuitively, these small non-distinctive changes may be better handled by the recent paradigm of the attention mechanism, which is the basic idea of this work. However, existing self-attention models require significant retraining cost per domain, so it is not directly applicable to robotics applications. We propose a new self-attention technique with an ability of unsupervised on-the-fly domain adaptation, which introduces an attention mask into the intermediate layer of an image change detection model, without modifying the input and output layers of the model. Experiments, in which an indoor robot aims to detect visually small changes in everyday navigation, demonstrate that our attention technique significantly boosts the state-of-the-art image change detection model. Our datset is available at https://github.com/KojiTakeda00/Small_object_change_detection Koji Takeda, Kanji Tanaka 0003, Yoshimasa Nakamura |
IROS | 2 |
| 2021 | Dark Reciprocal-Rank: Teacher-to-student Knowledge Transfer from Self-localization Model to Graph-convolutional Neural NetworkabstractIn visual robot self-localization, graph-based scene representation and matching have recently attracted research interest as robust and discriminative methods for self-localization. Although effective, their computational and storage costs do not scale well to large-size environments. To alleviate this problem, we formulate self-localization as a graph classification problem and attempt to use the graph convolutional neural network (GCN) as a graph classification engine. A straightforward approach is to use visual feature descriptors that are employed by state-of-the-art self-localization systems, directly as graph node features. However, their superior performance in the original self-localization system may not necessarily be replicated in GCN-based self-localization. To address this issue, we introduce a novel teacher-to-student knowledge-transfer scheme based on rank matching, in which the reciprocal-rank vector output by an off-the-shelf state-of-the-art teacher self-localization model is used as the dark knowledge to transfer. Experiments indicate that the proposed graph-convolutional self-localization network (GCLN) can significantly outperform state-of-the-art self-localization systems, as well as the teacher classifier. The code and dataset are available at https://github.com/KojiTakeda00/Reciprocal_rank_KT_GCN. Koji Takeda, Kanji Tanaka 0003 |
ICRA | 2 |
| 2020 | Deep Next-Best-View Planner for Cross-Season Visual Route ClassificationabstractThis paper addresses the problem of active visual place recognition (VPR) from a novel perspective of long-term autonomy. In our approach, a next-best-view (NBV) planner plans an optimal action-observation-sequence to maximize the expected cost-performance for a visual route classification task. A difficulty arises from the fact that the NBV planner is trained and tested in different domains (times of day, weather conditions, and seasons). Existing NBV methods may be confused and deteriorated by the domain-shifts, and require significant efforts for adapting them to a new domain. We address this issue by a novel deep convolutional neural network (DNN) -based NBV planner that does not require the adaptation step. Our main contributions in this paper are summarized as follows: (1) We present a novel domain-invariant NBV planner that is specifically tailored for DNN-based VPR. (2) We formulate the active VPR as a POMDP problem and present a feasible solution to address the inherent intractability. Specifically, the probability distribution vector (PDV) output by the available DNN is used as a domain-invariant observation model without the need to retrain it. (3) We verify efficacy of the proposed approach through challenging cross-season VPR experiments, where it is confirmed that the proposed approach clearly outperforms the previous single-view-based or multi-view-based VPR in terms of VPR accuracy and action-observation-cost. Kanya Kurauchi, Kanji Tanaka 0003 |
ICPR | 2 |
| 2020 | Self-Supervised Map-Segmentation by Mining Minimal-Map-SegmentsabstractIn visual place recognition (VPR), map segmentation (MS) is a preprocessing technique used to partition a given view-sequence map into place classes (i.e., map segments) so that each class has good place-specific training images for a visual place classifier (VPC). Existing approaches to MS implicitly/explicitly suppose that map segments have a certain size, or individual map segments are balanced in size. However, recent VPR systems showed that very small important map segments (minimal map segments) often suffice for VPC, and the remaining large unimportant portion of the map should be discarded to minimize map maintenance cost. Here, a new MS algorithm that can mine minimal map segments from a large view-sequence map is presented. To solve the inherently NP hard problem, MS is formulated as a video-segmentation problem and the recently-developed efficient point-trajectory based paradigm of video segmentation is used. The proposed map representation was implemented with three types of VPC: deep convolutional neural network, bag-of-words, and object class detector, and each was integrated into a Monte Carlo localization algorithm (MCL) within a topometric VPR framework. Experiments using the publicly available NCLT dataset thoroughly investigate the efficacy of MS in terms of VPR performance. Kanji Tanaka 0003 |
IV | 1 |
| 2019 | Detection-by-Localization: Maintenance-Free Change Object DetectorabstractRecent researches demonstrate that selflocalization performance is a very useful measure of likelihood-of-change (LoC) for change detection. In this paper, this “detection-by-localization” scheme is studied in a novel generalized task of object-level change detection. In our framework, a given query image is segmented into object-level subimages (termed “scene parts”), which are then converted to subimagelevel pixel-wise LoC maps via the detection-by-localization scheme. Our approach models a self-localization system as a ranking function, outputting a ranked list of reference images, without requiring relevance score. Thanks to this new setting, we can generalize our approach to a broad class of selflocalization systems. We further propose an aggregation of different self-localization results from different queries so as to achieve higher precision. Our ranking based self-localization model allows to fuse self-localization results from different modalities via an unsupervised rank fusion derived from a field of multi-modal information retrieval (MMR). Our framework does not rely on the raw-score-merging hypothesis. Challenging experiments of cross-season change detection using the publicly available North Campus Long-Term (NCLT) dataset validates the efficacy of our proposed method. Kanji Tanaka 0003 |
ICRA | 1 |
| 2019 | Deep Intersection Classification Using First and Third Person ViewsabstractWe explore the problem of intersection classification using monocular on-board passive vision, with the goal of classifying traffic scenes with respect to road topology. We divide the existing approaches into two broad categories according to the type of input data: (a) first person vision (FPV) approaches, which use an egocentric view sequence as the intersection is passed; and (b) third person vision (TPV) approaches, which use a single view immediately before entering the intersection. The FPV and TPV approaches each have advantages and disadvantages. Therefore, we aim to combine them into a unified deep learning framework. Experimental results show that the proposed FPV-TPV scheme outperforms previous methods and only requires minimal FPV/TPV measurements. Koji Takeda, Kanji Tanaka 0003 |
IV | 2 |
| 2018 | Leveraging Object Proposals for Object-Level Change DetectionabstractFeature-based image differencing is an efficient approach to image change detection, which performs fast enough for self-driving car and robotic applications. Extant approaches typically take local keypoint features as input to the differencing stage. In this study, we aim to extend the differencing stage to consider object-level features. Our object level approach is inspired by recent advances in two independent object-region proposal techniques: supervised object proposal (e.g., YOLO) and unsupervised object proposal (e.g., BING). A difficulty arises from the fact that even state-of-the-art object proposal techniques suffer from misdetections and false alarms. Our key concept is combining the supervised and unsupervised techniques into a common framework that evaluates the likelihood of change at the semantic object level. We address a challenging urban scenario using the publicly available Malaga dataset and experimentally verify that improved change detection performance can be obtained with our approach. Takuma Sugimoto, Kanji Tanaka 0003, Kousuke Yamaguchi |
Intelligent Vehicles Symposium | 2 |
| 2018 | An Experimental Study on Generative Adversarial Network and Visual Experience Mining for Domain Adaptive Change DetectionabstractThis paper addresses the problem of cross-domain change detection from a novel perspective of image-to-image translation. In general, change detection aims to identify interesting changes between a given query image and a reference image of the same scene taken at a different time. This problem becomes a challenging one when query and reference images involve different domains (e.g., time of the day, weather, and season) due to variations in object appearance and a limited amount of training examples. In this study, we address the above issue by leveraging a generative adversarial network (GAN). Our key concept is to use a limited amount of training data to train a GAN-based image translator that maps a reference image to a virtual image that cannot be discriminated from query domain images. This enables us to treat the cross-domain change detection task as an in-domain image comparison. This allows us to leverage the large body of literature on in-domain generic change detectors. As a part of our contribution, we investigate to what extent the usage of the translation GAN can alleviate the issue of domain shift by developing and evaluating several change detection algorithms. In addition, we also consider the use of visual place recognition as a method for mining more appropriate reference images over the space of virtual images. Experimental results on cross-season change detection using the publicly available North Campus Long-Term autonomy dataset validate the efficacy of the proposed approach. Kousuke Yamaguchi, Kanji Tanaka 0003, Yuusuke Kojima, Takuma Sugimoto |
SMC | 2 |
| 2017 | Grammar-based map parsing for view invariant map descriptorabstractMap retrieval, the problem of similarity search over a large collection of 2D pointset maps previously built by mobile robots, is crucial for autonomous navigation in indoor and outdoor environments. Bag-of-words (BoW) methods constitute a popular approach to map retrieval; however, these methods have extremely limited descriptive ability because they ignore the spatial layout information of the local features. The main contribution of this paper is an extension of the bag-of-words map retrieval method to enable the use of spatial information from local features. Our strategy is to explicitly model a unique viewpoint of an input local map; the pose of the local feature is defined with respect to this unique viewpoint, and can be viewed as an additional invariant feature for discriminative map retrieval. Specifically, we wish to determine a unique viewpoint that is invariant to moving objects, clutter, occlusions, and actual viewpoints. Hence, we perform scene parsing to analyze the scene structure, and consider the “center” of the scene structure to be the unique viewpoint. Our scene parsing is based on a Manhattan world grammar that imposes a quasi-Manhattan world constraint to enable the robust detection of a scene structure that is invariant to clutter and moving objects. Experimental results using the publicly available radish dataset validate the efficacy of the proposed approach. Enfu Liu, Kanji Tanaka 0003, Xiaoxiao Fei |
IPIN | 2 |
| 2016 | Self-localization from images with small overlapabstractWith the recent success of visual features from deep convolutional neural networks (DCNN) in visual robot self-localization, it has become important and practical to address more general self-localization scenarios. In this paper, we address the scenario of self-localization from images with small overlap. We explicitly introduce a localization difficulty index as a decreasing function of view overlap between query and relevant database images and investigate performance versus difficulty for challenging cross-view self-localization tasks. We then reformulate the self-localization as a scalable bag-of-visual-features (BoVF) scene retrieval and present an efficient solution called PCA-NBNN, aiming to facilitate fast and yet discriminative correspondence between partially overlapping images. The proposed approach adopts recent findings in discriminativity preserving encoding of DCNN features using principal component analysis (PCA) and cross-domain scene matching using naive Bayes nearest neighbor distance metric (NBNN). We experimentally demonstrate that the proposed PCA-NBNN framework frequently achieves comparable results to previous DCNN features and that the BoVF model is significantly more efficient. We further address an important alternative scenario of “self-localization from images with NO overlap” and report the result. Kanji Tanaka 0003 |
IROS | 1 |
| 2015 | Leveraging image-based prior in cross-season place recognitionabstractIn this paper, we address the challenging problem of single-view cross-season place recognition. A new approach is proposed for compact discriminative scene descriptor that helps in coping with changes in appearance in the environment. We focus on a simple effective strategy that uses objects whose appearance remain the same across seasons as valid landmarks. Unlike popular bag-of-words (BoW) scene descriptors that rely on a library of vector quantized visual features, our descriptor is based on a library of raw image data (e.g., visual experience shared by colleague robots, publicly available photo collections from Google StreetView), and directly mines it to identify landmarks (i.e., image patches) that effectively explain an input query/database image. The discovered landmarks are then compactly described by their pose and shape (i.e., library image ID, and bounding boxes) and used as a compact discriminative scene descriptor for the input image. We collected a dataset of single-view images across seasons with annotated ground truth, and evaluated the effectiveness of our scene description framework by comparing its performance to that of previous BoW approaches, and by applying an advanced Naive Bayes Nearest neighbor (NBNN) image-to-class distance measure. Masatoshi Ando, Yuuto Chokushi, Kanji Tanaka 0003, Kentaro Yanagihara |
ICRA | 3 |
| 2015 | Unsupervised part-based scene modeling for visual robot localizationabstractScene modeling is an important first stage in visual robot localization. In recent years, the bag-of-words (BoW) scene modeling approach has attracted considerable attention as a method for obtaining compact discriminative scene descriptors for map retrieval. However, a BoW scene descriptor alone cannot address partial view changes and often produces poor localization in practice. In this work, we address this issue by proposing a simple effective approach, “unsupervised part-based scene modeling,” in which a set of useful parts is discovered via scene parsing and the parts are used as additional queries for the map retrieval. We also address the issue of discovering useful parts in a scene, and present a solution that provides similar parts for similar scenes. The next contribution of this work is that we present a practical robot self-localization system that consists of three distinct steps: (1) robust hierarchical scene parsing to obtain multiple scene/part queries, (2) saliency-based selection of useful parts, and (3) aggregation of ranking results from multiple scene/part queries to obtain a reliable ranking result. For rank aggregation, we consider multiple search engines for multiple part queries and adopt the idea of unsupervised rank fusion. Experimental results obtained using a challenging outdoor scene dataset show that our approach is an improvement over previous approaches despite the fact that we do not rely on domain-specific scene/part models nor supervision. Kanji Tanaka 0003 |
ICRA | 1 |
| 2015 | Cross-season place recognition using NBNN scene descriptorabstractWe propose a discriminative compact scene descriptor for single-view cross-season place recognition. Unlike previous bag-of-words approaches which rely on a library of vector quantized visual features, the proposed scene descriptor is based on a library of raw image data (such as available visual experience, images shared by other colleague robots, and publicly available image data on the web) that is directly mined to find nearest neighbor (NN) visual features (i.e., landmarks) for effectively explaining the input image. Our scene matcher adopts naive Bayes nearest neighbor (NBNN) techniques, where (1) raw visual features are used without vector quantization, and (2) image-to-class (rather than image-to-image) distance is used for scene comparison. Finally, we acquire a challenging cross-season place recognition dataset and validate the effectiveness of the proposed scene descriptor. Kanji Tanaka 0003 |
IROS | 1 |
| 2014 | Mining visual phrases for long-term visual SLAMabstractWe propose a discriminative and compact scene descriptor for single-view place recognition that facilitates long-term visual SLAM in familiar, semi-dynamic and partially changing environments. In contrast to popular bag-of-words scene descriptors, which rely on a library of vector quantized visual features, our proposed scene descriptor is based on a library of raw image data (such as an available visual experience, images shared by other colleague robots, and publicly available image data on the web) and directly mine it to find visual phrases (VPs) that discriminatively and compactly explain an input query / database image. Our mining approach is motivated by recent success in the field of common pattern discovery-specifically mining of common visual patterns among scenes-and requires only a single library of raw images that can be acquired at different time or day. Experimental results show that even though our scene descriptor is significantly more compact than conventional descriptors it has a relatively higher recognition performance. Kanji Tanaka 0003, Yuuto Chokushi, Masatoshi Ando |
IROS | 1 |
| 2013 | PartSLAM: Unsupervised part-based scene modeling for fast succinct map matchingabstractIn this paper, we explore the challenging 1-to-N map matching problem, which exploits a compact description of map data, to improve the scalability of map matching techniques used by various robot vision tasks. We propose a first method explicitly aimed at fast succinct map matching, which consists only of map-matching subtasks. These tasks include offline map matching attempts to find a compact part-based scene model that effectively explains each map using fewer larger parts. The tasks also include an online map matching attempt to efficiently find correspondence between the part-based maps. Our part-based scene modeling approach is unsupervised and uses common pattern discovery (CPD) between the input and known reference maps. This enables a robot to learn a compact map model without human intervention. We also present a practical implementation that uses the state-of-the-art CPD technique of randomized visual phrases (RVP) with a compact bounding box (BB) based part descriptor, which consists of keypoint and descriptor BBs. The results of our challenging map-matching experiments, which use a publicly available radish dataset, show that the proposed approach achieves successful map matching with significant speedup and a compact description of map data that is tens of times more compact. Although this paper focuses on the standard 2D pointset map and the BB-based part representation, we believe our approach is sufficiently general to be applicable to a broad range of map formats, such as the 3D point cloud map, as well as to general bounding volumes and other compact part representations. Shogo Hanada, Kanji Tanaka 0003 |
IROS | 2 |
| 2011 | Dictionary-based map compression for sparse feature mapsabstractObtaining a compact representation of a large size feature map built by mapper robots is a critical issue in the context of lightweight information sharing as well as Kolmogorov complexity. This map compression problem is explored from a novel perspective of dictionary-based data compression techniques in the paper. The primary contribution of the paper is proposal of the dictionary-based map compression approach. A map compression system is developed using RANSAC map matching and sparse coding as building blocks. Experiments show promising results in terms of map compression ratio, compression speed as well as the retrieval performance of compressed/decompressed maps. Tomomi Nagasaka, Kanji Tanaka 0003 |
ICRA | 2 |
| 2011 | An incremental scheme for dictionary-based compressive SLAMabstractObtaining a compact representation of a largesize pointset map built by mapper robots is a critical issue for recent SLAM applications. This ¿map compression¿ problem is explored from a novel perspective of dictionary-based map compression in the paper. The primary contribution of the paper is proposal of an incremental scheme for simultaneous mapping and map-compression applications. An incremental map compressor is presented by employing a modified RANSAC map-matching scheme as well as the compact projection visual search. Experiments show promising results in terms of compression speed, compactness of data and structure, as well as an application to the compression distance. Tomomi Nagasaka, Kanji Tanaka 0003 |
IROS | 2 |
| 2010 | Visual robot localization using compact binary landmarksabstractThis paper is concerned with the problem of mobile robot localization using a novel compact representation of visual landmarks. With recent progress in lifelong map-learning as well as in information sharing networks, compact representation of a large-size landmark database has become crucial. In this paper, we propose a compact binary code (e.g. 32bit code) landmark representation by employing the semantic hashing technique from web-scale image retrieval. We show how well such a binary representation achieves compactness of a landmark database while maintaining efficiency of the localization system. In our contribution, we investigate the cost-performance, the semantic gap, the saliency evaluation using the presented techniques as well as challenge to further reduce the resources (#bits) per landmark. Experiments using a high-speed car-like robot show promising results. Kouichirou Ikeda, Kanji Tanaka 0003 |
ICRA | 2 |
| 2009 | LSH-RANSAC: An incremental scheme for scalable localizationabstractThis paper addresses the problem of feature-based robot localization in large-size environments. With recent progress in SLAM techniques, it has become crucial for a robot to estimate the self-position in real-time with respect to a large-size map that can be incrementally build by other mapper robots. Self-localization using large-size maps have been studied in literature, but most of them assume that a complete map is given prior to the self-localization task. In this paper, we present a novel scheme for robot localization as well as map representation that can successfully work with large-size and incremental maps. This work combines our two previous works on incremental methods, iLSH and iRANSAC, for appearance-based and position-based localization. Kenichi Saeki, Kanji Tanaka 0003, Takeshi Ueda |
ICRA | 2 |
| 2008 | On the scalability of robot localization using high-dimensional featuresabstractThis study provides an investigation of scalability of mobile robot localization. In recent years, inference algorithms based on map-matching have proved their superior performance in large-scale environments. In this paper, the scalability is augmented by an ANN retrieval of high-dimensional descriptive features. The proposed algorithm is then exhaustively evaluated using large-size real maps, including over 100K feature maps. Takeshi Ueda, Kanji Tanaka 0003 |
ICPR | 2 |