EDBT 2026 Demo / reviewers in the wild / expert
Asako Kanezaki
dblp:37/7634
· DBLP profile ↗
35ranked-venue papers
11as first author
14since 2021 · last 2025
0000-0003-3217-1405ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 3 since 2021Systems, architecture and hardware · 11 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Embodied Navigation with Auxiliary Task of Action Description Prediction
Haru Kondoh, Asako Kanezaki |
ICCV | 2 |
| 2025 | Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language ModelsabstractAchieving zero-shot peg insertion, where inserting an arbitrary peg into an unseen hole without task-specific training, remains a fundamental challenge in robotics. This task demands a highly generalizable perception system capable of detecting potential holes, selecting the correct mating hole from multiple candidates, estimating its precise pose, and executing insertion despite uncertainties. While learning-based methods have been applied to peg insertion, they often fail to generalize beyond the specific peg-hole pairs encountered during training. Recent advancements in Vision-Language Models (VLMs) offer a promising alternative, leveraging large-scale datasets to enable robust generalization across diverse tasks. Inspired by their success, we introduce a novel zero-shot peg insertion framework that utilizes a VLM to identify mating holes and estimate their poses without prior knowledge of their geometry. This approach assumes a known peg pose and a leveled surface for insertion. Extensive experiments demonstrate that our method achieves 90.2% accuracy, significantly outperforming baselines in identifying the correct mating hole across a wide range of previously unseen peg-hole pairs, including 3D-printed objects, toy puzzles, and industrial connectors. Furthermore, we validate the effectiveness of our approach in a real-world connector insertion task on a backpanel of a PC, where our system successfully detects holes, identifies the correct mating hole, estimates its pose, and completes the insertion with a success rate of 88.3%. These results highlight the potential of VLM-driven zero-shot reasoning for enabling robust and generalizable robotic assembly. Masaru Yajima, Kei Ota, Asako Kanezaki, Rei Kawakami |
IROS | 3 |
| 2025 | Enhancing multimodal-input object goal navigation by leveraging large language models for inferring room-object relationship knowledge
Leyuan Sun, Asako Kanezaki, Guillaume Caron, Yusuke Yoshiyasu |
Adv. Eng. Informatics | 2 |
| 2024 | OP-Align: Object-Level and Part-Level Alignment for Self-supervised Category-Level Articulated Object Pose Estimation
Yuchen Che, Asako Kanezaki |
ECCV (75) | 3 |
| 2024 | Tactile Estimation of Extrinsic Contact Patch for Stable PlacementabstractPrecise perception of contact interactions is essential for fine-grained manipulation skills for robots. In this paper, we present the design of feedback skills for robots that must learn to stack complex-shaped objects on top of each other (see Fig. 1). To design such a system, a robot should be able to reason about the stability of placement from very gentle contact interactions. Our results demonstrate that it is possible to infer the stability of object placement based on tactile readings during contact formation between the object and its environment. In particular, we estimate the contact patch between a grasped object and its environment using force and tactile observations to estimate the stability of the object during a contact formation. The contact patch could be used to estimate the stability of the object upon release of the grasp. The proposed method is demonstrated in various pairs of objects that are used in a very popular board game. Kei Ota, Devesh K. Jha, Krishna Murthy Jatavallabhula, Asako Kanezaki, Josh Tenenbaum |
ICRA | 4 |
| 2024 | A framework for training larger networks for deep Reinforcement learningabstractAbstract The success of deep learning in computer vision and natural language processing communities can be attributed to the training of very deep neural networks with millions or billions of parameters, which can then be trained with massive amounts of data. However, a similar trend has largely eluded the training of deep reinforcement learning (RL) algorithms where larger networks do not lead to performance improvement. Previous work has shown that this is mostly due to instability during the training of deep RL agents when using larger networks. In this paper, we make an attempt to understand and address the training of larger networks for deep RL. We first show that naively increasing network capacity does not improve performance. Then, we propose a novel method that consists of (1) wider networks with DenseNet connection, (2) decoupling representation learning from the training of RL, and (3) a distributed training method to mitigate overfitting problems. Using this three-fold technique, we show that we can train very large networks that result in significant performance gains. We present several ablation studies to demonstrate the efficacy of the proposed method and some intuitive understanding of the reasons for performance gain. We show that our proposed method outperforms other baseline algorithms on several challenging locomotion tasks. Kei Ota, Devesh K. Jha, Asako Kanezaki |
Mach. Learn. | 3 |
| 2023 | Cross-Level Distillation and Feature Denoising for Cross-Domain Few-Shot Classification
Runqi Wang, Jianzhuang Liu, Asako Kanezaki |
ICLR | 4 |
| 2023 | H-SAUR: Hypothesize, Simulate, Act, Update, and Repeat for Understanding Object Articulations from InteractionsabstractThe world is filled with articulated objects that are difficult to determine how to use from vision alone, e.g., a door might open inwards or outwards. Humans handle these objects with strategic trial-and-error: first pushing a door then pulling if that doesn't work. We enable these capabilities in autonomous agents by proposing “Hypothesize, Simulate, Act, Update, and Repeat” (H-SAUR), a probabilistic generative framework that simultaneously generates a distribution of hypotheses about how objects articulate given input observations, captures certainty over hypotheses over time, and infer plausible actions for exploration and goal-conditioned manipulation. We compare our model with existing work in manipulating objects after a handful of exploration actions, on the PartNet-Mobility dataset. We further propose a novel PuzzleBoxes benchmark that contains locked boxes that require multiple steps to solve. We show that the proposed model significantly outperforms the current state-of-the-art articulated object manipulation framework, despite using zero training data. We further improve the test-time efficiency of H-SAUR by integrating a learned prior from learning-based vision models. Kei Ota, Hsiao-Yu Fish Tung, Kevin A. Smith 0001, Anoop Cherian, Tim K. Marks, Alan Sullivan, Asako Kanezaki, Josh Tenenbaum |
ICRA | 7 |
| 2023 | Multi-Goal Audio-Visual Navigation Using Sound Direction MapabstractOver the past few years, there has been a great deal of research on navigation tasks in indoor environments using deep reinforcement learning agents. Most of these tasks use only visual information in the form of first-person images to navigate to a single goal. More recently, tasks that simultaneously use visual and auditory information to navigate to the sound source and even navigation tasks with multiple goals instead of one have been proposed. However, there has been no proposal for a generalized navigation task combining these two types of tasks and using both visual and auditory information in a situation where multiple sound sources are goals. In this paper, we propose a new framework for this generalized task: multi-goal audio-visual navigation. We first define the task in detail, and then we investigate the difficulty of the multi-goal audio-visual navigation task relative to the current navigation tasks by conducting experiments in various situations. The research shows that multi-goal audio-visual navigation has the difficulty of the implicit need to separate the sources of sound. Next, to mitigate the difficulties in this new task, we propose a method named sound direction map (SDM), which dynamically localizes multiple sound sources in a learning-based manner while making use of past memories. Experimental results show that the use of SDM significantly improves the performance of multiple baseline methods, regardless of the number of goals. Haru Kondoh, Asako Kanezaki |
IROS | 2 |
| 2023 | EvIs-Kitchen: Egocentric Human Activities Recognition with Video and Inertial Sensor Data
Yuzhe Hao, Kuniaki Uto, Asako Kanezaki, Ikuro Sato, Rei Kawakami, Koichi Shinoda |
MMM (1) | 3 |
| 2022 | Object Memory Transformer for Object Goal NavigationabstractThis paper presents a reinforcement learning method for object goal navigation (ObjNav) where an agent navigates in 3D indoor environments to reach a target object based on long-term observations of objects and scenes. To this end, we propose Object Memory Transformer (OMT) that consists of two key ideas: 1) Object-Scene Memory (OSM) that enables to store long-term scenes and object semantics, and 2) Transformer that attends to salient objects in the sequence of previously observed scenes and objects stored in OSM. This mechanism allows the agent to efficiently navigate in the indoor environment without prior knowledge about the environments, such as topological maps or 3D meshes. To the best of our knowledge, this is the first work that uses a long-term memory of object semantics in a goal-oriented navigation task. Experimental results conducted on the AI2-THOR dataset show that OMT outperforms previous approaches in navigating in unknown environments. In particular, we show that utilizing the long-term object semantics information improves the efficiency of navigation. Rui Fukushima, Kei Ota, Asako Kanezaki, Yoko Sasaki, Yusuke Yoshiyasu |
ICRA | 3 |
| 2022 | OPIRL: Sample Efficient Off-Policy Inverse Reinforcement Learning via Distribution MatchingabstractInverse Reinforcement Learning (IRL) is attractive in scenarios where reward engineering can be tedious. However, prior IRL algorithms use on-policy transitions, which require intensive sampling from the current policy for stable and optimal performance. This limits IRL applications in the real world, where environment interactions can become highly expensive. To tackle this problem, we present Off-Policy Inverse Reinforcement Learning (OPIRL), which (1) adopts off-policy data distribution instead of on-policy and enables significant reduction of the number of interactions with the environment, (2) learns a reward function that is transferable with high generalization capabilities on changing dynamics, and (3) leverages mode-covering behavior for faster convergence. We demonstrate that our method is considerably more sample efficient and generalizes to novel environments through the experiments. Our method achieves better or comparable results on policy performance baselines with significantly fewer interactions. Furthermore, we empirically show that the recovered reward function generalizes to different tasks where prior arts are prone to fail. Hana Hoshino, Kei Ota, Asako Kanezaki, Rio Yokota |
ICRA | 3 |
| 2021 | Path Planning using Neural A* SearchabstractWe present Neural A*, a novel data-driven search method for path planning problems. Despite the recent increasing attention to data-driven path planning, machine learning approaches to search-based planning are still challenging due to the discrete nature of search algorithms. In this work, we reformulate a canonical A* search algorithm to be differentiable and couple it with a convolutional encoder to form an end-to-end trainable neural network planner. Neural A* solves a path planning problem by encoding a problem instance to a guidance map and then performing the differentiable A* search with the guidance map. By learning to match the search results with ground-truth paths provided by experts, Neural A* can produce a path consistent with the ground truth accurately and efficiently. Our extensive experiments confirmed that Neural A* outperformed state-of-the-art data-driven planners in terms of the search optimality and efficiency trade-off. Furthermore, Neural A* successfully predicted realistic human trajectories by directly performing search-based planning on natural image inputs. Ryo Yonetani, Tatsunori Taniai, Mohammadamin Barekatain, Mai Nishimura, Asako Kanezaki |
ICML | 5 |
| 2021 | RotationNet for Joint Object Categorization and Unsupervised Pose Estimation from Multi-View ImagesabstractWe propose a Convolutional Neural Network (CNN)-based model "RotationNet," which takes multi-view images of an object as input and jointly estimates its pose and object category. Unlike previous approaches that use known viewpoint labels for training, our method treats the viewpoint labels as latent variables, which are learned in an unsupervised manner during the training using an unaligned object dataset. RotationNet uses only a partial set of multi-view images for inference, and this property makes it useful in practical scenarios where only partial views are available. Moreover, our pose alignment strategy enables one to obtain view-specific feature representations shared across classes, which is important to maintain high accuracy in both object categorization and pose estimation. Effectiveness of RotationNet is demonstrated by its superior performance to the state-of-the-art methods of 3D object classification on 10- and 40-class ModelNet datasets. We also show that RotationNet, even trained without known poses, achieves comparable performance to the state-of-the-art methods on an object pose estimation dataset. Furthermore, our object ranking method based on classification by RotationNet achieved the first prize in two tracks of the 3D Shape Retrieval Contest (SHREC) 2017. Finally, we demonstrate the performance of real-world applications of RotationNet trained with our newly created multi-view image dataset using a moving USB camera. Asako Kanezaki, Yasuyuki Matsushita, Yoshifumi Nishida |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Efficient Exploration in Constrained Environments with Goal-Oriented Reference PathabstractIn this paper, we consider the problem of building learning agents that can efficiently learn to navigate in constrained environments. The main goal is to design agents that can efficiently learn to understand and generalize to different environments using high-dimensional inputs (a 2D map), while following feasible paths that avoid obstacles in obstacle-cluttered environment. To achieve this, we make use of traditional path planning algorithms, supervised learning, and reinforcement learning algorithms in a synergistic way. The key idea is to decouple the navigation problem into planning and control, the former of which is achieved by supervised learning whereas the latter is done by reinforcement learning. Specifically, we train a deep convolutional network that can predict collision-free paths based on a map of the environment- this is then used by an reinforcement learning algorithm to learn to closely follow the path. This allows the trained agent to achieve good generalization while learning faster. We test our proposed method in the recently proposed Safety Gym suite that allows testing of safety-constraints during training of learning agents. We compare our proposed method with existing work and show that our method consistently improves the sample efficiency and generalization capability to novel environments. Kei Ota, Yoko Sasaki, Devesh K. Jha, Yusuke Yoshiyasu, Asako Kanezaki |
IROS | 5 |
| 2020 | Incremental multi-view object detection from a moving cameraabstractObject detection in a single image is a challenging problem due to clutters, occlusions, and a large variety of viewing locations. This task can benefit from integrating multi-frame information captured by a moving camera. In this paper, we propose a method to increment object detection scores extracted from multiple frames captured from different viewpoints. For each frame, we run an efficient end-to-end object detector that outputs object bounding boxes, each of which is associated with the scores of categories and poses. The scores of detected objects are then stored in grid locations in 3D space. After observing multiple frames, the object scores stored in each grid location are integrated based on the best object pose hypothesis. This strategy requires the consistency of object categories and poses among multiple frames, and thus it significantly suppresses miss detections. The performance of the proposed method is evaluated on our newly created multi-class object dataset captured in robot simulation and real environments, as well as on a public benchmark dataset. Takashi Konno, Ayako Amma, Asako Kanezaki |
MMAsia | 3 |
| 2020 | Unsupervised Learning of Image Segmentation Based on Differentiable Feature ClusteringabstractThe usage of convolutional neural networks (CNNs) for unsupervised image segmentation was investigated in this study. Similar to supervised image segmentation, the proposed CNN assigns labels to pixels that denote the cluster to which the pixel belongs. In unsupervised image segmentation, however, no training images or ground truth labels of pixels are specified beforehand. Therefore, once a target image is input, the pixel labels and feature representations are jointly optimized, and their parameters are updated by the gradient descent. In the proposed approach, label prediction and network parameter learning are alternately iterated to meet the following criteria: (a) pixels of similar features should be assigned the same label, (b) spatially continuous pixels should be assigned the same label, and (c) the number of unique labels should be large. Although these criteria are incompatible, the proposed approach minimizes the combination of similarity loss and spatial continuity loss to find a plausible solution of label assignment that balances the aforementioned criteria well. The contributions of this study are four-fold. First, we propose a novel end-to-end network of unsupervised image segmentation that consists of normalization and an argmax function for differentiable clustering. Second, we introduce a spatial continuity loss function that mitigates the limitations of fixed segment boundaries possessed by previous work. Third, we present an extension of the proposed method for segmentation with scribbles as user input, which showed better accuracy than existing methods while maintaining efficiency. Finally, we introduce another extension of the proposed method: unseen image segmentation by using networks pre-trained with a few reference images without re-training the networks. The effectiveness of the proposed approach was examined on several benchmark datasets of image segmentation. Wonjik Kim, Asako Kanezaki, Masayuki Tanaka 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Salient Object Detection on Hyperspectral Images Using Features Learned from Unsupervised Segmentation TaskabstractVarious saliency detection algorithms from color images have been proposed to mimic eye fixation or attentive object detection response of human observers for the same scenes. However, developments on hyperspectral imaging systems enable us to obtain redundant spectral information of the observed scenes from the reflected light source from objects. A few studies using low-level features on hyper-spectral images demonstrated that salient object detection can be achieved. In this work, we proposed a salient object detection model on hyperspectral images by applying manifold ranking (MR) on self-supervised Convolutional Neural Network (CNN) features (high-level features) from unsupervised image segmentation task. Self-supervision of CNN continues until clustering loss or saliency maps converges to a defined error between each iteration. Finally, saliency estimations is done as the saliency map at last iteration when the self-supervision procedure terminates with convergence. Experimental evaluations demonstrated that proposed saliency detection algorithm on hyperspectral images is outperforming state-of-the-arts hyperspectral saliency models including the original MR based saliency model. Nevrez Imamoglu, Guanqun Ding, Y. Fang, Asako Kanezaki, Toru Kouyama, Ryosuke Nakamura |
ICASSP | 4 |
| 2019 | A3C Based Motion Learning for an Autonomous Mobile Robot in CrowdsabstractThe paper proposes a motion planning method using a deep reinforcement learning algorithm, Asynchronous Advantage Actor-Critic (A3C). For mobile robot navigation tasks in crowds, existing path planning based approaches are limited because the surrounding environments change dynamically. The correct motion in such a dynamic environment is underspecified, and a reinforcement learning approach is suitable for generating applicable motion. We propose an A3C based motion planning method for acquiring robot motion for a robot moving through crowds. The proposed method is evaluated in simulated crowds of pedestrians. The experiment section shows the basic performance depending on training parameters and some generated motion examples in the simulator. The learning results using real pedestrian motion are also shown. Yoko Sasaki, Syusuke Matsuo, Asako Kanezaki, Hiroshi Takemura |
SMC | 3 |
| 2018 | "Change the Changeable" Framework for Implementation Research in Health
Mikiko Oono, Yoshifumi Nishida, Koji Kitamura, Asako Kanezaki, Tatsuhiro Yamanaka |
CSEDU (2) | 4 |
| 2018 | RotationNet: Joint Object Categorization and Pose Estimation Using Multiviews From Unsupervised ViewpointsabstractWe propose a Convolutional Neural Network (CNN)-based model "RotationNet," which takes multi-view images of an object as input and jointly estimates its pose and object category. Unlike previous approaches that use known viewpoint labels for training, our method treats the viewpoint labels as latent variables, which are learned in an unsupervised manner during the training using an unaligned object dataset. RotationNet is designed to use only a partial set of multi-view images for inference, and this property makes it useful in practical scenarios where only partial views are available. Moreover, our pose alignment strategy enables one to obtain view-specific feature representations shared across classes, which is important to maintain high accuracy in both object categorization and pose estimation. Effectiveness of RotationNet is demonstrated by its superior performance to the state-of-the-art methods of 3D object classification on 10- and 40-class ModelNet datasets. We also show that RotationNet, even trained without known poses, achieves the state-of-the-art performance on an object pose estimation dataset. Asako Kanezaki, Yasuyuki Matsushita, Yoshifumi Nishida |
CVPR | 1 |
| 2018 | Unsupervised Image Segmentation by BackpropagationabstractWe investigate the use of convolutional neural networks (CNNs) for unsupervised image segmentation. As in the case of supervised image segmentation, the proposed CNN assigns labels to pixels that denote the cluster to which the pixel belongs. In the unsupervised scenario, however, no training images or ground truth labels of pixels are given beforehand. Therefore, once when a target image is input, we jointly optimize the pixel labels together with feature representations while their parameters are updated by gradient descent. In the proposed approach, we alternately iterate label prediction and network parameter learning to meet the following criteria: (a) pixels of similar features are desired to be assigned the same label, (b) spatially continuous pixels are desired to be assigned the same label, and (c) the number of unique labels is desired to be large. Although these criteria are incompatible, the proposed approach finds a plausible solution of label assignment that balances well the above criteria' which demonstrates good performance on a benchmark dataset of image segmentation. Asako Kanezaki |
ICASSP | 1 |
| 2016 | Recognizing Activities of Daily Living with a Wrist-Mounted CameraabstractWe present a novel dataset and a novel algorithm for recognizing activities of daily living (ADL) from a first-person wearable camera. Handled objects are crucially important for egocentric ADL recognition. For specific examination of objects related to users' actions separately from other objects in an environment, many previous works have addressed the detection of handled objects in images captured from head-mounted and chest-mounted cameras. Nevertheless, detecting handled objects is not always easy because they tend to appear small in images. They can be occluded by a user's body. As described herein, we mount a camera on a user's wrist. A wrist-mounted camera can capture handled objects at a large scale, and thus it enables us to skip the object detection process. To compare a wrist-mounted camera and a head-mounted camera, we also developed a novel and publicly available dataset1that includes videos and annotations of daily activities captured simultaneously by both cameras. Additionally, we propose a discriminative video representation that retains spatial and temporal information after encoding the frame descriptors extracted by convolutional neural networks (CNN). Katsunori Ohnishi, Atsushi Kanehira, Asako Kanezaki, Tatsuya Harada |
CVPR | 3 |
| 2016 | IBC127: Video dataset for fine-grained bird classificationabstractBeyond general object recognition whereby general categories such as dogs and cats are estimated from images, Fine-grained visual categorization (FGVC) is a new trend that goes beyond general object recognition - where general categories such as cats and dogs are estimated from images - to classify fine-grained categories of objects (or animals) such as poodles or bulldogs. It is difficult to distinguish between categories with similar appearance, e.g., sparrows and hummingbirds, using image features alone. Consequently, motion features extracted from videos are effective for classifying similar animals. In this paper, we demonstrate the effectiveness of motion features for FGVC of animals. We use our novel dataset that consists of videos of fine-grained categories of birds collected from the internet. Our dataset is publicly available for academics. Tomoaki Saito, Asako Kanezaki, Tatsuya Harada |
ICME | 2 |
| 2015 | 3D Selective Search for obtaining object candidatesabstractWe propose a new method for obtaining object candidates in 3D space. Our method requires no learning, has no limitation of object properties such as compactness or symmetry, and therefore produces object candidates using a completely general approach. This method is a simple combination of Selective Search, which is a non-learning-based objectness detector working in 2D images, and a supervoxel segmentation method, which works with 3D point clouds. We made a small but non-trivial modification to supervoxel segmentation; it brings better “seeding” for supervoxels, which produces more proper object candidates as a result. Our experiments using a couple of publicly available RGB-D datasets demonstrated that our method outperformed state-of-the-art methods of generating object proposals in 2D images. Asako Kanezaki, Tatsuya Harada |
IROS | 1 |
| 2015 | Probabilistic Semi-Canonical Correlation Analysisabstractanonical Correlation Analysis (CCA) requires paired multimodal data to ascertain the relation between two variables. However, it is generally difficult to collect a sufficient amount of paired data of two variables as training samples. This fact leads individual samples of unpaired variables to be additional resources for learning CCA, which are not only able to increase the number of training samples; they are also effective to remove the learning bias caused by the variables' missing patterns. As described in this paper, we propose a novel model of probabilistic CCA by considering the mechanism of data missing. Our method enables widespread applications such as semi-supervised learning via partially labeled training samples and analysis of sensory data which are lacking under certain circumstances. We demonstrate the superior performance of parameter estimation as well as an application of image annotation, compared with existing methods. Chie Kamada, Asako Kanezaki, Tatsuya Harada |
ACM Multimedia | 2 |
| 2014 | Learning Similarities for Rigid and Non-rigid Object DetectionabstractIn this paper, we propose an optimization method for estimating the parameters that typically appear in graph-theoretical formulations of the matching problem for object detection. Although several methods have been proposed to optimize parameters for graph matching in a way to promote correct correspondences and to restrict wrong ones, our approach is novel in the sense that it aims at improving performance in the more general task of object detection. In our formulation, similarity functions are adjusted so as to increase the overall similarity among a reference model and the observed target, and at the same time reduce the similarity among reference and "non-target" objects. We evaluate the proposed method in two challenging scenarios, namely object detection using data captured with a Kinect sensor in a real environment, and intrinsic metric learning for deformable shapes, demonstrating substantial improvements in both settings. Asako Kanezaki, Emanuele Rodolà, Daniel Cremers, Tatsuya Harada |
3DV | 1 |
| 2014 | Mirror reflection invariant HOG descriptors for object detectionabstractHistogram of Oriented Gradients (HOG) [1] descriptors have been widely used for object detection. An important limitation is that these descriptors tend to vary considerably when objects are horizontally flipped, as is often the case. We propose novel MI-HOG descriptors that are obtained by transforming HOG descriptors to be invariant to mirror reflection. In their extraction process, we consider not only the transform of independent elements but also the combination of those in different location and in orientation, which yields better performance. We showed a greater than 10 % increase in average precision compared to HOG descriptors. Asako Kanezaki, Yusuke Mukuta, Tatsuya Harada |
ICIP | 1 |
| 2014 | Hard negative classes for multiple object detectionabstractWe propose an efficient method to train multiple object detectors simultaneously using a large scale image dataset. The one-vs-all approach that optimizes the boundary between positive samples from a target class and negative samples from the others has been the most standard approach for object detection. However, because this approach trains each object detector independently, the scores are not balanced between object classes. The proposed method combines ideas derived from both detection and classification in order to balance the scores across all object classes. We optimized the boundary between target classes and their “hard negative” samples, just as in detection, while simultaneously balancing the detector scores across object classes, as done in multi-class classification. We evaluated the performances on multi-class object detection using a subset of the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) 2011 dataset and showed our method outperformed a de facto standard method. Asako Kanezaki, Sho Inaba, Yoshitaka Ushiku, Yuya Yamashita, Hiroshi Muraoka, Yasuo Kuniyoshi, Tatsuya Harada |
ICRA | 1 |
| 2014 | Automatic Image Synthesis from Keywords Using Scene ContextabstractText is one of the simplest way to express one's idea, and an image is one of the most impactive way to do so. Therefore, if a system can synthesize an image from text without direct user manipulation, novel image synthesis applications will be opened to users without artistic skills. In such a system, which objects to synthesize will be declared in texts. However, information about positional relations and scale of objects is not much provided and must be estimated using common sense. As described in this paper, we develop a system that can automatically synthesize objects to an image, given the background image and class name of the target synthesizing object. With the inputs as the background image and keywords, images for synthesizing objects are searched automatically. Although some previously developed systems that can synthesize an image from sketches and paintings, this is the first system that can estimate the position, scale, and appearance of objects and automatically synthesize them to images without direct user input. We propose a scene context, which indicates the position, scale, and appearance of synthesizing objects. The contribution of this paper is twofold: (1) the scene context extraction method for automatic image synthesis and (2) application of automatic image synthesis using the scene context. Sho Inaba, Asako Kanezaki, Tatsuya Harada |
ACM Multimedia | 2 |
| 2014 | Clothing Retrieval Based on Local Similarity with Multiple ImagesabstractRecently, the online shopping market has been expanded, which has advanced studies of clothing retrieval via image search. For this study, we develop a novel clothing retrieval system considering local similarity, where users can retrieve their desired clothes which are globally similar to an image and partially similar to another image. We propose a method of coding global features by merging local descriptors extracted from multiple images. Furthermore, we design a system that re-evaluates output of similar image search by the similarity of local regions. We demonstrated that our method increased the probability of users finding their desired clothes from 39.7%-55.1%, compared to a standard similar image search system with global features of a single image. Statistical significance is proven using t-tests. Masaru Mizuochi, Asako Kanezaki, Tatsuya Harada |
ACM Multimedia | 2 |
| 2013 | Weakly-supervised multi-class object detection using multi-type 3D featuresabstractWe propose a weakly-supervised learning method for object detection using color and depth images of a real environment attached with object labels. The proposed method applies Multiple Instance Learning to find proper instances of the objects in training images. This method is novel in the sense that it learns multiple objects simultaneously in a way to balance the scores of each training sample across all object classes. Moreover, we combine 3D features considering different properties, that is, color texture, grayscale texture, and surface curvature, to improve the performance. We show that our method surpasses a conventional method using color and depth images. Furthermore, we evaluate its performance with our new dataset consisting of color and depth images with weak labels of 100 objects and various backgrounds. Asako Kanezaki, Yasuo Kuniyoshi, Tatsuya Harada |
ACM Multimedia | 1 |
| 2011 | Fast object detection for robots in a cluttered indoor environment using integral 3D feature tableabstractRealizing automatic object search by robots in an indoor environment is one of the most important and challenging topics in mobile robot research. If the target object does not exist in a nearby area, the obvious strategy is to go to the area in which it was last observed. We have developed a robot system that collects 3D-scene data in an indoor environment during automatic routine crawling, and also detects objects quickly through a global search of the collected 3D-scene data. The 3D-scene data can be obtained automatically by transforming color images and range images into a set of color voxel data using self-location information. To detect an object, the system moves the bounding box of the target object by a certain step in the color voxel data, extracts 3D features in each box region, and computes the similarity between these features and the target object's features, using an appropriate feature projection learned beforehand. Taking advantage of the additive property of our 3D features, both feature extraction and similarity calculation are considerably accelerated. In the object learning process, the system obtains the feature-projection matrix by weighting unique features of the target object rather than its common features, resulting in reducing object detection errors. Asako Kanezaki, Tatsuya Harada, Yasuo Kuniyoshi |
ICRA | 1 |
| 2010 | High-speed 3D object recognition using additive features in a linear subspaceabstractIn this paper we propose a method of high-speed 3D object recognition using linear subspace method and our 3D features. This method can be applied to partial models with any size in any posture. Although it is becoming easy to obtain textured 3D models by a 3D scanner, there are few methods for 3D object recognition which take into account both shape and textures of objects. Moreover, it is difficult to achieve high-speed processing of large 3D data. Our 3D features consider the co-occurrence of shape and colors of an object's surface. The additive property of these features makes it possible to calculate the similarity between a query part and the subspace of each object in a database without division, and therefore the time for recognition is quite short. In the experiments, we compare our method with conventional methods using Spin-Images and Textured Spin-Images. We show that our method is appropriate for 3D object recognition. Asako Kanezaki, Hideki Nakayama, Tatsuya Harada, Yasuo Kuniyoshi |
ICRA | 1 |
| 2010 | Partial matching of real textured 3D objects using color cubic higher-order local auto-correlation features
Asako Kanezaki, Tatsuya Harada, Yasuo Kuniyoshi |
Vis. Comput. | 1 |