Tohru Kamiya

dblp:275/2469 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-2872-9018ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Generalizable Zero-Shot Object Pose Estimation for Bin-Picking
abstract
Unordered grasping in industrial robotic manipulation requires precise six-degree-of-freedom (6D) pose estimation. However, existing methods often struggle with unknown objects and require retraining, limiting their practicality. Traditional 3D point-pair feature methods, while training-free, perform poorly with textured symmetric objects. We propose a generalizable approach for zero-shot 6 D pose estimation without retraining. Our method consists of two steps: generating CAD-based templates through real-time rendering for coarse pose estimation, and refining poses using semantic point-pair features aligned with the camera viewpoint. We conducted experiments on seven core datasets from the Benchmark for 6D Object Pose Estimation (BOP) challenge, and the results are publicly available on the BOP website. Integration into a robotic grasping system further highlights its high precision and fast execution, making it ideal for applications such as bin-picking. (GZS6D-BP) https://bop.felk.cvut.cz/leaderboards/.
Zijiang Zhang, Huimin Lu 0001, Jintong Cai, Tohru Kamiya, Seiichi Serikawa
ICRA4
2025 Classification of Respiratory Sounds Based on Hybrid Convolutional Recurrent Neural Network
abstract
Nearly 8 million people suffer and die from respiratory diseases every year. Therefore, to reduce the number of deaths which are caused by the diseases, early detection and early treatment of respiratory diseases are required as global issues, and several techniques have been proposed until now. Currently, the ICBHI (International Conference on Biomedical and Health Informatics) 2017 Challenge Dataset has been released for research on respiratory sound analysis, and respiratory sound classification methods using this dataset have been proposed worldwide. The authors also proposed a respiratory sound classification method using an improved CRNN (Convolutional Recurrent Neural Network), which is a combination of CNN and RNN (Recurrent Neural Network) with some modifications. However, it was still difficult to classify by image features alone due to noise such as voice in the respiratory sound data. To overcome this problem, we try to classify breath sounds automatically using a deep learning model that considers the features of the raw breath sound data. To extract the sound features of the raw respiratory sound data, we use a 1D-CRNN, which is a 1D-CNN reconstruction of an improved CRNN proposed in a previous study of ours. Then, it is combined with the deep features obtained by our previous improved CRNN (2D-CRNN) for the final classification. The proposed method achieves AUC (Area Under Curve) of 0.92, sensitivity of 0.75, specificity of 0.86, and ICBHI score of 0.80 based on the ROC (Receiver Operating Characteristic) analysis, respectively, which are the highest values compared to the other methods under the same experimental conditions.
Naoki Asatani, Tohru Kamiya, Shingo Mabu, Shoji Kido
SMC2
2023 Pose Estimation of Point Sets Using Residual MLP in Intelligent Transportation Infrastructure
abstract
6D pose estimation of arbitrary objects is a crucial topic for intelligent transportation infrastructure measurement. However, some external environmental factors and the characteristics of the object itself impact the accuracy of the object’s pose estimation in practical applications. In this paper, we propose a new multi-class dataset ICD-4 (Industrial car Components Dataset) for 6D object pose estimation, which mainly includes four component categories, and every category takes 20,000 different scenarios. ICD-4 dataset delivers quite a few research challenges involving the range of object pose transformations and has significant research value for small-scale pose estimation tasks. We also propose an innovative method PoseMLP, a pose estimation network that uses residual MLP (multilayer perceptron) modules to predict the 6D pose estimation directly. Simultaneously, the experimental results demonstrate the effectiveness and reliability of the proposed method.
Yujie Li 0001, Zhiyun Yin, Yuchao Zheng 0001, Huimin Lu 0001, Tohru Kamiya, Yoshihisa Nakatoh, Seiichi Serikawa
IEEE Trans. Intell. Transp. Syst.5
2023 Multidimensional Deformable Object Manipulation Based on DN-Transporter Networks
abstract
In the process of transportation, the handling and loading methods of rigid objects are becoming more and more perfect. However, whether in today’s transportation system or in daily life, such as packing objects or sorting cables before transportation, the manipulation of deformable objects has been always inevitable and has attracted more and more attention. Due to the super degrees of freedom and the unpredictable physical state of deformed objects. It is difficult for robots to complete tasks under the environment of the deformable object. Therefore, we present a method based on imitation learning. In the generated expert demonstration, the agent is offered to learn the state sequence, and then imitate the expert’s trajectory sequence which avoid the above-mentioned difficulties. In addition, compared with the baseline method, our proposed DN-Transporter Networks are more competitive in a simulation environment involving cloth, ropes or bags.
Yadong Teng, Huimin Lu 0001, Yujie Li 0001, Tohru Kamiya, Yoshihisa Nakatoh, Seiichi Serikawa, Pengxiang Gao
IEEE Trans. Intell. Transp. Syst.4
2022 Grasp Position Estimation from Depth Image Using Stacked Hourglass Network Structure
abstract
In recent years, robots have been used not only in factories. However, most robots currently used in such places can only perform the actions programmed to perform in a predefined space. For robots to become widespread in the future, not only in factories, distribution warehouses, and other places but also in homes and other environments where robots receive complex commands and their surroundings are constantly being updated, it is necessary to make robots intelligent. Therefore, this study proposed a deep learning grasp position estimation model using depth images to achieve intelligence in pick-and-place. This study used only depth images as the training data to build the deep learning model. Some previous studies have used RGB images and depth images. However, in this study, we used only depth images as training data because we expect the inference to be based on the object's shape, independent of the color information of the object. By performing inference based on the target object's shape, the deep learning model is expected to minimize the need for re-training when the target object package changes in the production line since it is not dependent on the RGB image. In this study, we propose a deep learning model that focuses on the stacked encoder-decoder structure of the Stacked Hourglass Network. We compared the proposed method with the baseline method in the same evaluation metrics and a real robot, which shows higher accuracy than other methods in previous studies.
Keisuke Hamamoto, Huimin Lu 0001, Yujie Li 0001, Tohru Kamiya, Yoshihisa Nakatoh, Seiichi Serikawa
COMPSAC4
2021 Weakly unsupervised conditional generative adversarial network for image-based prognostic prediction for COVID-19 patients based on chest CT
abstract
Because of the rapid spread and wide range of the clinical manifestations of the coronavirus disease 2019 (COVID-19), fast and accurate estimation of the disease progression and mortality is vital for the management of the patients. Currently available image-based prognostic predictors for patients with COVID-19 are largely limited to semi-automated schemes with manually designed features and supervised learning, and the survival analysis is largely limited to logistic regression. We developed a weakly unsupervised conditional generative adversarial network, called pix2surv, which can be trained to estimate the time-to-event information for survival analysis directly from the chest computed tomography (CT) images of a patient. We show that the performance of pix2surv based on CT images significantly outperforms those of existing laboratory tests and image-based visual and quantitative predictors in estimating the disease progression and mortality of COVID-19 patients. Thus, pix2surv is a promising approach for performing image-based prognostic predictions.
Tomoki Uemura, Janne Näppi, Chinatsu Watari, Toru Hironaka, Tohru Kamiya, Hiroyuki Yoshida
Medical Image Anal.5
2021 Construction of a Hierarchical Feature Enhancement Network and Its Application in Fault Recognition
abstract
Industrial Internet of Things (IIoT) provide significant support for observing and controlling industrial machinery. In this article, a novel hierarchical feature enhancement network (HFEN) is proposed by combining signal processing and representation learning. The signal processing block extracts features with definite physical significance. Then, the representability of the physical features is improved by connecting stacked denoising autoencoders and squeeze-and-excitation networks. A novel two-stream architecture is designed for HFEN to fuse two types of features. Consequently, HFEN can extract features that can be analyzed for physical significance and that are also representative in terms of recognizable patterns. The experimental results prove that the performance of HFEN is satisfactory in terms of accuracy and efficiency when compared to other methods. Finally, this article also aims to demonstrate the potential of a new pairing that fuses the model- and data-driven strategies for IIoT.
Zhe Chen 0004, Huimin Lu 0001, Shiqing Tian, Junlin Qiu, Tohru Kamiya, Seiichi Serikawa
IEEE Trans. Ind. Informatics5
2020 Deep Learning for Visual Segmentation: A Review
abstract
Big data-driven deep learning methods have been widely used in image or video segmentation. The main challenge is that a large amount of labeled data is required in training deep learning models, which is important in real-world applications. To the best of our knowledge, there exist few researches in the deep learning-based visual segmentation. To this end, this paper summarizes the algorithms and current situation of image or video segmentation technologies based on deep learning and point out the future trends. The characteristics of segmentation that based on semi-supervised or unsupervised learning, all of the recent novel methods are summarized in this paper. The principle, advantages and disadvantages of each algorithms are also compared and analyzed.
Yujie Li 0001, Huimin Lu 0001, Tohru Kamiya, Seiichi Serikawa
COMPSAC4