Toshinori Hosoi

dblp:06/6688 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MDCN-PS: Monocular-Depth-Guided Coarse Normal Attention for Robust Photometric Stereo
abstract
Photometric Stereo (PS) is a technique for estimating surface normals from images illuminated by multiple light sources. However, when the target object has a complex shape or the light sources are not appropriately arranged, certain regions may experience severe shadows, leading to insufficient information for accurate estimation. In this paper, we propose a Monocular-Depth-guided Coarse Normal attention for Photometric Stereo (MDCN-PS). The MDCN-PS can effectively combine monocular depth from a single image with PS with multiple light sources by a Photometric Stereo network Adaptor (PS Adaptor) with Coarse Normal Attention. The key is to use the coarse normals obtained from Monocular Depth Estimation as supplementary information, which can improve accuracy in regions where the light source is limited due to severe shadows or inhomogeneous light source distribution. Comprehensive experiments on real-world and synthetic datasets show that the proposed method achieved an accuracy improvement of 1.2 points in real-world datasets when limited to two input images and of 3.1 points in synthetic datasets in mean angular error compared to existing methods. Qualitative results also demonstrated that our method improves accuracy in areas with insufficient lighting patterns due to shadows.
Masahiro Yamaguchi 0001, Takashi Shibata 0001, Shoji Yachida, Keiko Yokoyama, Toshinori Hosoi
WACV5
2024 Robust 3D Semantic Segmentation With Incomplete Point Clouds Based on Sequential Frame Sampling
abstract
This paper proposes a method for learning 3D semantic segmentation robust to incomplete point clouds. Our method first generates pseud-incomplete point clouds from original 3D point clouds by sequential frame sampling that creates multiple subsets considering the continuity of an RGB-D sequence for reproducing incomplete areas in the point clouds. It then simultaneously learns completion networks and semantic segmentation networks with the pseud-incomplete point clouds. We evaluate our method on the 3D semantic segmentation task. Experimental results on ScanNet v2, an indoor environment, show that our method improves mIoU by 0.4 points for the original point clouds and 6.3 points for the incomplete point clouds compared with a conventional method. Experimental results on WorkPlace Dataset, an outdoor environment, show that our method improves mIoU by 6.5 points for the original point clouds and 11.1 points for the incomplete point clouds compared with the conventional method. These results improve the safety and operability of environmental awareness in applications such as robotics.
Masahiro Yamaguchi 0001, Kyota Higa, Toshinori Hosoi, Takashi Shibata 0001
ICIP3
2024 Cyclic Learning of a Frame Downsampler and a Recognion Model in High-Speed Camera Image Recognition
Shigeaki Namiki, Takuya Ogawa, Keiko Yokoyama, Shoji Yachida, Toshinori Hosoi
ICPR (17)5
2024 PostAugment: Adversarial Data Augmentation with Hard Sample Suppression by Incorrect Class Likelihood
Azusa Sawada, Takashi Shibata 0001, Keiko Yokoyama, Shoji Yachida, Toshinori Hosoi
ICPR (10)5
2024 Toward Micro Eye Movement Detection in Practice: Stand-alone Eye Tracker with High Resolution and Wide Measurement Range
abstract
Detecting the micromovements of eyes that reflect a person’s inner state can be an essential step in many applications, but most eye trackers need to fix the subject’s head using a chin rest to obtain sufficient data quality. We propose a stand-alone eye tracker that utilizes two 500 fps cameras, a pair of rotating mirrors for gaze control, a liquid lens for focus control and an intensity-controllable light source, and describe how the proposed system works in real-time. The experimental results show that our system covers more than twice as wide a measurement range in the depth direction as the conventional eye tracker while achieving sufficient data quality to analyze microsaccades with an amplitude of down to 0.2 deg. We also examine the practical use of the proposed system for microsaccade detection.
Keiko Yokoyama, Tomohiro Sueishi, Michiaki Inoue, Shoji Yachida, Toshinori Hosoi, Masatoshi Ishikawa
IROS5
2024 FRoG-MOT: Fast and Robust Generic Multiple-Object Tracking by IoU and Motion-State Associations
abstract
This paper proposes a generic multi-object tracking (MOT) algorithm that is robust to unexpected motion changes for generic objects. Deep learning has dramatically been improving MOT performances. Nevertheless, state-of-the-art tracking algorithms are still sensitive to unexpected motion changes and the generic object target beyond person tracking. This is because standard MOT benchmark datasets such as MOT17 mainly consist of persons in a crowd, often lacking unexpected shape and motion changes; thus, these issues have yet to be focused on. We propose a simple-yet-effective MOT framework that can dynamically improve tracking continuity by associating each target based on adaptively modified motion states. The keys are 1) to represent the target motions using multiple motion states that have weak correlations with each other and 2) to modify those states that have the lowest similarity to past states as outliers. Our approach can improve trajectory continuity and robustness to unexpected motion changes for generic objects. Comprehensive experiments have confirmed that our framework is comparable to existing state-of-the-art methods on a standard dataset and outperforms those algorithms on the GMOT dataset with an overall 2% improvement in IDF1, a measure of tracking continuity.
Takuya Ogawa, Takashi Shibata 0001, Toshinori Hosoi
WACV3
2023 ICCL: Self-Supervised Intra- and Cross-Modal Contrastive Learning with 2D-3D Pairs for 3D Scene Understanding
abstract
This paper proposes self-supervised intra- and cross-modal contrastive learning (ICCL) with 2D-3D pairs for 3D scene understanding. Learning from different modalities has produced substantial results in self-supervised learning. Our method learns a model with high transferability by minimizing contrastive losses based on 2D, 3D, and 2D-3D features. Compared with a conventional approach minimizing 3D and 2D-3D contrastive losses, our method minimizes a 2D contrastive loss in addition to them. It leads to learning a better feature representation. We evaluate the transferability by conducting three downstream tasks, including object classification and part segmentation. The results of the 3D object classification show that our approach achieves an accuracy of 91.7 and 85.4 (0.5 and 3.7 points higher than the conventional method). The results of the few-shot object classification and the part segmentation show that our accuracy is equal to or higher than conventional methods. With better feature representation for 2D images and 3D point clouds, transfer learning can be more accessible, enabling the implementation of various applications in many fields.
Kyota Higa, Masahiro Yamaguchi 0001, Toshinori Hosoi
ICIP3
2022 Multi Object Tracking Based on Uncertainty-Aware RE-ID
abstract
In multi-object tracking (MOT), many tracking-by-detection methods have been proposed, and intersection over union (IoU) is the most common re-identification (re-ID) method. However, IoU does not care about target motion and so that weak for large change of position, size, or disturbances that often occur in MOT scenes. In addition, only detections with a certain level of confidence are used for re-ID to avoid noise, which also exclude some positive detections and lead to fragmentation of a tracklet. In this paper, we propose a robust MOT method that represents the target motion by a combining multiple indices and performing re-ID by allowing uncertainty in each index which means one of them does not have to be satisfied for re-ID to deal with such changes and disturbances. Our re-ID also makes it possible to effectively utilize positive detections that were previously excluded, achieving top performance in comprehensive metrics MOTA, IDF1, and HOTA, and in robustness metrics such as AssA, Frag, and MT comparing to the state-of-the-art methods in MOT17 benchmark.
Takuya Ogawa, Shoji Yachida, Toshinori Hosoi
ICIP3
2022 Convolutional Neural Networks for Time-dependent Classification of Variable-length Time Series
abstract
Time series data are often obtained only within a limited time range due to interruptions during observation process. To classify such partial time series, we need to account for 1) the variable-length data drawn from 2) different timestamps. To address the first problem, existing convolutional neural networks use global pooling after convolutional layers to cancel the length differences. This architecture suffers from the trade-off between incorporating entire temporal correlations in long data and avoiding feature collapse for short data. To resolve this trade-off, we propose Adaptive Multi-scale Pooling, which aggregates features from an adaptive number of layers, i.e., only the first few layers for short data and more layers for long data. Furthermore, to address the second problem, we introduce Temporal Encoding, which embeds the observation timestamps into the intermediate features. Experiments on our private dataset and the UCR/UEA time series archive show that our modules improve classification accuracy especially on short data obtained as partial time series.
Azusa Sawada, Taiki Miyagawa, Akinori F. Ebihara, Shoji Yachida, Toshinori Hosoi
IJCNN5
2020 Reducing False Positives in Object Tracking with Siamese Network
abstract
We propose a robust long-term object tracking method that resolves the fundamental cause of the drift and loss of a target in visual object tracking. The proposed method consists of “sampling area extension”, which prevents a tracking result from drifting to other objects by learning false positive samples in advance (before they enter the search region of the target), and “adaptive search based on motion models”, which prevents a tracking result from drifting to other objects and avoids the loss of the target by using not only appearance features but also motion models to adaptively search for the target. Experiments conducted on long-term tracking dataset showed that our first technique improved robustness by 16.6% while the second technique improved robustness by 15.3%. By combining both, our method achieved 21.7% and 9.1% improvement for the robustness and precision, and the processing speed became 3.3 times faster. Additional experiments showed that our method achieved the top robustness among state-of-the-art methods on three long-term tracking datasets. These findings demonstrate that our method is effective for long-term object tracking and that its performance and speed are promising for use in practical applications of various technologies underlying object tracking.
Takuya Ogawa, Takashi Shibata 0001, Shoji Yachida, Toshinori Hosoi
ICPR4
2003 FAce MOUSe: A novel human-machine interface for controlling the position of a laparoscope
abstract
Robotic laparoscope positioners are now expected as assisting devices for solo surgery among endoscopic surgeons. In such robotic systems, the human-machine (surgeon-robot) interface is of paramount importance because it is the means by which the surgeon communicates with and controls the robotic camera assistant. We have designed a novel human-machine interface, called "FAce MOUSe", for controlling the position of a laparoscope. The proposed human interface is an image-based system which tracks the surgeon's facial motions robustly in real time and does not require the use of any body-contact devices, such as head-mounted sensing devices. The surgeon can easily and precisely control the motion of the laparoscope by simply making the appropriate face gesture, without hand or foot switches or voice input. Based on the FAce MOUSe interface, we have developed a new robotic laparoscope positioning system for solo surgery. Our system allows nonintrusive, nonverbal, hands off and feet off laparoscope operations, which seem more convenient for the surgeon. To evaluate the performance of the proposed system and its applicability in clinical use, we set up an in vivo experiment, in which the surgeon used the system to perform a laparoscopic cholecystectomy on a pig.
Atsushi Nishikawa, Toshinori Hosoi, Kengo Koara, Daiji Negoro, Ayae Hikita, Shuichi Asano, Haruhiko Kakutani, Fumio Miyazaki, Mitsugu Sekimoto, Masayoshi Yasui, Yasuhiro Miyake, Shuji Takiguchi, Morito Monden
IEEE Trans. Robotics Autom.2
2001 Real-Time Visual Tracking of the Surgeon's Face for Laparoscopic Surgery
Atsushi Nishikawa, Toshinori Hosoi, Kengo Koara, Daiji Negoro, Ayae Hikita, Shuichi Asano, Fumio Miyazaki, Mitsugu Sekimoto, Yasuhiro Miyake, Masayoshi Yasui, Morito Monden
MICCAI2