Makoto Terao

dblp:69/3180 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
8since 2021 · last 2025
0009-0001-9982-2463ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Action-Agnostic Point-Level Supervision for Temporal Action Detection
abstract
We propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the proposed scheme, a small portion of video frames is sampled in an unsupervised manner and presented to human annotators, who then label the frames with action categories. Unlike point-level supervision, which requires annotators to search for every action instance in an untrimmed video, frames to annotate are selected without human intervention in AAPL supervision. We also propose a detection model and learning method to effectively utilize the AAPL labels. Extensive experiments on the variety of datasets (THUMOS'14, FineAction, GTEA, BEOID, and ActivityNet 1.3) demonstrate that the proposed approach is competitive with or outperforms prior methods for video-level and point-level supervision in terms of the trade-off between the annotation cost and detection performance.
Shuhei M. Yoshida, Takashi Shibata 0001, Makoto Terao, Takayuki Okatani, Masashi Sugiyama
AAAI3
2025 MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval
abstract
Result diversification (RD) is a crucial technique in Text-to-Image Retrieval for enhancing the efficiency of a practical application. Conventional methods focus solely on increasing the diversity metric of image appearances. However, the diversity metric and its desired value vary depending on the application, which limits the applications of RD. This paper proposes a novel task called CDR-CA (Contextual Diversity Refinement of Composite Attributes). CDR-CA aims to refine the diversities of multiple attributes, according to the application's context. To address this task, we propose Multi-Source DPPs, a simple yet strong baseline that extends the Determinantal Point Process (DPP) to multi-sources. We model MS-DPP as a single DPP model with a unified similarity matrix based on a manifold representation. We also introduce Tangent Normalization to reflect contexts. Extensive experiments demonstrate the effectiveness of the proposed method.
Naoya Sogi, Takashi Shibata 0001, Makoto Terao, Masanori Suganuma, Takayuki Okatani
IJCAI3
2024 Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
Naoya Sogi, Takashi Shibata 0001, Makoto Terao
ECCV (79)3
2024 Task Success Classification with Final State of Future Prediction for Robot Control Planning
Taku Fujitomi, Naoya Sogi, Takashi Shibata 0001, Makoto Terao
ICPR (2)4
2024 Disaster Damage Visualization by VLM-Based Interactive Image Retrieval and Cross-View Image Geo-Localization
abstract
We propose a framework for quickly selecting images that show the disaster situation from many images, estimating their locations with high accuracy, and displaying them on a map. The proposed framework introduces interactive image retrieval based on the Vision and Language Model (VLM), which can retrieve images from many images that show the disaster situation according to the user’s intention. Using the correlation between language and images based on VLM and the similarity between images selected interactively enables more accurate retrieval. Next, for selected images for which the location of the affected area is unknown, the location of the image is estimated with street address-level accuracy by matching it with an overhead image covering a large area of the city and map data and then displayed on a map. We confirmed the effectiveness of the proposed method on publicly available datasets such as CrisisNLP.
Naoya Sogi, Takashi Shibata 0001, Makoto Terao, Kenta Senzaki, Masahiro Tani, Royston Rodrigues
IGARSS3
2024 Future Predictive Success-or-Failure Classification for Long-Horizon Robotic Tasks
abstract
Automating long-horizon tasks with a robotic arm has been a central research topic in robotics. Optimization-based action planning is an efficient approach for creating an action plan to complete a given task. Construction of a reliable planning method requires a design process of conditions, e.g., to avoid collision between objects. The design process, however, has two critical issues: 1) iterative trials–the design process is time-consuming due to the trial-and-error process of modifying conditions, and 2) manual redesign–it is difficult to cover all the necessary conditions manually. To tackle these issues, this paper proposes a future-predictive success-or-failure-classification method to obtain conditions automatically. The key idea behind the proposed method is an end-to-end approach for determining whether the action plan can complete a given task instead of manually redesigning the conditions. The proposed method uses a long-horizon future-prediction method to enable success-or-failure classification without the execution of an action plan. This paper also proposes a regularization term called transition consistency regularization to provide easy-to-predict feature distribution. The regularization term improves future prediction and classification performance. The effectiveness of our method is demonstrated through classification and robotic-manipulation experiments.
Naoya Sogi, Hiroyuki Oyama, Takashi Shibata 0001, Makoto Terao
IJCNN4
2024 Appearance-Based Curriculum for Semi-Supervised Learning with Multi-Angle Unlabeled Data
abstract
We propose an appearance-based curriculum (ABC) for a semi-supervised learning scenario where labeled images taken from limited angles and unlabeled ones taken from various angles are available for training. A common approach to semi-supervised learning relies on pseudo-labeling and data augmentation, but it struggles with large visual variations that cannot be covered by data augmentation. To solve this problem, ABC incrementally expands the pool of unlabeled images fed to a base semi-supervised learner so that newly added data are the ones most similar to those already in the pool. This way, the learner can assign pseudo-labels to the new data with high accuracy, keeping the quality of pseudo-labels higher than that when all the unlabeled data are processed at once, as customarily done in existing semi-supervised learning methods. We conducted extensive experiments and confirmed that our method outperforms the state-of-the-art semi-supervised learning methods in our scenario.
Shuhei M. Yoshida, Takashi Shibata 0001, Makoto Terao, Takayuki Okatani, Masashi Sugiyama
WACV4
2022 Non-Iterative Optimization of Pseudo-Labeling Thresholds for Training Object Detection Models from Multiple Datasets
abstract
We propose a non-iterative method to optimize pseudo-labeling thresholds for learning object detection from a collection of low-cost datasets, each of which is annotated for only a subset of all the object classes. A popular approach to this problem is first to train teacher models and then to use their confident predictions as pseudo ground-truth labels when training a student model. To obtain the best result, however, thresholds for prediction confidence must be adjusted. This process typically involves iterative search and repeated training of student models and is time-consuming. Therefore, we develop a method to optimize the thresholds without iterative optimization by maximizing the Fβ-score on a validation dataset, which measures the quality of pseudo labels and can be measured without training a student model. We experimentally demonstrate that our proposed method achieves an mAP comparable to that of grid search on the COCO and VOC datasets.
Shuhei M. Yoshida, Makoto Terao
ICIP3
2008 Open-vocabulary spoken-document retrieval based on query expansion using related web documents
Makoto Terao, Takafumi Koshinaka, Shinichi Ando, Ryosuke Isotani, Akitoshi Okumura
INTERSPEECH1
2002 Statistical language modeling with prosodic boundaries and its use for continuous speech recognition
Keikichi Hirose, Nobuaki Minematsu, Makoto Terao
INTERSPEECH3
2000 Controlling Effective Introns for Multi-Agent Learning by Genetic Programming
Hitoshi Iba, Makoto Terao
GECCO2