VLDB 2026 Research / reviewers in the wild / expert
Atsunori Moteki
dblp:92/10299
· DBLP profile ↗
8ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0003-3160-5275ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
Atsunori Moteki, Akiyoshi Uchida, Shoichi Masui, Fan Yang 0032, Kanji Uchino, Yueqi Song, Yonatan Bisk, Graham Neubig, Ikuo Kusajima, Yasuto Watanabe, Hiroyuki Ishida, Koki Nakagawa, Shan Jiang 0006 |
ICPR (4) | 2 |
| 2026 | Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human ActivitiesabstractPeriodic human activities with implicit workflows are common in manufacturing, sports, and daily life. While short-term periodic activities—characterized by simple structures and high-contrast patterns—have been widely studied, long-term periodic workflows with low-contrast patterns remain largely underexplored. To bridge this gap, we introduce the first benchmark comprising 580 multimodal human activity sequences featuring long-term periodic workflows. The benchmark supports three evaluation tasks aligned with real-world applications: unsupervised periodic workflow detection, task completion tracking, and procedural anomaly detection. We also propose a lightweight, training-free baseline for modeling diverse periodic workflow patterns. Experiments show that: (i) our benchmark presents significant challenges to both unsupervised periodic detection methods and zero-shot approaches based on powerful large language models (LLMs); (ii) our baseline outperforms competing methods by a substantial margin in all evaluation tasks; and (iii) in real-world applications, our baseline demonstrates deployment advantages on par with traditional supervised workflow detection approaches, eliminating the need for annotation and retraining. Our project page is https://sites.google.com/view/periodicworkflow. Fan Yang 0032, Quanting Xie, Atsunori Moteki, Shoichi Masui, Shan Jiang 0006, Kanji Uchino, Yonatan Bisk, Graham Neubig |
WACV | 3 |
| 2025 | YOWO: You Only Walk Once to Jointly Map an Indoor Scene and Register Ceiling-Mounted CamerasabstractUsing ceiling-mounted cameras (CMCs) for indoor visual capturing opens up a wide range of applications. However, registering CMCs to the target scene layout presents a challenging task. While manual registration with specialized tools is inefficient and costly, automatic registration with visual localization may yield poor results when visual ambiguity exists. To alleviate these issues, we propose a novel solution for jointly mapping an indoor scene and registering CMCs to the scene layout. Our approach involves equipping a mobile agent with a head-mounted RGB-D camera to traverse the entire scene once and synchronize CMCs to capture this mobile agent. The egocentric videos generate world-coordinate agent trajectories and the scene layout, while the videos of CMCs provide pseudo-scale agent trajectories and CMC relative poses. By correlating all the trajectories with their corresponding timestamps, the CMC relative poses can be aligned to the world-coordinate scene layout. Based on this initialization, a factor graph is customized to enable the joint optimization of ego-camera poses, scene layout, and CMC poses. We also develop a new dataset, setting the first benchmark for collaborative scene mapping and CMC registration. Experimental results indicate that our method not only effectively accomplishes two tasks within a unified framework, but also jointly enhances their performance. We thus provide a reliable tool to facilitate downstream position-aware applications. Fan Yang 0032, Sosuke Yamao, Ikuo Kusajima, Atsunori Moteki, Shoichi Masui, Shan Jiang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | MSD-CRFS: Multi-Scale Dual Aggregation Conditional Random Fields for Monocular Depth EstimationabstractWe address the problem of estimating a high-quality dense depth map from a single RGB input image. We first analyze Conditional Random Field (CRF) in combination with transformers and exploit the multi-head attention mechanism to compute a potential function. Then, we propose spatial window CRFs and channel-wise CRFs to observe information in spatial and channel dimensions, and fuse them with a two-way fusion module, which is called Dual aggregation CRFs (DCRFs). Finally, the information from the multi-scale features observed by DCRFs is used for internal scene clustering by slot-attention to obtain the depth map. We call our method as MSD-CRFs. Experiments demonstrate that our method improves the performance across all metrics on the KITTI, and outperforms current SOTA results on the main ranking metrics $A b s \_$Rel on NYU Depth-v2. Further, we explore the model generalization capability via zero-shot test. Xidan Zhang, Jianing Wei, Atsunori Moteki, Yoshie Kobayashi, Genta Suzuki, Zhiming Tan |
ICIP | 3 |
| 2024 | MSCC: Multi-Scale Transformers for Camera CalibrationabstractCamera calibration is very important for some vision tasks, like rendering 3D scenes, environment reconstruction, and self-localization, etc. In this paper, we propose a framework of multi-scale transformers for camera calibration. With the input of a single image, the multi-scale features output from the model’s backbone are utilized to estimate camera parameters. At the same time, we show that the way of coarse-to-fine is effective to locate global structures and detailed features in the image, by studying the attention response of horizon line estimation. Moreover, deep supervision is proven to get more precise results and accelerated training. Our method outperforms all the state-of-the-art methods by objective and subjective experiments on Google Street View dataset and Pano360. Xu Song, Hao Kang, Atsunori Moteki, Genta Suzuki, Yoshie Kobayashi, Zhiming Tan |
WACV | 3 |
| 2020 | Guideline and Tool for Designing an Assembly Task Support System Using Augmented RealityabstractAugmented reality (AR) systems support complex tasks like assembly by overlaying task-related content onto the real world. In recent years, the effort of designing and developing assembly task support systems in AR decreased with the availability of high potential head-mounted displays and provision of integrated development environments. Nevertheless, problems still arise when companies craft an effective AR task support system, particularly in the difficulty of selecting appropriate techniques and information-presentation methods, and the requirements that vary with each use case. In this study, we formulated a corresponding guideline, developed a selection aid tool that incorporates filtering based on the categorization of subtasks and the degree of freedom of available tracking, and evaluated their effectiveness in two experiments. First, to confirm effects on system design, we asked 18 participants to perform the design action of the AR system with the guideline for two tasks (PC assembly and rope work). Consequently, to verify the quality of the designed AR systems from Experiment 1, we asked another set of 20 participants to perform the same tasks with those systems. The results confirm that using the guideline can considerably lower efforts creating media and alleviate the error for a specific process. We envision our guideline and tool to be accessible as an online web page, assisting AR assembly task support system designer/developers worldwide. Keishi Tainaka, Yuichiro Fujimoto, Masayuki Kanbara, Hirokazu Kato 0001, Atsunori Moteki, Kensuke Kuraki, Kazuki Osamura, Toshiyuki Yoshitake, Toshiyuki Fukuoka |
ISMAR | 5 |
| 2018 | Handheld Guides in Inspection Tasks: Augmented Reality versus PictureabstractInspection tasks focus on observation of the environment and are required in many industrial domains. Inspectors usually execute these tasks by using a guide such as a paper manual, and directly observing the environment. The effort required to match the information in a guide with the information in an environment and the constant gaze shifts required between the two can severely lower the work efficiency of inspector in performing his/her tasks. Augmented reality (AR) allows the information in a guide to be overlaid directly on an environment. This can decrease the amount of effort required for information matching, thus increasing work efficiency. AR guides on head-mounted displays (HMDs) have been shown to increase efficiency. Handheld AR (HAR) is not as efficient as HMD-AR in terms of manipulability, but is more practical and features better information input and sharing capabilities. In this study, we compared two handheld guides: an AR interface that shows 3D registered annotations, that is, annotations having a fixed 3D position in the AR environment, and a non-AR picture interface that displays non-registered annotations on static images. We focused on inspection tasks that involve high information density and require the user to move, as well as to perform several viewpoint alignments. The results of our comparative evaluation showed that use of the AR interface resulted in lower task completion times, fewer errors, fewer gaze shifts, and a lower subjective workload. We are the first to present findings of a comparative study of an HAR and a picture interface when used in tasks that require the user to move and execute viewpoint alignments, focusing only on direct observation. Our findings can be useful for AR practitioners and psychology researchers. Jarkko Polvi, Takafumi Taketomi, Atsunori Moteki, Toshiyuki Yoshitake, Toshiyuki Fukuoka, Goshiro Yamamoto, Christian Sandor, Hirokazu Kato 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2016 | Fast and accurate relocalization for keyframe-based SLAM using geometric model selectionabstractIn this paper, we propose a relocalization method for keyframe-based SLAM that enables real-time and accurate recovery from tracking failures. To realize an AR-based application in a real world situation, not only accurate camera tracking but also fast and accurate relocalization from tracking failure is required. The previous keyframe-based relocalization methods have some drawbacks with regard to speed and accuracy. The proposed relocalization method selects two algorithms adaptively depending on the relative camera pose between a current frame and a target keyframe. In addition, it estimates a degree of false matches to speed up RANSAC-based model estimation. We present effectiveness of our method by an evaluation using public tracking dataset. Atsunori Moteki, Nobuyasu Yamaguchi, Ayu Karasudani, Toshiyuki Yoshitake |
VR | 1 |