VLDB 2026 Research / reviewers in the wild / expert
Satoshi Yamazaki
dblp:95/6974
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Object-Centric Framework for Video Moment RetrievalabstractMost existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such representations often fail to capture fine-grained object semantics and appearance, which are crucial for localizing moments described by object-oriented queries involving specific entities and their interactions. In particular, temporal dynamics at the object level have been largely overlooked, limiting the effectiveness of existing approaches in scenarios requiring detailed object-level reasoning. To address this limitation, we propose a novel object-centric framework for moment retrieval. Our method first extracts query-relevant objects using a scene graph parser and then generates scene graphs from video frames to represent these objects and their relationships. Based on the scene graphs, we construct object-level feature sequences that encode rich visual and semantic information. These sequences are processed by a relational tracklet transformer, which models spatio-temporal correlations among objects over time. By explicitly capturing object-level state changes, our framework enables more accurate localization of moments aligned with object-oriented queries. We evaluated our method on three benchmarks: Charades-STA, QVHighlights, and TACoS. Experimental results demonstrate that our method outperforms existing state-of-the-art methods across all benchmarks. Yongkang Wong, Satoshi Yamazaki, Jianquan Liu, Mohan Kankanhalli |
AAAI | 3 |
| 2026 | KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video UnderstandingabstractWe propose KFS-Bench, the first benchmark for key frame sampling in long video question answering (QA), featuring multi-scene annotations to enable direct and robust evaluation of sampling strategies. Key frame sampling is crucial for efficient long-form video understanding. In long video QA, selecting informative frames enables multimodal large language models (MLLMs) to improve both accuracy and efficiency. KFS-Bench addresses the limitation of prior works that only indirectly assess frame selection quality via QA accuracy. By providing ground-truth annotations of multiple disjoint scenes required per question, KFS-Bench allows us to directly analyze how different sampling approaches capture essential content across an entire long video. Using KFS-Bench, we conduct a comprehensive study of key frame sampling methods and identify that not only sampling precision but also scene coverage and sampling balance are the key factors influencing QA performance. Regarding all the factors, we design a novel sampling quality metric that correlates with QA accuracy. Furthermore, we develop a novel key frame sampling method that leverages question–video relevance to balance sampling diversity against question–frame similarity, thereby improving coverage of relevant scenes. Our adaptively balanced sampling approach achieves superior performance in both key frame sampling and QA performance. The benchmark is available at https://github.com/NEC-VID/KFS-Bench. Kengo Ishida, Satoshi Yamazaki, Xiaotong Ji, Jianquan Liu |
WACV | 3 |
| 2026 | Analytical Modeling of Energy Consumption for a LoRaWAN Class B Node
Satoshi Yamazaki, Hidenobu Isogai |
WCNC | 1 |
| 2025 | Dual-Stream Spatio-Temporal Accident Anticipation and DetectionabstractThis paper proposes a novel dual-stream spatio-temporal modeling method for spatio-temporal accident anticipation and detection in dashcam videos, a task that requires not only accident anticipation/detection but also localization of the accident participants. Previous methods have been developed based on temporal modeling with global visual embeddings and object visual embeddings, resulting in an insufficient ability to capture spatio-temporal correlations. Moreover, the object embeddings do not comprehensively incorporate visual, motion, and semantic information of the objects, which is all essential to the traffic understanding. To this end, our method uses a dual-stream framework to model the spatio-temporal correlations from both region and object perspectives. The region stream focuses on local visual contents, while the object stream traces the visual changes and motions of objects using a comprehensive object representation. In contrast to previous works, our method has been evaluated for both accident anticipation and detection on four benchmarks. The experimental results show that our method significantly outperforms previous methods on all benchmarks with a lead up to 10.1 in AUC and 6.1 in AP. Satoshi Yamazaki, Jianquan Liu |
ICIP | 2 |
| 2025 | Conversation Simulation and Behavioral Change Design for Evacuation Decision-MakingabstractThis study introduces multi-agent conversation models that utilize generative AI to overcome the challenges posed by traditional rule-based multi-agent simulations. It proposes an effective method for issuing appropriate evacuation instructions method through simulations considering psychological factors and information transmission during disasters. Using an earthquake as the disaster scenario, the correlation between agent conversations and evacuation behaviors was evaluated, and the effectiveness of various evacuation instructions were compared. Additionally, the evacuation instruction methods were analyzed, and the improvements in prediction accuracy were verified. This research enhances the understanding of evacuation behavior during disasters and contributes to the development of realistic disaster response simulations, offering insights for improving actual evacuation instruction methods. Fumio Tanabe, Kosuke Hayashi, Satoshi Yamazaki, Rie Sato, Hideki Murakoshi, Hisashi Hayashi |
KES | 3 |
| 2025 | Open-Vocabulary Scene Graph Generation via Synonym-Based Predicate Descriptor
Yuta Goto, Satoshi Yamazaki, Takashi Shibata 0001, Jianquan Liu |
MMM (3) | 2 |
| 2024 | A Parallel Transformer Framework for Video Moment RetrievalabstractIn the realm of video understanding, Video Moment Retrieval (VMR) is an important yet challenging task that aims to locate the boundary of a moment of interest within a long untrimmed video. Existing VMR methods often focus on the visual content extracted from the video only (or frame sequences), however, the rich semantic information at the object level that describes the image's content has not been explored yet. To overcome those limitations, we propose PaTF, an attention-based Parallel Transformer Framework that enriches the feature representations by exploring both low-level visual cues and high-level relational contexts of video-query pairs. Our framework consists of two parallel transformers: one for the visual-textual stream and the other for the semantic-textual stream. The visual-textual stream extracts the links between global visual features and textual information, while the semantic-textual stream emphasises the relations between objects via scene graph representations. Furthermore, our comprehensive experiment conducted on the Charades-STA dataset demonstrates that the proposed framework outperforms the state-of-the-art methods by a large margin, roughly 5% and 7% at Recall@1 with IoU = 0.5 and IoU = 0.7, respectively. Thao-Nhu Nguyen, Zongyao Li 0004, Satoshi Yamazaki, Jianquan Liu, Cathal Gurrin |
ICMR | 3 |
| 2023 | Client Selection Based on Channel Capacity for Federated Learning Under Wireless ChannelsabstractThis paper proposes a user selection scheme for federated learning (FL) over wireless networks to reduce communication time based on channel capacity. In particular, the edge server calculates the Shannon channel capacity of each client for each round, and clients with a certain threshold capacity are randomly selected to participate in FL. We show that the convergence time of the proposed scheme outperformed that of the conventional scheme through computer simulation based on an image processing task under a wireless channel with pass-loss, shadowing, and Rician flat-fading. Moreover, the superiority of FL to centralized learning (CL) regarding total time is demonstrated theoretically and validated through computer simulation. Satoshi Yamazaki, Takuma Furuki |
APCC | 1 |
| 2023 | NarSUM '23: The 2nd Workshop on User-Centric Narrative Summarization of Long VideosabstractWith video capture devices becoming widely popular, the amount of video data generated per day has seen a rapid increase over the past few years. Browsing through hours of video data to retrieve useful information is a tedious and boring task. Video Summarization technology has played a crucial role in addressing this issue. It is a well-researched topic in the multimedia community. However, the focus so far has been limited to creating summary to videos which are short (only a few minutes). This workshop aims to call for researchers on relevant background to focus on novel solutions for user-centric narrative summarization of long videos. This workshop will also cover important aspects of video summarization research like what is "important" in a video, how to evaluate the goodness of a created summary, open challenges in video summarization, etc. Mohan Kankanhalli, Ioannis Patras, Jianquan Liu, Yongkang Wong, Takahiro Komamizu, Satoshi Yamazaki, Karen Stephen, Kajal Kansal |
ACM Multimedia | 6 |
| 2023 | Energy Consumption Model for LoRaWAN via Field Experiments for Smart AgricultureabstractThis paper proposes a scheme to model the energy consumption of LoRaWAN, which is a popular example of low-power wide-area networks (LPWANs), nodes via the results of outdoor field experiments by assuming regional smart agriculture as a use case for internet of things (IoT). Specifically, we derive an experimental approximation formula to estimate the battery lifetime by introducing parameters such as spread factor and payload length. The validity of the proposed scheme is demonstrated by confirming that the results obtained by the approximate formula, the experimental and theoretical results generally agree, regardless of the node state and spread factor. Furthermore, we show that the obtained approximate formula can be used to identify the current consumption value in the sleep state that should be achieved to achieve the desired battery lifetime. Satoshi Yamazaki, Taiki Ogura |
PIMRC | 1 |
| 2020 | A Simple and Energy-Efficient Flooding Scheme for Wireless RoutingabstractIn many wireless networks, such as ad hoc, sensor, and delay-tolerant networks, the destination node is determined from the source node by the flooding process. Flooding efficiency is important for nodes, because they are driven by limited batteries. In this paper, we propose a simple flooding scheme to transmit a route request (RREQ) message based on the remaining power of its own node without using control packets and complex calculations. We applied the proposed scheme to ad hoc on-demand distance vector (AODV) routing protocol as an example and carried out computer simulations (ns3). The results showed that the proposed scheme was superior to conventional schemes in static and mobile scenarios. First, we showed the limit of node density that causes the decrease of throughput in the proposed scheme and that the proposed scheme was superior in terms of energy efficiency (bits/J), including throughput and energy consumption. Next, as the number of flooding times is made uniform in the proposed scheme, all nodes will have almost the same battery replacement time. As a result, when the nodes are static, the lifetime in the proposed scheme is longer than that in the conventional scheme. Satoshi Yamazaki, Yu Abiko, Hideki Mizuno |
Wirel. Commun. Mob. Comput. | 1 |
| 2015 | A fast method of verifying network routing with back-trace header space analysisabstractIt is a tough job for operators to make perfectly accurate configuration of many network elements in large networks. Erroneous configurations may cause critical incidents in network, on which many ICT systems are running. It may also result in a security hole as well as system incidents. There has been much work on preventing erroneous configurations, but this has taken a lot of time to verify routing with large networks. We propose a new method of verifying network routing. It only focuses on verifying isolation and reachability, but it can verify these properties with O(R2), where R is the number of flow entries, while the performance of an existing method of verification is O(R3). We also provide a proof of the correctness of our method. Toshio Tonouchi, Satoshi Yamazaki, Yutaka Yakuwa, Nobuyuki Tomizawa |
IM | 2 |
| 2002 | A Development of Classroom Design Simulator for Interactive Video teleconferenceabstract"Visual Field Record System" is aimed to support participant's behavior analysis in distance education using an interactive video teleconference system. An examination of the eye camera position is important to use the system effectively, because the camera's angle and zoom rate is considered participant's behavior in this system. Actually, there are some problems that it originates in the camera position of the system. We approached these problems by developing the eye camera install-simulator including arrangement of the classroom equipment. This simulator is called Classroom Design Simulator. This paper describes functions and operations of the simulator. After that, the system is evaluated by the application to a case. It was shown that this system was useful for a distance education planner as a result. Satoshi Yamazaki, Naoto Nakamura, Youzou Miyadera, Setsuo Yokoyama |
ICCE | 1 |
| 2000 | Eye Mark Pointer in Immersive Projection DisplayabstractWe developed a pointer in 3D virtual space, using an eye-tracking system as a sensor. The eye mark pointer is installed to a virtual environment system which provides stereoscopic vision with an immersive projection display. The circular-polarization stereoscopic vision enables us to use the eye-tracking system in the immersive projection display. The eye-tracking system obtains relative gaze directions with respect to the head, so the absolute position requires compensation of the user's head motion with a head tracker. We then compare the eye mark pointer with a joystick in an experiment with the virtual environment system. The experimental result indicates the pointing of the eye mark pointer is 9.8 times quicker than that of the joystick, and suggests that the eye mark pointer is available for pointing at the target in the virtual environment. Kikuo Asai, Noritaka Osawa, Hideaki Takahashi, Yuji Y. Sugimoto, Satoshi Yamazaki, Masahiro Samejima, Taiki Tanimae |
VR | 5 |