EDBT 2026 Demo / reviewers in the wild / expert
Koya Sakamoto
dblp:378/1384
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Robot navigation and mapping · 38% 3D vision · 38% Vision and language · 19% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › implicit neural representation
3d language field |
0.9 | 1 | 2025 | GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields · ICCV 2025 |
Computer vision › 3D vision
3d scene understanding |
0.9 | 1 | 2025 | GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields · ICCV 2025 |
Robotics › Robot navigation and mapping › mobile robot navigation › 3d navigation
aerial robot navigation |
0.9 | 1 | 2025 | CityNav: A Large-Scale Dataset for Real-World Aerial Navigation · ICCV 2025 |
Computer vision › Vision and language › visual reasoning
compositional visual reasoning |
0.9 | 1 | 2025 | GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields · ICCV 2025 |
Robotics › Robot navigation and mapping
visual navigation |
0.9 | 1 | 2025 | CityNav: A Large-Scale Dataset for Real-World Aerial Navigation · ICCV 2025 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
visual programming · 0.9large language model reasoning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CityNav: A Large-Scale Dataset for Real-World Aerial Navigation
Jungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Yutaka Matsuo, Nakamasa Inoue |
ICCV | 4 |
| 2025 | GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language FieldsabstractThe advancement of 3D language fields has enabled intuitive interactions with 3D scenes via natural language. However, existing approaches are typically limited to small-scale environments, lacking the scalability and compositional reasoning capabilities necessary for large, complex urban settings. To overcome these limitations, we propose GeoProg3D, a visual programming framework that enables natural language-driven interactions with city-scale high-fidelity 3D scenes. GeoProg3D consists of two key components: (i) a Geography-aware City-scale 3D Language Field (GCLF) that leverages a memory-efficient hierarchical 3D model to handle large-scale data, integrated with geographic information for efficiently filtering vast urban spaces using directional cues, distance measurements, elevation data, and landmark references; and (ii) Geographical Vision APIs (GV-APIs), specialized geographic vision tools such as area segmentation and object detection. Our framework employs large language models (LLMs) as reasoning engines to dynamically combine GV-APIs and operate GCLF, effectively supporting diverse geographic vision tasks. To assess performance in city-scale reasoning, we introduce GeoEval3D, a comprehensive benchmark dataset containing 952 query-answer pairs across five challenging tasks: grounding, spatial reasoning, comparison, counting, and measurement. Experiments demonstrate that GeoProg3D significantly outperforms existing 3D language fields and vision-language models across multiple tasks. To our knowledge, GeoProg3D is the first framework enabling compositional geographic reasoning in high-fidelity city-scale 3D environments via natural language. The code is available at https://snskysk.github.io/GeoProg3D/. Shunsuke Yasuki, Taiki Miyanishi, Nakamasa Inoue, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Masato Taki, Yutaka Matsuo |
ICCV | 5 |
| 2024 | Answerability Fields: Answerable Location Estimation via Diffusion ModelsabstractWe propose Answerability Fields (AnsFields), a novel approach for predicting the answerability of questions at different locations within indoor environments. AnsFields is represented as a map, where each grid’s score reflects how well a question can be answered using the panoramic image at that location. Using a 3D question-answering dataset, we construct comprehensive AnsFields covering diverse scenes from ScanNet. Additionally, we employ a diffusion model to infer AnsFields from a scene’s top-down view image and the question. We then conduct 3D question-answering using these predicted AnsFields and achieve a 24% improvement in accuracy over the standard 3D-QA method. Our results demonstrate the importance of object locations for answering questions in the environment, highlighting the potential of AnsFields for applications in robotics, augmented reality, and human-robot interaction. Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto, Motoaki Kawanabe |
IROS | 4 |
| 2024 | Map-based Modular Approach for Zero-shot Embodied Question AnsweringabstractEmbodied Question Answering (EQA) serves as a benchmark task to evaluate the capability of robots to navigate within novel environments and identify objects in response to human queries. However, existing EQA methods often rely on simulated environments and operate with limited vocabularies. This paper presents a map-based modular approach to EQA, enabling real-world robots to explore and map unknown environments. By leveraging foundation models, our method facilitates answering a diverse range of questions using natural language. We conducted extensive experiments in both virtual and real-world settings, demonstrating the robustness of our approach in navigating and comprehending queries within unknown environments. Koya Sakamoto, Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, Motoaki Kawanabe |
IROS | 1 |