EDBT 2026 Demo / reviewers in the wild / expert
Taiki Miyanishi
dblp:45/8008
· DBLP profile ↗
21ranked-venue papers
10as first author
9since 2021 · last 2025
0000-0001-9105-1601ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-authorSystems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Vision and language · 37% 3D vision · 31% Robot navigation and mapping · 18% | |
| Human-computer interaction and pervasive computing
3 papers |
Ubiquitous computing and smart environments · 57% Interaction techniques and input · 43% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 70% Knowledge graphs · 30% |
Topics — the 19 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d scene understanding |
1.4 | 2 | 2025 | GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields · ICCV 2025 ScanQA: 3D Question Answering for Spatial Scene Understanding · CVPR 2022 |
Computer vision › 3D vision › implicit neural representation
3d language field |
0.9 | 1 | 2025 | GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields · ICCV 2025 |
Robotics › Robot navigation and mapping › mobile robot navigation › 3d navigation
aerial robot navigation |
0.9 | 1 | 2025 | CityNav: A Large-Scale Dataset for Real-World Aerial Navigation · ICCV 2025 |
Computer vision › Vision and language › visual reasoning
compositional visual reasoning |
0.9 | 1 | 2025 | GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields · ICCV 2025 |
Robotics › Robot navigation and mapping
visual navigation |
0.9 | 1 | 2025 | CityNav: A Large-Scale Dataset for Real-World Aerial Navigation · ICCV 2025 |
Computer vision › 3D vision › 3d scene understanding
3d visual grounding |
0.7 | 1 | 2023 | CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data · NeurIPS 2023 |
Computer vision › Vision and language › 3d vision and language
language-guided 3d object localization |
0.7 | 1 | 2023 | CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data · NeurIPS 2023 |
Computer vision › Vision and language
visual grounding |
0.7 | 1 | 2023 | CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data · NeurIPS 2023 |
Computer vision › Vision and language › 3d vision and language
3d question answering |
0.6 | 1 | 2022 | ScanQA: 3D Question Answering for Spatial Scene Understanding · CVPR 2022 |
Computer vision › Vision and language
visual question answering |
0.6 | 1 | 2022 | ScanQA: 3D Question Answering for Spatial Scene Understanding · CVPR 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.3 | 1 | 2018 | Generating an Event Timeline About Daily Activities From a Semantic Concept Stream · AAAI 2018 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields · ICCV 2025 |
Information retrieval › multimedia analysis and retrieval
video retrieval |
0.2 | 1 | 2016 | Egocentric Video Search via Physical Interactions · AAAI 2016 |
Computer vision › Vision and language › visual grounding
object grounding |
0.2 | 1 | 2022 | ScanQA: 3D Question Answering for Spatial Scene Understanding · CVPR 2022 |
Information retrieval
query suggestion |
0.2 | 1 | 2013 | Time-aware structured query suggestion · SIGIR 2013 |
Ubiquitous computing and smart environments › context recognition
activity recognition |
0.1 | 1 | 2018 | Generating an Event Timeline About Daily Activities From a Semantic Concept Stream · AAAI 2018 |
Ubiquitous computing and smart environments
context-aware computing |
0.1 | 1 | 2016 | Selecting home appliances with smart glass based on contextual information · UbiComp 2016 |
Interaction techniques and input
gesture input |
0.1 | 1 | 2016 | Egocentric Video Search via Physical Interactions · AAAI 2016 |
Information retrieval › ranking › text ranking
document ranking |
0.0 | 1 | 2013 | Time-aware structured query suggestion · SIGIR 2013 |
Methods — techniques the papers use, named apart from their topics
multimodal learning · 1.3language encoding · 1.3visual programming · 0.9large language model reasoning · 0.9temporal reasoning · 0.7commonsense knowledge · 0.7sentence embeddings · 0.6multimodal fusion · 0.6bounding box regression · 0.63d object proposals · 0.6probabilistic modeling · 0.5canonical correlation analysis · 0.5non-parametric bayesian · 0.2multiple kernel learning · 0.2deep learning · 0.2query-URL bipartite graph · 0.2clustering · 0.2click counts · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CityNav: A Large-Scale Dataset for Real-World Aerial Navigation
Jungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Yutaka Matsuo, Nakamasa Inoue |
ICCV | 2 |
| 2025 | GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language FieldsabstractThe advancement of 3D language fields has enabled intuitive interactions with 3D scenes via natural language. However, existing approaches are typically limited to small-scale environments, lacking the scalability and compositional reasoning capabilities necessary for large, complex urban settings. To overcome these limitations, we propose GeoProg3D, a visual programming framework that enables natural language-driven interactions with city-scale high-fidelity 3D scenes. GeoProg3D consists of two key components: (i) a Geography-aware City-scale 3D Language Field (GCLF) that leverages a memory-efficient hierarchical 3D model to handle large-scale data, integrated with geographic information for efficiently filtering vast urban spaces using directional cues, distance measurements, elevation data, and landmark references; and (ii) Geographical Vision APIs (GV-APIs), specialized geographic vision tools such as area segmentation and object detection. Our framework employs large language models (LLMs) as reasoning engines to dynamically combine GV-APIs and operate GCLF, effectively supporting diverse geographic vision tasks. To assess performance in city-scale reasoning, we introduce GeoEval3D, a comprehensive benchmark dataset containing 952 query-answer pairs across five challenging tasks: grounding, spatial reasoning, comparison, counting, and measurement. Experiments demonstrate that GeoProg3D significantly outperforms existing 3D language fields and vision-language models across multiple tasks. To our knowledge, GeoProg3D is the first framework enabling compositional geographic reasoning in high-fidelity city-scale 3D environments via natural language. The code is available at https://snskysk.github.io/GeoProg3D/. Shunsuke Yasuki, Taiki Miyanishi, Nakamasa Inoue, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Masato Taki, Yutaka Matsuo |
ICCV | 2 |
| 2025 | LegalViz: Legal Text Visualization by Text To Diagram GenerationabstractEri Onami, Taiki Miyanishi, Koki Maeda, Shuhei Kurita. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Eri Onami, Taiki Miyanishi, Koki Maeda, Shuhei Kurita |
NAACL (Long Papers) | 2 |
| 2024 | Cross3DVG: Cross-Dataset 3D Visual Grounding on Different RGB-D ScansabstractWe present a novel task for cross-dataset visual grounding in 3D scenes (Cross3DVG), which overcomes limitations of existing 3D visual grounding models, specifically their restricted 3D resources and consequent tendencies of overfitting a specific 3D dataset. We created RIORefer, a large-scale 3D visual grounding dataset, to facilitate Cross3DVG. It includes more than 63 k diverse descriptions of 3D objects within 1,380 indoor RGB-D scans from 3RScan [48], with human annotations. After training the Cross3DVG model using the source 3D visual grounding dataset, we evaluate it without target labels using the target dataset with, e.g., different sensors, 3D reconstruction methods, and language annotators. Comprehensive experiments are conducted using established visual grounding models and with CLIP-based multi-view 2D and 3D integration designed to bridge gaps among 3D datasets. For Cross3DVG tasks, (i) cross-dataset 3D visual grounding exhibits significantly worse performance than learning and evaluation with a single dataset because of the 3D data and language variants across datasets. Moreover, (ii) better object detector and localization modules and fusing 3D data and multi-view CLIP-based image features can alleviate this lower performance. Our Cross3DVG task can provide a benchmark for developing robust 3D visual grounding models to handle diverse 3D scenes while leveraging deep language understanding.11Project page: https://github.com/ATR-DBI/Cross3DVG Taiki Miyanishi, Daichi Azuma, Shuhei Kurita, Motoaki Kawanabe |
3DV | 1 |
| 2024 | JDocQA: Japanese Document Question Answering Dataset for Generative Language ModelsabstractDocument question answering is a task of question answering on given documents such as reports, slides, pamphlets, and websites, and it is a truly demanding task as paper and electronic forms of documents are so common in our society. This is known as a quite challenging task because it requires not only text understanding but also understanding of figures and tables, and hence visual question answering (VQA) methods are often examined in addition to textual approaches. We introduce Japanese Document Question Answering (JDocQA), a large-scale document-based QA dataset, essentially requiring both visual and textual information to answer questions, which comprises 5,504 documents in PDF format and annotated 11,600 question-and-answer instances in Japanese. Each QA instance includes references to the document pages and bounding boxes for the answer clues. We incorporate multiple categories of questions and unanswerable questions from the document for realistic question-answering applications. We empirically evaluate the effectiveness of our dataset with text-based large language models (LLMs) and multimodal models. Incorporating unanswerable questions in finetuning may contribute to harnessing the so-called hallucination generation. Eri Onami, Shuhei Kurita, Taiki Miyanishi, Taro Watanabe |
LREC/COLING | 3 |
| 2024 | Answerability Fields: Answerable Location Estimation via Diffusion ModelsabstractWe propose Answerability Fields (AnsFields), a novel approach for predicting the answerability of questions at different locations within indoor environments. AnsFields is represented as a map, where each grid’s score reflects how well a question can be answered using the panoramic image at that location. Using a 3D question-answering dataset, we construct comprehensive AnsFields covering diverse scenes from ScanNet. Additionally, we employ a diffusion model to infer AnsFields from a scene’s top-down view image and the question. We then conduct 3D question-answering using these predicted AnsFields and achieve a 24% improvement in accuracy over the standard 3D-QA method. Our results demonstrate the importance of object locations for answering questions in the environment, highlighting the potential of AnsFields for applications in robotics, augmented reality, and human-robot interaction. Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto, Motoaki Kawanabe |
IROS | 2 |
| 2024 | Map-based Modular Approach for Zero-shot Embodied Question AnsweringabstractEmbodied Question Answering (EQA) serves as a benchmark task to evaluate the capability of robots to navigate within novel environments and identify objects in response to human queries. However, existing EQA methods often rely on simulated environments and operate with limited vocabularies. This paper presents a map-based modular approach to EQA, enabling real-world robots to explore and map unknown environments. By leveraging foundation models, our method facilitates answering a diverse range of questions using natural language. We conducted extensive experiments in both virtual and real-world settings, demonstrating the robustness of our approach in navigating and comprehending queries within unknown environments. Koya Sakamoto, Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, Motoaki Kawanabe |
IROS | 3 |
| 2023 | CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud DataabstractCity-scale 3D point cloud is a promising way to express detailed and complicated outdoor structures. It encompasses both the appearance and geometry features of segmented city components, including cars, streets, and buildings that can be utilized for attractive applications such as user-interactive navigation of autonomous vehicles and drones. However, compared to the extensive text annotations available for images and indoor scenes, the scarcity of text annotations for outdoor scenes poses a significant challenge for achieving these applications. To tackle this problem, we introduce the CityRefer dataset for city-level visual grounding. The dataset consists of 35k natural language descriptions of 3D objects appearing in SensatUrban city scenes and 5k landmarks labels synchronizing with OpenStreetMap. To ensure the quality and accuracy of the dataset, all descriptions and labels in the CityRefer dataset are manually verified. We also have developed a baseline system that can learn encoded language descriptions, 3D object instances, and geographical information about the city's landmarks to perform visual grounding on the CityRefer dataset. To the best of our knowledge, the CityRefer dataset is the largest city-level visual grounding dataset for localizing specific 3D objects. Taiki Miyanishi, Fumiya Kitamori, Shuhei Kurita, Jungdae Lee, Motoaki Kawanabe, Nakamasa Inoue |
NeurIPS | 1 |
| 2022 | ScanQA: 3D Question Answering for Spatial Scene UnderstandingabstractWe propose a new 3D spatial understanding task for 3D question answering (3D-QA). In the 3D-QA task, models receive visual information from the entire 3D scene of a rich RGB-D indoor scan and answer given textual questions about the 3D scene. Unlike the 2D-question answering of visual question answering, the conventional 2D-QA models suffer from problems with spatial understanding of object alignment and directions and fail in object localization from the textual questions in 3D-QA. We propose a baseline model for 3D-QA, called the ScanQA11https://github.com/ATR-DBI/ScanQA, which learns a fused descriptor from 3D object proposals and encoded sentence embeddings. This learned descriptor correlates language expressions with the underlying geometric features of the 3D scan and facilitates the regression of 3D bounding boxes to determine the described objects in textual questions. We collected human-edited question-answer pairs with free-form answers grounded in 3D objects in each 3D scene. Our new ScanQA dataset contains over 41k question-answer pairs from 800 indoor scenes obtained from the ScanNet dataset. To the best of our knowledge, ScanQA is the first large-scale effort to perform object-grounded question answering in 3D environments. Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, Motoaki Kawanabe |
CVPR | 2 |
| 2020 | Two-Stream Spatiotemporal Compositional Attention Network for VideoQA
Taiki Miyanishi, Takuya Maekawa, Motoaki Kawanabe |
BMVC | 1 |
| 2018 | Generating an Event Timeline About Daily Activities From a Semantic Concept StreamabstractRecognizing activities of daily living (ADLs) in the real world is an important task for understanding everyday human life. However, even though our life events consist of chronological ADLs with the corresponding places and objects (e.g., drinking coffee in the living room after making coffee in the kitchen and walking to the living room), most existing works focus on predicting individual activity labels from sensor data. In this paper, we introduce a novel framework that produces an event timeline of ADLs in a home environment. The proposed method combines semantic concepts such as action, object, and place detected by sensors for generating stereotypical event sequences with the following three real-world properties. First, we use temporal interactions among concepts to remove objects and places unrelated to each action. Second, we use commonsense knowledge mined from a language resource to find a possible combination of concepts in the real world. Third, we use temporal variations of events to filter repetitive events, since our daily life changes over time. We use cross-place validation to evaluate our proposed method on a daily-activities dataset with manually labeled event descriptions. The empirical evaluation demonstrates that our method using real-world properties improves the performance of generating an event timeline over diverse environments. Taiki Miyanishi, Takuya Maekawa, Motoaki Kawanabe |
AAAI | 1 |
| 2018 | A Sparse Coding Framework for Gaze Prediction in Egocentric VideoabstractTo efficiently process and understand a large amount of incoming visual information from first-person perspective (i.e. egocentric vision), predicting human gaze is important. However, even though people continuously gaze in noisy environments, most existing gaze prediction methods mainly use image saliency, which is sensitive to noise in the real-world. To address this issue, we propose a sparse coding-based saliency detection method for gaze prediction. Our model uses a cost function with the 10 norm as a sparse constraint that can control the area of visual saliency in response to the contents of egocentric vision in intuitive and consistent ways. Moreover, we use canonical correlation analysis (CCA) to combine different types of features for reducing noise and the computational complexity. We also utilize the temporal continuity of image frames when defining our saliency. Experiments using a real-world gaze dataset show that our proposed approach outperforms the state-of-the-art algorithms on gaze prediction in egocentric videos. Yujie Li 0002, Atsunori Kanemura, Hideki Asoh, Taiki Miyanishi, Motoaki Kawanabe |
ICASSP | 4 |
| 2018 | Answering Mixed Type Questions about Daily Living EpisodesabstractWe propose a physical-world question-answering (QA) method, where the system answers a text question about the physical world by searching a given sequence of sentences about daily-life episodes. To address various information needs in a physical world situation, the physical-world QA methods have to generate mixed-type responses (e.g. word sequence, word set, number, and time as well as a single word) according to the content of questions, after reading physical-world event stories. Most existing methods only provide words or choose answers from multiple candidates. In this paper, we use multiple decoders to generate a mixed-type answer encoding daily episodes with a memory architecture that can capture short- and long-term event dependencies. Results using house-activity stories show that the use of multiple decoders with memory components is effective for answering various physical-world QA questions. Taiki Miyanishi, Atsunori Kanemura, Motoaki Kawanabe |
IJCAI | 1 |
| 2017 | Extracting key frames from first-person videos in the common space of multiple sensorsabstractSelecting authentic scenes about activities of daily living (ADL) is useful to support our memory of everyday life. Key-frame extraction for first-person vision (FPV) videos is a core technology to realize such memory assistant. However, most existing key-frame extraction methods have mainly focused on stable scenes not related to ADL and only used visual signals of the image sequence even though the activities usually associate with our visual experience. To deal with dynamically changing scenes of FPV about daily activities, integrating motion and visual signals are essential. In this paper, we present a novel key-frame extraction method for ADL, which integrates multi-modal sensor signals to temper noise and detect salient activities. Our proposed method projects motion and visual features to a shared space by a probabilistic canonical correlation analysis and selects key frames there. The experimental results using ADL datasets collected in a house suggest that our key-frame extraction technique running in the shared space improves the precision of extracted key frames and the coverage of the entire video. Yujie Li 0002, Atsunori Kanemura, Hideki Asoh, Taiki Miyanishi, Motoaki Kawanabe |
ICIP | 4 |
| 2017 | Key frame extraction from first-person video with multi-sensor integrationabstractFirst-person videos (FPVs) in daily living help us to memorize our life experience and information systems to process daily activities. Summarizing FPVs into key frames that represent the entire data would allow us to remember our memory in the past and computers to efficiently process the data. However, most video summarization approaches only use visual information, even though our daily activities consist of multiple modalities such as movements and sounds. FPVs are not as stable as movies or sport scenes since the camera attached to the head shakes frequently, and key frame extraction methods rely only on video frames do not always produce satisfactory results. In this paper, we introduce a novel key frame extraction method for FPVs using multiple wearable sensors. To efficiently integrate multimodal sensor signals, our formulation uses sparse dictionary selection, which minimizes a reconstruction error with a subset (key frames) of the original data. We present experimental results with multimodal datasets captured by wearable sensors in a natural environment. The results suggest multi-sensor information improves the precision of extracted key frames as well as the coverage of an entire video sequence. Yujie Li 0002, Atsunori Kanemura, Hideki Asoh, Taiki Miyanishi, Motoaki Kawanabe |
ICME | 4 |
| 2016 | Egocentric Video Search via Physical InteractionsabstractRetrieving past egocentric videos about personal daily life is important to support and augment human memory. Most previous retrieval approaches have ignored the crucial feature of human-physical world interactions, which is greatly related to our memory and experience of daily activities. In this paper, we propose a gesture-based egocentric video retrieval framework, which retrieves past visual experience using body gestures as non-verbal queries. We use a probabilistic framework based on a canonical correlation analysis that models physical interactions through a latent space and uses them for egocentric video retrieval and re-ranking search results. By incorporating physical interactions into the retrieval models, we address the problems resulting from the variability of human motions. We evaluate our proposed method on motion and egocentric video datasets about daily activities in household settings and demonstrate that our egocentric video retrieval framework robustly improves retrieval performance when retrieving past videos from personal and even other persons' video archives. Taiki Miyanishi, Quan Kong, Takuya Maekawa, Hiroki Moriya, Takayuki Suyama |
AAAI | 1 |
| 2016 | Selecting home appliances with smart glass based on contextual informationabstractWe propose a method for selecting home appliances using a smart glass, which facilitates the control of network-connected appliances in a smart house. Our proposed method is image-based appliance selection and enables smart glass users to easily select a particular appliance by just looking at it. The main feature of our method is that it achieves high precision appliance selection using user contextual information such as position and activity, inferred from various sensor data in addition to camera images captured by the glass because such contextual information is greatly related in the home appliance that a user wants to control in her daily life. We design a state-of-the-art appliance selection method by fusing image features extracted by deep learning techniques and context information estimated by non-parametric Bayesian techniques within a framework of multiple kernel learning. Our experimental results, which use sensor data obtained in an actual house equipped with many network-connected appliances, show the effectiveness of our method. Quan Kong, Takuya Maekawa, Taiki Miyanishi, Takayuki Suyama |
UbiComp | 3 |
| 2014 | Time-Aware Latent Concept Expansion for Microblog Search
Taiki Miyanishi, Kazuhiro Seki, Kuniaki Uehara |
ICWSM | 1 |
| 2013 | Improving pseudo-relevance feedback via tweet selectionabstractQuery expansion methods using pseudo-relevance feedback have been shown effective for microblog search because they can solve vocabulary mismatch problems often seen in searching short documents such as Twitter messages (tweets), which are limited to 140 characters. Pseudo-relevance feedback assumes that the top ranked documents in the initial search results are relevant and that they contain topic-related words appropriate for relevance feedback. However, those assumptions do not always hold in reality because the initial search results often contain many irrelevant documents. In such a case, only a few of the suggested expansion words may be useful with many others being useless or even harmful. To overcome the limitation of pseudo-relevance feedback for microblog search, we propose a novel query expansion method based on two-stage relevance feedback that models search interests by manual tweet selection and integration of lexical and temporal evidence into its relevance model. Our experiments using a corpus of microblog data (the Tweets2011 corpus) demonstrate that the proposed two-stage relevance feedback approaches considerably improve search result relevance over almost all topics. Taiki Miyanishi, Kazuhiro Seki, Kuniaki Uehara |
CIKM | 1 |
| 2013 | Combining Recency and Topic-Dependent Temporal Variation for Microblog Search
Taiki Miyanishi, Kazuhiro Seki, Kuniaki Uehara |
ECIR | 1 |
| 2013 | Time-aware structured query suggestionabstractMost commercial search engines have a query suggestion feature, which is designed to capture various possible search intents behind the user's original query. However, even though different search intents behind a given query may have been popular at different time periods in the past, existing query suggestion methods neither utilize nor present such information. In this study, we propose Time-aware Structured Query Suggestion (TaSQS) which clusters query suggestions along a timeline so that the user can narrow down his search from a temporal point of view. Moreover, when a suggested query is clicked, TaSQS presents web pages from query-URL bipartite graphs after ranking them according to the click counts within a particular time period. Our experiments using data from a commercial search engine log show that the time-aware clustering and the time-aware document ranking features of TaSQS are both effective. Taiki Miyanishi, Tetsuya Sakai |
SIGIR | 1 |