Hiroki Ouchi

dblp:148/4520 · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
14since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 9 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction
abstract
Recent advances in multimodal large language models (MLLMs) have significantly enhanced video understanding capabilities, opening new possibilities for practical applications. Yet current video benchmarks focus largely on indoor scenes or short-range outdoor activities, leaving the challenges associated with long-distance travel largely unexplored. Mastering extended geospatial-temporal trajectories is critical for next-generation MLLMs, underpinning real-world tasks such as embodied-AI planning and navigation. To bridge this gap, we present VIR-Bench, a novel benchmark consisting of 200 travel videos that frames itinerary reconstruction as a challenging task designed to evaluate and push forward MLLMs' geospatial-temporal intelligence. Experimental results reveal that state-of-the-art MLLMs, including proprietary ones, struggle to achieve high scores, underscoring the difficulty of handling videos that span extended spatial and temporal scales. Moreover, we conduct an in-depth case study in which we develop a prototype travel-planning agent that leverages the insights gained from VIR-Bench. The agent’s markedly improved itinerary recommendations verify that our evaluation protocol not only benchmarks models effectively but also translates into concrete performance gains in user-facing applications.
Eiki Murata, Lingfang Zhang, Ayako Sato, So Fukuda, Keisuke Nakao, Yusuke Nakamura, Sebastian Zwirner, Yi-Chia Chen, Hiroyuki Otomo, Hiroki Ouchi, Daisuke Kawahara
AAAI13
2026 A Large-Scale Dataset for Linking-Based Geocoding
Hibiki Nakatani, Yuichiro Yasui, Ryosuke Wakamoto, Masayuki Ishii, Tetsuhisa Suizu, Hiroki Ouchi, Taro Watanabe
LREC6
2025 Graph-Structured Trajectory Extraction from Travelogues
abstract
Aitaro Yamamoto, Hiroyuki Otomo, Hiroki Ouchi, Shohei Higashiyama, Hiroki Teranishi, Hiroyuki Shindo, Taro Watanabe. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Aitaro Yamamoto, Hiroyuki Otomo, Hiroki Ouchi, Shohei Higashiyama, Hiroki Teranishi, Hiroyuki Shindo, Taro Watanabe
ACL (1)3
2025 A Text Embedding Model with Contrastive Example Mining for Point-of-Interest Geocoding
abstract
Geocoding is a fundamental technique that links location mentions to their geographic positions, which is important for understanding texts in terms of where the described events occurred. Unlike most geocoding studies that targeted coarse-grained locations, we focus on geocoding at a fine-grained point-of-interest (POI) level. To address the challenge of finding appropriate geo-database entries from among many candidates with similar POI names, we develop a text embedding-based geocoding model and investigate (1) entry encoding representations and (2) hard negative mining approaches suitable for enhancing the model’s disambiguation ability. Our experiments show that the second factor significantly impact the geocoding accuracy of the model.
Hibiki Nakatani, Hiroki Teranishi, Shohei Higashiyama, Yuya Sawada, Hiroki Ouchi, Taro Watanabe
COLING5
2025 AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising
abstract
Peinan Zhang, Yusuke Sakai, Masato Mita, Hiroki Ouchi, Taro Watanabe. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Peinan Zhang, Yusuke Sakai 0010, Masato Mita, Hiroki Ouchi, Taro Watanabe
NAACL (Long Papers)4
2024 Constructing Indonesian-English Travelogue Dataset
abstract
Research in low-resource language is often hampered due to the under-representation of how the language is being used in reality. This is particularly true for Indonesian language because there is a limited variety of textual datasets, and majority were acquired from official sources with formal writing style. All the more for the task of geoparsing, which could be implemented for navigation and travel planning applications, such datasets are rare, even in the high-resource languages, such as English. Being aware of the need for a new resource in both languages for this specific task, we constructed a new dataset comprising both Indonesian and English from personal travelogue articles. Our dataset consists of 88 articles, exactly half of them written in each language. We covered both named and nominal expressions of four entity types related to travel: location, facility, transportation, and line. We also conducted experiments by training classifiers to recognise named entities and their nominal expressions. The results of our experiments showed a promising future use of our dataset as we obtained F1-score above 0.9 for both languages.
Eunike Andriani Kardinata, Hiroki Ouchi, Taro Watanabe
LREC/COLING2
2024 Can Language Models Induce Grammatical Knowledge from Indirect Evidence?
abstract
What kinds of and how much data is necessary for language models to induce grammatical knowledge to judge sentence acceptability?Recent language models still have much room for improvement in their data efficiency compared to humans.This paper investigates whether language models efficiently use indirect data (indirect evidence), from which they infer sentence acceptability.In contrast, humans use indirect evidence efficiently, which is considered one of the inductive biases contributing to efficient language acquisition.To explore this question, we introduce the Wug In-Direct Evidence Test (WIDET), a dataset consisting of training instances inserted into the pre-training data and evaluation instances.We inject synthetic instances with newly coined wug words into pretraining data and explore the model's behavior on evaluation data that assesses grammatical acceptability regarding those words.We prepare the injected instances by varying their levels of indirectness and quantity.Our experiments surprisingly show that language models do not induce grammatical knowledge even after repeated exposure to instances with the same structure but differing only in lexical items from evaluation instances in certain language phenomena.Our findings suggest a potential direction for future research: developing models that use latent indirect evidence to induce grammatical knowledge.
Miyu Oba, Yohei Oseki, Akiyo Fukatsu, Akari Haga, Hiroki Ouchi, Taro Watanabe, Saku Sugawara
EMNLP5
2024 Text2Traj2Text: Learning-by-Synthesis Framework for Contextual Captioning of Human Movement Trajectories
abstract
This paper presents Text2Traj2Text, a novel learning-by-synthesis framework for captioning possible contexts behind shopper's trajectory data in retail stores.Our work will impact various retail applications that need better customer understanding, such as targeted advertising and inventory management.The key idea is leveraging large language models to synthesize a diverse and realistic collection of contextual captions as well as the corresponding movement trajectories on a store map.Despite learned from fully synthesized data, the captioning model can generalize well to trajectories/captions created by real human subjects.Our systematic evaluation confirmed the effectiveness of the proposed framework over competitive approaches in terms of ROUGE and BERT Score metrics.
Hikaru Asano, Ryo Yonetani, Taiki Sekii, Hiroki Ouchi
INLG4
2022 Iterative Span Selection: Self-Emergence of Resolving Orders in Semantic Role Labeling
abstract
Semantic Role Labeling (SRL) is the task of labeling semantic arguments for marked semantic predicates. Semantic arguments and their predicates are related in various distinct manners, of which certain semantic arguments are a necessity while others serve as an auxiliary to their predicates. To consider such roles and relations of the arguments in the labeling order, we introduce iterative argument identification (IAI), which combines global decoding and iterative identification for the semantic arguments. In experiments, we first realize that the model with random argument labeling orders outperforms other heuristic orders such as the conventional left-to-right labeling order. Combined with simple reinforcement learning, the proposed model spontaneously learns the optimized labeling orders that are different from existing heuristic orders. The proposed model with the IAI algorithm achieves competitive or outperforming results from the existing models in the standard benchmark datasets of span-based SRL: CoNLL-2005 and CoNLL-2012.
Shuhei Kurita, Hiroki Ouchi, Kentaro Inui, Satoshi Sekine
COLING2
2022 Law Retrieval with Supervised Contrastive Learning Using the Hierarchical Structure of Law
Jungmin Choi, Ukyo Honda, Taro Watanabe, Hiroki Ouchi, Kentaro Inui
PACLIC4
2022 N-best Response-based Analysis of Contradiction-awareness in Neural Response Generation Models
abstract
Avoiding the generation of responses that contradict the preceding context is a significant challenge in dialogue response generation.One feasible method is post-processing, such as filtering out contradicting responses from a resulting n-best response list.In this scenario, the quality of the n-best list considerably affects the occurrence of contradictions because the final response is chosen from this n-best list.This study quantitatively analyzes the contextual contradiction-awareness of neural response generation models using the consistency of the n-best lists.Particularly, we used polar questions as stimulus inputs for concise and quantitative analyses.Our tests illustrate the contradiction-awareness of recent neural response generation models and methodologies, followed by a discussion of their properties and limitations.
Shiki Sato, Reina Akama, Hiroki Ouchi, Ryoko Tokuhisa, Jun Suzuki 0001, Kentaro Inui
SIGDIAL3
2021 Pseudo Zero Pronoun Resolution Improves Zero Anaphora Resolution
abstract
Masked language models (MLMs) have contributed to drastic performance improvements with regard to zero anaphora resolution (ZAR).To further improve this approach, in this study, we made two proposals.The first is a new pretraining task that trains MLMs on anaphoric relations with explicit supervision, and the second proposal is a new finetuning method that remedies a notorious issue, the pretrainfinetune discrepancy.Our experiments on Japanese ZAR demonstrated that our two proposals boost the state-of-the-art performance, and our detailed analysis provides new insights on the remaining challenges.
Ryuto Konno, Shun Kiyono, Yuichiroh Matsubayashi, Hiroki Ouchi, Kentaro Inui
EMNLP (1)4
2021 Instance-Based Neural Dependency Parsing
abstract
Abstract Interpretable rationales for model predictions are crucial in practical applications. We develop neural models that possess an interpretable inference process for dependency parsing. Our models adopt instance-based inference, where dependency edges are extracted and labeled by comparing them to edges in a training set. The training edges are explicitly used for the predictions; thus, it is easy to grasp the contribution of each edge to the predictions. Our experiments show that our instance-based models achieve competitive accuracy with standard neural models and have the reasonable plausibility of instance-based explanations.
Hiroki Ouchi, Jun Suzuki 0001, Sosuke Kobayashi, Sho Yokoi, Tatsuki Kuribayashi, Masashi Yoshikawa, Kentaro Inui
Trans. Assoc. Comput. Linguistics1
2021 Corruption Is Not All Bad: Incorporating Discourse Structure Into Pre-Training via Corruption for Essay Scoring
abstract
Existing approaches for automated essay scoring and document representation learning typically rely on discourse parsers to incorporate discourse structure into text representation. However, the performance of parsers is not always adequate, especially when they are used on noisy texts, such as student essays. In this paper, we propose an unsupervised pre-training approach to capture discourse structure of essays in terms of coherence and cohesion that does not require any discourse parser or annotation. We introduce several types of token, sentence and paragraph-level corruption techniques for our proposed pre-training approach and augment masked language modeling pre-training with our pre-training method to leverage both contextualized and discourse information. Our proposed unsupervised approach achieves a new state-of-the-art result on the task of essay Organization scoring.
Farjana Sultana Mim, Naoya Inoue, Paul Reisert, Hiroki Ouchi, Kentaro Inui
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Instance-Based Learning of Span Representations: A Case Study through Named Entity Recognition
abstract
Interpretable rationales for model predictions play a critical role in practical applications.In this study, we develop models possessing interpretable inference process for structured prediction.Specifically, we present a method of instance-based learning that learns similarities between spans.At inference time, each span is assigned a class label based on its similar spans in the training set, where it is easy to understand how much each training instance contributes to the predictions.Through empirical analysis on named entity recognition, we demonstrate that our method enables to build models that have high interpretability without sacrificing performance.
Hiroki Ouchi, Jun Suzuki 0001, Sosuke Kobayashi, Sho Yokoi, Tatsuki Kuribayashi, Ryuto Konno, Kentaro Inui
ACL1
2020 Evaluating Dialogue Generation Systems via Response Selection
abstract
Existing automatic evaluation metrics for open-domain dialogue response generation systems correlate poorly with human evaluation.We focus on evaluating response generation systems via response selection.To evaluate systems properly via response selection, we propose a method to construct response selection test sets with well-chosen false candidates.Specifically, we propose to construct test sets filtering out some types of false candidates: (i) those unrelated to the ground-truth response and (ii) those acceptable as appropriate responses.Through experiments, we demonstrate that evaluating systems via response selection with the test set developed by our method correlates more strongly with human evaluation, compared with widely used automatic evaluation metrics such as BLEU.
Shiki Sato, Reina Akama, Hiroki Ouchi, Jun Suzuki 0001, Kentaro Inui
ACL3
2020 An Empirical Study of Contextual Data Augmentation for Japanese Zero Anaphora Resolution
abstract
One critical issue of zero anaphora resolution (ZAR) is the scarcity of labeled data.This study explores how effectively this problem can be alleviated by data augmentation.We adopt a state-ofthe-art data augmentation method, called the contextual data augmentation (CDA), that generates labeled training instances using a pretrained language model.The CDA has been reported to work well for several other natural language processing tasks, including text classification and machine translation (Kobayashi, 2018;Wu et al., 2019;Gao et al., 2019).This study addresses two underexplored issues on CDA, that is, how to reduce the computational cost of data augmentation and how to ensure the quality of the generated data.We also propose two methods to adapt CDA to ZAR: [MASK]-based augmentation and linguistically-controlled masking.Consequently, the experimental results on Japanese ZAR show that our methods contribute to both the accuracy gain and the computation cost reduction.Our closer analysis reveals that the proposed method can improve the quality of the augmented training data when compared to the conventional CDA.
Ryuto Konno, Yuichiroh Matsubayashi, Shun Kiyono, Hiroki Ouchi, Kentaro Inui
COLING4
2019 An Empirical Study of Span Representations in Argumentation Structure Parsing
abstract
Tatsuki Kuribayashi, Hiroki Ouchi, Naoya Inoue, Paul Reisert, Toshinori Miyoshi, Jun Suzuki, Kentaro Inui. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Tatsuki Kuribayashi, Hiroki Ouchi, Naoya Inoue, Paul Reisert, Toshinori Miyoshi, Jun Suzuki 0001, Kentaro Inui
ACL (1)2
2019 Transductive Learning of Neural Language Models for Syntactic and Semantic Analysis
abstract
Hiroki Ouchi, Jun Suzuki, Kentaro Inui. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Hiroki Ouchi, Jun Suzuki 0001, Kentaro Inui
EMNLP/IJCNLP (1)1
2018 Addressee and Response Selection for Multilingual Conversation
abstract
Developing conversational systems that can converse in many languages is an interesting challenge for natural language processing. In this paper, we introduce multilingual addressee and response selection. In this task, a conversational system predicts an appropriate addressee and response for an input message in multiple languages. A key to developing such multilingual responding systems is how to utilize high-resource language data to compensate for low-resource language data. We present several knowledge transfer methods for conversational systems. To evaluate our methods, we create a new multilingual conversation dataset. Experiments on the dataset demonstrate the effectiveness of our methods.
Motoki Sato, Hiroki Ouchi, Yuta Tsuboi
COLING2
2018 A Span Selection Model for Semantic Role Labeling
abstract
We present a simple and accurate span-based model for semantic role labeling (SRL).Our model directly takes into account all possible argument spans and scores them for each label.At decoding time, we greedily select higher scoring labeled spans.One advantage of our model is to allow us to design and use spanlevel features, that are difficult to use in tokenbased BIO tagging approaches.Experimental results demonstrate that our ensemble model achieves the state-of-the-art results, 87.4 F1 and 87.0 F1 on the CoNLL-2005 and 2012 datasets, respectively.
Hiroki Ouchi, Hiroyuki Shindo, Yuji Matsumoto 0001
EMNLP1
2018 Suspicious News Detection Using Micro Blog Text
Tsubasa Tagami, Hiroki Ouchi, Hiroki Asano, Kazuaki Hanawa, Kaori Uchiyama, Kaito Suzuki, Kentaro Inui, Atsushi Komiya, Atsuo Fujimura, Ryo Yamashita, Hitofumi Yanai, Akinori Machino
PACLIC2
2017 Neural Modeling of Multi-Predicate Interactions for Japanese Predicate Argument Structure Analysis
abstract
The performance of Japanese predicate argument structure (PAS) analysis has improved in recent years thanks to the joint modeling of interactions between multiple predicates.However, this approach relies heavily on syntactic information predicted by parsers, and suffers from error propagation.To remedy this problem, we introduce a model that uses grid-type recurrent neural networks.The proposed model automatically induces features sensitive to multi-predicate interactions from the word sequence information of a sentence.Experiments on the NAIST Text Corpus demonstrate that without syntactic information, our model outperforms previous syntax-dependent models.
Hiroki Ouchi, Hiroyuki Shindo, Yuji Matsumoto 0001
ACL (1)1
2016 Addressee and Response Selection for Multi-Party Conversation
abstract
To create conversational systems working in actual situations, it is crucial to assume that they interact with multiple agents.In this work, we tackle addressee and response selection for multi-party conversation, in which systems are expected to select whom they address as well as what they say.The key challenge of this task is to jointly model who is talking about what in a previous context.For the joint modeling, we propose two modeling frameworks: 1) static modeling and 2) dynamic modeling.To show benchmark results of our frameworks, we created a multi-party conversation corpus.Our experiments on the dataset show that the recurrent neural network based models of our frameworks robustly predict addressees and responses in conversations with a large number of agents.
Hiroki Ouchi, Yuta Tsuboi
EMNLP1
2016 Transition-Based Dependency Parsing Exploiting Supertags
abstract
Lexical information, including surface word form and part-of-speech (POS) information, plays a crucial role when predicting ambiguous dependency relationships in dependency parsing. However, for resolving dependency ambiguities, surface word information may be too sparse, while POS information may be too coarse. Supertags, which are lexical templates that represent rich syntactic information, have been shown to provide effective features at an intermediate level on the coarse-to-fine scale. In this work, we present a supertag design framework that allows us to instantiate various supertag sets based on the dependency structures. Using this framework, we instantiate various supertag sets and utilize them as features in transition-based dependency parsing systems. Performing experiments on the Penn Treebank and Universal Dependencies data sets, we show that our supertags are effective for transition-based parsers in multilingual parsing as well as English parsing. The comparison of the results of the different supertag sets shows that it is crucial to incorporate the head directionality, head labels, and dependent possession information in supertags to improve the parser performance.
Hiroki Ouchi, Kevin Duh, Hiroyuki Shindo, Yuji Matsumoto 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Joint Case Argument Identification for Japanese Predicate Argument Structure Analysis
abstract
Hiroki Ouchi, Hiroyuki Shindo, Kevin Duh, Yuji Matsumoto. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Hiroki Ouchi, Hiroyuki Shindo, Kevin Duh, Yuji Matsumoto 0001
ACL (1)1
2014 Improving Dependency Parsers with Supertags
abstract
Transition-based dependency parsing systems can utilize rich feature representations.However, in practice, features are generally limited to combinations of lexical tokens and part-of-speech tags.In this paper, we investigate richer features based on supertags, which represent lexical templates extracted from dependency structure annotated corpus.First, we develop two types of supertags that encode information about head position and dependency relations in different levels of granularity.Then, we propose a transition-based dependency parser that incorporates the predictions from a CRF-based supertagger as new features.On standard English Penn Treebank corpus, we show that our supertag features achieve parsing improvements of 1.3% in unlabeled attachment, 2.07% root attachment, and 3.94% in complete tree accuracy.
Hiroki Ouchi, Kevin Duh, Yuji Matsumoto 0001
EACL1