Tatsuya Ishigaki

dblp:212/0211 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0003-3278-2345ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 3 first-author · 21 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Assessing the Belief Consistency of Large Language Models on the Logical Conversation Process
abstract
Tomoki Tsujimura, Matīss Rikters, Masaki Asada, Shusaku Egami, Tatsuya Ishigaki, Ken Yano, Hiroya Takamura. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tomoki Tsujimura, Matiss Rikters, Masaki Asada, Shusaku Egami, Tatsuya Ishigaki, Ken Yano, Hiroya Takamura
ACL (1)5
2026 Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
Anum Afzal, Yuki Saito 0001, Hiroya Takamura, Katsuhito Sudoh, Shinnosuke Takamichi, Graham Neubig, Florian Matthes, Tatsuya Ishigaki
LREC8
2026 HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities
abstract
Large Language Models (LLMs) provide flexible natural language processing capabilities, while knowledge graphs (KGs) offer explicit and structured knowledge. Integrating these two in a complementary manner enables the development of reliable and verifiable AI systems. In particular, knowledge graph question answering (KGQA) has attracted attention as a means to reduce LLM hallucinations and to leverage knowledge beyond the training data. However, existing KGQA benchmark datasets are biased toward encyclopedic knowledge, limited to a single modality, and lack fine-grained spatiotemporal data, which limits their applicability to real-world scenarios targeted by Embodied AI. We introduce HOME-KGQA, a novel KGQA benchmark dataset built on a multimodal KG of daily household activities. HOME-KGQA consists of complex, multi-hop natural language questions paired with graph database query languages. Compared to existing benchmarks, it includes more challenging questions that involve multi-level spatiotemporal reasoning, multimodal grounding, and aggregate functions. Experimental results show that the LLM-based KGQA methods fail to achieve performance comparable to that on existing datasets when evaluated on HOME-KGQA. This highlights significant challenges that should be addressed for the real-world deployment of KGQA systems. Our dataset is available at https://github.com/aistairc/home-kgqa
Shusaku Egami, Aoi Ohta, Tomoki Tsujimura, Masaki Asada, Tatsuya Ishigaki, Ken Fukuda, Masahiro Hamasaki, Hiroya Takamura
LREC5
2026 Multimodal Task Interference: A Benchmark and Analysis of History-Target Mismatch in Multimodal LLMs
Masayuki Kawarada, Tatsuya Ishigaki, Hiroya Takamura
LREC2
2026 Evaluating Social Intelligence in LLMs via Japanese Honorifics in Email Generation: A Social Semiotic System Perspective
Muxuan Liu, Tatsuya Ishigaki, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001
LREC2
2025 Measuring Time Delay Tolerance in Third-Person Live Commentary for Super Smash Bros. Ultimate
abstract
This study proposes a methodology for measuring the acceptable delay tolerance for third-person game commentary. Third-person game commentary refers to commentary delivered by someone other than the player, with the role of helping viewers better understand the game and enhancing the viewing experience. With the recent advancement of AI, there has been increasing interest in automating such commentary using video understanding and audio generation. However, automating this process using video understanding and audio generation introduces delays, potentially affecting the naturalness of the commentary. In this context, since the extent to which such delays are acceptable to viewers remains unclear, we address this issue. The tolerance is modeled using an unnormalized Gaussian function. Through experiments on Super Smash Bros. Ultimate with 727 participants, we found that the average acceptable delay for this game is 3.71 seconds, with variations depending on different viewer attributes and gameplay contexts.
Ryosuke Matsushita, Ryosuke Sakai, Koki Fukuda, Shinnosuke Takamichi, Kota Iura, Yuki Saito 0001, Graham Neubig, Katsuhito Sudoh, Hiroya Takamura, Tatsuya Ishigaki
CoG10
2025 Evaluating LLMs' Ability to Understand Numerical Time Series for Text Generation
abstract
Data-to-text generation tasks often involve processing numerical time-series as input such as financial statistics or meteorological data. Although large language models (LLMs) are a powerful approach to data-to-text, we still lack a comprehensive understanding of how well they actually understand time-series data. We therefore introduce a benchmark with 18 evaluation tasks to assess LLMs’ abilities of interpreting numerical time-series, which are categorized into: 1) event detection—identifying maxima and minima; 2) computation—averaging and summation; 3) pairwise comparison—comparing values over time; and 4) inference—imputation and forecasting. Our experiments reveal five key findings: 1) even state-of-the-art LLMs struggle with complex multi-step reasoning; 2) tasks that require extracting values or performing computations within a specified range of the time-series significantly reduce accuracy; 3) instruction tuning offers inconsistent improvements for numerical interpretation; 4) reasoning-based models outperform standard LLMs in complex numerical tasks; and 5) LLMs perform interpolation better than forecasting. These results establish a clear baseline and serve as a wake-up call for anyone aiming to blend fluent language with trustworthy numeric precision in time-series scenarios.
Mizuki Arai, Tatsuya Ishigaki, Masayuki Kawarada, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001
INLG2
2025 QCoder Benchmark: Bridging Language Generation and Quantum Hardware through Simulator-Based Feedback
abstract
Large language models (LLMs) have increasingly been applied to automatic programming code generation. This task can be viewed as a language generation task that bridges natural language, human knowledge, and programming logic. However, it remains underexplored in domains that require interaction with hardware devices, such as quantum programming, where human coders write Python code that is executed on a quantum computer. To address this gap, we introduce QCoder Benchmark, an evaluation framework that assesses LLMs on quantum programming with feedback from simulated hardware devices. Our benchmark offers two key features. First, it supports evaluation using a quantum simulator environment beyond conventional Python execution, allowing feedback of domain-specific metrics such as circuit depth, execution time, and error classification, which can be used to guide better generation. Second, it incorporates human-written code submissions collected from real programming contests, enabling both quantitative comparisons and qualitative analyses of LLM outputs against human-written codes. Our experiments reveal that even advanced models like GPT-4o achieve only around 18.97% accuracy, highlighting the difficulty of the benchmark. In contrast, reasoning-based models such as o3 reach up to 78% accuracy, outperforming averaged success rates of human-written codes (39.98%). We release the QCoder Benchmark dataset and public evaluation API to support further research.
Taku Mikuriya, Tatsuya Ishigaki, Masayuki Kawarada, Shunya Minami, Tadashi Kadowaki, Yohichi Suzuki, Soshun Naito, Shunya Takada, Takumi Kato, Tamotsu Basseda, Reo Yamada, Hiroya Takamura
INLG2
2025 Live Football Commentary (LFC): A Large-Scale Dataset for Building Football Commentary Generation Models
abstract
Live football commentary brings the atmosphere and excitement of matches to fans in real time, but producing it requires costly professional announcers. We address this challenge by formulating commentary generation from player- and ball-tracking coordinates as a new language–generation task. To facilitate research on this problem we compile the Live Football Commentary (LFC) dataset, 12,440 time-stamped Japanese utterances aligned with tracking data for 40 J1 League matches ( 60 h). We benchmark three LLM-based baselines that receive the tracking data (i) as plain text, (ii) as pitch-map images, or (iii) in both modalities. Human evaluation shows that the text encoding already outperforms image and multimodal variants in both accuracy and relevance, indicating that current LLMs exploit structured coordinates more effectively than raw visuals. We release the LFC transcripts and evaluation code to establish a public test bed and spur future work on tracking-based commentary generation, saliency detection, and cross-modal integration.
Taiga Someya, Tatsuya Ishigaki, Hiroya Takamura
INLG2
2025 A Comparative Study of Demonstration Selection for Practical Large Language Models-Based Next POI Prediction
Ryo Nishida, Masayuki Kawarada, Tatsuya Ishigaki, Hiroya Takamura, Masaki Onishi
PRICAI3
2025 Exploring the Design of Multi-Agent LLM Dialogues for Research Ideation
abstract
Large language models (LLMs) are increasingly used to support creative tasks such as research idea generation. While recent work has shown that structured dialogues between LLMs can improve the novelty and feasibility of generated ideas, the optimal design of such interactions remains unclear. In this study, we conduct a comprehensive analysis of multi-agent LLM dialogues for scientific ideation. We compare different configurations of agent roles, number of agents, and dialogue depth to understand how these factors influence the novelty and feasibility of generated ideas. Our experimental setup includes settings where one agent generates ideas and another critiques them, enabling iterative improvement. Our results show that enlarging the agent cohort, deepening the interaction depth, and broadening agent persona heterogeneity each enrich the diversity of generated ideas. Moreover, specifically increasing critic-side diversity within the ideation–critique–revision loop further boosts the feasibility of the final proposals. Our findings offer practical guidelines for building effective multi-agent LLM systems for scientific ideation.
Keisuke Ueda, Wataru Hirota, Kosuke Takahashi, Takahiro Omi, Kosuke Arima, Tatsuya Ishigaki
SIGDIAL6
2024 Prompting for Numerical Sequences: A Case Study on Market Comment Generation
abstract
Large language models (LLMs) have been applied to a wide range of data-to-text generation tasks, including tables, graphs, and time-series numerical data-to-text settings. While research on generating prompts for structured data such as tables and graphs is gaining momentum, in-depth investigations into prompting for time-series numerical data are lacking. Therefore, this study explores various input representations, including sequences of tokens and structured formats such as HTML, LaTeX, and Python-style codes. In our experiments, we focus on the task of Market Comment Generation, which involves taking a numerical sequence of stock prices as input and generating a corresponding market comment. Contrary to our expectations, the results show that prompts resembling programming languages yield better outcomes, whereas those similar to natural languages and longer formats, such as HTML and LaTeX, are less effective. Our findings offer insights into creating effective prompts for tasks that generate text from numerical sequences.
Masayuki Kawarada, Tatsuya Ishigaki, Hiroya Takamura
LREC/COLING2
2024 Leveraging Plug-and-Play Models for Rhetorical Structure Control in Text Generation
abstract
We propose a method that extends a BARTbased language generator using the plug-andplay language model to control the rhetorical structure of generated text.Our approach considers rhetorical relations between clauses and generates sentences that reflect this structure using plug-and-play language models.We evaluated our method using the Newsela corpus, which consists of texts at various levels of English proficiency.Our experiments demonstrated that our method outperforms the vanilla BART in terms of the correctness of output discourse and rhetorical structures.In existing methods, the rhetorical structure tends to deteriorate when compared to the baseline, the vanilla BART, as measured by n-gram overlap metrics such as BLEU.However, our proposed method does not exhibit this significant deterioration, demonstrating its advantage.
Yuka Yokogawa, Tatsuya Ishigaki, Hiroya Takamura, Yusuke Miyao, Ichiro Kobayashi 0001
INLG2
2024 Evaluating LlaMA-2's Adaptation to Social Context in Japanese Emails via Fine-Tuning
Muxuan Liu, Tatsuya Ishigaki, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001
PACLIC2
2024 Pretraining and Updates of Domain-Specific LLM: A Case Study in the Japanese Business Domain
Kosuke Takahashi, Takahiro Omi, Kosuke Arima, Tatsuya Ishigaki
PACLIC4
2023 Constructing a Japanese Business Email Corpus Based on Social Situations
Muxuan Liu, Tatsuya Ishigaki, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001
PACLIC2
2023 Training Generative Question-Answering on Synthetic Data Obtained from an Instruct-tuned Model
Kosuke Takahashi, Takahiro Omi, Kosuke Arima, Tatsuya Ishigaki
PACLIC4
2022 Open-domain Video Commentary Generation
abstract
Edison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topić, Yusuke Miyao, Ichiro Kobayashi, Hiroya Takamura. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Edison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topic, Yusuke Miyao, Ichiro Kobayashi 0001, Hiroya Takamura
EMNLP3
2022 Automating Horizon Scanning in Future Studies
abstract
We introduce document retrieval and comment generation tasks for automating horizon scanning. This is an important task in the field of futurology that collects sufficient information for predicting drastic societal changes in the mid- or long-term future. The steps used are: 1) retrieving news articles that imply drastic changes, and 2) writing subjective comments on each article for others’ ease of understanding. As a first step in automating these tasks, we create a dataset that contains 2,266 manually collected news articles with comments written by experts. We analyze the collected documents and comments regarding characteristic words, the distance to general articles, and contents in the comments. Furthermore, we compare several methods for automating horizon scanning. Our experiments show that 1) manually collected articles are different from general articles regarding the words used and semantic distances, 2) the contents in the comment can be classified into several categories, and 3) a supervised model trained on our dataset achieves a better performance. The contributions are: 1) we propose document retrieval and comment generation tasks for horizon scanning, 2) create and analyze a new dataset, and 3) report the performance of several models and show that comment generation tasks are challenging.
Tatsuya Ishigaki, Suzuko Nishino, Sohei Washino, Hiroki Igarashi, Yukari Nagai, Yuichi Washida, Akihiko Murai
LREC1
2021 Generating Racing Game Commentary from Vision, Language, and Structured Data
abstract
We propose the task of automatically generating commentaries for races in a motor racing game, from vision, structured numerical, and textual data.Commentaries provide information to support spectators in understanding events in races.Commentary generation models need to interpret the race situation and generate the correct content at the right moment.We divide the task into two subtasks: utterance timing identification and utterance generation.Because existing datasets do not have such alignments of data in multiple modalities, this setting has not been explored in depth.In this study, we introduce a new large-scale dataset that contains aligned video data, structured numerical data, and transcribed commentaries that consist of 129,226 utterances in 1,389 races in a game.Our analysis reveals that the characteristics of commentaries change depending on time and viewpoints.Our experiments on the subtasks show that it is still challenging for a state-of-the-art vision encoder to capture useful information from videos to generate accurate commentaries.We make the dataset and baseline implementation publicly available for further research.1
Tatsuya Ishigaki, Goran Topic, Yumi Hamazono, Hiroshi Noji, Ichiro Kobayashi 0001, Yusuke Miyao, Hiroya Takamura
INLG1
2021 Unpredictable Attributes in Market Comment Generation
Yumi Hamazono, Tatsuya Ishigaki, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi 0001
PACLIC2
2021 Controlling contents in data-to-document generation with human-designed topic labels
Kasumi Aoki, Akira Miyazawa, Tatsuya Ishigaki, Tatsuya Aoki, Hiroshi Noji, Keiichi Goshima, Hiroya Takamura, Yusuke Miyao, Ichiro Kobayashi 0001
Comput. Speech Lang.3
2020 Learning with Contrastive Examples for Data-to-Text Generation
abstract
Yui Uehara, Tatsuya Ishigaki, Kasumi Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi, Hiroya Takamura, Yusuke Miyao. Proceedings of the 28th International Conference on Computational Linguistics. 2020.
Yui Uehara, Tatsuya Ishigaki, Kasumi Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi 0001, Hiroya Takamura, Yusuke Miyao
COLING2
2020 Neural Query-Biased Abstractive Summarization Using Copying Mechanism
Tatsuya Ishigaki, Hen-Hsen Huang, Hiroya Takamura, Hsin-Hsi Chen, Manabu Okumura
ECIR (2)1
2020 Distant Supervision for Extractive Question Summarization
Tatsuya Ishigaki, Kazuya Machida, Hayato Kobayashi, Hiroya Takamura, Manabu Okumura
ECIR (2)1
2020 Semi-supervised Extractive Question Summarization Using Question-Answer Pairs
Kazuya Machida, Tatsuya Ishigaki, Hayato Kobayashi, Hiroya Takamura, Manabu Okumura
ECIR (2)2
2019 Learning to Select, Track, and Generate for Data-to-Text
abstract
Hayate Iso, Yui Uehara, Tatsuya Ishigaki, Hiroshi Noji, Eiji Aramaki, Ichiro Kobayashi, Yusuke Miyao, Naoaki Okazaki, Hiroya Takamura. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Hayate Iso, Yui Uehara, Tatsuya Ishigaki, Hiroshi Noji, Eiji Aramaki, Ichiro Kobayashi 0001, Yusuke Miyao, Naoaki Okazaki, Hiroya Takamura
ACL (1)3
2019 Controlling Contents in Data-to-Document Generation with Human-Designed Topic Labels
abstract
Kasumi Aoki, Akira Miyazawa, Tatsuya Ishigaki, Tatsuya Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi, Hiroya Takamura, Yusuke Miyao. Proceedings of the 12th International Conference on Natural Language Generation. 2019.
Kasumi Aoki, Akira Miyazawa, Tatsuya Ishigaki, Tatsuya Aoki, Hiroshi Noji, Keiichi Goshima, Ichiro Kobayashi 0001, Hiroya Takamura, Yusuke Miyao
INLG3
2018 Generating Market Comments Referring to External Resources
abstract
Tatsuya Aoki, Akira Miyazawa, Tatsuya Ishigaki, Keiichi Goshima, Kasumi Aoki, Ichiro Kobayashi, Hiroya Takamura, Yusuke Miyao. Proceedings of the 11th International Conference on Natural Language Generation. 2018.
Tatsuya Aoki, Akira Miyazawa, Tatsuya Ishigaki, Keiichi Goshima, Kasumi Aoki, Ichiro Kobayashi 0001, Hiroya Takamura, Yusuke Miyao
INLG3
2017 Summarizing Lengthy Questions
abstract
In this research, we propose the task of question summarization. We first analyzed question-summary pairs extracted from a Community Question Answering (CQA) site, and found that a proportion of questions cannot be summarized by extractive approaches but requires abstractive approaches. We created a dataset by regarding the question-title pairs posted on the CQA site as question-summary pairs. By using the data, we trained extractive and abstractive summarization models, and compared them based on ROUGE scores and manual evaluations. Our experimental results show an abstractive method using an encoder-decoder model with a copying mechanism achieves better scores for both ROUGE-2 F-measure and the evaluations by human judges.
Tatsuya Ishigaki, Hiroya Takamura, Manabu Okumura
IJCNLP(1)1