Yusuke Sakai 0010

dblp:332/6403 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness
abstract
The hubness problem, in which hub embeddings are close to many unrelated examples, occurs often in high-dimensional embedding spaces and may pose a practical threat for purposes such as information retrieval and automatic evaluation metrics.In particular, since cross-modal similarity between text and images cannot be calculated by direct comparisons, such as string matching, cross-modal encoders that project different modalities into a shared space are helpful for various cross-modal applications, and thus, the existence of hubs may pose practical threats.To reveal the vulnerabilities of cross-modal encoders, we propose a method for identifying the hub embedding and its corresponding hub text.Experiments on image captioning evaluation in MSCOCO and nocaps along with image-to-text retrieval tasks in MSCOCO and Flickr30k showed that our method can identify a single hub text that unreasonably achieves comparable or higher similarity scores than human-written reference captions in many images, thereby revealing the vulnerabilities in cross-modal encoders.
Hiroyuki Deguchi 0002, Katsuki Chousa, Yusuke Sakai 0010
ACL (1)3
2026 HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
abstract
Recently, we have often observed hallucinated citations or references that do not correspond to any existing work in papers under review, preprints, or published papers.Such hallucinated citations pose a serious concern to scientific reliability.When they appear in accepted papers, they may also negatively affect the credibility of conferences.In this study, we refer to hallucinated citations as "HalluCitation" and systematically investigate their prevalence and impact.We analyze all papers published at ACL, NAACL, and EMNLP in 2024 and 2025, including main conference, Findings, and workshop papers.Our analysis reveals that over 300 papers contain at least one HalluCitation, most of which were published in 2025.Notably, half of these papers were identified at EMNLP 2025, the most recent conference, indicating that this issue is rapidly increasing.Moreover, more than 100 such papers were accepted as main conference and Findings papers at EMNLP 2025, affecting the credibility.
Yusuke Sakai 0010, Hidetaka Kamigaito, Taro Watanabe
ACL (1)1
2026 Grammatical Error Correction Evaluation by Optimally Transporting Edit Representation
abstract
Abstract Automatic evaluation in grammatical error correction (GEC) is crucial for selecting the best-performing systems. Currently, reference-based metrics are a popular choice, which basically measure the similarity between hypothesis and reference sentences. However, similarity measures based on embeddings, such as BERTScore, are often ineffective, since many words in the source sentences remain unchanged in both the hypothesis and the reference. This study focuses on edits specifically designed for GEC, i.e., ERRANT, and computes similarity measured over the edits from the source sentence. To this end, we propose edit vector, a representation for an edit, and introduce a new metric, UOT-ERRANT, which transports these edit vectors from hypothesis to reference using unbalanced optimal transport. Experiments with SEEDA meta-evaluation show that UOT-ERRANT improves evaluation performance, particularly in the +Fluency domain where many edits occur. Moreover, our method is highly interpretable because the transport plan can be interpreted as a soft edit alignment, making UOT-ERRANT a useful metric for both system ranking and analyzing GEC systems. Our code is available from https://github.com/gotutiyan/uot-errant.
Takumi Goto, Yusuke Sakai 0010, Taro Watanabe
Trans. Assoc. Comput. Linguistics2
2025 Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
abstract
In generative commonsense reasoning tasks such as CommonGen, generative large language models (LLMs) compose sentences that include all given concepts.However, when focusing on instruction-following capabilities, if a prompt specifies a concept order, LLMs must generate sentences that adhere to the specified order.To address this, we propose Ordered CommonGen, a benchmark designed to evaluate the compositional generalization and instruction-following abilities of LLMs.This benchmark measures ordered coverage to assess whether concepts are generated in the specified order, enabling a simultaneous evaluation of both abilities.We conducted a comprehensive analysis using 36 LLMs and found that, while LLMs generally understand the intent of instructions, biases toward specific concept order patterns often lead to low-diversity outputs or identical results even when the concept order is altered.Moreover, even the most instructioncompliant LLM achieved only about 75% ordered coverage, highlighting the need for improvements in both instruction-following and compositional generalization capabilities. Concepts Coverage (↑) Similarlity (↓)Diversity (↑) Perplexity (↓) w/o order w/ order Ordered Rate pBLEU pBLEURT Distinct Diverse Rate
Yusuke Sakai 0010, Hidetaka Kamigaito, Taro Watanabe
ACL (1)1
2025 Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding
abstract
Hidetaka Kamigaito, Hiroyuki Deguchi, Yusuke Sakai, Katsuhiko Hayashi, Taro Watanabe. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hidetaka Kamigaito, Hiroyuki Deguchi 0002, Yusuke Sakai 0010, Katsuhiko Hayashi 0001, Taro Watanabe
ACL (1)3
2025 IRR: Image Review Ranking Framework for Evaluating Vision-Language Models
abstract
Large-scale Vision-Language Models (LVLMs) process both images and text, excelling in multimodal tasks such as image captioning and description generation. However, while these models excel at generating factual content, their ability to generate and evaluate texts reflecting perspectives on the same image, depending on the context, has not been sufficiently explored. To address this, we propose IRR: Image Review Rank, a novel evaluation framework designed to assess critic review texts from multiple perspectives. IRR evaluates LVLMs by measuring how closely their judgments align with human interpretations. We validate it using a dataset of images from 15 categories, each with five critic review texts and annotated rankings in both English and Japanese, totaling over 2,000 data instances. Our results indicate that, although LVLMs exhibited consistent performance across languages, their correlation with human annotations was insufficient, highlighting the need for further advancements. These findings highlight the limitations of current evaluation methods and the need for approaches that better capture human reasoning in Vision & Language tasks.
Kazuki Hayashi, Kazuma Onishi, Toma Suzuki, Yusuke Ide, Seiji Gobara, Shigeki Saito, Yusuke Sakai 0010, Hidetaka Kamigaito, Katsuhiko Hayashi 0001, Taro Watanabe
COLING7
2025 SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
abstract
Ashmari Pramodya, Nirasha Nelki, Heshan Shalinda, Chamila Liyanage, Yusuke Sakai, Randil Pushpananda, Ruvan Weerasinghe, Hidetaka Kamigaito, Taro Watanabe. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Ashmari Pramodya, Nirasha Nelki, Heshan Shalinda, Chamila Liyanage, Yusuke Sakai 0010, Randil Pushpananda, Ruvan Weerasinghe, Hidetaka Kamigaito, Taro Watanabe
EMNLP5
2025 LoCt-Instruct: An Automatic Pipeline for Constructing Datasets of Logical Continuous Instructions
abstract
Hongyu Sun, Yusuke Sakai, Haruki Sakajo, Shintaro Ozaki, Kazuki Hayashi, Hidetaka Kamigaito, Taro Watanabe. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yusuke Sakai 0010, Haruki Sakajo, Shintaro Ozaki, Kazuki Hayashi, Hidetaka Kamigaito, Taro Watanabe
EMNLP2
2025 Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark
abstract
Multilingual machine translation (MT) benchmarks play a central role in evaluating the capabilities of modern MT systems.Among them, the FLORES+ benchmark is widely used, offering English-to-many translation data for over 200 languages, curated with strict quality control protocols.However, we study data in four languages (Asante Twi, Japanese, Jinghpaw, and South Azerbaijani) and uncover critical shortcomings in the benchmark's suitability for truly multilingual evaluation.Human assessments reveal that many translations fall below the claimed 90% quality standard, and the annotators report that source sentences are often too domain-specific and culturally biased toward the English-speaking world.We further demonstrate that simple heuristics, such as copying named entities, can yield non-trivial BLEU scores, suggesting vulnerabilities in the evaluation protocol.Notably, we show that MT models trained on high-quality, naturalistic data perform poorly on FLORES+ while achieving significant gains on our domain-relevant evaluation set.Based on these findings, we advocate for multilingual MT benchmarks that use domain-general and culturally neutral source texts rely less on named entities, in order to better reflect real-world translation challenges. 1 * Equal contribution.
Chihiro Taguchi, Seng Mai, Keita Kurabe, Yusuke Sakai 0010, Georgina Agyei, Soudabeh Eslami, David Chiang 0001
EMNLP4
2025 Multilingual Dialogue Generation and Localization with Dialogue Act Scripting
abstract
Non-English dialogue datasets are scarce, and models are often trained or evaluated on translations of English-language dialogues, an approach which can introduce artifacts that reduce their naturalness and cultural appropriateness.This work proposes Dialogue Act Script (DAS), a structured framework for encoding, localizing, and generating multilingual dialogues from abstract intent representations.Rather than translating dialogue utterances directly, DAS enables the generation of new dialogues in the target language that are culturally and contextually appropriate.By using structured dialogue act representations, DAS supports flexible localization across languages, mitigating translationese and enabling more fluent, naturalistic conversations.Human evaluations across Italian, German, and Chinese show that DASgenerated dialogues consistently outperform those produced by both machine and human translators on measures of cultural relevance, coherence, and situational appropriateness. 1
Justin Vasselli, Eunike Andriani Kardinata, Yusuke Sakai 0010, Taro Watanabe
EMNLP3
2025 How to Make the Most of LLMs' Grammatical Knowledge for Acceptability Judgments
abstract
Yusuke Ide, Yuto Nishida, Justin Vasselli, Miyu Oba, Yusuke Sakai, Hidetaka Kamigaito, Taro Watanabe. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yusuke Ide, Yuto Nishida, Justin Vasselli, Miyu Oba, Yusuke Sakai 0010, Hidetaka Kamigaito, Taro Watanabe
NAACL (Long Papers)5
2025 Tonguescape: Exploring Language Models Understanding of Vowel Articulation
abstract
Haruki Sakajo, Yusuke Sakai, Hidetaka Kamigaito, Taro Watanabe. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Haruki Sakajo, Yusuke Sakai 0010, Hidetaka Kamigaito, Taro Watanabe
NAACL (Long Papers)2
2025 AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising
abstract
Peinan Zhang, Yusuke Sakai, Masato Mita, Hiroki Ouchi, Taro Watanabe. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Peinan Zhang, Yusuke Sakai 0010, Masato Mita, Hiroki Ouchi, Taro Watanabe
NAACL (Long Papers)2
2024 Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair
abstract
In Simultaneous Machine Translation (SiMT), training with a simultaneous interpretation (SI) corpus is an effective method for achieving high-quality yet low-latency systems.However, constructing such a corpus is challenging due to high costs, and limitations in annotator capabilities, and as a result, existing SI corpora are limited.Therefore, we propose a method to convert existing speech translation (ST) corpora into interpretation-style corpora, maintaining the original word order and preserving the entire source content using Large Language Models (LLM-SI-Corpus).We demonstrated that fine-tuning SiMT models using the LLM-SI-Corpus reduces latencies while achieving better quality compared to models fine-tuned with other corpora in both speechto-text and text-to-text settings.
Yusuke Sakai 0010, Mana Makinae, Hidetaka Kamigaito, Taro Watanabe
EMNLP1
2024 Simul-MuST-C: Simultaneous Multilingual Speech Translation Corpus Using Large Language Model
abstract
Simultaneous Speech Translation (SiST) begins translating before the entire source input is received, making it crucial to balance quality and latency.In real interpreting situations, interpreters manage this simultaneity by breaking sentences into smaller segments and translating them while maintaining the source order as much as possible.SiST could benefit from this approach to balance quality and latency.However, current corpora used for simultaneous tasks often involve significant word reordering in translation, which is not ideal given that interpreters faithfully follow source syntax as much as possible.Inspired by conference interpreting by humans utilizing the salami technique, we introduce the Simul-MuST-C 1 , a dataset created by leveraging the Large Language Model (LLM), specifically GPT-4o, which aligns the target text as closely as possible to the source text by using minimal chunks that contain enough information to be interpreted.Experiments on three language pairs show that the effectiveness of segmentedbase monotonicity in training data varies with the grammatical distance between the source and the target, with grammatically distant language pairs benefiting the most in achieving quality while minimizing latency.
Mana Makinae, Yusuke Sakai 0010, Hidetaka Kamigaito, Taro Watanabe
EMNLP2
2024 Does Pre-trained Language Model Actually Infer Unseen Links in Knowledge Graph Completion?
abstract
Yusuke Sakai, Hidetaka Kamigaito, Katsuhiko Hayashi, Taro Watanabe. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yusuke Sakai 0010, Hidetaka Kamigaito, Katsuhiko Hayashi 0001, Taro Watanabe
NAACL-HLT1
2023 Universal Automatic Phonetic Transcription into the International Phonetic Alphabet
Chihiro Taguchi, Yusuke Sakai 0010, Parisa Haghani, David Chiang 0001
INTERSPEECH2