Chihiro Taguchi

dblp:332/6012 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Speech recognition and synthesis · 43% Language models and text generation · 32% Machine translation · 25%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation
0.912025
Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark · EMNLP 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-k · EMNLP 2025
Information retrieval › document retrieval
passage retrieval
0.912025
Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-k · EMNLP 2025
Information retrieval
pattern matching
0.912025
SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches · ICLR 2025
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.812024
Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't · ACL (1) 2024
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
multilingual speech recognition
0.812024
Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't · ACL (1) 2024
Natural language and speech › Language models and text generation › large language model evaluation
NLP evaluation
0.312025
Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark · EMNLP 2025
Information retrieval › indexing
inverted index
0.312025
SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches · ICLR 2025

Methods — techniques the papers use, named apart from their topics

similarity score distribution · 1.7word embeddings · 0.9human assessment · 0.9BLEU · 0.9fine-tuning · 0.8Wav2Vec2-XLSR-53 · 0.8
YearPublicationVenuePosition
2026 Automatic Speech Recognition for Documenting Endangered Languages: Case Study of Ikema Miyakoan
Chihiro Taguchi, Yukinori Takubo, David Chiang 0001
LREC1
2025 Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-k
abstract
Retrieval-augmented generation (RAG) and long-context language models (LCLMs) both address context limitations of LLMs in opendomain question answering (QA).However, optimal external context to retrieve remains an open problem: fixing the retrieval size risks either wasting tokens or omitting key evidence.Existing adaptive methods like Self-RAG and SELF-ROUTE rely on iterative LLM prompting and perform well on factoid QA, but struggle with aggregation QA, where the optimal context size is both unknown and variable.We present Adaptive-k retrieval, a simple and effective single-pass method that adaptively selects the number of passages based on the distribution of the similarity scores between the query and the candidate passages.It does not require model fine-tuning, extra LLM inferences or changes to existing retriever-reader pipelines.On both factoid and aggregation QA benchmarks, Adaptive-k matches or outperforms fixed-k baselines while using up to 10× fewer tokens than full-context input, yet still retrieves 70% of relevant passages.It improves accuracy across five LCLMs and two embedding models, highlighting that dynamically adjusting context size leads to more efficient and accurate QA. 1
Chihiro Taguchi, Seiji Maekawa, Nikita Bhutani
EMNLP1
2025 Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark
abstract
Multilingual machine translation (MT) benchmarks play a central role in evaluating the capabilities of modern MT systems.Among them, the FLORES+ benchmark is widely used, offering English-to-many translation data for over 200 languages, curated with strict quality control protocols.However, we study data in four languages (Asante Twi, Japanese, Jinghpaw, and South Azerbaijani) and uncover critical shortcomings in the benchmark's suitability for truly multilingual evaluation.Human assessments reveal that many translations fall below the claimed 90% quality standard, and the annotators report that source sentences are often too domain-specific and culturally biased toward the English-speaking world.We further demonstrate that simple heuristics, such as copying named entities, can yield non-trivial BLEU scores, suggesting vulnerabilities in the evaluation protocol.Notably, we show that MT models trained on high-quality, naturalistic data perform poorly on FLORES+ while achieving significant gains on our domain-relevant evaluation set.Based on these findings, we advocate for multilingual MT benchmarks that use domain-general and culturally neutral source texts rely less on named entities, in order to better reflect real-world translation challenges. 1 * Equal contribution.
Chihiro Taguchi, Seng Mai, Keita Kurabe, Yusuke Sakai 0010, Georgina Agyei, Soudabeh Eslami, David Chiang 0001
EMNLP1
2025 SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches
abstract
Researchers and practitioners in natural language processing and computational linguistics frequently observe and analyze the real language usage in large-scale corpora. For that purpose, they often employ off-the-shelf pattern-matching tools, such as grep, and keyword-in-context concordancers, which is widely used in corpus linguistics for gathering examples. Nonetheless, these existing techniques rely on surface-level string matching, and thus they suffer from the major limitation of not being able to handle orthographic variations and paraphrasing---notable and common phenomena in any natural language. In addition, existing continuous approaches such as dense vector search tend to be overly coarse, often retrieving texts that are unrelated but share similar topics. Given these challenges, we propose a novel algorithm that achieves soft (or semantic) yet efficient pattern matching by relaxing a surface-level matching with word embeddings. Our algorithm is highly scalable with respect to the size of the corpus text utilizing inverted indexes. We have prepared an efficient implementation, and we provide an accessible web tool. Our experiments demonstrate that the proposed method (i) can execute searches on billion-scale corpora in less than a second, which is comparable in speed to surface-level string matching and dense vector search; (ii) can extract harmful instances that semantically match queries from a large set of English and Japanese Wikipedia articles; and (iii) can be effectively applied to corpus-linguistic analyses of Latin, a language with highly diverse inflections.
Hiroyuki Deguchi 0002, Go Kamoda, Yusuke Matsushita 0002, Chihiro Taguchi, Kohei Suenaga, Masaki Waga, Sho Yokoi
ICLR4
2024 Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't
abstract
We investigate what linguistic factors affect the performance of Automatic Speech Recognition (ASR) models.We hypothesize that orthographic and phonological complexities both degrade accuracy.To examine this, we finetune the multilingual self-supervised pretrained model Wav2Vec2-XLSR-53 on 25 languages with 15 writing systems, and we compare their ASR accuracy, number of graphemes, unigram grapheme entropy, logographicity (how much word/morpheme-level information is encoded in the writing system), and number of phonemes.The results demonstrate that a high logographicity correlates with low ASR accuracy, while phonological complexity has no strong correlation.
Chihiro Taguchi, David Chiang 0001
ACL (1)1
2024 J-SNACS: Adposition and Case Supersenses for Japanese Joshi
abstract
Many languages use adpositions (prepositions or postpositions) to mark a variety of semantic relations, with different languages exhibiting both commonalities and idiosyncrasies in the relations grouped under the same lexeme. We present the first Japanese extension of the SNACS framework (Schneider et al., 2018), which has served as the basis for annotating adpositions in corpora from several languages. After establishing which of the set of particles (joshi) in Japanese qualify as case markers and adpositions as defined in SNACS, we annotate 10 chapters (≈10k tokens) of the Japanese translation of Le Petit Prince (The Little Prince), achieving high inter-annotator agreement. We find that, while a majority of the particles and their uses are captured by the existing and extended SNACS annotation guidelines from the previous work, some unique cases were observed. We also conduct experiments investigating the cross-lingual similarity of adposition and case marker supersenses, showing that the language-agnostic SNACS framework captures similarities not clearly observed in multilingual embedding space.
Tatsuya Aoyama, Chihiro Taguchi, Nathan Schneider 0001
LREC/COLING2
2024 Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information
abstract
This paper presents Killkan, the first dataset for automatic speech recognition (ASR) in the Kichwa language, an indigenous language of Ecuador. Kichwa is an extremely low-resource endangered language, and there have been no resources before Killkan for Kichwa to be incorporated in applications of natural language processing. The dataset contains approximately 4 hours of audio with transcription, translation into Spanish, and morphosyntactic annotation in the format of Universal Dependencies, all done in ELAN, the annotation software. The audio data was retrieved from a publicly available radio program in Kichwa. This paper also provides corpus-linguistic analyses of the dataset with a special focus on the agglutinative morphology of Kichwa and frequent code-switching with Spanish. The experiments show that the dataset makes it possible to develop the first ASR system for Kichwa with reliable quality despite its small dataset size. This dataset, the ASR model, and the code used to develop them will be publicly available. Thus, our study positively showcases resource building and its applications for low-resource languages and their community.
Chihiro Taguchi, Jefferson Saransig, Dayana Velásquez, David Chiang 0001
LREC/COLING1
2023 Universal Automatic Phonetic Transcription into the International Phonetic Alphabet
Chihiro Taguchi, Yusuke Sakai 0010, Parisa Haghani, David Chiang 0001
INTERSPEECH1