Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hiroyuki Deguchi 0002

dblp:17/6058-2 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0003-2127-6607ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 69% Vision and language · 16% Machine translation · 15%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 6 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
decoding
1.722025
Case-Based Decision-Theoretic Decoding with Quality Memories · EMNLP 2025
Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding · ACL (1) 2025
Natural language and speech › Language models and text generation › decoding
minimum bayes risk decoding
1.722025
Case-Based Decision-Theoretic Decoding with Quality Memories · EMNLP 2025
Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model inference
inference scaling laws
0.912025
Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding · ACL (1) 2025
Information retrieval
pattern matching
0.912025
SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches · ICLR 2025
Natural language and speech › Machine translation › neural machine translation
nearest neighbor machine translation
0.712023
Subset Retrieval Nearest Neighbor Machine Translation · ACL (1) 2023
Information retrieval › indexing
inverted index
0.312025
SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches · ICLR 2025

Methods — techniques the papers use, named apart from their topics

word embeddings · 0.9diversity analysis · 0.9case-based decision theory · 0.9nearest neighbor retrieval · 0.7
YearPublicationVenuePosition
2026 One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness
abstract
The hubness problem, in which hub embeddings are close to many unrelated examples, occurs often in high-dimensional embedding spaces and may pose a practical threat for purposes such as information retrieval and automatic evaluation metrics.In particular, since cross-modal similarity between text and images cannot be calculated by direct comparisons, such as string matching, cross-modal encoders that project different modalities into a shared space are helpful for various cross-modal applications, and thus, the existence of hubs may pose practical threats.To reveal the vulnerabilities of cross-modal encoders, we propose a method for identifying the hub embedding and its corresponding hub text.Experiments on image captioning evaluation in MSCOCO and nocaps along with image-to-text retrieval tasks in MSCOCO and Flickr30k showed that our method can identify a single hub text that unreasonably achieves comparable or higher similarity scores than human-written reference captions in many images, thereby revealing the vulnerabilities in cross-modal encoders.
Hiroyuki Deguchi 0002, Katsuki Chousa, Yusuke Sakai 0010
ACL (1)1
2025 Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding
abstract
Hidetaka Kamigaito, Hiroyuki Deguchi, Yusuke Sakai, Katsuhiko Hayashi, Taro Watanabe. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hidetaka Kamigaito, Hiroyuki Deguchi 0002, Yusuke Sakai 0010, Katsuhiko Hayashi 0001, Taro Watanabe
ACL (1)2
2025 Case-Based Decision-Theoretic Decoding with Quality Memories
abstract
Minimum Bayes risk (MBR) decoding is a decision rule of text generation, which selects the hypothesis that maximizes the expected utility and robustly generates higher-quality texts than maximum a posteriori (MAP) decoding.However, it depends on sample texts drawn from the text generation model; thus, it is difficult to find a hypothesis that correctly captures the knowledge or information of out-of-domain.To tackle this issue, we propose case-based decision-theoretic (CBDT) decoding, another method to estimate the expected utility using examples of domain data.CBDT decoding not only generates higher-quality texts than MAP decoding, but also the combination of MBR and CBDT decoding outperformed MBR decoding in seven domain De-En and Ja↔En translation tasks and image captioning tasks on MSCOCO and nocaps datasets.
Hiroyuki Deguchi 0002, Masaaki Nagata
EMNLP1
2025 SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches
abstract
Researchers and practitioners in natural language processing and computational linguistics frequently observe and analyze the real language usage in large-scale corpora. For that purpose, they often employ off-the-shelf pattern-matching tools, such as grep, and keyword-in-context concordancers, which is widely used in corpus linguistics for gathering examples. Nonetheless, these existing techniques rely on surface-level string matching, and thus they suffer from the major limitation of not being able to handle orthographic variations and paraphrasing---notable and common phenomena in any natural language. In addition, existing continuous approaches such as dense vector search tend to be overly coarse, often retrieving texts that are unrelated but share similar topics. Given these challenges, we propose a novel algorithm that achieves soft (or semantic) yet efficient pattern matching by relaxing a surface-level matching with word embeddings. Our algorithm is highly scalable with respect to the size of the corpus text utilizing inverted indexes. We have prepared an efficient implementation, and we provide an accessible web tool. Our experiments demonstrate that the proposed method (i) can execute searches on billion-scale corpora in less than a second, which is comparable in speed to surface-level string matching and dense vector search; (ii) can extract harmful instances that semantically match queries from a large set of English and Japanese Wikipedia articles; and (iii) can be effectively applied to corpus-linguistic analyses of Latin, a language with highly diverse inflections.
Hiroyuki Deguchi 0002, Go Kamoda, Yusuke Matsushita 0002, Chihiro Taguchi, Kohei Suenaga, Masaki Waga, Sho Yokoi
ICLR1
2024 Detector-Corrector: Edit-Based Automatic Post Editing for Human Post Editing
abstract
Post-editing is crucial in the real world because neural machine translation (NMT) sometimes makes errors.Automatic post-editing (APE) attempts to correct the outputs of an MT model for better translation quality.However, many APE models are based on sequence generation, and thus their decisions are harder to interpret for actual users.In this paper, we propose “detector–corrector”, an edit-based post-editing model, which breaks the editing process into two steps, error detection and error correction.The detector model tags each MT output token whether it should be corrected and/or reordered while the corrector model generates corrected words for the spans identified as errors by the detector.Experiments on the WMT’20 English–German and English–Chinese APE tasks showed that our detector–corrector improved the translation edit rate (TER) compared to the previous edit-based model and a black-box sequence-to-sequence APE model, in addition, our model is more explainable because it is based on edit operations.
Hiroyuki Deguchi 0002, Masaaki Nagata, Taro Watanabe
EAMT (1)1
2023 Subset Retrieval Nearest Neighbor Machine Translation
abstract
Hiroyuki Deguchi, Taro Watanabe, Yusuke Matsui, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hiroyuki Deguchi 0002, Taro Watanabe, Yusuke Matsui 0001, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita
ACL (1)1
2020 Bilingual Subword Segmentation for Neural Machine Translation
abstract
This paper proposed a new subword segmentation method for neural machine translation, "Bilingual Subword Segmentation," which tokenizes sentences to minimize the difference between the number of subword units in a sentence and that of its translation.While existing subword segmentation methods tokenize a sentence without considering its translation, the proposed method tokenizes a sentence by using subword units induced from bilingual sentences; this method could be more favorable to machine translation.Evaluations on WAT Asian Scientific Paper Excerpt Corpus (ASPEC) English-to-Japanese and Japanese-to-English translation tasks and WMT14 English-to-German and German-to-English translation tasks show that our bilingual subword segmentation improves the performance of Transformer neural machine translation (up to +0.81 BLEU).
Hiroyuki Deguchi 0002, Masao Utiyama, Akihiro Tamura, Takashi Ninomiya, Eiichiro Sumita
COLING1