VLDB 2026 Research / reviewers in the wild / expert
Hiroyuki Deguchi 0002
dblp:17/6058-2
· DBLP profile ↗
7ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0003-2127-6607ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 69% Vision and language · 16% Machine translation · 15% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 6 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
decoding |
1.7 | 2 | 2025 | Case-Based Decision-Theoretic Decoding with Quality Memories · EMNLP 2025 Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding · ACL (1) 2025 |
Natural language and speech › Language models and text generation › decoding
minimum bayes risk decoding |
1.7 | 2 | 2025 | Case-Based Decision-Theoretic Decoding with Quality Memories · EMNLP 2025 Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model inference
inference scaling laws |
0.9 | 1 | 2025 | Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding · ACL (1) 2025 |
Information retrieval
pattern matching |
0.9 | 1 | 2025 | SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches · ICLR 2025 |
Natural language and speech › Machine translation › neural machine translation
nearest neighbor machine translation |
0.7 | 1 | 2023 | Subset Retrieval Nearest Neighbor Machine Translation · ACL (1) 2023 |
Information retrieval › indexing
inverted index |
0.3 | 1 | 2025 | SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
word embeddings · 0.9diversity analysis · 0.9case-based decision theory · 0.9nearest neighbor retrieval · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via HubnessabstractThe hubness problem, in which hub embeddings are close to many unrelated examples, occurs often in high-dimensional embedding spaces and may pose a practical threat for purposes such as information retrieval and automatic evaluation metrics.In particular, since cross-modal similarity between text and images cannot be calculated by direct comparisons, such as string matching, cross-modal encoders that project different modalities into a shared space are helpful for various cross-modal applications, and thus, the existence of hubs may pose practical threats.To reveal the vulnerabilities of cross-modal encoders, we propose a method for identifying the hub embedding and its corresponding hub text.Experiments on image captioning evaluation in MSCOCO and nocaps along with image-to-text retrieval tasks in MSCOCO and Flickr30k showed that our method can identify a single hub text that unreasonably achieves comparable or higher similarity scores than human-written reference captions in many images, thereby revealing the vulnerabilities in cross-modal encoders. Hiroyuki Deguchi 0002, Katsuki Chousa, Yusuke Sakai 0010 |
ACL (1) | 1 |
| 2025 | Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk DecodingabstractHidetaka Kamigaito, Hiroyuki Deguchi, Yusuke Sakai, Katsuhiko Hayashi, Taro Watanabe. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Hidetaka Kamigaito, Hiroyuki Deguchi 0002, Yusuke Sakai 0010, Katsuhiko Hayashi 0001, Taro Watanabe |
ACL (1) | 2 |
| 2025 | Case-Based Decision-Theoretic Decoding with Quality MemoriesabstractMinimum Bayes risk (MBR) decoding is a decision rule of text generation, which selects the hypothesis that maximizes the expected utility and robustly generates higher-quality texts than maximum a posteriori (MAP) decoding.However, it depends on sample texts drawn from the text generation model; thus, it is difficult to find a hypothesis that correctly captures the knowledge or information of out-of-domain.To tackle this issue, we propose case-based decision-theoretic (CBDT) decoding, another method to estimate the expected utility using examples of domain data.CBDT decoding not only generates higher-quality texts than MAP decoding, but also the combination of MBR and CBDT decoding outperformed MBR decoding in seven domain De-En and Ja↔En translation tasks and image captioning tasks on MSCOCO and nocaps datasets. Hiroyuki Deguchi 0002, Masaaki Nagata |
EMNLP | 1 |
| 2025 | SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus SearchesabstractResearchers and practitioners in natural language processing and computational linguistics frequently observe and analyze the real language usage in large-scale corpora.
For that purpose, they often employ off-the-shelf pattern-matching tools, such as grep, and keyword-in-context concordancers, which is widely used in corpus linguistics for gathering examples.
Nonetheless, these existing techniques rely on surface-level string matching, and thus they suffer from the major limitation of not being able to handle orthographic variations and paraphrasing---notable and common phenomena in any natural language.
In addition, existing continuous approaches such as dense vector search tend to be overly coarse, often retrieving texts that are unrelated but share similar topics.
Given these challenges, we propose a novel algorithm that achieves soft (or semantic) yet efficient pattern matching by relaxing a surface-level matching with word embeddings.
Our algorithm is highly scalable with respect to the size of the corpus text utilizing inverted indexes.
We have prepared an efficient implementation, and we provide an accessible web tool.
Our experiments demonstrate that the proposed method
(i) can execute searches on billion-scale corpora in less than a second, which is comparable in speed to surface-level string matching and dense vector search;
(ii) can extract harmful instances that semantically match queries from a large set of English and Japanese Wikipedia articles;
and (iii) can be effectively applied to corpus-linguistic analyses of Latin, a language with highly diverse inflections. Hiroyuki Deguchi 0002, Go Kamoda, Yusuke Matsushita 0002, Chihiro Taguchi, Kohei Suenaga, Masaki Waga, Sho Yokoi |
ICLR | 1 |
| 2024 | Detector-Corrector: Edit-Based Automatic Post Editing for Human Post EditingabstractPost-editing is crucial in the real world because neural machine translation (NMT) sometimes makes errors.Automatic post-editing (APE) attempts to correct the outputs of an MT model for better translation quality.However, many APE models are based on sequence generation, and thus their decisions are harder to interpret for actual users.In this paper, we propose “detector–corrector”, an edit-based post-editing model, which breaks the editing process into two steps, error detection and error correction.The detector model tags each MT output token whether it should be corrected and/or reordered while the corrector model generates corrected words for the spans identified as errors by the detector.Experiments on the WMT’20 English–German and English–Chinese APE tasks showed that our detector–corrector improved the translation edit rate (TER) compared to the previous edit-based model and a black-box sequence-to-sequence APE model, in addition, our model is more explainable because it is based on edit operations. Hiroyuki Deguchi 0002, Masaaki Nagata, Taro Watanabe |
EAMT (1) | 1 |
| 2023 | Subset Retrieval Nearest Neighbor Machine TranslationabstractHiroyuki Deguchi, Taro Watanabe, Yusuke Matsui, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hiroyuki Deguchi 0002, Taro Watanabe, Yusuke Matsui 0001, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita |
ACL (1) | 1 |
| 2020 | Bilingual Subword Segmentation for Neural Machine TranslationabstractThis paper proposed a new subword segmentation method for neural machine translation, "Bilingual Subword Segmentation," which tokenizes sentences to minimize the difference between the number of subword units in a sentence and that of its translation.While existing subword segmentation methods tokenize a sentence without considering its translation, the proposed method tokenizes a sentence by using subword units induced from bilingual sentences; this method could be more favorable to machine translation.Evaluations on WAT Asian Scientific Paper Excerpt Corpus (ASPEC) English-to-Japanese and Japanese-to-English translation tasks and WMT14 English-to-German and German-to-English translation tasks show that our bilingual subword segmentation improves the performance of Transformer neural machine translation (up to +0.81 BLEU). Hiroyuki Deguchi 0002, Masao Utiyama, Akihiro Tamura, Takashi Ninomiya, Eiichiro Sumita |
COLING | 1 |