VLDB 2026 Research / reviewers in the wild / expert
Ukyo Honda
dblp:220/2007
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-4894-9886ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 79% Generative modeling · 16% Trustworthy machine learning · 5% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | Exploring Explanations Improves the Robustness of In-Context Learning · ACL (1) 2025 |
Natural language and speech › Language models and text generation
multimodal language model |
0.9 | 1 | 2025 | Multimodal Markup Document Models for Graphic Design Completion · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation › in-context learning
robustness of in-context learning |
0.9 | 1 | 2025 | Exploring Explanations Improves the Robustness of In-Context Learning · ACL (1) 2025 |
Natural language and speech › Language models and text generation
decoding |
0.8 | 1 | 2024 | Model-Based Minimum Bayes Risk Decoding for Text Generation · ICML 2024 |
Natural language and speech › Language models and text generation › decoding
minimum bayes risk decoding |
0.8 | 1 | 2024 | Model-Based Minimum Bayes Risk Decoding for Text Generation · ICML 2024 |
Machine learning › Trustworthy machine learning › interpretability
explanation-based learning |
0.3 | 1 | 2025 | Exploring Explanations Improves the Robustness of In-Context Learning · ACL (1) 2025 |
Visual content generation and editing › graphic design
graphic design generation |
0.3 | 1 | 2025 | Multimodal Markup Document Models for Graphic Design Completion · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation
text generation |
0.2 | 1 | 2024 | Model-Based Minimum Bayes Risk Decoding for Text Generation · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
instruction tuning · 1.7fill-in-the-middle training · 1.7discrete image tokenization · 1.7monte carlo estimation · 0.8beam search · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring Explanations Improves the Robustness of In-Context LearningabstractIn-context learning (ICL) has emerged as a successful paradigm for leveraging large language models (LLMs). However, it often struggles to generalize beyond the distribution of the provided demonstrations. A recent advancement in enhancing robustness is ICL with explanations (X-ICL), which improves prediction reliability by guiding LLMs to understand and articulate the reasoning behind correct labels. Building on this approach, we introduce an advanced framework that extends X-ICL by systematically exploring explanations for all possible labels (X^2-ICL), thereby enabling more comprehensive and robust decision-making. Experimental results on multiple natural language understanding datasets validate the effectiveness of X^2-ICL, demonstrating significantly improved robustness to out-of-distribution data compared to the existing ICL approaches. Ukyo Honda, Tatsushi Oka |
ACL (1) | 1 |
| 2025 | Multimodal Markup Document Models for Graphic Design CompletionabstractWe introduce MarkupDM, a multimodal markup document model that represents graphic design as an interleaved multimodal document consisting of both markup language and images. Unlike existing holistic approaches that rely on an element-by-attribute grid representation, our representation accommodates variable-length elements, type-dependent attributes, and text content. Inspired by fill-in-the-middle training in code generation, we train the model to complete the missing part of a design document from its surrounding context, allowing it to treat various design tasks in a unified manner. Our model also supports image generation by predicting discrete image tokens through a specialized tokenizer with support for image transparency. We evaluate MarkupDM on three tasks, attribute value, image, and text completion, and demonstrate that it can produce plausible designs consistent with the given context. To further illustrate the flexibility of our approach, we evaluate our approach on a new instruction-guided design completion task where our instruction-tuned MarkupDM compares favorably to state-of-the-art image editing models, especially in textual completion. These findings suggest that multimodal language models with our document representation can serve as a versatile foundation for broad design automation. Kotaro Kikuchi, Ukyo Honda, Naoto Inoue, Mayu Otani, Edgar Simo-Serra, Kota Yamaguchi |
ACM Multimedia | 2 |
| 2024 | CAMERA³: An Evaluation Dataset for Controllable Ad Text Generation in JapaneseabstractAd text generation is the task of creating compelling text from an advertising asset that describes products or services, such as a landing page. In advertising, diversity plays an important role in enhancing the effectiveness of an ad text, mitigating a phenomenon called “ad fatigue,” where users become disengaged due to repetitive exposure to the same advertisement. Despite numerous efforts in ad text generation, the aspect of diversifying ad texts has received limited attention, particularly in non-English languages like Japanese. To address this, we present CAMERA³, an evaluation dataset for controllable text generation in the advertising domain in Japanese. Our dataset includes 3,980 ad texts written by expert annotators, taking into account various aspects of ad appeals. We make CAMERA³ publicly available, allowing researchers to examine the capabilities of recent NLG models in controllable text generation in a real-world scenario. Go Inoue, Akihiko Kato, Masato Mita, Ukyo Honda, Peinan Zhang |
LREC/COLING | 4 |
| 2024 | A Single Linear Layer Yields Task-Adapted Low-Rank MatricesabstractLow-Rank Adaptation (LoRA) is a widely used Parameter-Efficient Fine-Tuning (PEFT) method that updates an initial weight matrix W_0 with a delta matrix \Delta W consisted by two low-rank matrices A and B. A previous study suggested that there is correlation between W_0 and \Delta W. In this study, we aim to delve deeper into relationships between W_0 and low-rank matrices A and B to further comprehend the behavior of LoRA. In particular, we analyze a conversion matrix that transform W_0 into low-rank matrices, which encapsulates information about the relationships. Our analysis reveals that the conversion matrices are similar across each layer. Inspired by these findings, we hypothesize that a single linear layer, which takes each layer’s W_0 as input, can yield task-adapted low-rank matrices. To confirm this hypothesis, we devise a method named Conditionally Parameterized LoRA (CondLoRA) that updates initial weight matrices with low-rank matrices derived from a single linear layer. Our empirical results show that CondLoRA maintains a performance on par with LoRA, despite the fact that the trainable parameters of CondLoRA are fewer than those of LoRA. Therefore, we conclude that “a single linear layer yields task-adapted low-rank matrices.” The code used in our experiments is available at https://github.com/CyberAgentAILab/CondLoRA. Hwichan Kim, Shota Sasaki, Sho Hoshino, Ukyo Honda |
LREC/COLING | 4 |
| 2024 | Model-Based Minimum Bayes Risk Decoding for Text GenerationabstractMinimum Bayes Risk (MBR) decoding has been shown to be a powerful alternative to beam search decoding in a variety of text generation tasks. MBR decoding selects a hypothesis from a pool of hypotheses that has the least expected risk under a probability model according to a given utility function. Since it is impractical to compute the expected risk exactly over all possible hypotheses, two approximations are commonly used in MBR. First, it integrates over a sampled set of hypotheses rather than over all possible hypotheses. Second, it estimates the probability of each hypothesis using a Monte Carlo estimator. While the first approximation is necessary to make it computationally feasible, the second is not essential since we typically have access to the model probability at inference time. We propose model-based MBR (MBMBR), a variant of MBR that uses the model probability itself as the estimate of the probability distribution instead of the Monte Carlo estimate. We show analytically and empirically that the model-based estimate is more promising than the Monte Carlo estimate in text generation tasks. Our experiments show that MBMBR outperforms MBR in several text generation tasks, both with encoder-decoder models and with language models. Yuu Jinnai, Tetsuro Morimura, Ukyo Honda, Kaito Ariu, Kenshi Abe |
ICML | 3 |
| 2024 | Not Eliminate but Aggregate: Post-Hoc Control over Mixture-of-Experts to Address Shortcut Shifts in Natural Language UnderstandingabstractAbstract Recent models for natural language understanding are inclined to exploit simple patterns in datasets, commonly known as shortcuts. These shortcuts hinge on spurious correlations between labels and latent features existing in the training data. At inference time, shortcut-dependent models are likely to generate erroneous predictions under distribution shifts, particularly when some latent features are no longer correlated with the labels. To avoid this, previous studies have trained models to eliminate the reliance on shortcuts. In this study, we explore a different direction: pessimistically aggregating the predictions of a mixture-of-experts, assuming each expert captures relatively different latent features. The experimental results demonstrate that our post-hoc control over the experts significantly enhances the model’s robustness to the distribution shift in shortcuts. Additionally, we show that our approach has some practical advantages. We also analyze our model and provide results to support the assumption.1 Ukyo Honda, Tatsushi Oka, Peinan Zhang, Masato Mita |
Trans. Assoc. Comput. Linguistics | 1 |
| 2023 | Switching to Discriminative Image Captioning by Relieving a Bottleneck of Reinforcement LearningabstractDiscriminativeness is a desirable feature of image captions: captions should describe the characteristic details of input images. However, recent high-performing captioning models, which are trained with reinforcement learning (RL), tend to generate overly generic captions despite their high performance in various other criteria. First, we investigate the cause of the unexpectedly low discriminativeness and show that RL has a deeply rooted side effect of limiting the output words to high-frequency words. The limited vocabulary is a severe bottleneck for discriminativeness as it is difficult for a model to describe the details beyond its vocabulary. Then, based on this identification of the bottleneck, we drastically recast discriminative image captioning as a much simpler task of encouraging low-frequency word generation. Hinted by long-tail classification and debiasing methods, we propose methods that easily switch off-the-shelf RL models to discriminativeness-aware models with only a single-epoch fine-tuning on the part of the parameters. Extensive experiments demonstrate that our methods significantly enhance the discriminative-ness of off-the-shelf RL models and even outperform previous discriminativeness-aware methods with much smaller computational costs. Detailed analysis and human evaluation also verify that our methods boost the discriminativeness without sacrificing the overall quality of captions.1 Ukyo Honda, Taro Watanabe, Yuji Matsumoto 0001 |
WACV | 1 |
| 2022 | Law Retrieval with Supervised Contrastive Learning Using the Hierarchical Structure of Law
Jungmin Choi, Ukyo Honda, Taro Watanabe, Hiroki Ouchi, Kentaro Inui |
PACLIC | 2 |
| 2022 | Characterization of Pulmonary Nodules in Computed Tomography Images Based on Pseudo-Labeling Using Radiology ReportsabstractA computer-aided diagnosis (CAD) system that characterizes nodules in medical images can help radiologists determine its malignancy. Preparing large volumes of labeled data for CAD systems, however, requires advanced medical knowledge. This makes it extremely difficult to develop such systems, despite their growing demand. In this paper, we propose a new training method to build an image classifier for characterization of nodules utilizing pseudo-labels, i.e., image labels automatically retrieved from radiology reports. A radiology report is a type of record in which radiologists present a summary of lesion characteristics and diagnosis. Labeling radiology reports is much easier than labeling radiology images, and can be done without high expertise. Using several thousand labeled reports, we constructed a hierarchical attention network-based text classifier to assign pseudo-labels of the characteristics of pulmonary nodules with high accuracy (macro F1-score of 0.941). Experimental results show that the image classifier trained with the pseudo-labels can achieve almost the same performance as the one trained with the labels annotated by radiologists: AUC 0.848 for the model trained with the pseudo-labels on 3,000 computed tomography (CT) images and 0.847 for the model trained with the manual labels on 800 CT images. Yohei Momoki, Akimichi Ichinose, Yutaro Shigeto, Ukyo Honda, Keigo Nakamura, Yuji Matsumoto 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Removing Word-Level Spurious Alignment between Images and Pseudo-Captions in Unsupervised Image CaptioningabstractUkyo Honda, Yoshitaka Ushiku, Atsushi Hashimoto, Taro Watanabe, Yuji Matsumoto. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Ukyo Honda, Yoshitaka Ushiku, Atsushi Hashimoto 0001, Taro Watanabe, Yuji Matsumoto 0001 |
EACL | 1 |