VLDB 2026 Research / reviewers in the wild / expert
Yuu Jinnai
dblp:178/8539
· DBLP profile ↗
16ranked-venue papers
10as first author
6since 2021 · last 2025
0000-0002-8756-644XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 10 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
14 papers |
Language models and text generation · 40% Reinforcement learning · 36% Planning, search and constraint satisfaction · 16% | |
| Theoretical computer science
2 papers |
Computational complexity · 44% Approximation and online algorithms · 44% Coding theory · 13% |
Topics — the 27 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
decoding |
2.5 | 3 | 2025 | Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport · ACL (1) 2025 Theoretical Guarantees for Minimum Bayes Risk Decoding · ACL (1) 2025 Model-Based Minimum Bayes Risk Decoding for Text Generation · ICML 2024 |
Natural language and speech › Language models and text generation › decoding
minimum bayes risk decoding |
2.5 | 3 | 2025 | Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport · ACL (1) 2025 Theoretical Guarantees for Minimum Bayes Risk Decoding · ACL (1) 2025 Model-Based Minimum Bayes Risk Decoding for Text Generation · ICML 2024 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery |
1.2 | 3 | 2020 | Exploration in Reinforcement Learning with Deep Covering Options · ICLR 2020 Discovering Options for Exploration by Minimizing Cover Time · ICML 2019 Finding Options that Minimize Planning Time · ICML 2019 |
Natural language and speech › Machine translation
document-level machine translation |
0.9 | 1 | 2025 | Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport · ACL (1) 2025 |
Machine learning › Reinforcement learning › non-stationary reinforcement learning
continual reinforcement learning |
0.8 | 2 | 2021 | Lipschitz Lifelong Reinforcement Learning · AAAI 2021 Policy and Value Transfer in Lifelong Reinforcement Learning · ICML 2018 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search |
0.8 | 3 | 2017 | Learning to Avoid Dominated Action Sequences in Planning for Black-Box Domains · AAAI 2017 Learning to Prune Dominated Action Sequences in Online Black-Box Planning · AAAI 2017 Abstract Zobrist Hashing: An Efficient Work Distribution Method for Parallel Best-First Search · AAAI 2016 |
Machine learning › Reinforcement learning
exploration |
0.8 | 2 | 2020 | Exploration in Reinforcement Learning with Deep Covering Options · ICLR 2020 Discovering Options for Exploration by Minimizing Cover Time · ICML 2019 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Filtered Direct Preference Optimization · EMNLP 2024 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.8 | 1 | 2024 | Filtered Direct Preference Optimization · EMNLP 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
black-box planning |
0.6 | 2 | 2017 | Learning to Avoid Dominated Action Sequences in Planning for Black-Box Domains · AAAI 2017 Learning to Prune Dominated Action Sequences in Online Black-Box Planning · AAAI 2017 |
Machine learning › Reinforcement learning
transfer learning in reinforcement learning |
0.5 | 1 | 2021 | Lipschitz Lifelong Reinforcement Learning · AAAI 2021 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.4 | 1 | 2020 | Exploration in Reinforcement Learning with Deep Covering Options · ICLR 2020 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search |
0.4 | 1 | 2020 | Neural Architecture Search Using Deep Neural Networks and Monte Carlo Tree Search · AAAI 2020 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.4 | 1 | 2020 | Neural Architecture Search Using Deep Neural Networks and Monte Carlo Tree Search · AAAI 2020 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
hierarchical planning |
0.4 | 1 | 2019 | Finding Options that Minimize Planning Time · ICML 2019 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.4 | 1 | 2019 | State Abstraction as Compression in Apprenticeship Learning · AAAI 2019 |
Machine learning › Reinforcement learning
sparse reward |
0.4 | 1 | 2019 | Discovering Options for Exploration by Minimizing Cover Time · ICML 2019 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state abstraction |
0.4 | 1 | 2019 | State Abstraction as Compression in Apprenticeship Learning · AAAI 2019 |
Approximation and online algorithms
approximation algorithms |
0.4 | 1 | 2019 | Finding Options that Minimize Planning Time · ICML 2019 |
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
policy transfer |
0.3 | 1 | 2018 | Policy and Value Transfer in Lifelong Reinforcement Learning · ICML 2018 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search
best-first search |
0.2 | 1 | 2016 | Abstract Zobrist Hashing: An Efficient Work Distribution Method for Parallel Best-First Search · AAAI 2016 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search › search-based problem solving
parallel search |
0.2 | 1 | 2016 | Abstract Zobrist Hashing: An Efficient Work Distribution Method for Parallel Best-First Search · AAAI 2016 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.2 | 1 | 2024 | Filtered Direct Preference Optimization · EMNLP 2024 |
Natural language and speech › Language models and text generation
text generation |
0.2 | 1 | 2024 | Model-Based Minimum Bayes Risk Decoding for Text Generation · ICML 2024 |
Machine learning › Reinforcement learning › sample efficiency
PAC-MDP |
0.1 | 1 | 2021 | Lipschitz Lifelong Reinforcement Learning · AAAI 2021 |
Coding theory › source coding
rate-distortion theory |
0.1 | 1 | 2019 | State Abstraction as Compression in Apprenticeship Learning · AAAI 2019 |
Machine learning › Reinforcement learning › deep reinforcement learning
atari game playing |
0.1 | 1 | 2017 | Learning to Prune Dominated Action Sequences in Online Black-Box Planning · AAAI 2017 |
Methods — techniques the papers use, named apart from their topics
wasserstein distance · 0.9statistical learning theory · 0.9optimal transport · 0.9concentration bounds · 0.9reward model · 0.8monte carlo estimation · 0.8filtered direct preference optimization · 0.8beam search · 0.8dominance pruning · 0.6lipschitz continuity · 0.5value iteration · 0.4rate-distortion theory · 0.4information bottleneck · 0.4blahut-arimoto algorithm · 0.4approximation algorithm · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Theoretical Guarantees for Minimum Bayes Risk DecodingabstractMinimum Bayes Risk (MBR) decoding optimizes output selection by maximizing the expected utility value of an underlying human distribution.While prior work has shown the effectiveness of MBR decoding through empirical evaluation, few studies have analytically investigated why the method is effective.As a result of our analysis, we show that, given the size n of the reference hypothesis set used in computation, MBR decoding approaches the optimal solution with high probability at a rate of O n -1 2 , under certain assumptions, even though the language space Y is significantly larger |Y| ≫ n.This result helps to theoretically explain the strong performance observed in several prior empirical studies on MBR decoding.In addition, we provide the performance gap for maximum-a-posteriori (MAP) decoding and compare it to MBR decoding.The result of this paper indicates that MBR decoding tends to converge to the optimal solution faster than MAP decoding in several cases. Yuki Ichihara, Yuu Jinnai, Kaito Ariu, Tetsuro Morimura, Eiji Uchibe |
ACL (1) | 2 |
| 2025 | Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal TransportabstractDocument-level text generation tasks are known to be more difficult than sentence-level text generation tasks as they require the understanding of longer context to generate highquality texts.In this paper, we investigate the adaption of Minimum Bayes Risk (MBR) decoding for document-level text generation tasks.MBR decoding makes use of a utility function to estimate the output with the highest expected utility from a set of candidate outputs.Although MBR decoding is shown to be effective in a wide range of sentencelevel text generation tasks, its performance on document-level text generation tasks is limited as many of the utility functions are designed for evaluating the utility of sentences.To this end, we propose MBR-OT, a variant of MBR decoding using Wasserstein distance to compute the utility of a document using a sentence-level utility function.The experimental result shows that the performance of MBR-OT outperforms that of the standard MBR in document-level machine translation, text simplification, and dense image captioning tasks.Our code is available at https://github. com/jinnaiyuu/mbr-optimal-transport. Yuu Jinnai |
ACL (1) | 1 |
| 2025 | Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model AlignmentabstractYuu Jinnai, Tetsuro Morimura, Kaito Ariu, Kenshi Abe. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yuu Jinnai, Tetsuro Morimura, Kaito Ariu, Kenshi Abe |
NAACL (Long Papers) | 1 |
| 2024 | Filtered Direct Preference OptimizationabstractReinforcement learning from human feedback (RLHF) plays a crucial role in aligning language models with human preferences.While the significance of dataset quality is generally recognized, explicit investigations into its impact within the RLHF framework, to our knowledge, have been limited.This paper addresses the issue of text quality within the preference dataset by focusing on direct preference optimization (DPO), an increasingly adopted reward-model-free RLHF method.We confirm that text quality significantly influences the performance of models optimized with DPO more than those optimized with reward-modelbased RLHF.Building on this new insight, we propose an extension of DPO, termed filtered direct preference optimization (fDPO).fDPO uses a trained reward model to monitor the quality of texts within the preference dataset during DPO training.Samples of lower quality are discarded based on comparisons with texts generated by the model being optimized, resulting in a more accurate dataset.Experimental results demonstrate that fDPO enhances the final model performance.Our code is available at https://github.com/CyberAgentAILab/ filtered-dpo. Tetsuro Morimura, Mitsuki Sakamoto, Yuu Jinnai, Kenshi Abe, Kaito Ariu |
EMNLP | 3 |
| 2024 | Model-Based Minimum Bayes Risk Decoding for Text GenerationabstractMinimum Bayes Risk (MBR) decoding has been shown to be a powerful alternative to beam search decoding in a variety of text generation tasks. MBR decoding selects a hypothesis from a pool of hypotheses that has the least expected risk under a probability model according to a given utility function. Since it is impractical to compute the expected risk exactly over all possible hypotheses, two approximations are commonly used in MBR. First, it integrates over a sampled set of hypotheses rather than over all possible hypotheses. Second, it estimates the probability of each hypothesis using a Monte Carlo estimator. While the first approximation is necessary to make it computationally feasible, the second is not essential since we typically have access to the model probability at inference time. We propose model-based MBR (MBMBR), a variant of MBR that uses the model probability itself as the estimate of the probability distribution instead of the Monte Carlo estimate. We show analytically and empirically that the model-based estimate is more promising than the Monte Carlo estimate in text generation tasks. Our experiments show that MBMBR outperforms MBR in several text generation tasks, both with encoder-decoder models and with language models. Yuu Jinnai, Tetsuro Morimura, Ukyo Honda, Kaito Ariu, Kenshi Abe |
ICML | 1 |
| 2021 | Lipschitz Lifelong Reinforcement LearningabstractWe consider the problem of knowledge transfer when an agent is facing a series of Reinforcement Learning (RL) tasks. We introduce a novel metric between Markov Decision Processes and establish that close MDPs have close optimal value functions. Formally, the optimal value functions are Lipschitz continuous with respect to the tasks space. These theoretical results lead us to a value-transfer method for Lifelong RL, which we use to build a PAC-MDP algorithm with improved convergence rate. Further, we show the method to experience no negative transfer with high probability. We illustrate the benefits of the method in Lifelong RL experiments. Erwan Lecarpentier, David Abel, Kavosh Asadi, Yuu Jinnai, Emmanuel Rachelson, Michael L. Littman |
AAAI | 4 |
| 2020 | Neural Architecture Search Using Deep Neural Networks and Monte Carlo Tree SearchabstractNeural Architecture Search (NAS) has shown great success in automating the design of neural networks, but the prohibitive amount of computations behind current NAS methods requires further investigations in improving the sample efficiency and the network evaluation cost to get better results in a shorter time. In this paper, we present a novel scalable Monte Carlo Tree Search (MCTS) based NAS agent, named AlphaX, to tackle these two aspects. AlphaX improves the search efficiency by adaptively balancing the exploration and exploitation at the state level, and by a Meta-Deep Neural Network (DNN) to predict network accuracies for biasing the search toward a promising region. To amortize the network evaluation cost, AlphaX accelerates MCTS rollouts with a distributed design and reduces the number of epochs in evaluating a network by transfer learning, which is guided with the tree structure in MCTS. In 12 GPU days and 1000 samples, AlphaX found an architecture that reaches 97.84% top-1 accuracy on CIFAR-10, and 75.5% top-1 accuracy on ImageNet, exceeding SOTA NAS methods in both the accuracy and sampling efficiency. Particularly, we also evaluate AlphaX on NASBench-101, a large scale NAS dataset; AlphaX is 3x and 2.8x more sample efficient than Random Search and Regularized Evolution in finding the global optimum. Finally, we show the searched architecture improves a variety of vision applications from Neural Style Transfer, to Image Captioning and Object Detection. Linnan Wang, Yuu Jinnai, Yuandong Tian, Rodrigo Fonseca |
AAAI | 3 |
| 2020 | Exploration in Reinforcement Learning with Deep Covering Options
Yuu Jinnai, Jee Won Park, Marlos C. Machado, George Dimitri Konidaris |
ICLR | 1 |
| 2019 | State Abstraction as Compression in Apprenticeship LearningabstractState abstraction can give rise to models of environments that are both compressed and useful, thereby enabling efficient sequential decision making. In this work, we offer the first formalism and analysis of the trade-off between compression and performance made in the context of state abstraction for Apprenticeship Learning. We build on Rate-Distortion theory, the classic Blahut-Arimoto algorithm, and the Information Bottleneck method to develop an algorithm for computing state abstractions that approximate the optimal tradeoff between compression and performance. We illustrate the power of this algorithmic structure to offer insights into effective abstraction, compression, and reinforcement learning through a mixture of analysis, visuals, and experimentation. David Abel, Dilip Arumugam, Kavosh Asadi, Yuu Jinnai, Michael L. Littman, Lawson L. S. Wong |
AAAI | 4 |
| 2019 | Finding Options that Minimize Planning TimeabstractWe formalize the problem of selecting the optimal set of options for planning as that of computing the smallest set of options so that planning converges in less than a given maximum of value-iteration passes. We first show that the problem is $\NP$-hard, even if the task is constrained to be deterministic—the first such complexity result for option discovery. We then present the first polynomial-time boundedly suboptimal approximation algorithm for this setting, and empirically evaluate it against both the optimal options and a representative collection of heuristic approaches in simple grid-based domains. Yuu Jinnai, David Abel, D. Ellis Hershkowitz, Michael L. Littman, George Dimitri Konidaris |
ICML | 1 |
| 2019 | Discovering Options for Exploration by Minimizing Cover TimeabstractOne of the main challenges in reinforcement learning is solving tasks with sparse reward. We show that the difficulty of discovering a distant rewarding state in an MDP is bounded by the expected cover time of a random walk over the graph induced by the MDP’s transition dynamics. We therefore propose to accelerate exploration by constructing options that minimize cover time. We introduce a new option discovery algorithm that diminishes the expected cover time by connecting the most distant states in the state-space graph with options. We show empirically that the proposed algorithm improves learning in several domains with sparse rewards. Yuu Jinnai, Jee Won Park, David Abel, George Dimitri Konidaris |
ICML | 1 |
| 2018 | Policy and Value Transfer in Lifelong Reinforcement LearningabstractWe consider the problem of how best to use prior experience to bootstrap lifelong learning, where an agent faces a series of task instances drawn from some task distribution. First, we identify the initial policy that optimizes expected performance over the distribution of tasks for increasingly complex classes of policy and task distributions. We empirically demonstrate the relative performance of each policy class’ optimal element in a variety of simple task distributions. We then consider value-function initialization methods that preserve PAC guarantees while simultaneously minimizing the learning required in two learning algorithms, yielding MaxQInit, a practical new method for value-function-based transfer. We show that MaxQInit performs well in simple lifelong RL experiments. David Abel, Yuu Jinnai, Yue Guo 0003, George Dimitri Konidaris, Michael L. Littman |
ICML | 2 |
| 2017 | Learning to Prune Dominated Action Sequences in Online Black-Box PlanningabstractBlack-box domains where the successor states generated by applying an action are generated by a completely opaque simulator pose a challenge for domain-independent planning. The main computational bottleneck in search-based planning for such domains is the number of calls to the black-box simulation. We propose a method for significantly reducing the number of calls to the simulator by the search algorithm by detecting and pruning sequences of actions which are dominated by others. We apply our pruning method to Iterated Width and breadth-first search in domain-independent black-box planning for Atari 2600 games in the Arcade Learning Environment (ALE), adding our pruning method significantly improves upon the baseline algorithms. Yuu Jinnai, Alex S. Fukunaga |
AAAI | 1 |
| 2017 | Learning to Avoid Dominated Action Sequences in Planning for Black-Box Domains
Yuu Jinnai, Alex S. Fukunaga |
AAAI | 1 |
| 2017 | On Hash-Based Work Distribution Methods for Parallel Best-First SearchabstractParallel best-first search algorithms such as Hash Distributed A* (HDA*) distribute work among the processes using a global hash function. We analyze the search and communication overheads of state-of-the-art hash-based parallel best-first search algorithms, and show that although Zobrist hashing, the standard hash function used by HDA*, achieves good load balance for many domains, it incurs significant communication overhead since almost all generated nodes are transferred to a different processor than their parents. We propose Abstract Zobrist hashing, a new work distribution method for parallel search which, instead of computing a hash value based on the raw features of a state, uses a feature projection function to generate a set of abstract features which results in a higher locality, resulting in reduced communications overhead. We show that Abstract Zobrist hashing outperforms previous methods on search domains using hand-coded, domain specific feature projection functions. We then propose GRAZHDA*, a graph-partitioning based approach to automatically generating feature projection functions. GRAZHDA* seeks to approximate the partitioning of the actual search space graph by partitioning the domain transition graph, an abstraction of the state space graph. We show that GRAZHDA* outperforms previous methods on domain-independent planning. Yuu Jinnai, Alex S. Fukunaga |
J. Artif. Intell. Res. | 1 |
| 2016 | Abstract Zobrist Hashing: An Efficient Work Distribution Method for Parallel Best-First Search
Yuu Jinnai, Alex S. Fukunaga |
AAAI | 1 |