Yuu Jinnai

dblp:178/8539 · DBLP profile ↗
← Back
16ranked-venue papers
10as first author
6since 2021 · last 2025
0000-0002-8756-644XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 10 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Language models and text generation · 40% Reinforcement learning · 36% Planning, search and constraint satisfaction · 16%
Theoretical computer science
2 papers
Computational complexity · 44% Approximation and online algorithms · 44% Coding theory · 13%

Topics — the 27 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
decoding
2.532025
Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport · ACL (1) 2025
Theoretical Guarantees for Minimum Bayes Risk Decoding · ACL (1) 2025
Model-Based Minimum Bayes Risk Decoding for Text Generation · ICML 2024
Natural language and speech › Language models and text generation › decoding
minimum bayes risk decoding
2.532025
Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport · ACL (1) 2025
Theoretical Guarantees for Minimum Bayes Risk Decoding · ACL (1) 2025
Model-Based Minimum Bayes Risk Decoding for Text Generation · ICML 2024
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery
1.232020
Exploration in Reinforcement Learning with Deep Covering Options · ICLR 2020
Discovering Options for Exploration by Minimizing Cover Time · ICML 2019
Finding Options that Minimize Planning Time · ICML 2019
Natural language and speech › Machine translation
document-level machine translation
0.912025
Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport · ACL (1) 2025
Machine learning › Reinforcement learning › non-stationary reinforcement learning
continual reinforcement learning
0.822021
Lipschitz Lifelong Reinforcement Learning · AAAI 2021
Policy and Value Transfer in Lifelong Reinforcement Learning · ICML 2018
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search
0.832017
Learning to Avoid Dominated Action Sequences in Planning for Black-Box Domains · AAAI 2017
Learning to Prune Dominated Action Sequences in Online Black-Box Planning · AAAI 2017
Abstract Zobrist Hashing: An Efficient Work Distribution Method for Parallel Best-First Search · AAAI 2016
Machine learning › Reinforcement learning
exploration
0.822020
Exploration in Reinforcement Learning with Deep Covering Options · ICLR 2020
Discovering Options for Exploration by Minimizing Cover Time · ICML 2019
Natural language and speech › Language models and text generation
alignment
0.812024
Filtered Direct Preference Optimization · EMNLP 2024
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.812024
Filtered Direct Preference Optimization · EMNLP 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
black-box planning
0.622017
Learning to Avoid Dominated Action Sequences in Planning for Black-Box Domains · AAAI 2017
Learning to Prune Dominated Action Sequences in Online Black-Box Planning · AAAI 2017
Machine learning › Reinforcement learning
transfer learning in reinforcement learning
0.512021
Lipschitz Lifelong Reinforcement Learning · AAAI 2021
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.412020
Exploration in Reinforcement Learning with Deep Covering Options · ICLR 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.412020
Neural Architecture Search Using Deep Neural Networks and Monte Carlo Tree Search · AAAI 2020
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.412020
Neural Architecture Search Using Deep Neural Networks and Monte Carlo Tree Search · AAAI 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
hierarchical planning
0.412019
Finding Options that Minimize Planning Time · ICML 2019
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.412019
State Abstraction as Compression in Apprenticeship Learning · AAAI 2019
Machine learning › Reinforcement learning
sparse reward
0.412019
Discovering Options for Exploration by Minimizing Cover Time · ICML 2019
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state abstraction
0.412019
State Abstraction as Compression in Apprenticeship Learning · AAAI 2019
Approximation and online algorithms
approximation algorithms
0.412019
Finding Options that Minimize Planning Time · ICML 2019
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
policy transfer
0.312018
Policy and Value Transfer in Lifelong Reinforcement Learning · ICML 2018
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search
best-first search
0.212016
Abstract Zobrist Hashing: An Efficient Work Distribution Method for Parallel Best-First Search · AAAI 2016
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search › search-based problem solving
parallel search
0.212016
Abstract Zobrist Hashing: An Efficient Work Distribution Method for Parallel Best-First Search · AAAI 2016
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.212024
Filtered Direct Preference Optimization · EMNLP 2024
Natural language and speech › Language models and text generation
text generation
0.212024
Model-Based Minimum Bayes Risk Decoding for Text Generation · ICML 2024
Machine learning › Reinforcement learning › sample efficiency
PAC-MDP
0.112021
Lipschitz Lifelong Reinforcement Learning · AAAI 2021
Coding theory › source coding
rate-distortion theory
0.112019
State Abstraction as Compression in Apprenticeship Learning · AAAI 2019
Machine learning › Reinforcement learning › deep reinforcement learning
atari game playing
0.112017
Learning to Prune Dominated Action Sequences in Online Black-Box Planning · AAAI 2017

Methods — techniques the papers use, named apart from their topics

wasserstein distance · 0.9statistical learning theory · 0.9optimal transport · 0.9concentration bounds · 0.9reward model · 0.8monte carlo estimation · 0.8filtered direct preference optimization · 0.8beam search · 0.8dominance pruning · 0.6lipschitz continuity · 0.5value iteration · 0.4rate-distortion theory · 0.4information bottleneck · 0.4blahut-arimoto algorithm · 0.4approximation algorithm · 0.4
YearPublicationVenuePosition
2025 Theoretical Guarantees for Minimum Bayes Risk Decoding
abstract
Minimum Bayes Risk (MBR) decoding optimizes output selection by maximizing the expected utility value of an underlying human distribution.While prior work has shown the effectiveness of MBR decoding through empirical evaluation, few studies have analytically investigated why the method is effective.As a result of our analysis, we show that, given the size n of the reference hypothesis set used in computation, MBR decoding approaches the optimal solution with high probability at a rate of O n -1 2 , under certain assumptions, even though the language space Y is significantly larger |Y| ≫ n.This result helps to theoretically explain the strong performance observed in several prior empirical studies on MBR decoding.In addition, we provide the performance gap for maximum-a-posteriori (MAP) decoding and compare it to MBR decoding.The result of this paper indicates that MBR decoding tends to converge to the optimal solution faster than MAP decoding in several cases.
Yuki Ichihara, Yuu Jinnai, Kaito Ariu, Tetsuro Morimura, Eiji Uchibe
ACL (1)2
2025 Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport
abstract
Document-level text generation tasks are known to be more difficult than sentence-level text generation tasks as they require the understanding of longer context to generate highquality texts.In this paper, we investigate the adaption of Minimum Bayes Risk (MBR) decoding for document-level text generation tasks.MBR decoding makes use of a utility function to estimate the output with the highest expected utility from a set of candidate outputs.Although MBR decoding is shown to be effective in a wide range of sentencelevel text generation tasks, its performance on document-level text generation tasks is limited as many of the utility functions are designed for evaluating the utility of sentences.To this end, we propose MBR-OT, a variant of MBR decoding using Wasserstein distance to compute the utility of a document using a sentence-level utility function.The experimental result shows that the performance of MBR-OT outperforms that of the standard MBR in document-level machine translation, text simplification, and dense image captioning tasks.Our code is available at https://github. com/jinnaiyuu/mbr-optimal-transport.
Yuu Jinnai
ACL (1)1
2025 Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment
abstract
Yuu Jinnai, Tetsuro Morimura, Kaito Ariu, Kenshi Abe. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yuu Jinnai, Tetsuro Morimura, Kaito Ariu, Kenshi Abe
NAACL (Long Papers)1
2024 Filtered Direct Preference Optimization
abstract
Reinforcement learning from human feedback (RLHF) plays a crucial role in aligning language models with human preferences.While the significance of dataset quality is generally recognized, explicit investigations into its impact within the RLHF framework, to our knowledge, have been limited.This paper addresses the issue of text quality within the preference dataset by focusing on direct preference optimization (DPO), an increasingly adopted reward-model-free RLHF method.We confirm that text quality significantly influences the performance of models optimized with DPO more than those optimized with reward-modelbased RLHF.Building on this new insight, we propose an extension of DPO, termed filtered direct preference optimization (fDPO).fDPO uses a trained reward model to monitor the quality of texts within the preference dataset during DPO training.Samples of lower quality are discarded based on comparisons with texts generated by the model being optimized, resulting in a more accurate dataset.Experimental results demonstrate that fDPO enhances the final model performance.Our code is available at https://github.com/CyberAgentAILab/ filtered-dpo.
Tetsuro Morimura, Mitsuki Sakamoto, Yuu Jinnai, Kenshi Abe, Kaito Ariu
EMNLP3
2024 Model-Based Minimum Bayes Risk Decoding for Text Generation
abstract
Minimum Bayes Risk (MBR) decoding has been shown to be a powerful alternative to beam search decoding in a variety of text generation tasks. MBR decoding selects a hypothesis from a pool of hypotheses that has the least expected risk under a probability model according to a given utility function. Since it is impractical to compute the expected risk exactly over all possible hypotheses, two approximations are commonly used in MBR. First, it integrates over a sampled set of hypotheses rather than over all possible hypotheses. Second, it estimates the probability of each hypothesis using a Monte Carlo estimator. While the first approximation is necessary to make it computationally feasible, the second is not essential since we typically have access to the model probability at inference time. We propose model-based MBR (MBMBR), a variant of MBR that uses the model probability itself as the estimate of the probability distribution instead of the Monte Carlo estimate. We show analytically and empirically that the model-based estimate is more promising than the Monte Carlo estimate in text generation tasks. Our experiments show that MBMBR outperforms MBR in several text generation tasks, both with encoder-decoder models and with language models.
Yuu Jinnai, Tetsuro Morimura, Ukyo Honda, Kaito Ariu, Kenshi Abe
ICML1
2021 Lipschitz Lifelong Reinforcement Learning
abstract
We consider the problem of knowledge transfer when an agent is facing a series of Reinforcement Learning (RL) tasks. We introduce a novel metric between Markov Decision Processes and establish that close MDPs have close optimal value functions. Formally, the optimal value functions are Lipschitz continuous with respect to the tasks space. These theoretical results lead us to a value-transfer method for Lifelong RL, which we use to build a PAC-MDP algorithm with improved convergence rate. Further, we show the method to experience no negative transfer with high probability. We illustrate the benefits of the method in Lifelong RL experiments.
Erwan Lecarpentier, David Abel, Kavosh Asadi, Yuu Jinnai, Emmanuel Rachelson, Michael L. Littman
AAAI4
2020 Neural Architecture Search Using Deep Neural Networks and Monte Carlo Tree Search
abstract
Neural Architecture Search (NAS) has shown great success in automating the design of neural networks, but the prohibitive amount of computations behind current NAS methods requires further investigations in improving the sample efficiency and the network evaluation cost to get better results in a shorter time. In this paper, we present a novel scalable Monte Carlo Tree Search (MCTS) based NAS agent, named AlphaX, to tackle these two aspects. AlphaX improves the search efficiency by adaptively balancing the exploration and exploitation at the state level, and by a Meta-Deep Neural Network (DNN) to predict network accuracies for biasing the search toward a promising region. To amortize the network evaluation cost, AlphaX accelerates MCTS rollouts with a distributed design and reduces the number of epochs in evaluating a network by transfer learning, which is guided with the tree structure in MCTS. In 12 GPU days and 1000 samples, AlphaX found an architecture that reaches 97.84% top-1 accuracy on CIFAR-10, and 75.5% top-1 accuracy on ImageNet, exceeding SOTA NAS methods in both the accuracy and sampling efficiency. Particularly, we also evaluate AlphaX on NASBench-101, a large scale NAS dataset; AlphaX is 3x and 2.8x more sample efficient than Random Search and Regularized Evolution in finding the global optimum. Finally, we show the searched architecture improves a variety of vision applications from Neural Style Transfer, to Image Captioning and Object Detection.
Linnan Wang, Yuu Jinnai, Yuandong Tian, Rodrigo Fonseca
AAAI3
2020 Exploration in Reinforcement Learning with Deep Covering Options
Yuu Jinnai, Jee Won Park, Marlos C. Machado, George Dimitri Konidaris
ICLR1
2019 State Abstraction as Compression in Apprenticeship Learning
abstract
State abstraction can give rise to models of environments that are both compressed and useful, thereby enabling efficient sequential decision making. In this work, we offer the first formalism and analysis of the trade-off between compression and performance made in the context of state abstraction for Apprenticeship Learning. We build on Rate-Distortion theory, the classic Blahut-Arimoto algorithm, and the Information Bottleneck method to develop an algorithm for computing state abstractions that approximate the optimal tradeoff between compression and performance. We illustrate the power of this algorithmic structure to offer insights into effective abstraction, compression, and reinforcement learning through a mixture of analysis, visuals, and experimentation.
David Abel, Dilip Arumugam, Kavosh Asadi, Yuu Jinnai, Michael L. Littman, Lawson L. S. Wong
AAAI4
2019 Finding Options that Minimize Planning Time
abstract
We formalize the problem of selecting the optimal set of options for planning as that of computing the smallest set of options so that planning converges in less than a given maximum of value-iteration passes. We first show that the problem is $\NP$-hard, even if the task is constrained to be deterministic—the first such complexity result for option discovery. We then present the first polynomial-time boundedly suboptimal approximation algorithm for this setting, and empirically evaluate it against both the optimal options and a representative collection of heuristic approaches in simple grid-based domains.
Yuu Jinnai, David Abel, D. Ellis Hershkowitz, Michael L. Littman, George Dimitri Konidaris
ICML1
2019 Discovering Options for Exploration by Minimizing Cover Time
abstract
One of the main challenges in reinforcement learning is solving tasks with sparse reward. We show that the difficulty of discovering a distant rewarding state in an MDP is bounded by the expected cover time of a random walk over the graph induced by the MDP’s transition dynamics. We therefore propose to accelerate exploration by constructing options that minimize cover time. We introduce a new option discovery algorithm that diminishes the expected cover time by connecting the most distant states in the state-space graph with options. We show empirically that the proposed algorithm improves learning in several domains with sparse rewards.
Yuu Jinnai, Jee Won Park, David Abel, George Dimitri Konidaris
ICML1
2018 Policy and Value Transfer in Lifelong Reinforcement Learning
abstract
We consider the problem of how best to use prior experience to bootstrap lifelong learning, where an agent faces a series of task instances drawn from some task distribution. First, we identify the initial policy that optimizes expected performance over the distribution of tasks for increasingly complex classes of policy and task distributions. We empirically demonstrate the relative performance of each policy class’ optimal element in a variety of simple task distributions. We then consider value-function initialization methods that preserve PAC guarantees while simultaneously minimizing the learning required in two learning algorithms, yielding MaxQInit, a practical new method for value-function-based transfer. We show that MaxQInit performs well in simple lifelong RL experiments.
David Abel, Yuu Jinnai, Yue Guo 0003, George Dimitri Konidaris, Michael L. Littman
ICML2
2017 Learning to Prune Dominated Action Sequences in Online Black-Box Planning
abstract
Black-box domains where the successor states generated by applying an action are generated by a completely opaque simulator pose a challenge for domain-independent planning. The main computational bottleneck in search-based planning for such domains is the number of calls to the black-box simulation. We propose a method for significantly reducing the number of calls to the simulator by the search algorithm by detecting and pruning sequences of actions which are dominated by others. We apply our pruning method to Iterated Width and breadth-first search in domain-independent black-box planning for Atari 2600 games in the Arcade Learning Environment (ALE), adding our pruning method significantly improves upon the baseline algorithms.
Yuu Jinnai, Alex S. Fukunaga
AAAI1
2017 Learning to Avoid Dominated Action Sequences in Planning for Black-Box Domains
Yuu Jinnai, Alex S. Fukunaga
AAAI1
2017 On Hash-Based Work Distribution Methods for Parallel Best-First Search
abstract
Parallel best-first search algorithms such as Hash Distributed A* (HDA*) distribute work among the processes using a global hash function. We analyze the search and communication overheads of state-of-the-art hash-based parallel best-first search algorithms, and show that although Zobrist hashing, the standard hash function used by HDA*, achieves good load balance for many domains, it incurs significant communication overhead since almost all generated nodes are transferred to a different processor than their parents. We propose Abstract Zobrist hashing, a new work distribution method for parallel search which, instead of computing a hash value based on the raw features of a state, uses a feature projection function to generate a set of abstract features which results in a higher locality, resulting in reduced communications overhead. We show that Abstract Zobrist hashing outperforms previous methods on search domains using hand-coded, domain specific feature projection functions. We then propose GRAZHDA*, a graph-partitioning based approach to automatically generating feature projection functions. GRAZHDA* seeks to approximate the partitioning of the actual search space graph by partitioning the domain transition graph, an abstraction of the state space graph. We show that GRAZHDA* outperforms previous methods on domain-independent planning.
Yuu Jinnai, Alex S. Fukunaga
J. Artif. Intell. Res.1
2016 Abstract Zobrist Hashing: An Efficient Work Distribution Method for Parallel Best-First Search
Yuu Jinnai, Alex S. Fukunaga
AAAI1