Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mitsuki Sakamoto

dblp:243/6951 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 57% Mathematical optimization · 43%
Artificial intelligence
1 paper
Language models and text generation · 87% Reinforcement learning · 13%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Algorithmic game theory and mechanism design
equilibrium computation
0.912025
Boosting Perturbed Gradient Ascent for Last-Iterate Convergence in Games · ICLR 2025
Mathematical optimization › continuous optimization › convex optimization
first-order methods
0.912025
Boosting Perturbed Gradient Ascent for Last-Iterate Convergence in Games · ICLR 2025
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
gradient ascent
0.912025
Boosting Perturbed Gradient Ascent for Last-Iterate Convergence in Games · ICLR 2025
Algorithmic game theory and mechanism design › game dynamics › equilibrium convergence
last-iterate convergence
0.912025
Boosting Perturbed Gradient Ascent for Last-Iterate Convergence in Games · ICLR 2025
Natural language and speech › Language models and text generation
alignment
0.812024
Filtered Direct Preference Optimization · EMNLP 2024
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.812024
Filtered Direct Preference Optimization · EMNLP 2024
Algorithmic game theory and mechanism design
learning in games
0.812024
Adaptively Perturbed Mirror Descent for Learning in Games · ICML 2024
Mathematical optimization › continuous optimization › convex optimization › first-order methods
mirror descent
0.812024
Adaptively Perturbed Mirror Descent for Learning in Games · ICML 2024
Algorithmic game theory and mechanism design › game dynamics › equilibrium convergence
nash equilibrium convergence
0.812024
Adaptively Perturbed Mirror Descent for Learning in Games · ICML 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.212024
Filtered Direct Preference Optimization · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

payoff perturbation · 1.6strong convexity · 0.9reward model · 0.8optimistic mirror descent · 0.8mirror descent · 0.8filtered direct preference optimization · 0.8
YearPublicationVenuePosition
2025 Boosting Perturbed Gradient Ascent for Last-Iterate Convergence in Games
abstract
This paper presents a payoff perturbation technique, introducing a strong convexity to players' payoff functions in games. This technique is specifically designed for first-order methods to achieve last-iterate convergence in games where the gradient of the payoff functions is monotone in the strategy profile space, potentially containing additive noise. Although perturbation is known to facilitate the convergence of learning algorithms, the magnitude of perturbation requires careful adjustment to ensure last-iterate convergence. Previous studies have proposed a scheme in which the magnitude is determined by the distance from a periodically re-initialized anchoring or reference strategy. Building upon this, we propose Gradient Ascent with Boosting Payoff Perturbation, which incorporates a novel perturbation into the underlying payoff function, maintaining the periodically re-initializing anchoring strategy scheme. This innovation empowers us to provide faster last-iterate convergence rates against the existing payoff perturbed algorithms, even in the presence of additive noise.
Kenshi Abe, Mitsuki Sakamoto, Kaito Ariu, Atsushi Iwasaki
ICLR2
2024 Filtered Direct Preference Optimization
abstract
Reinforcement learning from human feedback (RLHF) plays a crucial role in aligning language models with human preferences.While the significance of dataset quality is generally recognized, explicit investigations into its impact within the RLHF framework, to our knowledge, have been limited.This paper addresses the issue of text quality within the preference dataset by focusing on direct preference optimization (DPO), an increasingly adopted reward-model-free RLHF method.We confirm that text quality significantly influences the performance of models optimized with DPO more than those optimized with reward-modelbased RLHF.Building on this new insight, we propose an extension of DPO, termed filtered direct preference optimization (fDPO).fDPO uses a trained reward model to monitor the quality of texts within the preference dataset during DPO training.Samples of lower quality are discarded based on comparisons with texts generated by the model being optimized, resulting in a more accurate dataset.Experimental results demonstrate that fDPO enhances the final model performance.Our code is available at https://github.com/CyberAgentAILab/ filtered-dpo.
Tetsuro Morimura, Mitsuki Sakamoto, Yuu Jinnai, Kenshi Abe, Kaito Ariu
EMNLP2
2024 Adaptively Perturbed Mirror Descent for Learning in Games
abstract
This paper proposes a payoff perturbation technique for the Mirror Descent (MD) algorithm in games where the gradient of the payoff functions is monotone in the strategy profile space, potentially containing additive noise. The optimistic family of learning algorithms, exemplified by optimistic MD, successfully achieves *last-iterate* convergence in scenarios devoid of noise, leading the dynamics to a Nash equilibrium. A recent re-emerging trend underscores the promise of the perturbation approach, where payoff functions are perturbed based on the distance from an anchoring, or *slingshot*, strategy. In response, we propose *Adaptively Perturbed MD* (APMD), which adjusts the magnitude of the perturbation by repeatedly updating the slingshot strategy at a predefined interval. This innovation empowers us to find a Nash equilibrium of the underlying game with guaranteed rates. Empirical demonstrations affirm that our algorithm exhibits significantly accelerated convergence.
Kenshi Abe, Kaito Ariu, Mitsuki Sakamoto, Atsushi Iwasaki
ICML3
2023 Last-Iterate Convergence with Full and Noisy Feedback in Two-Player Zero-Sum Games
abstract
This paper proposes Mutation-Driven Multiplicative Weights Update (M2WU) for learning an equilibrium in two-player zero-sum normal-form games and proves that it exhibits the last-iterate convergence property in both full and noisy feedback settings. In the former, players observe their exact gradient vectors of the utility functions. In the latter, they only observe the noisy gradient vectors. Even the celebrated Multiplicative Weights Update (MWU) and Optimistic MWU (OMWU) algorithms may not converge to a Nash equilibrium with noisy feedback. On the contrary, M2WU exhibits the last-iterate convergence to a stationary point near a Nash equilibrium in both feedback settings. We then prove that it converges to an exact Nash equilibrium by iteratively adapting the mutation term. We empirically confirm that M2WU outperforms MWU and OMWU in exploitability and convergence rates.
Kenshi Abe, Kaito Ariu, Mitsuki Sakamoto, Kentaro Toyoshima, Atsushi Iwasaki
AISTATS3
2022 Mutation-driven follow the regularized leader for last-iterate convergence in zero-sum games
abstract
In this study, we consider a variant of the Follow the Regularized Leader (FTRL) dynamics in two-player zero-sum games. FTRL is guaranteed to converge to a Nash equilibrium when time-averaging the strategies, while a lot of variants suffer from the issue of limit cycling behavior, i.e., lack the last-iterate convergence guarantee. To this end, we propose mutant FTRL (M-FTRL), an algorithm that introduces mutation for the perturbation of action probabilities. We then investigate the continuous-time dynamics of M-FTRL and provide the strong convergence guarantees toward stationary points that approximate Nash equilibria under full-information feedback. Furthermore, our simulation demonstrates that M-FTRL can enjoy faster convergence rates than FTRL and optimistic FTRL under full-information feedback and surprisingly exhibits clear convergence under bandit feedback.
Kenshi Abe, Mitsuki Sakamoto, Atsushi Iwasaki
UAI2