Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yi-Shan Wu 0003

dblp:138/4357-3 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-7949-0115ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Learning theory · 88% Kernel, tree and ensemble methods · 8% Probabilistic and Bayesian machine learning · 4%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
generalization bounds
1.432024
Recursive PAC-Bayes: A Frequentist Approach to Sequential Prior Updates with No Information Loss · NeurIPS 2024
Chebyshev-Cantelli PAC-Bayes-Bennett Inequality for the Weighted Majority Vote · NeurIPS 2021
Split-kl and PAC-Bayes-split-kl Inequalities for Ternary Random Variables · NeurIPS 2022
Machine learning › Learning theory › generalization bounds
PAC-Bayes bounds
1.122022
Split-kl and PAC-Bayes-split-kl Inequalities for Ternary Random Variables · NeurIPS 2022
Chebyshev-Cantelli PAC-Bayes-Bennett Inequality for the Weighted Majority Vote · NeurIPS 2021
Machine learning › Learning theory
PAC-Bayesian analysis
0.812024
Recursive PAC-Bayes: A Frequentist Approach to Sequential Prior Updates with No Information Loss · NeurIPS 2024
Machine learning › Learning theory
concentration inequalities
0.612022
Split-kl and PAC-Bayes-split-kl Inequalities for Ternary Random Variables · NeurIPS 2022
Machine learning › Learning theory
concentration of measure
0.512021
Chebyshev-Cantelli PAC-Bayes-Bennett Inequality for the Weighted Majority Vote · NeurIPS 2021
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.512021
Chebyshev-Cantelli PAC-Bayes-Bennett Inequality for the Weighted Majority Vote · NeurIPS 2021
Machine learning › Learning theory › excess risk bounds
oracle inequality
0.512021
Chebyshev-Cantelli PAC-Bayes-Bennett Inequality for the Weighted Majority Vote · NeurIPS 2021
Machine learning › Learning theory
weighted majority vote
0.512021
Chebyshev-Cantelli PAC-Bayes-Bennett Inequality for the Weighted Majority Vote · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.212024
Recursive PAC-Bayes: A Frequentist Approach to Sequential Prior Updates with No Information Loss · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

split-kl inequality · 0.8PAC-Bayes · 0.8kl inequality · 0.6empirical bernstein inequality · 0.6chebyshev-cantelli inequality · 0.5bennett's inequality · 0.5PAC-Bayesian bounding · 0.5
YearPublicationVenuePosition
2025 Deep Exploration with PAC-Bayes
abstract
Reinforcement learning (RL) for continuous control under delayed rewards is an under-explored problem despite its significance in real-world applications. Many complex skills are based on intermediate ones as prerequisites. For instance, a humanoid locomotor must learn how to stand before it can learn to walk. To cope with delayed reward, an agent must perform deep exploration. However, existing deep exploration methods are designed for small discrete action spaces, and their generalization to state-of-the-art continuous control remains unproven. We address the deep exploration problem for the first time from a PAC-Bayesian perspective in the context of actor-critic learning. To do this, we quantify the error of the Bellman operator through a PAC-Bayes bound, where a bootstrapped ensemble of critic networks represents the posterior distribution, and their targets serve as a data-informed function-space prior. We derive an objective function from this bound and use it to train the critic ensemble. Each critic trains an individual soft actor network, implemented as a shared trunk and critic-specific heads. The agent performs deep exploration by acting epsilon-softly on a randomly chosen actor head. Our proposed algorithm, named PAC-Bayesian Actor-Critic (PBAC), is the only algorithm to consistently discover delayed rewards on continuous control tasks with varying difficulty.
Bahareh Tasdighi, Manuel Haußmann, Nicklas Werge, Yi-Shan Wu 0003, Melih Kandemir
ECAI4
2025 Improving Actor-Critic Training with Steerable Action-Value Approximation Errors
abstract
Off-policy actor-critic algorithms have shown strong potential in deep reinforcement learning for continuous control tasks. Their success primarily comes from leveraging pessimistic state-action value function updates, which reduce function approximation errors and stabilize learning. However, excessive pessimism can limit exploration, preventing the agent from effectively refining its policies. Conversely, optimism can encourage exploration but may lead to high-risk behaviors and unstable learning if not carefully managed. To address this trade-off, we propose Utility Soft Actor-Critic (USAC), a novel framework that allows independent, interpretable control of pessimism and optimism for both the actor and the critic. USAC dynamically adapts its exploration strategy based on the uncertainty of critics using a utility function, enabling a task-specific balance between optimism and pessimism. This approach goes beyond binary choices of pessimism or optimism, making the method both theoretically meaningful and practically feasible. Experiments across a variety of continuous control tasks show that adjusting the degree of pessimism or optimism significantly impacts performance. When configured appropriately, USAC consistently outperforms state-of-the-art algorithms, demonstrating its practical utility and feasibility.
Bahareh Tasdighi, Nicklas Werge, Yi-Shan Wu 0003, Melih Kandemir
ECAI3
2024 Recursive PAC-Bayes: A Frequentist Approach to Sequential Prior Updates with No Information Loss
abstract
PAC-Bayesian analysis is a frequentist framework for incorporating prior knowledge into learning. It was inspired by Bayesian learning, which allows sequential data processing and naturally turns posteriors from one processing step into priors for the next. However, despite two and a half decades of research, the ability to update priors sequentially without losing confidence information along the way remained elusive for PAC-Bayes. While PAC-Bayes allows construction of data-informed priors, the final confidence intervals depend only on the number of points that were not used for the construction of the prior, whereas confidence information in the prior, which is related to the number of points used to construct the prior, is lost. This limits the possibility and benefit of sequential prior updates, because the final bounds depend only on the size of the final batch. We present a novel and, in retrospect, surprisingly simple and powerful PAC-Bayesian procedure that allows sequential prior updates with no information loss. The procedure is based on a novel decomposition of the expected loss of randomized classifiers. The decomposition rewrites the loss of the posterior as an excess loss relative to a downscaled loss of the prior plus the downscaled loss of the prior, which is bounded recursively. As a side result, we also present a generalization of the split-kl and PAC-Bayes-split-kl inequalities to discrete random variables, which we use for bounding the excess losses, and which can be of independent interest. In empirical evaluation the new procedure significantly outperforms state-of-the-art.
Yi-Shan Wu 0003, Badr-Eddine Chérief-Abdellatif, Yevgeny Seldin
NeurIPS1
2022 Split-kl and PAC-Bayes-split-kl Inequalities for Ternary Random Variables
abstract
We present a new concentration of measure inequality for sums of independent bounded random variables, which we name a split-kl inequality. The inequality combines the combinatorial power of the kl inequality with ability to exploit low variance. While for Bernoulli random variables the kl inequality is tighter than the Empirical Bernstein, for random variables taking values inside a bounded interval and having low variance the Empirical Bernstein inequality is tighter than the kl. The proposed split-kl inequality yields the best of both worlds. We discuss an application of the split-kl inequality to bounding excess losses. We also derive a PAC-Bayes-split-kl inequality and use a synthetic example and several UCI datasets to compare it with the PAC-Bayes-kl, PAC-Bayes Empirical Bernstein, PAC-Bayes Unexpected Bernstein, and PAC-Bayes Empirical Bennett inequalities.
Yi-Shan Wu 0003, Yevgeny Seldin
NeurIPS1
2021 Chebyshev-Cantelli PAC-Bayes-Bennett Inequality for the Weighted Majority Vote
abstract
We present a new second-order oracle bound for the expected risk of a weighted majority vote. The bound is based on a novel parametric form of the Chebyshev-Cantelli inequality (a.k.a. one-sided Chebyshev’s), which is amenable to efficient minimization. The new form resolves the optimization challenge faced by prior oracle bounds based on the Chebyshev-Cantelli inequality, the C-bounds [Germain et al., 2015], and, at the same time, it improves on the oracle bound based on second order Markov’s inequality introduced by Masegosa et al. [2020]. We also derive a new concentration of measure inequality, which we name PAC-Bayes-Bennett, since it combines PAC-Bayesian bounding with Bennett’s inequality. We use it for empirical estimation of the oracle bound. The PAC-Bayes-Bennett inequality improves on the PAC-Bayes-Bernstein inequality of Seldin et al. [2012]. We provide an empirical evaluation demonstrating that the new bounds can improve on the work of Masegosa et al. [2020]. Both the parametric form of the Chebyshev-Cantelli inequality and the PAC-Bayes-Bennett inequality may be of independent interest for the study of concentration of measure in other domains.
Yi-Shan Wu 0003, Andrés R. Masegosa, Stephan Sloth Lorenzen, Christian Igel, Yevgeny Seldin
NeurIPS1