Yuqi Yun

dblp:348/4761 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2024
0009-0006-4397-1538ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 61% Probabilistic and Bayesian machine learning · 30% Generative modeling · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
dirichlet process mixture model
0.812024
Context-Based Meta-Reinforcement Learning With Bayesian Nonparametric Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Reinforcement learning
meta-reinforcement learning
0.812024
Context-Based Meta-Reinforcement Learning With Bayesian Nonparametric Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
task representation learning
0.812024
Context-Based Meta-Reinforcement Learning With Bayesian Nonparametric Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Generative modeling
variational autoencoder
0.212024
Context-Based Meta-Reinforcement Learning With Bayesian Nonparametric Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024

Methods — techniques the papers use, named apart from their topics

zero-shot adaptation · 0.8recurrence-based context encoding · 0.8deep clustering · 0.8
YearPublicationVenuePosition
2024 Context-Based Meta-Reinforcement Learning With Bayesian Nonparametric Models
abstract
Deep reinforcement learning agents usually need to collect a large number of interactions to solve a single task. In contrast, meta-reinforcement learning (meta-RL) aims to quickly adapt to new tasks using a small amount of experience by leveraging the knowledge from training on a set of similar tasks. State-of-the-art context-based meta-RL algorithms use the context to encode the task information and train a policy conditioned on the inferred latent task encoding. However, most recent works are limited to parametric tasks, where a handful of variables control the full variation in the task distribution, and also failed to work in non-stationary environments due to the few-shot adaptation setting. To address those limitations, we propose MEta-reinforcement Learning with Task Self-discovery (MELTS), which adaptively learns qualitatively different nonparametric tasks and adapts to new tasks in a zero-shot manner. We introduce a novel deep clustering framework (DPMM-VAE) based on an infinite mixture of Gaussians, which combines the Dirichlet process mixture model (DPMM) and the variational autoencoder (VAE), to simultaneously learn task representations and cluster the tasks in a self-adaptive way. Integrating DPMM-VAE into MELTS enables it to adaptively discover the multi-modal structure of the nonparametric task distribution, which previous methods using isotropic Gaussian random variables cannot model. In addition, we propose a zero-shot adaptation mechanism and a recurrence-based context encoding strategy to improve the data efficiency and make our algorithm applicable in non-stationary environments. On various continuous control tasks with both parametric and nonparametric variations, our algorithm produces a more structured and self-adaptive task latent space and also achieves superior sample efficiency and asymptotic performance compared with state-of-the-art meta-RL algorithms.
Zhenshan Bing, Yuqi Yun, Kai Huang 0001, Alois C. Knoll
IEEE Trans. Pattern Anal. Mach. Intell.2