EDBT 2026 Demo / reviewers in the wild / expert
Yu Chen 0074
dblp:87/1254-74
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2025
0009-0006-9503-6613ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 72% Learning theory · 19% Generative modeling · 5% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 87% Mathematical optimization · 13% | |
| Computer networks
1 paper |
Network optimization and economics · 87% Cellular and mobile networks · 13% |
Topics — the 28 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning |
2.3 | 3 | 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024 Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation · ICML 2024 Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human Feedback · ICLR 2024 |
Machine learning › Learning theory › online learning
regret bounds |
1.5 | 2 | 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024 Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation · ICML 2024 |
Machine learning › Reinforcement learning › markov decision process › low-rank MDP
linear MDP |
1.2 | 2 | 2023 | Towards Minimax Optimal Reward-free Reinforcement Learning in Linear MDPs · ICLR 2023 Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation · ICML 2022 |
Machine learning › Learning theory
minimax optimality |
1.2 | 2 | 2023 | Towards Minimax Optimal Reward-free Reinforcement Learning in Linear MDPs · ICLR 2023 Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation · ICML 2022 |
Machine learning › Optimization for machine learning
convergence analysis |
0.9 | 1 | 2025 | Finite-Time Analysis of Discrete-Time Stochastic Interpolants · ICML 2025 |
Machine learning › Generative modeling › generative model › continuous-time generative model
stochastic interpolants |
0.9 | 1 | 2025 | Finite-Time Analysis of Discrete-Time Stochastic Interpolants · ICML 2025 |
Algorithmic game theory and mechanism design
multi-armed bandit |
0.9 | 1 | 2025 | uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABs · ICLR 2025 |
Algorithmic game theory and mechanism design
regret minimization |
0.9 | 1 | 2025 | uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABs · ICLR 2025 |
Machine learning › Reinforcement learning
exploration |
0.8 | 2 | 2023 | Towards Minimax Optimal Reward-free Reinforcement Learning in Linear MDPs · ICLR 2023 Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation · ICML 2022 |
Machine learning › Reinforcement learning
actor-critic methods |
0.8 | 1 | 2024 | Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation · ICML 2024 |
Machine learning › Reinforcement learning › safe reinforcement learning › risk-sensitive reinforcement learning
conditional value-at-risk |
0.8 | 1 | 2024 | Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human Feedback · ICLR 2024 |
Machine learning › Reinforcement learning › value-based reinforcement learning
distributional reinforcement learning |
0.8 | 1 | 2024 | Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation · ICML 2024 |
Machine learning › Reinforcement learning › safe reinforcement learning › risk-sensitive reinforcement learning
entropic risk measure |
0.8 | 1 | 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024 |
Machine learning › Reinforcement learning
partially observable reinforcement learning |
0.8 | 1 | 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024 |
Machine learning › Reinforcement learning › reinforcement learning theory
Provably efficient RL |
0.8 | 1 | 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.8 | 1 | 2024 | Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human Feedback · ICLR 2024 |
Machine learning › Learning theory
sample complexity |
0.8 | 1 | 2024 | Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation · ICML 2024 |
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning |
0.8 | 1 | 2024 | Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human Feedback · ICLR 2024 |
Network optimization and economics
delay-constrained scheduling |
0.8 | 1 | 2024 | Multi-User Delay-Constrained Scheduling With Deep Recurrent Reinforcement Learning · IEEE/ACM Trans. Netw. 2024 |
Network optimization and economics
resource allocation |
0.8 | 1 | 2024 | Multi-User Delay-Constrained Scheduling With Deep Recurrent Reinforcement Learning · IEEE/ACM Trans. Netw. 2024 |
Machine learning › Reinforcement learning › unsupervised reinforcement learning
reward-free reinforcement learning |
0.7 | 1 | 2023 | Towards Minimax Optimal Reward-free Reinforcement Learning in Linear MDPs · ICLR 2023 |
Machine learning › Reinforcement learning › function approximation
linear function approximation |
0.6 | 1 | 2022 | Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation · ICML 2022 |
Machine learning › Reinforcement learning
regret minimization |
0.6 | 1 | 2022 | Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation · ICML 2022 |
Machine learning › Reinforcement learning
reinforcement learning theory |
0.6 | 1 | 2022 | Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation · ICML 2022 |
Mathematical optimization
online optimization |
0.3 | 1 | 2025 | uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABs · ICLR 2025 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.2 | 1 | 2024 | Multi-User Delay-Constrained Scheduling With Deep Recurrent Reinforcement Learning · IEEE/ACM Trans. Netw. 2024 |
Cellular and mobile networks
multiuser scheduling |
0.2 | 1 | 2024 | Multi-User Delay-Constrained Scheduling With Deep Recurrent Reinforcement Learning · IEEE/ACM Trans. Netw. 2024 |
Machine learning › Reinforcement learning › bandit
upper confidence bound |
0.2 | 1 | 2022 | Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
deep reinforcement learning · 1.5stochastic differential equation · 0.9skipping-clipping · 0.9ordinary differential equation · 0.9log-barrier analysis · 0.9adaptive learning rate scheduling · 0.9recurrent neural network · 0.8maximum likelihood estimation · 0.8least squares regression · 0.8lagrangian dual · 0.8function approximation · 0.8dual gradient descent · 0.8change-of-measure · 0.8beta vectors · 0.8augmented MDP · 0.8POMDP · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABsabstractIn this paper, we present a novel algorithm, `uniINF`, for the Heavy-Tailed Multi-Armed Bandits (HTMAB) problem, demonstrating robustness and adaptability in both stochastic and adversarial environments. Unlike the stochastic MAB setting where loss distributions are stationary with time, our study extends to the adversarial setup, where losses are generated from heavy-tailed distributions that depend on both arms and time. Our novel algorithm `uniINF` enjoys the so-called Best-of-Both-Worlds (BoBW) property, performing optimally in both stochastic and adversarial environments *without* knowing the exact environment type. Moreover, our algorithm also possesses a Parameter-Free feature, *i.e.*, it operates *without* the need of knowing the heavy-tail parameters $(\sigma, \alpha)$ a-priori.
To be precise, `uniINF` ensures nearly-optimal regret in both stochastic and adversarial environments, matching the corresponding lower bounds when $(\sigma, \alpha)$ is known (up to logarithmic factors). To our knowledge, `uniINF` is the first parameter-free algorithm to achieve the BoBW property for the heavy-tailed MAB problem. Technically, we develop innovative techniques to achieve BoBW guarantees for Parameter-Free HTMABs, including a refined analysis for the dynamics of log-barrier, an auto-balancing learning rate scheduling scheme, an adaptive skipping-clipping loss tuning technique, and a stopping-time analysis for logarithmic regret. Yu Chen 0074, Jiatai Huang, Yan Dai 0002, Longbo Huang |
ICLR | 1 |
| 2025 | Finite-Time Analysis of Discrete-Time Stochastic InterpolantsabstractThe stochastic interpolant framework offers a powerful approach for constructing generative models based on ordinary differential equations (ODEs) or stochastic differential equations (SDEs) to transform arbitrary data distributions. However, prior analyses of this framework have primarily focused on the continuous-time setting, assuming perfect solution of the underlying equations. In this work, we present the first discrete-time analysis of the stochastic interpolant framework, where we introduce a innovative discrete-time sampler and derive a finite-time upper bound on its distribution estimation error. Our result provides a novel quantification on how different factors, including
the distance between source and target distributions and estimation accuracy, affect the convergence rate and also offers a new principled way to design efficient schedule for convergence acceleration. Finally, numerical experiments are conducted on the discrete-time sampler to corroborate our theoretical findings. Yu Chen 0074, Longbo Huang |
ICML | 2 |
| 2024 | Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human FeedbackabstractRisk-sensitive reinforcement learning (RL) aims to optimize policies that balance the expected reward and risk. In this paper, we present a novel risk-sensitive RL framework that employs an Iterated Conditional Value-at-Risk (CVaR) objective under both linear and general function approximations, enriched by human feedback. These new formulations provide a principled way to guarantee safety in each decision making step throughout the control process. Moreover, integrating human feedback into risk-sensitive RL framework bridges the gap between algorithmic decision-making and human participation, allowing us to also guarantee safety for human-in-the-loop systems. We propose provably sample-efficient algorithms for this Iterated CVaR RL and provide rigorous theoretical analysis. Furthermore, we establish a matching lower bound to corroborate the optimality of our algorithms in a linear context. Yu Chen 0074, Yihan Du, Pihe Hu, Siwei Wang 0002, Desheng Dash Wu, Longbo Huang |
ICLR | 1 |
| 2024 | Provable Risk-Sensitive Distributional Reinforcement Learning with General Function ApproximationabstractIn the realm of reinforcement learning (RL), accounting for risk is crucial for making decisions under uncertainty, particularly in applications where safety and reliability are paramount. In this paper, we introduce a general framework on Risk-Sensitive Distributional Reinforcement Learning (RS-DisRL), with static Lipschitz Risk Measures (LRM) and general function approximation. Our framework covers a broad class of risk-sensitive RL, and facilitates analysis of the impact of estimation functions on the effectiveness of RSRL strategies and evaluation of their sample complexity. We design two innovative meta-algorithms: RS-DisRL-M, a model-based strategy for model-based function approximation, and RS-DisRL-V, a model-free approach for general value function approximation. With our novel estimation techniques via Least Squares Regression (LSR) and Maximum Likelihood Estimation (MLE) in distributional RL with augmented Markov Decision Process (MDP), we derive the first $\widetilde{\mathcal{O}}(\sqrt{K})$ dependency of the regret upper bound for RSRL with static LRM, marking a pioneering contribution towards statistically efficient algorithms in this domain. Yu Chen 0074, Xiangcheng Zhang, Siwei Wang 0002, Longbo Huang |
ICML | 1 |
| 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight ObservationabstractThis work pioneers regret analysis of risk-sensitive reinforcement learning in partially observable environments with hindsight observation, addressing a gap in theoretical exploration. We introduce a novel formulation that integrates hindsight observations into a Partially Observable Markov Decision Process (POMDP) framework, where the goal is to optimize accumulated reward under the entropic risk measure. We develop the first provably efficient RL algorithm tailored for this setting. We also prove by rigorous analysis that our algorithm achieves polynomial regret $\tilde{O}\left(\frac{e^{|{\gamma}|H}-1}{|{\gamma}|H}H^2\sqrt{KHS^2OA}\right)$, which outperforms or matches existing upper bounds when the model degenerates to risk-neutral or fully observable settings. We adopt the method of change-of-measure and develop a novel analytical tool of beta vectors to streamline mathematical derivations. These techniques are of particular interest to the theoretical study of reinforcement learning. Tonghe Zhang, Yu Chen 0074, Longbo Huang |
ICML | 2 |
| 2024 | Multi-User Delay-Constrained Scheduling With Deep Recurrent Reinforcement LearningabstractMulti-user delay-constrained scheduling is a crucial challenge in various real-world applications, such as wireless communication, live streaming, and cloud computing. The scheduler must make real-time decisions to guarantee both delay and resource constraints simultaneously, without prior information on system dynamics that can be time-varying and challenging to estimate. Additionally, many practical scenarios suffer from partial observability issues due to sensing noise or hidden correlation. To address these challenges, we propose a deep reinforcement learning (DRL) algorithm called Recurrent Softmax Delayed Deep Double Deterministic Policy Gradient ($\mathtt{RSD4}$) (https://github.com/hupihe/RSD4), which is a data-driven method based on a Partially Observed Markov Decision Process (POMDP) formulation.$\mathtt{RSD4}$guarantees resource and delay constraints by Lagrangian dual and delay-sensitive queues, respectively. It also efficiently handles partial observability with a memory mechanism enabled by the recurrent neural network (RNN). Moreover, it introduces user-level decomposition and node-level merging to support large-scale multihop scenarios. Extensive experiments on simulated and real-world datasets demonstrate that$\mathtt{RSD4}$is robust to system dynamics and partially observable environments and achieves superior performance over existing methods. Pihe Hu, Yu Chen 0074, Ling Pan, Zhixuan Fang, Fu Xiao 0001, Longbo Huang |
IEEE/ACM Trans. Netw. | 2 |
| 2023 | Towards Minimax Optimal Reward-free Reinforcement Learning in Linear MDPs
Pihe Hu, Yu Chen 0074, Longbo Huang |
ICLR | 2 |
| 2022 | Nearly Minimax Optimal Reinforcement Learning with Linear Function ApproximationabstractWe study reinforcement learning with linear function approximation where the transition probability and reward functions are linear with respect to a feature mapping $\boldsymbol{\phi}(s,a)$. Specifically, we consider the episodic inhomogeneous linear Markov Decision Process (MDP), and propose a novel computation-efficient algorithm, LSVI-UCB$^+$, which achieves an $\widetilde{O}(Hd\sqrt{T})$ regret bound where $H$ is the episode length, $d$ is the feature dimension, and $T$ is the number of steps. LSVI-UCB$^+$ builds on weighted ridge regression and upper confidence value iteration with a Bernstein-type exploration bonus. Our statistical results are obtained with novel analytical tools, including a new Bernstein self-normalized bound with conservatism on elliptical potentials, and refined analysis of the correction term. To the best of our knowledge, this is the first minimax optimal algorithm for linear MDPs up to logarithmic factors, which closes the $\sqrt{Hd}$ gap between the best known upper bound of $\widetilde{O}(\sqrt{H^3d^3T})$ in \cite{jin2020provably} and lower bound of $\Omega(Hd\sqrt{T})$ for linear MDPs. Pihe Hu, Yu Chen 0074, Longbo Huang |
ICML | 2 |
| 2022 | Effective multi-user delay-constrained scheduling with deep recurrent reinforcement learningabstractMulti-user delay constrained scheduling is important in many real-world applications including wireless communication, live streaming, and cloud computing. Yet, it poses a critical challenge since the scheduler needs to make real-time decisions to guarantee the delay and resource constraints simultaneously without prior information of system dynamics, which can be time-varying and hard to estimate. Moreover, many practical scenarios suffer from partial observability issues, e.g., due to sensing noise or hidden correlation. To tackle these challenges, we propose a deep reinforcement learning (DRL) algorithm, named Recurrent Softmax Delayed Deep Double Deterministic Policy Gradient (RSD4)1, which is a data-driven method based on a Partially Observed Markov Decision Process (POMDP) formulation. RSD4 guarantees resource and delay constraints by Lagrangian dual and delay-sensitive queues, respectively. It also efficiently tackles partial observability with a memory mechanism enabled by the recurrent neural network (RNN) and introduces user-level decomposition and node-level merging to ensure scalability. Extensive experiments on simulated/real-world datasets demonstrate that RSD4 is robust to system dynamics and partially observable environments, and achieves superior performances over existing DRL and non-DRL-based methods. Pihe Hu, Ling Pan, Yu Chen 0074, Zhixuan Fang, Longbo Huang |
MobiHoc | 3 |
| 2019 | Catching Escapers: A Detection Method for Advanced Persistent Escapers in Industry Internet of Things Based on Identity-based Broadcast Encryption (IBBE)abstractAs the Industry 4.0 or Internet of Things (IoT) era begins, security plays a key role in the Industry Internet of Things (IIoT) due to various threats, which include escape or Distributed Denial of Service (DDoS) attackers in the virtualization layer and vulnerability exploiters in the device layer. A successful cross-VM escape attack in the virtualization layer combined with cross-layer penetration in the device layer, which we define as an Advanced Persistent Escaper (APE), poses a great threat. Therefore, the development of detection and rejection methods for APEs across multiple layers in IIoT is an open issue. To the best of our knowledge, less effective methods are established, especially for vulnerability exploitation in the virtualization layer and backdoor leverage in the device layer. On the basis of this, we propose Escaper Cops (EscaperCOP), a detection method for cross-VM escapers in the virtualization layer and cross-layer penetrators in the device layer. In particular, a new detection method for guest-to-host escapers is proposed for the virtualization layer. Finally, a novel encryption method based on Identity-based Broadcast Encryption (IBBE) is proposed to protect the critical components in EscaperCOP, detection library, and control command library. To verify our method, experimental tests are performed for a large number of APEs in an IIoT framework. The test results have demonstrated the proposed method is effective with an acceptable level of detection ratio. Letian Sha, Fu Xiao 0001, Haiping Huang, Yu Chen 0074, Ruchuan Wang 0001 |
ACM Trans. Embed. Comput. Syst. | 4 |