VLDB 2026 Research / reviewers in the wild / expert
Yuankun Jiang
dblp:296/8284
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-2863-3646ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MERINA+: Improving Generalization for Neural Video Adaptation via Information-Theoretic Meta-Reinforcement LearningabstractAdaptive bitrate (ABR) streaming is a popular technique used to improve the quality of experience (QoE) for users who watch videos online, which, for example, can provide a smoother video playback by dynamically adjusting the requested video quality with associated bitrate according to the constrained yet diverse network conditions. Recently, learning-based ABR algorithms have achieved a notable performance gain with lower inference overhead than the conventional heuristic or model-based baselines. However, their performance may degrade significantly in an unseen network environment with time-varying and heterogeneous throughput dynamics. For a better generalization, in this paper, we propose a meta-reinforcement learning (meta-RL)-based neural ABR algorithm that is able to quickly adapt its policy to these unseen throughput dynamics. Specifically, we propose a model-free system framework comprising an inference network and a policy network. The inference network infers distribution of the latent representation for underlying dynamics based on the recent throughout context, while the policy network is trained to quickly adapt to the changing throughout dynamics with the sampled latent representation. To effectively learn the inference network and meta-policy on mixed dynamics of the practical ABR scenarios, we further design a variational information bottleneck theory-based loss function for training the inference and policy networks, whose objective is to strike a trade-off between brevity of the latent representation and expressiveness of the meta-policy. We also derive a theoretically necessary condition for the bitrate versions that yield higher long-term QoE, based on which a dynamic action pruning strategy is further developed for practical implementation. This pruning strategy can not only prevent unsafe policy outputs in midst of unseen throughput dynamics, but may also reduce the computational complexity of model-based ABR algorithms. Finally, the meta-training and meta-adaptation procedures of our proposed algorithm are implemented across a range of throughput dynamics. The empirical evaluations on various datasets containing real-world network traces verify that our algorithm surpasses the state-of-the-art ABR algorithms, particularly in terms of the average chunk QoE and fast adaptation across out-of-distribution throughput traces. Nuowen Kan, Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong, Laura Toni |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Stabilizing and Accelerating Autofocus with Expert Trajectory Regularized Deep Reinforcement LearningabstractAutofocus is a crucial component of modern digital cameras. While recent learning-based methods achieve state-of-the-art in focus prediction accuracy, they unfortunately ignore the potential focus hunting phenomenon of back-and-forth lens movement in the multi-step focusing procedure. To address this, in this paper, we propose an expert regularized deep reinforcement learning (DRL)-based approach for autofocus, which can utilize the sequential information of lens movement trajectory to both enhance the multi-step in-focus prediction accuracy and reduce the chance of focus hunting. Our method generally follows an actor-critic framework. To accelerate the DRL’s training with a higher sample efficiency, we initialize the policy with a pre-trained single-step prediction network, where the network is further improved by modifying the output of absolute in-focus position distribution to the relative lens movement distribution to establish a better mapping between input images and lens movement. To further stabilize DRL’s training with a lower occurrence of focus hunting in the resulting lens movement trajectory, we generate some offline trajectories based on prior knowledge to avoid focus hunting, which are then leveraged as an offline dataset of expert trajectories to regularize the actor network’s training. Empirical evaluations show that our method outperforms those learning-based methods on public benchmarks, with higher single- and multi-step prediction accuracy, and a significant reduction of focus hunting rate. Shouhang Zhu, Yuankun Jiang, Nuowen Kan, Wenrui Dai, Junni Zou, Hongkai Xiong |
CVPR | 3 |
| 2024 | Adversarially Robust Neural Lyapunov ControlabstractState-of-the-art learning-based stability control methods for nonlinear robotic systems suffer from the issue of reality gap, which stems from discrepancy of the system dynamics between training and target (test) environments. To mitigate this gap, we propose an adversarially robust neural Lyapunov control (ARNLC) method to improve the robustness and generalization capabilities for Lyapunov theory-based stability control. Specifically, inspired by adversarial learning, we introduce an adversary to simulate the dynamics discrepancy, which is learned through deep reinforcement learning to generate the worst-case perturbations during the controller’s training. By alternatively updating the controller to minimize the perturbed Lyapunov risk and the adversary to deviate the controller from its objective, the learned control policy enjoys a theoretical guarantee of stability. Empirical evaluations on five stability control tasks with the uniform and worst-case perturbations demonstrate that ARNLC not only accelerates the convergence to asymptotic stability, but can generalize better in the entire perturbation space. Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong |
ECAI | 2 |
| 2024 | Deep Reinforcement Learning-Based Camera Autofocus with Gaussian Process RegressionabstractAutofocus (AF) is a fundamental task in digital image capturing, yet current approaches often exhibit a poor performance in terms of speed or accuracy. In this paper, we propose a deep reinforcement learning (DRL)-based framework for camera AF, by incorporating a pre-trained MobileNet-v2 encoder for image feature extraction and a Gaussian process regression (GPR) module for in-focus prediction. Specially, we formulate AF as a multi-step process, in which the lens movement trajectory is dynamically generated from a DRL-based policy. By leveraging image features extracted from the MobileNet-v2 encoder, the DRL-based policy can efficiently get close to the in-focus position, which enables the focus search in a fast speed, thus overcoming the drawback of slow focusing procedure suffered by traditional search-based AF methods. A GPR-based in-focus predictor is further designed to terminate the DRL’s exploration once it is able to provide a reliable prediction on the in-focus position, by exploiting the sharpness metric values of image patches captured at the DRL-generated trajectory of lens positions. In further contrast with the one-step prediction-based methods, we leverage the sequential information brought additionally by the multi-step exploration, enabling to gain a higher prediction accuracy. Experimental results demonstrate that our approach provides a significant improvement compared to both the search-based and prediction-based AF baselines. Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong |
VCIP | 2 |
| 2024 | Variance Reduced Domain Randomization for Reinforcement Learning With Policy GradientabstractBy introducing randomness on the environments, domain randomization (DR) imposes diversity to the policy training of deep reinforcement learning, and thus improves its capability of generalization. The randomization of environments, however, introduces another source of variability for the estimate of policy gradients, in addition to the already high variance incurred by trajectory sampling. Therefore, with standard state-dependent baselines, the policy gradient methods may still suffer high variance, causing a low sample efficiency during the training of DR. In this paper, we theoretically derive a bias-free and state/environment-dependent optimal baseline for DR, and analytically show its ability to achieve further variance reduction over the standard constant and state-dependent baselines for DR. Based on our theory, we further propose a variance reduced domain randomization (VRDR) approach for policy gradient methods, to strike a tradeoff between the variance reduction and computational complexity for the practical implementation. By dividing the entire space of environments into some subspaces and then estimating the state/subspace-dependent baseline, VRDR enjoys a theoretical guarantee of variance reduction and faster convergence than the state-dependent baselines. Empirical evaluations on six robot control tasks with randomized dynamics demonstrate that VRDR not only accelerates the convergence of policy training, but can consistently achieve a better eventual policy with improved training stability. Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Successor Feature-Based Transfer Reinforcement Learning for Video Rate Adaptation With Heterogeneous QoE PreferencesabstractIn adaptive video streaming, the design of an adaptive bitrate (ABR) strategy is critical for the quality-of-experience (QoE) perceived by users. Though current learning-based ABR algorithms achieve state-of-the-art performance for users with a given QoE metric setting for training, they may unfortunately suffer the poor generalization issue for other users with different QoE preferences. Besides, how to quantitatively characterize the distinct QoE preference for a user has also not been extensively studied yet. In this paper, we propose STEER, a successor feature-based transfer reinforcement learning framework for fast learning the ABR strategies on heterogeneous QoE preferences. Specifically, we first develop a QoE preference analysis scheme to infer the personal QoE preference of a single user based on the user's actual viewing history. We then formulate the personalized QoE maximization problem as a reinforcement learning (RL) task, which optimizes the ABR strategy to maximize the overall QoE perceived by the user. Further, we model the QoE maximization problem for multiple users with heterogeneous QoE preferences as a multi-task RL problem, with each task distinguished by the user-distinct QoE preference. To efficiently address this problem, the proposed STEER solves for each RL-based ABR task by learning its optimal successor feature (SF) function, which can be exploited as shared knowledge across tasks to facilitate the transfer between tasks. With SF functions, STEER can quickly evaluate the optimal policies of previously learned tasks on a new task, and further use the generalized policy improvement operation to obtain a jumpstart policy. Both theoretically and empirically, we show that this jumpstart policy is a good initial policy with a performance guarantee for better generalization in the new task, and can also lead to a faster convergence to the optimal policy of the new task. Kexin Tang, Nuowen Kan, Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong |
IEEE Trans. Multim. | 3 |
| 2023 | Doubly Robust Augmented Transfer for Meta-Reinforcement LearningabstractMeta-reinforcement learning (Meta-RL), though enabling a fast adaptation to learn new skills by exploiting the common structure shared among different tasks, suffers performance degradation in the sparse-reward setting. Current hindsight-based sample transfer approaches can alleviate this issue by transferring relabeled trajectories from other tasks to a new task so as to provide informative experience for the target reward function, but are unfortunately constrained with the unrealistic assumption that tasks differ only in reward functions. In this paper, we propose a doubly robust augmented transfer (DRaT) approach, aiming at addressing the more general sparse reward meta-RL scenario with both dynamics mismatches and varying reward functions across tasks. Specifically, we design a doubly robust augmented estimator for efficient value-function evaluation, which tackles dynamics mismatches with the optimal importance weight of transition distributions achieved by minimizing the theoretically derived upper bound of mean squared error (MSE) between the estimated values of transferred samples and their true values in the target task. Due to its intractability, we then propose an interval-based approximation to this optimal importance weight, which is guaranteed to cover the optimum with a constrained and sample-independent upper bound on the MSE approximation error. Based on our theoretical findings, we finally develop a DRaT algorithm for transferring informative samples across tasks during the training of meta-RL. We implement DRaT on an off-policy meta-RL baseline, and empirically show that it significantly outperforms other hindsight-based approaches on various sparse-reward MuJoCo locomotion tasks with varying dynamics and reward functions. Yuankun Jiang, Nuowen Kan, Wenrui Dai, Junni Zou, Hongkai Xiong |
NeurIPS | 1 |
| 2022 | Adaptive Task Sampling and Variance Reduction for Gradient-Based Meta-Learning
Zhuoqun Liu, Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong |
BMVC | 2 |
| 2022 | Improving Generalization for Neural Adaptive Video Streaming via Meta Reinforcement LearningabstractIn this paper, we present a meta reinforcement learning (Meta-RL)-based neural adaptive bitrate streaming (ABR) algorithm that is able to rapidly adapt its control policy to the changing network throughput dynamics. Specifically, to allow rapid adaptation, we discuss the necessity of detaching the inference of throughput dynamics with the universal control mechanism that is in essence shared by all potential throughput dynamics for neural ABR algorithms. To meta-learn the ABR policy, we then build up a model-free system framework, composed of a probabilistic latent encoder that infers the underlying dynamics from the recent throughput context, and a policy network that is conditioned on latent variable and learns to quickly adapt to new environments. Additionally, to address the difficulties caused by training the policy on mixed dynamics, on-policy RL (or imitation learning) algorithms are suggested for policy training, with a mutual information-based regularization to make the latent variable more informative about the policy. Finally, we implement our algorithm's meta-training and meta-adaptation procedures under a variety of throughput dynamics. Empirical evaluations on different QoE metrics and multiple datasets containing real-world network traces demonstrate that our algorithm outperforms state-of-the-art ABR algorithms, in terms of the performance on the average chunk QoE, consistency and fast adaptation across a wide range of throughput patterns. Nuowen Kan, Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong |
ACM Multimedia | 2 |
| 2021 | Monotonic Robust Policy Optimization with Model DiscrepancyabstractState-of-the-art deep reinforcement learning (DRL) algorithms tend to overfit due to the model discrepancy between source and target environments. Though applying domain randomization during training can improve the average performance by randomly generating a sufficient diversity of environments in simulator, the worst-case environment is still neglected without any performance guarantee. Since the average and worst-case performance are both important for generalization in RL, in this paper, we propose a policy optimization approach for concurrently improving the policy’s performance in the average and worst-case environment. We theoretically derive a lower bound for the worst-case performance of a given policy by relating it to the expected performance. Guided by this lower bound, we formulate an optimization problem to jointly optimize the policy and sampling distribution, and prove that by iteratively solving it the worst-case performance is monotonically improved. We then develop a practical algorithm, named monotonic robust policy optimization (MRPO). Experimental evaluations in several robot control tasks demonstrate that MRPO can generally improve both the average and worst-case performance in the source environments for training, and facilitate in all cases the learned policy with a better generalization capability in some unseen testing environments. Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong |
ICML | 1 |