Xiaoteng Ma

dblp:238/3249 · DBLP profile ↗
← Back
28ranked-venue papers
8as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 5 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Computer networks · 3 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Word to World: Can Large Language Models be Implicit Text-based World Models?
abstract
Yixia Li, Hongru Wang, Jiahao Qiu, Zhenfei Yin, Dongdong Zhang, Cheng Qian, Zeping Li, Xiaoteng Ma, Guanhua Chen, Heng Ji. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yixia Li, Hongru Wang 0003, Jiahao Qiu, Zhenfei Yin, Dongdong Zhang 0001, Cheng Qian 0008, Zeping Li, Xiaoteng Ma, Guanhua Chen 0001, Heng Ji 0001
ACL (1)8
2025 Episodic Novelty Through Temporal Distance
abstract
Exploration in sparse reward environments remains a significant challenge in reinforcement learning, particularly in Contextual Markov Decision Processes (CMDPs), where environments differ across episodes. Existing episodic intrinsic motivation methods for CMDPs primarily rely on count-based approaches, which are ineffective in large state spaces, or on similarity-based methods that lack appropriate metrics for state comparison. To address these shortcomings, we propose Episodic Novelty Through Temporal Distance (ETD), a novel approach that introduces temporal distance as a robust metric for state similarity and intrinsic reward computation. By employing contrastive learning, ETD accurately estimates temporal distances and derives intrinsic rewards based on the novelty of states within the current episode. Extensive experiments on various benchmark tasks demonstrate that ETD significantly outperforms state-of-the-art methods, highlighting its effectiveness in enhancing exploration in sparse reward CMDPs.
Yuhua Jiang, Qihan Liu, Yiqin Yang, Xiaoteng Ma, Dianyu Zhong, Hao Hu 0006, Jun Yang 0028, Bin Liang 0001, Bo Xu 0002, Chongjie Zhang, Qianchuan Zhao
ICLR4
2025 Cross-Domain Offline Policy Adaptation with Optimal Transport and Dataset Constraint
abstract
We explore cross-domain offline reinforcement learning (RL) where offline datasets from another domain can be accessed to facilitate policy learning. However, the underlying environments of the two datasets may have dynamics mismatches, incurring inferior performance when simply merging the data of two domains. Existing methods mitigate this issue by training domain classifiers, using contrastive learning methods, etc. Nevertheless, they still rely on a large amount of target domain data to function well. Instead, we address this problem by establishing a concrete performance bound of a policy given datasets from two domains. Motivated by the theoretical insights, we propose to align transitions in the two datasets using optimal transport and selectively share source domain samples, without training any neural networks. This enables reliable data filtering even given a few target domain data. Additionally, we introduce a dataset regularization term that ensures the learned policy remains within the scope of the target domain dataset, preventing it from being biased towards the source domain data. Consequently, we propose the Optimal Transport Data Filtering (dubbed OTDF) method and examine its effectiveness by conducting extensive experiments across various dynamics shift conditions (e.g., gravity shift), given limited target domain data. It turns out that OTDF exhibits superior performance on many tasks and dataset qualities, often surpassing prior strong baselines by a large margin.
Jiafei Lyu, Mengbei Yan, Zhongjian Qiao, Runze Liu 0002, Xiaoteng Ma, Deheng Ye, Zongqing Lu 0002, Xiu Li 0001
ICLR5
2025 DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
abstract
We present Distributional Soft Actor-Critic (DSAC), a distributional reinforcement learning (RL) algorithm that combines the strengths of distributional information of accumulated rewards and entropy-driven exploration from Soft Actor-Critic (SAC) algorithm. DSAC models the randomness in both action and rewards, surpassing baseline performances on various continuous control tasks. Unlike standard approaches that solely maximize expected rewards, we propose a unified framework for risk-sensitive learning, one that optimizes the risk-related objective while balancing entropy to encourage exploration. Extensive experiments demonstrate DSAC’s effectiveness in enhancing agent performances for both risk-neutral and risk-sensitive control tasks.
Xiaoteng Ma, Junyao Chen, Jun Yang 0028, Qianchuan Zhao, Zhengyuan Zhou
J. Artif. Intell. Res.1
2025 CVaR-Constrained Policy Optimization for Safe Reinforcement Learning
abstract
Current constrained reinforcement learning (RL) methods guarantee constraint satisfaction only in expectation, which is inadequate for safety-critical decision problems. Since a constraint satisfied in expectation remains a high probability of exceeding the cost threshold, solving constrained RL problems with high probabilities of satisfaction is critical for RL safety. In this work, we consider the safety criterion as a constraint on the conditional value-at-risk (CVaR) of cumulative costs, and propose the CVaR-constrained policy optimization algorithm (CVaR-CPO) to maximize the expected return while ensuring agents pay attention to the upper tail of constraint costs. According to the bound on the CVaR-related performance between two policies, we first reformulate the CVaR-constrained problem in augmented state space using the state extension procedure and the trust-region method. CVaR-CPO then derives the optimal update policy by applying the Lagrangian method to the constrained optimization problem. In addition, CVaR-CPO utilizes the distribution of constraint costs to provide an efficient quantile-based estimation of the CVaR-related value function. We conduct experiments on constrained control tasks to show that the proposed method can produce behaviors that satisfy safety constraints, and achieve comparable performance to most safe RL (SRL) methods.
Shu Leng, Xiaoteng Ma, Qihan Liu, Xueqian Wang 0001, Bin Liang 0001, Yu Liu 0036, Jun Yang 0028
IEEE Trans. Neural Networks Learn. Syst.3
2024 Learning Diverse Risk Preferences in Population-Based Self-Play
abstract
Among the remarkable successes of Reinforcement Learning (RL), self-play algorithms have played a crucial role in solving competitive games. However, current self-play RL methods commonly optimize the agent to maximize the expected win-rates against its current or historical copies, resulting in a limited strategy style and a tendency to get stuck in local optima. To address this limitation, it is important to improve the diversity of policies, allowing the agent to break stalemates and enhance its robustness when facing with different opponents. In this paper, we present a novel perspective to promote diversity by considering that agents could have diverse risk preferences in the face of uncertainty. To achieve this, we introduce a novel reinforcement learning algorithm called Risk-sensitive Proximal Policy Optimization (RPPO), which smoothly interpolates between worst-case and best-case policy learning, enabling policy learning with desired risk preferences. Furthermore, by seamlessly integrating RPPO with population-based self-play, agents in the population optimize dynamic risk-sensitive objectives using experiences gained from playing against diverse opponents. Our empirical results demonstrate that our method achieves comparable or superior performance in competitive games and, importantly, leads to the emergence of diverse behavioral modes. Code is available at https://github.com/Jackory/RPBT.
Yuhua Jiang, Qihan Liu, Xiaoteng Ma, Chenghao Li 0002, Yiqin Yang, Jun Yang 0028, Bin Liang 0001, Qianchuan Zhao
AAAI3
2024 Efficient Multi-agent Reinforcement Learning by Planning
abstract
Multi-agent reinforcement learning (MARL) algorithms have accomplished remarkable breakthroughs in solving large-scale decision-making tasks. Nonetheless, most existing MARL algorithms are model-free, limiting sample efficiency and hindering their applicability in more challenging scenarios. In contrast, model-based reinforcement learning (MBRL), particularly algorithms integrating planning, such as MuZero, has demonstrated superhuman performance with limited data in many tasks. Hence, we aim to boost the sample efficiency of MARL by adopting model-based approaches. However, incorporating planning and search methods into multi-agent systems poses significant challenges. The expansive action space of multi-agent systems often necessitates leveraging the nearly-independent property of agents to accelerate learning. To tackle this issue, we propose the MAZero algorithm, which combines a centralized model with Monte Carlo Tree Search (MCTS) for policy search. We design an ingenious network structure to facilitate distributed execution and parameter sharing. To enhance search efficiency in deterministic environments with sizable action spaces, we introduce two novel techniques: Optimistic Search Lambda (OS($\lambda$)) and Advantage-Weighted Policy Optimization (AWPO). Extensive experiments on the SMAC benchmark demonstrate that MAZero outperforms model-free approaches in terms of sample efficiency and provides comparable or better performance than existing model-based methods in terms of both sample and computational efficiency.
Qihan Liu, Jianing Ye, Xiaoteng Ma, Jun Yang 0028, Bin Liang 0001, Chongjie Zhang
ICLR3
2024 SEABO: A Simple Search-Based Method for Offline Imitation Learning
abstract
Offline reinforcement learning (RL) has attracted much attention due to its ability in learning from static offline datasets and eliminating the need of interacting with the environment. Nevertheless, the success of offline RL relies heavily on the offline transitions annotated with reward labels. In practice, we often need to hand-craft the reward function, which is sometimes difficult, labor-intensive, or inefficient. To tackle this challenge, we set our focus on the offline imitation learning (IL) setting, and aim at getting a reward function based on the expert data and unlabeled data. To that end, we propose a simple yet effective search-based offline IL method, tagged SEABO. SEABO allocates a larger reward to the transition that is close to its closest neighbor in the expert demonstration, and a smaller reward otherwise, all in an unsupervised learning manner. Experimental results on a variety of D4RL datasets indicate that SEABO can achieve competitive performance to offline RL algorithms with ground-truth rewards, given only a single expert trajectory, and can outperform prior reward learning and offline IL methods across many tasks. Moreover, we demonstrate that SEABO also works well if the expert demonstrations contain only observations. Our code is publicly available at https://github.com/dmksjfl/SEABO.
Jiafei Lyu, Xiaoteng Ma, Le Wan, Runze Liu 0002, Xiu Li 0001, Zongqing Lu 0002
ICLR2
2024 Single-Trajectory Distributionally Robust Reinforcement Learning
abstract
To mitigate the limitation that the classical reinforcement learning (RL) framework heavily relies on identical training and test environments, Distributionally Robust RL (DRRL) has been proposed to enhance performance across a range of environments, possibly including unknown test environments. As a price for robustness gain, DRRL involves optimizing over a set of distributions, which is inherently more challenging than optimizing over a fixed distribution in the non-robust case. Existing DRRL algorithms are either model-based or fail to learn from a single sample trajectory. In this paper, we design a first fully model-free DRRL algorithm, called distributionally robust Q-learning with single trajectory (DRQ). We delicately design a multi-timescale framework to fully utilize each incrementally arriving sample and directly learn the optimal distributionally robust policy without modeling the environment, thus the algorithm can be trained along a single trajectory in a model-free fashion. Despite the algorithm’s complexity, we provide asymptotic convergence guarantees by generalizing classical stochastic approximation tools.Comprehensive experimental results demonstrate the superior robustness and sample complexity of our proposed algorithm, compared to non-robust methods and other robust RL algorithms.
Xiaoteng Ma, Jose H. Blanchet, Jun Yang 0028, Jiheng Zhang, Zhengyuan Zhou
ICML2
2024 Smart Data-Driven Proactive Push to Edge Network for User-Generated Videos
abstract
To reduce costs and improve performance, video Content Delivery Networks (CDNs) have started to incorporate lightweight edge nodes, e.g., WiFi access points. Because of this, it is necessary for CDNs to intelligently select which video files should be placed at their core data centers vs. these edge nodes. This is more complex than traditional CDN management, as lightweight edge nodes are much more numerous and unstable than data centers. With this in mind, we present SDPush —- a system for managing content placement in edge CDNs. SDPush tackles two problems. First, it is necessary for SDPush to select which files to proactive push. To address this, we build a file popularity prediction model that effectively identifies video files that will receive many views. Second, SDPush should determine how many replicas of each file to push. To address this, we design a model to predict the benefits of pushing particular files (regarding traffic savings) and then formulate the replica decision problem as a lightweight problem, which is solvable within seconds, even for platforms that accommodate millions of daily active users. Through a trace-driven evaluation and a live deployment on a real video platform, we validate SDPush’s effectiveness, offloading peak-period traffic by 12.1% to 23.9% from the data center to edge nodes, thereby reducing the CDN costs.
Xiaoteng Ma, Qing Li 0006, Junkun Peng, Gareth Tyson, Ziwen Ye, Shisong Tang, Shengbin Meng, Gabriel-Miro Muntean
INFOCOM1
2024 NeuralPlane: An Efficiently Parallelizable Platform for Fixed-wing Aircraft Control with Reinforcement Learning
abstract
Reinforcement learning (RL) demonstrates superior potential over traditional flight control methods for fixed-wing aircraft, particularly under extreme operational conditions. However, the high demand for training samples and the lack of efficient computation in existing simulators hinder its further application. In this paper, we introduce NeuralPlane, the first benchmark platform for large-scale parallel simulations of fixed-wing aircraft. NeuralPlane significantly boosts high-fidelity simulation via GPU-accelerated Flight Dynamics Model (FDM) computation, achieving a single-step simulation time of just 0.2 seconds at a parallel scale of $10^{6}$, far exceeding current platforms. We also provide clear code templates, comprehensive evaluation/visualization tools and hierarchical frameworks for integrating RL and traditional control methods. We believe that NeuralPlane can accelerate the development of RL-based fixed-wing flight control and serve as a new challenging benchmark for the RL community. Our NeuralPlane is open-source and accessible at https://github.com/xuecy22/NeuralPlane.
Chuanyi Xue, Qihan Liu, Xiaoteng Ma, Xinyao Qin, Gui Ning, Jinsheng Ren
NeurIPS3
2024 KEPC-Push: A Knowledge-Enhanced Proactive Content Push Strategy for Edge-Assisted Video Feed Streaming
Ziwen Ye, Qing Li 0006, Chunyu Qiao, Xiaoteng Ma, Yong Jiang 0001, Shengbin Meng, Zhenhui Yuan, Zili Meng
USENIX ATC4
2023 Uncertainty-Driven Trajectory Truncation for Data Augmentation in Offline Reinforcement Learning
abstract
Equipped with the trained environmental dynamics, model-based offline reinforcement learning (RL) algorithms can often successfully learn good policies from fixed-sized datasets, even some datasets with poor quality. Unfortunately, however, it can not be guaranteed that the generated samples from the trained dynamics model are reliable (e.g., some synthetic samples may lie outside of the support region of the static dataset). To address this issue, we propose Trajectory Truncation with Uncertainty (TATU), which adaptively truncates the synthetic trajectory if the accumulated uncertainty along the trajectory is too large. We theoretically show the performance bound of TATU to justify its benefits. To empirically show the advantages of TATU, we first combine it with two classical model-based offline RL algorithms, MOPO and COMBO. Furthermore, we integrate TATU with several off-the-shelf model-free offline RL algorithms, e.g., BCQ. Experimental results on the D4RL benchmark show that TATU significantly improves their performance, often by a large margin. Code is available here.
Jiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Jun Yang 0028, Le Wan, Xiu Li 0001
ECAI3
2023 What is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL?
abstract
Offline goal-conditioned RL (GCRL) offers a way to train general-purpose agents from fully offline datasets. In addition to being conservative within the dataset, the generalization ability to achieve unseen goals is another fundamental challenge for offline GCRL. However, to the best of our knowledge, this problem has not been well studied yet. In this paper, we study out-of-distribution (OOD) generalization of offline GCRL both theoretically and empirically to identify factors that are important. In a number of experiments, we observe that weighted imitation learning enjoys better generalization than pessimism-based offline RL method. Based on this insight, we derive a theory for OOD generalization, which characterizes several important design choices. We then propose a new offline GCRL method, Generalizable Offline goAl-condiTioned RL (GOAT), by combining the findings from our theoretical and empirical studies. On a new benchmark containing 9 independent identically distributed (IID) tasks and 17 OOD tasks, GOAT outperforms current state-of-the-art methods by a large margin.
Rui Yang 0010, Lin Yong, Xiaoteng Ma, Hao Hu 0006, Chongjie Zhang, Tong Zhang 0001
ICML3
2023 Mean-Semivariance Policy Optimization via Risk-Averse Reinforcement Learning (Extended Abstract)
abstract
Keeping risk under control is often more crucial than maximizing expected rewards in real-world decision-making situations, such as finance, robotics, autonomous driving, etc. The most natural choice of risk measures is variance, while it penalizes the upside volatility as much as the downside part. Instead, the (downside) semivariance, which captures negative deviation of a random variable under its mean, is more suitable for risk-averse proposes. This paper aims at optimizing the mean-semivariance (MSV) criterion in reinforcement learning w.r.t. steady reward distribution. Since semivariance is time-inconsistent and does not satisfy the standard Bellman equation, the traditional dynamic programming methods are inapplicable to MSV problems directly. To tackle this challenge, we resort to Perturbation Analysis (PA) theory and establish the performance difference formula for MSV. We reveal that the MSV problem can be solved by iteratively solving a sequence of RL problems with a policy-dependent reward function. Further, we propose two on-policy algorithms based on the policy gradient theory and the trust region method. Finally, we conduct diverse experiments from simple bandit problems to continuous control tasks in MuJoCo, which demonstrate the effectiveness of our proposed methods.
Xiaoteng Ma, Qianchuan Zhao
IJCAI1
2023 Cross-Domain Policy Adaptation via Value-Guided Data Filtering
abstract
Generalizing policies across different domains with dynamics mismatch poses a significant challenge in reinforcement learning. For example, a robot learns the policy in a simulator, but when it is deployed in the real world, the dynamics of the environment may be different. Given the source and target domain with dynamics mismatch, we consider the online dynamics adaptation problem, in which case the agent can access sufficient source domain data while online interactions with the target domain are limited. Existing research has attempted to solve the problem from the dynamics discrepancy perspective. In this work, we reveal the limitations of these methods and explore the problem from the value difference perspective via a novel insight on the value consistency across domains. Specifically, we present the Value-Guided Data Filtering (VGDF) algorithm, which selectively shares transitions from the source domain based on the proximity of paired value targets across the two domains. Empirical results on various environments with kinematic and morphology shifts demonstrate that our method achieves superior performance compared to prior approaches.
Chenjia Bai, Xiaoteng Ma, Dong Wang 0028, Bin Zhao 0001, Zhen Wang 0004, Xuelong Li 0001, Wei Li 0235
NeurIPS3
2022 Efficient Continuous Control with Double Actors and Regularized Critics
abstract
How to obtain good value estimation is a critical problem in Reinforcement Learning (RL). Current value estimation methods in continuous control, such as DDPG and TD3, suffer from unnecessary over- or under- estimation. In this paper, we explore the potential of double actors, which has been neglected for a long time, for better value estimation in the continuous setting. First, we interestingly find that double actors improve the exploration ability of the agent. Next, we uncover the bias alleviation property of double actors in handling overestimation with single critic, and underestimation with double critics respectively. Finally, to mitigate the potentially pessimistic value estimate in double critics, we propose to regularize the critics under double actors architecture. Together, we present Double Actors Regularized Critics (DARC) algorithm. Extensive experiments on challenging continuous control benchmarks, MuJoCo and PyBullet, show that DARC significantly outperforms current baselines with higher average return and better sample efficiency.
Jiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu Li 0001
AAAI2
2022 Offline Reinforcement Learning with Value-based Episodic Memory
Xiaoteng Ma, Yiqin Yang, Hao Hu 0006, Jun Yang 0028, Chongjie Zhang, Qianchuan Zhao, Bin Liang 0001, Qihan Liu
ICLR1
2022 Mildly Conservative Q-Learning for Offline Reinforcement Learning
abstract
Offline reinforcement learning (RL) defines the task of learning from a static logged dataset without continually interacting with the environment. The distribution shift between the learned policy and the behavior policy makes it necessary for the value function to stay conservative such that out-of-distribution (OOD) actions will not be severely overestimated. However, existing approaches, penalizing the unseen actions or regularizing with the behavior policy, are too pessimistic, which suppresses the generalization of the value function and hinders the performance improvement. This paper explores mild but enough conservatism for offline learning while not harming generalization. We propose Mildly Conservative Q-learning (MCQ), where OOD actions are actively trained by assigning them proper pseudo Q values. We theoretically show that MCQ induces a policy that behaves at least as well as the behavior policy and no erroneous overestimation will occur for OOD actions. Experimental results on the D4RL benchmarks demonstrate that MCQ achieves remarkable performance compared with prior work. Furthermore, MCQ shows superior generalization ability when transferring from offline to online, and significantly outperforms baselines. Our code is publicly available at https://github.com/dmksjfl/MCQ.
Jiafei Lyu, Xiaoteng Ma, Xiu Li 0001, Zongqing Lu 0002
NeurIPS2
2022 Exploit Reward Shifting in Value-Based Deep-RL: Optimistic Curiosity-Based Exploration and Conservative Exploitation via Linear Reward Shaping
abstract
In this work, we study the simple yet universally applicable case of reward shaping in value-based Deep Reinforcement Learning (DRL). We show that reward shifting in the form of a linear transformation is equivalent to changing the initialization of the $Q$-function in function approximation. Based on such an equivalence, we bring the key insight that a positive reward shifting leads to conservative exploitation, while a negative reward shifting leads to curiosity-driven exploration. Accordingly, conservative exploitation improves offline RL value estimation, and optimistic value estimation improves exploration for online RL. We validate our insight on a range of RL tasks and show its improvement over baselines: (1) In offline RL, the conservative exploitation leads to improved performance based on off-the-shelf algorithms; (2) In online continuous control, multiple value functions with different shifting constants can be used to tackle the exploration-exploitation dilemma for better sample efficiency; (3) In discrete control tasks, a negative reward shifting yields an improvement over the curiosity-based exploration method.
Lei Han 0001, Rui Yang 0010, Xiaoteng Ma, Bolei Zhou
NeurIPS4
2022 RORL: Robust Offline Reinforcement Learning via Conservative Smoothing
abstract
Offline reinforcement learning (RL) provides a promising direction to exploit massive amount of offline data for complex decision-making tasks. Due to the distribution shift issue, current offline RL algorithms are generally designed to be conservative in value estimation and action selection. However, such conservatism can impair the robustness of learned policies when encountering observation deviation under realistic conditions, such as sensor errors and adversarial attacks. To trade off robustness and conservatism, we propose Robust Offline Reinforcement Learning (RORL) with a novel conservative smoothing technique. In RORL, we explicitly introduce regularization on the policy and the value function for states near the dataset, as well as additional conservative value estimation on these states. Theoretically, we show RORL enjoys a tighter suboptimality bound than recent theoretical results in linear MDPs. We demonstrate that RORL can achieve state-of-the-art performance on the general offline RL benchmark and is considerably robust to adversarial observation perturbations.
Rui Yang 0010, Chenjia Bai, Xiaoteng Ma, Zhaoran Wang 0001, Chongjie Zhang, Lei Han 0001
NeurIPS3
2022 MagNet: Cooperative Edge Caching by Automatic Content Congregating
abstract
Nowadays, the surge of Internet contents and the need for high Quality of Experience (QoE) put the backbone network under unprecedented pressure. The emerging edge caching solutions help ease the pressure by caching contents closer to users. However, these solutions suffer from two challenges: 1) a low hit ratio due to edges’ high density and small coverages. 2) unbalanced edges’ workloads caused by dynamic requests and heterogeneous edge capacities. In this paper, we formulate a typical cooperative edge caching problem and propose the MagNet, a decentralized and cooperative edge caching system to address these two challenges. The proposed MagNet system consists of two innovative mechanisms: 1) the Automatic Content Congregating (ACC), which utilizes a neural embedding algorithm to capture underlying patterns of historical traces to cluster contents into some types. The ACC then can guide requests to their optimal edges according to their types so that contents congregate automatically in different edges by type. This process forms a virtuous cycle between edges and requests, driving a high hit ratio. 2) the Mutual Assistance Group (MAG), which lets idle edges share overloaded edges’ workloads by forming temporary groups promptly. To evaluate the performance of MagNet, we conduct experiments to compare it with classical, Machine Learning (ML)-based and cooperative caching solutions using the real-world trace. The results show that the MagNet can improve the hit ratio from 40% and 60% to 75% for non-cooperative and cooperative solutions, respectively, and significantly improve the balance of edges’ workloads.
Junkun Peng, Qing Li 0006, Xiaoteng Ma, Yong Jiang 0001, Yutao Dong, Chuang Hu, Meng Chen 0005
WWW3
2022 Knowledge-based Temporal Fusion Network for Interpretable Online Video Popularity Prediction
abstract
Predicting the popularity of online videos has many real-world applications, such as recommendation, precise advertising, and edge caching strategies. Despite many efforts have been dedicated to the online video popularity prediction, there still exist several challenges: (1) The meta-data from online videos is usually sparse and noisy, which makes it difficult to learn a stable and robust representation. (2) The influence of content features and temporal features in different life cycles of online videos is dynamically changing, so it is necessary to build a model that can capture the dynamics. (3) Besides, there is a great need to interpret the predictive behavior of the model to assist administrators of video platforms in the subsequent decision-making.
Shisong Tang, Qing Li 0006, Xiaoteng Ma, Ci Gao, Dingmin Wang, Yong Jiang 0001, Aoyang Zhang, Hechang Chen
WWW3
2022 Mean-Semivariance Policy Optimization via Risk-Averse Reinforcement Learning
abstract
Keeping risk under control is often more crucial than maximizing expected reward in real-world decision-making situations, such as finance, robotics, autonomous driving, etc. The most natural choice of risk measures is variance, while it penalizes the upside volatility as much as the downside part. Instead, the (downside) semivariance, which captures the negative deviation of a random variable under its mean, is more suitable for risk-averse proposes. This paper aims at optimizing the mean-semivariance (MSV) criterion in reinforcement learning w.r.t. steady rewards. Since semivariance is time-inconsistent and does not satisfy the standard Bellman equation, the traditional dynamic programming methods are inapplicable to MSV problems directly. To tackle this challenge, we resort to the Perturbation Analysis (PA) theory and establish the performance difference formula for MSV. We reveal that the MSV problem can be solved by iteratively solving a sequence of RL problems with a policy-dependent reward function. Further, we propose two on-policy algorithms based on the policy gradient theory and the trust region method. Finally, we conduct diverse experiments from simple bandit problems to continuous control tasks in MuJoCo, which demonstrate the effectiveness of our proposed methods.
Xiaoteng Ma, Qianchuan Zhao
J. Artif. Intell. Res.1
2022 Learning-Based Joint QoE Optimization for Adaptive Video Streaming Based on Smart Edge
abstract
The latest increase in HTTP-based adaptive video streaming over the Internet enables a growing number of clients to compete for a shared bottleneck bandwidth. This competition may affect users’ Quality of Experience (QoE) negatively, especially in terms of fairness and stability. This paper presentsFlex-Steward, a solution that performs multi-client joint QoE optimization for adaptive video streaming during bottleneck bandwidth sharing. Joint QoE optimization refers to improving QoE fairness among clients with various video devices and availing from differentiated services with different priorities. Flex-Steward deploys an adaptive bitrate delivery algorithm based on Neural Networks (NN) and reinforcement learning at the network edge. It relies on a trained NN model to make appropriate bitrate recommendations in terms of video chunks to be requested by clients sharing the same bottleneck bandwidth. Flex-Steward is assessed in comparison with alternative state-of-the-art algorithms under different network conditions using a real-life prototype. Results show how Flex-Steward reduces the unfairness in terms of joint QoE optimization with between 10.9% and 41.7%.
Xiaoteng Ma, Qing Li 0006, Yong Jiang 0001, Gabriel-Miro Muntean, Longhao Zou
IEEE Trans. Netw. Serv. Manag.1
2021 Average-Reward Reinforcement Learning with Trust Region Methods
abstract
Most of reinforcement learning algorithms optimize the discounted criterion which is beneficial to accelerate the convergence and reduce the variance of estimates. Although the discounted criterion is appropriate for certain tasks such as financial related problems, many engineering problems treat future rewards equally and prefer a long-run average criterion. In this paper, we study the reinforcement learning problem with the long-run average criterion. Firstly, we develop a unified trust region theory with discounted and average criteria. With the average criterion, a novel performance bound within the trust region is derived with the Perturbation Analysis (PA) theory. Secondly, we propose a practical algorithm named Average Policy Optimization (APO), which improves the value estimation with a novel technique named Average Value Constraint. To the best of our knowledge, our work is the first one to study the trust region approach with the average criterion and it complements the framework of reinforcement learning beyond the discounted criterion. Finally, experiments are conducted in the continuous control environment MuJoCo. In most tasks, APO performs better than the discounted PPO, which demonstrates the effectiveness of our approach.
Xiaoteng Ma, Xiaohang Tang, Jun Yang 0028, Qianchuan Zhao
IJCAI1
2021 Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement Learning
abstract
Learning from datasets without interaction with environments (Offline Learning) is an essential step to apply Reinforcement Learning (RL) algorithms in real-world scenarios.However, compared with the single-agent counterpart, offline multi-agent RL introduces more agents with the larger state and action space, which is more challenging but attracts little attention. We demonstrate current offline RL algorithms are ineffective in multi-agent systems due to the accumulated extrapolation error. In this paper, we propose a novel offline RL algorithm, named Implicit Constraint Q-learning (ICQ), which effectively alleviates the extrapolation error by only trusting the state-action pairs given in the dataset for value estimation. Moreover, we extend ICQ to multi-agent tasks by decomposing the joint-policy under the implicit constraint. Experimental results demonstrate that the extrapolation error is successfully controlled within a reasonable range and insensitive to the number of agents. We further show that ICQ achieves the state-of-the-art performance in the challenging multi-agent offline tasks (StarCraft II). Our code is public online at https://github.com/YiqinYang/ICQ.
Yiqin Yang, Xiaoteng Ma, Chenghao Li 0002, Zewu Zheng, Gao Huang 0001, Jun Yang 0028, Qianchuan Zhao
NeurIPS2
2019 Steward: smart edge based joint QoE optimization for adaptive video streaming
abstract
With the increase of HTTP-based adaptive video streaming over the Internet, multiple clients may compete for a shared bottleneck bandwidth, which brings some damage to the fairness and stability of Quality of Experience (QoE). This paper presents Steward, a system that enforces multi-client joint QoE optimization for bottleneck bandwidth sharing. Joint QoE optimization refers to improving QoE fairness among clients with various video devices and providing differentiated service for clients with different priorities. Steward deploys the adaptive bitrate (ABR) algorithm based on neural networks (NN) and reinforcement learning at the network edge. The ABR agent trains the NN model through experience and makes appropriate bitrate guidance for video chunks to be requested by clients sharing the same bottleneck bandwidth. We compare Steward with state-of-the-art algorithms under different network conditions. Compared with all considered algorithms and conditions, Steward reduces 30%~85% QoE unfairness under the premise of differentiated service.
Xiaoteng Ma, Qing Li 0006, Jimeng Chai, Xi Xiao 0001, Shutao Xia, Yong Jiang 0001
NOSSDAV1