Xingyuan Hua

dblp:352/6715 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 41% Efficient and distributed learning · 37% Language models and text generation · 7%
Databases, data mining, and information retrieval
1 paper
Spatial and temporal data management · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
federated learning
2.432025
Momentum-Based Contextual Federated Reinforcement Learning · IEEE Trans. Netw. 2025
Federated Offline Policy Optimization with Dual Regularization · INFOCOM 2024
Momentum-Based Federated Reinforcement Learning with Interaction and Communication Efficiency · INFOCOM 2024
Machine learning › Efficient and distributed learning › federated learning › federated sequential learning
federated reinforcement learning
1.622025
Momentum-Based Contextual Federated Reinforcement Learning · IEEE Trans. Netw. 2025
Momentum-Based Federated Reinforcement Learning with Interaction and Communication Efficiency · INFOCOM 2024
Machine learning › Reinforcement learning
imitation learning
1.522024
OLLIE: Imitation Learning from Offline Pretraining to Online Finetuning · ICML 2024
How to Leverage Diverse Demonstrations in Offline Imitation Learning · ICML 2024
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning
0.812024
How to Leverage Diverse Demonstrations in Offline Imitation Learning · ICML 2024
Natural language and speech › Language models and text generation › in-context learning
demonstration selection
0.812024
How to Leverage Diverse Demonstrations in Offline Imitation Learning · ICML 2024
Machine learning › Reinforcement learning › imitation learning
offline imitation learning
0.812024
How to Leverage Diverse Demonstrations in Offline Imitation Learning · ICML 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Federated Offline Policy Optimization with Dual Regularization · INFOCOM 2024
Machine learning › Deep learning architectures and training
regularization
0.812024
Federated Offline Policy Optimization with Dual Regularization · INFOCOM 2024
Machine learning › Optimization for machine learning
stochastic optimization
0.812024
Momentum-Based Federated Reinforcement Learning with Interaction and Communication Efficiency · INFOCOM 2024
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.712023
Air-Ground Spatial Crowdsourcing with UAV Carriers by Geometric Graph Convolutional Multi-Agent Deep Reinforcement Learning · ICDE 2023
Spatial and temporal data management
spatial crowdsourcing
0.712023
Air-Ground Spatial Crowdsourcing with UAV Carriers by Geometric Graph Convolutional Multi-Agent Deep Reinforcement Learning · ICDE 2023
Machine learning › Graph learning › graph neural network
graph convolutional network
0.212023
Air-Ground Spatial Crowdsourcing with UAV Carriers by Geometric Graph Convolutional Multi-Agent Deep Reinforcement Learning · ICDE 2023

Methods — techniques the papers use, named apart from their topics

momentum · 1.6importance sampling · 1.6attention mechanism · 1.3attention-based contextual representation · 0.9state-action similarity · 0.8resultant state criterion · 0.8policy pretraining · 0.8policy optimization · 0.8dual regularization · 0.8discriminator alignment · 0.8e-comm · 0.7MC-GCN · 0.7GARL · 0.7
YearPublicationVenuePosition
2025 Momentum-Based Contextual Federated Reinforcement Learning
abstract
Federated Reinforcement Learning (FRL) is an attractive edge learning paradigm for decision-making applications, which has garnered significant interest recently. However, owing to the inherent spatio-temporal non-stationarity of local state-action distributions, current FRL approaches typically suffer from high interaction and communication costs. In this paper, we introduce a new FRL method, which incorporates momentum, importance sampling, and server-side adjustments, capable of controlling the gradient shifts induced by the non-stationary data. We prove that by proper selection of momentum parameters and interaction frequency, it can achieve$\tilde {\mathcal {O}}(H N^{-1}\epsilon ^{-3/2})$and$\tilde {\mathcal {O}}(\epsilon ^{-1})$interaction and communication complexities (N represents the agent number), where the interaction complexity achieves linear speedup with the number of agents, and the communication complexity aligns with the best achievable among existing first-order FL algorithms. Further, we leverage attention-based contextual representation extraction to enable the learning policy to adapt to heterogeneous tasks and environments. Extensive experiments demonstrate that our proposed method significantly outperforms existing baselines on a range of complex, high-dimensional single-task and multi-task benchmarks.
Sheng Yue 0001, Xingyuan Hua, Yongheng Deng, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Netw.2
2024 How to Leverage Diverse Demonstrations in Offline Imitation Learning
abstract
Offline Imitation Learning (IL) with imperfect demonstrations has garnered increasing attention owing to the scarcity of expert data in many real-world domains. A fundamental problem in this scenario is *how to extract positive behaviors from noisy data*. In general, current approaches to the problem select data building on state-action similarity to given expert demonstrations, neglecting precious information in (potentially abundant) *diverse* state-actions that deviate from expert ones. In this paper, we introduce a simple yet effective data selection method that identifies positive behaviors based on their *resultant states* - a more informative criterion enabling explicit utilization of dynamics information and effective extraction of both expert and beneficial diverse behaviors. Further, we devise a lightweight behavior cloning algorithm capable of leveraging the expert and selected data correctly. In the experiments, we evaluate our method on a suite of complex and high-dimensional offline IL benchmarks, including continuous-control and vision-based tasks. The results demonstrate that our method achieves state-of-the-art performance, outperforming existing methods on **20/21** benchmarks, typically by **2-5x**, while maintaining a comparable runtime to Behavior Cloning (BC).
Sheng Yue 0001, Jiani Liu 0005, Xingyuan Hua, Ju Ren 0001, Sen Lin 0001, Junshan Zhang, Yaoxue Zhang
ICML3
2024 OLLIE: Imitation Learning from Offline Pretraining to Online Finetuning
abstract
In this paper, we study offline-to-online Imitation Learning (IL) that pretrains an imitation policy from static demonstration data, followed by fast finetuning with minimal environmental interaction. We find the naive combination of existing offline IL and online IL methods tends to behave poorly in this context, because the initial discriminator (often used in online IL) operates randomly and discordantly against the policy initialization, leading to misguided policy optimization and *unlearning* of pretraining knowledge. To overcome this challenge, we propose a principled offline-to-online IL method, named OLLIE, that simultaneously learns a near-expert policy initialization along with an *aligned discriminator initialization*, which can be seamlessly integrated into online IL, achieving smooth and fast finetuning. Empirically, OLLIE consistently and significantly outperforms the baseline methods in **20** challenging tasks, from continuous control to vision-based domains, in terms of performance, demonstration efficiency, and convergence speed. This work may serve as a foundation for further exploration of pretraining and finetuning in the context of IL.
Sheng Yue 0001, Xingyuan Hua, Ju Ren 0001, Sen Lin 0001, Junshan Zhang, Yaoxue Zhang
ICML2
2024 Momentum-Based Federated Reinforcement Learning with Interaction and Communication Efficiency
abstract
Federated Reinforcement Learning (FRL) has garnered increasing attention recently. However, due to the intrinsic spatio-temporal non-stationarity of data distributions, the current approaches typically suffer from high interaction and communication costs. In this paper, we introduce a new FRL algorithm, named MFPO, that utilizes momentum, importance sampling, and additional server-side adjustment to control the shift of stochastic policy gradients and enhance the efficiency of data utilization. We prove that by proper selection of momentum parameters and interaction frequency, MFPO can achieve $\widetilde {\mathcal{O}}\left({H{N^{ - 1}}{\varepsilon ^{ - 3/2}}}\right)$ and $\widetilde {\mathcal{O}}\left({{\varepsilon ^{ - 1}}}\right)$ interaction and communication complexities (N represents the number of agents), where the interaction complexity achieves linear speedup with the number of agents, and the communication complexity aligns the best achievable of existing first-order FL algorithms. Extensive experiments corroborate the substantial performance gains of MFPO over existing methods on a suite of complex and high-dimensional benchmarks.
Sheng Yue 0001, Xingyuan Hua, Ju Ren 0001
INFOCOM2
2024 Federated Offline Policy Optimization with Dual Regularization
abstract
Federated Reinforcement Learning (FRL) has been deemed as a promising solution for intelligent decision-making in the era of Artificial Internet of Things. However, existing FRL approaches often entail repeated interactions with the environment during local updating, which can be prohibitively expensive or even infeasible in many real-world domains. To overcome this challenge, this paper proposes a novel offline federated policy optimization algorithm, named DRPO, which enables distributed agents to collaboratively learn a decision policy only from private and static data without further environmental interactions. DRPO leverages dual regularization, incorporating both the local behavioral policy and the global aggregated policy, to judiciously cope with the intrinsic two-tier distributional shifts in offline FRL. Theoretical analysis characterizes the impact of the dual regularization on performance, demonstrating that by achieving the right balance thereof, DRPO can effectively counteract distributional shifts and ensure strict policy improvement in each federative learning round. Extensive experiments validate the significant performance gains of DRPO over baseline methods.
Sheng Yue 0001, Zerui Qin, Xingyuan Hua, Yongheng Deng, Ju Ren 0001
INFOCOM3
2023 Air-Ground Spatial Crowdsourcing with UAV Carriers by Geometric Graph Convolutional Multi-Agent Deep Reinforcement Learning
abstract
Spatial Crowdsourcing (SC) has been proved as an effective paradigm for data acquisition in urban environments. Apart from using human participants, with the rapid development of unmanned vehicles (UVs) technologies, unmanned aerial or ground vehicles (UAVs, UGVs) are equipped with various high-precision sensors, enabling them to become new types of data collectors. However, UGVs’ operational range is constrained by the road network, and UAVs are limited by power supply, it is thus natural to use UGVs and UAVs together as a coalition, and more precisely, UGVs behave as the UAV carriers for range extensions to achieve complicated air-ground SC tasks. In this paper, we propose a novel communication-based multi-agent deep reinforcement learning method called "GARL", which consists of a multi-center attention-based graph convolutional network (GCN) to accurately extract UGV specific features from UGV stop network called "MC-GCN", and a novel GNN-based communication mechanism called "E-Comm" to make the cooperation among UGVs adaptive to constant changing of geometric shapes formed by UGVs. Extensive simulation results on two campuses of KAIST and UCLA campuses show that GARL consistently outperforms eight other baselines in terms of overall efficiency.
Yu Wang 0115, Jingfei Wu, Xingyuan Hua, Chi Harold Liu, Guozheng Li 0002, Jianxin Zhao 0001, Ye Yuan 0001, Guoren Wang
ICDE3