VLDB 2026 Research / reviewers in the wild / expert
Jingqing Ruan
dblp:304/3544
· DBLP profile ↗
17ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-4857-9053ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular DataabstractFengXian Dong, Zhi Zheng, Xiao Han, Wei Chen, Jingqing Ruan, Tong Xu, Yong Chen, Enhong Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Fengxian Dong, Zhi Zheng 0008, Wei Chen 0156, Jingqing Ruan, Tong Xu 0001, Enhong Chen |
ACL (1) | 5 |
| 2026 | PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward ModelingabstractReward models (RMs) are central to reinforcement learning from human feedback (RLHF), providing the critical supervision signals that align large language models (LLMs) with human preferences.Generative reward models (GRMs) provide greater interpretability than traditional scalar RMs, but they come with a critical trade-off: pairwise methods are hindered by a training-inference mismatch, while pointwise methods require expensive absolute annotations.To bridge this gap, we propose the Preference-aware Task-adaptive Reward Model (PaTaRM).Unlike prior approaches, PaTaRM enables robust pointwise training using readily available pairwise data via a novel Preference-Aware Reward (PAR) mechanism, eliminating the need for explicit rating labels.Furthermore, it incorporates a taskadaptive rubric system that dynamically generates instance-specific criteria for precise evaluation.Extensive experiments demonstrate that PaTaRM achieves an average relative improvement of 8.7% over the corresponding base models on RewardBench and RMBench across the Qwen3-8B and Qwen3-14B backbones.Crucially, when used as a reward model for downstream RLHF, it yields an average relative improvement of 13.6% over the corresponding base policies on IFEval and InfoBench, validating its effectiveness for policy alignment.Our code, data, and checkpoints are available at https://huggingface.co/AIJian/PaTaRM. Ai Jian, Jingqing Ruan, Dailin Li |
ACL (1) | 2 |
| 2026 | Adaptive Graph Coordination Strategy in Multiagent Reinforcement LearningabstractMany real-world applications involve a team of agents who must coordinate each other's policies in real-time to achieve a shared goal. Previous studies mainly focus on decentralized control to maximize common rewards, with little consideration of coordination between control policies, which is critical in dynamic and complicated environments. Viewing this issue, we propose a novel adaptive graph coordination strategy that factorizes the joint policy into an adaptive graph generator and a graph-based coordinated policy. We employ a difference aware module to control when to generate graphs and an encoder-decoder module to acquire the underlying decision graph structure. Moreover, we introduce DAGness- and DAG depth constrained optimization to adjust the graph structure and strike a balance between efficiency and performance. We also present a graph-based coordinated policy to make asynchronous decisions based on the inter-agent coordination dependencies implied in the generated graph. Empirical evaluations on some cooperative multi-agent environments demonstrate the superiority of the proposed method, with faster convergence and more efficient coordinated policies. Zhongwei Yu, Jingqing Ruan, Dengpeng Xing |
IEEE Trans. Games | 2 |
| 2025 | An End-to-End Deep Reinforcement Learning Based Modular Task Allocation Framework for Autonomous Mobile SystemsabstractIntelligent decision-making systems that can solve task allocation problems are critical for multi-robot systems to conduct industrial applications in a collaborative and automated way, such as warehouse inspection using mobile robots, hydrographic surveying using unmanned surface vehicles, etc. This paper, therefore, aims to address the task allocation problem for multi-agent autonomous mobile systems to autonomously and intelligently allocate multiple tasks to a fleet of robots. Such a problem is normally regarded as an independent decision-making process decoupled from the following task planning for the member robots. To avoid the sub-optimal allocation caused by the decoupling, an end-to-end task allocation framework is proposed to tackle this combinatorial optimisation problem while taking the succeeding task planning into account during the optimisation process. The problem is formulated as a special variant of the multi-depot multiple travelling salesmen problem (mTSP). The proposed end-to-end task allocation framework employs deep reinforcement learning methods to replace the handcrafted heuristics used in previous works. The proposed framework features a modular design of the reinforcement learning agent which can be customised for various applications. Moreover, a real-robot implementation setup based on the Robot Operating System 2 is presented to fulfil the simulation-to-reality gap. A warehouse inspection mission is executed to validate the training outcome of the proposed framework. The framework has been cross-validated via both simulated and real-robot tests with various parameter settings, where adaptability and performance are well demonstrated.Note to Practitioners—This paper is motivated by the problem of dispatching a fleet of autonomous mobile robots to tackle a mission that can be resolved into multiple waypoint-following tasks. An end-to-end modular framework is proposed, making task allocation decisions based on the given waypoint information. By using the reinforcement learning technique, the deep neural network could learn sophisticated policies for allocating tasks. The policies are trained in a specific pattern which ensures their joint optimisation for a solver that outputs the near optimal task execution sequences in an efficient way. This leads to a multiple travelling salesmen problem (mTSP) solution. Pre-trained policies are tested in several industrial scenarios reflecting the applications of search and rescue, maritime surveying, and warehouse automation, among others. A hardware implementation configuration based on the Robot Operating System 2 is also presented to support the practical deployment the framework. Jingqing Ruan, Yali Du 0001, Richard Bucknall, Yuanchang Liu |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | MaDE: Multi-Scale Decision Enhancement for Multi-Agent Reinforcement LearningabstractIn the domain of multi-agent reinforcement learning (MARL), the limited information availability, complex agent interactions, and individual capabilities among agents often pose a bottleneck for effective decision-making. Previous studies frequently fall short due to insufficient consideration of these multi-dimensional challenges. Thus, this paper introduces a novel methodology, termed Multi-scale Decision Enhancement (MaDE), anchored by a dual-wise bisimulation framework for pre-training agent encoders. The MaDE framework aims to facilitate decision-making across three pivotal dimensions: macroscale awareness, mesoscale coordination, and microscale insight. At the macro level, a pretrained global encoder captures a situational awareness map to guide overall strategies. At the meso level, specialized local encoders generate cluster-based representations to promote inter-agent cooperation. At the micro level, individual agents focus on the accurate decision-making process. Empirical evaluations validate that MaDE outperforms state-of-the-art methods in various multi-agent environments, which shows the potential to tackle the intricate challenges of MARL, enabling agents to make more informed, coordinated, and adaptive decisions. Code is available at https://github.com/paper2023/MaDE. Jingqing Ruan, Runpeng Xie, Xuantang Xiong, Bo Xu 0002 |
ICASSP | 1 |
| 2024 | UNeC: Unsupervised Exploring In Controllable SpaceabstractIn unsupervised reinforcement learning, agents traverse a reward-free environment, aiming for rapid generalisation to subsequent tasks. This strategy offers a compelling resolution to the challenges of sample efficiency. Nevertheless, environments are frequently saturated with excessive information. The omnipresence of uncontrollable states, akin to noise, can impede the efficacy of unsupervised exploration. To address this issue, we introduce the UNeC framework, standing for UNsupervised Exploring in Controllable Space. This approach leverages the intricate dependencies between states and actions, establishing constraints on controllable state representation. Building on this foundation, we employ particle entropy to evaluate states, adeptly directing agent exploration. Empirical evaluations on the unsupervised RL benchmark affirm the superiority of our method, with a particular emphasis on its dominance in noise-affected scenarios. Xuantang Xiong, Linghui Meng 0001, Jingqing Ruan, Bo Xu 0002 |
ICASSP | 3 |
| 2024 | Learning Causal Dynamics Models in Object-Oriented EnvironmentsabstractCausal dynamics models (CDMs) have demonstrated significant potential in addressing various challenges in reinforcement learning. To learn CDMs, recent studies have performed causal discovery to capture the causal dependencies among environmental variables. However, the learning of CDMs is still confined to small-scale environments due to computational complexity and sample efficiency constraints. This paper aims to extend CDMs to large-scale object-oriented environments, which consist of a multitude of objects classified into different categories. We introduce the Object-Oriented CDM (OOCDM) that shares causalities and parameters among objects belonging to the same class. Furthermore, we propose a learning method for OOCDM that enables it to adapt to a varying number of objects. Experiments on large-scale tasks indicate that OOCDM outperforms existing CDMs in terms of causal discovery, prediction accuracy, generalization, and computational efficiency. Zhongwei Yu, Jingqing Ruan, Dengpeng Xing |
ICML | 2 |
| 2024 | X-Light: Cross-City Traffic Signal Control Using Transformer on Transformer as Meta Multi-Agent Reinforcement Learner
Haoyuan Jiang, Ziyue Li 0002, Hua Wei 0001, Xuantang Xiong, Jingqing Ruan, Jiaming Lu, Hangyu Mao, Rui Zhao 0001 |
IJCAI | 5 |
| 2024 | CoSLight: Co-optimizing Collaborator Selection and Decision-making to Enhance Traffic Signal ControlabstractEffective multi-intersection collaboration is pivotal for reinforcement-learning-based traffic signal control to alleviate congestion. Existing work mainly chooses neighboring intersections as collaborators. However, quite a lot of congestion, even some wide-range congestion, is caused by non-neighbors failing to collaborate. To address these issues, we propose to separate the collaborator selection as a second policy to be learned, concurrently being updated with the original signal-controlling policy. Specifically, the selection policy in real-time adaptively selects the best teammates according to phase- and intersection-level features. Empirical results on both synthetic and real-world datasets provide robust validation for the superiority of our approach, offering significant improvements over existing state-of-the-art methods. Code is available at https://github.com/bonaldli/CoSLight. Jingqing Ruan, Ziyue Li 0002, Hua Wei 0001, Haoyuan Jiang, Jiaming Lu, Xuantang Xiong, Hangyu Mao, Rui Zhao 0001 |
KDD | 1 |
| 2023 | Learning to Collaborate by Grouping: A Consensus-Oriented Strategy for Multi-Agent Reinforcement LearningabstractMulti-agent systems require effective coordination between groups and individuals to achieve common goals. However, current multi-agent reinforcement learning (MARL) methods primarily focus on improving individual policies and do not adequately address group-level policies, which leads to weak cooperation. To address this issue, we propose a novel Consensus-oriented Strategy (CoS) that emphasizes group and individual policies simultaneously. Specifically, CoS comprises two main components: (a) the vector quantized group consensus module, which extracts discrete latent embeddings that represent the stable and discriminative group consensus, and (b) the group consensus-oriented strategy, which integrates the group policy using a hypernet and the individual policies using the group consensus, thereby promoting coordination at both the group and individual levels. Through empirical experiments on cooperative navigation tasks with both discrete and continuous spaces, as well as google research football, we demonstrate that CoS outperforms state-of-the-art MARL algorithms and achieves better collaboration, thus providing a promising solution for achieving effective coordination in multi-agent systems. Jingqing Ruan, Xiaotian Hao, Dong Li 0016, Hangyu Mao |
ECAI | 1 |
| 2023 | Cardsformer: Grounding Language to Learn a Generalizable Policy in HearthstoneabstractHearthstone is a widely played collectible card game that challenges players to strategize using cards with various effects described in natural language. While human players can easily comprehend card descriptions and make informed decisions, artificial agents struggle to understand the game’s inherent rules and are unable to generalize their policies through natural language. To address this issue, we propose Cardsformer, a method capable of acquiring linguistic knowledge and learning a generalizable policy in Hearthstone. Cardsformer consists of a Prediction Model trained with offline trajectories to predict state transitions based on card descriptions and a Policy Model capable of generalizing its policy on unseen cards. To our knowledge, this is the first work to consider language knowledge in a card game. Experiments show that our approach significantly improves data efficiency and outperforms the state-of-the-art in Hearthstone even when there are untrained cards in the deck, inspiring a new perspective of tackling problems as such with knowledge representation from large language models. As the game constantly releases new cards along with new descriptions and new effects, the challenge in Hearthstone remains. To encourage further research, we make our code publicly available and publish PyStone, the code base of Hearthstone on which we conducted our experiments, as an open benchmark. Wannian Xia, Jingqing Ruan, Dengpeng Xing, Bo Xu 0002 |
ECAI | 3 |
| 2023 | Task-Prompt Generalised World Model in Multi-Environment Offline Reinforcement LearningabstractOffline reinforcement learning (RL) circumvents costly interactions with the environment by utilising historical trajectories. Incorporating a world model into this method could substantially enhance the transfer performance of various tasks without expensive calculations from scratch. However, due to the complexity arising from different types of generalisation, previous works have focused almost exclusively on single-environment tasks. In this study, we introduce a multi-environment offline RL setting to investigate whether a generalised world model can be learned from large, diverse datasets and serve as a good surrogate for policy learning in different tasks. Inspired by the success of multi-task prompt methods, we propose the Task-prompt Generalised World Model (TGW) framework, which demonstrates notable performance in this setting. TGW comprises three modules: a task-state prompter, a generalised dynamics module, and a reward module. We implement the generalised dynamics module as a transformer-based recurrent state-space model and employ prompts to provide task-specific instructions, enabling TGW to address the internal stochasticity of the generalised world model. On the MuJoCo control benchmarks, TGW significantly outperforms previous offline RL algorithms in multi-environment setting. Xuantang Xiong, Linghui Meng 0001, Jingqing Ruan, Qingyang Zhang 0004, Guoqi Li 0002, Dengpeng Xing, Bo Xu 0002 |
ECAI | 3 |
| 2023 | Efficient Hierarchical Reinforcement Learning via Mutual Information Constrained Subgoal Discovery
Kaishen Wang, Jingqing Ruan, Qingyang Zhang 0004, Dengpeng Xing |
ICONIP (7) | 2 |
| 2023 | Explainable Reinforcement Learning via a Causal World ModelabstractGenerating explanations for reinforcement learning (RL) is challenging as actions may produce long-term effects on the future. In this paper, we develop a novel framework for explainable RL by learning a causal world model without prior knowledge of the causal structure of the environment. The model captures the influence of actions, allowing us to interpret the long-term effects of actions through causal chains, which present how actions influence environmental variables and finally lead to rewards. Different from most explanatory models which suffer from low accuracy, our model remains accurate while improving explainability, making it applicable in model-based learning. As a result, we demonstrate that our causal model can serve as the bridge between explainability and learning. Zhongwei Yu, Jingqing Ruan, Dengpeng Xing |
IJCAI | 2 |
| 2023 | Balancing Exploration and Exploitation in Hierarchical Reinforcement Learning via Latent Landmark GraphsabstractGoal-Conditioned Hierarchical Reinforcement Learning (GCHRL) is a promising paradigm to address the exploration-exploitation dilemma in reinforcement learning. It decomposes the source task into sub goal conditional subtasks and conducts exploration and exploitation in the subgoal space. The effectiveness of GCHRL heavily relies on sub goal representation functions and sub goal selection strategy. However, existing works often overlook the temporal coherence in GCHRL when learning latent sub goal representations and lack an efficient sub goal selection strategy that balances exploration and exploitation. This paper proposes HIerarchical reinforcement learning via dynamically building Latent Landmark graphs (HILL) to overcome these limitations. HILL learns latent subgoal representations that satisfy temporal coherence using a contrastive representation learning objective. Based on these representations, HILL dynamically builds latent landmark graphs and employs a novelty measure on nodes and a utility measure on edges. Finally, HILL develops a subgoal selection strategy that balances exploration and exploitation by jointly considering both measures. Experimental results demonstrate that HILL outperforms state-of-the-art baselines on continuous control tasks with sparse rewards in sample efficiency and asymptotic performance. Our code is available at https://github.com/papercode2022/HILL. Qingyang Zhang 0004, Jingqing Ruan, Xuantang Xiong, Dengpeng Xing, Bo Xu 0002 |
IJCNN | 3 |
| 2023 | DRL-Based Adaptive Sharding for Blockchain-Based Federated LearningabstractBlockchain-based Federated Learning (FL) technology enables vehicles to make smart decisions, improving vehicular services and enhancing the driving experience through a secure and privacy-preserving manner in Intelligent Transportation Systems (ITS). Many existing works exploit two-layer blockchain-based FL frameworks consisting of a mainchain and subchains for data interactions among intelligent vehicles, which resolve the limited throughput issue of single blockchain-based vehicular networks. However, the existing two-layer frameworks still suffer from a) strong dependency on predetermined and fixed parameters of vehicular blockchains which limit blockchain throughput and reliability; and b) high communication costs incurred by interactions among intelligent vehicles between the mainchain and subchains. To address the above challenges, we first design an adaptive blockchain-enabled FL framework for ITS based on blockchain sharding to facilitate decentralized vehicular data flows among intelligent vehicles. A streamline-based shard transmission mechanism is proposed to ensure communication efficiency almost without compromising the FL accuracy. We further formulate the proposed framework and propose an adaptive sharding mechanism using Deep Reinforcement Learning to automate the selection of parameters of vehicular shards. Numerical results clearly show that the proposed framework and mechanisms achieve adaptive, communication-efficient, credible, and scalable data interactions among intelligent vehicles. Yijing Lin, Zhipeng Gao 0001, Hongyang Du 0001, Jiawen Kang 0001, Dusit Niyato, Qian Wang 0015, Jingqing Ruan, Shaohua Wan 0001 |
IEEE Trans. Commun. | 7 |
| 2022 | Learning in Bi-level Markov GamesabstractAlthough multi-agent reinforcement learning (MARL) has demonstrated remarkable progress in tackling sophisticated cooperative tasks, the assumption that agents take simultaneous actions still limits the applicability of MARL for many real-world problems. In this work, we relax the assumption by proposing the framework of the bi-level Markov game (BMG). BMG breaks the simultaneity by assigning two players with a leader-follower relationship in which the leader considers the policy of the follower who is taking the best response based on the leader's actions. We propose two provably convergent algorithms to solve BMG: BMG-1 and BMG-2. The former uses the standard Q-learning, while the latter relieves solving the local Stackelberg equilibrium in BMG-1 with the further two-step transition to estimate the state value. For both methods, we consider temporal difference learning techniques with both tabular and neural network representations. To verify the effectiveness of our BMG framework, we test on a series of games, including Seeker, Cooperative Navigation, and Football, that are challenging to existing MARL solvers find challenging to solve: Seeker, Cooperative Navigation, and Football. Experimental results show that our BMG methods achieve competitive advantages in terms of better performance and lower variance. Linghui Meng 0001, Jingqing Ruan, Dengpeng Xing, Bo Xu 0002 |
IJCNN | 2 |