VLDB 2026 Research / reviewers in the wild / expert
Shanqi Liu
dblp:277/9570
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0003-0583-2423ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MCMC: Multi-Constrained Model Compression via One-Stage Envelope Reinforcement LearningabstractModel compression methods are being developed to bridge the gap between the massive scale of neural networks and the limited hardware resources on edge devices. Since most real-world applications deployed on resource-limited hardware platforms typically have multiple hardware constraints simultaneously, most existing model compression approaches that only consider optimizing one single hardware objective are ineffective. In this article, we propose an automated pruning method called multi-constrained model compression (MCMC) that allows for the optimization of multiple hardware targets, such as latency, floating point operations (FLOPs), and memory usage, while minimizing the impact on accuracy. Specifically, we propose an improved multi-objective reinforcement learning (MORL) algorithm, the one-stage envelope deep deterministic policy gradient (DDPG) algorithm, to determine the pruning strategy for neural networks. Our improved one-stage envelope DDPG algorithm reduces exploration time and offers greater flexibility in adjusting target priorities, enhancing its suitability for pruning tasks. For instance, on the visual geometry group (VGG)-16 network, our method achieved an 80% reduction in FLOPs, a reduction in memory usage, and a acceleration, with an accuracy improvement of 0.09% compared with the baseline. For larger datasets, such as ImageNet, we reduced FLOPs by 50% for MobileNet-V1, resulting in a faster speed and memory compression, while maintaining the same accuracy. When applied to edge devices, such as JETSON XAVIER NX, our method resulted in a 71% reduction in FLOPs for MobileNet-V1, leading to a faster speed, memory compression, and an accuracy improvement. Siqi Li 0009, Jun Chen 0023, Shanqi Liu, Chengrui Zhu, Guanzhong Tian, Yong Liu 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Exploit Spatiotemporal Contextual Information for 3D Single Object Tracking via Memory NetworksabstractThe point cloud-based 3D single object tracking plays an indispensable role in autonomous driving. However, the application of 3D object tracking in the real world is still challenging due to the inherent sparsity and self-occlusion of point cloud data. Therefore, it is necessary to exploit as much useful information from limited data as we can. Since 3D object tracking is a video-level task, the appearance of objects changes gradually over time, and there is rich spatiotemporal contextual information among historical frames. However, existing methods do not fully utilize this information. To address this, we propose a new method called SCTrack, which utilizes a memory-based paradigm to exploit spatiotemporal contextual information. SCTrack incorporates both long-term and short-term memory banks to store the spatiotemporal features of targets from historical frames. By doing so, the tracker can benefit from the entire video sequence and make more informed predictions. Additionally, SCTrack extracts the mask prior to augmenting the target representation, improving the target-background discriminability. Extensive experiments on KITTI, nuScenes, and Waymo Open datasets verify the effectiveness of our proposed method. Jongwon Ra, Mengmeng Wang 0005, Jianbiao Mei, Shanqi Liu, Yu Yang 0001, Yong Liu 0007 |
3DV | 4 |
| 2024 | Solving Homogeneous and Heterogeneous Cooperative Tasks with Greedy Sequential ExecutionabstractCooperative multi-agent reinforcement learning (MARL) is extensively used for solving complex cooperative tasks, and value decomposition methods are a prevalent approach for this domain. However, these methods have not been successful in addressing both homogeneous and heterogeneous tasks simultaneously which is a crucial aspect for the practical application of cooperative agents.
On one hand, value decomposition methods demonstrate superior performance in homogeneous tasks. Nevertheless, they tend to produce agents with similar policies, which is unsuitable for heterogeneous tasks. On the other hand, solutions based on personalized observation or assigned roles are well-suited for heterogeneous tasks. However, they often lead to a trade-off situation where the agent's performance in homogeneous scenarios is negatively affected due to the aggregation of distinct policies. An alternative approach is to adopt sequential execution policies, which offer a flexible form for learning both types of tasks. However, learning sequential execution policies poses challenges in terms of credit assignment, and the limited information about subsequently executed agents can lead to sub-optimal solutions, which is known as the relative over-generalization problem. To tackle these issues, this paper proposes Greedy Sequential Execution (GSE) as a solution to learn the optimal policy that covers both scenarios. In the proposed GSE framework, we introduce an individual utility function into the framework of value decomposition to consider the complex interactions between agents.
This function is capable of representing both the homogeneous and heterogeneous optimal policies. Furthermore, we utilize greedy marginal contribution calculated by the utility function as the credit value of the sequential execution policy to address the credit assignment and relative over-generalization problem. We evaluated GSE in both homogeneous and heterogeneous scenarios. The results demonstrate that GSE achieves significant improvement in performance across multiple domains, especially in scenarios involving both homogeneous and heterogeneous tasks. Shanqi Liu, Dong Xing, Pengjie Gu, Xinrun Wang, Bo An 0001 |
ICLR | 1 |
| 2024 | True Knowledge Comes from Practice: Aligning Large Language Models with Embodied Environments via Reinforcement LearningabstractDespite the impressive performance across numerous tasks, large language models (LLMs) often fail in solving simple decision-making tasks due to the misalignment of the knowledge in LLMs with environments. On the contrary, reinforcement learning (RL) agents learn policies from scratch, which makes them always align with environments but difficult to incorporate prior knowledge for efficient explorations. To narrow the gap, we propose TWOSOME, a novel general online framework that deploys LLMs as decision-making agents to efficiently interact and align with embodied environments via RL without requiring any prepared datasets or prior knowledge of the environments. Firstly, we query the joint probabilities of each valid action with LLMs to form behavior policies. Then, to enhance the stability and robustness of the policies, we propose two normalization methods and summarize four prompt design principles. Finally, we design a novel parameter-efficient training architecture where the actor and critic share one frozen LLM equipped with low-rank adapters (LoRA) updated by PPO. We conduct extensive experiments to evaluate TWOSOME. i) TWOSOME exhibits significantly better sample efficiency and performance compared to the conventional RL method, PPO, and prompt tuning method, SayCan, in both classical decision-making environment, Overcooked, and simulated household environment, VirtualHome. ii) Benefiting from LLMs' open-vocabulary feature, TWOSOME shows superior generalization ability to unseen tasks. iii) Under our framework, there is no significant loss of the LLMs' original ability during online PPO finetuning. Weihao Tan, Wentao Zhang 0007, Shanqi Liu, Longtao Zheng, Xinrun Wang, Bo An 0001 |
ICLR | 3 |
| 2024 | Learning Multi-Agent Cooperation via Considering Actions of TeammatesabstractRecently value-based centralized training with decentralized execution (CTDE) multi-agent reinforcement learning (MARL) methods have achieved excellent performance in cooperative tasks. However, the most representative method among these methods, Q-network MIXing (QMIX), restricts the joint action Q values to be a monotonic mixing of each agent's utilities. Furthermore, current methods cannot generalize to unseen environments or different agent configurations, which is known as ad hoc team play situation. In this work, we propose a novel Q values decomposition that considers both the return of an agent acting on its own and cooperating with other observable agents to address the nonmonotonic problem. Based on the decomposition, we propose a greedy action searching method that can improve exploration and is not affected by changes in observable agents or changes in the order of agents' actions. In this way, our method can adapt to ad hoc team play situation. Furthermore, we utilize an auxiliary loss related to environmental cognition consistency and a modified prioritized experience replay (PER) buffer to assist training. Our extensive experimental results show that our method achieves significant performance improvements in both challenging monotonic and nonmonotonic domains, and can handle the ad hoc team play situation perfectly. Shanqi Liu, Wenzhou Chen, Guanzhong Tian, Jun Chen 0023, Yong Liu 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Controlling Type Confounding in Ad Hoc Teamwork with Instance-wise Teammate Feedback RectificationabstractAd hoc teamwork requires an agent to cooperate with unknown teammates without prior coordination. Many works propose to abstract teammate instances into high-level representation of types and then pre-train the best response for each type. However, most of them do not consider the distribution of teammate instances within a type. This could expose the agent to the hidden risk of type confounding. In the worst case, the best response for an abstract teammate type could be the worst response for all specific instances of that type. This work addresses the issue from the lens of causal inference. We first theoretically demonstrate that this phenomenon is due to the spurious correlation brought by uncontrolled teammate distribution. Then, we propose our solution, CTCAT, which disentangles such correlation through an instance-wise teammate feedback rectification. This operation reweights the interaction of teammate instances within a shared type to reduce the influence of type confounding. The effect of CTCAT is evaluated in multiple domains, including classic ad hoc teamwork tasks and real-world scenarios. Results show that CTCAT is robust to the influence of type confounding, a practical issue that directly hazards the robustness of our trained agents but was unnoticed in previous works. Dong Xing, Pengjie Gu, Xinrun Wang, Shanqi Liu, Longtao Zheng, Bo An 0001, Gang Pan 0001 |
ICML | 5 |
| 2023 | Lightweight Cryptography Implementation for Internet of Things Network on FPGAabstractWith the development of modern communication technology, traditional Internet of Things systems could not provide sufficient support for large data flow, hardware resource, power consumption, and security during transmission. Especially when it comes to the Industrial Internet of Things (IIoT), where the limitations of hardware resource and power consumption are more strict; the requirements of network security and hardware security are much higher. This paper implements a User Datagram Protocol (UDP) communication system, which is encrypted with one of the Lightweight Cryptography, Xoodyak. We aim to design a lightweight encrypted communication platform. Compared with a public key encryption system, our implementation costs much less hardware resource and power consumption, while providing excellent side-channel attack (SCA) protection. Besides, when there is a large burst length of data flow, two asynchronous FIFO can restore those data respectively, so our design can maintain throughput and data integrity in extreme cases. Based on the above properties, our design is relatively ideal in IIoT encryption scenarios. Guanzhong Tian, Longhua Ma, Zhishan Li, Shanqi Liu |
SMC | 5 |
| 2023 | Expert demonstrations guide reward decomposition for multi-agent cooperation
Shanqi Liu, Yudi Ruan, Kexin Zhang 0007, Yong Liu 0007 |
Neural Comput. Appl. | 3 |
| 2022 | CICC: Channel Pruning via the Concentration of Information and Contributions of Channels
Zhishan Li, Yingqing Yang, Lei Xie 0007, Yong Liu 0007, Longhua Ma, Shanqi Liu, Guanzhong Tian |
BMVC | 7 |
| 2021 | One-shot Face Reenactment Using Appearance Adaptive NormalizationabstractThe paper proposes a novel generative adversarial network for one-shot face reenactment, which can animate a single face image to a different pose-and-expression (provided by a driving image) while keeping its original appearance. The core of our network is a novel mechanism called appearance adaptive normalization, which can effectively integrate the appearance information from the input image into our face generator by modulating the feature maps of the generator using the learned adaptive parameters. Furthermore, we specially design a local net to reenact the local facial components (i.e., eyes, nose and mouth) first, which is a much easier task for the network to learn and can in turn provide explicit anchors to guide our face generator to learn the global appearance and pose-and-expression. Extensive quantitative and qualitative experiments demonstrate the significant efficacy of our model compared with prior one-shot methods. Guangming Yao, Yi Yuan 0002, Tianjia Shao, Shuang Li 0008, Shanqi Liu, Yong Liu 0007, Mengmeng Wang 0005, Kun Zhou 0001 |
AAAI | 5 |
| 2021 | Moving Forward in Formation: A Decentralized Hierarchical Learning Approach to Multi-Agent Moving TogetherabstractMulti-agent path finding in formation has many potential real-world applications like mobile warehouse robotics. However, previous multi-agent path finding (MAPF) methods hardly take formation into consideration. Further-more, they are usually centralized planners and require the whole state of the environment. Other decentralized partially observable approaches to MAPF are reinforcement learning (RL) methods. However, these RL methods encounter difficulties when learning path finding and formation problems at the same time. In this paper, we propose a novel decentralized partially observable RL algorithm that uses a hierarchical structure to decompose the multi-objective task into unrelated ones. It also calculates a theoretical weight that makes each tasks reward has equal influence on the final RL value function. Additionally, we introduce a communication method that helps agents cooperate with each other. Experiments in simulation show that our method outperforms other end-to-end RL methods and our method can naturally scale to large world sizes where centralized planner struggles. We also deploy and validate our method in a real-world scenario. Shanqi Liu, Licheng Wen, Jinhao Cui, Xuemeng Yang, Yong Liu 0007 |
IROS | 1 |
| 2021 | Self-play reinforcement learning with comprehensive critic in computer games
Shanqi Liu, Yujie Wang 0005, Wenzhou Chen, Yong Liu 0007 |
Neurocomputing | 1 |