Zongkai Liu

dblp:214/0917 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 86% Multi-agent systems · 10% Knowledge representation and reasoning · 2%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 67% Mathematical optimization · 33%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.422025
Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization · AAAI 2025
A Unified Diversity Measure for Multiagent Reinforcement Learning · NeurIPS 2022
Algorithmic game theory and mechanism design
game solving
1.422025
Rapid Learning in Constrained Minimax Games with Negative Momentum · AAAI 2025
A Unified Diversity Measure for Multiagent Reinforcement Learning · NeurIPS 2022
Machine learning › Reinforcement learning › offline reinforcement learning
conservative value estimation
0.912025
Conservative Offline Goal-Conditioned Implicit V-Learning · ICML 2025
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.912025
Conservative Offline Goal-Conditioned Implicit V-Learning · ICML 2025
Machine learning › Reinforcement learning › off-policy reinforcement learning › experience replay
hindsight experience replay
0.912025
Conservative Offline Goal-Conditioned Implicit V-Learning · ICML 2025
Machine learning › Reinforcement learning › offline reinforcement learning
offline policy optimization
0.912025
Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization · AAAI 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Conservative Offline Goal-Conditioned Implicit V-Learning · ICML 2025
Machine learning › Reinforcement learning
value function estimation
0.912025
Conservative Offline Goal-Conditioned Implicit V-Learning · ICML 2025
Algorithmic game theory and mechanism design › non-cooperative game
extensive-form games
0.912025
Rapid Learning in Constrained Minimax Games with Negative Momentum · AAAI 2025
Mathematical optimization
minimax optimization
0.912025
Rapid Learning in Constrained Minimax Games with Negative Momentum · AAAI 2025
Machine learning › Reinforcement learning
constrained reinforcement learning
0.812024
An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning
multi-objective reinforcement learning
0.812024
An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning
safe reinforcement learning
0.812024
An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning
behavioral diversity
0.612022
A Unified Diversity Measure for Multiagent Reinforcement Learning · NeurIPS 2022
Knowledge, reasoning and agents › Multi-agent systems
equilibrium computation
0.612022
A Unified Diversity Measure for Multiagent Reinforcement Learning · NeurIPS 2022
Knowledge, reasoning and agents › Multi-agent systems › game theory
nash equilibrium
0.612022
A Unified Diversity Measure for Multiagent Reinforcement Learning · NeurIPS 2022
Mathematical optimization
convergence analysis
0.312025
Rapid Learning in Constrained Minimax Games with Negative Momentum · AAAI 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › nonmonotonic reasoning › preference handling › preference reasoning
preference inference
0.212024
An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement Learning · NeurIPS 2024
Machine learning › Optimization for machine learning › hyperparameter optimization
population-based training
0.212022
A Unified Diversity Measure for Multiagent Reinforcement Learning · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

fictitious play · 1.1sequential policy optimization · 0.9quasimetric framework · 0.9quantal response equilibrium · 0.9negative momentum · 0.9momentum buffer updating · 0.9game-solver algorithms · 0.9conservative penalty · 0.9offline reinforcement learning · 0.8multi-objective reinforcement learning · 0.8policy-space response oracle · 0.6policy space response oracles · 0.6
YearPublicationVenuePosition
2025 Rapid Learning in Constrained Minimax Games with Negative Momentum
abstract
In this paper, we delve into the utilization of the negative momentum technique in constrained minimax games. From an intuitive mechanical standpoint, we introduce a novel framework for momentum buffer updating, which extends the findings of negative momentum from the unconstrained setting to the constrained setting and provides a universal enhancement to the classic game-solver algorithms. Additionally, we provide theoretical guarantees of convergence for our momentum-augmented learning algorithms. We then extend these algorithms to their extensive-form counterparts. Experimental results on both Normal Form Games (NFGs) and Extensive Form Games (EFGs) demonstrate that our momentum techniques can significantly improve algorithm performance, surpassing both their original versions and the SOTA baselines by a large margin.
Zijian Fang, Zongkai Liu, Chao Yu 0004, Chaohao Hu
AAAI2
2025 Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization
abstract
Offline Multi-Agent Reinforcement Learning (MARL) is an emerging field that aims to learn optimal multi-agent policies from pre-collected datasets. Compared to single-agent case, multi-agent setting involves a large joint state-action space and coupled behaviors of multiple agents, which bring extra complexity to offline policy optimization. In this work, we revisit the existing offline MARL methods and show that in certain scenarios they can be problematic, leading to uncoordinated behaviors and out-of-distribution (OOD) joint actions. To address these issues, we propose a new offline MARL algorithm, named In-Sample Sequential Policy Optimization (InSPO). InSPO sequentially updates each agent's policy in an in-sample manner, which not only avoids selecting OOD joint actions but also carefully considers teammates' updated policies to enhance coordination. Additionally, by thoroughly exploring low-probability actions in the behavior policy, InSPO can well address the issue of premature convergence to sub-optimal solutions. Theoretically, we prove InSPO guarantees monotonic policy improvement and converges to quantal response equilibrium (QRE). Experimental results demonstrate the effectiveness of our method compared to current state-of-the-art offline MARL methods.
Zongkai Liu, Chao Yu 0004, Xiawei Wu, Yile Liang, Xuetao Ding
AAAI1
2025 Conservative Offline Goal-Conditioned Implicit V-Learning
abstract
Offline goal-conditioned reinforcement learning (GCRL) learns a goal-conditioned value function to train policies for diverse goals with pre-collected datasets. Hindsight experience replay addresses the issue of sparse rewards by treating intermediate states as goals but fails to complete goal-stitching tasks where achieving goals requires stitching different trajectories. While cross-trajectory sampling is a potential solution that associates states and goals belonging to different trajectories, we demonstrate that this direct method degrades performance in goal-conditioned tasks due to the overestimation of values on unconnected pairs. To this end, we propose Conservative Goal-Conditioned Implicit Value Learning (CGCIVL), a novel algorithm that introduces a penalty term to penalize value estimation for unconnected state-goal pairs and leverages the quasimetric framework to accurately estimate values for connected pairs. Evaluations on OGBench, a benchmark for offline GCRL, demonstrate that CGCIVL consistently surpasses state-of-the-art methods across diverse tasks.
Kaiqiang Ke, Zongkai Liu, Shenghong He, Chao Yu 0004
ICML3
2025 Federated Distillation With Lightweight Generative Adversarial Network for Servo Motor Bearing Fault Diagnosis in Heterogeneous Data
abstract
The importance of data privacy has made federated learning a research focal point in the field of fault diagnosis. However, the application of existing methods to the diagnosis of critical components in servo motors is hindered by the heterogeneity of samples across clients, meaning they are not independently and identically distributed, i.e., non-IID. To address this, a federated distillation with generative adversarial network, i.e., FDGAN approach is designed for fault diagnosis under data heterogeneity. The design of FDGAN is based on the concept of knowledge distillation, facilitating the transfer of knowledge from generated data (teacher) to raw data (student). Specifically, for imbalanced datasets, a lightweight generative adversarial network, i.e., LGAN is employed to enhance the raw data. Then, a similarity measurement strategy is devised to uncover the correlation between the raw and generated data, and an attention measurement strategy is implemented to extract critical dependencies between these two types of data. Finally, within the federated framework, the acquired knowledge is used to improve the performance of the global model. The proposed method is comprehensively validated using two sets of rolling bearing data. Experimental results demonstrate that the method effectively addresses fault diagnosis under data heterogeneity while preserving data privacy.
Zongkai Liu, Haidong Shao, Rao Xin
IEEE Internet Things J.1
2024 RegFTRL: Efficient Equilibrium Learning in Two-Player Zero-Sum Games
abstract
Recent literature has witnessed a rising interest in learning Nash equilibrium with a guarantee of last-iterate convergence.In this paper, we introduce a novel approach called Regularized Followthe-Regularized-Leader (RegFTRL), which is an efficient variant of FTRL enriched with an adaptive regularization, for the purpose of learning equilibria in two-player zero-sum games.In the context of normal-form games (NFGs), our proposed RegFTRL algorithm exhibits desirable property of last-iterate linear convergence towards an approximated equilibrium, and converges to an exact Nash equilibrium through adaptive adjustments of the regularization.Moreover, we extend our method to extensive-form games (EFGs) and propose FollowMu, a practical implementation of RegFTRL with a neural network as the function approximator, for model-free learning in sequential non-stationary environments.Finally, empirical results substantiate the theoretical properties of RegFTRL, and demonstrate that FollowMu can achieve favorable performance in EFGs.
Zijian Fang, Zongkai Liu, Chao Yu 0004
DAI2
2024 An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement Learning
abstract
In recent years, significant progress has been made in multi-objective reinforcement learning (RL) research, which aims to balance multiple objectives by incorporating preferences for each objective. In most existing studies, specific preferences must be provided during deployment to indicate the desired policies explicitly. However, designing these preferences depends heavily on human prior knowledge, which is typically obtained through extensive observation of high-performing demonstrations with expected behaviors. In this work, we propose a simple yet effective offline adaptation framework for multi-objective RL problems without assuming handcrafted target preferences, but only given several demonstrations to implicitly indicate the preferences of expected policies. Additionally, we demonstrate that our framework can naturally be extended to meet constraints on safety-critical objectives by utilizing safe demonstrations, even when the safety thresholds are unknown. Empirical results on offline multi-objective and safe tasks demonstrate the capability of our framework to infer policies that align with real preferences while meeting the constraints implied by the provided demonstrations.
Zongkai Liu, Danying Mo, Chao Yu 0004
NeurIPS2
2024 Reinforced fuzzy domain adaptation: Revolutionizing data-unaccessible rotating machinery fault diagnosis across multiple domains
Zongkai Liu, Haidong Shao, Yifan Wan
Expert Syst. Appl.1
2022 A Unified Diversity Measure for Multiagent Reinforcement Learning
abstract
Promoting behavioural diversity is of critical importance in multi-agent reinforcement learning, since it helps the agent population maintain robust performance when encountering unfamiliar opponents at test time, or, when the game is highly non-transitive in the strategy space (e.g., Rock-Paper-Scissor). While a myriad of diversity metrics have been proposed, there are no widely accepted or unified definitions in the literature, making the consequent diversity-aware learning algorithms difficult to evaluate and the insights elusive. In this work, we propose a novel metric called the Unified Diversity Measure (UDM) that offers a unified view for existing diversity metrics. Based on UDM, we design the UDM-Fictitious Play (UDM-FP) and UDM-Policy Space Response Oracle (UDM-PSRO) algorithms as efficient solvers for normal-form games and open-ended games. In theory, we prove that UDM-based methods can enlarge the gamescape by increasing the response capacity of the strategy pool, and have convergence guarantee to two-player Nash equilibrium. We validate our algorithms on games that show strong non-transitivity, and empirical results show that our algorithms achieve better performances than strong PSRO baselines in terms of the exploitability and population effectivity.
Zongkai Liu, Chao Yu 0004, Yaodong Yang 0002, Zifan Wu
NeurIPS1
2022 A Distributed Algorithm for Task Offloading in Vehicular Networks With Hybrid Fog/Cloud Computing
abstract
Fog computing has been an effective paradigm of real-time applications in the IoT area, which enables task offloading at network edge devices. Particularly, many emerging vehicular applications require real-time interaction between the terminal users and computation servers, which can be implemented in fog-based architecture. However, it is still challenging to apply fog computing in vehicular networks due to high mobility of vehicles and uneven distribution of vehicle density, which may result in performance degradation, such as unbalanced workload and unexpected task failure. In this article, we investigate a new service scenario of task offloading under a three-layer service architecture, where the resources of vehicular fog (VF), fog server (FS), and central cloud (CC) are utilized in a cooperative way. On this basis, we formulate the probabilistic task offloading (PTO) problem by synthesizing task transmission, computation, and result retrieval, as well as characterizing the heterogeneity of computation servers. The objective of the PTO is to minimize the weighted sum of execution delay, energy consumption, and payment cost. To resolve the PTO problem, we propose a comprehensive task offloading algorithm by combining the alternating direction method of multipliers (ADMMs) and particle swarm optimization (PSO), called ADMM-PSO. The basic idea of the ADMM-PSO is to divide the PTO problem into multiple unconstrained subproblems and achieve the optimal solution in the form of an iterative coordination process. For each iteration, the solution is achieved by solving each subproblem with the PSO and updated based on a designed rule, which is able to converge to the optimal solution when the stop criterion is satisfied. Finally, we build the simulation model and implement the proposed algorithm for performance evaluation. The simulation results demonstrate the superiority of the proposed algorithm under a wide range of service scenarios.
Zongkai Liu, Penglin Dai, Huanlai Xing, Zhaofei Yu, Wei Zhang 0161
IEEE Trans. Syst. Man Cybern. Syst.1
2017 FPGA-based high-performance time-to-digital converters by utilizing multi-channels looped carry chains
abstract
Time-to-digital converters (TDCs) are core components in many applications and numerous works on this theme have been conducted in recent years. For field programmable gate array (FPGA) based TDCs, their overall performance are still not satisfying when compared with application specific integrated circuit (ASIC) based TDCs. We propose multi-channels looped carry chain TDC architecture in this paper in order to narrow down the performance gap between FPGAs and ASICs. An example TDC prototype implemented on a Stratix III FPGA chip by using the proposed method achieves the resolution below 20 ps, the precision root mean square (RMS) below 15 ps, and the differential non-linearity (DNL) and integral non-linearity (INL) within the range of 2 least significant bit (LSB) peak-to-peak value. This performance is very competitive among all existing FPGA-based designs and close to some ASIC-based TDCs.
Ke Cui, Zongkai Liu, Rihong Zhu, Xiangyu Li 0005
FPT2