Qi Zhang 0038

dblp:52/323-38 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0002-8562-5987ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 71% Multi-agent systems · 12% Knowledge representation and reasoning · 10%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 100%

Topics — the 15 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
2.442024
E(3)-Equivariant Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning · ICML 2024
Context-Aware Bayesian Network Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning · ICML 2023
Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games · ICLR 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning
1.632023
Context-Aware Bayesian Network Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning · ICML 2023
Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games · ICLR 2022
Learning to Communicate and Solve Visual Blocks-World Tasks · AAAI 2019
Machine learning › Reinforcement learning
actor-critic methods
1.322024
E(3)-Equivariant Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning · ICML 2024
Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games · ICLR 2022
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.712023
Context-Aware Bayesian Network Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning › multi-agent reinforcement learning
markov games
0.612022
Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games · ICLR 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
uncertainty reasoning
0.412020
Modeling Probabilistic Commitments for Maintenance Is Inherently Harder than for Achievement · AAAI 2020
Knowledge, reasoning and agents › Multi-agent systems
emergent communication
0.412019
Learning to Communicate and Solve Visual Blocks-World Tasks · AAAI 2019
Computer vision › Vision and language › visual grounding
language grounding
0.412019
Learning to Communicate and Solve Visual Blocks-World Tasks · AAAI 2019
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
decision making under uncertainty
0.212016
Commitment Semantics for Sequential Decision Making under Reward Uncertainty · IJCAI 2016
Machine learning › Reinforcement learning
reward uncertainty
0.212016
Commitment Semantics for Sequential Decision Making under Reward Uncertainty · IJCAI 2016
Network optimization and economics › mechanism design › incentive mechanism
crowdsourcing incentive mechanism
0.212015
Incentivize crowd labeling under budget constraint · INFOCOM 2015
Algorithmic game theory and mechanism design › auction theory › auction mechanism
reverse auction
0.212015
Incentivize crowd labeling under budget constraint · INFOCOM 2015
Algorithmic game theory and mechanism design › mechanism design
truthful mechanism
0.212015
Incentivize crowd labeling under budget constraint · INFOCOM 2015
Data mining › crowdsourcing
crowdsourced annotation
0.112015
Incentivize crowd labeling under budget constraint · INFOCOM 2015
Data mining › crowdsourcing
label aggregation
0.112015
Incentivize crowd labeling under budget constraint · INFOCOM 2015

Methods — techniques the papers use, named apart from their topics

actor-critic · 2.0probabilistic reasoning · 1.0probabilistic analysis · 0.9approximate modeling strategies · 0.9e(3) equivariance · 0.8differentiable DAG · 0.7bayesian network · 0.7sequential bayesian aggregation · 0.7reverse auction · 0.7communication-efficient learning · 0.6reinforcement learning · 0.4emergent communication · 0.4
YearPublicationVenuePosition
2024 E(3)-Equivariant Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning
Dingyang Chen 0001, Qi Zhang 0038
ICML2
2024 On the Spatio-Temporal Analysis and Optimization of AoI in Cell-Free IIoT Networks
abstract
Cell-free massive multiple-input multiple-output (mMIMO) architecture is a promising solution for Industrial Internet of Things (IIoT) because it not only provides massive connectivity but also eliminates the traditional cell edges. Considering the heterogeneous traffic and requirements in the industry, in this paper, we propose a device priority-aware resource allocation policy under cell-free mMIMO IIoT networks. Specifically, we design a priority-aware frame structure that can be used to provide differentiated age of information (AoI) guarantees for devices of different priorities and locations. To characterize the proposed policy, we develop a general analysis framework to evaluate the signal-to-interference ratio meta distribution and the average AoI of a generic device. The framework captures multiple main features under wireless IIoT networks, including cell-free mMIMO architecture, frame structure, finite-sized geographic areas, densely deployed devices, device priority, retransmission, and interaction among different transmission links. The analytical framework is validated by simulations. Based on the analysis, we study a mean-variance optimization problem to improve the network average AoI, while guaranteeing the average AoI per device. Numerical results show that the proposed frame structure works effectively in enhancing the AoI performance of cell-free IIoT networks.
Meiyan Song, Hangguan Shan, Yu Cheng 0003, Weihua Zhuang, Xinyu Li 0001, Qi Zhang 0038, Xianhua He
IEEE Trans. Wirel. Commun.6
2023 Context-Aware Bayesian Network Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning
abstract
Executing actions in a correlated manner is a common strategy for human coordination that often leads to better cooperation, which is also potentially beneficial for cooperative multi-agent reinforcement learning (MARL). However, the recent success of MARL relies heavily on the convenient paradigm of purely decentralized execution, where there is no action correlation among agents for scalability considerations. In this work, we introduce a Bayesian network to inaugurate correlations between agents’ action selections in their joint policy. Theoretically, we establish a theoretical justification for why action dependencies are beneficial by deriving the multi-agent policy gradient formula under such a Bayesian network joint policy and proving its global convergence to Nash equilibria under tabular softmax policy parameterization in cooperative Markov games. Further, by equipping existing MARL algorithms with a recent method of differentiable directed acyclic graphs (DAGs), we develop practical algorithms to learn the context-aware Bayesian network policies in scenarios with partial observability and various difficulty. We also dynamically decrease the sparsity of the learned DAG throughout the training process, which leads to weakly or even purely independent policies for decentralized execution. Empirical results on a range of MARL benchmarks show the benefits of our approach.
Dingyang Chen 0001, Qi Zhang 0038
ICML2
2023 Risk-aware analysis for interpretations of probabilistic achievement and maintenance commitments
Qi Zhang 0038, Edmund H. Durfee, Satinder Singh 0001
Artif. Intell.1
2022 Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games
Dingyang Chen 0001, Yile Li, Qi Zhang 0038
ICLR3
2022 Ensemble Policy Distillation with Reduced Data Distribution Mismatch
abstract
Policy distillation is a method for model compression for deep reinforcement learning, which is typically applied onto mobile devices to reduce power consumption and inference time. However, achieving full and stable distilled policies is challenging, which impedes higher compression ratios. In this work, we develop two policy distillation algorithms to address this problem. Our first algorithm, Ensemble Policy Distillation (EPD), incorporates the idea from supervised learning distillation that uses an ensemble of teacher networks to provide diverse supervision for a compact student policy network. In the Deep Q-Network (DQN) framework, our experiments verify that highly compressed student networks distilled using EPD even outperform teachers for numerous Atari games. Additionally, we analyze how the issue of data distribution mismatch caused by the teacher ensemble in EPD negatively impacts teachers' learning, and introduce the second algorithm, Double Policy Distillation (DPD), as a novel method to mitigate the distribution mismatch. Empirical results show that DPD improves both the teachers' learning and the student's distillation in Atari games and continuous control tasks.
Qi Zhang 0038
IJCNN2
2021 Efficient Querying for Cooperative Probabilistic Commitments
Qi Zhang 0038, Edmund H. Durfee, Satinder Singh 0001
AAAI1
2021 Knowledge Infused Policy Gradients with Upper Confidence Bound for Relational Bandits
Kaushik Roy 0009, Qi Zhang 0038, Manas Gaur, Amit P. Sheth
ECML/PKDD (1)2
2020 Modeling Probabilistic Commitments for Maintenance Is Inherently Harder than for Achievement
abstract
Most research on probabilistic commitments focuses on commitments to achieve enabling preconditions for other agents. Our work reveals that probabilistic commitments to instead maintain preconditions for others are surprisingly harder to use well than their achievement counterparts, despite strong semantic similarities. We isolate the key difference as being not in how the commitment provider is constrained, but rather in how the commitment recipient can locally use the commitment specification to approximately model the provider's effects on the preconditions of interest. Our theoretic analyses show that we can more tightly bound the potential suboptimality due to approximate modeling for achievement than for maintenance commitments. We empirically evaluate alternative approximate modeling strategies, confirming that probabilistic maintenance commitments are qualitatively more challenging for the recipient to model well, and indicating the need for more detailed specifications that can sacrifice some of the agents' autonomy.
Qi Zhang 0038, Edmund H. Durfee, Satinder Singh 0001
AAAI1
2020 Semantics and algorithms for trustworthy commitment achievement under model uncertainty
Qi Zhang 0038, Edmund H. Durfee, Satinder Singh 0001
Auton. Agents Multi Agent Syst.1
2019 Learning to Communicate and Solve Visual Blocks-World Tasks
Qi Zhang 0038, Richard L. Lewis, Satinder Singh 0001, Edmund H. Durfee
AAAI1
2016 Commitment Semantics for Sequential Decision Making under Reward Uncertainty
Qi Zhang 0038, Edmund H. Durfee, Satinder Singh 0001, Anna Chen, Stefan J. Witwicki
IJCAI1
2015 Incentivize crowd labeling under budget constraint
abstract
Crowdsourcing systems allocate tasks to a group of workers over the Internet, which have become an effective paradigm for human-powered problem solving such as image classification, optical character recognition and proofreading. In this paper, we focus on incentivizing crowd workers to label a set of binary tasks under strict budget constraint. We properly profile the tasks' difficulty levels and workers' quality in crowdsourcing systems, where the collected labels are aggregated with sequential Bayesian approach. To stimulate workers to undertake crowd labeling tasks, the interaction between workers and the platform is modeled as a reverse auction. We reveal that the platform utility maximization could be intractable, for which an incentive mechanism that determines the winning bid and payments with polynomial-time computation complexity is developed. Moreover, we theoretically prove that our mechanism is truthful, individually rational and budget feasible. Through extensive simulations, we demonstrate that our mechanism utilizes budget efficiently to achieve high platform utility with polynomial computation complexity.
Qi Zhang 0038, Yutian Wen, Xiaohua Tian, Xiaoying Gan, Xinbing Wang
INFOCOM1