EDBT 2026 Demo / reviewers in the wild / expert
Qi Zhang 0038
dblp:52/323-38
· DBLP profile ↗
13ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0002-8562-5987ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 71% Multi-agent systems · 12% Knowledge representation and reasoning · 10% | |
| Theoretical computer science
2 papers |
Algorithmic game theory and mechanism design · 100% |
Topics — the 15 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
2.4 | 4 | 2024 | E(3)-Equivariant Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning · ICML 2024 Context-Aware Bayesian Network Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning · ICML 2023 Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games · ICLR 2022 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning |
1.6 | 3 | 2023 | Context-Aware Bayesian Network Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning · ICML 2023 Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games · ICLR 2022 Learning to Communicate and Solve Visual Blocks-World Tasks · AAAI 2019 |
Machine learning › Reinforcement learning
actor-critic methods |
1.3 | 2 | 2024 | E(3)-Equivariant Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning · ICML 2024 Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games · ICLR 2022 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.7 | 1 | 2023 | Context-Aware Bayesian Network Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning · ICML 2023 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
markov games |
0.6 | 1 | 2022 | Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games · ICLR 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
uncertainty reasoning |
0.4 | 1 | 2020 | Modeling Probabilistic Commitments for Maintenance Is Inherently Harder than for Achievement · AAAI 2020 |
Knowledge, reasoning and agents › Multi-agent systems
emergent communication |
0.4 | 1 | 2019 | Learning to Communicate and Solve Visual Blocks-World Tasks · AAAI 2019 |
Computer vision › Vision and language › visual grounding
language grounding |
0.4 | 1 | 2019 | Learning to Communicate and Solve Visual Blocks-World Tasks · AAAI 2019 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
decision making under uncertainty |
0.2 | 1 | 2016 | Commitment Semantics for Sequential Decision Making under Reward Uncertainty · IJCAI 2016 |
Machine learning › Reinforcement learning
reward uncertainty |
0.2 | 1 | 2016 | Commitment Semantics for Sequential Decision Making under Reward Uncertainty · IJCAI 2016 |
Network optimization and economics › mechanism design › incentive mechanism
crowdsourcing incentive mechanism |
0.2 | 1 | 2015 | Incentivize crowd labeling under budget constraint · INFOCOM 2015 |
Algorithmic game theory and mechanism design › auction theory › auction mechanism
reverse auction |
0.2 | 1 | 2015 | Incentivize crowd labeling under budget constraint · INFOCOM 2015 |
Algorithmic game theory and mechanism design › mechanism design
truthful mechanism |
0.2 | 1 | 2015 | Incentivize crowd labeling under budget constraint · INFOCOM 2015 |
Data mining › crowdsourcing
crowdsourced annotation |
0.1 | 1 | 2015 | Incentivize crowd labeling under budget constraint · INFOCOM 2015 |
Data mining › crowdsourcing
label aggregation |
0.1 | 1 | 2015 | Incentivize crowd labeling under budget constraint · INFOCOM 2015 |
Methods — techniques the papers use, named apart from their topics
actor-critic · 2.0probabilistic reasoning · 1.0probabilistic analysis · 0.9approximate modeling strategies · 0.9e(3) equivariance · 0.8differentiable DAG · 0.7bayesian network · 0.7sequential bayesian aggregation · 0.7reverse auction · 0.7communication-efficient learning · 0.6reinforcement learning · 0.4emergent communication · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | E(3)-Equivariant Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning
Dingyang Chen 0001, Qi Zhang 0038 |
ICML | 2 |
| 2024 | On the Spatio-Temporal Analysis and Optimization of AoI in Cell-Free IIoT NetworksabstractCell-free massive multiple-input multiple-output (mMIMO) architecture is a promising solution for Industrial Internet of Things (IIoT) because it not only provides massive connectivity but also eliminates the traditional cell edges. Considering the heterogeneous traffic and requirements in the industry, in this paper, we propose a device priority-aware resource allocation policy under cell-free mMIMO IIoT networks. Specifically, we design a priority-aware frame structure that can be used to provide differentiated age of information (AoI) guarantees for devices of different priorities and locations. To characterize the proposed policy, we develop a general analysis framework to evaluate the signal-to-interference ratio meta distribution and the average AoI of a generic device. The framework captures multiple main features under wireless IIoT networks, including cell-free mMIMO architecture, frame structure, finite-sized geographic areas, densely deployed devices, device priority, retransmission, and interaction among different transmission links. The analytical framework is validated by simulations. Based on the analysis, we study a mean-variance optimization problem to improve the network average AoI, while guaranteeing the average AoI per device. Numerical results show that the proposed frame structure works effectively in enhancing the AoI performance of cell-free IIoT networks. Meiyan Song, Hangguan Shan, Yu Cheng 0003, Weihua Zhuang, Xinyu Li 0001, Qi Zhang 0038, Xianhua He |
IEEE Trans. Wirel. Commun. | 6 |
| 2023 | Context-Aware Bayesian Network Actor-Critic Methods for Cooperative Multi-Agent Reinforcement LearningabstractExecuting actions in a correlated manner is a common strategy for human coordination that often leads to better cooperation, which is also potentially beneficial for cooperative multi-agent reinforcement learning (MARL). However, the recent success of MARL relies heavily on the convenient paradigm of purely decentralized execution, where there is no action correlation among agents for scalability considerations. In this work, we introduce a Bayesian network to inaugurate correlations between agents’ action selections in their joint policy. Theoretically, we establish a theoretical justification for why action dependencies are beneficial by deriving the multi-agent policy gradient formula under such a Bayesian network joint policy and proving its global convergence to Nash equilibria under tabular softmax policy parameterization in cooperative Markov games. Further, by equipping existing MARL algorithms with a recent method of differentiable directed acyclic graphs (DAGs), we develop practical algorithms to learn the context-aware Bayesian network policies in scenarios with partial observability and various difficulty. We also dynamically decrease the sparsity of the learned DAG throughout the training process, which leads to weakly or even purely independent policies for decentralized execution. Empirical results on a range of MARL benchmarks show the benefits of our approach. Dingyang Chen 0001, Qi Zhang 0038 |
ICML | 2 |
| 2023 | Risk-aware analysis for interpretations of probabilistic achievement and maintenance commitments
Qi Zhang 0038, Edmund H. Durfee, Satinder Singh 0001 |
Artif. Intell. | 1 |
| 2022 | Communication-Efficient Actor-Critic Methods for Homogeneous Markov Games
Dingyang Chen 0001, Yile Li, Qi Zhang 0038 |
ICLR | 3 |
| 2022 | Ensemble Policy Distillation with Reduced Data Distribution MismatchabstractPolicy distillation is a method for model compression for deep reinforcement learning, which is typically applied onto mobile devices to reduce power consumption and inference time. However, achieving full and stable distilled policies is challenging, which impedes higher compression ratios. In this work, we develop two policy distillation algorithms to address this problem. Our first algorithm, Ensemble Policy Distillation (EPD), incorporates the idea from supervised learning distillation that uses an ensemble of teacher networks to provide diverse supervision for a compact student policy network. In the Deep Q-Network (DQN) framework, our experiments verify that highly compressed student networks distilled using EPD even outperform teachers for numerous Atari games. Additionally, we analyze how the issue of data distribution mismatch caused by the teacher ensemble in EPD negatively impacts teachers' learning, and introduce the second algorithm, Double Policy Distillation (DPD), as a novel method to mitigate the distribution mismatch. Empirical results show that DPD improves both the teachers' learning and the student's distillation in Atari games and continuous control tasks. Qi Zhang 0038 |
IJCNN | 2 |
| 2021 | Efficient Querying for Cooperative Probabilistic Commitments
Qi Zhang 0038, Edmund H. Durfee, Satinder Singh 0001 |
AAAI | 1 |
| 2021 | Knowledge Infused Policy Gradients with Upper Confidence Bound for Relational Bandits
Kaushik Roy 0009, Qi Zhang 0038, Manas Gaur, Amit P. Sheth |
ECML/PKDD (1) | 2 |
| 2020 | Modeling Probabilistic Commitments for Maintenance Is Inherently Harder than for AchievementabstractMost research on probabilistic commitments focuses on commitments to achieve enabling preconditions for other agents. Our work reveals that probabilistic commitments to instead maintain preconditions for others are surprisingly harder to use well than their achievement counterparts, despite strong semantic similarities. We isolate the key difference as being not in how the commitment provider is constrained, but rather in how the commitment recipient can locally use the commitment specification to approximately model the provider's effects on the preconditions of interest. Our theoretic analyses show that we can more tightly bound the potential suboptimality due to approximate modeling for achievement than for maintenance commitments. We empirically evaluate alternative approximate modeling strategies, confirming that probabilistic maintenance commitments are qualitatively more challenging for the recipient to model well, and indicating the need for more detailed specifications that can sacrifice some of the agents' autonomy. Qi Zhang 0038, Edmund H. Durfee, Satinder Singh 0001 |
AAAI | 1 |
| 2020 | Semantics and algorithms for trustworthy commitment achievement under model uncertainty
Qi Zhang 0038, Edmund H. Durfee, Satinder Singh 0001 |
Auton. Agents Multi Agent Syst. | 1 |
| 2019 | Learning to Communicate and Solve Visual Blocks-World Tasks
Qi Zhang 0038, Richard L. Lewis, Satinder Singh 0001, Edmund H. Durfee |
AAAI | 1 |
| 2016 | Commitment Semantics for Sequential Decision Making under Reward Uncertainty
Qi Zhang 0038, Edmund H. Durfee, Satinder Singh 0001, Anna Chen, Stefan J. Witwicki |
IJCAI | 1 |
| 2015 | Incentivize crowd labeling under budget constraintabstractCrowdsourcing systems allocate tasks to a group of workers over the Internet, which have become an effective paradigm for human-powered problem solving such as image classification, optical character recognition and proofreading. In this paper, we focus on incentivizing crowd workers to label a set of binary tasks under strict budget constraint. We properly profile the tasks' difficulty levels and workers' quality in crowdsourcing systems, where the collected labels are aggregated with sequential Bayesian approach. To stimulate workers to undertake crowd labeling tasks, the interaction between workers and the platform is modeled as a reverse auction. We reveal that the platform utility maximization could be intractable, for which an incentive mechanism that determines the winning bid and payments with polynomial-time computation complexity is developed. Moreover, we theoretically prove that our mechanism is truthful, individually rational and budget feasible. Through extensive simulations, we demonstrate that our mechanism utilizes budget efficiently to achieve high platform utility with polynomial computation complexity. Qi Zhang 0038, Yutian Wen, Xiaohua Tian, Xiaoying Gan, Xinbing Wang |
INFOCOM | 1 |