EDBT 2026 Demo / reviewers in the wild / expert
Kuangqi Zhou
dblp:267/5502
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 74% Learning paradigms · 13% Representation and self-supervised learning · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reward design
reward shaping |
1.2 | 2 | 2023 | Reachability-Aware Laplacian Representation in Reinforcement Learning · ICML 2023 Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph Drawing · ICML 2021 |
Machine learning › Reinforcement learning
exploration |
0.7 | 1 | 2023 | Revisiting Intrinsic Reward for Exploration in Procedurally Generated Environments · ICLR 2023 |
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.7 | 1 | 2023 | Revisiting Intrinsic Reward for Exploration in Procedurally Generated Environments · ICLR 2023 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation |
0.7 | 1 | 2023 | Reachability-Aware Laplacian Representation in Reinforcement Learning · ICML 2023 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.6 | 1 | 2022 | Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental Learning · CVPR 2022 |
Machine learning › Learning paradigms › continual learning
class-incremental learning |
0.6 | 1 | 2022 | Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental Learning · CVPR 2022 |
Machine learning › Representation and self-supervised learning › representation learning
feature decorrelation |
0.6 | 1 | 2022 | Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental Learning · CVPR 2022 |
Machine learning › Reinforcement learning › robust reinforcement learning
robust markov decision process |
0.6 | 1 | 2022 | The Geometry of Robust Value Functions · ICML 2022 |
Machine learning › Reinforcement learning
robust reinforcement learning |
0.6 | 1 | 2022 | The Geometry of Robust Value Functions · ICML 2022 |
Machine learning › Reinforcement learning
value function |
0.6 | 1 | 2022 | The Geometry of Robust Value Functions · ICML 2022 |
Recommender systems › reinforcement-learning-based recommendation
offline reinforcement learning for recommendation |
0.6 | 1 | 2022 | Value Penalized Q-Learning for Recommender Systems · SIGIR 2022 |
Recommender systems
reinforcement-learning-based recommendation |
0.6 | 1 | 2022 | Value Penalized Q-Learning for Recommender Systems · SIGIR 2022 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › spectral manifold learning
laplacian embedding |
0.5 | 1 | 2021 | Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph Drawing · ICML 2021 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery |
0.5 | 1 | 2021 | Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph Drawing · ICML 2021 |
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning |
0.5 | 1 | 2021 | Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph Drawing · ICML 2021 |
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
procedural environment generation |
0.2 | 1 | 2023 | Revisiting Intrinsic Reward for Exploration in Procedurally Generated Environments · ICLR 2023 |
Machine learning › Reinforcement learning › offline reinforcement learning
conservative q-learning |
0.2 | 1 | 2022 | Value Penalized Q-Learning for Recommender Systems · SIGIR 2022 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.2 | 1 | 2022 | Value Penalized Q-Learning for Recommender Systems · SIGIR 2022 |
Methods — techniques the papers use, named apart from their topics
uncertainty estimation · 1.1q-learning · 1.1spectral methods · 0.7laplacian representation · 0.7intrinsic motivation · 0.7uncertainty sets · 0.6representation regularization · 0.6conic hypersurface analysis · 0.6spectral graph drawing · 0.5eigenvector decomposition · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Online Learning for Non-monotone DR-Submodular Maximization: From Full Information to Bandit FeedbackabstractIn this paper, we revisit the online non-monotone continuous DR-submodular maximization problem over a down-closed convex set, which finds wide real-world applications in the domain of machine learning, economics, and operations research. At first, we present the Meta-MFW algorithm achieving a $1/e$-regret of $O(\sqrt{T})$ at the cost of $T^{3/2}$ stochastic gradient evaluations per round. As far as we know, Meta-MFW is the first algorithm to obtain $1/e$-regret of $O(\sqrt{T})$ for the online non-monotone continuous DR-submodular maximization problem over a down-closed convex set. Furthermore, in sharp contrast with ODC algorithm (Thang $&$ Srivastav, 2021), Meta-MFW relies on the simple online linear oracle without discretization, lifting, or rounding operations. Considering the practical restrictions, we then propose the Mono-MFW algorithm, which reduces the per-function stochastic gradient evaluations from $T^{3/2}$ to 1 and achieves a $1/e$-regret bound of $O(T^{4/5})$. Next, we extend Mono-MFW to the bandit setting and propose the Bandit-MFW algorithm which attains a $1/e$-regret bound of $O(T^{8/9})$. To the best of our knowledge, Mono-MFW and Bandit-MFW are the first sublinear-regret algorithms to explore the one-shot and bandit setting for online non-monotone continuous DR-submodular maximization problem over a down-closed convex set, respectively. Finally, we conduct numerical experiments on both synthetic and real-world datasets to verify the effectiveness of our methods. Qixin Zhang 0001, Zengde Deng, Zaiyi Chen, Kuangqi Zhou, Haoyuan Hu, Yu Yang 0001 |
AISTATS | 4 |
| 2023 | Revisiting Intrinsic Reward for Exploration in Procedurally Generated Environments
Kuangqi Zhou, Bingyi Kang, Jiashi Feng, Shuicheng Yan |
ICLR | 2 |
| 2023 | Reachability-Aware Laplacian Representation in Reinforcement LearningabstractIn Reinforcement Learning (RL), Laplacian Representation (LapRep) is a task-agnostic state representation that encodes the geometry of the environment. A desirable property of LapRep stated in prior works is that the Euclidean distance in the LapRep space roughly reflects the reachability between states, which motivates the usage of this distance for reward shaping. However, we find that LapRep does not necessarily have this property in general: two states having a small distance under LapRep can actually be far away in the environment. Such a mismatch would impede the learning process in reward shaping. To fix this issue, we introduce a Reachability-Aware Laplacian Representation (RA-LapRep), by properly scaling each dimension of LapRep. Despite the simplicity, we demonstrate that RA-LapRep can better capture the inter-state reachability as compared to LapRep, through both theoretical explanations and experimental results. Additionally, we show that this improvement yields a significant boost in reward shaping performance and benefits bottleneck state discovery. Kuangqi Zhou, Jiashi Feng, Bryan Hooi, Xinchao Wang |
ICML | 2 |
| 2022 | Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental LearningabstractClass Incremental Learning (CIL) aims at learning a classifier in a phase-by-phase manner, in which only data of a subset of the classes are provided at each phase. Previous works mainly focus on mitigating forgetting in phases after the initial one. However, we find that improving CIL at its initial phase is also a promising direction. Specifically, we experimentally show that directly encouraging CIL Learner at the initial phase to output similar representations as the model jointly trained on all classes can greatly boost the CIL performance. Motivated by this, we study the differ-ence between a naively-trained initial-phase model and the oracle model. Specifically, since one major difference be-tween these two models is the number of training classes, we investigate how such difference affects the model rep-resentations. We find that, with fewer training classes, the data representations of each class lie in a long and narrow region; with more training classes, the representations of each class scatter more uniformly. Inspired by this obser-vation, we propose Class-wise Decorrelation (CwD) that ef-fectively regularizes representations of each class to scatter more uniformly, thus mimicking the model jointly trained with all classes (i.e., the oracle model). Our CwD is simple to implement and easy to plug into existing methods. Ex-tensive experiments on various benchmark datasets show that CwD consistently and significantly improves the per-formance of existing state-of-the-art methods by around 1% to 3%. Code: https://github.com/Yujun-Shi/CwD. Yujun Shi, Kuangqi Zhou, Jian Liang 0001, Zihang Jiang, Jiashi Feng, Philip Torr 0001, Song Bai 0001, Vincent Y. F. Tan |
CVPR | 2 |
| 2022 | The Geometry of Robust Value FunctionsabstractThe space of value functions is a fundamental concept in reinforcement learning. Characterizing its geometric properties may provide insights for optimization and representation. Existing works mainly focus on the value space for Markov Decision Processes (MDPs). In this paper, we study the geometry of the robust value space for the more general Robust MDPs (RMDPs) setting, where transition uncertainties are considered. Specifically, since we find it hard to directly adapt prior approaches to RMDPs, we start with revisiting the non-robust case, and introduce a new perspective that enables us to characterize both the non-robust and robust value space in a similar fashion. The key of this perspective is to decompose the value space, in a state-wise manner, into unions of hypersurfaces. Through our analysis, we show that the robust value space is determined by a set of conic hypersurfaces, each of which contains the robust values of all policies that agree on one state. Furthermore, we find that taking only extreme points in the uncertainty set is sufficient to determine the robust value space. Finally, we discuss some other aspects about the robust value space, including its non-convexity and policy agreement on multiple states. Navdeep Kumar, Kuangqi Zhou, Bryan Hooi, Jiashi Feng, Shie Mannor |
ICML | 3 |
| 2022 | Value Penalized Q-Learning for Recommender SystemsabstractScaling reinforcement learning (RL) to recommender systems (RS) is promising since maximizing the expected cumulative rewards for RL agents meets the objective of RS, i.e., improving customers' long-term satisfaction. A key approach to this goal is offline RL, which aims to learn policies from logged data rather than expensive online interactions. In this paper, we propose Value Penalized Q-learning (VPQ), a novel uncertainty-based offline RL algorithm that penalizes the unstable Q-values in the regression target using uncertainty-aware weights, achieving the conservative Q-function without the need of estimating the behavior policy, suitable for RS with a large number of items. Experiments on two real-world datasets show the proposed method serves as a gain plug-in for existing RS models. Chengqian Gao, Ke Xu 0002, Kuangqi Zhou, Lanqing Li, Xueqian Wang 0001, Bo Yuan 0008, Peilin Zhao |
SIGIR | 3 |
| 2021 | Understanding and Resolving Performance Degradation in Deep Graph Convolutional NetworksabstractA Graph Convolutional Network (GCN) stacks several layers and in each layer performs a PROPagation operation~(PROP) and a TRANsformation operation~(TRAN) for learning node representations over graph-structured data. Though powerful, GCNs tend to suffer performance drop when the model gets deep. Previous works focus on PROPs to study and mitigate this issue, but the role of TRANs is barely investigated. In this work, we study performance degradation of GCNs by experimentally examining how stacking only TRANs or PROPs works. We find that TRANs contribute significantly, or even more than PROPs, to declining performance, and moreover that they tend to amplify node-wise feature variance in GCNs, causing variance inflammation that we identify as a key factor for causing performance drop. Motivated by such observations, we propose a variance-controlling technique termed Node Normalization (NodeNorm), which scales each node's features using its own standard deviation. Experimental results validate the effectiveness of NodeNorm on addressing performance degradation of GCNs. Specifically, it enables deep GCNs to outperform shallow ones in cases where deep models are needed, and to achieve comparable results with shallow ones on 6 benchmark datasets. NodeNorm is a generic plug-in and can well generalize to other GNN architectures. Code is publicly available at https://github.com/miafei/NodeNorm. Kuangqi Zhou, Yanfei Dong, Wee Sun Lee, Bryan Hooi, Huan Xu 0001, Jiashi Feng |
CIKM | 1 |
| 2021 | Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph DrawingabstractThe Laplacian representation recently gains increasing attention for reinforcement learning as it provides succinct and informative representation for states, by taking the eigenvectors of the Laplacian matrix of the state-transition graph as state embeddings. Such representation captures the geometry of the underlying state space and is beneficial to RL tasks such as option discovery and reward shaping. To approximate the Laplacian representation in large (or even continuous) state spaces, recent works propose to minimize a spectral graph drawing objective, which however has infinitely many global minimizers other than the eigenvectors. As a result, their learned Laplacian representation may differ from the ground truth. To solve this problem, we reformulate the graph drawing objective into a generalized form and derive a new learning objective, which is proved to have eigenvectors as its unique global minimizer. It enables learning high-quality Laplacian representations that faithfully approximate the ground truth. We validate this via comprehensive experiments on a set of gridworld and continuous control environments. Moreover, we show that our learned Laplacian representations lead to more exploratory options and better reward shaping. Kuangqi Zhou, Qixin Zhang 0001, Jie Shao 0006, Bryan Hooi, Jiashi Feng |
ICML | 2 |