Kuangqi Zhou

dblp:267/5502 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 74% Learning paradigms · 13% Representation and self-supervised learning · 12%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › reward design
reward shaping
1.222023
Reachability-Aware Laplacian Representation in Reinforcement Learning · ICML 2023
Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph Drawing · ICML 2021
Machine learning › Reinforcement learning
exploration
0.712023
Revisiting Intrinsic Reward for Exploration in Procedurally Generated Environments · ICLR 2023
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.712023
Revisiting Intrinsic Reward for Exploration in Procedurally Generated Environments · ICLR 2023
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation
0.712023
Reachability-Aware Laplacian Representation in Reinforcement Learning · ICML 2023
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.612022
Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental Learning · CVPR 2022
Machine learning › Learning paradigms › continual learning
class-incremental learning
0.612022
Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental Learning · CVPR 2022
Machine learning › Representation and self-supervised learning › representation learning
feature decorrelation
0.612022
Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental Learning · CVPR 2022
Machine learning › Reinforcement learning › robust reinforcement learning
robust markov decision process
0.612022
The Geometry of Robust Value Functions · ICML 2022
Machine learning › Reinforcement learning
robust reinforcement learning
0.612022
The Geometry of Robust Value Functions · ICML 2022
Machine learning › Reinforcement learning
value function
0.612022
The Geometry of Robust Value Functions · ICML 2022
Recommender systems › reinforcement-learning-based recommendation
offline reinforcement learning for recommendation
0.612022
Value Penalized Q-Learning for Recommender Systems · SIGIR 2022
Recommender systems
reinforcement-learning-based recommendation
0.612022
Value Penalized Q-Learning for Recommender Systems · SIGIR 2022
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › spectral manifold learning
laplacian embedding
0.512021
Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph Drawing · ICML 2021
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery
0.512021
Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph Drawing · ICML 2021
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
0.512021
Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph Drawing · ICML 2021
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
procedural environment generation
0.212023
Revisiting Intrinsic Reward for Exploration in Procedurally Generated Environments · ICLR 2023
Machine learning › Reinforcement learning › offline reinforcement learning
conservative q-learning
0.212022
Value Penalized Q-Learning for Recommender Systems · SIGIR 2022
Machine learning › Reinforcement learning
offline reinforcement learning
0.212022
Value Penalized Q-Learning for Recommender Systems · SIGIR 2022

Methods — techniques the papers use, named apart from their topics

uncertainty estimation · 1.1q-learning · 1.1spectral methods · 0.7laplacian representation · 0.7intrinsic motivation · 0.7uncertainty sets · 0.6representation regularization · 0.6conic hypersurface analysis · 0.6spectral graph drawing · 0.5eigenvector decomposition · 0.5
YearPublicationVenuePosition
2023 Online Learning for Non-monotone DR-Submodular Maximization: From Full Information to Bandit Feedback
abstract
In this paper, we revisit the online non-monotone continuous DR-submodular maximization problem over a down-closed convex set, which finds wide real-world applications in the domain of machine learning, economics, and operations research. At first, we present the Meta-MFW algorithm achieving a $1/e$-regret of $O(\sqrt{T})$ at the cost of $T^{3/2}$ stochastic gradient evaluations per round. As far as we know, Meta-MFW is the first algorithm to obtain $1/e$-regret of $O(\sqrt{T})$ for the online non-monotone continuous DR-submodular maximization problem over a down-closed convex set. Furthermore, in sharp contrast with ODC algorithm (Thang $&$ Srivastav, 2021), Meta-MFW relies on the simple online linear oracle without discretization, lifting, or rounding operations. Considering the practical restrictions, we then propose the Mono-MFW algorithm, which reduces the per-function stochastic gradient evaluations from $T^{3/2}$ to 1 and achieves a $1/e$-regret bound of $O(T^{4/5})$. Next, we extend Mono-MFW to the bandit setting and propose the Bandit-MFW algorithm which attains a $1/e$-regret bound of $O(T^{8/9})$. To the best of our knowledge, Mono-MFW and Bandit-MFW are the first sublinear-regret algorithms to explore the one-shot and bandit setting for online non-monotone continuous DR-submodular maximization problem over a down-closed convex set, respectively. Finally, we conduct numerical experiments on both synthetic and real-world datasets to verify the effectiveness of our methods.
Qixin Zhang 0001, Zengde Deng, Zaiyi Chen, Kuangqi Zhou, Haoyuan Hu, Yu Yang 0001
AISTATS4
2023 Revisiting Intrinsic Reward for Exploration in Procedurally Generated Environments
Kuangqi Zhou, Bingyi Kang, Jiashi Feng, Shuicheng Yan
ICLR2
2023 Reachability-Aware Laplacian Representation in Reinforcement Learning
abstract
In Reinforcement Learning (RL), Laplacian Representation (LapRep) is a task-agnostic state representation that encodes the geometry of the environment. A desirable property of LapRep stated in prior works is that the Euclidean distance in the LapRep space roughly reflects the reachability between states, which motivates the usage of this distance for reward shaping. However, we find that LapRep does not necessarily have this property in general: two states having a small distance under LapRep can actually be far away in the environment. Such a mismatch would impede the learning process in reward shaping. To fix this issue, we introduce a Reachability-Aware Laplacian Representation (RA-LapRep), by properly scaling each dimension of LapRep. Despite the simplicity, we demonstrate that RA-LapRep can better capture the inter-state reachability as compared to LapRep, through both theoretical explanations and experimental results. Additionally, we show that this improvement yields a significant boost in reward shaping performance and benefits bottleneck state discovery.
Kuangqi Zhou, Jiashi Feng, Bryan Hooi, Xinchao Wang
ICML2
2022 Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental Learning
abstract
Class Incremental Learning (CIL) aims at learning a classifier in a phase-by-phase manner, in which only data of a subset of the classes are provided at each phase. Previous works mainly focus on mitigating forgetting in phases after the initial one. However, we find that improving CIL at its initial phase is also a promising direction. Specifically, we experimentally show that directly encouraging CIL Learner at the initial phase to output similar representations as the model jointly trained on all classes can greatly boost the CIL performance. Motivated by this, we study the differ-ence between a naively-trained initial-phase model and the oracle model. Specifically, since one major difference be-tween these two models is the number of training classes, we investigate how such difference affects the model rep-resentations. We find that, with fewer training classes, the data representations of each class lie in a long and narrow region; with more training classes, the representations of each class scatter more uniformly. Inspired by this obser-vation, we propose Class-wise Decorrelation (CwD) that ef-fectively regularizes representations of each class to scatter more uniformly, thus mimicking the model jointly trained with all classes (i.e., the oracle model). Our CwD is simple to implement and easy to plug into existing methods. Ex-tensive experiments on various benchmark datasets show that CwD consistently and significantly improves the per-formance of existing state-of-the-art methods by around 1% to 3%. Code: https://github.com/Yujun-Shi/CwD.
Yujun Shi, Kuangqi Zhou, Jian Liang 0001, Zihang Jiang, Jiashi Feng, Philip Torr 0001, Song Bai 0001, Vincent Y. F. Tan
CVPR2
2022 The Geometry of Robust Value Functions
abstract
The space of value functions is a fundamental concept in reinforcement learning. Characterizing its geometric properties may provide insights for optimization and representation. Existing works mainly focus on the value space for Markov Decision Processes (MDPs). In this paper, we study the geometry of the robust value space for the more general Robust MDPs (RMDPs) setting, where transition uncertainties are considered. Specifically, since we find it hard to directly adapt prior approaches to RMDPs, we start with revisiting the non-robust case, and introduce a new perspective that enables us to characterize both the non-robust and robust value space in a similar fashion. The key of this perspective is to decompose the value space, in a state-wise manner, into unions of hypersurfaces. Through our analysis, we show that the robust value space is determined by a set of conic hypersurfaces, each of which contains the robust values of all policies that agree on one state. Furthermore, we find that taking only extreme points in the uncertainty set is sufficient to determine the robust value space. Finally, we discuss some other aspects about the robust value space, including its non-convexity and policy agreement on multiple states.
Navdeep Kumar, Kuangqi Zhou, Bryan Hooi, Jiashi Feng, Shie Mannor
ICML3
2022 Value Penalized Q-Learning for Recommender Systems
abstract
Scaling reinforcement learning (RL) to recommender systems (RS) is promising since maximizing the expected cumulative rewards for RL agents meets the objective of RS, i.e., improving customers' long-term satisfaction. A key approach to this goal is offline RL, which aims to learn policies from logged data rather than expensive online interactions. In this paper, we propose Value Penalized Q-learning (VPQ), a novel uncertainty-based offline RL algorithm that penalizes the unstable Q-values in the regression target using uncertainty-aware weights, achieving the conservative Q-function without the need of estimating the behavior policy, suitable for RS with a large number of items. Experiments on two real-world datasets show the proposed method serves as a gain plug-in for existing RS models.
Chengqian Gao, Ke Xu 0002, Kuangqi Zhou, Lanqing Li, Xueqian Wang 0001, Bo Yuan 0008, Peilin Zhao
SIGIR3
2021 Understanding and Resolving Performance Degradation in Deep Graph Convolutional Networks
abstract
A Graph Convolutional Network (GCN) stacks several layers and in each layer performs a PROPagation operation~(PROP) and a TRANsformation operation~(TRAN) for learning node representations over graph-structured data. Though powerful, GCNs tend to suffer performance drop when the model gets deep. Previous works focus on PROPs to study and mitigate this issue, but the role of TRANs is barely investigated. In this work, we study performance degradation of GCNs by experimentally examining how stacking only TRANs or PROPs works. We find that TRANs contribute significantly, or even more than PROPs, to declining performance, and moreover that they tend to amplify node-wise feature variance in GCNs, causing variance inflammation that we identify as a key factor for causing performance drop. Motivated by such observations, we propose a variance-controlling technique termed Node Normalization (NodeNorm), which scales each node's features using its own standard deviation. Experimental results validate the effectiveness of NodeNorm on addressing performance degradation of GCNs. Specifically, it enables deep GCNs to outperform shallow ones in cases where deep models are needed, and to achieve comparable results with shallow ones on 6 benchmark datasets. NodeNorm is a generic plug-in and can well generalize to other GNN architectures. Code is publicly available at https://github.com/miafei/NodeNorm.
Kuangqi Zhou, Yanfei Dong, Wee Sun Lee, Bryan Hooi, Huan Xu 0001, Jiashi Feng
CIKM1
2021 Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph Drawing
abstract
The Laplacian representation recently gains increasing attention for reinforcement learning as it provides succinct and informative representation for states, by taking the eigenvectors of the Laplacian matrix of the state-transition graph as state embeddings. Such representation captures the geometry of the underlying state space and is beneficial to RL tasks such as option discovery and reward shaping. To approximate the Laplacian representation in large (or even continuous) state spaces, recent works propose to minimize a spectral graph drawing objective, which however has infinitely many global minimizers other than the eigenvectors. As a result, their learned Laplacian representation may differ from the ground truth. To solve this problem, we reformulate the graph drawing objective into a generalized form and derive a new learning objective, which is proved to have eigenvectors as its unique global minimizer. It enables learning high-quality Laplacian representations that faithfully approximate the ground truth. We validate this via comprehensive experiments on a set of gridworld and continuous control environments. Moreover, we show that our learned Laplacian representations lead to more exploratory options and better reward shaping.
Kuangqi Zhou, Qixin Zhang 0001, Jie Shao 0006, Bryan Hooi, Jiashi Feng
ICML2