VLDB 2026 Research / reviewers in the wild / expert
Shangzhe Li
dblp:216/8524
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Language Model Distillation: A Temporal Difference Imitation Learning PerspectiveabstractLarge language models have led to significant progress across many NLP tasks, although their massive sizes often incur substantial computational costs. Distillation has become a common practice to compress these large and highly capable models into smaller, more efficient ones. Many existing language model distillation methods can be viewed as behavior cloning from the perspective of imitation learning or inverse reinforcement learning. This viewpoint has inspired subsequent studies that leverage (inverse) reinforcement learning techniques, including variations of behavior cloning and temporal difference learning methods. Rather than proposing yet another specific temporal difference method, we introduce a general framework for temporal difference-based distillation by exploiting the distributional sparsity of the teacher model. Specifically, it is often observed that language models assign most probability mass to a small subset of tokens. Motivated by this observation, we design a temporal difference learning framework that operates on a reduced action space (a subset of vocabulary), and demonstrate how practical algorithms can be derived and the resulting performance improvements. Zishun Yu, Shangzhe Li |
AAAI | 2 |
| 2026 | Uncovering capabilities of hash function in graph classification
Yingke Liu, Shangzhe Li, Bowen Shi 0001, Junran Wu |
Pattern Recognit. | 2 |
| 2025 | Reward-free World Models for Online Imitation LearningabstractImitation learning (IL) enables agents to acquire skills directly from expert demonstrations, providing a compelling alternative to reinforcement learning. However, prior online IL approaches struggle with complex tasks characterized by high-dimensional inputs and complex dynamics. In this work, we propose a novel approach to online imitation learning that leverages reward-free world models. Our method learns environmental dynamics entirely in latent spaces without reconstruction, enabling efficient and accurate modeling. We adopt the inverse soft-Q learning objective, reformulating the optimization process in the Q-policy space to mitigate the instability associated with traditional optimization in the reward-policy space. By employing a learned latent dynamics model and planning for control, our approach consistently achieves stable, expert-level performance in tasks with high-dimensional observation or action spaces and intricate dynamics. We evaluate our method on a diverse set of benchmarks, including DMControl, MyoSuite, and ManiSkill2, demonstrating superior empirical performance compared to existing approaches. Shangzhe Li, Zhiao Huang, Hao Su 0001 |
ICML | 1 |
| 2025 | ChartNet: Reducing Subjectivity in Stock Prediction Through Unified Technical Chart RepresentationabstractABSTRACT Technical analysis, which includes technical indicators and charts derived from specific rules, has proven effective and widely used for stock movement prediction. However, technical chart evaluation is often limited by subjectivity, arising from sparse chart types and substantial information loss due to rigid rules. While pattern recognition algorithms have been developed to address this issue, they still rely on manual chart labelling and primarily focus on closing prices, leaving much of the chart's broader information untapped. To overcome these limitations, we propose a novel framework called ChartNet, designed to extract general information from technical charts and reduce subjectivity in chart analysis. ChartNet employs a unified representation for charts across financial series with varying simplification levels and leverages a chart triplet loss function for unsupervised training, eliminating the need for labelled data. Compared with several state‐of‐the‐art baselines, our framework has reached the best prediction accuracy on CSI‐300, SZ‐50 components and Dow Jones Index in 2022: 65.91%, 63.70% and 64.96% respectively. In backtesting using actual stock data, our framework achieves the highest average return of 1.12 and 1.15. Furthermore, we highlight the interpretability of ChartNet through two case studies, some important charts and failure cases, illustrating its capability to uncover meaningful insights from charts. This research contributes to advancing the objective evaluation of technical charts and promoting a more comprehensive understanding of chart‐based stock prediction performance. Shangzhe Li, Yingke Liu, Fanglei Cheng, Junran Wu, Ke Xu 0001 |
Expert Syst. J. Knowl. Eng. | 1 |
| 2025 | Molecular graph contrastive learning with line graph
Xueyuan Chen, Shangzhe Li, Ruomei Liu, Bowen Shi 0001, Junran Wu, Ke Xu 0001 |
Pattern Recognit. | 2 |
| 2025 | SuperMPFL: A Supermask-Based Mechanism for Personalized Federated LearningabstractPersonalized federated learning (PFL) is a specialized application of the federated learning paradigm designed to support personalized use cases. Unlike traditional federated learning, which aims to train a high-quality global model, the goal of PFL is to tailor a model that best fits each individual user. Most existing PFL approaches adopt training architectures similar to those used in traditional federated learning, relying on global or partial model sharing during training. While this helps improve model personalization across clients, it also introduces a range of challenges, including risks of data leakage and increased communication overhead. To address these challenges, we propose a novel personalized federated learning (PFL) framework called SuperMPFL, which leverages supermasks to effectively tackle issues related to accuracy, privacy, and efficiency. In particular, the SuperMPFL technique utilizes masking and ranking strategies to obscure the true gradient information. By converting gradients into ranked numerical representations, this approach enhances privacy protection during the training process. Furthermore, this approach reduces communication overhead by transmitting significantly less information compared to conventional methods. In SuperMPFL, each client receives the global model and then emphasizes its personalized parameters, particularly at the model’s edges. This design not only improves accuracy but also strengthens robustness against privacy attacks. Evaluations on standard federated learning benchmarks demonstrate the superiority of our approach, which outperforms state-of-the-art methods in terms of accuracy, privacy, and efficiency. Zhe Sun 0005, Shangzhe Li, Lihua Yin, Yahong Chen, Aohai Zhang, Meifan Zhang, Yuanyuan He 0002 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2024 | IPM: Information Lossless Pre-training Strategy for Molecular Property PredictionabstractGiven the pivotal role of molecular property prediction in drug development and material science, graph self-supervised learning has been implemented in molecular representation learning to compensate for the shortage of labeled molecules. However, current proposed methods often focus on designing data augmentation schemes and leveraging domain knowledge to improve performance, which inevitably leads to molecular semantics loss and limited generalization capability. To the end, we propose IPM, an Information lossless Pretraining strategy for Molecular property prediction that leverages the information of both the original graph and line graph of molecules. Specifically, by contrasting the given graph with the corresponding line graph, the graph encoder can fully learn the generic molecular semantic representation without profound domain knowledge. We also design a new message-passing scheme that retains information consistency during message passing between two kinds of graphs. Additionally, we present two graph contrastive losses for performance fixing and over-smoothing prevention during the learning process. Experimental results on multiple regression tasks for molecular property prediction demonstrate the effectiveness of IPM against state-of-the-art (SOTA) methods. Ruomei Liu, Shangzhe Li, Xingyu Peng, Haitao Yuan 0002, Junran Wu, Ke Xu 0001 |
BIBM | 2 |
| 2024 | Uncovering Capabilities of Model Pruning in Graph Contrastive LearningabstractGraph contrastive learning has achieved great success in pre-training graph neural networks without ground-truth labels. Leading graph contrastive learning follows the classical scheme of contrastive learning, forcing model to identify the essential information from augmented views. However, general augmented views are produced via random corruption or learning, which inevitably leads to semantics alteration. Although domain knowledge guided augmentations alleviate this issue, the generated views are domain specific and undermine the generalization. In this work, motivated by the firm representation ability of sparse model from pruning, we reformulate the problem of graph contrastive learning via contrasting different model versions rather than augmented views. We first theoretically reveal the superiority of model pruning in contrast to data augmentations. In practice, we take original graph as input and dynamically generate a perturbed graph encoder to contrast with the original encoder by pruning its transformation weights. Furthermore, considering the integrity of node embedding in our method, we are capable of developing a local contrastive loss to tackle the hard negative samples that disturb the model training. We extensively validate our method on various benchmarks regarding graph classification via unsupervised and transfer learning. Compared to the state-of-the-art (SOTA) works, better performance can always be obtained by the proposed method. Junran Wu, Xueyuan Chen, Shangzhe Li |
ACM Multimedia | 3 |
| 2024 | HILL: Hierarchy-aware Information Lossless Contrastive Learning for Hierarchical Text ClassificationabstractHe Zhu, Junran Wu, Ruomei Liu, Yue Hou, Ze Yuan, Shangzhe Li, Yicheng Pan, Ke Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Junran Wu, Ruomei Liu, Ze Yuan, Shangzhe Li, Yicheng Pan 0001, Ke Xu 0001 |
NAACL-HLT | 6 |
| 2024 | Forecasting Turning Points in Stock Price by Integrating Chart Similarity and MultipersistenceabstractForecasting financial data plays a crucial role in financial market. Relying solely on prices or price trends as prediction targets often leads to a vast of invalid transactions. As a result, researchers have increasingly turned their attention to turning points as the prediction target. Surprisingly, existing methods have largely overlooked the role of technical charts, despite turning points being closely related to the technical charts. Recently, several researchers have attempted to utilize chart information via converting price sequences into images for turning point forecasting, but robustness and convergence problems arise. To address these challenges and enhance the turning point predictions, this article introduces a new method known as MPCNet. Specifically, we first transform the price series into a graph structure using chart similarity to robustly extract valuable information from technical charts. Additionally, we introduce the multipersistence topology tool to accurately predict stock turning points and provide convergence guarantee. Experimental results demonstrate the significant superiority of our proposed model over existing methods. Furthermore, based on additional performance evaluations using real stock data, MPCNet consistently achieves the highest average return during the transaction backtesting period. Meanwhile, we provide empirical validation of robustness and theoretical analysis to confirm its convergence, establishing it as a superior tool for financial forecasting. Shangzhe Li, Yingke Liu, Xueyuan Chen, Junran Wu, Ke Xu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | SEGA: Structural Entropy Guided Anchor View for Graph Contrastive LearningabstractIn contrastive learning, the choice of "view" controls the information that the representation captures and influences the performance of the model. However, leading graph contrastive learning methods generally produce views via random corruption or learning, which could lead to the loss of essential information and alteration of semantic information. An anchor view that maintains the essential information of input graphs for contrastive learning has been hardly investigated. In this paper, based on the theory of graph information bottleneck, we deduce the definition of this anchor view; put differently, the anchor view with essential information of input graph is supposed to have the minimal structural uncertainty. Furthermore, guided by structural entropy, we implement the anchor view, termed SEGA, for graph contrastive learning. We extensively validate the proposed anchor view on various benchmarks regarding graph classification under unsupervised, semi-supervised, and transfer learning and achieve significant performance boosts compared to the state-of-the-art methods. Junran Wu, Xueyuan Chen, Bowen Shi 0001, Shangzhe Li, Ke Xu 0001 |
ICML | 4 |
| 2022 | Structural Entropy Guided Graph Hierarchical PoolingabstractFollowing the success of convolution on non-Euclidean space, the corresponding pooling approaches have also been validated on various tasks regarding graphs. However, because of the fixed compression ratio and stepwise pooling design, these hierarchical pooling methods still suffer from local structure damage and suboptimal problem. In this work, inspired by structural entropy, we propose a hierarchical pooling approach, SEP, to tackle the two issues. Specifically, without assigning the layer-specific compression ratio, a global optimization algorithm is designed to generate the cluster assignment matrices for pooling at once. Then, we present an illustration of the local structure damage from previous methods in reconstruction of ring and grid synthetic graphs. In addition to SEP, we further design two classification models, SEP-G and SEP-N for graph classification and node classification, respectively. The results show that SEP outperforms state-of-the-art graph pooling methods on graph classification benchmarks and obtains superior performance on node classifications. Junran Wu, Xueyuan Chen, Ke Xu 0001, Shangzhe Li |
ICML | 4 |
| 2022 | A Simple yet Effective Method for Graph ClassificationabstractIn deep neural networks, better results can often be obtained by increasing the complexity of previously developed basic models. However, it is unclear whether there is a way to boost performance by decreasing the complexity of such models. Intuitively, given a problem, a simpler data structure comes with a simpler algorithm. Here, we investigate the feasibility of improving graph classification performance while simplifying the learning process. Inspired by structural entropy on graphs, we transform the data sample from graphs to coding trees, which is a simpler but essential structure for graph data. Furthermore, we propose a novel message passing scheme, termed hierarchical reporting, in which features are transferred from leaf nodes to root nodes by following the hierarchical structure of coding trees. We then present a tree kernel and a convolutional network to implement our scheme for graph classification. With the designed message passing scheme, the tree kernel and convolutional network have a lower runtime complexity of O(n) than Weisfeiler-Lehman subtree kernel and other graph neural networks of at least O(hm). We empirically validate our methods with several graph classification benchmarks and demonstrate that they achieve better performance and lower computational consumption than competing approaches. Junran Wu, Shangzhe Li, Yicheng Pan 0001, Ke Xu 0001 |
IJCAI | 2 |
| 2022 | Price graphs: Utilizing the structural information of financial time series for stock prediction
Junran Wu, Ke Xu 0001, Xueyuan Chen, Shangzhe Li, Jichang Zhao |
Inf. Sci. | 4 |
| 2022 | Chart GCN: Learning chart information with a graph convolutional network for stock movement prediction
Shangzhe Li, Junran Wu, Xin Jiang 0008, Ke Xu 0001 |
Knowl. Based Syst. | 1 |