VLDB 2026 Research / reviewers in the wild / expert
Kexin Zhang 0007
dblp:119/0668-7
· DBLP profile ↗
16ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0003-2678-8556ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Survey of Large Language Models for Text-Guided Molecular Discovery: From Molecule Generation to OptimizationabstractZiqing Wang, Kexin Zhang, Zihan Zhao, Yibo Wen, Abhishek Pandey, Han Liu, Kaize Ding. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kexin Zhang 0007, Yibo Wen, Kaize Ding |
ACL (1) | 2 |
| 2026 | A Survey of Deep Graph Learning under Distribution Shifts: From Graph Out-of-Distribution Generalization to AdaptationabstractDistribution shifts on graphs—the discrepancies in data distribution between training and employing a graph machine learning model—are ubiquitous and often unavoidable in real-world applications. These shifts may severely deteriorate model performance, posing significant challenges for reliable graph machine learning. In recent years, there has been a surge in research on graph machine learning specifically designed to tackle such distribution shifts, aiming to train models to achieve satisfactory performance on Out-of-Distribution (OOD) test data. This survey provides an up-to-date and forward-looking review of deep graph learning under distribution shifts. We categorize the field into three primary scenarios: graph OOD generalization, training-time graph OOD adaptation, and test-time graph OOD adaptation. We begin by formally formulating the problems and discussing various types of distribution shifts that can affect graph learning, such as covariate shifts and concept shifts. To provide a structured understanding of the literature, we introduce a systematic taxonomy that classifies existing methods into model-centric and data-centric approaches, investigating the techniques used in each category. We also summarize commonly used datasets in this research area to facilitate further investigation. Finally, we point out promising research directions and the corresponding challenges to encourage further study in this vital domain. Additionally, we provide a continuously updated reading list at https://github.com/kaize0409/Awesome-Graph-OOD . Kexin Zhang 0007, Song Wang 0013, Weili Shi, Chen Chen 0022, Pan Li 0005, Sheng Li 0001, Jundong Li, Kaize Ding |
ACM Trans. Knowl. Discov. Data | 1 |
| 2026 | Source-Free Time-Series Domain Adaptation With Prior Evaluation of Model SalienceabstractSource-free domain adaptation (SFDA) is a challenging, yet valuable task within unsupervised domain adaptation (UDA), which adapts pretrained models to diverse unlabeled target domains while safeguarding the data security of the source domain. However, existing SFDA methods primarily focus on computer vision applications, often overlooking the unique characteristics of time series, such as temporal dependencies and sequential nature. Moreover, the fine-tuning paradigm of current SFDA methods is typically limited to posterior adaptation, focusing solely on constraining the statistical properties of model outputs. We argue that this black-box paradigm lacks semantic interpretability and risks aligning with spurious contextual noise, leading to negative transfer. This necessitates a paradigm evolution from blind statistical adaptation to interpretable adaptation. To this end, we introduce model salience as a quantifiable proxy of semantic interpretability, representing the importance weights a trained model assigns to specific temporal fragments. Accordingly, we propose a novel fine-tuning paradigm for time-series SFDA, termed PrEPoA, which integrates Prior Evaluation of model salience with Posterior Adaptation. In the prior evaluation stage, a key pattern reconstruction (KPR) module based on a sensitive masking mechanism is designed to quantify the model salience, while a novel interpattern triplet loss is introduced to calibrate it. In the posterior adaptation stage, robust prototype clustering (RPC) generates trustworthy reference labels as pseudo-ground truth for adaptation. Comprehensive experiments on the wireless sensor data mining (WISDM), human activity recognition (HAR), heterogeneity HAR (HHAR), machine fault diagnosis (MFD), and sleep stage classification (SSC) datasets demonstrate the superiority of our PrEPoA framework compared to nine UDA and seven SFDA methods. Furthermore, we experimentally validate that PrEPoA serves as a plug-and-play module that effectively incorporated into other SFDA methods. Rongyao Cai, Ming Jin 0005, Qingsong Wen, Kexin Zhang 0007, Yong Liu 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Combinatorial Optimization Perspective based Framework for Multi-behavior RecommendationabstractIn real-world recommendation scenarios, users engage with items through various types of behaviors. Leveraging diversified user behavior information for learning can enhance the recommendation of target behaviors (e.g., buy), as demonstrated by recent multi-behavior methods. The mainstream multi-behavior recommendation framework consists of two steps: fusion and prediction. Recent approaches utilize graph neural networks for multi-behavior fusion and employ multi-task learning paradigms for joint optimization in the prediction step, achieving significant success. However, these methods have limited perspectives on multi-behavior fusion, which leads to inaccurate capture of user behavior patterns in the fusion step. Moreover, when using multi-task learning for prediction, the relationship between the target task and auxiliary tasks is not sufficiently coordinated, resulting in negative information transfer. To address these problems, we propose a novel multi-behavior recommendation framework based on the combinatorial optimization perspective, named COPF. Specifically, we treat multi-behavior fusion as a combinatorial optimization problem, imposing different constraints at various stages of each behavior to restrict the solution space, thus significantly enhancing fusion efficiency (COGCN). In the prediction step, we improve both forward and backward propagation during the generation and aggregation of multiple experts to mitigate negative transfer caused by differences in both feature and label distributions (DFME). Comprehensive experiments on three real-world datasets indicate the superiority of COPF. Further analyses also validate the effectiveness of the COGCN and DFME modules. Our code is available at https://github.com/1918190/COPF. Chenhao Zhai, Chang Meng, Yu Yang 0015, Kexin Zhang 0007, Xuhao Zhao 0001, Xiu Li 0001 |
KDD (1) | 4 |
| 2025 | Glocal Information Bottleneck for Time Series ImputationabstractTime Series Imputation (TSI), which aims to recover missing values in temporal data, remains a fundamental challenge due to the complex and often high-rate missingness in real-world scenarios. Existing models typically optimize the point-wise reconstruction loss, focusing on recovering numerical values (local information). However, we observe that under high missing rates, these models still perform well in the training phase yet produce poor imputations and distorted latent representation distributions (global information) in the inference phase. This reveals a critical optimization dilemma: current objectives lack global guidance, leading models to overfit local noise and fail to capture global information of the data. To address this issue, we propose a new training paradigm, **Glocal** **I**nformation **B**ottleneck (**Glocal-IB**). Glocal-IB is model-agnostic and extends the standard IB framework by introducing a Global Alignment loss, derived from a tractable mutual information approximation. This loss aligns the latent representations of masked inputs with those of their originally observed counterparts. It helps the model retain global structure and local details while suppressing noise caused by missing values, giving rise to better generalization under high missingness. Extensive experiments on nine datasets confirm that Glocal-IB leads to consistently improved performance and aligned latent representations under missingness. Our code implementation is available in [https://github.com/Muyiiiii/NeurIPS-25-Glocal-IB](https://github.com/Muyiiiii/NeurIPS-25-Glocal-IB). Kexin Zhang 0007, Guibin Zhang, Philip S. Yu, Kaize Ding |
NeurIPS | 2 |
| 2025 | Cross-Domain Conditional Diffusion Models for Time Series Imputation
Kexin Zhang 0007, Baoyu Jing, K. Selçuk Candan, Dawei Zhou 0003, Qingsong Wen, Kaize Ding |
ECML/PKDD (8) | 1 |
| 2025 | Fusion Matters: Learning Fusion in Deep Click-through Rate Prediction ModelsabstractThe evolution of previous Click-Through Rate (CTR) models has mainly been driven by proposing complex components, whether shallow or deep, that are adept at modeling feature interactions. However, there has been less focus on improving fusion design. Instead, two naive solutions, stacked and parallel fusion, are commonly used. Both solutions rely on pre-determined fusion connections and fixed fusion operations. It has been repetitively observed that changes in fusion design may result in different performances, highlighting the critical role that fusion plays in CTR models. While there have been attempts to refine these basic fusion strategies, these efforts have often been constrained to specific settings or dependent on specific components. Neural architecture search has also been introduced to partially deal with fusion design, but it comes with limitations. The complexity of the search space can lead to inefficient and ineffective results. To bridge this gap, we introduce OptFusion, a method that automates the learning of fusion, encompassing both the connection learning and the operation selection. We have proposed a one-shot learning algorithm tackling these tasks concurrently. Our experiments are conducted over three large-scale datasets. Extensive experiments prove both the effectiveness and efficiency of OptFusion in improving CTR model performance. Our code implementation is available here https://github.com/kexin-kxzhang/OptFusion. Kexin Zhang 0007, Fuyuan Lyu, Xing Tang 0007, Dugang Liu, Chen Ma 0001, Kaize Ding, Xiuqiang He 0001, Xue (Steve) Liu |
WSDM | 1 |
| 2024 | DCS: Debiased Contrastive Learning with Weak Supervision for Time Series ClassificationabstractSelf-supervised contrastive learning (SSCL) has performed excellently on time series classification tasks. Most SSCL- based classification algorithms generate positive and negative samples in the time or frequency domains, focusing on mining similarities between them. However, two issues are not well addressed in the SSCL framework: the sampling bias and the task-agnostic representation problems. Sampling bias indicates fake negative sample selection in SSCL, and task- agnostic representation results in the unknown correlation between the extracted feature and downstream tasks. To address the issues, we propose Debiased Contrastive learning with weak Supervision framework, abbreviated as DCS. It employs the clustering operation to remove fake negative samples and introduces weak supervisory signals into the SSCL framework to guide feature extraction. Additionally, we propose a channel augmentation method that allows the DCS to extract features from local and global perspectives simultaneously. The comprehensive experiments on the widely used datasets show that DCS achieves performance superior to state-of-the-art methods on the widely used popular benchmark datasets. Rongyao Cai, Linpeng Peng, Zhengming Lu, Kexin Zhang 0007, Yong Liu 0007 |
ICASSP | 4 |
| 2024 | Skip-Step Contrastive Predictive Coding for Time Series Anomaly DetectionabstractSelf-supervised learning (SSL) shows impressive performance in many tasks lacking sufficient labels. In this paper, we study SSL in time series anomaly detection (TSAD) by incorporating the characteristics of time series data. Specifically, we build an anomaly detection algorithm consisting of global pattern learning and local association learning. The global pattern learning module builds encoder and decoder to reconstruct the raw time series data to detect global anomalies. To complement the limitation of the global pattern learning that ignores local associations between anomaly points and their adjacent windows, we design a local association learning module, which leverages contrastive predictive coding (CPC) to transform the identification of anomaly points into positive pairs identification. Motivated by the observation that adjusting the distance between the history window and the time point to be detected directly impacts the detection performance in the CPC framework, we further propose a skip-step CPC scheme in the local association learning module which adjusts the distance for better construction of the positive pairs and detection results. The experimental results show that the proposed algorithm achieves superior performance on SMD and PSM datasets in comparison with 12 state-of-the-art algorithms. Kexin Zhang 0007, Qingsong Wen, Chaoli Zhang 0001, Liang Sun 0001, Yong Liu 0007 |
ICASSP | 1 |
| 2024 | Position: What Can Large Language Models Tell Us about Time Series AnalysisabstractTime series analysis is essential for comprehending the complexities inherent in various real-world systems and applications. Although large language models (LLMs) have recently made significant strides, the development of artificial general intelligence (AGI) equipped with time series analysis capabilities remains in its nascent phase. Most existing time series models heavily rely on domain knowledge and extensive model tuning, predominantly focusing on prediction tasks. In this paper, we argue that current LLMs have the potential to revolutionize time series analysis, thereby promoting efficient decision-making and advancing towards a more universal form of time series analytical intelligence. Such advancement could unlock a wide range of possibilities, including time series modality switching and question answering. We encourage researchers and practitioners to recognize the potential of LLMs in advancing time series analysis and emphasize the need for trust in these related efforts. Furthermore, we detail the seamless integration of time series analysis with existing LLM technologies and outline promising avenues for future research. Ming Jin 0005, Yifan Zhang 0004, Wei Chen 0070, Kexin Zhang 0007, Yuxuan Liang 0002, Bin Yang 0002, Jindong Wang 0001, Shirui Pan, Qingsong Wen |
ICML | 4 |
| 2024 | IncMSR: An Incremental Learning Approach for Multi-Scenario RecommendationabstractFor better performance and less resource consumption, multi-scenario recommendation (MSR) is proposed to train a unified model to serve all scenarios by leveraging data from multiple scenarios. Current works in MSR focus on designing effective networks for better information transfer among different scenarios. However, they omit two important issues when applying MSR models in industrial situations. The first is the efficiency problem brought by mixed data, which delays the update of models and further leads to performance degradation. The second is that MSR models are insensitive to the changes of distribution over time, resulting in suboptimal effectiveness in the incoming data. In this paper, we propose an incremental learning approach for MSR (IncMSR), which can not only improve the training efficiency but also perceive changes in distribution over time. Specifically, we first quantify the pair-wise distance between representations from scenario, time and time-scenario dimensions respectively. Then, we decompose the MSR model into scenario-shared and scenario-specific parts and apply fine-grained constraints on the distances quantified with respect to the two different parts. Finally, all constraints are fused in an elegant way using a metric learning framework as a supplementary penalty term to the original MSR loss function. Offline experiments on two real-world datasets are conducted to demonstrate the superiority and compatibility of our proposed approach. Kexin Zhang 0007, Yichao Wang 0002, Xiu Li 0001, Ruiming Tang, Rui Zhang 0003 |
WSDM | 1 |
| 2024 | Self-Supervised Learning for Time Series Analysis: Taxonomy, Progress, and ProspectsabstractSelf-supervised learning (SSL) has recently achieved impressive performance on various time series tasks. The most prominent advantage of SSL is that it reduces the dependence on labeled data. Based on the pre-training and fine-tuning strategy, even a small amount of labeled data can achieve high performance. Compared with many published self-supervised surveys on computer vision and natural language processing, a comprehensive survey for time series SSL is still missing. To fill this gap, we review current state-of-the-art SSL methods for time series data in this article. To this end, we first comprehensively review existing surveys related to SSL and time series, and then provide a new taxonomy of existing time series SSL methods by summarizing them from three perspectives: generative-based, contrastive-based, and adversarial-based. These methods are further divided into ten subcategories with detailed reviews and discussions about their key intuitions, main frameworks, advantages and disadvantages. To facilitate the experiments and validation of time series SSL methods, we also summarize datasets commonly used in time series forecasting, classification, anomaly detection, and clustering tasks. Finally, we present the future directions of SSL for time series analysis. Kexin Zhang 0007, Qingsong Wen, Chaoli Zhang 0001, Rongyao Cai, Ming Jin 0005, Yong Liu 0007, James Y. Zhang, Yuxuan Liang 0002, Guansong Pang, Dongjin Song, Shirui Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Debiased Contrastive Learning With Supervision Guidance for Industrial Fault DetectionabstractThe time series self-supervised contrastive learning framework has succeeded significantly in industrial fault detection scenarios. It typically consists of pretraining on abundant unlabeled data and fine-tuning on limited annotated data. However, the two-phase framework faces three challenges: Sampling bias, task-agnostic representation issue, and angular-centricity issue. These challenges hinder further development in industrial applications. This article introduces a debiased contrastive learning with supervision guidance (DCLSG) framework and applies it to industrial fault detection tasks. First, DCLSG employs channel augmentation to integrate temporal and frequency domain information. Pseudolabels based on momentum clustering operation are assigned to extracted representations, thereby mitigating the sampling bias raised by the selection of positive pairs. Second, the generated supervisory signal guides the pretraining phase, tackling the task-agnostic representation issue. Third, the angular-centricity issue is addressed using the proposed Gaussian distance metric measuring the radial distribution of representations. The experiments conducted on three industrial datasets (ISDB, CWRU, and practical datasets) validate the superior performance of the DCLSG compared to other fault detection methods. Rongyao Cai, Wang Gao 0001, Linpeng Peng, Zhengming Lu, Kexin Zhang 0007, Yong Liu 0007 |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Debiased Contrastive Learning for Time-Series Representation Learning and Fault DetectionabstractBuilding reliable fault detection systems through deep neural networks is an appealing topic in industrial scenarios. In these contexts, the representations extracted by neural networks on available labeled time-series data can reflect system states. However, this endeavor remains challenging due to the necessity of labeled data. Self-supervised contrastive learning (SSCL) is one of the effective approaches to deal with this challenge, but existing SSCL-based models suffer from sampling bias and representation bias problems. This article introduces a debiased contrastive learning framework for time-series data and applies it to industrial fault detection tasks. This framework first develops the multigranularity augmented view generation method to generate augmented views at different granularities. It then introduces the momentum clustering contrastive learning strategy and the expert knowledge guidance mechanism to mitigate sampling bias and representation bias, respectively. Finally, the experiments on a public bearing fault detection dataset and a widely used valve stiction detection dataset show the effectiveness of the proposed feature learning framework. Kexin Zhang 0007, Rongyao Cai, Chunlin Zhou, Yong Liu 0007 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | HCL4QC: Incorporating Hierarchical Category Structures Into Contrastive Learning for E-commerce Query ClassificationabstractQuery classification plays a crucial role in e-commerce, where the goal is to assign user queries to appropriate categories within a hierarchical product category taxonomy. However, existing methods rely on a limited number of words from the category description and often neglect the hierarchical structure of the category tree, resulting in suboptimal category representations. To overcome these limitations, we propose a novel approach named hierarchical contrastive learning framework for query classification (HCL4QC), which leverages the hierarchical category tree structure to improve the performance of query classification. Specifically, HCL4QC is designed as a plugin module that consists of two innovative losses, namely local hierarchical contrastive loss (LHCL) and global hierarchical contrastive loss (GHCL). LHCL adjusts representations of categories according to their positional relationship in the hierarchical tree, while GHCL ensures the semantic consistency between the parent category and its child categories. Our proposed method can be adapted to any query classification tasks that involve a hierarchical category structure. We conduct experiments on two real-world datasets to demonstrate the superiority of our hierarchical contrastive learning. The results demonstrate significant improvements in the query classification task, particularly for long-tail categories with sparse supervised information. Lvxing Zhu, Kexin Zhang 0007, Hao Chen 0122, Chao Wei 0010, Weiru Zhang, Haihong Tang, Xiu Li 0001 |
CIKM | 2 |
| 2023 | Expert demonstrations guide reward decomposition for multi-agent cooperation
Shanqi Liu, Yudi Ruan, Kexin Zhang 0007, Yong Liu 0007 |
Neural Comput. Appl. | 5 |