VLDB 2026 Research / reviewers in the wild / expert
Xi Chen 0073
dblp:16/3283-73
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0009-6180-4524ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenDis: Generative-Discriminative Dual-View Co-Training for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) aims to identify both known and novel categories from partially labeled data, reflecting more realistic open-world learning scenarios.However, most existing methods rely solely on one-hot discriminative supervision, leading to overfitting on seen classes and poor generalization to unseen ones.Recent advances introduce large language models (LLMs) to incorporate external semantics, yet they often suffer from semantic-label misalignment and weak semantic integration during training.We propose GenDis, a Generative-Discriminative Dual-View Co-Training framework that unifies discriminative classification and semantic label generation within an LLM.Discriminative pseudo-labels guide the formation of a separable generative latent space, enabling semantically meaningful supervision for novel classes.To ensure consistency between the two views, we employ Canonical Correlation Analysis (CCA)-based alignment and a curriculumguided, dispersion-aware pseudo-labeling strategy for iterative refinement.Extensive experiments on five GCD benchmarks demonstrate that GenDis substantially outperforms prior methods, validating the effectiveness of dual-view co-training with semantically enriched supervision.Code is available at https: //github.com/cx9941/GenDis. Xi Chen 0073, Chuan Qin 0002, Shasha Hu, Chao Wang 0086, Hengshu Zhu, Hui Xiong 0001 |
ACL (1) | 1 |
| 2026 | TLSA: LLM-Guided Text-Label Space Alignment with Contrastive Learning for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) aims to classify data from partially labeled datasets by jointly recognizing known categories and discovering novel ones.Despite recent advances, existing methods still suffer from weak text-label alignment, inconsistent objectives across known and novel categories, and poor discrimination of semantically similar clusters.To mitigate these issues, we propose TLSA, a unified framework that enforces contrastive alignment between text and label representations within a shared semantic space.Specifically, we first design a label-semantic aware dual-encoder equipped with a symmetric contrastive objective to achieve text-label alignment.Then, we leverage LLM-based label induction to generate explicit and semantically meaningful names for previously unseen categories, followed by a graph-based refinement strategy that disambiguates semantically overlapping clusters through forced renaming.Finally, a confidence-aware sampling strategy ensures balanced learning across both easy and hard instances.Extensive experiments on four benchmark datasets show that TLSA consistently outperforms state-of-theart GCD methods.The code is available at https://github.com/Wenxi-Xu/TLSA. Wenxi Xu, Chuan Qin 0002, Xi Chen 0073, Chuyu Fang, Yuanchun Zhou, Hengshu Zhu |
ACL (1) | 3 |
| 2025 | Enhancing Dual-Target Cross-Domain Recommendation via Similar User BridgingabstractDual-target cross-domain recommendation aims to mitigate data sparsity and enables mutual enhancement via bidirectional knowledge transfer. Most existing methods rely on overlapping users to build cross-domain connections. However, in many real-world scenarios, overlapping data is extremely limited-or even entirely absent-significantly diminishing the effectiveness of these methods. To address this challenge, we propose SUBCDR, a novel framework that leverages large language models (LLMs) to bridge similar users across domains, thereby enhancing dual-target cross-domain recommendation. Specifically, we introduce a Multi-Interests-Aware Prompt Learning mechanism that enables LLMs to generate comprehensive user profiles, disentangling domain-invariant interest points while capturing fine-grained preferences. Then, we construct intra-domain bipartite graphs from user-item interactions and an inter-domain heterogeneous graph that links similar users across domains. Subsequently, to facilitate effective knowledge transfer, we employ Graph Convolutional Networks (GCNs) for intra-domain relationship modeling and design an Inter-domain Hierarchical Attention Network (InterHAN) to facilitate inter-domain knowledge transfer through similar users, learning both shared and specific user representations. Extensive experiments on seven public datasets demonstrate that SUBCDR outperforms state-of-the-art cross-domain recommendation algorithms and single-domain recommendation methods. Our code is publicly available at https://github.com/97z/SUBCDR.git. Xi Chen 0073, Chuyu Fang, Jianji Wang 0001, Chuan Qin 0002, Fuzhen Zhuang |
CIKM | 2 |
| 2025 | SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language ModelsabstractIn recent years, the rapid advancement of Artificial Intelligence (AI) technologies, particularly Large Language Models (LLMs), has revolutionized the paradigm of scientific discovery, establishing AI-for-Science (AI4Science) as a dynamic and evolving field. However, there is still a lack of an effective framework for the overall assessment of AI4Science, particularly from a holistic perspective on data quality and model capability. Therefore, in this study, we propose SciHorizon, a comprehensive assessment framework designed to benchmark the readiness of AI4Science from both scientific data and LLM perspectives. First, we introduce a generalizable framework for assessing AI-ready scientific data, encompassing four key dimensions-Quality, FAIRness, Explainability, and Compliance-which are subdivided into 15 sub-dimensions. Drawing on data resource papers published between 2018 and 2023 in peer-reviewed journals, we present recommendation lists of AI-ready datasets for Earth, Life, and Materials Sciences, making a novel and original contribution to the field. Concurrently, to assess the capabilities of LLMs across multiple scientific disciplines, we establish 16 assessment dimensions based on five core indicators-Knowledge, Understanding, Reasoning, Multimodality, and Values-spanning Mathematics, Physics, Chemistry, Life Sciences, and Earth and Space Sciences. Using the developed benchmark datasets, we have conducted a comprehensive evaluation of over 50 representative open-source and closed-source LLMs. All the results are publicly available and can be accessed online at www.scihorizon.cn/en. Chuan Qin 0002, Pengmin Wu, Xi Chen 0073, Yihang Cheng 0001, Meng Xiao 0001, Xiangchao Dong, Qingqing Long, Boya Pan, Han Wu 0002, Chengzan Li, Yuanchun Zhou, Hui Xiong 0001, Hengshu Zhu |
KDD (2) | 5 |
| 2025 | From Missteps to Mastery: Enhancing Low-Resource Dense Retrieval through Adaptive Query GenerationabstractDocument retrieval, designed to recall query-relevant documents from expansive collections, is essential for information-seeking tasks, such as web search and open-domain question-answering. Advances in representation learning and pretrained language models (PLMs) have driven a paradigm shift from traditional sparse retrieval methods to more effective dense retrieval approaches, forging enhanced semantic connections between queries and documents and establishing new performance benchmarks. However, reliance on extensive annotated document-query pairs limits their competitiveness in low-resource scenarios. Recent research efforts employing the few-shot capabilities of large language models (LLMs) and prompt engineering for synthetic data generation have emerged as a promising solution. Nonetheless, these approaches are hindered by the generation of lower-quality data within the conventional dense retrieval training process. To this end, in this paper, we introduce iGFT, a framework aimed at enhancing low-resource dense retrieval by integrating a three-phase process --- Generation, Filtering, and Tuning --- coupled with an iterative optimization strategy. Specifically, we first employ supervised fine-tuning on limited ground truth data, enabling an LLM to function as the generator capable of producing potential queries from given documents. Subsequently, we present a multi-stage filtering module to minimize noise in the generated data while retaining samples poised to significantly improve the dense retrieval model's performance in the follow-up fine-tuning process. Furthermore, we design a novel iterative optimization strategy that dynamically optimizes the query generator for producing more informative queries, thereby enhancing the efficacy of the entire framework. Finally, extensive experiments conducted on a series of publicly available retrieval benchmark datasets have demonstrated the effectiveness of the proposed iGFT. Zhenyu Tong, Chuan Qin 0002, Chuyu Fang, Kaichun Yao, Xi Chen 0073, Jingshuai Zhang, Chen Zhu 0003, Hengshu Zhu |
KDD (1) | 5 |
| 2025 | A Comprehensive Survey of Artificial Intelligence Techniques for Talent AnalyticsabstractIn today’s competitive and fast-evolving business environment, it is critical for organizations to rethink how to make talent-related decisions in a quantitative manner. Indeed, the recent development of big data and artificial intelligence (AI) techniques has revolutionized human resource management (HRM). The availability of large-scale talent and management-related data provides unparalleled opportunities for business leaders to comprehend organizational behaviors and gain tangible knowledge from a data science perspective, which, in turn, delivers intelligence for real-time decision-making and effective talent management for their organizations. In the last decade, talent analytics has emerged as a promising field in applied data science for HRM, garnering significant attention from AI communities and inspiring numerous research efforts. To this end, we present an up-to-date and comprehensive survey on AI technologies used for talent analytics in the field of HRM. Specifically, we first provide the background knowledge of talent analytics and categorize various pertinent data. Subsequently, we offer a comprehensive taxonomy of relevant research efforts, categorized based on three distinct application-driven scenarios at different levels: talent management, organization management, and labor market analysis. In conclusion, we summarize the open challenges and potential prospects for future research directions in the domain of AI-driven talent analytics. Chuan Qin 0002, Le Zhang 0010, Yihang Cheng 0001, Rui Zha, Dazhong Shen, Qi Zhang 0053, Xi Chen 0073, Ying Sun 0006, Chen Zhu 0003, Hengshu Zhu, Hui Xiong 0001 |
Proc. IEEE | 7 |
| 2024 | Towards Efficient Resume Understanding: A Multi-Granularity Multi-Modal Pre-Training ApproachabstractIn the contemporary era of widespread online recruitment, resume understanding has been widely acknowledged as a fundamental and crucial task, which aims to extract structured information from resume documents automatically. Compared to the traditional rule-based approaches, the utilization of recently proposed pre-trained document understanding models can greatly enhance the effectiveness of resume understanding. The present approaches have, however, disregarded the hierarchical relations within the structured information presented in resumes, and have difficulty parsing resumes in an efficient manner. To this end, in this paper, we propose a novel model, namely ERU, to achieve efficient resume understanding. Specifically, we first introduce a layout-aware multi-modal fusion transformer for encoding the segments in the resume with integrated textual, visual, and layout information. Then, we design three self-supervised tasks to pre-train this module via a large number of unlabeled resumes. Next, we fine-tune the model with a multi-granularity sequence labeling task to extract structured information from resumes. Finally, extensive experiments on a real-world dataset clearly demonstrate the effectiveness of ERU. Feihu Jiang, Chuan Qin 0002, Jingshuai Zhang, Kaichun Yao, Xi Chen 0073, Dazhong Shen, Chen Zhu 0003, Hengshu Zhu, Hui Xiong 0001 |
ICME | 5 |
| 2024 | Pre-DyGAE: Pre-training Enhanced Dynamic Graph Autoencoder for Occupational Skill Demand Forecasting
Xi Chen 0073, Chuan Qin 0002, Zhigaoyuan Wang, Yihang Cheng 0001, Chao Wang 0086, Hengshu Zhu, Hui Xiong 0001 |
IJCAI | 1 |
| 2024 | Job-SDF: A Multi-Granularity Dataset for Job Skill Demand Forecasting and BenchmarkingabstractIn a rapidly evolving job market, skill demand forecasting is crucial as it enables policymakers and businesses to anticipate and adapt to changes, ensuring that workforce skills align with market needs, thereby enhancing productivity and competitiveness. Additionally, by identifying emerging skill requirements, it directs individuals towards relevant training and education opportunities, promoting continuous self-learning and development. However, the absence of comprehensive datasets presents a significant challenge, impeding research and the advancement of this field. To bridge this gap, we present Job-SDF, a dataset designed to train and benchmark job-skill demand forecasting models. Based on millions of public job advertisements collected from online recruitment platforms, this dataset encompasses monthly recruitment demand.Our dataset uniquely enables evaluating skill demand forecasting models at various granularities, including occupation, company, and regional levels. We benchmark a range of models on this dataset, evaluating their performance in standard scenarios, in predictions focused on lower value ranges, and in the presence of structural breaks, providing new insights for further research. Our code and dataset are publicly accessible via the https://github.com/Job-SDF/benchmark. Xi Chen 0073, Chuan Qin 0002, Chuyu Fang, Chao Wang 0086, Chen Zhu 0003, Fuzhen Zhuang, Hengshu Zhu, Hui Xiong 0001 |
NeurIPS | 1 |