Chuyu Fang

dblp:327/9420 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0007-8353-7979ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TLSA: LLM-Guided Text-Label Space Alignment with Contrastive Learning for Generalized Category Discovery
abstract
Generalized Category Discovery (GCD) aims to classify data from partially labeled datasets by jointly recognizing known categories and discovering novel ones.Despite recent advances, existing methods still suffer from weak text-label alignment, inconsistent objectives across known and novel categories, and poor discrimination of semantically similar clusters.To mitigate these issues, we propose TLSA, a unified framework that enforces contrastive alignment between text and label representations within a shared semantic space.Specifically, we first design a label-semantic aware dual-encoder equipped with a symmetric contrastive objective to achieve text-label alignment.Then, we leverage LLM-based label induction to generate explicit and semantically meaningful names for previously unseen categories, followed by a graph-based refinement strategy that disambiguates semantically overlapping clusters through forced renaming.Finally, a confidence-aware sampling strategy ensures balanced learning across both easy and hard instances.Extensive experiments on four benchmark datasets show that TLSA consistently outperforms state-of-theart GCD methods.The code is available at https://github.com/Wenxi-Xu/TLSA.
Wenxi Xu, Chuan Qin 0002, Xi Chen 0073, Chuyu Fang, Yuanchun Zhou, Hengshu Zhu
ACL (1)4
2025 Enhancing Dual-Target Cross-Domain Recommendation via Similar User Bridging
abstract
Dual-target cross-domain recommendation aims to mitigate data sparsity and enables mutual enhancement via bidirectional knowledge transfer. Most existing methods rely on overlapping users to build cross-domain connections. However, in many real-world scenarios, overlapping data is extremely limited-or even entirely absent-significantly diminishing the effectiveness of these methods. To address this challenge, we propose SUBCDR, a novel framework that leverages large language models (LLMs) to bridge similar users across domains, thereby enhancing dual-target cross-domain recommendation. Specifically, we introduce a Multi-Interests-Aware Prompt Learning mechanism that enables LLMs to generate comprehensive user profiles, disentangling domain-invariant interest points while capturing fine-grained preferences. Then, we construct intra-domain bipartite graphs from user-item interactions and an inter-domain heterogeneous graph that links similar users across domains. Subsequently, to facilitate effective knowledge transfer, we employ Graph Convolutional Networks (GCNs) for intra-domain relationship modeling and design an Inter-domain Hierarchical Attention Network (InterHAN) to facilitate inter-domain knowledge transfer through similar users, learning both shared and specific user representations. Extensive experiments on seven public datasets demonstrate that SUBCDR outperforms state-of-the-art cross-domain recommendation algorithms and single-domain recommendation methods. Our code is publicly available at https://github.com/97z/SUBCDR.git.
Xi Chen 0073, Chuyu Fang, Jianji Wang 0001, Chuan Qin 0002, Fuzhen Zhuang
CIKM3
2025 From Missteps to Mastery: Enhancing Low-Resource Dense Retrieval through Adaptive Query Generation
abstract
Document retrieval, designed to recall query-relevant documents from expansive collections, is essential for information-seeking tasks, such as web search and open-domain question-answering. Advances in representation learning and pretrained language models (PLMs) have driven a paradigm shift from traditional sparse retrieval methods to more effective dense retrieval approaches, forging enhanced semantic connections between queries and documents and establishing new performance benchmarks. However, reliance on extensive annotated document-query pairs limits their competitiveness in low-resource scenarios. Recent research efforts employing the few-shot capabilities of large language models (LLMs) and prompt engineering for synthetic data generation have emerged as a promising solution. Nonetheless, these approaches are hindered by the generation of lower-quality data within the conventional dense retrieval training process. To this end, in this paper, we introduce iGFT, a framework aimed at enhancing low-resource dense retrieval by integrating a three-phase process --- Generation, Filtering, and Tuning --- coupled with an iterative optimization strategy. Specifically, we first employ supervised fine-tuning on limited ground truth data, enabling an LLM to function as the generator capable of producing potential queries from given documents. Subsequently, we present a multi-stage filtering module to minimize noise in the generated data while retaining samples poised to significantly improve the dense retrieval model's performance in the follow-up fine-tuning process. Furthermore, we design a novel iterative optimization strategy that dynamically optimizes the query generator for producing more informative queries, thereby enhancing the efficacy of the entire framework. Finally, extensive experiments conducted on a series of publicly available retrieval benchmark datasets have demonstrated the effectiveness of the proposed iGFT.
Zhenyu Tong, Chuan Qin 0002, Chuyu Fang, Kaichun Yao, Xi Chen 0073, Jingshuai Zhang, Chen Zhu 0003, Hengshu Zhu
KDD (1)3
2024 Enhancing Question Answering for Enterprise Knowledge Bases using Large Language Models
Feihu Jiang, Chuan Qin 0002, Kaichun Yao, Chuyu Fang, Fuzhen Zhuang, Hengshu Zhu, Hui Xiong 0001
DASFAA (4)4
2024 Job-SDF: A Multi-Granularity Dataset for Job Skill Demand Forecasting and Benchmarking
abstract
In a rapidly evolving job market, skill demand forecasting is crucial as it enables policymakers and businesses to anticipate and adapt to changes, ensuring that workforce skills align with market needs, thereby enhancing productivity and competitiveness. Additionally, by identifying emerging skill requirements, it directs individuals towards relevant training and education opportunities, promoting continuous self-learning and development. However, the absence of comprehensive datasets presents a significant challenge, impeding research and the advancement of this field. To bridge this gap, we present Job-SDF, a dataset designed to train and benchmark job-skill demand forecasting models. Based on millions of public job advertisements collected from online recruitment platforms, this dataset encompasses monthly recruitment demand.Our dataset uniquely enables evaluating skill demand forecasting models at various granularities, including occupation, company, and regional levels. We benchmark a range of models on this dataset, evaluating their performance in standard scenarios, in predictions focused on lower value ranges, and in the presence of structural breaks, providing new insights for further research. Our code and dataset are publicly accessible via the https://github.com/Job-SDF/benchmark.
Xi Chen 0073, Chuan Qin 0002, Chuyu Fang, Chao Wang 0086, Chen Zhu 0003, Fuzhen Zhuang, Hengshu Zhu, Hui Xiong 0001
NeurIPS3
2023 RecruitPro: A Pretrained Language Model with Skill-Aware Prompt Learning for Intelligent Recruitment
abstract
Recent years have witnessed the rapid development of machine-learning-based intelligent recruitment services. Along this line, a large number of emerging models have been proposed, achieving remarkable performance in various tasks, such as person-job fit, job classification and salary prediction. However, existing studies are usually domain/task specific, which significantly hinders the adaptation of models for different industries/tasks with limited training data. To this end, in this paper, we propose a novel skill-aware prompt-based pretraining framework, namely RecruitPro, which is capable of learning unified representations on the recruitment data and adapting for various downstream tasks of intelligent recruitment services. To be specific, we first present a contextualized embedding model that is pretrained on a large-scale recruitment dataset. Then, we construct 13 downstream benchmark tasks that are representative in the recruitment process. Along this line, we propose a skill-aware prompt learning module to enhance the adaptability of the pretrained model on downstream tasks. This module includes a skill-related prompt, which is designed to explore key semantic information (i.e., skills) from recruitment text, and a task-related prompt, which is designed to bridge the gap between the pretrained model and different downstream tasks. Moreover, we propose a strategy for extracting potential skills to further improve the performance of our skill-aware prompt learning module. Finally, extensive experiments have clearly demonstrated the effectiveness of RecruitPro. In addition, a case study has been presented to discuss the privacy preserving issue of our RecruitPro.
Chuyu Fang, Chuan Qin 0002, Qi Zhang 0053, Kaichun Yao, Jingshuai Zhang, Hengshu Zhu, Fuzhen Zhuang, Hui Xiong 0001
KDD1
2023 Hierarchical Neural Topic Model with Embedding Cluster and Neural Variational Inference
abstract
Compared to flat topic models, hierarchical topic models not only exploit inherent structural information in the corpus but detect better semantic topics with the help of hierarchy knowledge. Recently, Neural-Variational-Inference (NVI) based hierarchical neural topic models have achieved better performance. However, existing NVI-based models learn topics of different levels with the same strategy, i.e., word co-occurrence patterns, which causes that topics of different levels cannot be distinguished from a semantic perspective and topics of the first level degenerate into some meaningless common words. To address the above problems, we propose a novel Hierarchical Neural Topic Model with embedding cluster and neural variational inference (C-HNTM). Specifically, C-HNTM adopts Gaussian Mixture Model (GMM) to learn topics of the first level based on word embeddings, which can capture the global semantic information of the whole corpus and generate more meaningful and global semantic topics. Then, the NVI-based method is adopted to learn topics of the second level with Bag-of-Word from a document perspective, which can generate local and more detailed topics. Third, we simultaneously learn global and local topic distributions and dependency matrix by using Stochastic Gradient Variational Bayes (SGVB) estimator. Finally, we provide the detailed inference of variational lower bound and extensive experiments on three real-world datasets to validate the effectiveness of our model.
Ningjing Wang, Chenguang Du, Chuyu Fang, Fuzhen Zhuang
SDM5
2022 Improving biomedical named entity recognition by dynamic caching inter-sentence information
abstract
MOTIVATION: Biomedical Named Entity Recognition (BioNER) aims to identify biomedical domain-specific entities (e.g. gene, chemical and disease) from unstructured texts. Despite deep learning-based methods for BioNER achieving satisfactory results, there is still much room for improvement. Firstly, most existing methods use independent sentences as training units and ignore inter-sentence context, which usually leads to the labeling inconsistency problem. Secondly, previous document-level BioNER works have approved that the inter-sentence information is essential, but what information should be regarded as context remains ambiguous. Moreover, there are still few pre-training-based BioNER models that have introduced inter-sentence information. Hence, we propose a cache-based inter-sentence model called BioNER-Cache to alleviate the aforementioned problems. RESULTS: We propose a simple but effective dynamic caching module to capture inter-sentence information for BioNER. Specifically, the cache stores recent hidden representations constrained by predefined caching rules. And the model uses a query-and-read mechanism to retrieve similar historical records from the cache as the local context. Then, an attention-based gated network is adopted to generate context-related features with BioBERT. To dynamically update the cache, we design a scoring function and implement a multi-task approach to jointly train our model. We build a comprehensive benchmark on four biomedical datasets to evaluate the model performance fairly. Finally, extensive experiments clearly validate the superiority of our proposed BioNER-Cache compared with various state-of-the-art intra-sentence and inter-sentence baselines. AVAILABILITYAND IMPLEMENTATION: Code will be available at https://github.com/zgzjdx/BioNER-Cache. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yiqi Tong, Fuzhen Zhuang, Chuyu Fang, Yu Zhao 0019, Deqing Wang 0001, Hengshu Zhu, Bin Ni
Bioinform.4