EDBT 2026 Demo / reviewers in the wild / expert
Xukai Liu
dblp:351/5535
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generative Query Augmentation with Dual-view Contrastive Learning for Dense Retrieval in Conversational Search
Wenyu Yan, Aoran Gan, Xukai Liu, Yanjiang Chen, Kai Zhang 0038, Qi Liu 0003 |
DASFAA (6) | 3 |
| 2026 | Learn to Understand: Knowledge Exemplification via Multi-Agent Cooperation for Science Question AnsweringabstractScience Question Answering (SQA) is an important task for evaluating models' capability to reason with scientific knowledge. However, the extensive availability of scientific information (e.g., basic concepts in biology, physics, and chemistry) in pre-trained corpora may cause large language models (LLMs) to rely more on memorized information rather than actual reasoning when answering questions. This reliance persists even with techniques like Chain-of-Thought prompting, resulting in shallow understanding and limited reasoning based on scientific knowledge. Therefore, to enhance LLMs' capacity to comprehend and apply scientific knowledge, we propose a framework calledMulti-AgentCooperation-basedKnowledgeExemplification (MCKE). Specifically, MCKE leverages knowledge alongside questions to create exemplified knowledge, promoting deeper understanding through innovative knowledge representation. To better evaluate the model's ability to reason and apply knowledge, we introduceNovSciQA, a multiple-choice question answering dataset based on newly created scientific knowledge. This dataset covers multi-subject scientific knowledge and questions that do not exist in reality, making it impossible for the model to rely on memorized answer-related information to answer questions. Experimental results show that the MCKE framework outperforms baselines, and the NovSciQA dataset effectively assesses models' knowledge understanding and application. Our code and dataset are available inhttps://anonymous.4open.science/r/MCKE-NovSciQA. Meikai Bao, Kai Zhang 0038, Xukai Liu, Qi Liu 0003, Hongke Zhao, Enhong Chen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Learnable Relational Knowledge Distillation For Language Model Compression
Feng Hu 0005, Kai Zhang 0038, Ye Liu 0011, Meikai Bao, Xukai Liu, Yanjiang Chen, Qi Liu 0003 |
DASFAA (6) | 5 |
| 2025 | Detect, Investigate, Judge and Determine: A Knowledge-Guided Framework for Few-Shot Fake News DetectionabstractFew-Shot Fake News Detection (FS-FND) aims to distinguish inaccurate news from real ones in extremely lowresource scenarios. This task has garnered increased attention due to the widespread dissemination and harmful impact of fake news on social media. Large Language Models (LLMs) have demonstrated competitive performance with the help of their rich prior knowledge and excellent in-context learning abilities. However, existing methods face significant limitations, such as the Understanding Ambiguity and Information Scarcity, which significantly undermine the potential of LLMs. To address these shortcomings, we propose a Dual-perspective Knowledge-guided Fake News Detection (DKFND) model, designed to enhance LLMs from both inside and outside perspectives. Specifically, DKFND first identifies the knowledge concepts of each news article through a Detection Module. Subsequently, DKFND creatively designs an Investigation Module to retrieve inside and outside valuable information concerning to the current news, followed by another Judge Module to evaluate the relevance and confidence of them. Finally, a Determination Module further derives two respective predictions and obtain the final result. Extensive experiments on two public datasets show the efficacy of our proposed method, particularly in low-resource settings. Ye Liu 0011, Xukai Liu, Haoyu Tang 0001, Yanghai Zhang, Kai Zhang 0038, Xiaofang Zhou 0001, Enhong Chen |
ICDM | 3 |
| 2025 | Learn while Unlearn: An Iterative Unlearning Framework for Generative Language ModelsabstractRecent advances in machine learning, particularly in Natural Language Processing (NLP), have produced powerful models trained on vast datasets. However, these models risk leaking sensitive information, raising privacy concerns. In response, regulatory measures such as the European Union's General Data Protection Regulation (GDPR) have driven increasing interest in Machine Unlearning techniques, which enable models to selectively forget specific data entries. Early unlearning approaches primarily relied on pre-processing methods, while more recent research has shifted towards training-based solutions. Despite their effectiveness, a key limitation persists: most methods require access to original training data, which is often unavailable. Additionally, directly applying unlearning techniques bears the cost of undermining the model's expressive capabilities. To address these challenges, we introduce the Iterative Contrastive Unlearning (ICU) framework, which consists of three core components: A Knowledge Unlearning Induction module designed to target specific knowledge for removal using an unlearning loss; A Contrastive Learning Enhancement module to preserve the model's expressive capabilities against the pure unlearning goal; And an Iterative Unlearning Refinement module that dynamically adjusts the unlearning process through ongoing evaluation and updates. Experimental results demonstrate the efficacy of our ICU method in unlearning sensitive information while maintaining the model's overall performance, offering a promising solution for privacy-conscious machine learning applications. Haoyu Tang 0001, Ye Liu 0011, Xi Zhao 0006, Xukai Liu, Yanghai Zhang, Kai Zhang 0038, Xiaofang Zhou 0001, Enhong Chen |
ICDM | 4 |
| 2025 | An Efficient Certificateless Key-Insulated Anonymous Signature Scheme Based on Smart Contract for Data Sharing in Industrial Internet of ThingsabstractIn the Industrial Internet of Things (IIoT) environment, a multitude of sensing devices continually gather critical data. These data are indispensable for the operations and advancements across diverse industries. However, the sharing of these data poses privacy threats, with attackers exploiting channel analysis and physical device attacks to access sensitive data. To address this, we propose a privacy protection scheme that combines smart contracts (SCs), key-insulated technology, and certificateless anonymous signature (CLBS). This scheme aims to ensure data privacy during sharing and maintain user anonymity. By leveraging SCs, our scheme enables fair and automated key distribution, replacing traditional key generation centers. Key-insulated technology ensures that the signer’s key changes periodically, enhancing system stability. The security of our solution is validated through a random oracle model, and we have optimized an elliptic curve point to reduce signature length and minimize communication overhead. Our scheme outperforms other CLBS schemes in terms of computational and communication efficiency. Nana Kong, Zhifeng Wan, Cui Xu, Xukai Liu, Yixin Yuan |
IEEE Internet Things J. | 4 |
| 2025 | Semantic-Aligned Code Summarization: Bridging the Gap Between Code and Natural Language Through Data Flow AnalysisabstractCode summarization is designed to generate descriptive natural language for code snippets, facilitating understanding and increasing productivity for developers. Previous research often overlooks the semantic connection between code and its natural language description, resulting in a noticeable gap and suboptimal solution. To address this issue, we introduce a semantic-aligned code summarization framework that leverages crucial data flow information from code for semantic analysis, ensuring alignment between code and summaries. Specifically, we utilize a semantic extraction module (SEM) to decipher the meaning of code and align it with natural language through a semantic alignment module. In the SEM, we construct a code graph that includes data flow edges using static program analysis techniques. Then, on this well-constructed code graph, we innovatively adopt a walking algorithm guided by data flow to extract the semantics of the code. This walking algorithm understands code semantics by analyzing the information transfer between variables during the program execution process. In the semantic alignment module, we integrate a contrastive learning loss mechanism for semantic alignment, which cohesively maps the semantic domains of code and natural language into a unified vector space. We further theoretically analyzed that the data-flow-guided walking algorithm can ensure capturing semantically highly related nodes in shorter paths. Extensive experiments on two benchmark datasets demonstrate the efficacy and broad applicability of the framework. Yuze Zhao, Zhenya Huang, Kai Zhang 0038, Weibo Gao, Qi Liu 0003, Xukai Liu, Fangzhou Yao, Enhong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | OneNet: A Fine-Tuning Free Framework for Few-Shot Entity Linking via Large Language Model PromptingabstractEntity Linking (EL) is the process of associating ambiguous textual mentions to specific entities in a knowledge base.Traditional EL methods heavily rely on large datasets to enhance their performance, a dependency that becomes problematic in the context of few-shot entity linking, where only a limited number of examples are available for training.To address this challenge, we present OneNet, an innovative framework that utilizes the few-shot learning capabilities of Large Language Models (LLMs) without the need for fine-tuning.To the best of our knowledge, this marks a pioneering approach to applying LLMs to few-shot entity linking tasks.OneNet is structured around three key components prompted by LLMs: (1) an entity reduction processor that simplifies inputs by summarizing and filtering out irrelevant entities, (2) a dual-perspective entity linker that combines contextual cues and prior knowledge for precise entity linking, and (3) an entity consensus judger that employs a unique consistency algorithm to alleviate the hallucination in the entity linking reasoning.Comprehensive evaluations across seven benchmark datasets reveal that OneNet outperforms current stateof-the-art entity linking methods. Xukai Liu, Ye Liu 0011, Kai Zhang 0038, Kehang Wang, Qi Liu 0003, Enhong Chen |
EMNLP | 1 |
| 2024 | Puncturable-based broadcast encryption with tracking for preventing malicious encryptors in cloud file sharing
Yingzi Hu, Xu An Wang 0014, Xukai Liu, Yuqing Yin |
J. Inf. Secur. Appl. | 4 |