Ziyu Shang

dblp:243/1450 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0002-3737-4634ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 14 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Balanced Knowledge Distillation for Large Language Models with Mix-of-Experts
abstract
Mixture-of-Experts (MoE) architectures have recently become a more prevalent choice for large language models (LLMs) than dense architectures due to their superior performance. However, billions of parameters bring MoE LLMs a huge cost for deployment and inference. To address these issues, knowledge distillation (KD) has become a widely adopted technique to compress LLMs. Existing KD methods for LLMs can be divided into dense-to-dense and moe-to-dense distillation. Dense-to-dense distillation transfers knowledge between single dense LLMs, while moe-to-dense distillation attempts to transfer knowledge between the MoE LLMs and the dense LLMs. However, the architectural mismatch prevents the student from fully absorbing knowledge when distilling MoE LLMs. To address this limitation, we investigate a new distillation setting, moe-to-moe, which aims to fully leverage expert knowledge of teachers and enable the student to absorb it more effectively. Compared to dense-to-dense and moe-to-dense, moe-to-moe suffers from two imbalance issues. First, expert-coverage deficiency reflects an imbalanced knowledge transfer of teacher experts: traditional distillation utilizes only the few experts activated by the teacher router. Second, routing imbalance appears when the student routing distribution drifts from the teacher, which makes it difficult for students to learn how to distribute different experts. To overcome these issues, we propose a novel distillation framework for moe-to-moe, Balanced Distillation (B-Distill), which equally spreads teacher expertise across student experts while regularizing the student router toward teacher-consistent balance. First, to mitigate expert-coverage deficiency, we introduce Monte Carlo exploration, which stochastically perturbs router probabilities so every teacher and student expert is sampled without enlarging the search space. Second, to correct routing imbalance and avert load collapse, we propose an entropy-aware router distillation mechanism that aligns the student router with the teacher while curbing over-concentration. Experiments show that B-Distill outperforms baselines by up to 6.6% in Rouge-L.
Jiajun Liu 0005, Wenjun Ke 0002, Peng Wang 0004, Ziyu Shang, Zijie Xu 0003
AAAI5
2026 Optimizing LoRA Allocation of MoE with the Alignment of Topic Correlation
abstract
Mixture of experts (MoE) dynamically routes inputs to specialized expert networks to scale model capacity with low inference overhead. However, the excessive parameter growth in MoE models poses challenges in low-resource settings. To address these issues, MoE with parameter-efficient fine-tuning (PEFT) methods have emerged as a lightweight adaptation paradigm that distributes knowledge among experts via multiple LoRA blocks. Existing MoE-PEFT methods can be broadly categorized into External and Internal PEFT methods. External PEFT methods incorporate lightweight models into existing MoE architectures without modifying their routing, which limits the model’s parameter efficiency. To overcome these issues, Internal PEFT methods integrate MoE architectures into PEFT, enabling minimal parameter overhead. However, they still face two major challenges: (1) lack of expert functional differentiation, resulting in overlapping specialization across modules, and (2) absence of a structured attribution mechanism to guide expert selection based on semantic relevance. To alleviate these challenges, we propose TopicLoRA, a novel three-stage framework that leverages topic knowledge as semantic anchors to guide expert allocation. Specifically, (1) to address expert redundancy, we construct a topic-level prior graph using Graph Neural Network-enhanced representation learning over Big-Bench categories, enforcing structural separation among expert embeddings, and (2) to introduce semantic attribution, we design a dual-loss training mechanism that softly aligns input-query relevance with topic-guided routing distributions via KL divergence. Extensive experiments on representative datasets (e.g., MMLU, GSM8K, Flanv2) demonstrate that TopicLoRA outperforms state-of-the-art PEFT baselines by 2.40% on average in accuracy. Notably, the maximum improvement is 4.21%. Furthermore, ablation studies demonstrate that our framework's robustness to intricate topics and input sequence variations, which stems from the dual-loss training mechanism.
Hengyuan Xu, Wenjun Ke 0002, Jiajun Liu 0005, Dong Nie, Peng Wang 0004, Ziyu Shang, Zijie Xu 0003
AAAI7
2026 Benchmarking and Enhancing Rule Knowledge-Driven Reasoning of Large Language Models
abstract
Large Language Models (LLMs) have demonstrated strong capabilities across diverse tasks under the example-driven learning paradigm. However, in high-stakes domains such as emergency response and industrial safety, historical incidents are scarce, confidential, or both, while concise rule books are abundant. We formalize this underexplored setting as rule knowledge-driven reasoning and ask: Can LLMs reason reliably when rules are plentiful but examples are nearly absent? To study this question, we introduce RULER, an automatic benchmark that generates 32K rigorously verified questions from 1K expert-curated emergency response rules to probe three core abilities: rule memorization, single-rule application, and multi-rule complex reasoning. RULER is further equipped with a hallucination-aware evaluation suite and novel relational metrics. A comprehensive empirical study of five representative LLMs and five enhancement strategies shows that, even when models achieve reliable performance on rule memorization and single-rule application, multi-rule complex reasoning plateaus at 5.4 on a 10-point scale. To address this limitation, we propose RAMPS, a Rule knowledge-Aware Monte Carlo Tree Search Process-reward Supervision framework. RAMPS injects rule knowledge priors into MCTS, distills 12K step-level traces without human annotation, and trains an advantage-based reward model that scores candidate reasoning paths during beam search inference. Experimental results show that RAMPS significantly improves multi-rule complex reasoning performance to 7.7.
Zijie Xu 0003, Wenjun Ke 0002, Peng Wang 0004, Qingjian Ni, Jiajun Liu 0005, Ziyu Shang
AAAI7
2026 On the Role of Discriminative Models in Generative Relation Extraction
abstract
Relation extraction (RE) identifies semantic relations between entities in text, with existing methods falling into two main paradigms: discriminative and generative.Discriminative models encode sentences and entities into relation representations and classify the most likely relation, whereas generative models directly produce relation labels through sequence generation.Although the latter have benefited from recent advances in large language models (LLMs), their performance remains limited by bottlenecks.In this work, we present the systematic investigation of how discriminative models can support generative RE.We propose the Discriminative-to-Generative (D2G) framework, which first leverages discriminative models to produce a top-k set of candidate relations, and then integrates this knowledge into generative models via in-context or prompt learning.Extensive experiments on five benchmarks demonstrate that D2G consistently achieves state-of-the-art performance, with notable gains on long-tailed relation classes.
Peng Wang 0004, Zijie Xu 0003, Jiajun Liu 0005, Ziyu Shang
ACL (1)6
2026 Unlearning of Knowledge Graph Embedding via Preference Optimization
abstract
Existing knowledge graphs (KGs) inevitably contain outdated or erroneous knowledge that needs to be removed from knowledge graph embedding (KGE) models. To address this challenge, knowledge unlearning can be applied to eliminate specific information while preserving the integrity of the remaining knowledge in KGs. Existing unlearning methods can generally be categorized into exact unlearning and approximate unlearning. However, exact unlearning requires high training costs, while approximate unlearning faces two issues when applied to KGs due to the inherent connectivity of triples: (1) It fails to fully remove targeted information, as forgetting triples can still be inferred from remaining ones. (2) It focuses on local data for specific removal, which weakens the remaining knowledge in the forgetting boundary. To address these issues, we propose GraphDPO, a novel approximate unlearning framework based on direct preference optimization (DPO). Firstly, to effectively remove forgetting triples, we reframe unlearning as a preference optimization problem, where the model is trained by DPO to prefer reconstructed alternatives over the original forgetting triples. This formulation penalizes reliance on forgettable knowledge, mitigating incomplete forgetting caused by KG connectivity. Moreover, we introduce an out-boundary sampling strategy to construct preference pairs with minimal semantic overlap, weakening the connection between forgetting and retained knowledge. Secondly, to preserve boundary knowledge, we introduce a boundary recall mechanism that replays and distills relevant information both within and across time steps. We construct eight unlearning datasets across four popular KGs with varying unlearning rates. Experiments show that GraphDPO outperforms state-of-the-art baselines by up to 10.1% in MRR_Avg and 14.0% in MRR_F1. Further analysis confirms that GraphDPO more effectively removes target knowledge while preserving surrounding context.
Jiajun Liu 0005, Wenjun Ke 0002, Peng Wang 0004, Ziyu Shang, Zijie Xu 0003, Ke Ji
WWW5
2025 LLM-Guided Semantic-Aware Clustering for Topic Modeling
abstract
Topic modeling aims to discover the distribution of topics within a corpus. The advanced comprehension and generative capabilities of large language models (LLMs) have introduced new avenues for topic modeling, particularly by prompting LLMs to generate topics and refine them by merging similar ones. However, this approach necessitates that LLMs generate topics with consistent granularity, thus relying on the exceptional instruction-following capabilities of closed-source LLMs (such as GPT-4) or requiring additional training. Moreover, merging based only on topic words and neglecting the fine-grained semantics within documents might fail to fully uncover the underlying topic structure. In this work, we propose a semi-supervised topic modeling method, LiSA, that combines LLMs with clustering to improve topic generation and distribution. Specifically, we begin with prompting LLMs to generate a candidate topic word for each document, thereby constructing a topic-level semantic space. To further utilize the mutual complementarity between them, we first cluster documents and candidate topic words, and then establish a mapping from document to topic in the LLM-guided assignment stage. Subsequently, we introduce a collaborative enhancement strategy to align the two semantic spaces and establish a better topic distribution. Experimental results demonstrate that LiSA outperforms state-of-the-art methods that utilize GPT-4 on topic alignment, and exhibits competitive performance compared to Neural Topic Models on topic quality. The codes are available at https://github.com/ljh986/LiSA.
Jianghan Liu, Ziyu Shang, Wenjun Ke 0002, Peng Wang 0004, Zhizhao Luo, Jiajun Liu 0005
ACL (1)2
2025 Acquisition and Application of Novel Knowledge in Large Language Models
abstract
Ziyu Shang, Jianghan Liu, Zhizhao Luo, Peng Wang, Wenjun Ke, Jiajun Liu, Zijie Xu, Guozheng Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ziyu Shang, Jianghan Liu, Zhizhao Luo, Peng Wang 0004, Wenjun Ke 0002, Jiajun Liu 0005, Zijie Xu 0003
ACL (1)1
2025 Learning Multi-Granularity and Adaptive Representation for Knowledge Graph Reasoning
abstract
Knowledge graph reasoning (KGR) seeks to infer new factual triples from existing knowledge graphs (KGs). Recent methods have unified transductive and inductive reasoning by learning entity-independent representations through local neighboring structures. Nevertheless, these methods often encounter inefficiencies and rely on elaborate local structures without directly modeling the correlations between queries and various structures within KGs. In this paper, we propose a novel framework MulGA, which is designed to learn multi-granularity and adaptive embeddings for KGR. MulGA first employs connectivity subgraphs to uniformly and hierarchically represent query-related structures within KGs, such as triples, relation paths, and subgraphs, establishing the hierarchical relationship between structures at different granularities. Subsequently, we design a graph neural network-based multi-granularity embedding propagation module that unifies the message-passing process with the connectivity subgraph construction. This module obtains the query-related structural representations by all entities at multiple granularities, eliminating the need to explicitly extract any graph elements, thus addressing inefficiency issues. Moreover, we develop a structure-aware adaptive merging mechanism that assigns weights to different granularities and integrates them into cohesive subgraph-granularity representations for reasoning. The systematic experiments have been conducted on 15 benchmarks and MulGA achieves a significant improvement in MRR by an average of 0.5%-1.1% on transductive tasks and 0.2%-7.3% on inductive tasks than existing state-of-the-art methods. Moreover, MulGA exhibits faster convergence speed, smaller number of parameters, competitive inference time, and alleviates the over-smoothing prevalent in graph neural networks.
Ziyu Shang, Peng Wang 0004, Jianghan Liu, Jiajun Liu 0005, Zijie Xu 0003, Zhizhao Luo, Xiye Chen, Wenjun Ke 0002
IEEE Trans. Knowl. Data Eng.1
2024 Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
abstract
Current image-text retrieval methods have demonstrated impressive performance in recent years. However, they still face two problems: the inter-modal matching missing problem and the intra-modal semantic loss problem. These problems can significantly affect the accuracy of image-text retrieval. To address these challenges, we propose a novel method called Cross-modal and Uni-modal Soft-label Alignment (CUSA). Our method leverages the power of uni-modal pre-trained models to provide soft-label supervision signals for the image-text retrieval model. Additionally, we introduce two alignment techniques, Cross-modal Soft-label Alignment (CSA) and Uni-modal Soft-label Alignment (USA), to overcome false negatives and enhance similarity recognition between uni-modal samples. Our method is designed to be plug-and-play, meaning it can be easily applied to existing image-text retrieval models without changing their original architectures. Extensive experiments on various image-text retrieval models and datasets, we demonstrate that our method can consistently improve the performance of image-text retrieval and achieve new state-of-the-art results. Furthermore, our method can also boost the uni-modal retrieval performance of image-text retrieval models, enabling it to achieve universal retrieval. The code and supplementary files can be found at https://github.com/lerogo/aaai24_itr_cusa.
Hailang Huang, Zhijie Nie, Ziyu Shang
AAAI4
2024 Towards Continual Knowledge Graph Embedding via Incremental Distillation
abstract
Traditional knowledge graph embedding (KGE) methods typically require preserving the entire knowledge graph (KG) with significant training costs when new knowledge emerges. To address this issue, the continual knowledge graph embedding (CKGE) task has been proposed to train the KGE model by learning emerging knowledge efficiently while simultaneously preserving decent old knowledge. However, the explicit graph structure in KGs, which is critical for the above goal, has been heavily ignored by existing CKGE methods. On the one hand, existing methods usually learn new triples in a random order, destroying the inner structure of new KGs. On the other hand, old triples are preserved with equal priority, failing to alleviate catastrophic forgetting effectively. In this paper, we propose a competitive method for CKGE based on incremental distillation (IncDE), which considers the full use of the explicit graph structure in KGs. First, to optimize the learning order, we introduce a hierarchical strategy, ranking new triples for layer-by-layer learning. By employing the inter- and intra-hierarchical orders together, new triples are grouped into layers based on the graph structure features. Secondly, to preserve the old knowledge effectively, we devise a novel incremental distillation mechanism, which facilitates the seamless transfer of entity representations from the previous layer to the next one, promoting old knowledge preservation. Finally, we adopt a two-stage training paradigm to avoid the over-corruption of old knowledge influenced by under-trained new knowledge. Experimental results demonstrate the superiority of IncDE over state-of-the-art baselines. Notably, the incremental distillation mechanism contributes to improvements of 0.2%-6.5% in the mean reciprocal rank (MRR) score. More exploratory experiments validate the effectiveness of IncDE in proficiently learning new knowledge while preserving old knowledge across all time steps.
Jiajun Liu 0005, Wenjun Ke 0002, Peng Wang 0004, Ziyu Shang, Jinhua Gao, Ke Ji, Yanhe Liu
AAAI4
2024 Unify Named Entity Recognition Scenarios via Contrastive Real-Time Updating Prototype
abstract
Supervised named entity recognition (NER) aims to classify entity mentions into a fixed number of pre-defined types. However, in real-world scenarios, unknown entity types are continually involved. Naive fine-tuning will result in catastrophic forgetting on old entity types. Existing continual methods usually depend on knowledge distillation to alleviate forgetting, which are less effective on long task sequences. Moreover, most of them are specific to the class-incremental scenario and cannot adapt to the online scenario, which is more common in practice. In this paper, we propose a unified framework called Contrastive Real-time Updating Prototype (CRUP) that can handle different scenarios for NER. Specifically, we train a Gaussian projection model by a regularized contrastive objective. After training on each batch, we store the mean vectors of representations belong to new entity types as their prototypes. Meanwhile, we update existing prototypes belong to old types only based on representations of the current batch. The final prototypes will be used for the nearest class mean classification. In this way, CRUP can handle different scenarios through its batch-wise learning. Moreover, CRUP can alleviate forgetting in continual scenarios only with current data instead of old data. To comprehensively evaluate CRUP, we construct extensive benchmarks based on various datasets. Experimental results show that CRUP significantly outperforms baselines in continual scenarios and is also competitive in the supervised scenario.
Yanhe Liu, Peng Wang 0004, Wenjun Ke 0002, Xiye Chen, Jiteng Zhao, Ziyu Shang
AAAI7
2024 OntoFact: Unveiling Fantastic Fact-Skeleton of LLMs via Ontology-Driven Reinforcement Learning
abstract
Large language models (LLMs) have demonstrated impressive proficiency in information retrieval, while they are prone to generating incorrect responses that conflict with reality, a phenomenon known as intrinsic hallucination. The critical challenge lies in the unclear and unreliable fact distribution within LLMs trained on vast amounts of data. The prevalent approach frames the factual detection task as a question-answering paradigm, where the LLMs are asked about factual knowledge and examined for correctness. However, existing studies primarily focused on deriving test cases only from several specific domains, such as movies and sports, limiting the comprehensive observation of missing knowledge and the analysis of unexpected hallucinations. To address this issue, we propose OntoFact, an adaptive framework for detecting unknown facts of LLMs, devoted to mining the ontology-level skeleton of the missing knowledge. Specifically, we argue that LLMs could expose the ontology-based similarity among missing facts and introduce five representative knowledge graphs (KGs) as benchmarks. We further devise a sophisticated ontology-driven reinforcement learning (ORL) mechanism to produce error-prone test cases with specific entities and relations automatically. The ORL mechanism rewards the KGs for navigating toward a feasible direction for unveiling factual errors. Moreover, empirical efforts demonstrate that dominant LLMs are biased towards answering Yes rather than No, regardless of whether this knowledge is included. To mitigate the overconfidence of LLMs, we leverage a hallucination-free detection (HFD) strategy to tackle unfair comparisons between baselines, thereby boosting the result robustness. Experimental results on 5 datasets, using 32 representative LLMs, reveal a general lack of fact in current LLMs. Notably, ChatGPT exhibits fact error rates of 51.6% on DBpedia and 64.7% on YAGO, respectively. Additionally, the ORL mechanism demonstrates promising error prediction scores, with F1 scores ranging from 70% to 90% across most LLMs. Compared to the exhaustive testing, ORL achieves an average recall of 80% while reducing evaluation time by 35.29% to 63.12%.
Ziyu Shang, Wenjun Ke 0002, Nana Xiu, Peng Wang 0004, Jiajun Liu 0005, Zhizhao Luo, Ke Ji
AAAI1
2024 Unlocking Instructive In-Context Learning with Tabular Prompting for Relational Triple Extraction
abstract
The in-context learning (ICL) for relational triple extraction (RTE) has achieved promising performance, but still encounters two key challenges: (1) how to design effective prompts and (2) how to select proper demonstrations. Existing methods, however, fail to address these challenges appropriately. On the one hand, they usually recast RTE task to text-to-text prompting formats, which is unnatural and results in a mismatch between the output format at the pre-training time and the inference time for large language models (LLMs). On the other hand, they only utilize surface natural language features and lack consideration of triple semantics in sample selection. These issues are blocking improved performance in ICL for RTE, thus we aim to tackle prompt designing and sample selection challenges simultaneously. To this end, we devise a tabular prompting for RTE (TableIE) which frames RTE task into a table generation task to incorporate explicit structured information into ICL, facilitating conversion of outputs to RTE structures. Then we propose instructive in-context learning (I^2CL) which only selects and annotates a few samples considering internal triple semantics in massive unlabeled samples. Specifically, we first adopt off-the-shelf LLMs to perform schema-agnostic pre-extraction of triples in unlabeled samples using TableIE. Then we propose a novel triple-level similarity metric considering triple semantics between these samples and train a sample retrieval model based on calculated similarities in pre-extracted unlabeled data. We also devise three different sample annotation strategies for various scenarios. Finally, the annotated samples are considered as few-shot demonstrations in ICL for RTE. Experimental results on two RTE benchmarks show that I^2CL with TableIE achieves state-of-the-art performance compared to other methods under various few-shot RTE settings.
Wenjun Ke 0002, Peng Wang 0004, Zijie Xu 0003, Ke Ji, Jiajun Liu 0005, Ziyu Shang, Qiqing Luo
LREC/COLING7
2024 Domain-Hierarchy Adaptation via Chain of Iterative Reasoning for Few-shot Hierarchical Text Classification
Ke Ji, Peng Wang 0004, Wenjun Ke 0002, Jiajun Liu 0005, Jingsheng Gao, Ziyu Shang
IJCAI7
2024 Recall, Retrieve and Reason: Towards Better In-Context Relation Extraction
Peng Wang 0004, Wenjun Ke 0002, Yikai Guo, Ke Ji, Ziyu Shang, Jiajun Liu 0005, Zijie Xu 0003
IJCAI6
2024 Meta In-Context Learning Makes Large Language Models Better Zero and Few-Shot Relation Extractors
Peng Wang 0004, Jiajun Liu 0005, Yikai Guo, Ke Ji, Ziyu Shang, Zijie Xu 0003
IJCAI6
2024 Empirical Analysis of Dialogue Relation Extraction with Large Language Models
Zijie Xu 0003, Ziyu Shang, Jiajun Liu 0005, Ke Ji, Yikai Guo
IJCAI3
2024 Fast and Continual Knowledge Graph Embedding via Incremental LoRA
Jiajun Liu 0005, Wenjun Ke 0002, Peng Wang 0004, Jinhua Gao, Ziyu Shang, Zijie Xu 0003, Ke Ji
IJCAI6
2024 Learning Multi-Granularity and Adaptive Representation for Knowledge Graph Reasoning
Ziyu Shang, Peng Wang 0004, Wenjun Ke 0002, Jiajun Liu 0005, Hailang Huang, Chenxiao Wu, Jianghan Liu, Xiye Chen
IJCAI1
2024 Unveiling factuality and injecting knowledge for LLMs via reinforcement learning and data proportion
Wenjun Ke 0002, Ziyu Shang, Zhizhao Luo, Peng Wang 0004, Yikai Guo, Qi Liu 0056
Sci. China Inf. Sci.2
2023 IterDE: An Iterative Knowledge Distillation Framework for Knowledge Graph Embeddings
abstract
Knowledge distillation for knowledge graph embedding (KGE) aims to reduce the KGE model size to address the challenges of storage limitations and knowledge reasoning efficiency. However, current work still suffers from the performance drops when compressing a high-dimensional original KGE model to a low-dimensional distillation KGE model. Moreover, most work focuses on the reduction of inference time but ignores the time-consuming training process of distilling KGE models. In this paper, we propose IterDE, a novel knowledge distillation framework for KGEs. First, IterDE introduces an iterative distillation way and enables a KGE model to alternately be a student model and a teacher model during the iterative distillation process. Consequently, knowledge can be transferred in a smooth manner between high-dimensional teacher models and low-dimensional student models, while preserving good KGE performances. Furthermore, in order to optimize the training process, we consider that different optimization objects between hard label loss and soft label loss can affect the efficiency of training, and then we propose a soft-label weighting dynamic adjustment mechanism that can balance the inconsistency of optimization direction between hard and soft label loss by gradually increasing the weighting of soft label loss. Our experimental results demonstrate that IterDE achieves a new state-of-the-art distillation performance for KGEs compared to strong baselines on the link prediction task. Significantly, IterDE can reduce the training time by 50% on average. Finally, more exploratory experiments show that the soft-label weighting dynamic adjustment mechanism and more fine-grained iterations can improve distillation performance.
Jiajun Liu 0005, Peng Wang 0004, Ziyu Shang, Chenxiao Wu
AAAI3
2023 A Two-Stage Label Rectification Framework for Noisy Event Extraction
Zijie Xu 0003, Peng Wang 0004, Ziyu Shang, Jiajun Liu 0005
DASFAA (3)3
2023 ASKRL: An Aligned-Spatial Knowledge Representation Learning Framework for Open-World Knowledge Graph
Ziyu Shang, Peng Wang 0004, Yuzhang Liu, Jiajun Liu 0005, Wenjun Ke 0002
ISWC1