Simin Niu

dblp:362/0996 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
15since 2021 · last 2026
0009-0009-1862-8959ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 SEAP: Sparse Expert Activation Pruning Unlocks the Brainpower of Large Language Models
abstract
Pruning is a promising approach to reduce the high inference cost of large language models (LLMs), but it often comes at the expense of performance. Motivated by the "functional localization" theory in neuroscience, we hypothesize that LLMs contain task-specific expert activation paths, where specific subsets of neurons are co-activated for particular tasks. This structure allows selective activation to preserve task performance while improving inference efficiency. We introduce Sparse Expert Activation Pruning (SEAP), a training-free pruning method for large language models. SEAP identifies task-relevant activation paths by analyzing the clustering patterns of hidden states and neuron activations on a multi-task calibration dataset. Cross-task transfer evaluations confirm the existence of such expert activation structures. SEAP constructs task-aware pruning masks by leveraging a task-expert calibration dataset, which provides representative samples across diverse tasks to reveal their activation signatures. It then employs a lightweight task router to dynamically select relevant computation paths based on the input task. This design significantly reduces inference cost without compromising accuracy. Experimental results show that SEAP retains model performance with only a 1.5% drop on most tasks at 20% sparsity, and at 50% sparsity, it surpasses strong pruning baselines such as WandA and FLAP by over 20%. These results highlight SEAP as a scalable and effective solution for efficient LLM inference.
Xun Liang 0001, Huayi Lai, Simin Niu, Shichao Song, Jihao Zhao, Feiyu Xiong, Bo Tang 0018
AAAI4
2026 Safe RAG by RAG: Untying the Bell That RAG Rang with the RAG Hand
abstract
Retrieval-augmented generation (RAG) is widely adopted for knowledge-intensive tasks, but unverified external knowledge can pose risks such as data injection and retrieval pollution, leading to unexpected generation. Existing defenses rely on patch-based fixes, which limit generalization and increase system latency. To address these issues, we propose RAG2RAG, a framework-level security solution specifically designed for RAG. Inspired by human intuition to reason about what can and cannot be said during RAG phase, RAG2RAG augments the main RAG module with a lightweight RAG-based security expert module composed of two components: (1) a Detective that dynamically retrieves supporting evidence, and (2) a Judge that makes final decisions based on retrieved context. The main and expert modules operate in parallel without causing noticeable delays. Experiments across two languages, six domains, and seven types of poisoning attacks demonstrate that RAG2RAG overall achieves higher accuracy and lower attack success rates than seven mainstream baselines. Furthermore, it integrates seamlessly with various RAG architectures, offering efficient protection across diverse threat scenarios.
Xun Liang 0001, Mengwei Wang, Yuefeng Ma, Simin Niu
AAAI4
2026 Empowering Large Language Models to Set Up Knowledge Retrieval Indexing via Self-Learning
abstract
Retrieval-augmented generation (RAG) provides an efficient solution for expanding the knowledge boundaries of large language models (LLMs), where the indexing serves as a compass to guide LLMs in locating query-relevant external knowledge. Nevertheless, current indexing methods commonly encounter a critical challenge: native indexing is convenient to construct, but it usually disrupts contextual associations and constrains the expressive capacity of rich knowledge. Conversely, knowledge indexing can structure contextual knowledge, but it is often based on preset schemas that limit its generalizability. To address it, we propose a universal and flexible knowledge indexing called pseudo-graph (PG) indexing. During the indexing construction phase, we use the advanced LLMs to transform the knowledge of each raw text into a concise and structured mind map, organizing intra-document knowledge. Subsequently, independent mind maps are linked by associating highly relevant topics or consistent facts across documents, thereby establishing inter-document knowledge connections. Eventually, using the resulting knowledge network PG as the knowledge indexing can circumvent the challenges associated with schema design reliant on preset knowledge and relationship types. During the knowledge retrieval phase, we develop a PG knowledge retriever to mimic human note-reviewing, adaptively navigating and recalling query-relevant knowledge from PG. Experimental results demonstrate that retrieving relevant pseudo-subgraphs from the PG via PG indexing and retriever significantly improves performance in fact-based Q&A, hallucination correction, and two multi-document Q&A tasks, achieving$F1_{QE}$improvements of 15.85%, 8.12%, 3.34%, and 5.73%, respectively, and outperforming the state-of-the-art baseline KGP-LLaMA. Our code is available at:https://github.com/IAAR-Shanghai/PGRAG.
Simin Niu, Mengwei Wang, Xun Liang 0001, Sensen Zhang, Shichao Song, Feiyu Xiong, Chenyang Xi
IEEE Trans. Knowl. Data Eng.1
2025 Integrating Large Language Models and Möbius Group Transformations for Temporal Knowledge Graph Embedding on the Riemann Sphere
abstract
The significance of Temporal Knowledge Graphs (TKGs) in Artificial Intelligence (AI) lies in their capacity to incorporate time-dimensional information, support complex reasoning and prediction, optimize decision-making processes, enhance the accuracy of recommendation systems, promote multimodal data integration, and strengthen knowledge management and updates. This provides a robust foundation for various AI applications. To effectively learn and apply both static and dynamic temporal patterns for reasoning, a range of embedding methods and large language models (LLMs) have been proposed in the literature. However, these methods often rely on a single underlying embedding space, whose geometric properties severely limit their ability to model intricate temporal patterns, such as hierarchical and ring structures. To address this limitation, this paper proposes embedding TKGs into projective geometric space and leverages LLMs technology to extract crucial temporal node information, thereby constructing the 5EL model. By embedding TKGs into projective geometric space and utilizing Möbius Group transformations, we effectively model various temporal patterns. Subsequently, LLMs technology is employed to process the trained TKGs. We adopt a parameter-efficient fine-tuning strategy to align LLMs with specific task requirements, thereby enhancing the model's ability to recognize structural information of key nodes in historical chains and enriching the representation of central entities. Experimental results on five advanced TKG datasets demonstrate that our proposed 5EL model significantly outperforms existing models.
Sensen Zhang, Xun Liang 0001, Simin Niu, Zhendong Niu, Bo Wu 0026, Gengxin Hua, Zhenyu Guan 0003, Xuan Zhang 0009, Yuefeng Ma
AAAI3
2025 SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model
abstract
Xun Liang, Simin Niu, Zhiyu Li, Sensen Zhang, Hanyu Wang, Feiyu Xiong, Zhaoxin Fan, Bo Tang, Jihao Zhao, Jiawei Yang, Shichao Song, Mengwei Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xun Liang 0001, Simin Niu, Sensen Zhang, Feiyu Xiong, Jason Zhaoxin Fan, Bo Tang 0011, Jihao Zhao, Shichao Song, Mengwei Wang
ACL (1)2
2025 QAEncoder: Towards Aligned Representation Learning in Question Answering Systems
abstract
Zhengren Wang, Qinhan Yu, Shida Wei, Zhiyu Li, Feiyu Xiong, Xiaoxing Wang, Simin Niu, Hao Liang, Wentao Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhengren Wang, Qinhan Yu, Shida Wei, Feiyu Xiong, Xiaoxing Wang, Simin Niu, Hao Liang 0017, Wentao Zhang 0001
ACL (1)7
2025 GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and Reasoning
abstract
The evaluation of large language models (LLMs) has traditionally relied on static benchmarks, a paradigm that poses two major limitations: (1) predefined test sets lack adaptability to diverse application domains, and (2) standardized evaluation protocols often fail to capture fine-grained assessments of domain-specific knowledge and contextual reasoning abilities. To overcome these challenges, we propose GuessArena, an adaptive evaluation framework grounded in adversarial game-based interactions. Inspired by the interactive structure of the Guess Who I Am? game, our framework seamlessly integrates dynamic domain knowledge modeling with progressive reasoning assessment to improve evaluation fidelity. Empirical studies across five vertical domains-finance, healthcare, manufacturing, information technology, and education-demonstrate that GuessArena effectively distinguishes LLMs in terms of domain knowledge coverage and reasoning chain completeness. Compared to conventional benchmarks, our method provides substantial advantages in interpretability, scalability, and scenario adaptability.
Qingchen Yu 0001, Zifan Zheng, Simin Niu, Bo Tang 0011, Feiyu Xiong
ACL (1)4
2025 MoC: Mixtures of Text Chunking Learners for Retrieval-Augmented Generation System
abstract
Retrieval-Augmented Generation (RAG), while serving as a viable complement to large language models (LLMs), often overlooks the crucial aspect of text chunking within its pipeline. This paper initially introduces a dual-metric evaluation method, comprising Boundary Clarity and Chunk Stickiness, to enable the direct quantification of chunking quality. Leveraging this assessment method, we highlight the inherent limitations of traditional and semantic chunking in handling complex contextual nuances, thereby substantiating the necessity of integrating LLMs into chunking process. To address the inherent trade-off between computational efficiency and chunking precision in LLM-based approaches, we devise the granularity-aware Mixture-of-Chunkers (MoC) framework, which consists of a three-stage processing mechanism. Notably, our objective is to guide the chunker towards generating a structured list of chunking regular expressions, which are subsequently employed to extract chunks from the original text. Extensive experiments demonstrate that both our proposed metrics and the MoC framework effectively settle challenges of the chunking task, revealing the chunking kernel while enhancing the performance of the RAG system.
Jihao Zhao, Zhaoxin Fan, Simin Niu, Bo Tang 0011, Feiyu Xiong
ACL (1)5
2025 Retrieval-Augmented Multilingual Citation Generation
abstract
Retrieval-augmented citation generation (RACG) helps users trust the large language model output by retrieving evidence from reliable sources. However, most current RACG research focuses on single-language tasks, particularly in English, and overlooks the need for cross-lingual evidence retrieval and utilization in real-world applications. To address this issue, we introduce a plug-and-play Retrieval-Augmented Multilingual Citation Generation method (RAMCG) which uses a multilingual retriever to identify relevant evidence from a multilingual knowledge base. The evidence is then combined with the query and processed by a multilingual citation generator. The result is citations that are both accurate and comprehensive. Experiments show that RAMCG outperforms baseline methods in multilingual citation generation and is well-suited for practical use.
Xun Liang 0001, Simin Niu, Sensen Zhang, Xuan Zhang 0009, Bo Wu 0026, Feiyu Xiong, Bo Tang 0011, Shichao Song, Mengwei Wang
ICASSP2
2025 When Sparse Graph Representation Learning Falls into Domain Shift: Feature Augmentation for Cross-Domain Graph Meta-Learning
abstract
Graph Meta-learning methods have improved the performance of few-shot node classification by means of applying meta-learning to the data in non-Euclidean domains. However, most works focus on adopting a single domain, ignoring the fact that tasks in various domains may be distinct, which can cause overfitting problems and thus limit generalizability. To tackle this challenge, we propose a novel Graph Meta-learning framework called Feature-Enhanced Cross-domain Graph Meta-learning that consists of two crucial modules: 1) A feature information extraction module that aims to capture discriminative node importance and simulate various node feature distributions under distinct domains; 2) A heterogeneous graph encoder module that leverages the enhanced node features and topological information to generate task-specific node embeddings with simple fine-tuning. Moreover, we meta-learn the parameters involved to ensure the generalizability in the unseen domains. Results show that our method markedly outperforms the existing state-of-the-art methods.
Simin Niu, Xun Liang 0001, Sensen Zhang, Xuan Zhang 0009, Wu Bo, Shichao Song, Mengwei Wang
ICASSP1
2025 CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models
abstract
Retrieval-augmented generation (RAG) is a technique that enhances the capabilities of large language models (LLMs) by incorporating external knowledge sources. This method addresses common LLM limitations, including outdated information and the tendency to produce inaccurate “hallucinated” content. However, evaluating RAG systems is a challenge. Most benchmarks focus primarily on question-answering applications, neglecting other potential scenarios where RAG could be beneficial. Accordingly, in the experiments, these benchmarks often assess only the LLM components of the RAG pipeline or the retriever in knowledge-intensive scenarios, overlooking the impact of external knowledge base construction and the retrieval component on the entire RAG pipeline in non-knowledge-intensive scenarios. To address these issues, this article constructs a large-scale and more comprehensive benchmark and evaluates all the components of RAG systems in various RAG application scenarios. Specifically, we refer to the CRUD actions that describe interactions between users and knowledge bases and also categorize the range of RAG applications into four distinct types—create, read, update, and delete (CRUD). “Create” refers to scenarios requiring the generation of original, varied content. “Read” involves responding to intricate questions in knowledge-intensive situations. “Update” focuses on revising and rectifying inaccuracies or inconsistencies in pre-existing texts. “Delete” pertains to the task of summarizing extensive texts into more concise forms. For each of these CRUD categories, we have developed different datasets to evaluate the performance of RAG systems. We also analyze the effects of various components of the RAG system, such as the retriever, context length, knowledge base construction, and LLM. Finally, we provide useful insights for optimizing the RAG technology for different scenarios. The source code is available at GitHub: https://github.com/IAAR-Shanghai/CRUD_RAG .
Yuanjie Lyu, Simin Niu, Feiyu Xiong, Bo Tang 0018, Wenjin Wang 0003, Hao Wu 0022, Huanyong Liu, Tong Xu 0001, Enhong Chen
ACM Trans. Inf. Syst.3
2024 When Sparse Graph Representation Learning Falls into Domain Shift: Data Augmentation for Cross-Domain Graph Meta-Learning (Student Abstract)
abstract
Cross-domain Graph Meta-learning (CGML) has shown its promise, where meta-knowledge is extracted from few-shot graph data in multiple relevant but distinct domains. However, several recent efforts assume target data available, which commonly does not established in practice. In this paper, we devise a novel Cross-domain Data Augmentation for Graph Meta-Learning (CDA-GML), which incorporates the superiorities of CGML and Data Augmentation, has addressed intractable shortcomings of label sparsity, domain shift, and the absence of target data simultaneously. Specifically, our method simulates instance-level and task-level domain shift to alleviate the cross-domain generalization issue in conventional graph meta-learning. Experiments show that our method outperforms the existing state-of-the-art methods.
Simin Niu, Xun Liang 0001, Sensen Zhang, Shichao Song, Xuan Zhang 0009
AAAI1
2024 Biomedical Knowledge Graph Embedding with Householder Projection (Student Abstract)
abstract
Researchers have applied knowledge graph embedding (KGE) techniques with advanced neural network techniques, such as capsule networks, for predicting drug-drug interactions (DDIs) and achieved remarkable results. However, most ignore molecular structure and position features between drug pairs. They cannot model the biomedical field's significant relational mapping properties (RMPs,1-N, N-1, N-N) relation. To solve these problems, we innovatively propose CDHse that consists of two crucial modules: 1) Entity embedding module, we obtain position feature obtained by PubMedBERT and Convolutional Neural Network (CNN), obtain molecular structure feature with Graphic Nuaral Network (GNN), obtain entity embedding feature of drug pairs, and then incorporate these features into one synthetic feature. 2) Knowledge graph embedding module, the synthetic feature is Householder projections and then embedded in the complex vector space for training. In this paper, we have selected several advanced models for the DDIs task and performed experiments on three standard BioKG to validate the effectiveness of CDHse.
Sensen Zhang, Xun Liang 0001, Simin Niu, Xuan Zhang 0009, Yuefeng Ma
AAAI3
2024 UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained Generation
abstract
Xun Liang, Shichao Song, Simin Niu, Zhiyu Li, Feiyu Xiong, Bo Tang, Yezhaohui Wang, Dawei He, Cheng Peng, Zhonghao Wang, Haiying Deng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xun Liang 0001, Shichao Song, Simin Niu, Feiyu Xiong, Bo Tang 0011, Yezhaohui Wang, Dawei He, Haiying Deng
ACL (1)3
2024 Temporal Knowledge Graph Embedding using Householder Transformations
abstract
The rapid development of Knowledge Graph (KG) technology has led to the emergence of Temporal Knowledge Graphs (TKGs), which hold significant research importance and value. Temporal Knowledge Graph Embedding (TKGE) techniques complement TKGs and predict links within them. The efficacy of TKGE hinges upon its capability to effectively model intrinsic temporal relation patterns. However, existing methodologies often need to capture temporal relation patterns adequately or establish intrinsic connections between evolving relations. We propose a highly robust TKGE framework called HouRP to address this limitation and incorporate a more comprehensive range of relational information. This framework introduces two types of Householder transformations into TKGE. Specifically, the Householder projections enable the generation of temporal-specific representations for each entity. At the same time, the relational Householder rotations facilitate high-dimensional rotations between projected entities, thereby capturing relation-specific properties. Our proposed model’s effectiveness is demonstrated through extensive experimentation on four widely used TKGs datasets.
Sensen Zhang, Xun Liang 0001, Simin Niu, Junlan Feng, Mengwei Wang
ICASSP3