Yuming Shang

dblp:254/1955 · also Yu-Ming Shang · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 19 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 TruthfulRAG: Resolving Factual-level Conflicts in Retrieval-Augmented Generation with Knowledge Graphs
abstract
Retrieval-Augmented Generation (RAG) has emerged as a powerful framework for enhancing the capabilities of Large Language Models (LLMs) by integrating retrieval-based methods with generative models. As external knowledge repositories continue to expand and the parametric knowledge within models becomes outdated, a critical challenge for RAG systems is resolving conflicts between retrieved external information and LLMs' internal knowledge, which can significantly compromise the accuracy and reliability of generated content. However, existing approaches to conflict resolution typically operate at the token or semantic level, often leading to fragmented and partial understanding of factual discrepancies between LLMs' knowledge and context, particularly in knowledge-intensive tasks. To address this limitation, we propose TruthfulRAG, the first framework that leverages Knowledge Graphs (KGs) to resolve factual-level knowledge conflicts in RAG systems. Specifically, TruthfulRAG constructs KGs by systematically extracting triples from retrieved content, utilizes query-based graph retrieval to identify relevant knowledge, and employs entropy-based filtering mechanisms to precisely locate conflicting elements and mitigate factual inconsistencies, thereby enabling LLMs to generate faithful and accurate responses. Extensive experiments reveal that TruthfulRAG outperforms existing methods, effectively alleviating knowledge conflicts and improving the robustness and trustworthiness of RAG systems.
Shuyi Liu, Yuming Shang, Xi Zhang 0008
AAAI2
2026 When Reasoning Leaks Membership: Membership Inference Attack on Black-box Large Reasoning Models
abstract
Large Reasoning Models (LRMs) have rapidly gained prominence for their strong performance in solving complex tasks. Many modern black-box LRMs expose the intermediate reasoning traces through APIs to improve transparency (e.g., Gemini-2.5 and Claude-sonnet). Despite their benefits, we find that these traces can leak membership signals, creating a new privacy threat even without access to token logits used in prior attacks. In this work, we initiate the first systematic exploration of Membership Inference Attacks (MIAs) on black-box LRMs. Our preliminary analysis shows that LRMs produce confident, recall-like reasoning traces on familiar training member samples but more hesitant, inference-like reasoning traces on non-members. The representations of these traces are continuously distributed in the semantic latent space, spanning from familiar to unfamiliar samples. Building on this observation, we propose BlackSpectrum, the first membership inference attack framework targeting the black-box LRMs. The key idea is to construct a recall–inference axis in the semantic latent space, based on representations derived from the exposed traces. By locating where a query sample falls along this axis, the attacker can obtain a membership score and predict how likely it is to be a member of the training data. Additionally, to address the limitations of outdated datasets unsuited to modern LRMs, we provide two new datasets to support future research, arXivReasoning and BookReasoning. Empirically, exposing reasoning traces greatly increases the vulnerability of LRMs to MIAs, boosting attack accuracy by up to 23.8%, AUC by 29.9%, and nearly doubling TPR@5%FPR. Our findings highlight the need for LRM companies to balance transparency in intermediate reasoning traces with privacy preservation.
Ruihan Hu, Yuming Shang, Wei Luo 0016, Xi Zhang 0008
WWW2
2026 Fusing generation task data to enhance generalization of LLMs in zero-shot relational triplet extraction
Xingpeng Si, Yuhao Ye, Yuming Shang, Mengyuan Cao, Yu Bai 0018, Jiawei Li 0020, Yang Gao 0016
Expert Syst. Appl.3
2026 Discovering new intents via spatio-temporal pseudo-label denoising
Yuming Shang, Wei Huang 0013, Sanchuan Guo, Jinhu Chen, Xi Zhang 0008, Philip S. Yu
Inf. Process. Manag.2
2026 LoTTA: Low-rank test-time adaptation for unsupervised tabular anomaly detection
Yuming Shang, Wenzhi Peng, Wei Huang 0013, Ninglun Gu, Kailai Zhang
Inf. Process. Manag.1
2026 Enhancing text representation with frequency-domain features for robust fake news detection
Zetao Fei, Yuming Shang, Rouxi Wang, Yong Ma 0004
Knowl. Based Syst.2
2026 VideoEvent: Hierarchical and adaptive event modelling for complex video understanding
Xin Sun 0029, Jianfei Zhao, Yuming Shang
Knowl. Based Syst.4
2025 VEEF-Multi-LLM: Effective Vocabulary Expansion and Parameter Efficient Finetuning Towards Multilingual Large Language Models
abstract
Large Language Models(LLMs) have brought significant transformations to various aspects of human life and productivity. However, the heavy reliance on vast amounts of data in developing these models has resulted in a notable disadvantage for low-resource languages, such as Nuosu and others, which lack large datasets. Moreover, many LLMs exhibit significant performance discrepancies between high-and lowresource languages, thereby restricting equitable access to technological advances for all linguistic communities. To address these challenges, this paper propose a low-resource multilingual large language model, termed VEEF-Multi-LLM, constructed through effective vocabulary expansion and parameter-efficient fine-tuning. We introduce a series of innovative methods to address challenges in low-resource languages. First, we adopt Byte-level Byte-Pair Encoding to expand the vocabulary for broader multilingual support. We separate input and output embedding weights to boost performance, and apply RoPE for long-context handling, as well as RMSNorm for efficient training. To generate high-quality supervised fine-tuning (SFT) data, we use self-training and selective translation, and refine the resulting dataset with the assistance of native speakers to ensure cultural and linguistic accuracy. Our model, VEEF-Multi-LLM-8B, is trained on 600 billion tokens across 50 natural and 16 programming languages. Experimental results show that the model excels in multilingual instruction-following tasks, particularly in translation, outperforming competing models in benchmarks such as XCOPA and XStoryCloze. Although it lags slightly behind English-centric models in some tasks (e.g., m-MMLU), it prioritizes safety, reliability, and inclusivity, making it valuable for diverse linguistic communities. We open-source our models on GitHub and Huggingface.
Jiu Sha, Mengxiao Zhu 0004, Chong Feng 0001, Yuming Shang
COLING4
2025 SCCD: A Session-based Dataset for Chinese Cyberbullying Detection
abstract
The rampant spread of cyberbullying content poses a growing threat to societal well-being. However, research on cyberbullying detection in Chinese remains underdeveloped, primarily due to the lack of comprehensive and reliable datasets. Notably, no existing Chinese dataset is specifically tailored for cyberbullying detection. Moreover, while comments play a crucial role within sessions, current session-based datasets often lack detailed, fine-grained annotations at the comment level. To address these limitations, we present a novel Chinese cyberbullying dataset, termed SCCD, which consists of 677 session-level samples sourced from a major social media platform Weibo. Moreover, each comment within the sessions is annotated with fine-grained labels rather than conventional binary class labels. Empirically, we evaluate the performance of various baseline methods on SCCD, highlighting the challenges for effective Chinese cyberbullying detection.
Qingpo Yang, Yakai Chen, Zihui Xu, Yuming Shang, Sanchuan Guo, Xi Zhang 0008
COLING4
2025 Automated Detection of Pre-training Text in Black-box LLMs
abstract
Detecting whether a given text is a member in the pre-training data of Large Language Models (LLMs) is crucial for ensuring data privacy and copyright protection. Most existing methods rely on the LLM's hidden information (e.g., model parameters or token probabilities), making them ineffective in the black-box setting, where only input and output texts are accessible. Although some methods have been proposed for the black-box setting, they rely on massive manual efforts such as designing complicated questions or instructions. To address these issues, we propose VeilProbe, the first framework for automatically detecting LLMs' pre-training texts in a black-box setting without human intervention. VeilProbe utilizes a sequence-to-sequence mapping model to infer the latent mapping feature between the input text and the corresponding output suffix generated by the LLM. Then it performs the key token perturbations to obtain more distinguishable membership features. Additionally, considering real-world scenarios where the ground-truth training text samples are limited, a prototype-based membership classifier is introduced to alleviate the overfitting issue. Extensive evaluations on three widely used datasets demonstrate that our framework is effective and superior in the black-box setting.
Ruihan Hu, Yuming Shang, Jiankun Peng, Wei Luo 0016, Yazhe Wang, Xi Zhang 0008
IJCAI2
2025 Bayesian Network-Based Adaptive Prompt Learning for Emotion-Cause Pair Extraction
Hongyan Xie, Yuming Shang
NLPCC (3)2
2025 JailBench: A Comprehensive Chinese Security Assessment Benchmark for Large Language Models
Shuyi Liu, Simiao Cui, Haoran Bu, Yuming Shang, Xi Zhang 0008
PAKDD (5)4
2025 From local to global: Leveraging document graph for named entity recognition
Yuming Shang, Hongli Mao, Heyan Huang, Xianling Mao
Knowl. Based Syst.1
2025 DynImpt: A Dynamic Data Selection Method for Improving Model Training Efficiency
abstract
Selecting key data subsets for model training is an effective way to improve training efficiency. Existing methods generally utilize a well-trained model to evaluate samples and select crucial subsets, ignoring the fact that the sample importance changes dynamically during model training, resulting in the selected subset only being critical in a specific training epoch rather than a changing training phase. To address this issue, we attempt to evaluate the significant changes in sample importance during dynamic training and propose a novel data selection method to improve model training efficiency. Specifically, the temporal changes in sample importance are considered from three perspectives: (i) loss, the difference between the predicted labels and the true labels of samples in the current training epoch; (ii) instability, the dispersion of sample importance in the recent training phase; and (iii) inconsistency, the comparison of the changing trend in the importance of an individual sample relative to the average importance of all samples in the recent training phase. Extensive experiments demonstrate that dynamic data selection can reduce computational costs and improve model training efficiency. Additionally, we find that the difficulty level of the training task influences the data selection strategy.
Wei Huang 0013, Shangmin Guo, Yuming Shang, Xiangling Fu
IEEE Trans. Knowl. Data Eng.4
2025 Toward Balanced Denoising: Building a Structural and Textual Denoiser for Table Understanding
abstract
Recently, large language models (LLMs) have made remarkable progress in table understanding, yet they remain vulnerable to the structural noise (SN) and the textual noise (TN). Existing methods usually employ biased denoising strategies such as structural matching and textual filtering, or overzealous denoising strategies such as introducing supplementary tasks like text-to-SQL and table-to-text to reduce these two types of noise. However, these methods either neglect one type of noise or introduce substantial external noise. Therefore, how to simultaneously mitigate the structural and textual noise without introducing extra noise and improve the performance of LLMs in table understanding is still an unresolved issue. In this paper, we rethink the bottlenecks in table understanding from the perspective of noise reduction and propose a novel dual-denoiser-reasoner model, called TabDDR, for balanced and effective denoising. Specially, our model consists of a structural-and-textual denoiser and a task-adaptive reasoner. The former removes two types of noise via triplet alignment and planning extraction to seek an interpretable balance between breaking structural barriers and preserving structural characteristics, eliminating textual noise and retaining maximal information; the latter ensures a simple but effective reasoning process which can adapt to various downstream tasks. To highlight the presence and impact of the structural and textual noise, we construct the WTQ-SN and WTQ-TN datasets based on the WikiTableQuestion (WTQ) dataset. Extensive experiments on these self-constructed datasets and two other public datasets demonstrate that our proposed method performs better than state-of-the-art baselines.
Shu-Xun Yang, Xianling Mao, Yuming Shang, Heyan Huang
IEEE Trans. Knowl. Data Eng.3
2024 Span Graph Transformer for Document-Level Named Entity Recognition
abstract
Named Entity Recognition (NER), which aims to identify the span and category of entities within text, is a fundamental task in natural language processing. Recent NER approaches have featured pre-trained transformer-based models (e.g., BERT) as a crucial encoding component to achieve state-of-the-art performance. However, due to the length limit for input text, these models typically consider text at the sentence-level and cannot capture the long-range contextual dependency within a document. To address this issue, we propose a novel Span Graph Transformer (SGT) method for document-level NER, which constructs long-range contextual dependencies at both the token and span levels. Specifically, we first retrieve relevant contextual sentences in the document for each target sentence, and jointly encode them by BERT to capture token-level dependencies. Then, our proposed model extracts candidate spans from each sentence and integrates these spans into a document-level span graph, where nested spans within sentences and identical spans across sentences are connected. By leveraging the power of Graph Transformer and well-designed position encoding, our span graph can fully exploit span-level dependencies within the document. Extensive experiments on both resource-rich nested and flat NER datasets, as well as low-resource distantly supervised NER datasets, demonstrate that proposed SGT model achieves better performance than previous state-of-the-art models.
Hongli Mao, Xianling Mao, Hanlin Tang 0001, Yuming Shang, Heyan Huang
AAAI4
2024 Span-based Unified Named Entity Recognition Framework via Contrastive Learning
Hongli Mao, Xianling Mao, Hanlin Tang 0001, Yuming Shang, Xiaoyan Gao 0001, Ao-Jie Ma, Heyan Huang
IJCAI4
2024 A novel prompt-tuning method: Incorporating scenario-specific concepts into a verbalizer
Yong Ma 0004, Senlin Luo, Yuming Shang, Zhengjun Li
Expert Syst. Appl.3
2024 HCUKE: A Hierarchical Context-aware approach for Unsupervised Keyphrase Extraction
Xianling Mao, Cheng-Xin Xin, Yuming Shang, Tian-Yi Che, Hongli Mao, Heyan Huang
Knowl. Based Syst.4
2024 Don't Be Misled by Emotion! Disentangle Emotions and Semantics for Cross-Language and Cross-Domain Rumor Detection
abstract
Cross-language and cross-domain rumor detection is a crucial research topic for maintaining a healthy social media environment. Previous studies reveal that the emotions expressed in posts are important features for rumor detection. However, existing studies typically leverage the entangled representation of semantics and emotions, ignoring the fact that different languages and domains have different emotions toward rumors. Therefore, it inevitably leads to a biased adaptation of the features learned from the source to the target language and domain. To address this issue, this paper proposes a novel approach to adapt the knowledge obtained from the source to the target dataset by disentangling the emotional and semantic features of the datasets. Specifically, the proposed method mainly consists of three steps: (1) disentanglement, which encodes rumors into two separate semantic and emotional spaces to prevent emotional interference; (2) adaptation, merging semantics with the emotions from another language and domain for contrastive alignment to ensure effective adaptation; (3) joint training strategy, which enables the above two steps to work in synergy and mutually promote each other. Extensive experimental results demonstrate that the proposed method outperforms state-of-the-art baselines.
Xi Zhang 0008, Yuming Shang
IEEE Trans. Big Data3
2023 Clean-label Poisoning Attack against Fake News Detection Models
abstract
Researching data poisoning attacks against fake news detection models is crucial for bolstering their robustness and curbing the dissemination of fake news. Existing textual data poisoning attacks necessitate control over both the content and labels of news samples, making them impractical for real attack scenarios. In this paper, we propose COMCP, a novel clean-label poisoning attack model aimed at fake news detection models. Diverging from existing methods, COMCP ensures the poison samples are accurately labeled, while crafting stealthy poison comments without modifying the headlines or content, thereby enhancing the feasibility of the attack. Furthermore, COMCP generates poison comments by appending stealthy characters to ensure the stealthiness of the attack. Comprehensive experimental evaluations on three benchmark datasets illustrate that our proposal outperforms SOTA baselines in terms of attack success rate and text quality, while maintaining the accuracy of detecting clean samples.
Jiayi Liang, Xi Zhang 0008, Yuming Shang, Sanchuan Guo, Chaozhuo Li
IEEE Big Data3
2023 Enhanced CGSN System for Machine Reading Comprehension
Liwen Zheng, Hongyan Xie, Xi Zhang 0008, Yuming Shang
NLPCC (3)5
2023 Learning Relation Ties with a Force-Directed Graph in Distant Supervised Relation Extraction
abstract
Relation ties, defined as the correlation and mutual exclusion between different relations, are critical for distant supervised relation extraction. Previous studies usually obtain this property by greedily learning the local connections between relations. However, they are essentially limited because of failing to capture the global topology structure of relation ties and may easily fall into a locally optimal solution. To address this issue, we propose a novel force-directed graph to comprehensively learn relation ties. Specifically, we first construct a graph according to the global co-occurrence of all relations. Then, we borrow the idea of Coulomb’s law from physics and introduce the concept of attractive force and repulsive force into this graph to learn correlation and mutual exclusion between relations. Finally, the obtained relation representations are applied as an inter-dependent relation classifier. Extensive experimental results demonstrate that our method is capable of modeling global correlation and mutual exclusion between relations, and outperforms the state-of-the-art baselines. In addition, the proposed force-directed graph can be used as a module to augment existing relation extraction systems and improve their performance.
Yuming Shang, Heyan Huang, Xin Sun 0029, Wei Wei 0002, Xianling Mao
ACM Trans. Inf. Syst.1
2022 OneRel: Joint Entity and Relation Extraction with One Module in One Step
abstract
Joint entity and relation extraction is an essential task in natural language processing and knowledge graph construction. Existing approaches usually decompose the joint extraction task into several basic modules or processing steps to make it easy to conduct. However, such a paradigm ignores the fact that the three elements of a triple are interdependent and indivisible. Therefore, previous joint methods suffer from the problems of cascading errors and redundant information. To address these issues, in this paper, we propose a novel joint entity and relation extraction model, named OneRel, which casts joint extraction as a fine-grained triple classification problem. Specifically, our model consists of a scoring-based classifier and a relation-specific horns tagging strategy. The former evaluates whether a token pair and a relation belong to a factual triple. The latter ensures a simple but effective decoding process. Extensive experimental results on two widely used datasets demonstrate that the proposed method performs better than the state-of-the-art baselines, and delivers consistent performance gain on complex scenarios of various overlapping patterns and multiple triples.
Yuming Shang, Heyan Huang, Xianling Mao
AAAI1
2022 Relational Triple Extraction: One Step is Enough
abstract
Extracting relational triples from unstructured text is an essential task in natural language processing and knowledge graph construction. Existing approaches usually contain two fundamental steps: (1) finding the boundary positions of head and tail entities; (2) concatenating specific tokens to form triples. However, nearly all previous methods suffer from the problem of error accumulation, i.e., the boundary recognition error of each entity in step (1) will be accumulated into the final combined triples. To solve the problem, in this paper, we introduce a fresh perspective to revisit the triple extraction task and propose a simple but effective model, named DirectRel. Specifically, the proposed model first generates candidate entities through enumerating token sequences in a sentence, and then transforms the triple extraction task into a linking problem on a ``head -> tail" bipartite graph. By doing so, all triples can be directly extracted in only one step. Extensive experimental results on two widely used datasets demonstrate that the proposed model performs better than the state-of-the-art baselines.
Yuming Shang, Heyan Huang, Xin Sun 0029, Wei Wei 0002, Xianling Mao
IJCAI1
2022 A pattern-aware self-attention network for distant supervised relation extraction
Yuming Shang, Heyan Huang, Xin Sun 0029, Wei Wei 0002, Xianling Mao
Inf. Sci.1
2022 Three birds, one stone: A novel translation based framework for joint entity and relation extraction
Heyan Huang, Yuming Shang, Xin Sun 0029, Wei Wei 0002, Xianling Mao
Knowl. Based Syst.2
2021 ESRE: handling repeated entities in distant supervised relation extraction
Xin Sun 0029, Jinghu Jiang, Yuming Shang
Neural Comput. Appl.3
2020 Are Noisy Sentences Useless for Distant Supervised Relation Extraction?
abstract
The noisy labeling problem has been one of the major obstacles for distant supervised relation extraction. Existing approaches usually consider that the noisy sentences are useless and will harm the model's performance. Therefore, they mainly alleviate this problem by reducing the influence of noisy sentences, such as applying bag-level selective attention or removing noisy sentences from sentence-bags. However, the underlying cause of the noisy labeling problem is not the lack of useful information, but the missing relation labels. Intuitively, if we can allocate credible labels for noisy sentences, they will be transformed into useful training data and benefit the model's performance. Thus, in this paper, we propose a novel method for distant supervised relation extraction, which employs unsupervised deep clustering to generate reliable labels for noisy sentences. Specifically, our model contains three modules: a sentence encoder, a noise detector and a label generator. The sentence encoder is used to obtain feature representations. The noise detector detects noisy sentences from sentence-bags, and the label generator produces high-confidence relation labels for noisy sentences. Extensive experimental results demonstrate that our model outperforms the state-of-the-art baselines on a popular benchmark dataset, and can indeed alleviate the noisy labeling problem.
Yuming Shang, Heyan Huang, Xianling Mao, Xin Sun 0029, Wei Wei 0002
AAAI1