Shumin Shi

dblp:78/10884 · DBLP profile ↗
← Back
23ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0003-3436-7575ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Do Retrieval Augmented Language Models Know When They Don't Know?
abstract
Existing large language models (LLMs) occasionally generate plausible yet factually incorrect responses, known as hallucinations. Two main approaches have been proposed to mitigate hallucinations: retrieval-augmented language models (RALMs) and refusal post-training. However, current research predominantly focuses on their individual effectiveness while overlooking the evaluation of the refusal capability of RALMs. Ideally, if RALMs know when they do not know, they should refuse to answer. In this study, we ask the fundamental question: Do RALMs know when they don’t know? Specifically, we investigate three questions. First, are RALMs well calibrated with respect to different internal and external knowledge states? We examine the influence of various factors. Contrary to expectations, when all retrieved documents are irrelevant, RALMs still tend to refuse questions they could have answered correctly. Next, given the model's pronounced over-refusal behavior, we raise a second question: How does a RALM's refusal ability align with its calibration quality? Our results show that the over-refusal problem can be mitigated through in-context fine-tuning. However, we observe that improved refusal behavior does not necessarily imply better calibration or higher overall accuracy. Finally, we ask: Can we combine refusal-aware RALMs with uncertainty-based answer abstention to mitigate over-refusal? We develop a simple yet effective refusal mechanism for refusal-post-trained RALMs that improves their overall answer quality by balancing refusal and correct answers. Our study provides a more comprehensive understanding of the factors influencing RALM behavior. Meanwhile, we emphasize that uncertainty estimation for RALMs remains an open problem deserving deeper investigation.
Youchao Zhou, Heyan Huang, Xinglin Wang, Shumin Shi, Yang Deng 0002
AAAI7
2026 Would LLMs be Good Historical Linguists and Chinese Dialect Learners?
abstract
Large language models (LLMs) perform well on Standard Chinese but struggle with lowresource Chinese dialects due to substantial phonological divergence.We investigate whether incorporating Middle Chinese, the common historical ancestor of most of the modern Chinese dialects, can improve dialectal pronunciation modeling in a linguistically interpretable manner.We focus on two specific task variants: (1) conditional sound change rule induction (a variant of Sound Law Induction, SLI), where models infer executable phonological transformation rules from Middle Chinese to modern dialects, and (2) sentence-level dialectal pronunciation transcription (a variant of Grapheme-to-Phoneme, G2P), requiring dialect-specific International Phonetic Alphabet (IPA) generation.We construct a multi-source dataset covering Middle Chinese and 12 modern Chinese dialects, including character-level correspondences, rule exemplars, and sentencelevel IPA transcription.We adopt a parameterefficient training framework combining LoRAbased supervised fine-tuning and reinforcement learning via Group Relative Policy Optimization (GRPO) for the first task.Across both tasks and a wide range of dialects and evaluation metrics, our approach achieves overall improvements over strong baselines, including DeepSeek-V3.2 and ChatGPT-5.2,while revealing variation across dialects.These results demonstrate the value of leveraging historical linguistic knowledge for modeling low-resource Chinese dialects.
Shumin Shi, Youchao Zhou
ACL (1)2
2026 Selective and contrastive mechanism for distantly supervised relation extraction
Danjie Han, Heyan Huang, Shumin Shi, Cunhan Guo, Yanghao Zhou, Changsen Yuan
Neurocomputing3
2025 CARE: Contextual Augmentation with Retrieval Enhancement for Relation Extraction in Large Language Models
Danjie Han, Heyan Huang, Shumin Shi, Cunhan Guo, Yanghao Zhou, Changsen Yuan
NLPCC (1)3
2025 Distantly Supervised relation extraction with multi-level contextual information integration
Danjie Han, Heyan Huang, Shumin Shi, Changsen Yuan, Cunhan Guo
Neurocomputing3
2024 OSTOD: One-Step Task-Oriented Dialogue with activated state and retelling response
Heyan Huang, Puhai Yang, Wei Wei 0002, Shumin Shi, Xianling Mao
Knowl. Based Syst.4
2024 Alleviating repetitive tokens in non-autoregressive machine translation with unlikelihood training
Shuheng Wang, Shumin Shi, Heyan Huang
Soft Comput.2
2024 STN4DST: A Scalable Dialogue State Tracking Based on Slot Tagging Navigation
abstract
Dialogue state tracking plays a key role in tracking user intentions in task-oriented dialogue systems. Traditional dialogue state tracking methods usually rely on selecting slot values from a fixed ontology to represent the dialogue state. In recent years, more flexible open vocabulary based approaches have become the mainstream focus which are mainly divided into two categories: generative methods and span extraction methods. Among them, the span extraction method is favored for its outstanding ability to predict unknown slot values. However, the span extraction method only focuses on the predicted slot values, but ignores other potential slot values in the utterance, which leads to insufficient semantic understanding of the utterance and difficulty in dealing with complex utterance scenarios, such as more or longer unknown slot values. To tackle the above drawbacks, in this paper, we propose a novel scalable dialogue state tracking method, which employs slot tagging to locate all potential slot values in the utterances and jointly learns slot pointers to select the predicted slot value from them. Specifically, our STN4DST (Slot Tagging Navigation for Dialogue State Tracking) model not only adopts the above joint learning strategy, which we call slot tagging navigation, to extract slot values from utterances, but also uses previous dialogue states as dialogue contexts to track the change of slot values, and introduces appendix slot values to predict special slot values that cannot be extracted. Extensive experiments show that in the open vocabulary setting, STN4DST achieves the state-of-the-art joint goal accuracy of 85.4% and 96.5% on Sim-M and Sim-R datasets with a large number of unknown slot values, and is also comparable to other state-of-the-art models in the absence of token-level slot annotations for all potential slot values.
Puhai Yang, Heyan Huang, Shumin Shi, Xianling Mao
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Interest Aware Dual-Channel Graph Contrastive Learning for Session-Based Recommendation
Sichen Liu 0003, Shumin Shi
NLPCC (1)2
2023 Incorporating history and future into non-autoregressive machine translation
Shuheng Wang, Heyan Huang, Shumin Shi
Comput. Speech Lang.3
2023 Better Localness for Non-Autoregressive Transformer
abstract
The Non-Autoregressive Transformer, due to its low inference latency, has attracted much attention from researchers. Although, the performance of the non-autoregressive transformer has been significantly improved in recent years, there is still a gap between the non-autoregressive transformer and the autoregressive transformer. Considering the success of localness on the autoregressive transformer, in this work, we consider incorporating localness into the non-autoregressive transformer. Specifically, we design a dynamic mask matrix according to the query tokens, key tokens, and relative distance, and unify the localness module for self-attention and cross-attention module. We conduct experiments on several benchmark tasks, and the results show that our model can significantly improve the performance of the non-autoregressive transformer.
Shuheng Wang, Heyan Huang, Shumin Shi
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2022 Low-resource Neural Machine Translation: Methods and Trends
abstract
Neural Machine Translation (NMT) brings promising improvements in translation quality, but until recently, these models rely on large-scale parallel corpora. As such corpora only exist on a handful of language pairs, the translation performance is far from the desired effect in the majority of low-resource languages. Thus, developing low-resource language translation techniques is crucial and it has become a popular research field in neural machine translation. In this article, we make an overall review of existing deep learning techniques in low-resource NMT. We first show the research status as well as some widely used low-resource datasets. Then, we categorize the existing methods and show some representative works detailedly. Finally, we summarize the common characters among them and outline the future directions in this field.
Shumin Shi, Rihai Su, Heyan Huang
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2022 Improving Neural Machine Translation by Transferring Knowledge from Syntactic Constituent Alignment Learning
abstract
Statistical machine translation (SMT) models rely on word-, phrase-, and syntax-level alignments. But neural machine translation (NMT) models rarely explicitly learn the phrase- and syntax-level alignments. In this article, we propose to improve NMT by explicitly learning the bilingual syntactic constituent alignments. Specifically, we first utilize syntactic parsers to induce syntactic structures of sentences, and then we propose two ways to utilize the syntactic constituents in a perceptual (not adversarial) generator-discriminator training framework. One way is to use them to measure the alignment score of sentence-level training examples, and the other is to directly score the alignments of constituent-level examples generated with an algorithm based on word-level alignments from SMT. In our generator-discriminator framework, the discriminator is pre-trained to learn constituent alignments and distinguish the ground-truth translation from the fake ones, while the generative translation model is fine-tuned to receive the alignment knowledge and to generate translations that best approximate the true ones. Experiments and analysis show that the learned constituent alignments can help improve the translation results.
Chao Su 0002, Heyan Huang, Shumin Shi, Ping Jian
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2021 Improving Non-autoregressive Machine Translation with Soft-Masking
Shuheng Wang, Shumin Shi, Heyan Huang
NLPCC (1)2
2021 Enhanced encoder for non-autoregressive machine translation
Shuheng Wang, Shumin Shi, Heyan Huang
Mach. Transl.2
2021 Multi-level Chunk-based Constituent-to-Dependency Treebank Transformation for Tibetan Dependency Parsing
abstract
Dependency parsing is an important task for Natural Language Processing (NLP). However, a mature parser requires a large treebank for training, which is still extremely costly to create. Tibetan is a kind of extremely low-resource language for NLP, there is no available Tibetan dependency treebank, which is currently obtained by manual annotation. Furthermore, there are few related kinds of research on the construction of treebank. We propose a novel method of multi-level chunk-based syntactic parsing to complete constituent-to-dependency treebank conversion for Tibetan under scarce conditions. Our method mines more dependencies of Tibetan sentences, builds a high-quality Tibetan dependency tree corpus, and makes fuller use of the inherent laws of the language itself. We train the dependency parsing models on the dependency treebank obtained by the preliminary transformation. The model achieves 86.5% accuracy, 96% LAS, and 97.85% UAS, which exceeds the optimal results of existing conversion methods. The experimental results show that our method has the potential to use a low-resource setting, which means we not only solve the problem of scarce Tibetan dependency treebank but also avoid needless manual annotation. The method embodies the regularity of strong knowledge-guided linguistic analysis methods, which is of great significance to promote the research of Tibetan information processing.
Shumin Shi, Congjun Long, Heyan Huang
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2020 Dynamic Attention Aggregation with BERT for Neural Machine Translation
abstract
The recently proposed BERT has demonstrated great power in various natural language processing tasks. However, the model does not perform effectively on cross-lingual tasks, especially on machine translation. In this work, we propose three methods to introduce pre-trained BERT into neural machine translation without fine-tuning. Our approach consists of a) a linear-attention aggregation that leverages a parameter matrix to capture the key knowledge of BERT, b) a self-attention aggregation which aims to learn what is vital for input and output, and c) a switch-gate aggregation to dynamically control the balance of the information flowing from the pre-trained BERT or the NMT model. We conduct experiments on several translation benchmarks and substantially improve over 2 BELU points on the IWSLT'14 English - German task with switch-gate aggregation method compared to a strong baseline, while our proposed model also performs remarkably on the other tasks.
Jiarui Zhang 0003, Hongzheng Li, Shumin Shi, Heyan Huang, Yue Hu 0002, Xiangpeng Wei
IJCNN3
2020 Neural machine translation with Gumbel Tree-LSTM based encoder
Chao Su 0002, Heyan Huang, Shumin Shi, Ping Jian, Xuewen Shi 0001
J. Vis. Commun. Image Represent.3
2019 Combining External Sentiment Knowledge for Emotion Cause Detection
Jiaxing Hu, Shumin Shi, Heyan Huang
NLPCC (1)2
2019 Mapping sentences to concept transferred space for semantic textual similarity
Heyan Huang, Hao Wu 0066, Xiaochi Wei, Yang Gao 0016, Shumin Shi
Knowl. Inf. Syst.5
2017 A Parallel Recurrent Neural Network for Language Modeling with POS Tags
Chao Su 0002, Heyan Huang, Shumin Shi, Yuhang Guo 0001, Hao Wu 0066
PACLIC3
2017 Incorporating target language semantic roles into a string-to-tree translation model
abstract
The string-to-tree model is one of the most successful syntax-based statistical machine translation (SMT) models. It models the grammaticality of the output via target-side syntax. However, it does not use any semantic information and tends to produce translations containing semantic role confusions and error chunk sequences. In this paper, we propose two methods to use semantic roles to improve the performance of the string-to-tree translation model: (1) adding role labels in the syntax tree; (2) constructing a semantic role tree, and then incorporating the syntax information into it. We then perform string-to-tree machine translation using the newly generated trees. Our methods enable the system to train and choose better translation rules using semantic information. Our experiments showed significant improvements over the state-of-the-art string-to-tree translation system on both spoken and news corpora, and the two proposed methods surpass the phrase-based system on large-scale training data.
Chao Su 0002, Yuhang Guo 0001, Heyan Huang, Shumin Shi, Chong Feng 0001
Frontiers Inf. Technol. Electron. Eng.4
2014 A Method of Polarity Computation of Chinese Sentiment Words Based on Gaussian Distribution
Ruijing Li, Shumin Shi, Heyan Huang, Chao Su 0002, Tianhang Wang
CICLing (2)2