Bo Li 0099

dblp:50/3402-99 · DBLP profile ↗
← Back
15ranked-venue papers
11as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 10 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Modeling Uncertainty Trends for Timely Retrieval in Dynamic RAG
abstract
Dynamic retrieval-augmented generation (RAG) allows large language models (LLMs) to fetch external knowledge on demand, offering greater adaptability than static RAG. A central challenge in this setting lies in determining the optimal timing for retrieval. Existing methods often trigger retrieval based on low token-level confidence, which may lead to delayed intervention after errors have already propagated. We introduce Entropy-Trend Constraint (ETC), a training-free method that determines optimal retrieval timing by modeling the dynamics of token-level uncertainty. Specifically, ETC utilizes first- and second-order differences of the entropy sequence to detect emerging uncertainty trends, enabling earlier and more precise retrieval. Experiments on six QA benchmarks with three LLM backbones demonstrate that ETC consistently outperforms strong baselines while reducing retrieval frequency. ETC is particularly effective in domain-specific scenarios, exhibiting robust generalization capabilities. Ablation studies and qualitative analyses further confirm that trend-aware uncertainty modeling yields more effective retrieval timing. The method is plug-and-play, model-agnostic, and readily integrable into existing decoding pipelines. Implementation code is included in the supplementary materials.
Bo Li 0099, Zhenghua Xu 0001, Shikun Zhang, Wei Ye 0004
AAAI1
2026 Language Drift in Multilingual Retrieval-Augmented Generation: Characterization and Decoding-Time Mitigation
abstract
Multilingual Retrieval-Augmented Generation (RAG) enables large language models (LLMs) to perform knowledge-intensive tasks in multilingual settings by leveraging retrieved documents as external evidence. However, when the retrieved evidence differs in language from the user query and in-context exemplars, the model often exhibits language drift by generating responses in an unintended language. This phenomenon is especially pronounced during reasoning-intensive decoding, such as Chain-of-Thought (CoT) generation, where intermediate steps introduce further language instability. In this paper, we systematically study output language drift in multilingual RAG across multiple datasets, languages, and LLM backbones. Our controlled experiments reveal that the drift results not from comprehension failure but from decoder-level collapse, where dominant token distributions and high-frequency English patterns dominate the intended generation language. We further observe that English serves as a semantic attractor under cross-lingual conditions, emerging as both the strongest interference source and the most frequent fallback language. To mitigate this, we propose Soft Constrained Decoding (SCD), a lightweight, training-free decoding strategy that gently steers generation toward the target language by penalizing non-target-language tokens. SCD is model-agnostic and can be applied to any generation algorithm without modifying the architecture or requiring additional data. Experiments across three multilingual datasets and multiple typologically diverse languages show that SCD consistently improves language alignment and task performance, providing an effective and generalizable solution in multilingual RAG.
Bo Li 0099, Zhenghua Xu 0001, Rui Xie 0003
AAAI1
2026 Retrieval as Generation: A Unified Framework with Self-Triggered Information Planning
abstract
We revisit retrieval-augmented generation (RAG) by embedding retrieval control directly into generation.Instead of treating retrieval as an external intervention, we express retrieval decisions within token-level decoding, enabling end-to-end coordination without additional controllers or classifiers.Under the paradigm of Retrieval as Generation, we propose GRIP (Generation-guided Retrieval with Information Planning), a unified framework in which the model regulates retrieval behavior through control-token emission.Central to GRIP is Self-Triggered Information Planning, which allows the model to decide when to retrieve, how to reformulate queries, and when to terminate, all within a single autoregressive trajectory.This design tightly couples retrieval and reasoning and supports dynamic multi-step inference with on-the-fly evidence integration.To supervise these behaviors, we construct a structured training set covering answerable, partially answerable, and multi-hop queries, each aligned with specific token patterns.Experiments on five QA benchmarks show that GRIP surpasses strong RAG baselines and is competitive with GPT-4o while using substantially fewer parameters.
Bo Li 0099, Gexiang Fang, Shikun Zhang, Wei Ye 0004
ACL (1)1
2026 Instruction Data Selection via Answer Divergence
abstract
Instruction tuning relies on large instruction-response corpora whose quality and composition strongly affect downstream performance.We propose Answer Divergence-Guided Selection (ADG), which selects instruction data based on the geometric structure of multi-sample outputs.ADG draws several high-temperature generations per instruction, maps responses into an embedding space, and computes an output divergence score that jointly encodes dispersion magnitude and shape anisotropy.High scores correspond to instructions whose answers are both far apart and multi-modal, rather than clustered paraphrases along a single direction.Across two backbones and three public instruction pools, finetuning on only 10K ADG-selected examples consistently outperforms strong selectors on six benchmarks spanning reasoning, knowledge, and coding.Analyses further show that both dispersion magnitude and shape anisotropy are necessary, supporting answer divergence as a practical signal for instruction data selection.Code and appendix are included in the supplementary materials.
Bo Li 0099, Shikun Zhang, Wei Ye 0004
ACL (1)1
2026 Advancing federated semi-supervised medical image segmentation: A duo of interactive denoising pseudo-labels and convolutional contrastive learning
Zhenghua Xu 0001, Bo Li 0099, Gaoxi Zhou, Xianglin Lu, Thomas Lukasiewicz
Medical Image Anal.3
2026 You Need Glimpse Before Segmentation: Stochastic Detector-Actor-Critic for Medical Image Segmentation
abstract
Medical images often contain more redundant background areas than natural images, potentially introducing noise and degrading image segmentation performance. Inspired by doctors' diagnostic processes, where they identify the lesion area before conducting a detailed analysis, we introduce a novel Stochastic Detector-Actor-Critic (SDAC) framework to tackle this challenge. SDAC initially glimpses the entire image using a detector network and policy gradient algorithms to filter out irrelevant background regions and focus on crucial, smaller areas for segmentation. The Actor-Critic algorithm then dynamically creates segmentation masks pixel by pixel without user intervention or coarse masks, forming a robust segmentation module. Both processes are trained jointly to reduce error propagation and ensure stability and ease of implementation. Our experiments on two commonly used medical image segmentation datasets demonstrate that SDAC achieves competitive results comparable to state-of-the-art methods while using 10x fewer parameters than the best-performing baseline in terms of DICE and IoU metrics. We also conduct detailed ablation studies to enhance understanding and facilitate practical use. Furthermore, SDAC performs well in low-resource settings (i.e., 50-shot or 100-shot), making it ideal for real-world scenarios. Its lightweight design make SDAC an excellent baseline for medical image segmentation tasks.
Zhenghua Xu 0001, Bo Li 0099, Weipeng Liu, Thomas Lukasiewicz
IEEE J. Biomed. Health Informatics4
2024 Labels Need Prompts Too: Mask Matching for Natural Language Understanding Tasks
abstract
Textual label names (descriptions) are typically semantically rich in many natural language understanding (NLU) tasks. In this paper, we incorporate the prompting methodology, which is widely used to enrich model input, into the label side for the first time. Specifically, we propose a Mask Matching method, which equips an input with a prompt and its label with another, and then makes predictions by matching their mask representations. We evaluate our method extensively on 8 NLU tasks with 14 datasets. The experimental results show that Mask Matching significantly outperforms its counterparts of fine-tuning and conventional prompt-tuning, setting up state-of-the-art performances in several datasets. Mask Matching is particularly good at handling NLU tasks with large label counts and informative label names. As pioneering efforts that investigate the label-side prompt, we also discuss open issues for future study.
Bo Li 0099, Wei Ye 0004, Quansen Wang, Shikun Zhang
AAAI1
2023 Reviewing Labels: Label Graph Network with Top-k Prediction Set for Relation Extraction
abstract
The typical way for relation extraction is fine-tuning large pre-trained language models on task-specific datasets, then selecting the label with the highest probability of the output distribution as the final prediction. However, the usage of the Top-k prediction set for a given sample is commonly overlooked. In this paper, we first reveal that the Top-k prediction set of a given sample contains useful information for predicting the correct label. To effectively utilizes the Top-k prediction set, we propose Label Graph Network with Top-k Prediction Set, termed as KLG. Specifically, for a given sample, we build a label graph to review candidate labels in the Top-k prediction set and learn the connections between them. We also design a dynamic k selection mechanism to learn more powerful and discriminative relation representation. Our experiments show that KLG achieves the best performances on three relation extraction datasets. Moreover, we observe thatKLG is more effective in dealing with long-tailed classes.
Bo Li 0099, Wei Ye 0004, Shikun Zhang
AAAI1
2023 Sequence Generation with Label Augmentation for Relation Extraction
abstract
Sequence generation demonstrates promising performance in recent information extraction efforts, by incorporating large-scale pre-trained Seq2Seq models. This paper investigates the merits of employing sequence generation in relation extraction, finding that with relation names or synonyms as generation targets, their textual semantics and the correlation (in terms of word sequence pattern) among them affect model performance. We then propose Relation Extraction with Label Augmentation (RELA), a Seq2Seq model with automatic label augmentation for RE. By saying label augmentation, we mean prod semantically synonyms for each relation name as the generation target. Besides, we present an in-depth analysis of the Seq2Seq model's behavior when dealing with RE. Experimental results show that RELA achieves competitive results compared with previous methods on four RE datasets.
Bo Li 0099, Dingyao Yu, Wei Ye 0004, Shikun Zhang
AAAI1
2021 Multi-view Inference for Relation Extraction with Uncertain Knowledge
abstract
Knowledge graphs (KGs) are widely used to facilitate relation extraction (RE) tasks. While most previous RE methods focus on leveraging deterministic KGs, uncertain KGs, which assign a confidence score for each relation instance, can provide prior probability distributions of relational facts as valuable external knowledge for RE models. This paper proposes to exploit uncertain knowledge to improve relation extraction. Specifically, we introduce ProBase, an uncertain KG that indicates to what extent a target entity belongs to a concept, into our RE architecture. We then design a novel multi-view inference framework to systematically integrate local context and global knowledge across three views: mention-, entity- and concept-view. The experiment results show that our model achieves competitive performances on both sentence- and document-level relation extraction, which verifies the effectiveness of introducing uncertain knowledge and the multi-view inference framework that we design.
Bo Li 0099, Wei Ye 0004, Canming Huang, Shikun Zhang
AAAI1
2021 Point, Disambiguate and Copy: Incorporating Bilingual Dictionaries for Neural Machine Translation
abstract
Tong Zhang, Long Zhang, Wei Ye, Bo Li, Jinan Sun, Xiaoyu Zhu, Wen Zhao, Shikun Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Tong Zhang 0001, Long Zhang 0012, Wei Ye 0004, Bo Li 0099, Jinan Sun, Shikun Zhang
ACL/IJCNLP (1)4
2020 Graph Enhanced Dual Attention Network for Document-Level Relation Extraction
abstract
Document-level relation extraction requires inter-sentence reasoning capabilities to capture local and global contextual information for multiple relational facts.To improve inter-sentence reasoning, we propose to characterize the complex interaction between sentences and potential relation instances via a Graph Enhanced Dual Attention network (GEDA).In GEDA, sentence representation generated by the sentence-to-relation (S2R) attention is refined and synthesized by a Heterogeneous Graph Convolutional Network before being fed into the relation-to-sentence (R2S) attention .We further design a simple yet effective regularizer based on the natural duality of the S2R and R2S attention, whose weights are also supervised by the supporting evidence of relation instances during training.An extensive set of experiments on an existing large-scale dataset show that our model achieves competitive performance, especially for the inter-sentence relation extraction, while the neural predictions can also be interpretable and easily observed.
Bo Li 0099, Wei Ye 0004, Zhonghao Sheng, Rui Xie 0003, Xiangyu Xi, Shikun Zhang
COLING1
2020 Sliding Hierarchical Recurrent Neural Networks for Sequence Classification
abstract
Hierarchical Recurrent Neural Networks (HRNN) is an important advance in improving efficiency and performance of sequence classification in recent years. The intuition behind this approach is to slice long sequences into many short sub-sequences and process them in parallel, then capturing the long-term dependencies between those sub-sequences by deeper layers of the networks. In this paper, we propose a novel architecture called Sliding Hierarchical Recurrent Neural Network (SHRNN). We introduce a new sliding mechanism on the input sequence of each layer, named recursive block, so that SHRNN can process the input sequence effectively. We also introduce layer-wise attention and multi-layer regularization for further improvements. We perform large-scale experiments in sequence classification task of both text and image on 8 datasets. As result, we not only achieve new start-of-the-art performance on all datasets by SHRNN, but also investigate effects of different components of SHRNN systematically and thoroughly, which provides best practice for the usage of SHRNN.
Bo Li 0099, Zhonghao Sheng, Wei Ye 0004, Shikun Zhang
IJCNN1
2019 Exploiting Entity BIO Tag Embeddings and Multi-task Learning for Relation Extraction with Imbalanced Data
abstract
In practical scenario, relation extraction needs to first identify entity pairs that have relation and then assign a correct relation class.However, the number of non-relation entity pairs in context (negative instances) usually far exceeds the others (positive instances), which negatively affects a model's performance.To mitigate this problem, we propose a multitask architecture which jointly trains a model to perform relation identification with crossentropy loss and relation classification with ranking loss.Meanwhile, we observe that a sentence may have multiple entities and relation mentions, and the patterns in which the entities appear in a sentence may contain useful semantic information that can be utilized to distinguish between positive and negative instances.Thus we further incorporate the embeddings of character-wise/word-wise BIO tag from the named entity recognition task into character/word embeddings to enrich the input representation.Experiment results show that our proposed approach can significantly improve the performance of a baseline model with more than 10% absolute increase in F1-score, and outperform the state-of-theart models on ACE 2005 Chinese and English corpus.Moreover, BIO tag embeddings are particularly effective and can be used to improve other models as well.* indicates equal contribution.
Wei Ye 0004, Bo Li 0099, Rui Xie 0003, Zhonghao Sheng, Shikun Zhang
ACL (1)2
2019 Long Text Analysis Using Sliced Recurrent Neural Networks with Breaking Point Information Enrichment
abstract
Sliced recurrent neural networks (SRNNs) are the state-of-the-art efficient solution for long text analysis tasks; however, their slicing operations inevitably result in long-term dependency loss in lower-level networks and thus limit their accuracy. Therefore, we propose a breaking point information enrichment mechanism to strengthen dependencies between sliced subsequences without hindering parallelization. Then, the resulting BPIE-SRNN model is further extended to a bidirectional model, BPIE-BiSRNN, to utilize the dependency information in not only the previous but also the following contexts. Experiments on four large public real-world datasets demonstrate that the BPIE-SRNN and BPIE-BiSRNN models always achieve a much better accuracy than SRNNs and BiSRNNs, while maintaining a superior training efficiency.
Bo Li 0099, Zehua Cheng, Zhenghua Xu 0001, Wei Ye 0004, Thomas Lukasiewicz, Shikun Zhang
ICASSP1