Xingnan Jin

dblp:311/0443 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Efficient Tuning Framework for Resource- Constrained Biomedical Question Answering
abstract
Automatic question-answering systems demonstrate valuable utility in the biomedical domain, improving the precision and efficiency of clinical decision-making significantly. Despite large-scale language models achieving notable success in general domains, even outperforming human-level performance in certain aspects, they are still faced with challenges such as data privacy and scarcity in the special domain. This study explores the method for efficient fine-tuning under resource-constrained conditions in the biomedical field. We propose a multi-stage fine-tuning approach that effectively improves the performance of pre-trained language models in biomedical question-answering tasks. Specially, A multi-prompt-based contrastive learning strategy and a multi-prompt self-consistency voting module are introduced, which improve the accuracy of QA tasks. The experiments on the PubMedQA dataset under reasoning-required settings indicate that our approach outperforms domain-specific pre-training models and achieves comparable performance with GPT-4, while the number of fine-tuned parameters is much less than the total parameters of the base model.
Yongping Du, Xingnan Jin, Zikai Wang 0009
IEEE Trans. Comput. Biol. Bioinform.3
2025 A Contextualized Semantic Alignment Framework for Biomedical QA with Controllable Evidence Utilization
abstract
Biomedical retrieval-augmented generation (RAG) remains challenged by fragmented context and weak alignment between retrieved evidence and generative reasoning, leading to incoherent or hallucinated outputs. We propose BioCoRE, a continuity-aware retrieval and evidence-aligned generation framework that enables precise, controllable, and parameter-free evidence utilization. By preserving positional continuity, modeling inter-sentence dependencies, and integrating explicit instructions with latent anchor embeddings, BioCoRE bridges the gap between noisy retrieval and coherent generation. Experiments on biomedical question-answering (QA) benchmarks indicate consistent improvements in factual accuracy over domain-specific finetuned models, general-purpose large language models (LLMs), and state-of-the-art RAG methods. Ablation and visualization analyses further demonstrate that BioCoRE more effectively prioritizes and integrates retrieved evidence, achieving substantial gains in evidence-grounded biomedical reasoning.
Xingnan Jin, Hejun Yang, Yongping Du
BIBM1
2025 FoVer: First-Order Logic Verification for Natural Language Reasoning
Yongping Du, Xingnan Jin
Trans. Assoc. Comput. Linguistics3
2024 A Novel RAG Framework with Knowledge-Enhancement for Biomedical Question Answering
abstract
The biomedical question-answering system usually provide accurate and real-time responses, which is crucial for clinical decision-making and scientific research. Although large language models achieve remarkable results in general question-answering tasks, they still face challenges in specialized fields. This paper proposes a novel framework called RAG-Chain, which aims to enhance the performance of general-domain large models on special biomedical reasoning and question-answering tasks. The RAG-Chain framework improves the knowledge retrieval and generation abilities of general models by a multi-stage processing of external knowledge and automatic construction of chain-of-thought templates combined with self-consistency validation process of choice shuffling. The experimental results show that RAG-Chain improves the accuracy of the baseline model by an average of 6.9% on the MedQA dataset without the need for pre-training or fine-tuning in biomedical fields, verifying its strong adaptability and effectiveness in different large language models.
Yongping Du, Zikai Wang 0009, Xingnan Jin
BIBM4
2024 Cross-biased contrastive learning for answer selection with dual-tower structure
Xingnan Jin, Yongping Du
Neurocomputing1
2024 Multi-stage knowledge distillation for sequential recommendation with interest knowledge
Yongping Du, Jinyu Niu, Xingnan Jin
Inf. Sci.4
2023 Low-Resource Efficient Multi-Stage Tuning Strategy for Biomedical Question Answering Task
abstract
The automated question-answering system plays a crucial role in improving the accuracy and efficiency of clinical decision-making. While large-scale language models perform prominently in the general domain, even surpassing human performance in certain aspects, challenges such as data privacy and the massive training costs affect the broader adoption of LLMs in biomedical domain. This study explores the strategies of efficient fine-tuning and optimization methods in biomedical domain. We introduce a multi-stage fine-tuning strategy that improves the accuracy of medical question-answering tasks significantly. Specifically, a contrastive learning technique based on multi-prompts is proposed, and a self-consistency voting approach is used to improve the accuracy of reasoning-required tasks. Experimental results on PubMedQA dataset reveal that even fine-tuning only 0.152% of the baseline’s parameters, our method still improves its performance, making it outperform the domain-specific pre-trained models and achieve performance comparable to GPT-4.
Yongping Du, Xingnan Jin
BIBM3
2023 Sentiment enhanced answer generation and information fusing for product-related question answering
Yongping Du, Xingnan Jin, Jingya Yan
Inf. Sci.2
2023 Improving Biomedical Question Answering by Data Augmentation and Model Weighting
abstract
Biomedical Question Answering aims to extract an answer to the given question from a biomedical context. Due to the strong professionalism of specific domain, it's more difficult to build large-scale datasets for specific domain question answering. Existing methods are limited by the lack of training data, and the performance is not as good as in open-domain settings, especially degrading when facing to the adversarial sample. We try to resolve the above issues. First, effective data augmentation strategies are adopted to improve the model training, including slide window, summarization and round-trip translation. Second, we propose a model weighting strategy for the final answer prediction in biomedical domain, which combines the advantage of two models, open-domain model QANet and BioBERT pre-trained in biomedical domain data. Finally, we give adversarial training to reinforce the robustness of the model. The public biomedical dataset collected from PubMed provided by BioASQ challenge is used to evaluate our approach. The results show that the model performance has been improved significantly compared to the single model and other models participated in BioASQ challenge. It can learn richer semantic expression from data augmentation and adversarial samples, which is beneficial to solve more complex question answering problems in biomedical domain.
Yongping Du, Jingya Yan, Yuxuan Lu 0003, Yiliang Zhao, Xingnan Jin
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 Gated attention fusion network for multimodal sentiment classification
Yongping Du, Zhi Peng, Xingnan Jin
Knowl. Based Syst.4
2021 Dual Model Weighting Strategy and Data Augmentation in Biomedical Question Answering
abstract
Biomedical Question Answering aims to extract an answer to the given question from a biomedical context. Due to the strong professionalism of specific domain, it’s more difficult to build large-scale datasets for specific domain question answering. Existing methods are limited by the lack of training data, and the performance is not as good as in open-domain settings. We propose a model weighting strategy for the final answer prediction in biomedical domain, which combines the advantage of two models, open-domain model QANet and BioBERT pretrained in biomedical domain data. Especially, we adopt effective data augmentation strategies to improve the model performance, including round-trip translation and summarization. The public biomedical dataset collected from PubMed provided by BioASQ is used to evaluate our approach. The results show that the model performance has been improved significantly on BioASQ 6B, 7B and 8B datasets compared to the single model.
Yongping Du, Jingya Yan, Yiliang Zhao, Yuxuan Lu 0003, Xingnan Jin
BIBM5