Guoqing Zhang 0006

dblp:27/5832-6 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0001-8827-7546ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Retrieval-Augmented Generation Enhanced Domain-Adaptive Question Answering System for Radiation Biology
abstract
Large language models (LLMs) enhanced with retrieval-augmented generation (RAG) techniques still face difficulties delivering trustworthy question answering (QA) in specialized areas like radiation biology. This field is closely linked to human health and disease risks, making interpretable and accurate QA systems particularly important for scientific research and public understanding. To address these challenges, we present BioRadRAG (https://www.biosino.org/radiation/), a domain-specific RAG-based QA system designed to provide evidence-grounded answers in radiation biology. BioRadRAG constructs a curated knowledge base of$\mathbf{1 0, 6 6 9}$publications and adopts a parent-child chunking strategy to improve retrieval granularity and contextual coherence. We propose a multi-stage reasoning pipeline combining question classification, semantic rewriting, and reranking, all guided by structured prompts for response generation, thereby enhancing factual grounding and traceability. Two benchmark datasets were curated for evaluation. Experimental results demonstrate that BioRadRAG improves F1 scores by$\mathbf{7 \% - 1 1 \%}$on objective QA tasks compared to general-purpose LLMs and achieves better performance in subjective question evaluations. Ablation analysis further shows that the semantic rewriting and reranking modules contribute complementary gains in recall and precision, validating the multi-stage pipeline design. Furthermore, the system supports knowledge source filtering by publication year and journal impact factor, paragraph-level localization, mind map generation, and multi-turn interaction, enhancing accessibility and transparency in the exploration of scientific information. These findings highlight BioRadRAG's potential to advance domainspecific, explainable QA in radiation biology research.
Wanting Hu, Xinhao Zhuang, Guoping Zhao, Peng Zhang 0047, Guoqing Zhang 0006
BIBM6
2024 CyclicPepedia: a knowledge base of natural and synthetic cyclic peptides
abstract
Cyclic peptides offer a range of notable advantages, including potent antibacterial properties, high binding affinity and specificity to target molecules, and minimal toxicity, making them highly promising candidates for drug development. However, a comprehensive database that consolidates both synthetically derived and naturally occurring cyclic peptides is conspicuously absent. To address this void, we introduce CyclicPepedia (https://www.biosino.org/iMAC/cyclicpepedia/), a pioneering database that encompasses 8744 known cyclic peptides. This repository, structured as a composite knowledge network, offers a wealth of information encompassing various aspects of cyclic peptides, such as cyclic peptides' sources, categorizations, structural characteristics, pharmacokinetic profiles, physicochemical properties, patented drug applications, and a collection of crucial publications. Supported by a user-friendly knowledge retrieval system and calculation tools specifically designed for cyclic peptides, CyclicPepedia will be able to facilitate advancements in cyclic peptide drug development.
Lei Liu 0033, Suqi Cao, Guoqing Zhang 0006, Ruixin Zhu, Dingfeng Wu
Briefings Bioinform.6
2023 A dual-channel deep learning approach to continuous prediction of acute kidney injury in the intensive care unit
abstract
Acute Kidney Injury (AKI) often occurs in the intensive care units (ICU), where it is associated with high morbidity and mortality. Early and continuous prediction of AKI plays a crucial role in preventing AKI events, providing timely treatment, and reducing mortality. Continuously monitoring the high dimensional vital signs and lab measurements of the patients has always been a challenging task. In this study, we proposed a novel dual-channel deep learning approach named DC-AKI to address this challenge. DC-AKI extracts feature of temporal variables at different granularity, comprehensively and accurately obtains effective information, and finally continuous predicting risks of the patients. One channel of the model calculates corresponding local interactions of the variables with convolution kernels. The other channel uses the gated recurrent units (GRU) network and attention mechanism to obtains global interactions information among the variables. Finally, the model uses full connection layer to combine the representation vectors of the two channels, and uses Sigmoid classifier for prediction. We trained and tested our model on the MIMIC IV dataset, and the AUC ROC value of DC-AKI for predicting the risk of AKI in ICU patients 48 hours in advance was 0.9527, outperforming other methods. And our model was externally validated on the eICU dataset with an AUC ROC of 0.9054. These results demonstrate the robustness and accuracy of the model on different datasets. This finding provides clinicians with an opportunity to identify and manage patients with AKI in a timely manner, thereby improving patient prognosis. The implemented code is available online at https://github.com/BioMedBigDataCenter/DC-AKI.
Xuetong Kong, Peng Zhang 0047, Yunchao Ling, Daqing Lv, Guoqing Zhang 0006
BIBM6
2023 A domain adaptive pre-training language model for sentence classification of Chinese electronic medical record
abstract
Accurately extracting and classifying Chinese electronic medical record (EMR), which contain huge amounts of valuable medical information, have promising practical application and medical value in the health care of China. While the pivotal issue has gathered escalating attention, the bulk of current research is directed towards operations conducted at the document or entity level within medical records. Only a restricted body of work addresses these concerns at the sentence level, a critical aspect for downstream tasks like medical information retrieval, diagnosis normalization, and question answering. In this paper, we present a domain adaptive pre-training language model named CEMR-LM for sentence classification of Chinese EMRs. CEMR-LM acquires Chinese medical domain knowledge through the utilization of copious unlabeled clinical corpus for pre-training the language model. This is fortified by combining fine-tuning strategy and a dual-channel mechanism, which collectively contribute to the model’s heightened performance. Experiments on the benchmark dataset and real world hospital dataset both demonstrate that CEMR-LM is superior to the state-of-the-art methods. Furthermore, CEMR-LM possesses the capability to elucidate indicative elements within medical records by visualizing of the attention weights embedded within the model. The implemented code and experimental datasets are available online at https://github.com/BioMedBigDataCenter/CEMR-LM.
Yilin Zou, Peng Zhang 0047, Yunchao Ling, Daqing Lv, Shun Lu 0003, Guoqing Zhang 0006
BIBM7
2022 Coronavirus GenBrowser for monitoring the transmission and evolution of SARS-CoV-2
abstract
Genomic epidemiology is important to study the COVID-19 pandemic, and more than two million severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) genomic sequences were deposited into public databases. However, the exponential increase of sequences invokes unprecedented bioinformatic challenges. Here, we present the Coronavirus GenBrowser (CGB) based on a highly efficient analysis framework and a node-picking rendering strategy. In total, 1,002,739 high-quality genomic sequences with the transmission-related metadata were analyzed and visualized. The size of the core data file is only 12.20 MB, highly efficient for clean data sharing. Quick visualization modules and rich interactive operations are provided to explore the annotated SARS-CoV-2 evolutionary tree. CGB binary nomenclature is proposed to name each internal lineage. The pre-analyzed data can be filtered out according to the user-defined criteria to explore the transmission of SARS-CoV-2. Different evolutionary analyses can also be easily performed, such as the detection of accelerated evolution and ongoing positive selection. Moreover, the 75 genomic spots conserved in SARS-CoV-2 but non-conserved in other coronaviruses were identified, which may indicate the functional elements specifically important for SARS-CoV-2. The CGB was written in Java and JavaScript. It not only enables users who have no programming skills to analyze millions of genomic sequences, but also offers a panoramic vision of the transmission and evolution of SARS-CoV-2.
Dalang Yu, Bixia Tang, Yi-Hsuan Pan, Guangya Duan, Zi-Qian Hao, Hailong Mu, Long Dai, Wangjie Hu, Mochen Zhang, Tong Jin 0004, Cuiping Li 0004, Xiao Su 0004, Guoqing Zhang 0006, Wenming Zhao
Briefings Bioinform.18
2021 Dr AFC: drug repositioning through anti-fibrosis characteristic
abstract
Fibrosis is a key component in the pathogenic mechanism of a variety of diseases. These diseases involving fibrosis may share common mechanisms and therapeutic targets, and therefore common intervention strategies and medicines may be applicable for these diseases. For this reason, deliberately introducing anti-fibrosis characteristics into predictive modeling may lead to more success in drug repositioning. In this study, anti-fibrosis knowledge base was first built by collecting data from multiple resources. Both structural and biological profiles were then derived from the knowledge base and used for constructing machine learning models including Structural Profile Prediction Model (SPPM) and Biological Profile Prediction Model (BPPM). Three external public data sets were employed for validation purpose and further exploration of potential repositioning drugs in wider chemical space. The resulting SPPM and BPPM models achieve area under the receiver operating characteristic curve (area under the curve) of 0.879 and 0.972 in the training set, and 0.814 and 0.874 in the testing set. Additionally, our results also demonstrate that substantial amount of multi-targeting natural products possess notable anti-fibrosis characteristics and might serve as encouraging candidates in fibrosis treatment and drug repositioning. To leverage our methodology and findings, we developed repositioning prediction platform, drug repositioning based on anti-fibrosis characteristic that is freely accessible via https://www.biosino.org/drafc.
Dingfeng Wu, Wenxing Gao, Chuan Tian, Na Jiao, Sa Fang, Lixin Zhu, Guoqing Zhang 0006, Ruixin Zhu
Briefings Bioinform.10
2020 EpiDISH web server: Epigenetic Dissection of Intra-Sample-Heterogeneity with online GUI
abstract
SUMMARY: It is well recognized that cell-type heterogeneity hampers the interpretation of Epigenome-Wide Association Studies (EWAS). Many tools have emerged to address this issue, including several R/Bioconductor packages that infer cell-type composition. Here we present a web application for cell-type deconvolution, which offers the functionality of our EpiDISH Bioconductor/R package in a user-friendly GUI environment. Users can upload their data to infer cell-type composition and differentially methylated cytosines in individual cell-types (DMCTs) for a range of different tissues. AVAILABILITY AND IMPLEMENTATION: EpiDISH web server is implemented with Shiny in R, and is freely available at https://www.biosino.org/EpiDISH/.
Shijie C. Zheng, Charles E. Breeze, Stephan Beck 0002, Danyue Dong, Liang-Xiao Ma, Guoqing Zhang 0006, Andrew E. Teschendorff
Bioinform.8
2019 Knowledge Graph Embedding by Bias Vectors
abstract
Knowledge graph completion can predict the possible relation between entities. Previous work such as TransE, TransR, TransPES and GTrans embed knowledge graph into vector space and treat relations between entities as translations. In most cases, the more complex the algorithm is, the better the result will be, but it is difficult to apply to large-scale knowledge graphs. Therefore, we propose TransB, an efficient model, in this paper. We avoid the complex matrix or vector multiplication operation. Meanwhile, we make the representation of entities not too simple, which can satisfy the operation in the case of non-one-to-one relation. We use link prediction to evaluate the performance of our model in the experiment. The experimental results show that our model is valid and has low time complexity.
Minjie Ding, Weiqin Tong, Xuehai Ding, Xiaoli Zhi, Xiao Wang 0023, Guoqing Zhang 0006
ICTAI6
2019 A New Method for Complex Triplet Extraction of Biomedical Texts
Xiao Wang 0023, Qing Li 0011, Xuehai Ding, Guoqing Zhang 0006, Linhong Weng, Minjie Ding
KSEM (2)4
2008 GORouter: an RDF model for providing semantic query and inference services for Gene Ontology and its associations
abstract
BACKGROUND: The most renowned biological ontology, Gene Ontology (GO) is widely used for annotations of genes and gene products of different organisms. However, there are shortcomings in the Resource Description Framework (RDF) data file provided by the GO consortium: 1) Lack of sufficient semantic relationships between pairs of terms coming from the three independent GO sub-ontologies, that limit the power to provide complex semantic queries and inference services based on it. 2) The term-centric view of GO annotation data and the fact that all information is stored in a single file. This makes attempts to retrieve GO annotations based on big volume datasets unmanageable. 3) No support of GOSlim. RESULTS: We propose a RDF model, GORouter, which encodes heterogeneous original data in a uniform RDF format, creates additional ontology mappings between GO terms, and introduces a set of inference rulebases. Furthermore, we use the Oracle Network Data Model (NDM) as the native RDF data repository and the table function RDF_MATCH to seamlessly combine the result of RDF queries with traditional relational data. As a result, the scale of GORouter is minimized; information not directly involved in semantic inference is put into relational tables. CONCLUSION: Our work demonstrates how to use multiple semantic web tools and techniques to provide a mixture of semantic query and inference solutions of GO and its associations. GORouter is licensed under Apache License Version 2.0, and is accessible via the website: http://www.scbit.org/gorouter/.
Qingwei Xu, Yixiang Shi, Guoqing Zhang 0006, Qingming Luo
BMC Bioinform.4