EDBT 2026 Demo / reviewers in the wild / expert
Cong Sun 0004
dblp:45/5103-4
· DBLP profile ↗
10ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0003-0240-1070ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A multimodal approach for few-shot biomedical named entity recognition in low-resource languages
Leilei Su, Mingquan Lin, Yifan Peng 0002, Cong Sun 0004 |
J. Biomed. Informatics | 6 |
| 2024 | Deep learning with noisy labels in medical prediction problems: a scoping reviewabstractOBJECTIVES: Medical research faces substantial challenges from noisy labels attributed to factors like inter-expert variability and machine-extracted labels. Despite this, the adoption of label noise management remains limited, and label noise is largely ignored. To this end, there is a critical need to conduct a scoping review focusing on the problem space. This scoping review aims to comprehensively review label noise management in deep learning-based medical prediction problems, which includes label noise detection, label noise handling, and evaluation. Research involving label uncertainty is also included. METHODS: Our scoping review follows the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. We searched 4 databases, including PubMed, IEEE Xplore, Google Scholar, and Semantic Scholar. Our search terms include "noisy label AND medical/healthcare/clinical," "uncertainty AND medical/healthcare/clinical," and "noise AND medical/healthcare/clinical." RESULTS: A total of 60 papers met inclusion criteria between 2016 and 2023. A series of practical questions in medical research are investigated. These include the sources of label noise, the impact of label noise, the detection of label noise, label noise handling techniques, and their evaluation. Categorization of both label noise detection methods and handling techniques are provided. DISCUSSION: From a methodological perspective, we observe that the medical community has been up to date with the broader deep-learning community, given that most techniques have been evaluated on medical data. We recommend considering label noise as a standard element in medical research, even if it is not dedicated to handling noisy labels. Initial experiments can start with easy-to-implement methods, such as noise-robust loss functions, weighting, and curriculum learning. Yishu Wei, Cong Sun 0004, Mingquan Lin, Hongmei Jiang, Yifan Peng 0002 |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | Demonstration-based learning for few-shot biomedical named entity recognition under machine reading comprehension
Leilei Su, Yifan Peng 0002, Cong Sun 0004 |
J. Biomed. Informatics | 4 |
| 2022 | MRC4BioER: Joint extraction of biomedical entities and relations in the machine reading comprehension framework
Cong Sun 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
J. Biomed. Informatics | 1 |
| 2021 | Deep learning with language models improves named entity recognition for PharmaCoNERabstractBACKGROUND: The recognition of pharmacological substances, compounds and proteins is essential for biomedical relation extraction, knowledge graph construction, drug discovery, as well as medical question answering. Although considerable efforts have been made to recognize biomedical entities in English texts, to date, only few limited attempts were made to recognize them from biomedical texts in other languages. PharmaCoNER is a named entity recognition challenge to recognize pharmacological entities from Spanish texts. Because there are currently abundant resources in the field of natural language processing, how to leverage these resources to the PharmaCoNER challenge is a meaningful study. METHODS: Inspired by the success of deep learning with language models, we compare and explore various representative BERT models to promote the development of the PharmaCoNER task. RESULTS: The experimental results show that deep learning with language models can effectively improve model performance on the PharmaCoNER dataset. Our method achieves state-of-the-art performance on the PharmaCoNER dataset, with a max F1-score of 92.01%. CONCLUSION: For the BERT models on the PharmaCoNER dataset, biomedical domain knowledge has a greater impact on model performance than the native language (i.e., Spanish). The BERT models can obtain competitive performance by using WordPiece to alleviate the out of vocabulary limitation. The performance on the BERT model can be further improved by constructing a specific vocabulary based on domain knowledge. Moreover, the character case also has a certain impact on model performance. Cong Sun 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
BMC Bioinform. | 1 |
| 2021 | Biomedical named entity recognition using BERT in the machine reading comprehension framework
Cong Sun 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
J. Biomed. Informatics | 1 |
| 2020 | Chemical-protein interaction extraction via Gaussian probability distribution and external biomedical knowledgeabstractMOTIVATION: The biomedical literature contains a wealth of chemical-protein interactions (CPIs). Automatically extracting CPIs described in biomedical literature is essential for drug discovery, precision medicine, as well as basic biomedical research. Most existing methods focus only on the sentence sequence to identify these CPIs. However, the local structure of sentences and external biomedical knowledge also contain valuable information. Effective use of such information may improve the performance of CPI extraction. RESULTS: In this article, we propose a novel neural network-based approach to improve CPI extraction. Specifically, the approach first employs BERT to generate high-quality contextual representations of the title sequence, instance sequence and knowledge sequence. Then, the Gaussian probability distribution is introduced to capture the local structure of the instance. Meanwhile, the attention mechanism is applied to fuse the title information and biomedical knowledge, respectively. Finally, the related representations are concatenated and fed into the softmax function to extract CPIs. We evaluate our proposed model on the CHEMPROT corpus. Our proposed model is superior in performance as compared with other state-of-the-art models. The experimental results show that the Gaussian probability distribution and external knowledge are complementary to each other. Integrating them can effectively improve the CPI extraction performance. Furthermore, the Gaussian probability distribution can effectively improve the extraction performance of sentences with overlapping relations in biomedical relation extraction tasks. AVAILABILITY AND IMPLEMENTATION: Data and code are available at https://github.com/CongSun-dlut/CPI_extraction. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cong Sun 0004, Leilei Su, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
Bioinform. | 1 |
| 2020 | Attention guided capsule networks for chemical-protein interaction extraction
Cong Sun 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
J. Biomed. Informatics | 1 |
| 2018 | Hierarchical Recurrent Convolutional Neural Network for Chemical-protein Relation Extraction from Biomedical Literature
Cong Sun 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Liang Yang 0003, Kan Xu, Yi-Jia Zhang 0001 |
BIBM | 1 |
| 2017 | A hybrid protein-protein interaction triple extraction method for biomedical literatureabstractProtein-protein interaction extraction research can be widely applied to the field of life science research. However, most of the machine learning based methods focus on binary PPI relation extraction, which loses rich relationship type information that is critical to the PPIs study. The rule based open information extraction methods can extract the PPI triple (i.e. “protein1, interaction word, protein2”), but suffers from low recall rate problem. In this paper, we propose a hybrid protein-protein interaction triple extraction method. In this method, firstly, machine learning techniques are used to recognize protein entities and extract relational protein pairs. Then, the syntactic patterns and a dictionary are employed to find out corresponding interaction words that represent the relationships between two proteins. This method obtains an F-score of 40.18% on the AImed corpus, which is much higher than the result achieved by the rule based Stanford open information extraction method. Zhehuan Zhao, Cong Sun 0004, Lei Wang 0085, Hongfei Lin |
BIBM | 3 |