EDBT 2026 Demo / reviewers in the wild / expert
Yu Wang 0072
dblp:02/5889-72
· DBLP profile ↗
12ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0003-1606-6795ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FunLoc: A Novel Function-level Bug Localization Framework Enhanced by Contrastive and Active Learning StrategiesabstractThe increasing complexity of software systems has made them more prone to bugs, prompting the development of automated bug localization techniques to ensure software reliability. Despite these techniques having demonstrated notable success at the file level, their application and optimization at the function level often encounter serious performance cliffs. This limitation underscores the urgent need for a dedicated framework for function-level bug localization, which we address through FunLoc, a novel framework that takes coarse-grained source files as input units and identifies fine-grained buggy functions as output. To address the critical challenges of handling domain-specific bug reports and managing vast function-level sample space, we introduce two key innovations that are seamlessly integrated into FunLoc. First, we design a contrastive learning-based domain-adaptive language model to enhance the framework's ability to process and interpret specialized bug reports effectively. Second, we propose an active learning-based dynamic negative sampling strategy to address the scalability issues arising from the extensive function-level sample space. To evaluate the effectiveness of our approach, we extend and release a function-level bug localization dataset derived from large-scale real-world projects. Extensive experiments demonstrate that our approach outperforms state-of-the-art techniques. Ziye Zhu, Liangliang Peng, Yu Wang 0072, Yun Li 0009, Xianzhong Long |
CIKM | 3 |
| 2025 | Enhancing biomedical named entity recognition with parallel boundary detection and category classificationabstractBACKGROUND: Named entity recognition is a fundamental task in natural language processing. Recognizing entities in biomedical text, known as the BioNER, is particularly crucial for cutting-edge applications. However, BioNER poses greater challenges compared to traditional NER due to (1) nested structures and (2) category correlations inherent in biomedical entities. Recently, various BioNER models have been developed based on region classification or large language models. Despite being successful, these models still struggle to balance handling nested structures and capturing category knowledge. RESULTS: We present a novel parallel BioNER model, BEAN, designed to address the unique properties of biomedical entities while achieving a reasonable balance between handling nested structures and incorporating category correlations. Extensive experiments on five public NER datasets, including four biomedical datasets, demonstrate that BEAN achieves state-of-the-art performance. CONCLUSIONS: The proposed BEAN is elaborately designed to achieve two key objectives of the BioNER task: clearly detecting entity boundaries and correctly classifying entity categories. It is the first BioNER model to handle nested structures and category correlations in parallel. We exploit head, tail, and contextualized features to efficiently detect entity boundaries via a triaffine model. To the best of our knowledge, we are the first to introduce a multi-label classification model for the BioNER task to extract entity category information without boundary guidance. Yu Wang 0072, Hanghang Tong, Ziye Zhu, Fengzhen Hou, Yun Li 0009 |
BMC Bioinform. | 1 |
| 2023 | BL-GAN: Semi-Supervised Bug Localization via Generative Adversarial NetworkabstractVarious automated bug localization technologies have recently emerged that require adequate bug-fix records available to train a predictive model. However, many projects in practice might not provide these necessities, especially for new projects in the first release, due to the expensive human effort for constructing a large amount of bug-fix records. Aiming to capture the potential relevance distribution between the bug report and code file from a limited number of available bug-fix records, we present the first semi-supervised bug localization model named BL-GAN in this paper. For this purpose, the promising Generative Adversarial Network is introduced in BL-GAN, in which synthetic bug-fix records close to the real ones are constructed by searching the project directory tree to generate file paths instead of traversing the contents of all code files. For processing bug reports, the proposed BL-GAN adopts an attention-based Transformer architecture to capture semantic and sequence information. In order to capture the proprietary structural information in code files, BL-GAN incorporates a novel multilayer Graph Convolutional Network to process the source code in a graphical view. Extensive experiments on large-scale real-world datasets reveal that our model BL-GAN significantly outperforms the state-of-the-art on all evaluation measures. Ziye Zhu, Hanghang Tong, Yu Wang 0072, Yun Li 0009 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Enhancing bug localization with bug report decomposition and code hierarchical network
Ziye Zhu, Hanghang Tong, Yu Wang 0072, Yun Li 0009 |
Knowl. Based Syst. | 3 |
| 2022 | Nested Named Entity Recognition: A SurveyabstractWith the rapid development of text mining, many studies observe that text generally contains a variety of implicit information, and it is important to develop techniques for extracting such information. Named Entity Recognition (NER), the first step of information extraction, mainly identifies names of persons, locations, and organizations in text. Although existing neural-based NER approaches achieve great success in many language domains, most of them normally ignore the nested nature of named entities. Recently, diverse studies focus on the nested NER problem and yield state-of-the-art performance. This survey attempts to provide a comprehensive review on existing approaches for nested NER from the perspectives of the model architecture and the model property, which may help readers have a better understanding of the current research status and ideas. In this survey, we first introduce the background of nested NER, especially the differences between nested NER and traditional (i.e., flat) NER. We then review the existing nested NER approaches from 2002 to 2020 and mainly classify them into five categories according to the model architecture, including early rule-based, layered-based, region-based, hypergraph-based, and transition-based approaches. We also explore in greater depth the impact of key properties unique to nested NER approaches from the model property perspective, namely entity dependency, stage framework, error propagation, and tag scheme. Finally, we summarize the open challenges and point out a few possible future directions in this area. This survey would be useful for three kinds of readers: (i) Newcomers in the field who want to learn about NER, especially for nested NER. (ii) Researchers who want to clarify the relationship and advantages between flat NER and nested NER. (iii) Practitioners who just need to determine which NER technique (i.e., nested or not) works best in their applications. Yu Wang 0072, Hanghang Tong, Ziye Zhu, Yun Li 0009 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2021 | TroBo: A Novel Deep Transfer Model for Enhancing Cross-Project Bug Localization
Ziye Zhu, Yu Wang 0072, Yun Li 0009 |
KSEM | 2 |
| 2021 | A deep multimodal model for bug localization
Ziye Zhu, Yun Li 0009, Yu Wang 0072, Yaojing Wang, Hanghang Tong |
Data Min. Knowl. Discov. | 3 |
| 2020 | HIT: Nested Named Entity Recognition via Head-Tail Pair and Token InteractionabstractNamed Entity Recognition (NER) is a fundamental task in natural language processing.In order to identify entities with nested structure, many sophisticated methods have been recently developed based on either the traditional sequence labeling approaches or directed hypergraph structures.Despite being successful, these methods often fall short in striking a good balance between the expression power for nested structure and the model complexity.To address this issue, we present a novel nested NER model named HIT.Our proposed HIT model leverages two key properties pertaining to the (nested) named entity, including (1) explicit boundary tokens and (2) tight internal connection between tokens within the boundary.Specifically, we design (1) Head-Tail Detector based on the multi-head selfattention mechanism and bi-affine classifier to detect boundary tokens, and (2) Token Interaction Tagger based on traditional sequence labeling approaches to characterize the internal token connection within the boundary.Experiments on three public NER datasets demonstrate that the proposed HIT achieves state-ofthe-art performance. Yu Wang 0072, Yun Li 0009, Hanghang Tong, Ziye Zhu |
EMNLP (1) | 1 |
| 2020 | CooBa: Cross-project Bug Localization via Adversarial Transfer LearningabstractBug localization plays an important role in software quality control. Many supervised machine learning models have been developed based on historical bug-fix information. Despite being successful, these methods often require sufficient historical data (i.e., labels), which is not always available especially for newly developed software projects. In response, cross-project bug localization techniques have recently emerged whose key idea is to transferring knowledge from label-rich source project to locate bugs in the target project. However, a major limitation of these existing techniques lies in that they fail to capture the specificity of each individual project, and are thus prone to negative transfer. To address this issue, we propose an adversarial transfer learning bug localization approach, focusing on only transferring the common characteristics (i.e., public information) across projects. Specifically, our approach (CooBa) learns the indicative public information from cross-project bug reports through a shared encoder, and extracts the private information from code files by an individual feature extractor for each project. CooBa further incorporates adversarial learning mechanism to ensure that public information shared between multiple projects could be effectively extracted. Extensive experiments on four large-scale real-world data sets demonstrate that the proposed CooBa significantly outperforms the state of the art techniques. Ziye Zhu, Yun Li 0009, Hanghang Tong, Yu Wang 0072 |
IJCAI | 4 |
| 2020 | Adversarial Named Entity Recognition with POS label embeddingabstractNamed Entity Recognition (NER) is dedicated to recognizing different types of named entity. Previous works have shown that part-of-speech, as an important feature, provides complementary syntactical information to NER systems. However, these studies suffer from two limitations: (i) the previous models do not consider the noise from part-of-speech; (ii) the previous models need to re-extract features from token representations. In this paper, we propose a novel approach that can alleviate the above issues as well as make full use of part-of-speech features via attention mechanism and adversarial training. We evaluate our model on three NER datasets, and the experimental results demonstrate that our model achieves a state-of-the-art F1-score of Twitter dataset while matching a state-of-the-art performance on the CoNLL-2003 and Weibo datasets. Yu Wang 0072, Bin Xia 0003, Yun Li 0009, Ziye Zhu |
IJCNN | 2 |
| 2020 | Adversarial Learning for Multi-Task Sequence Labeling With Attention MechanismabstractWith the requirements of natural language applications, multi-task sequence labeling methods have some immediate benefits over the single-task sequence labeling methods. Recently, many state-of-the-art multi-task sequence labeling methods were proposed, while still many issues to be resolved including (C1) exploring a more general relationship between tasks, (C2) extracting the task-shared knowledge purely and (C3) merging the task-shared knowledge for each task appropriately. To address the above challenges, we propose MTAA, a symmetric multi-task sequence labeling model, which performs an arbitrary number of tasks simultaneously. Furthermore, MTAA extracts the shared knowledge among tasks by adversarial learning and integrates the proposed multi-representation fusion attention mechanism for merging feature representations. We evaluate MTAA on two widely used data sets: CoNLL2003 and OntoNotes5.0. Experimental results show that our proposed model outperforms the latest methods on the named entity recognition and the syntactic chunking task by a large margin, and achieves state-of-the-art results on the part-of-speech tagging task. Yu Wang 0072, Yun Li 0009, Ziye Zhu, Hanghang Tong |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | SC-NER: A Sequence-to-Sequence Model with Sentence Classification for Named Entity Recognition
Yu Wang 0072, Yun Li 0009, Ziye Zhu, Bin Xia 0003, Zheng Liu 0001 |
PAKDD (1) | 1 |