Wei Tao 0003

dblp:17/6159-3 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-1800-1904ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations
abstract
Spoken Dialogue Models (SDMs) have recently attracted significant attention for their ability to generate voice responses directly to users’ spoken queries. Despite their increasing popularity, there exists a gap in research focused on comprehensively understanding their practical effectiveness in comprehending and emulating human conversations. This is especially true compared to text-based Large Language Models (LLMs), which benefit from extensive benchmarking. Human voice interactions are inherently more complex than text due to characteristics unique to spoken dialogue. Ambiguity poses one challenge, stemming from semantic factors like polysemy, as well as phonological aspects such as heterograph, heteronyms, and stress patterns. Additionally, context-dependency, like omission, coreference, and multi-turn interaction, adds further complexity to human conversational dynamics. To illuminate the current state of SDM development and to address these challenges, we present a benchmark dataset in this paper, which comprises 1,079 instances in English and Chinese. Accompanied by an LLM-based evaluation method that closely aligns with human judgment, this dataset facilitates a comprehensive exploration of the performance of SDMs in tackling these practical challenges.
Chengqian Ma, Wei Tao 0003, Steven Y. Guo
EMNLP2
2024 MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution
abstract
In software development, resolving the emergent issues within GitHub repositories is a complex challenge that involves not only the incorporation of new code but also the maintenance of existing code. Large Language Models (LLMs) have shown promise in code generation but face difficulties in resolving Github issues, particularly at the repository level. To overcome this challenge, we empirically study the reason why LLMs fail to resolve GitHub issues and analyze the major factors. Motivated by the empirical findings, we propose a novel LLM-based **M**ulti-**A**gent framework for **G**itHub **I**ssue re**S**olution, **MAGIS**, consisting of four agents customized for software evolution: Manager, Repository Custodian, Developer, and Quality Assurance Engineer agents. This framework leverages the collaboration of various agents in the planning and coding process to unlock the potential of LLMs to resolve GitHub issues. In experiments, we employ the SWE-bench benchmark to compare MAGIS with popular LLMs, including GPT-3.5, GPT-4, and Claude-2. MAGIS can resolve **13.94%** GitHub issues, significantly outperforming the baselines. Specifically, MAGIS achieves an eight-fold increase in resolved ratio over the direct application of GPT-4, the advanced LLM.
Wei Tao 0003, Yucheng Zhou 0001, Yanlin Wang 0001, Hongyu Zhang 0002, Yu Cheng 0001
NeurIPS1
2024 KADEL: Knowledge-Aware Denoising Learning for Commit Message Generation
abstract
Commit messages are natural language descriptions of code changes, which are important for software evolution such as code understanding and maintenance. However, previous methods are trained on the entire dataset without considering the fact that a portion of commit messages adhere to good practice (i.e., good-practice commits), while the rest do not. On the basis of our empirical study, we discover that training on good-practice commits significantly contributes to the commit message generation. Motivated by this finding, we propose a novel knowledge-aware denoising learning method called KADEL. Considering that good-practice commits constitute only a small proportion of the dataset, we align the remaining training samples with these good-practice commits. To achieve this, we propose a model that learns the commit knowledge by training on good-practice commits. This knowledge model enables supplementing more information for training samples that do not conform to good practice. However, since the supplementary information may contain noise or prediction errors, we propose a dynamic denoising training method. This method composes a distribution-aware confidence function and a dynamic distribution list, which enhances the effectiveness of the training process. Experimental results on the whole MCMD dataset demonstrate that our method overall achieves state-of-the-art performance compared with previous methods.
Wei Tao 0003, Yucheng Zhou 0001, Yanlin Wang 0001, Hongyu Zhang 0002, Haofen Wang
ACM Trans. Softw. Eng. Methodol.1
2022 RACE: Retrieval-augmented Commit Message Generation
abstract
Commit messages are important for software development and maintenance.Many neural network-based approaches have been proposed and shown promising results on automatic commit message generation.However, the generated commit messages could be repetitive or redundant.In this paper, we propose RACE, a new retrieval-augmented neural commit message generation method, which treats the retrieved similar commit as an exemplar and leverages it to generate an accurate commit message.As the retrieved commit message may not always accurately describe the content/intent of the current code diff, we also propose an exemplar guider, which learns the semantic similarity between the retrieved and current code diff and then guides the generation of commit message based on the similarity.We conduct extensive experiments on a large public dataset with five programming languages.Experimental results show that RACE can outperform all baselines.Furthermore, RACE can boost the performance of existing Seq2Seq models in commit message generation.
Ensheng Shi, Yanlin Wang 0001, Wei Tao 0003, Lun Du, Hongyu Zhang 0002, Shi Han, Dongmei Zhang 0001, Hongbin Sun 0001
EMNLP3
2022 Type-Aware Medical Visual Question Answering
abstract
Medical Visual Question Answering (Med-VQA) helps answer medical questions raised by patients automatically so as to relieve the shortage of experienced doctors. Cross-modal feature alignment is a major challenge of Med-VQA. Moreover, it is critical to exploit sufficient semantic features with the consideration of characteristic of medical images and language. In this paper, we propose a novel From Image type point To Sentence (FITS) method to tackle the above challenge. In particular, the type of the medical images is represented as a type point which is further considered in the question sentence representation. The combined representation aims to optimize the feature distribution in an embedding space and thus enhances the ability of semantic alignment. Type point is also used in two feature extraction modules for medical questions and images respectively, which can efficiently improve the reasoning ability of different modalities, and further enhance the applicability of the fusion method for Med-VQA. The experimental results show that FITS outperforms all the previous approaches in terms of accuracy especially in open-ended questions significantly.
Anda Zhang, Wei Tao 0003, Haofen Wang
ICASSP2
2022 A large-scale empirical study of commit message generation: models, datasets and evaluation
Wei Tao 0003, Yanlin Wang 0001, Ensheng Shi, Lun Du, Shi Han, Hongyu Zhang 0002, Dongmei Zhang 0001
Empir. Softw. Eng.1
2021 Triple Sequence Generative Adversarial Nets for Unsupervised Image Captioning
abstract
Labelling image-sentence is expensive and some unsupervised image captioning methods show promising results on caption generation. However, the generated captions are not very relevant to images due to the excessive dependence on the corpus. In order to overcome that drawback, we focus on the correspondence between image and sentence to construct an image caption with better mapping relation. In this paper, we present a novel triple sequence generative adversarial net including an image generator, a discriminator, and a sentence generator. The image generator is used to generate the image regions for words. Meanwhile, the sentence corpus guides the sentence generator based on the generated image regions. The discriminator judges the relevance between the words in the sentence and the generated image regions. In the experiments, we use a large number of unpaired images and sentences to train our model on the unsupervised and unpaired setting. The experimental results demonstrate that our method achieves significant improvements as compared to all baselines.
Yucheng Zhou 0001, Wei Tao 0003
ICASSP2
2021 On the Evaluation of Commit Message Generation Models: An Experimental Study
abstract
Commit messages are natural language descriptions of code changes, which are important for program understanding and maintenance. However, writing commit messages manually is time-consuming and laborious, especially when the code is updated frequently. Various approaches utilizing generation or retrieval techniques have been proposed to automatically generate commit messages. To achieve a better understanding of how the existing approaches perform in solving this problem, this paper conducts a systematic and in-depth analysis of the state-of-the-art models and datasets. We find that: (1) Different variants of the BLEU metric are used in previous works, which affects the evaluation and understanding of existing methods. (2) Most existing datasets are crawled only from Java repositories while repositories in other programming languages are not sufficiently explored. (3) Dataset splitting strategies can influence the performance of existing models by a large margin. Some models show better performance when the datasets are split by commit, while other models perform better when the datasets are split by timestamp or by project. Based on our findings, we conduct a human evaluation and find the BLEU metric that best correlates with the human scores for the task. We also collect a large-scale, information-rich, and multi-language commit message dataset MCMD and evaluate existing models on this dataset. Furthermore, we conduct extensive experiments under different dataset splitting strategies and suggest the suitable models under different scenarios. Based on the experimental results and findings, we provide feasible suggestions for comprehensively evaluating commit message generation models and discuss possible future research directions. We believe this work can help practitioners and researchers better evaluate and select models for automatic commit message generation. Our source code and data are available at https://github.com/DeepSoftwareAnalytics/CommitMsgEmpirical.
Wei Tao 0003, Yanlin Wang 0001, Ensheng Shi, Lun Du, Shi Han, Hongyu Zhang 0002, Dongmei Zhang 0001
ICSME1