VLDB 2026 Research / reviewers in the wild / expert
Pingyi Zhou
dblp:178/6404
· DBLP profile ↗
11ranked-venue papers
3as first author
5since 2021 · last 2023
0000-0002-1305-7992ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | History, Present and Future: Enhancing Dialogue Generation with Few-Shot History-Future PromptabstractDialogue history and response in open-domain dialogue are loosely coupled. Generating informative responses solely based on the original dialogue history is not easy, as dialogue history may not contain enough information or it may contain irrelevant noises. Intuitively, if a generation model can foresee possible dialogue future, or obtain real useful histories, it could generate more informative responses. In this paper, we propose a novel lightweight dialogue generation framework named few-shot history-future prompt that utilizes useful histories and simulated futures to help generate informative responses, without the need for fine-tuning or adding extra parameters. To obtain useful histories, we retrieve and combine relevant utterances from noisy multi- turn histories. Then we adopt a retrieval-generation hybrid approach to obtain diversified simulated futures. Such that our model could learn to condition on history combinations and simulated futures via few-shot learning. Experiments over publicly available datasets demonstrate that our method can help models generate better responses. Yasheng Wang, Fei Mi, Pingyi Zhou, Jin Liu 0016, Xin Jiang 0002, Qun Liu 0001 |
ICASSP | 5 |
| 2022 | Modeling Hierarchical Syntax Structure with Triplet Position for Source Code SummarizationabstractAutomatic code summarization, which aims to describe the source code in natural language, has become an essential task in software maintenance.Our fellow researchers have attempted to achieve such a purpose through various machine learning-based approaches.One key challenge keeping these approaches from being practical lies in the lacking of retaining the semantic structure of source code, which has unfortunately been overlooked by the stateof-the-art methods.Existing approaches resort to representing the syntax structure of code by modeling the Abstract Syntax Trees (ASTs).However, the hierarchical structures of ASTs have not been well explored.In this paper, we propose CODESCRIBE to model the hierarchical syntax structure of code by introducing a novel triplet position for code summarization.Specifically, CODESCRIBE leverages the graph neural network and Transformer to preserve the structural and sequential information of code, respectively.In addition, we propose a pointer-generator network that pays attention to both the structure and sequential tokens of code for a better summary generation.Experiments on two real-world datasets in Java and Python demonstrate the effectiveness of our proposed approach when compared with several state-of-the-art baselines 1 . Juncai Guo 0003, Jin Liu 0016, Yao Wan 0001, Li Li 0029, Pingyi Zhou |
ACL (1) | 5 |
| 2022 | Pan More Gold from the Sand: Refining Open-domain Dialogue Training with Noisy Self-Retrieval GenerationabstractReal human conversation data are complicated, heterogeneous, and noisy, from which building open-domain dialogue systems remains a challenging task. In fact, such dialogue data still contains a wealth of information and knowledge, however, they are not fully explored. In this paper, we show existing open-domain dialogue generation methods that memorize context-response paired data with autoregressive or encode-decode language models underutilize the training data. Different from current approaches, using external knowledge, we explore a retrieval-generation training framework that can take advantage of the heterogeneous and noisy training data by considering them as “evidence”. In particular, we use BERTScore for retrieval, which gives better qualities of the evidence and generation. Experiments over publicly available datasets demonstrate that our method can help models generate better responses, even such training data are usually impressed as low-quality data. Such performance gain is comparable with those improved by enlarging the training set, even better. We also found that the model performance has a positive correlation with the relevance of the retrieved evidence. Moreover, our method performed well on zero-shot experiments, which indicates that our method can be more robust to real-world data. Yasheng Wang, Fei Mi, Pingyi Zhou, Xin Wang 0114, Jin Liu 0016, Xin Jiang 0002, Qun Liu 0001 |
COLING | 5 |
| 2022 | Test-Driven Multi-Task Learning with Functionally Equivalent Code Transformation for Neural Code GenerationabstractAutomated code generation is a longstanding challenge in both communities of software engineering and artificial intelligence. Currently, some works have started to investigate the functional correctness of code generation, where a code snippet is considered correct if it passes a set of test cases. However, most existing works still model code generation as text generation without considering program-specific information, such as functionally equivalent code snippets and test execution feedback. To address the above limitations, this paper proposes a method combining program analysis with deep learning for neural code generation, where functionally equivalent code snippets and test execution feedback will be considered at the training stage. Concretely, we firstly design several code transformation heuristics to produce different variants of the code snippet satisfying the same functionality. In addition, we employ the test execution feedback and design a test-driven discriminative task to train a novel discriminator, aiming to let the model distinguish whether the generated code is correct or not. The preliminary results on a newly published dataset demonstrate the effectiveness of our proposed framework for code generation. Particularly, in terms of the [email protected] metric, we achieve 8.81 and 11.53 gains compared with CodeGPT and CodeT5, respectively. Xin Wang 0114, Xiao Liu 0004, Pingyi Zhou, Qixia Liu, Jin Liu 0016, Hao Wu 0010, Xiaohui Cui |
ASE | 3 |
| 2021 | ServiceBERT: A Pre-trained Model for Web Service Tagging and Recommendation
Xin Wang 0114, Pingyi Zhou, Yasheng Wang, Xiao Liu 0004, Jin Liu 0016, Hao Wu 0010 |
ICSOC | 2 |
| 2019 | Is deep learning better than traditional approaches in tag recommendation for software information sites?
Pingyi Zhou, Jin Liu 0016, Xiao Liu 0004, Zijiang Yang 0006, John C. Grundy |
Inf. Softw. Technol. | 1 |
| 2018 | FastTagRec: fast tag recommendation for software information sites
Jin Liu 0016, Pingyi Zhou, Zijiang Yang 0006, Xiao Liu 0004, John C. Grundy |
Autom. Softw. Eng. | 2 |
| 2017 | Constructing Drug Ingredient Interaction Network to Ensure Medication SecurityabstractWith the fast development of big data, we can do better in modern digital health.The rational use of drugs is main threat of medication security.With big data mining, we analyze the drug instructions to automatically detect the adverse drug interaction in drug combination.The proposed method synthetically employs the NLP and complex network to automatically construct drug ingredient interaction network for mining and inference adverse drug interaction in drug combination.First, many drug instructions are collected by crawler from professional medicine websites.Next, NLP is utilized to automatically extract effective drug ingredient for building drug ingredient library.Then, drug ingredient interactions are extracted from drug instructions.The last, build drug ingredient interaction network.Because the drug ingredient interaction network is kind of complex, the principle of complex network is used to analyze the drug ingredient interaction network.Then we automatically generate a report to make knowledge visualize.The constructed drug ingredient interaction network is verified in experiment.The experiment results indicate that the validity and effectiveness of our proposed method. Zhiren Mao, Pingyi Zhou, Jiaxiang Zhong, Luojia Jiang, Chongzhi Deng |
SEKE | 2 |
| 2017 | Improving Bug Triage with Relevant SearchabstractBug triage is a process where bugs are assigned to developers.In large open source projects such as Mozilla and Eclipse, bug triage is time-consuming because numerous bugs are submitted everyday.To improve bug triage, many studies have proposed automatic approaches to recommend proper developers for resolving bugs.These approaches are based on machine learning algorithms, which treat bug triage like text classification.Although they are effective, the accuracy of them can be further improved.Our goal is to propose a method not only has good performance but also is simple.We propose a method based on relevant search technique to recommend developers for the given bugs.First, we construct an index for bugs to make them searchable.Then, for a given bug to be assigned, we utilize the index to search for the bugs related to it.Finally, we analyze these related bugs and recommend developers based on them.We conduct experiments on bugs of Mozilla and Eclipse to evaluate our method.The results indicate that our method has a good performance and outperforms machine learning algorithms like Naïve Bayes and SVM. Xinyu Peng, Pingyi Zhou |
SEKE | 2 |
| 2017 | Scalable tag recommendation for software information sitesabstractSoftware developers can search, share and learn development experience, solutions, bug fixes and open source projects in software information sites such as StackOverflow and Freecode. Many software information sites rely on tags to classify their contents, i.e. software objects, in order to improve the performance and accuracy of various operations on the sites. The quality of tags thus has a significant impact on the usefulness of these sites. High quality tags are expected to be concise and can describe the most important features of the software objects. Unfortunately tagging is inherently an uncoordinated process. The choice of tags made by individual software developers is dependent not only on a developer's understanding of the software object but also on the developer's English skills and preferences. As a result, the number of different tags grows rapidly along with continuous addition of software objects. With thousands of different tags, many of which introduce noise, software objects become poorly classified. Such phenomenon affects negatively the speed and accuracy of developers' queries. In this paper, we propose a tool called TagMulRec to automatically recommend tags and classify software objects in evolving large-scale software information sites. Given a new software object, TagMulRec locates the software objects that are semantically similar to the new one and exploit their tags. We have evaluated TagMulRec on four software information sites, StackOverflow, AskUbuntu, AskDifferent and Freecode. According to our empirical study, TagMulRec is not only accurate but also scalable that can handle a large-scale software information site with millions of software objects and thousands of tags. Pingyi Zhou, Jin Liu 0016, Zijiang Yang 0006, Guangyou Zhou |
SANER | 1 |
| 2016 | Automatically constructing course dependence graph based on association semantic link model
Pingyi Zhou, Jin Liu 0016, Xianzhao Yang, Xiaohui Cui, Liang Chang 0003, Shunxiang Zhang |
Pers. Ubiquitous Comput. | 1 |