Xueying Du

dblp:184/9459 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2024
0009-0005-0004-9183ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Evaluating Large Language Models in Class-Level Code Generation
abstract
Recently, many large language models (LLMs) have been proposed, showing advanced proficiency in code generation. Meanwhile, many efforts have been dedicated to evaluating LLMs on code generation benchmarks such as HumanEval. Although being very helpful for comparing different LLMs, existing evaluation focuses on a simple code generation scenario (i.e., function-level or statement-level code generation), which mainly asks LLMs to generate one single code unit (e.g., a function or a statement) for the given natural language description. Such evaluation focuses on generating independent and often small-scale code units, thus leaving it unclear how LLMs perform in real-world software development scenarios.
Xueying Du, Mingwei Liu 0002, Yixuan Chen 0012, Chaofeng Sha, Xin Peng 0001, Yiling Lou
ICSE1
2023 Knowledge Graph based Explainable Question Retrieval for Programming Tasks
abstract
Developers often seek solutions for their programming problems by retrieving existing questions on technical Q&A sites such as Stack Overflow. In many cases, they fail to find relevant questions due to the knowledge gap between the questions and the queries or feel it hard to choose the desired questions from the returned results due to the lack of explanations about the relevance. In this paper, we propose KGXQR, a knowledge graph based explainable question retrieval approach for programming tasks. It uses BERT-based sentence similarity to retrieve candidate Stack Overflow questions that are relevant to a given query. To bridge the knowledge gap and enhance the performance of question retrieval, it constructs a software development related concept knowledge graph and trains a question relevance prediction model to re-rank the candidate questions. The model is trained based on a combined sentence representation of BERT-based sentence embedding and graph-based concept embedding. To help understand the relevance of the returned Stack Overflow questions, KGXQR further generates explanations based on the association paths between the concepts involved in the query and the Stack Overflow questions. The evaluation shows that KGXQR outperforms the baselines in terms of accuracy, recall, MRR, and MAP and the generated explanations help the users to find the desired questions faster and more accurately.
Mingwei Liu 0002, Simin Yu, Xin Peng 0001, Xueying Du, Tianyong Yang, Huanjun Xu, Gaoyang Zhang
ICSME4
2023 CodeGen4Libs: A Two-Stage Approach for Library-Oriented Code Generation
abstract
Automated code generation has been extensively studied in recent literature. In this work, we first survey 66 participants to motivate a more pragmatic code generation scenario, i.e., library-oriented code generation, where the generated code should implement the functionally of the natural language query with the given library. We then revisit existing learning-based code generation techniques and find they have limited effectiveness in such a library-oriented code generation scenario. To address this limitation, we propose a novel library-oriented code generation technique, CodeGen4Libs, which incorporates two stages: import generation and code generation. The import generation stage generates import statements for the natural language query with the given third-party libraries, while the code generation stage generates concrete code based on the generated imports and the query. To evaluate the effectiveness of our approach, we conduct extensive experiments on a dataset of 403,780 data items. Our results demonstrate that CodeGen4Libs outperforms baseline models in both import generation and code generation stages, achieving improvements of up to 97.4% on EM (Exact Match), 54.5% on BLEU, and 53.5% on Hit@All. Overall, our proposed CodeGen4Libs approach shows promising results in generating high-quality code with specific third-party libraries, which can improve the efficiency and effectiveness of software development.
Mingwei Liu 0002, Tianyong Yang, Yiling Lou, Xueying Du, Xin Peng 0001
ASE4
2023 Recommending Analogical APIs via Knowledge Graph Embedding
abstract
Library migration, which replaces the current library with a different one to retain the same software behavior, is common in software evolution. An essential part of this is finding an analogous API for the desired functionality. However, due to the multitude of libraries/APIs, manually finding such an API is time-consuming and error-prone. Researchers created automated analogical API recommendation techniques, notably documentation-based methods. Despite potential, these methods have limitations, e.g., incomplete semantic understanding in documentation and scalability issues. In this study, we present KGE4AR, a novel documentation-based approach using knowledge graph (KG) embedding for recommending analogical APIs during library migration. KGE4AR introduces a unified API KG to comprehensively represent documentation knowledge, capturing high-level semantics. It further embeds this unified API KG into vectors for efficient, scalable similarity calculation. We assess KGE4AR with 35,773 Java libraries in two scenarios, with and without target libraries. KGE4AR notably outperforms state-of-the-art techniques (e.g., 47.1%-143.0% and 11.7%-80.6% MRR improvements), showcasing scalability with growing library counts.
Mingwei Liu 0002, Yiling Lou, Xin Peng 0001, Zhong Zhou, Xueying Du, Tianyong Yang
ESEC/SIGSOFT FSE6
2023 KG4CraSolver: Recommending Crash Solutions via Knowledge Graph
abstract
Fixing crashes is challenging, and developers often discuss their encountered crashes and refer to similar crashes and solutions on online Q&A forums (e.g., Stack Overflow). However, a crash often involves very complex contexts, which includes different contextual elements, e.g., purposes, environments, code, and crash traces. Existing crash solution recommendation or general solution recommendation techniques only use an incomplete context or treat the entire context as pure texts to search relevant solutions for a given crash, resulting in inaccurate recommendation results. In this work, we propose a novel crash solution knowledge graph (KG) to summarize the complete crash context and its solution with a graph-structured representation. To construct the crash solution KG automatically, we propose to leverage prompt learning to construct the KG from SO threads with a small set of labeled data. Based on the constructed KG, we further propose a novel KG-based crash solution recommendation technique KG4CraSolver, which precisely finds the relevant SO thread for an encountered crash by finely analyzing and matching the complete crash context based on the crash solution KG. The evaluation results show that the constructed KG is of high quality and KG4CraSolver outperforms baselines in terms of all metrics (e.g., 13.4%-113.4% MRR improvements). Moreover, we perform a user study and find that KG4CraSolver helps participants find crash solutions 34.4% faster and 63.3% more accurately.
Xueying Du, Yiling Lou, Mingwei Liu 0002, Xin Peng 0001, Tianyong Yang
ESEC/SIGSOFT FSE1
2021 Fabricate-Vanish: An Effective And Transferable Black-Box Adversarial Attack Incorporating Feature Distortion
abstract
Adversarial examples have emerged as increasingly severe threats for deep neural networks. Recent works have revealed that these malicious samples can transfer across different neural networks, and effectively attack other models. The state-of-the-art methodologies leverage Fast Gradient Sign Method to generate obstructing textures, which can cause neural networks to make incorrect inferences. However, the over-reliance on task-specific loss functions makes the adversarial examples less transferable across networks. Moreover, recent de-noising based adaptive defences provide promising performance against aforementioned attacks. Therefore, to achieve better transferability and attack effectiveness, we propose a novel attack, referred to as the Fabricate-Vanish (FV) attack, which is able to erase benign representations and generate obstruction textures simultaneously. The proposed FV attack treats the adversarial example transferability as latent contribution for each layer of deep neural networks, and maximizes the attack performance by balancing transferability and task specific loss function. Our experimental results on ImageNet show that the proposed FV attack achieves the best attack performance and better transferability by degrading the accuracy of classifiers 3.8% more on average compared to the state-of-the-art attacks.
Yantao Lu, Xueying Du, Bingkun Sun, Haining Ren, Senem Velipasalar
ICIP2
2021 Learning Diverse Policies in MOBA Games via Macro-Goals
abstract
Recently, many researchers have made successful progress in building the AI systems for MOBA-game-playing with deep reinforcement learning, such as on Dota 2 and Honor of Kings. Even though these AI systems have achieved or even exceeded human-level performance, they still suffer from the lack of policy diversity. In this paper, we propose a novel Macro-Goals Guided framework, called MGG, to learn diverse policies in MOBA games. MGG abstracts strategies as macro-goals from human demonstrations and trains a Meta-Controller to predict these macro-goals. To enhance policy diversity, MGG samples macro-goals from the Meta-Controller prediction and guides the training process towards these goals. Experimental results on the typical MOBA game Honor of Kings demonstrate that MGG can execute diverse policies in different matches and lineups, and also outperform the state-of-the-art methods over 102 heroes.
Yiming Gao 0007, Bei Shi, Xueying Du, Liang Wang 0015, Guangwei Chen, Zhenjie Lian, Fuhao Qiu, Guoan Han, Deheng Ye, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang
NeurIPS3
2019 Nonrigid Image Registration Using Spatially Region-Weighted Correlation Ratio and GPU-Acceleration
abstract
OBJECTIVE: Nonrigid image registration with high accuracy and efficiency remains a challenging task for medical image analysis. In this paper, we present the spatially region-weighted correlation ratio (SRWCR) as a novel similarity measure to improve the registration performance. METHODS: SRWCR is rigorously deduced from a three-dimension joint probability density function combining the intensity channels with an extra spatial information channel. SRWCR estimates the optimal functional dependence between the intensities for each spatial bin, in which the spatial distribution modeled by a cubic B-spline function is used to differentiate the contribution of voxels. We also analytically derive the gradient of SRWCR with respect to the transformation parameters and optimize it using a quasi-Newton approach. Furthermore, we propose a GPU-based parallel mechanism to accelerate the computation of SRWCR and its derivatives. RESULTS: The experiments on synthetic images, public four-dimensional thoracic computed tomography (CT) dataset, retinal optical coherence tomography data, and clinical CT and positron emission tomography images confirm that SRWCR significantly outperforms some state-of-the-art techniques such as spatially encoded mutual information and Robust PaTch-based cOrrelation Ration. CONCLUSION: This study demonstrates the advantages of SRWCR in tackling the practical difficulties due to distinct intensity changes, serious speckle noise, or different imaging modalities. SIGNIFICANCE: The proposed registration framework might be more reliable to correct the nonrigid deformations and more potential for clinical applications.
Lun Gong, Luwen Duan, Xueying Du, Hanqiu Liu, Xinjian Chen 0001, Jian Zheng 0001
IEEE J. Biomed. Health Informatics4
2016 Bayesian relevance feedback based Chinese calligraphy character synthesis
abstract
Sometimes calligraphy lovers want to generate a calligraphic plaque in style of some famous calligraphers, but some characters hadn't been written or were damaged in the long history of Chinese calligraphy. It will be a significant thing to use computer-aided synthesis technology to create calligraphic characters in the particular style. Though such kinds of research work have been done, the synthesized results are not so satisfying. So, in this paper, a novel calligraphy synthesis framework is designed to support the generation of calligraphic character in a particular style. Firstly, image relevance index is constructed based on relevance feedback. Secondly, a modified calligraphic components selection algorithm basing on Bayes classifier is proposed. Finally, experimental results are given, showing that our proposed approach is effective.
Xueying Du, Jiangqin Wu
ICME1