VLDB 2026 Research / reviewers in the wild / expert
Beijun Shen
dblp:83/2622
· DBLP profile ↗
77ranked-venue papers
0as first author
31since 2021 · last 2026
0000-0001-8370-3956ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 71 · 26 since 2021Artificial intelligence and machine learning · 16 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Anti-adversarial Learning: Desensitizing Prompts for Large Language ModelabstractWith the widespread use of LLMs, preserving privacy in user prompts has become crucial, as prompts risk exposing private and sensitive data to cloud LLMs. Conventional techniques like homomorphic encryption (HE), secure multi-party computation, and federated learning (FL) are not well-suited to this scenario due to the lack of control over user participation in remote model interactions. In this paper, we propose PromptObfus, a novel method for desensitizing LLM prompts. The core idea of PromptObfus is "anti-adversarial" learning, which perturbs sensitive words in the prompt to obscure private information while retaining the stability of model predictions. Specifically, PromptObfus frames prompt desensitization as a masked language modeling task, replacing privacy-sensitive terms with a [MASK] token. A desensitization model is utilized to generate candidate replacements for each masked position. These candidates are subsequently selected based on gradient feedback from a surrogate model, ensuring minimal disruption to the task output. We demonstrate the effectiveness of our approach on three NLP tasks. Results show that PromptObfus effectively prevents privacy inference from remote LLMs while preserving task performance. Xiaodong Gu 0002, Beijun Shen |
AAAI | 4 |
| 2026 | Synthetic Malware at Scale: Malicious Code Generation With Code TransplantingabstractMalicious code detection is one of the most essential tasks in safeguarding against security breaches, data compromise, and related threats. While machine learning has emerged as a predominant method for pattern detection, the training process is intricate due to the severe scarcity of malicious code samples. Consequently, machine learning detectors often encounter malicious patterns in limited and isolated scenarios, hindering their ability to generalize effectively across diverse threat landscapes. In this paper, we introduce MalCoder, a novel method for synthesizing malicious code samples. MalCoder enlarges the quantity and diversity of malicious instances by transplanting a set of malicious prototypes into a vast pool of benign code, thereby crafting a diverse array of malicious instances tailored to various application scenarios. For each malware prototype, MalCoder treats it as an incomplete code fragment and crafts its preceding and subsequent contexts through right-to-left and left-to-right code completion respectively. By leveraging GPTs with various sampling strategies, we can instantiate a large number of code samples bearing the malware prototype. Subsequently, MalCoder masks the original prototypes within the transplanted samples and fine-tunes an LLM code generator to reconstruct the original prototype. This process enables the model to seamlessly transplant malicious code fragments into benign code. During inference, MalCoder can automatically insert malicious fragments into benign samples at random positions, transforming benign code into malicious code. We apply MalCoder to a large pool of benign code in CodeSearchNet and craft over 50,000 malicious samples stemming from 39 malicious prototypes. Both qualitative and quantitative analyses show that the generated samples maintain key characteristics of malicious code while blending seamlessly with benign code, which helps in creating realistic and varied training data. Additionally, by using the generated samples as augmented training data, we witness a remarkable surge in malicious code detection capabilities. Specifically, the F1-score experiences a significant increase compared to utilizing only the original prototype samples. Guangzhan Wang, Diwei Chen, Xiaodong Gu 0002, Yuting Chen 0001, Beijun Shen |
IEEE Trans. Software Eng. | 5 |
| 2025 | ApiRAT: Integrating Multi-source API Knowledge for Enhanced Code Translation with LLMsabstractCode translation is an essential task in software migration, multilingual development, and system refactoring. Recent advancements in large language models (LLMs) have demonstrated significant potential in this task. However, prior studies have highlighted that LLMs often struggle with domain-specific code, particularly in resolving cross-lingual API mappings. To tackle this challenge, we propose ApiRAT, a novel code translation method that integrates multi-source API knowledge. ApiRAT employs three API knowledge augmentation techniques, including API sequence retrieval, API sequence back-translation, and API mapping, to guide LLMs to translating code, ensuring both the correct structure of API sequences and the accurate usage of individual APIs. Extensive experiments on two public datasets, CodeNet and AVATAR, indicate that ApiRAT significantly surpasses existing LLM-based methods, achieving improvements in computational accuracy ranging from 4% to 15.1%. Additionally, our evaluation across different LLMs showcases the generalizability of ApiRAT. An ablation study further confirms the individual contributions of each API knowledge component, underscoring the effectiveness of our approach. Guanjie Qiu, Xiaodong Gu 0002, Beijun Shen |
COMPSAC | 4 |
| 2025 | Empowering AI to Generate Better AI Code: Guided Generation of Deep Learning Projects with LLMsabstractWhile large language models (LLMs) have been widely applied to code generation, they struggle with generating entire deep learning projects, which are characterized by complex structures, longer functions, and stronger reliance on domain knowledge than general-purpose code. An open-domain LLM often lacks coherent contextual guidance and domain expertise for specific projects, making it challenging to produce complete code that fully meets user requirements. In this paper, we propose a novel planning-guided code generation method, DLCodeGen, tailored for generating deep learning projects. DLCodeGen predicts a structured solution plan, offering global guidance for LLMs to generate the project. The generated plan is then leveraged to retrieve semantically analogous code samples and subsequently abstract a code template. To effectively integrate these multiple retrieval-augmented techniques, a comparative learning mechanism is designed to generate the final code. We validate the effectiveness of our approach on a dataset we build for deep learning code generation. Experimental results demonstrate that DLCodeGen outperforms other baselines, achieving improvements of 9.7% in CodeBLEU and 3.6% in human evaluation metrics. Mingsheng Jiao, Xiaodong Gu 0002, Beijun Shen |
COMPSAC | 4 |
| 2025 | Transplant Then Regenerate: A New Paradigm for Text Data AugmentationabstractData augmentation is a critical technique in deep learning.Traditional methods like Backtranslation typically focus on lexical-level rephrasing, which primarily produces variations with the same semantics.While large language models (LLMs) have enhanced text augmentation by their "knowledge emergence" capability, controlling the style and structure of these outputs remains challenging and requires meticulous prompt engineering.In this paper, we propose LMTransplant, a novel text augmentation paradigm leveraging LLMs.The core idea of LMTransplant is transplant-thenregenerate: incorporating seed text into a context expanded by LLM, and asking the LLM to regenerate a variant based on the expanded context.This strategy allows the model to create more diverse and creative content-level variants by fully leveraging the knowledge embedded in LLMs, while preserving the core attributes of the original text.We evaluate LMTransplant across various text-related tasks, demonstrating its superior performance over existing text augmentation methods.Moreover, LMTransplant demonstrates exceptional scalability as the size of augmented data grows. Guangzhan Wang, Hongyu Zhang 0002, Beijun Shen, Xiaodong Gu 0002 |
EMNLP | 3 |
| 2025 | LongCodeZip: Compress Long Context for Code Language ModelsabstractCode generation under long contexts is becoming increasingly critical as Large Language Models (LLMs) are required to reason over extensive information in the code-base. While recent advances enable code LLMs to process long inputs, high API costs and generation latency remain substantial bottlenecks. Existing context pruning techniques, such as LLMLingua, achieve promising results for general text but overlook code-specific structures and dependencies, leading to suboptimal performance in programming tasks. In this paper, we propose LongCodeZip, a novel plug-and-play code compression framework designed specifically for code LLMs. LongCodeZip employs a dual-stage strategy: (1) coarse-grained compression, which identifies and ranks function-level chunks using conditional perplexity with respect to the instruction, retaining only the most relevant functions; and (2) fine-grained compression, which segments retained functions into blocks based on perplexity and selects an optimal subset under an adaptive token budget to maximize relevance. Evaluations across multiple tasks, including code completion, summarization, and question answering, show that LongCodeZip consistently outperforms baseline methods, achieving up to a 5.6× compression ratio without degrading task performance. By effectively reducing context size while preserving essential information, LongCodeZip enables LLMs to better scale to real-world, large-scale code scenarios, advancing the efficiency and capability of code intelligence applications1. Yuling Shi, Yichun Qian, Hongyu Zhang 0002, Beijun Shen, Xiaodong Gu 0002 |
ASE | 4 |
| 2025 | Just-in-time software defect prediction via bi-modal change representation learning
Yuze Jiang, Beijun Shen, Xiaodong Gu 0002 |
J. Syst. Softw. | 2 |
| 2024 | Unraveling the Potential of Large Language Models in Code Translation: How Far are We?abstractWhile large language models (LLMs) exhibit state-of-the-art performance in various tasks, recent studies have revealed their struggle for code translation. This is because they haven't been extensively pre-trained with parallel multilingual code, which code translation heavily depends on. Moreover, existing benchmarks only cover a limited subset of common programming languages, and thus cannot reflect the full potential of LLMs in code translation. In this paper, we conduct a large-scale empirical study to exploit the capabilities and incapabilities of LLMs in code translation tasks. We first craft a novel benchmark called PolyHumanEval by extending HumanEval to a multilingual benchmark of 14 languages. With PolyHumanEval, we then perform over 110,000 translations with bleeding-edge code LLMs. The result shows LLMs' suboptimal performance on Python to other languages and the negligible impact of widely adopted LLM optimization techniques such as conventional pre-training and instruction tuning on code translation. To further uncover the potential of LLMs in code translation, we propose two methods: (1) intermediary translation which selects an intermediary language between the source and target ones; and (2) self-training which fine-tunes LLMs on self-generated parallel data. Evaluated with CodeLlama-13B, our approach yields an average improvement of 11.7% computation accuracy on Python-to-other translations. Notably, we interestingly find that Go can serve as a lingua franca for translating between any two studied languages. Qingxiao Tao, Tingrui Yu, Xiaodong Gu 0002, Beijun Shen |
APSEC | 4 |
| 2024 | Few-shot code translation via task-adapted prompt learning
Xiaodong Gu 0002, Yuting Chen 0001, Beijun Shen |
J. Syst. Softw. | 5 |
| 2024 | Project-specific code summarization with in-context learning
Shangbo Yun, Shuhuai Lin, Xiaodong Gu 0002, Beijun Shen |
J. Syst. Softw. | 4 |
| 2023 | Lejacon: A Lightweight and Efficient Approach to Java Confidential Computing on SGXabstractIntel's SGX is a confidential computing technique. It allows key functionalities of C/C++/native applications to be confidentially executed in hardware enclaves. However, numerous cloud applications are written in Java. For supporting their confidential computing, state-of-the-art approaches deploy Java Virtual Machines (JVMs) in enclaves and perform confidential computing on JVMs. Meanwhile, these JVM-in-enclave solutions still suffer from serious limitations, such as heavy overheads of running JVMs in enclaves, large attack surfaces, and deep computation stacks. To mitigate the above limitations, we for-malize a Secure Closed-World (SCW) principle and then propose Lejacon, a lightweight and efficient approach to Java confidential computing. The key idea is, given a Java application, to (1) separately compile its confidential computing tasks into a bundle of Native Confidential Computing (NCC) services; (2) run the NCC services in enclaves on the Trusted Execution Environment (TEE) side, and meanwhile run the non-confidential code on a JVM on the Rich Execution Environment (REE) side. The two sides interact with each other, protecting confidential computing tasks and as well keeping the Trusted Computing Base (TCB) size small. We implement Lejacon and evaluate it against OcclumJ (a state-of-the-art JVM-in-enclave solution) on a set of benchmarks using the BouncyCastle cryptography library. The evaluation results clearly show the strengths of Lejacon: it achieves compet-itive performance in running Java confidential code in enclaves; compared with OcclumJ, Lejacon achieves speedups by up to 16.2x in running confidential code and also reduces the TCB sizes by 90+% on average. Xinyuan Miao, Sanhong Li, Pengbo Nie, Yuting Chen 0001, Beijun Shen, He Jiang 0001 |
ICSE | 9 |
| 2023 | On the Evaluation of Neural Code Translation: Taxonomy and BenchmarkabstractIn recent years, neural code translation has gained increasing attention. While most of the research focuses on improving model architectures and training processes, we notice that the evaluation process and benchmark for code translation models are severely limited: they primarily treat source code as natural languages and provide a holistic accuracy score while disregarding the full spectrum of model capabilities across different translation types and complexity. In this paper, we present a comprehensive investigation of four state-of-the-art models and analyze in-depth the advantages and limitations of three existing benchmarks. Based on the empirical results, we develop a taxonomy that categorizes code translation tasks into four primary types according to their complexity and knowledge dependence: token level (type 1), syntactic level (type 2), library level (type 3), and algorithm level (type 4). We then conduct a thorough analysis of how existing approaches perform across these four categories. Our findings indicate that while state-of-the-art code translation models excel in type-1 and type-2 translations, they struggle with knowledge-dependent ones such as type-3 and type-4. Existing benchmarks are biased towards trivial translations, such as keyword mapping. To overcome these limitations, we construct G-TransEval, a new benchmark by manually curating type-3 and type-4 translation pairs and unit test cases. Results on our new benchmark suggest that G-TransEval can exhibit more comprehensive and finer-grained capability of code translation models and thus provide a more rigorous evaluation. Our studies also provide more insightful findings and suggestions for future research, such as building type-3 and type-4 training data and ensembling multiple pretraining approaches. Mingsheng Jiao, Tingrui Yu, Guanjie Qiu, Xiaodong Gu 0002, Beijun Shen |
ASE | 6 |
| 2023 | InfeRE: Step-by-Step Regex Generation via Chain of InferenceabstractAutomatically generating regular expressions (abbrev. regexes) from natural language description (NL2RE) has been an emerging research area. Prior studies treat regex as a linear sequence of tokens and generate the final expressions autoregressively in a single pass. They did not take into account the step-by-step internal text-matching processes behind the final results. This significantly hinders the efficacy and interpretability of regex generation by neural language models. In this paper, we propose a new paradigm called InfeRE, which decomposes the generation of regexes into chains of step-bystep inference. To enhance the robustness, we introduce a self-consistency decoding mechanism that ensembles multiple outputs sampled from different models. We evaluate InfeRE on two publicly available datasets, NL-RX-Turk and KB13, and compare the results with state-of-the-art approaches and the popular tree-based generation approach TRANX. Experimental results show that InfeRE substantially outperforms previous baselines, yielding 16.3% and 14.7% improvement in DFA@5 accuracy on two datasets, respectively. Xiaodong Gu 0002, Yuting Chen 0001, Beijun Shen |
ASE | 4 |
| 2023 | DuReSE: Rewriting Incomplete Utterances via Neural Sequence Editing
Xiaodong Gu 0002, Yuting Chen 0001, Beijun Shen |
Neural Process. Lett. | 4 |
| 2022 | Code Question Answering via Task-Adaptive Sequence-to-Sequence Pre-trainingabstractThe development of a question answering (QA) system for code can greatly facilitate programs understanding for developers. Recently, pre-trained language models (PLMs) have shown promising results in the code QA task. However, directly applying PLMs to code QA often causes suboptimal performance due to the large discrepancy between pre-training and the downstream QA task. While code PLMs are pre-trained on largescale unlabeled code corpora, there is often a scarce availability of annotated QA pairs for fine-tuning. Existing code PLMs simply reuse the code representation part and require to train the QA part from scratch, which causes the model to overfit QA data. In this paper, we propose CodeMaster, a novel pre-training based approach for automatically answering code questions via task adaptation. CodeMaster employs CodeT5, a popular PLM for source code. In order to mitigate the gap between pretraining and QA, CodeMaster continually pre-trains CodeT5 on multiple self-supervised learning tasks such as partial comment completion and noun-phrase prediction. Experimental results on the CodeQA benchmark show that CodeMaster achieves state-of-the-art performance, and highlight the effectiveness of our approach. Tingrui Yu, Xiaodong Gu 0002, Beijun Shen |
APSEC | 3 |
| 2022 | Cross-Domain Deep Code Search with Meta LearningabstractRecently, pre-trained programming language models such as CodeBERT have demonstrated substantial gains in code search. Despite their success, they rely on the availability of large amounts of parallel data to fine-tune the semantic mappings between queries and code. This restricts their practicality in domain-specific languages with relatively scarce and expensive data. In this paper, we propose CDCS, a novel approach for domain-specific code search. CDCS employs a transfer learning framework where an initial program representation model is pre-trained on a large corpus of common programming languages (such as Java and Python), and is further adapted to domain-specific languages such as Solidity and SQL. Unlike cross-language CodeBERT, which is directly fine-tuned in the target language, CDCS adapts a few-shot meta-learning algorithm called MAML to learn the good initialization of model parameters, which can be best reused in a domain-specific language. We evaluate the proposed approach on two domain-specific languages, namely Solidity and SQL, with model transferred from two widely used languages (Python and Java). Experimental results show that CDCS significantly outperforms conventional pre-trained code models that are directly fine-tuned in domain-specific languages, and it is particularly effective for scarce data. Yitian Chai, Hongyu Zhang 0002, Beijun Shen, Xiaodong Gu 0002 |
ICSE | 3 |
| 2022 | Zero-shot program representation learningabstractLearning program representations has been the core prerequisite of code intelligence tasks (e.g., code search and code clone detection). The state-of-the-art pre-trained models such as CodeBERT require the availability of large-scale code corpora. However, gathering training samples can be costly and infeasible for domain-specific languages such as Solidity for smart contracts. In this paper, we propose Zecoler, a zero-shot learning approach for code representations. Zecoler is built upon a pre-trained programming language model. In order to elicit knowledge from the pre-trained models efficiently, Zecoler casts the downstream tasks to the same form of pre-training tasks by inserting trainable prompts into the original input. Then, it employs the prompt learning technique to optimize the pre-trained model by merely adjusting the original input. This enables the representation model to efficiently fit the scarce task-specific data while reusing pre-trained knowledge. We evaluate Zecoler in three code intelligence tasks in two programming languages that have no training samples, namely, Solidity and Go, with model trained in corpora of common languages such as Java. Experimental results show that our approach significantly outperforms baseline models in both zero-shot and few-shot settings. Nan Cui, Yuze Jiang, Xiaodong Gu 0002, Beijun Shen |
ICPC | 4 |
| 2022 | Self-supervised learning of smart contract representationsabstractLearning smart contract representations can greatly facilitate the development of smart contracts in many tasks such as bug detection and clone detection. Existing approaches for learning program representations are difficult to apply to smart contracts which have insufficient data and significant homogenization. To overcome these challenges, in this paper, we propose SRCL, a novel, self-supervised approach for learning smart contract representations. Unlike existing supervised methods, which are tied on task-specific data labels, SRCL leverages large-scale unlabeled data by self-supervised learning of both local and global information of smart contracts. It automatically extracts structural sequences from abstract syntax trees (ASTs). Then, two discriminators are designed to guide the Transformer encoder to learn local and global semantic features of smart contracts. We evaluate SRCL on a dataset of 75,006 smart contracts collected from Etherscan. Experimental results show that SRCL considerably outperforms the state-of-the-art code representation models on three downstream tasks. Shouliang Yang, Xiaodong Gu 0002, Beijun Shen |
ICPC | 3 |
| 2022 | Answering Software Deployment Questions via Neural Machine Reading at ScaleabstractAs software systems continue to grow in complexity and scale, deploying and delivering them becomes increasingly difficult. In this work, we develop DeployQA, a novel QA bot that automatically answers software deployment questions over user manuals and Stack Overflow posts. DeployQA is built upon RoBERTa. To bridge the gap between natural language and the domain of software deployment, we propose three adaptations in terms of vocabulary, pre-training, and fine-tuning, respectively. We evaluate our approach on our constructed DeQuAD dataset. The results show that DeployQA remarkably outperforms baseline methods by leveraging the three domain adaptation strategies. Guanjie Qiu, Diwei Chen, Yitian Chai, Xiaodong Gu 0002, Beijun Shen |
ASE | 6 |
| 2022 | Diet code is healthy: simplifying programs for pre-trained models of codeabstractPre-trained code representation models such as CodeBERT have demonstrated superior performance in a variety of software engineering tasks, yet they are often heavy in complexity, quadratically with the length of the input sequence. Our empirical analysis of CodeBERT's attention reveals that CodeBERT pays more attention to certain types of tokens and statements such as keywords and data-relevant statements. Based on these findings, we propose DietCode, which aims at lightweight leverage of large pre-trained models for source code. DietCode simplifies the input program of CodeBERT with three strategies, namely, word dropout, frequency filtering, and an attention-based strategy that selects statements and tokens that receive the most attention weights during pre-training. Hence, it gives a substantial reduction in the computational cost without hampering the model performance. Experimental results on two downstream tasks show that DietCode provides comparable results to CodeBERT with 40% less computational cost in fine-tuning and testing. Hongyu Zhang 0002, Beijun Shen, Xiaodong Gu 0002 |
ESEC/SIGSOFT FSE | 3 |
| 2022 | Clean and Learn: Improving Robustness to Spurious Solutions in API Question AnsweringabstractThe development of a question answering (QA) system for application programming interface (API) documentation can greatly facilitate developers in API-related tasks. However, when applying deep learning technology, API QA systems suffer from the spurious solution problem. That is, the answer can literally appear in multiple positions (i.e. start-end indices) in the API documentation, though only one of them (called golden solution) correctly solves the question given its context. The other incorrect candidates (called spurious solutions) hinder the neural network model to learn reasonable solutions or correct answers. In this work, we propose Clean-and-Learn, an effective and robust method for API QA over documents. In order to reduce the spuriousness of candidate solutions used for training, we design several scoring functions to rank the candidate occurrences (clean). Only high-quality (top-[Formula: see text]) candidate solutions are involved in training. Then, we perform multi-task learning by weighing the losses computed from the top-k occurrences (learn). We evaluate our method on the constructed APIQASet dataset. The experiment results show that Clean-and-Learn achieves a ROUGE-L score of 75.8 and accuracy of 70.5% in API QA, which significantly outperforms state-of-the-art approaches. Haozhe Qin, Xiaodong Gu 0002, Beijun Shen |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2022 | Automatically repairing tensor shape faults in deep learning programs
Dangwei Wu, Beijun Shen, Yuting Chen 0001, He Jiang 0001, Lei Qiao 0002 |
Inf. Softw. Technol. | 2 |
| 2021 | Mining API Constraints from Library and Client to Detect API MisusesabstractCalling Application Programming Interfaces (APIs) shall follow various constraints (e.g., call orders). If these con-straints are violated, API misuses are introduced to code, and such misuses can cause severe bugs. To effectively detect API misuses, most prior approaches mine constraints from client code, and assume that the violations of constraints are potential misuses. However, as client code only illustrates a small portion of API usages, constraints mined from client code are typically incomplete. As a result, when mined constraints are used to detect bugs, many violations of constraints turn out to be false positives. In this paper, our research purpose is to find more misuses and to reduce false positives. As library code contains many details on APIs, we propose an approach that mines API constraints from both client and library code. From client code, our approach builds API usage graphs and uses a frequent subgraph mining algorithm to mine frequent usage patterns as API constraints. From library code, our approach derives various types of constraints with our predefined strategies. With constraints from both sources, our graph matching algorithm can detect API misuses. As a result, our approach takes advantage from both the comprehensiveness and informativeness of library-based constraints and the accuracy of client-based patterns. We compared our approach with MuDetect on the MuBench dataset. Our results show that it significantly improves the detection effectiveness of MuBench from 39.5% to 50.2% of the recall, and from 30.6% to 41.7% of the precision. Hushuang Zeng, Jingxin Chen, Beijun Shen, Hao Zhong 0001 |
APSEC | 3 |
| 2021 | Learning to Match Workers and Tasks via a Multi-View Graph Attention NetworkabstractThe worker-task matching problem brings up unique characteristics that are not present in traditional matching scenarios, i.e., the huge flow of tasks with short lifespans, the importance of workers’ capabilities, and the quality of the completed tasks. These characteristics further pose significant challenges of data sparsity and comprehensive modeling.To address the two challenges, this paper proposes MvkGAN, a multi-view attention network on a bi-collaborative knowledge graph (BicKG). The core ideas of our work are 1) building BicKG from the data of workers, tasks, their interactions, and domain knowledge, and then leveraging it to reveal the latent interactions between workers and tasks to mitigate the data sparsity challenge; and 2) designing a multi-view knowledge graph attention network (MvkGAN) which learns to match workers and tasks, to meet the comprehensive modeling challenge. In this network, different features are organized as multiple views and these views are further connected by the attention mechanism.We have implemented MvkGAN and evaluated it against five state-of-the-art approaches (Wide&Deep, DeepFM, KGAT, Crow-dRex and PJFNN) on two real-world datasets. The evaluation results show that MvkGAN improves the accuracy by 6.90% and F1-score by 5.52% on average, and also has the ability of generating reasonable explanations. Nan Cui, Chunqi Chen, Beijun Shen, Yuting Chen 0001 |
COMPSAC | 3 |
| 2021 | ApproxiFuzzer: Fuzzing towards Deep Code Snippets in Java ProgramsabstractA real-world, complex software system can contain a number of code snippets. Many snippets are deep, surrounded by complicated triggering conditions and/or hidden in functions less frequently invoked. Fuzzing and symbolic execution are two mainstreams for exploring input spaces and increasing code coverage of complicated software systems. Meanwhile, it remains a challenge to determine whether a deep code snippet is reachable, and if it is reachable, which test(s) can reach it.This paper presents ApproxiFuzzer, an effective, demand-driven approach to fuzzing towards deep code snippets in Java programs. Given a program P, a target deep code snippet tcs, and a set of seeding test inputs, the key idea behind ApproxiFuzzer is to selectively mutate the test inputs and collect their execution traces such that the execution traces gradually approximate tcs; several measures are designed for measuring the distances between execution traces and the code snippet and directing the fuzzing process towards generating test inputs reaching tcs.We have implemented ApproxiFuzzer and evaluated it against Kelinci (an AFL-based fuzzer) and JDart (a concolic execution tool) on a set of real-world benchmarks. The evaluation clearly demonstrates the strengths of ApproxiFuzzer—ApproxiFuzzer outperforms Kelinci by 36× in efficiently generating test inputs, obtaining up to 18.2% higher code coverage; ApproxiFuzzer also outperforms JDart by 46.2∼96.2% in hitting deep code snippets. Xintian Yu, Enze Ma, Pengbo Nie, Beijun Shen, Yuting Chen 0001 |
COMPSAC | 4 |
| 2021 | Tensfa: Detecting and Repairing Tensor Shape Faults in Deep Learning SystemsabstractSoftware developers frequently invoke deep learning (DL) APIs to incorporate learning solutions into software systems. However, misuses of these APIs can cause various DL faults, such as tensor shape faults. Tensor shape faults occur when restriction conditions of operations are not met; they are prevalent in practice, leading to many system crashes. Meanwhile, researchers and engineers still face a strong challenge in detecting tensor shape faults ─ static techniques incur heavy overheads in defining detection rules, and the only dynamic technique requires human engineers to rewrite APIs for tracking shape changes. To address the above challenge, we conduct a deep empirical study on crashing tensor shape faults (i.e., those causing programs to crash), categorizing them into four types and revealing twelve repair patterns. We then propose and implement Tensfa, an approach to detecting and repairing crashing tensor shape faults. Tensfa takes a machine learning method to learn from crash messages and employs decision trees in detecting tensor shape faults. Tensfa also provides the first automated solution to repairing the detected faults: it tracks shape properties by a customized Python debugger, analyzes their data dependences, and uses the twelve patterns to generate patches. We construct SFData, a set of 146 buggy programs with crashing tensor shape faults. Our Tensfa has been implemented and evaluated on SFData and IslamData (another dataset of tensor shape faults). The results clearly show the effectiveness of Tensfa. In particular, Tensfa achieves the state-of-the-art results: it reaches an F1-score of 96.88% in detecting the faults and repairs 80 out of 146 buggy programs in SFData. Dangwei Wu, Beijun Shen, Yuting Chen 0001, He Jiang 0001, Lei Qiao 0002 |
ISSRE | 2 |
| 2021 | Locating Faulty Methods with a Mixed RNN and Attention ModelabstractIR-based fault localization approaches achieves promising results when locating faulty files by comparing a bug report with source code. Unfortunately, they become less effective to locate faulty methods. We conduct a preliminary study to explore its challenges, and identify three problems: the semantic gap problem, the representation sparseness problem, and the single revision problem.To tackle these problems, we propose MRAM, a mixed RNN and attention model, which combines bug-fixing features and method structured features to explore both implicit and explicit relevance between methods and bug reports for method level fault localization task. The core ideas of our model are: (1) constructing code revision graphs from code, commits and past bug reports, which reveal the latent relations among methods to augment short methods and as well provide all revisions of code and past fixes to train more accurate models; (2) embedding three method structured features (token sequences, API invocation sequences, and comments) jointly with RNN and soft attention to represent source methods and obtain their implicit relevance with bug reports; and (3) integrating multi-revision bug-fixing features, which provide the explicit relevance between bug reports and methods, to improve the performance.We have implemented MRAM and conducted a controlled experiment on five open-source projects. Comparing with state-of-the-art approaches, our MRAM improves MRR values by 3.8-5.1% (3.7-5.4%) when the dataset contains (does not contain) localized bug reports. Our statistics test shows that our improvements are significant. Shouliang Yang, Junming Cao, Hushuang Zeng, Beijun Shen, Hao Zhong 0001 |
ICPC | 4 |
| 2021 | ConLAR: Learning to Allocate Resources to Docker Containers under Time-Varying WorkloadsabstractCloud platforms are increasingly using containers for lightweight virtualization. However, the mainstream operating systems are currently limited in their capabilities in customizing containers’ resource management. There remains two main challenges in resource allocations. First, the application workloads can be time-varying, leading to a problem of resource over- or under-allocations. Second, it becomes difficult to minimize the resource provisioning cost while guaranteeing the SLO (service level objective). To address these challenges, we propose ConLAR, a learning-to-allocate approach that predicts and allocates resources to Docker containers under time-varying workloads. ConLAR efficiently reduces over-provisioning cost with the SLO guarantees by taking an Observing-Predicting-Allocating-Executing paradigm: given a container, it observes the running of the online application and its environment, leverages the LSTM (long short term memory) model to predict its future workload, adaptively learns to construct resource allocation strategies with two objectives through RL (reinforcement learning), and executes them to scale container resources dynamically. We have evaluated ConLAR on two real-world workloads of ClarkNet and GoogleClusterData. The results clearly show the effectiveness of ConLAR. In particular, ConLAR achieves a resource over-provisioning cost of less than 16.5% and an SLO violations rate of 8.9%; it also shows good flexibility to learn different adaption policies. Diwei Chen, Beijun Shen, Yuting Chen 0001 |
QRS | 2 |
| 2021 | Weakly-Supervised Question Answering with Effective Rank and Weighted Loss over CandidatesabstractWe study the weakly supervised question answering problem. Weak-ly supervised question answering aims to learn how the questions should be answered directly from the pairs without golden solutions/evidences, which makes question answering models much easier to scale to many domains. However, in weak supervision setup, a question typically involves many candidate solutions and the spuriousness of candidate solutions will hurt the performance of the question answering models. In this paper, we present an effective method to learn a question answering model in a weak supervision way. Specifically, in order to reduce the spuriousness of candidate solutions used for training, we design several simple yet effective scoring functions to rank the candidate solutions. Despite its simplicity, this ranking process can improve the quality of the training data significantly with fewer spurious candidates left. Then, different from previous approaches that either treat all candidates equally for training or only select the candidate with the largest likelihood in each iteration, we formulate this problem as a multi-task learning problem by weighing the losses computed from top-k candidates. Experimental results show that, our method1 can outperform previous approaches on both semantic parsing and machine reading comprehension tasks. Haozhe Qin, Jiangang Zhu, Beijun Shen |
WWW | 3 |
| 2021 | Context-Aware Conversational Recommendation of Trigger-Action Rules in IoT ProgrammingabstractTrigger-action (TA) programming is a programming paradigm that allows end-users to automate and connect IoT devices and online services using if-trigger-then-action rules. Early studies have demonstrated this paradigms usability, but more recent work has also highlighted complexities that arise in realistic scenarios. To facilitate end-users in TA programming, we propose AutoTAR, a context-aware conversational recommendation technique for recommending TA rules. AutoTAR leverages a TA knowledge graph to encode semantic features and abstract functionalities of rules, and then takes a two-phase method to recommend TA rules to end-users: during the context-aware recommendation phase, it elicits user preferences from programming context and recommends the top-N rules using a mixed content and collaborative technique; during the conversational recommendation phase, it justifies recommendations by iteratively raising questions and collecting feedback from end-users. We evaluate AutoTAR on Mturk and real data collected from the IFTTT community. The results show that our method outperforms state-of-the-arts significantly — its context-aware recommendation outperforms RecRules by 26% on R@5 and 21% on NDCG@5; its conversational recommendation outperforms LARecommender (a conversational recommender with the LA model) by 67.64% on accuracy. In addition, AutoTAR is effective in solving three problems frequently occurring in TA rule recommendations, i.e., the cold-start problem, the repeat-consumption problem, and the incomplete-intent problem. Mingxin Zhao, Qinyue Wu, Enze Ma, Beijun Shen, Yuting Chen 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2021 | DeFiHap: Detecting and Fixing HiveQL Anti-PatternsabstractThe emergence of Hive greatly facilitates the management of massive data stored in various places. Meanwhile, data scientists face challenges during HiveQL programming - they may not use correct and/or efficient HiveQL statements in their programs; developers may also introduce anti-patterns indeliberately into HiveQL programs, leading to poor performance, low maintainability, and/or program crashes. This paper presents an empirical study on HiveQL programming, in which 38 HiveQL anti-patterns are revealed. We then design and implement DeFiHap, the first tool for automatically detecting and fixing HiveQL anti-patterns. DeFiHap detects HiveQL anti-patterns via analyzing the abstract syntax trees of HiveQL statements and Hive configurations, and generates fix suggestions by rule-based rewriting and performance tuning techniques. The experimental results show that DeFiHap is effective. In particular, DeFiHap detects 25 anti-patterns and generates fix suggestions for 17 of them. Yuetian Mao, Nan Cui, Tianjiao Du, Beijun Shen, Yuting Chen 0001 |
Proc. VLDB Endow. | 5 |
| 2020 | Learning Code-Query Interaction for Enhancing Code SearchesabstractCode search plays an important role in software development and maintenance. In recent years, deep learning (DL) has achieved a great success in this domain-several DL-based code search methods, such as DeepCS and UNIF, have been proposed for exploring deep, semantic correlations between code and queries; each method usually embeds source code and natural language queries into real vectors followed by computing their vector distances representing their semantic correlations. Meanwhile, deep learning-based code search still suffers from three main problems, i.e., the OOV (Out of Vocabulary) problem, the independent similarity matching problem, and the small training dataset problem. To tackle the above problems, we propose CQIL, a novel, deep learning-based code search method. CQIL learns code-query interactions and uses a CNN (Convolutional Neural Network) to compute semantic correlations between queries and code snippets. In particular, CQIL employs a hybrid representation to model code-query correlations, which solves the OOV problem. CQIL also deeply learns the code-query interaction for enhancing code searches, which solves the independent similarity matching and the small training dataset problems. We evaluate CQIL on two datasets (CODEnn and CosBench). The evaluation results show the strengths of CQIL-it achieves the MAP@1 values, 0.694 and 0.574, on CODEnn and CosBench, respectively. In particular, it outperforms DeepCS and UNIF, two state-of-the-art code search methods, by 13.6% and 18.1% in MRR, respectively, when the training dataset is insufficient. Wei Li 0254, Haozhe Qin, Shuhan Yan, Beijun Shen, Yuting Chen 0001 |
ICSME | 4 |
| 2020 | Learning to Recommend Trigger-Action Rules for End-User Development - A Knowledge Graph Based Approach
Qinyue Wu, Beijun Shen, Yuting Chen 0001 |
ICSR | 2 |
| 2020 | BugPecker: Locating Faulty Methods with Deep Learning on Revision GraphsabstractGiven a bug report of a project, the task of locating the faults of the bug report is called fault localization. To help programmers in the fault localization process, many approaches have been proposed, and have achieved promising results to locate faulty files. However, it is still challenging to locate faulty methods, because many methods are short and do not have sufficient details to determine whether they are faulty. In this paper, we present BugPecker, a novel approach to locate faulty methods based on its deep learning on revision graphs. Its key idea includes (1) building revision graphs and capturing the details of past fixes as much as possible, and (2) discovering relations inside our revision graphs to expand the details for methods and calculating various features to assist our ranking. We have implemented BugPecker, and evaluated it on three open source projects. The early results show that BugPecker achieves a mean average precision (MAP) of 0.263 and mean reciprocal rank (MRR) of 0.291, which improve the prior approaches significantly. For example, BugPecker improves the MAP values of all three projects by five times, compared with two recent approaches such as DNNLoc-m and BLIA 1.5. Junming Cao, Shouliang Yang, Hushuang Zeng, Beijun Shen, Hao Zhong 0001 |
ASE | 5 |
| 2020 | Are the Code Snippets What We Are Searching for? A Benchmark and an Empirical Study on Code Search with Natural-Language QueriesabstractCode search methods, especially those that allow programmers to raise queries in a natural language, plays an important role in software development. It helps to improve programmers' productivity by returning sample code snippets from the Internet and/or source-code repositories for their natural-language queries. Meanwhile, there are many code search methods in the literature that support natural-language queries. Difficulties exist in recognizing the strengths and weaknesses of each method and choosing the right one for different usage scenarios, because (1) the implementations of those methods and the datasets for evaluating them are usually not publicly available, and (2) some methods leverage different training datasets or auxiliary data sources and thus their effectiveness cannot be fairly measured and may be negatively affected in practical uses. To build a common ground for measuring code search methods, this paper builds CosBench, a dataset that consists of 1000 projects, 52 code-independent natural-language queries with ground truths, and a set of scripts for calculating four metrics on code research results. We have evaluated four IR (Information Retrieval)-based and two DL (Deep Learning)-based code search methods on CosBench. The empirical evaluation results clearly show the usefulness of the CosBench dataset and various strengths of each code search method. We found that DL-based methods are more suitable for queries on reusing code, and IR-based ones for queries on resolving bugs and learning API uses. Shuhan Yan, Yuting Chen 0001, Beijun Shen, Lingxiao Jiang |
SANER | 4 |
| 2020 | Semantic Service Search in IT Crowdsourcing Platform: A Knowledge Graph-Based ApproachabstractUnderstanding user’s search intent in vertical websites like IT service crowdsourcing platform relies heavily on domain knowledge. Meanwhile, searching for services accurately on crowdsourcing platforms is still difficult, because these platforms do not contain enough information to support high-performance search. To solve these problems, we build and leverage a knowledge graph named ITServiceKG to enhance search performance of crowdsourcing IT services. The main ideas are to (1) build an IT service knowledge graph from Wikipedia, Baidupedia, CN-DBpedia, StuQ and data in IT service crowdsourcing platforms, (2) use properties and relations of entities in the knowledge graph to expand user query and service information, and (3) apply a listwise approach with relevance features and topic features to re-rank the search results. The results of our experiments indicate that our approach outperforms the traditional search approaches. Qinyue Wu, Duankang Fu, Beijun Shen, Yuting Chen 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2020 | Feedback2Code: A Deep Learning Approach to Identifying User-Feedback-Related Source Code FilesabstractUsers frequently raise feedback when using software products. Feedback from users regarding their experiences and expectations and software defects they found adds values to software maintenance and evolution — software managers collect user feedback and then dispatch feedback issues that developers (and/or maintainers) need to track and process. Feedback tracking is often supported by open source platforms and collaborative software systems. Meanwhile, there still exists a gap between feedback issues and source code: since user feedback is usually informal and arbitrary, engineers have to spend much effort on comprehending issues and identifying which source code files need to be improved or fixed. This paper introduces a deep learning approach, Feedback2Code , which facilitates identification of user-feedback-related source code files. The core idea is to (1) explore latent semantics of user feedback and source code using several deep learning techniques such as Multi-Layer Perceptron (MLP), Convolutional Neutral Network (CNN) and skip-gram and (2) establish a multi-correlation model to explore linkages between feedback issues and source code files. Given a feedback issue, the linkages then allow engineers to identify source code files that are highly relevant to the issue. We have implemented Feedback2Code and evaluated it against ChangeAdvisor (a state-of-the-art approach) on 24 open source projects. The evaluation results clearly show the strength of Feedback2Code : for 103793 feedback issues, Feedback2Code successfully established 101190 feedback-code linkages and achieved a precision that is [Formula: see text] higher than that of ChangeAdvisor . Feedback2Code also achieved an MRR and an MAP that are [Formula: see text] and [Formula: see text] higher than those of ChangeAdvisor , respectively. Furthermore, we also found that a Feedback2Code -trained model can be easily transferred, allowing feedback-code linkages to be established in new projects with a little history data. Shuhan Yan, Tianjiao Du, Beijun Shen, Yuting Chen 0001, Zhilei Ren |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2019 | Reinforcement Learning of Code Search SessionsabstractSearching and reusing online code is a common activity in software development. Meanwhile, like many general-purposed searches, code search also faces the session search problem: in a code search session, the user needs to iteratively search for code snippets, exploring new code snippets that meet his/her needs and/or making some results highly ranked. This paper presents Cosoch, a reinforcement learning approach to session search of code documents (code snippets with textual explanations). Cosoch is aimed at generating a session that reveals user intentions, and correspondingly searching and reranking the resulting documents. More specifically, Cosoch casts a code search session into a Markov decision process, in which rewards measuring the relevances between the queries and the resulting code documents guide the whole session search. We have built a dataset, say CosoBe, from StackOverflow, containing 103 code search sessions with 378 pieces of user feedback. We have also evaluated Cosoch on CosoBe. The evaluation results show that Cosoch achieves an average NDCG@3 score of 0.7379, outperforming StackOverflow by 21.3%. Wei Li 0254, Shuhan Yan, Beijun Shen, Yuting Chen 0001 |
APSEC | 3 |
| 2019 | An Adaptive Approach to Recommending Obfuscation Rules for Java Bytecode ObfuscatorsabstractBytecode obfuscation is an essential technique for protecting intellectual property and defending against Man-AtThe-End (MATE) attacks to Java/Android applications. Several bytecode obfuscators have been developed for modifying or refactoring Java bytecode (.class) so that it becomes hard to understand but remains fully functional. These obfuscators usually integrate a variety of obfuscation rules, allowing obfuscation algorithms to be combined and enforced on the applications. Meanwhile, it still remains a difficulty: Given a bytecode file f, which obfuscation rule(s) need to be applied such that f can get obfuscated sufficiently? This paper presents ORChooser (Obfuscation Rule Chooser), an adaptive approach to recommending a small number obfuscation rules for Java bytecode obfuscators. The key idea of ORChooser is, given a bytecode obfuscator, to (1)randomly select/unselect obfuscation rules for the obfuscator, and (2)calculate the obfuscation distance between the bytecode before and after obfuscation. Furthermore, ORChooser takes an iterative process to adaptively obfuscate the bytecode file f such that the obfuscated code is far away from f. We have implemented ORChooser and evaluated it on a state-of-the-art bytecode obfuscators: Android R8. The evaluation results clearly show the strength of ORChooser. In particular, within 5 iterations, ORChooser chose about 25% of obfuscation rules for R8, reducing more than 29% of the bytecode size. The similarity between the bytecode files before and after obfuscation is less than 27%, indicating that the ORChooser-supported obfuscators have obfuscated bytecode sufficiently and reduce its comprehensibility significantly. Yanru Peng, Yuting Chen 0001, Beijun Shen |
COMPSAC (1) | 3 |
| 2019 | CocoQa: Question Answering for Coding Conventions Over Knowledge GraphsabstractCoding convention plays an important role in guaranteeing software quality. However, coding conventions are usually informally presented and inconvenient for programmers to use. In this paper, we present CocoQa, a system that answers programmer's questions about coding conventions. CocoQa answers questions by querying a knowledge graph for coding conventions. It employs 1) a subgraph matching algorithm that parses the question into a SPARQL query, and 2) a machine comprehension algorithm that uses an end-to-end neural network to detect answers from searched paragraphs. We have implemented CocoQa, and evaluated it on a coding convention QA dataset. The results show that CocoQa can answer questions about coding conventions precisely. In particular, CocoQa can achieve a precision of 82.92% and a recall of 91.10%. Repository: https://github.com/14dtj/CocoQa/ Video: https://youtu.be/VQaXi1WydAU. Tianjiao Du, Junming Cao, Qinyue Wu, Wei Li 0254, Beijun Shen, Yuting Chen 0001 |
ASE | 5 |
| 2019 | Lancer: Your Code Tell Me What You NeedabstractProgramming is typically a difficult and repetitive task. Programmers encounter endless problems during programming, and they often need to write similar code over and over again. To prevent programmers from reinventing wheels thus increase their productivity, we propose a context-aware code-to-code recommendation tool named Lancer. With the support of a Library-Sensitive Language Model (LSLM) and the BERT model, Lancer is able to automatically analyze the intention of the incomplete code and recommend relevant and reusable code samples in real-time. A video demonstration of Lancer can be found at https://youtu.be/tO9nhqZY35g. Lancer is open source and the code is available at https://github.com/sfzhou5678/Lancer. Shufan Zhou, Beijun Shen, Hao Zhong 0001 |
ASE | 2 |
| 2019 | Constructing a Knowledge Base of Coding Conventions from Online ResourcesabstractCoding conventions are a set of coding guidelines used by software developers to improve the readability of source code, increase software maintainability, and promote the reuse of coding patterns.In this paper, we introduce CCBase, a knowledge base of coding conventions, that was constructed from online resources.Specifically, CCBase was constructed as follows.We designed the ontology of the coding convention domain, crawled data related to coding conventions from a variety of online resources, and then extracted entities and relations using an NLP-enabled rule matching method.To uncover the latent relations, we further proposed a similarity metric to reveal the similar-to and relate-to relations, and developed a RCE algorithm to establish a unified type hierarchy of coding conventions.The resulting knowledge base contains 3139 coding conventions for Java and C++, with 3761 entities and 767 relations.Furthermore, we have extended the usability of CCBase by developing a question answering system on the base.We have conducted experiments to evaluate CCBase.The experimental results show that CCBase has a wide coverage on entities and relations in coding conventions domain, and the QA system achieves an F1 score of 84.5% on 214 questions raised in StackOverflow. Junming Cao, Tianjiao Du, Beijun Shen, Wei Li 0254, Qinyue Wu, Yuting Chen 0001 |
SEKE | 3 |
| 2019 | Generating SQL Statements from Natural Language Queries: A Multitask Learning Approach (S)abstractNL2SQL advocates an idea of helping engineers and/or end users generate SQL statements from natural language queries.However, it still remains a strong challenge in improving its precision and scalability.This paper introduces MultiSQL, a multitask deep learning approach to performing NL2SQL.MultiSQL unifies the task representations and trains a model in parallel on multiple tasks, including NL2SQL, machine translation, etc.It employs a multitask question-answering network for jointly learning all tasks and transferring knowledge among tasks.We have evaluated MultiSQL on two query datasets: WikiSQL (an open sourced dataset) and CnSQL (a Chinese dataset we created).The evaluation results clearly show the effectiveness of MultiSQL.In particular, the accuracies achieved by MultiSQL approximate those achieved by the state-of-the-art NL2SQL methods on WikiSQL, and its accuracy is 78%, which is 17% higher than the "Chinese2English + NL2SQL" method on CnSQL. Chunqi Chen, Yunxiang Xiong, Beijun Shen, Yuting Chen 0001 |
SEKE | 3 |
| 2019 | Enhancing Semantic Search of Crowdsourcing IT Services using Knowledge GraphabstractMining search intents in vertical websites like IT service crowdsourcing platform relies heavily on domain knowledge.Meanwhile, it still remains a difficulty of searching services in crowdsourcing platforms, as these platforms do contain much insufficient information, for example, users tend to use images describing IT services for the purpose of advertisements.To solve these problems, we build and leverage a knowledge graph to enhance searching of crowdsourcing IT services.The key idea is to (1) build an IT service knowledge graph from StackOverflow tag synonym system, Wikipedia, StuQ and data in IT service crowdsourcing platforms, (2) plug two activities into the basic search processterm expansion and service re-ranking, (3) use superordinates, hypernyms, synonyms, descriptions and relations of entities in the knowledge graph to expand user query and service information, and (4) apply a learning-to-rank model with four features to re-rank the search results, enforcing those more relevant services have the higher-ranking position.We have conducted several experiments to evaluate our approach.The results show that our approach achieves an MRR 34.9% higher and a Recall@15 11% higher than those of a basic search approach. Duankang Fu, Shufan Zhou, Beijun Shen, Yuting Chen 0001 |
SEKE | 3 |
| 2019 | CrowDevBot: A Task-Oriented Conversational Bot for Software Crowdsourcing Platform (S)abstractWith the trends of developing software on the Internet, many software crowdsourcing platforms are emerging.They attract a lot of developers to bid for crowdsourced projects and develop software systems collaboratively.In this paper, we present CrowDevBot, a task-oriented conversational bot for software crowdsourcing platform, that aims to assist online users in completing crowdsourcing-related tasks in a more natural manner.The key idea of CrowDevBot is to: (1) combine a rulebased method and an SVM-NaiveBayes-C4.5 integrated learning method to discover users' intention; (2) employ an integrated CRF (conditional random field) method with novel features to improve the performance of slot filling; and (3) leverage a software service knowledge base to unify entity names and predefine the key slots of user query.We implement CrowDevBot and integrate it into JointForce, an IT software crowdsourcing platform in China.To the best of our knowledge, this is the first time that a task-oriented conversational bot is practically used in software crowdsourcing platform(s).We evaluated our approach on real data set from JointForce.The results show that our intention detecting method achieves F1-score of 87% on the limited training data.For the slot filling, the F1-score of our integrated CRF model reaches 82%, 8% higher than that of the normal CRF model. Zeyu Ni, Beijun Shen, Yuting Chen 0001, Zhangyuan Meng, Junming Cao |
SEKE | 2 |
| 2018 | SLAMPA: Recommending Code Snippets with Statistical Language ModelabstractProgramming is typically a difficult and repetitive task. Programmers will encounter endless problems during programming, and they often need to write similar code over and over again. Over the years, many tools have been proposed to support programming. However, to the best of our knowledge, these approaches require high-quality queries or programming contexts, which are often difficult to be built or even unavailable. To address this challenge, we propose SLAMPA, a novel tool which takes advantage of statistical language model and clone detection techniques to recommend code snippets during programming. Given a piece of incomplete code, SLAMPA first infers its intention using a neural language model. Then it retrieves code snippets from codebase with the support of an efficient clone detection technology Hybrid-CD we proposed. Finally, it recommends the most similar code snippets to programmers. Our evaluation results demonstrate that Hybrid-CD precisely detects similar code snippets and it outperforms previous techniques. Our results also show that the snippets recommended by SLAMPA catch the intention of programmers and SLAMPA is capable of finding potential code reuse opportunities during programming. Shufan Zhou, Hao Zhong 0001, Beijun Shen |
APSEC | 3 |
| 2017 | Transfer Learning for Cross-Platform Software Crowdsourcing RecommendationabstractRecently, with the development of software crowdsourcing industry, an increasing number of users joined the software crowdsourcing platforms to publish software project tasks or to seek proper work opportunities. One of competitive functions of these platforms is to recommend proficient projects to developers. However, in such recommender system, there exists a serious platform cold-start problem, especially for new software crowdsourcing platforms, as they usually have too little cumulative data to support accurate model training and prediction. This paper focuses on solving the platform cold-start problem in software crowdsourcing recommendation system by transfer learning technologies. We proposed a novel cross-platform recommendation method for new software crowdsourcing platforms, whose idea is trying to transfer data and knowledge from other mature software crowdsourcing platforms (source domains) to solve the insufficient recommendation model training problem in a new platform (target domain). The proposed method maps different kinds of features both in the source domain and the target domain after a certain transformation and combination to a latent space by learning the correspondences between features. Specifically, our method is an instance of content-based recommendation, which uses tags and keywords extracted from project description in crowdsourcing platforms as features, and then set weights for each feature to reflect its importance. Then, Weight-SCL is proposed to merge and distinguish tag features and keyword features before doing feature mapping and data migration to implement knowledge transformation. Finally, we use the data from two famous software crowdsourcing platform as dataset, and a series of experiments are conducted to evaluate the performance of the multi-source recommendation system in comparison with the baseline methods, and get 1.2X performance promotion. Shuhan Yan, Beijun Shen, Wenkai Mo |
APSEC | 2 |
| 2017 | How Do Programmers Maintain Concurrent Code?abstractConcurrent programming is pervasive in nowadays software development. Many programmers believe that concurrent programming is difficult, and maintaining concurrency code is error-prone. Although researchers have conducted empirical studies to understand concurrent programming, they still rarely study how programmers maintain concurrent code. To the best of our knowledge, only a recent study explored the modifications on critical sections, and many related questions are still open. In this paper, we conduct an empirical study to explore how programmers maintain concurrent code. We analyze more concurrency-related commits and explore more issues such as the change patterns of maintaining concurrent code than the previous study. We summarize five change patterns according to our analysis on 696 concurrency-related commits. We apply our change patterns to three open source projects, and synthesize three pull requests. Until now, two of them have been accepted. Our results can be useful for programmers to maintain concurrent code and for researchers to implement treating techniques. Feiyue Yu, Hao Zhong 0001, Beijun Shen |
APSEC | 3 |
| 2017 | CRSearcher: Searching Code Database for Repairing BugsabstractWith the exponentially rising of software development in the past decades, millions of software products have been created. Existing empirical studies show that many code snippets are similar. Although there exist many difficulties in maintaining these similar code snippets, we believe that it is feasible to leverage the similarity to enhance program repairs, as bugs may have already been repaired in many other similar code snippets. Yingyi Wang, Yuting Chen 0001, Beijun Shen, Hao Zhong 0001 |
Internetware | 3 |
| 2017 | A GQM-based Approach for Software Process Patterns RecommendationabstractA good software process can help project manager manage software development effectively and control development risks.For this reason, theory and experts' experience are concluded and put into process patterns.But it still requires human skills to search for appropriate process patterns in practice.To tackle this challenge, this paper proposes a Goal-Question-Metric (GQM) based approach to recommending software process patterns.The essential idea of this approach is to use a GQM method to design scenario questions for software process patterns, elicit the requirement of new project by answering these questions, and then recommend the optimal matching patterns to the project.In particular, we use a Latent Dirichlet Allocation model on the scenario descriptions of software process patterns to achieve a text-topic distribution, and then apply the K-means method to do text clustering, which facilitate scenario questions design a lot.We evaluate the performance of our topic clustering method by comparing it with that of the statistics method based on TF-IDF.The evaluation results show that our method contributes a high F-score which is 11.6% higher than that of the traditional TF-IDF approach.Furthermore, the average precision of recommendation can reach 57%. Zhangyuan Meng, Beijun Shen, Yin Wei |
SEKE | 3 |
| 2017 | Mining Developer Behavior Across GitHub and StackOverflowabstractNowadays, software developers are increasingly involved in GitHub and StackOverflow, creating a lot of valuable data in the two communities.Researchers mine the information in these software communities to understand developer behaviors, while previous work mainly focuses on mining data within a single community.In this paper, we propose a novel approach to mining developer behaviors across GitHub and StackOverflow.This approach links the accounts from two communities using a CART decision tree, leveraging the features from usernames, user behaviors and writing styles.Then, it explores cross-site developer behaviors through T-graph analysis, LDA-based topics clustering and cross-site tagging.We conducted several experiments to evaluate this approach.The results show that the precision and F-Score of our identity linkage method are higher than previous methods in software communities.Especially, we discovered that (1) active issue committers are also active question askers; (2) for most developers, the topics of their contents in GitHub are similar to that of their questions and answers in StackOverflow;(3) developers' concerns in StackOverflow shift over the time of their current participating projects in GitHub; (4) developers' concerns in GitHub are more relevant to their answers than questions and comments in StackOverflow. Yunxiang Xiong, Zhangyuan Meng, Beijun Shen |
SEKE | 3 |
| 2017 | Cold-Start Developer Recommendation in Software Crowdsourcing: A Topic Sampling ApproachabstractRecently, software crowdsourcing platforms, which provide paid tasks for developers, become attractive to both employers and developers.Developers expect to find tasks that match their interests and capabilities via crowdsourcing platforms, and thus recommender systems play important roles in these platforms.However, we still face several challenges when building a recommender system for a crowdsourcing platform.A major challenge is how to recommend tasks to cold-start developers whose task interaction data is not available.This paper presents a novel, topic sampling approach to tackling with the cold-start developer recommendation problem.First, it employs a general method for modeling developers and tasks, which solves the data heterogeneous issue across different platforms.After that, it casts the cold-start developer recommendation problem into a multi-optimization problem, and takes a topic-sampling based genetic algorithm to recommend tasks.More specifically, our approach is different from traditional solutions in that it leverages task descriptions and popularity-to-be, allowing new tasks to be recommended to cold-start developers.To evaluate the effectiveness of the proposed approach, we have conducted experiments on a large dataset crawled from three real-world software crowdsourcing platforms.Compared with other state-ofthe-art recommendation solutions, the experimental results show that the proposed approach improves 75% of precision and recall on average. Wenkai Mo, Beijun Shen, Yuting Chen 0001 |
SEKE | 3 |
| 2017 | Developer Identity Linkage and Behavior Mining Across GitHub and StackOverflowabstractNowadays, software developers are increasingly involved in GitHub and StackOverflow, creating a lot of valuable data in the two communities. Researchers mine the information in these software communities to understand developer behaviors, while previous works mainly focus on mining data within a single community. In this paper, we propose a novel approach to developer identity linkage and behavior mining across GitHub and StackOverflow. This approach links the accounts from two communities using a CART decision tree, leveraging the features from usernames, user behaviors and writing styles. Then, it explores cross-site developer behaviors through [Formula: see text]-graph analysis, LDA-based topics clustering and cross-site tagging. We conducted several experiments to evaluate this approach. The results show that the precision and [Formula: see text]-score of our identity linkage method are higher than previous methods in software communities. Especially, we discovered that (1) active issue committers are also active question askers; (2) for most developers, the topics of their contents in GitHub are similar to those of those questions and answers in StackOverflow; (3) developers’ concerns in StackOverflow shift over the time of their current participating projects in GitHub; (4) developers’ concerns in GitHub are more relevant to their answers than questions and comments in StackOverflow. Yunxiang Xiong, Zhangyuan Meng, Beijun Shen |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2016 | Task Recommendation with Developer Social Network in Software CrowdsourcingabstractRecently, crowdsourcing has been increasingly used in software industry to lower costs and increase innovations, by utilizing experiences, labor, or creativity of developers worldwide. In software crowdsourcing platforms, developers expect to find suitable tasks for their interests and abilities. So it is significant for software crowdsourcing to build a recommender system to match developers with suitable tasks. However, there are a significant number of inactive developers who have very sparse historical behavior records in the platform, and thus state-of-the-art recommendation approaches in software crowdsourcing, such as collaborative filtering, suffer from this cold-start problem. In this paper, a social influence-based method is proposed to recommend suitable tasks for both active and inactive developers. The essential idea of the novel method is (1) to construct developer social network from developer behaviors, such as browsing and bidding for tasks, (2) to calculate the influence degrees between developers using developer social network, (3) to recommend tasks to active developers using SiSVD, and (4) to recommend tasks to inactive developers by combining the recommended tasks of their friends. We have evaluated our method on a large real data set from the JointForce, a popular software crowdsourcing platform in China. The results show that our method is feasible and practical for recommendation in software crowdsourcing. In particular, the F1-Measure of our method for inactive developers with task-bidding friends is increased by 16.7% than other previous approaches averagely. Wenkai Mo, Beijun Shen |
APSEC | 3 |
| 2016 | EXPSOL: Recommending Online Threads for Exception-Related Bug ReportsabstractAn exception-related bug is a kind of program bug which causes exceptions. During software maintenance, when programmers repair exception-related bugs, they typically analyze thrown exceptions to understand the root causes of such bugs. When encountering unfamiliar thrown exceptions, programmers often refer to online forum threads (e.g. StackOverflow) to understand how to fix them. Although some general search engines are available and some research tools are proposed, they are insufficient to recommend threads for exception-related bugs from large-scale online resources. In this paper, we propose an approach, named EXPSOL, which recommends online threads as solutions for a newly reported exception-related bug with a model trained by support vector machines. We conduct two evaluations on thousands of threads from StackOverflow and fixed issues from GitHub. The results of our first evaluation show the significance of our internal features and highlight the importance of integrating different features. The results of our second evaluation show that, EXPSOL performs better in mean average precision, mean reciprocal rank and recall than those of the Google search engine, the internal search engine of StackOverflow, and other existing approaches. Beijun Shen, Hao Zhong 0001, Jiangang Zhu |
APSEC | 2 |
| 2016 | Heterogeneous Cross-Company Effort Estimation through Transfer LearningabstractSoftware effort estimation is vital but challenging activity during software development. In many small or medium-sized companies, such challenges are stemmed from historical data shortage. The problem can be solved by leveraging cross-company data for effort estimation. While in practice, cross-company effort estimation may not be easy to take because the cross-company data for effort estimation can be heterogenous. In this paper, we propose a novel approach named Mixture of Canonical Correlation Analysis and Restricted Boltzmann Machines (MCR) to address data heterogeneity issue in cross-company effort estimation. The essential ideas in MCR are (1) to present a unified metric representing heterogenous effort estimation data; and (2) to combine Canonical Correlation Analysis and Restricted Boltzmann Machines method to estimate effort in heterogenous cross-company effort estimation. The MCR approach is evaluated on 5 public datasets in PROMISE repository. The evaluation results show that: (1) for estimations with partially different metrics, the MCR approach outperforms within-company effort estimator KNN with a decrease in MMRE by 0.60, an increase in PRED(25) by 0.16, and a decrease in MdMRE by 0.19; (2) for estimations with totally different metrics, the MCR approach outperforms within-company effort estimator KNN with a decrease in MMRE by 0.49, an increase in PRED(25) by 0.08, and a decrease in MdMRE by 0.10. Shensi Tong, Yuting Chen 0001, Beijun Shen |
APSEC | 5 |
| 2016 | GRETA: Graph-Based Tag Assignment for GitHub RepositoriesabstractGitHub is a well-known software community where a large number of software repositories are hosted. Since large amounts of documents and code in GitHub repositories are in a mess, users cannot search or understand them efficiently. One solution is to employ a tag system, which annotates each repository with several tags. Thus, the GitHub repositories can be more efficiently accessed and understood. However, GitHub does not provide any automated tools of tagging repositories. In order to tackle this problem, we propose GRETA, a novel graph-based approach to assigning tags for repositories on GitHub. The core insight of GRETA is (1) to construct an Entity-Tag Graph (ETG) for GitHub using the domain knowledge from StackOverflow, and (2) to assign tags for repositories by taking a random walk algorithm. We have implemented GRETA and also developed a repository search engine for GitHub using the tag assignment results of GRETA. We have evaluated GRETA against several baseline methods to investigate its effectiveness of tagging GitHub repositories. The results show GRETA achieves up to 35% of F-Measure, outperforming the baseline methods. Besides, the GRETA-based search engine gains a higher NDCG value than the search engine provided by GitHub, indicating that it significantly enhances the search ability on GitHub with the tagged repositories. Xuyang Cai, Jiangang Zhu, Beijun Shen, Yuting Chen 0001 |
COMPSAC | 3 |
| 2016 | Software Defect Prediction Using Semi-Supervised Learning with Change Burst InformationabstractSoftware defect prediction is an important software quality assurance technique. It utilizes historical project data and previously discovered defects to predict potential defects. However, most of existing methods assume that large amounts of labeled historical data are available for prediction, while in the early stage of the life cycle, projects may lack the data needed for building such predictors. In addition, most of existing techniques use static code metrics as predictors, while they omit change information that may introduce risks into software development. In this paper, we take these two issues into consideration, and propose a semi-supervised based defect prediction approach - extRF. extRF extends the classical supervised Random Forest algorithm by self-training paradigm. It also employs change burst information for improving accuracy of software defect prediction. We also conduct an experiment to evaluate extRF against three other supervised machine learners (i.e. Logistic Regression, Naive Bayes, Random Forest) and compare the effectiveness of code metrics, change burst metrics, and a combination of them. Experimental results show that extRF trained with a small size of labeled dataset achieves comparable performance to some supervised learning approaches trained with a larger size of labeled dataset. When only 2% of Eclipse 2.0 data are used for training, extRF can achieve F-measure about 0.562, approximate to that of LR (a supervised learning approach) at labeled sampling rate of 50%. Besides, change burst metrics outperform code metrics in that F-measure rises to a peak value of 0.75 for Eclipse 3.0 and JDT.Core. Beijun Shen, Yuting Chen 0001 |
COMPSAC | 2 |
| 2016 | SOLinker: Constructing Semantic Links between Tags and URLs on StackOverflowabstractThanks to the strength of crowdsourcing, there is a lot of useful information on StackOverflow, the most popular Question and Answer (Qaamp;A) platform in software engineering area. This information can be treated as numerous URLs (Uniform Resource Locators), which can be categorized into URLs of Qaamp;As and URLs in Qaamp;As. The domain of former ones is Stack-Overflow itself, while domains of latter ones are miscellaneous, such as some personal blogs and so on. Although each Qaamp;A has been manually assigned tags, relations between URLs and tags are not clear enough. In this paper, we propose SOLinker, a method to build semantic links between various URLs and tags. Firstly, SOLinker identifies proper relations from a predefined relation set between tags and URLs, which is modeled as a text classification problem. Features are extracted from content of Qaamp;A, the URL and the tag list, and classification algorithms are Logistic Regression and Gradient Boosting Decision Tree, depending on the category of URLs. Secondly, there exists a partial tagging problem, which means for a URL in a Qaamp;A, there are only a part of tags of the Qaamp;A relating to the URL. To address this problem, we propose a semantic analysis method to analyze context of this URL and the URL itself from both implicit and explicit aspects. Then SOLinker will infer proper tags by the label propagation technique. Results show that our method is feasible and practical in constructing semantic links between tags and URLs of/in Qaamp;As. In particular, the F-Score of semantic relation identification is around 78%, 5% higher than the other existing method, and F-Score of partial tagging solving is around 88%. Wenkai Mo, Jiangang Zhu, Zhenzheng Qian, Beijun Shen |
COMPSAC | 4 |
| 2016 | SatiIndicator: Leveraging User Reviews to Evaluate User Satisfaction of SourceForge ProjectsabstractQuality of software (QoS) is important for users, as it may lead to high cost when a user or a company happens to pick up a software project with low quality. In recent years, many software quality assessment models take user satisfaction as an important metric for measuring software quality. However, user satisfaction on a software project is usually not precisely evaluated. In this paper, we propose a novel, automated approach called SatiIndicator to evaluate user satisfaction of a software project by analyzing user reviews with user opinions and emotions. The essential idea of SatiIndicator is to (1) use a topic model to cluster all aspects in a software genre into different topics and compute the weight of each topic, (2) take sentiment analysis and calculate the sentiment strength of every aspect and pure attitude reviews, and (3) evaluate the user satisfaction score for a software project. Wilson Interval is applied to punish the software projects with insufficient reviews in order to keep fairness. We have evaluated SatiIndicator on ten software genres in SourceForge. The evaluation results show that when software projects have sufficient reviews, SatiIndicator performs 35% higher than baselines at p@3, 15% higher than baselines at p@15 and over 85% Spearman Coefficient with the ground truth. When software projects have insufficient reviews, SatiIndicator performs 30% higher than baselines at p@3, 15% higher than baselines at p@15 and over 60% Spearman Coefficient with the ground truth. Zhenzheng Qian, Beijun Shen, Wenkai Mo, Yuting Chen 0001 |
COMPSAC | 2 |
| 2016 | ESSE: an early software size estimation method based on auto-extracted requirements featuresabstractSoftware size estimation is a crucial step in project management. According to the Standish Chaos Report, 65% of software projects are over budget or deadline; therefore, a good size estimation method is very important. However, existing estimation methods are complicated and human-effort consuming. In many industrial projects, project technical leads (PTLs) do not use these methods but just give a rough estimation based on their experience. To decrease human effort, we propose an early software size estimation (ESSE) method, which can extract semantic features from natural language requirements automatically, and build size estimation models for project. Firstly, ESSE makes a two-level semantic analysis of requirements specification documents by information extraction and activation spreading. Then, complexity-related features are extracted from the results of semantic analysis. Finally, a size estimation model is trained to predict size of new projects by regression algorithms. Experiments in real industrial datasets show that our method is effective and can be applied to real industrial projects. Shensi Tong, Wenkai Mo, Yong Xia 0009, Beijun Shen |
Internetware | 6 |
| 2016 | Building a Domain Knowledge Base from Wikipedia: a Semi-supervised ApproachabstractKnowledge bases are becoming indispensable to software engineering and knowledge engineering.However, the existing domain knowledge bases are always artificially constructed and small-scale.In this paper, we propose a semi-supervised approach to domain concepts detection and software engineering knowledge base construction from Wikipedia.First, the approach selects domain relevant tags from Stackoverflow.Then, it matches Wikipedia entities and expands the concept set through an improved label propagation algorithm.A rule-based method is designed to discover semantic relations including relate, subclassOf and equal by analyzing structural information of Wikipedia.A relation derivation mechanism is presented to optimize the relation set.We finally construct SEBase, a domainspecific knowledge base of software engineering.Experimental results show the high accuracy of the integrated concepts and relations.Compared with other knowledge bases, SEBase has the widest coverage of concepts and relations in software engineering. Xiang Dong, Jiangang Zhu, Beijun Shen |
SEKE | 4 |
| 2016 | Learning to Discover Subsumptions between Software Engineering Concepts in WikipediaabstractWikipedia contains large-scale concepts and rich semantic information.A number of knowledge base construction projects such as WikiTaxonomy, DBpedia, and YAGO have acquired data from Wikipedia.Despite the huge amount of relations in Wikipedia, the semantic relations (i.e.subsumptions) between domain concepts are rather sparse, especially in software engineering (SE) area.Hence, it is difficult to derive a software engineering knowledge base directly from Wikipedia.Meanwhile, domain knowledge base has become indispensable to a growing number of applications in software engineering.So the discovery of missing semantic relations between software engineering concepts in Wikipedia is essential.In this paper, we propose an approach to automatically discovering the missing subsumption relations between software engineering concepts.Specifically, we extract the SE domain concepts from Wikipedia firstly.And secondly, we design a machine learning based algorithm with some novel features to calculate the semantic relevancy between concepts.Thirdly, we offer and utilize a semi-supervised model to incorporate the features, which discovers the SE subsumptions.Experimental results show that our approach can effectively find the missing subsumption relations between software engineering concepts.Finally, we build a taxonomy which contains 193,593 concepts together with 357,662 subsumption relations.Compared with the taxonomies which are extracted from general-purpose knowledge bases such as WikiTaxonomy, YAGO and Schema.org,our dataset has a larger scale in software engineering domain. Xiang Dong, Jiangang Zhu, Beijun Shen |
SEKE | 4 |
| 2016 | CPDScorer: Modeling and Evaluating Developer Programming Ability across Software CommunitiesabstractSince developer ability is recognized as a determinant of better software project performance, it is a critical step to model and evaluate the programming ability of developers.However, most existing approaches require manual assessment, like 360 degree performance evaluation.With the emergence of social networking sites such as StackOverflow and Github, a vast amount of developer information is created on a daily basis.Such personal and social context data has huge potential to support automatic and effective developer ability evaluation.In this paper, we propose CPDScorer, a novel approach to modeling and scoring the programming ability of developer through mining heterogeneous information from both Community Question Answering (CQA) sites and Open-Source Software (OSS) communities.CPDScorer analyzes the answers posted in CQA sites and evaluates the projects submitted in OSS communities to assign expertise scores to developers, considering both the quantitative and qualitative factors.When modeling the programming ability of developer, a programming ability term extraction algorithm is also designed based on set covering.We have conducted experiments on StackOverflow and Github to measure the effectiveness of CPDScorer.The results show that our approach is feasible and practical in user programming ability modeling.In particular, the precision of our approach reaches 80%. Weizhi Huang, Wenkai Mo, Beijun Shen |
SEKE | 3 |
| 2016 | Operational pattern based code generation for management information system: An industrial case studyabstractCode generation technology can significantly improve productivity and software quality. However, due to limited financial and human resources in most of small and medium software enterprises, there are many challenges when leveraging code generation approaches to large-scale software development. In this paper, an operational pattern based code generation approach is proposed for rapid development of domain-specific management information system. We demonstrate the approach with details: (I) semi-automatically extracting operational patterns from requirement documents, (II) building feature models to manage the commonalities and variability of each operational pattern, (III) mapping operational patterns into skeleton code with a template-based code generation technique, etc. Then we conduct an industrial case study in asset information management domain at CancoSoft Company for about 2 years, to analyze its feasibility and efficiency. 14 operational patterns are successfully extracted from 355 initial key phrases, and a code generator is implemented and applied to develop new Web applications. Preliminary findings show that the software development based on our approach yields a nearly 30% higher productivity as compared to traditional software development. Through code analysis, we find that around 70% of code can be automatically generated, and the generated code is also effective. Fagui Mao, Xuyang Cai, Beijun Shen, Yong Xia 0009 |
SNPD | 3 |
| 2016 | Automatically Modeling Developer Programming Ability and Interest Across Software CommunitiesabstractDeveloper profile plays an important role in software project planning, developer recommendation, personnel training, and other tasks. Modeling the ability and interest of developers is its key issue. However, most existing approaches require manual assessment, like 360[Formula: see text] performance evaluation. With the emergence of social networking sites such as StackOverflow and Github, a vast amount of developer information is created on a daily basis. Such personal and social context data has huge potential to support automatic and effective developer ability evaluation and interest mining. In this paper, we propose CPDScorer, a novel approach for modeling and scoring the programming ability and interest of developers through mining heterogeneous information from both community question answering (CQA) sites and open-source software (OSS) communities. CPDScorer analyzes the questions and answers posted in CQA sites, and evaluates the projects submitted in OSS communities to assign expertise scores as well as interest scores to developers, considering both the quantitative and qualitative factors. When profiling developer's ability and interest, a programming term extraction algorithm is also designed based on set covering. We have conducted experiments on StackOverflow and Github to measure the effectiveness of CPDScorer. The results show that our approach is feasible and practical in user programming ability and interest modeling. In particular, the precision of our approach reaches 80%. Weizhi Huang, Wenkai Mo, Beijun Shen |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2015 | TBIL: A Tagging-Based Approach to Identity Linkage Across Software CommunitiesabstractNowadays, developers can be involved in several software developer communities like StackOverflow and Github. Meanwhile, accounts from different communities are usually less connected. Linking these accounts, which is called identity linkage, is a prerequisite of many interesting studies such as investigating activities of one developer in two or more communities. Many researches have been performed on social networks, but very few of them can be adapted to software communities, as information of users provided in these communities has a huge difference to that in social networks. We tackle with the problem by introducing TBIL, a novel tagging-based approach to identity linkage among software communities. The essential idea of this approach is to employ skills (measured by tags), usernames and concerned topics of developers as hints, and to use a decision tree-based algorithm and another heuristic greedy matching algorithm to link user identities. We measure the effectiveness of TBIL on two well-known software communities, i.e., StackOverflow and Github. The results show that our method is feasible and practical in linking developer identities. In particular, the F-Score of our method is 0.15 higher than previous identity linkage methods in software communities. Wenkai Mo, Beijun Shen, Yuting Chen 0001, Jiangang Zhu |
APSEC | 2 |
| 2015 | A Learning to Rank Framework for Developer Recommendation in Software CrowdsourcingabstractRecently, crowdsourcing has been widely used in many tasks that computers are not good at such as image recognition, entity resolution or some question answering tasks. A key feature of these tasks is that they are all simple tasks even decision making tasks. People can deal with these tasks with common sense knowledge. However, different from crowdsourcing in a general domain, software crowdsourcing is more complex. Only people with software developing skills can finish these tasks which could take a long time. Thus, an essential component of building a successful software crowdsourcing platform is effective developer recommendation, which aims to match a given task to the right crowdworkers. In order to solve this problem, in this paper, we propose a learning to rank framework. Specifically, we first propose a CRF-based model to extract criterias (i.e. skills and locations) from descriptions. Task characteristics learned from their descriptions and developer' characteristic distributions extracted from their historical tasks are fed into our learning to rank algorithms for developer recommendation together with some other features such as topic-based features. We have evaluated our approach on a large dataset crawled from a real-world software crowdsourcing platform. The experimental results show that our approach is feasible and effective. Jiangang Zhu, Beijun Shen, Fanghuai Hu |
APSEC | 2 |
| 2015 | Code Bad Smell Detection through Evolutionary Data MiningabstractThe existence of code bad smell has a severe impact on the software quality. Numerous researches show that ignoring code bad smells can lead to failure of a software system. Thus, the detection of bad smells has drawn the attention of many researchers and practitioners. Quite a few approaches have been proposed to detect code bad smells. Most approaches are solely based on structural information extracted from source code. However, we have observed that some code bad smells have the evolutionary property, and thus propose a novel approach to detect three code bad smells by mining software evolutionary data: duplicated code, shotgun surgery, and divergent change. It exploits association rules mined from change history of software systems, upon which we define heuristic algorithms to detect the three bad smells. The experimental results on five open source projects demonstrate that the proposed approach achieves higher precision, recall and F-measure. Shizhe Fu, Beijun Shen |
ESEM | 2 |
| 2015 | GEMiner: Mining Social and Programming Behaviors to Identify Experts in GithubabstractHosting over 10 million repositories, GitHub becomes the largest open source community in the world. Besides sharing code, Github is also a social network, in which developers can follow others or keep track of their interested projects. Considering the multi-roles of Github, integrating heterogenous data of each developer to identify experts is a challenging task. In this paper, we propose GEMiner, a novel approach to identify experts for some specific programming languages in Github. Different from previous approaches, GEMiner analyzes the social behaviors and programming behaviors of a developer to determine the expertise of the developer. When modeling social behaviors of developers, to integrate heterogenous social networks in Github, GEMiner implements a Multi-Sources PageRank algorithm. Also, GEMiner analyzes the behaviors of developers when they are programming (e.g., their commit activities and their preferred programming languages) to model programming behaviors of them. Based on our expertise models and our extracted programming languages data, GEMiner can then identify experts for some specific programming languages in Github. We conducted experiments on a real data set, and our results show that GEMiner identifies experts with 60% accuracy higher than the state-of-the-art algorithms. Wenkai Mo, Beijun Shen, Yuming He, Hao Zhong 0001 |
Internetware | 2 |
| 2015 | Cross-Project Software Defect Prediction Using Feature-Based Transfer LearningabstractCross-project defect prediction is taken as an effective means of predicting software defects when the data shortage exists in the early phase of software development. Unfortunately, the precision of cross-project defect prediction is usually poor, largely because of the differences between the reference and the target projects. Having realized the project differences, this paper proposes CPDP, a feature-based transfer learning approach to cross-project defect prediction. The core insight of CPDP is to (1) filter and transfer highly-correlated data based on data samples in the target projects, and (2) evaluate and choose learning schemas for transferring data sets. Models are then built for predicting defects in the target projects. We have also conducted an evaluation of the proposed approach on PROMISE datasets. The evaluation results show that, the proposed approach adapts to cross-project defect prediction in that f-measure of 81.8% of projects can get improved, and AUC of 54.5% projects improved. It also achieves similar f-measure and AUC as some inner-project defect prediction approaches. He Qing, Biwen Li, Beijun Shen, Yong Xia 0009 |
Internetware | 3 |
| 2015 | Towards Effective Developer Recommendation in Software CrowdsourcingabstractCrowdsourcing has attracted increasing attention from both industry and academia since it was proposed.Now a lot of work is finished by crowdsourcing, such as logo design, website promotion, industrial design, copywriting, software development, translation and image annotation.Although software crowdsourcing achieves positive results in practice, we still face a challenge of assigning suitable developers to specific tasks.In this paper, we propose a novel approach that recommends developers.In particular, our approach supports: comprehensively measuring the tasks and developers in software crowdsourcing, and recommending developers on the basis of the developer-task competence, task-task similarity, and soft power. Shixiong Zhao, Beijun Shen, Yuting Chen 0001, Hao Zhong 0001 |
SEKE | 2 |
| 2015 | Building a Large-scale Software Programming Taxonomy from StackoverflowabstractTaxonomy is becoming indispensable to a growing number of applications in software engineering such as software repository mining and defect prediction.However, the existing related taxonomies are always manually constructed.The sizes of these taxonomies are small and their depths are limited.In order to show the full potential of taxonomies in software engineering applications, in this paper, we present the first large-scale software programming taxonomy which is more comprehensive than any existing ones.It contains 38,205 concepts and 68,098 subsumption relations.Instead of learning from a open domain, we focus on taxonomy construction from Stackoverflow which is one of the largest QA websites about software programming.We propose a machine learning based method with novel features to create a taxonomy that captures the hierarchical semantic structure of tags in Stackoverflow.This method executes iteratively to find as many relations as possible.Experimental results show that our approach achieves much better accuracy than baselines.Compared with taxonomies related to software programming which are extracted from the general-purpose taxonomies such as WikiTaxonomy, Yago Taxonomy and Schema.org,our taxonomy has the widest coverage of concepts, contains the largest number of subsumption relations, and runs up to the deepest semantic hierarchy. Jiangang Zhu, Beijun Shen, Xuyang Cai, Haofen Wang |
SEKE | 2 |
| 2014 | Mining Developer Mailing List to Predict Software DefectsabstractIt has been studied that the communication among software stakeholders can be used to predict potential software defects. Yet researchers have rarely studied the relations between the software and the mailing lists of the developers. In this paper, we research on how to predict software defects by mining the mailing lists of the software developers. First, we extract both the structural and the unstructured information from mailing lists as metrics. The structural information is calculated through analyzing the social network hidden in the mailing lists, and the unstructured information is obtained through taking topical and textual analysis of the lists. Second, we design a mailing list-based approach to predicting software defects. We have also analyzed the software repository of several open source projects by linking their bug tracking data-bases to the mailing list archives. The experimental results provide empirical evidence that the mailing list metrics are related to software quality and can be used as predictors of defect-proneness. Furthermore, we found that (1) messages having certain structures may indicate some defect related files, (2) the sentiment and some topic-specific mailing models are of strong correlations with the software defects. Beijun Shen, Yuting Chen 0001 |
APSEC (1) | 2 |
| 2014 | A Scenario-Based Approach to Predicting Software Defects Using Compressed C4.5 ModelabstractDefect prediction approaches use software metrics and fault data to learn which software properties are associated with what kinds of software faults in programs. One trend of existing techniques is to predict the software defects in a program construct (file, class, method, and so on) rather than in a specific function scenario, while the latter is important for assessing software quality and tracking the defects in software functionalities. However, it still remains a challenge in that how a functional scenario is derived and how a defect prediction technique should be applied to a scenario. In this paper, we propose a scenario-based approach to defect prediction using compressed C4.5 model. The essential idea of this approach is to use a k-medoids algorithm to cluster functions followed by deriving functional scenarios, and then to use the C4.5 model to predict the fault in the scenarios. We have also conducted an experiment to evaluate the scenario-based approach and compared it with a file-based prediction approach. The experimental results show that the scenario-based approach provides with high performance by reducing the size of the decision tree by 52.65% on average and also slightly increasing the accuracy. Biwen Li, Beijun Shen, Yuting Chen 0001, Jinshuang Wang |
COMPSAC | 2 |
| 2013 | Mining GitHub: Why Commit Stops - Exploring the Relationship between Developer's Commit Pattern and File Version EvolutionabstractUsing the freeware in GitHub, we are often confused by a phenomenon: the new version of GitHub freeware usually was released in an indefinite frequency, and developers often committed nothing for a long time. This evolution phenomenon interferes with our own development plan and architecture design. Why do these updates happen at that time? Can we predict GitHub software version evolution by developers' activities? This paper aims to explore the developer commit patterns in GitHub, and try to mine the relationship between these patterns (if exists) and code evolution. First, we define four metrics to measure commit activity and code evolution: the changes in each commit, the time between two commits, the author of each changes, and the source code dependency. Then, we adopt visualization techniques to explore developers' commit activity and code evolution. Visual techniques are used to describe the progress of the given project and the authors' contributions. To analyze the commit logs in GitHub software repository automatically, Commits Analysis Tool (CAT) is designed and implemented. Finally, eight open source projects in GitHub are analyzed using CAT, and we find that: 1) the file changes in the previous versions may affect the file depend on it in the next version, 2) the average days around "huge commit" is 3 times of that around normal commit. Using these two patterns and developer's commit model, we can predict when his next commit comes and which file may be changed in that commit. Such information is valuable for project planning of both GitHub projects and other projects which use GitHub freeware to develop software. Weicheng Yang, Beijun Shen, Ben Xu |
APSEC (2) | 2 |
| 2011 | A semantic unification approach for M2M applications based on ontologyabstractM2M (Machine-to-Machine) technology makes it possible to network all kinds of terminal devices and their corresponding enterprise application servers. Nowadays, M2M platform, working as a mid-workstation, collecting raw data from terminals through 3G wireless techs and then transferring them to enterprise applications, has been popularly utilized by several enterprise applications. However, it still has difficulty in information sharing among different terminals and enterprise applications when transferring data because of the distinguished usage of business vocabularies. Hence, this paper puts forward a core business vocabulary (CBV) mapping approach based on Multi-Domain Ontologies that can lead to semantic unification when M2M platform dispatches the data among terminal devices and enterprise application servers, in order to make it easier to share information among each other. Moreover, we implement a CBV mapping module which is integrated into the work M2M Middleware and then put in practice. Beijun Shen |
WiMob | 2 |