VLDB 2026 Research / reviewers in the wild / expert
Hongyu Kuang
dblp:124/2236
· DBLP profile ↗
33ranked-venue papers
8as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 23 · 6 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fusion Is Not A Simple Ensemble! Towards The Evolving Views in Insider Threat DetectionabstractInsider threat detection (ITD) is notoriously difficult: malicious actions are rare, context-dependent, and deliberately hidden within massive volumes of legitimate user behavior. Existing ITD methods rely on single- or fused-view models, which lack extensibility and therefore fail to leverage the supervisory signals from newly introduced complementary views. While ensembling is a natural next step, its direct application to ITD confronts three core obstacles: scalability bottlenecks from independently trained sub - models, semantic misalignment across heterogeneous feature spaces, and view imbalance, where strong views overshadow weaker yet informative ones. In this work, we propose Insight-LLM, the first extensible multi-view fusion framework tailored for ITD. Insight-LLM encodes each view with frozen pre-trained backbones and aligns heterogeneous representations into a unified semantic space via a lightweight ViewAdapter, enabling coherent cross-view reasoning without incurring additional training overhead. A context-adaptive fusion module dynamically re-weights views to emphasize subtle yet semantically consistent threat signals, and the fused representation is integrated with task prompts for lightweight LLM fine-tuning. Experiments on CERT datasets show that Insight-LLM improves F1 by up to 4.8% and reduces false positives by 61%, while decreasing training time per newly added view by up to 83.2% compared with the simple Ensemble method. Chengyu Song, Lin Yang 0031, Jianming Zheng, Jingjing Zhang 0005, Hongyu Kuang, Jinzhi Liao, Mengchun Zhao |
WWW | 5 |
| 2026 | Understanding Users' Affective States During Issue Resolution in Open Source Software Projects
Yu-Qian Zhuang, Liang Wang 0001, Ke-Xin Sun, Hongyu Kuang, Xian-Ping Tao |
J. Comput. Sci. Technol. | 4 |
| 2026 | UntCC: Untangling Composite Commits Using Structural and Semantic InformationabstractSmall and focused commits are highly valued in modern software development. However, developers sometimes submit a commit with more than one concern, represented by several lines of code changes for a specific purpose, e.g., adding new features or fixing bugs. Such composite commits confuse developers during code reviews as well as other software activities, resulting in various issues. Existing studies predominantly leverage code structure to untangle composite commits, but without considering code semantics that have been demonstrated to be important in many related studies. In this article, we propose UNTCC, a new approach that uses structural and semantic information forUNTanglingCompositeCommits. To achieve structural information, we propose the code change graph, a fine-grained, text-attributed graph representation of a commit, incorporating before-change and after-change code dependencies; and UNTCC employs the graph autoencoder to learn its structural representation. To achieve semantic information, UNTCC leverages a large language model (Llama-3.2-3B) to learn joint embeddings of the raw commit and its aligned graph representation, which guide the division of different concerns within a composite commit. The experimental evaluation using 27,853 composite commits from 9 C# and 10 Java projects shows that in terms of Accuracya/Accuracyc, UNTCC achieves 94%/74% in C# and 77%/54% in Java, outperforming state-of-the-art approaches by 2%—623%/32%—573% in C# and 22%—285%/35%—286% in Java. The results indicate that UNTCC can effectively untangle composite commits. Yuzhe Jin, Lanxin Yang, He Zhang 0001, Gongyuan Li, Bohan Liu 0003, Xin Zhou 0016, Hongyu Kuang, Liming Dong 0001 |
IEEE Trans. Software Eng. | 8 |
| 2026 | Does AI Code Review Lead to Code Changes? A Case Study of GitHub Actions
Hongyu Kuang, Sebastian Baltes, Xin Zhou 0016, He Zhang 0001, Xiaoxing Ma, Guoping Rong, Dong Shao, Christoph Treude |
IEEE Trans. Software Eng. | 2 |
| 2025 | VulPelican: An LLM and Interactive Static Analysis Tool Based Vulnerability Detection Framework
Hongquan Xu, Hongyu Kuang |
ICIC (4) | 2 |
| 2025 | Poster: Boosting Inter-Procedural Vulnerability Detection via Retrieval-Augmented GenerationabstractTraditional LLM-based vulnerability detection methods face challenges like hallucinations and high false positive rates. In order to overcome these constraints, we propose an innovative RAG-based method for inter-procedural vulnerability detection named IPVRAG. The main innovation design of our IPVRAG is its multi-level feature extraction strategy: during knowledge base construction, it not only extracts functional semantics and vulnerability causes but also stores pruned Data Flow Graph structures and semantic identifier information. Evaluated on a widely-used dataset, IPVRAG outperformed many of LLM-based baselines. It achieved optimal overall performance in inter-procedural vulnerability detection, particularly demonstrating superior precision-recall balance that effectively reduced false negatives. Linru Ma, Hongquan Xu, Hongyu Kuang, Boyu Deng |
ICPADS | 3 |
| 2025 | Brevity is the Soul of Wit: Condensing Code Changes to Improve Commit Message GenerationabstractCommit messages are valuable resources for describing why code changes are committed to repositories in version control systems (e.g., Git).They effectively help developers understand code changes and better perform software maintenance tasks.Unfortunately, developers often neglect to write high-quality commit messages in practice.Therefore, a growing body of work is proposed to generate commit messages automatically.These works all demonstrated that how to organize and represent code changes is vital in generating good commit messages, including the use of fine-grained graphs or embeddings to better represent code changes.In this study, we choose an alternative way to condense code changes before generation, i.e., proposing brief yet concise text templates consisting of the following three parts: (1) summarized code changes, (2) elicited comments, and (3) emphasized code identifiers.Specifically, we first condense code changes by using our proposed templates with the help of a heuristic-based tool named ChangeScribe, and then fine-tune CodeLlama-7B on the pairs of our proposed templates and corresponding commit messages.Our proposed templates better utilize pre-trained language models, while being naturally brief and readable to complement generated commit messages for developers. Hongyu Kuang, Xin Zhou 0016, Wesley K. G. Assunção, Xiaoxing Ma, Dong Shao, Guoping Rong, He Zhang 0001 |
Internetware | 1 |
| 2025 | AUCAD: Automated Construction of Alignment Dataset from Log-Related Issues for Enhancing LLM-based Log GenerationabstractLog statements have become an integral part of modern software systems.Prior research efforts have focused on supporting the decisions of placing log statements, such as where/what to log.With the increasing adoption of Large Language Models (LLMs) for coderelated tasks such as code completion or generation, automated approaches for generating log statements have gained much momentum.However, the performance of these approaches still has a long way to go.This paper explores enhancing the performance of LLM-based solutions for automated log statement generation by post-training LLMs with a purpose-built dataset.Thus the primary contribution is a novel approach called AUCAD, which automatically constructs such a dataset with information extracting from log-related issues.Researchers have long noticed that a significant portion of the issues in the open-source community are related to log statements.However, distilling this portion of data requires manual efforts, which is labor-intensive and costly, rendering it impractical.Utilizing our approach, we automatically extract logrelated issues from 1,537 entries of log data across 88 projects and identify 808 code snippets (i.e., methods) with retrievable source code both before and after modification of each issue (including log statements) to construct a dataset.Each entry in the dataset consists of a data pair representing high-quality and problematic log statements, respectively.With this dataset, we proceed to post-train multiple LLMs (primarily from the Llama series) for automated * Corresponding author. Hao Zhang 0210, Dongjun Yu, Lei Zhang 0160, Guoping Rong, Yongda Yu, Haifeng Shen, He Zhang 0001, Dong Shao, Hongyu Kuang |
Internetware | 9 |
| 2024 | TRIAD: Automated Traceability Recovery based on Biterm-enhanced Deduction of Transitive Links among ArtifactsabstractTraceability allows stakeholders to extract and comprehend the trace links among software artifacts introduced across the software life cycle, to provide significant support for software engineering tasks. Despite its proven benefits, software traceability is challenging to recover and maintain manually. Hence, plenty of approaches for automated traceability have been proposed. Most rely on textual similarities among software artifacts, such as those based on Information Retrieval (IR). However, artifacts in different abstraction levels usually have different textual descriptions, which can greatly hinder the performance of IR-based approaches (e.g., a requirement in natural language may have a small textual similarity to a Java class). In this work, we leverage the consensual biterms and transitive relationships (i.e., inner- and outer-transitive links) based on intermediate artifacts to improve IR-based traceability recovery. We first extract and filter biterms from all source, intermediate, and target artifacts. We then use the consensual biterms from the intermediate artifacts to enrich the texts of both source and target artifacts, and finally deduce outer and inner-transitive links to adjust text similarities between source and target artifacts. We conducted a comprehensive empirical evaluation based on five systems widely used in other literature to show that our approach can outperform four state-of-the-art approaches in Average Precision over 15% and Mean Average Precision over 10% on average. Hongyu Kuang, Wesley K. G. Assunção, Christoph Mayr-Dorn, Guoping Rong, He Zhang 0001, Xiaoxing Ma, Alexander Egyed |
ICSE | 2 |
| 2024 | Mining Pull Requests to Detect Process Anomalies in Open Source Software DevelopmentabstractTrustworthy Open Source Software (OSS) development processes are the basis that secures the long-term trustworthiness of software projects and products. With the aim to investigate the trustworthiness of the Pull Request (PR) process, the common model of collaborative development in OSS community, we exploit process mining to identify and analyze the normal and anomalous patterns of PR processes, and propose our approach to identifying anomalies from both control-flow and semantic aspects, and then to analyze and synthesize the root causes of the identified anomalies. We analyze 17531 PRs of 18 OSS projects on GitHub, extracting 26 root causes of control-flow anomalies and 19 root causes of semantic anomalies. We find that most PRs can hardly contain both semantic anomalies and control-flow anomalies, and the internal custom rules in projects may be the key causes for the identified anomalous PRs. We further discover and analyze the patterns of normal PR processes. We find that PRs in the non-fork model (42%) are far more likely than the fork model (5%) to bypass the review process, indicating a higher potential risk. Besides, we analyzed nine poisoned projects whose PR practices were indeed worse. Given the complex and diverse PR processes in OSS community, the proposed approach can help identify and understand not only anomalous PRs but also normal PRs, which offers early risk indications of suspicious incidents (such as poisoning) to OSS supply chain. Bohan Liu 0003, He Zhang 0001, Weigang Ma, Hongyu Kuang, Jinwei Xu, Shan Gao 0009 |
ICSE | 4 |
| 2024 | AVIATE: Exploiting Translation Variants of Artifacts to Improve IR-based Traceability Recovery in Bilingual Software ProjectsabstractTraceability plays a vital role in facilitating various software development activities by establishing the traces between different types of artifacts (e.g., issues and commits in software repositories). Among the explorations for automated traceability recovery, the IR (Information Retrieval)-based approaches leverage textual similarity to measure the likelihood of traces between artifacts and show advantages in many scenarios. However, the globalization of software development has introduced new challenges, such as the possible multilingualism on the same concept (e.g., "[SEE PDF]" vs. "attribute") in the artifact texts, thus significantly hampering the performance of IR-based approaches. Existing research has shown that machine translation can help address the term inconsistency in bilingual projects. However, the translation can also bring in synonymous terms that are not consistent with those in the bilingual projects (e.g., another translation of "[SEE PDF]" as "property"). Therefore, we propose an enhancement strategy called AVIATE that exploits translation variants from different translators by utilizing the word pairs that appear simultaneously across the translation variants from different kinds artifacts (a.k.a. consensual biterms). We use these biterms to first enrich the artifact texts, and then to enhance the calculated IR values for improving IR-based trace-ability recovery for bilingual software projects. The experiments on 17 bilingual projects (involving English and 4 other languages) demonstrate that AVIATE significantly outperformed the IR-based approach with machine translation (the state-of-the-art in this field) with an average increase of 16.67 in Average Precision (31.43%) and 8.38 (11.22%) in Mean Average Precision, indicating its effectiveness in addressing the challenges of multilingual traceability recovery. Yiding Ren, Hongyu Kuang, Xiaoxing Ma, Guoping Rong, Dong Shao, He Zhang 0001 |
ASE | 3 |
| 2024 | VulCausal: Robust Vulnerability Detection Using Neural Network Models from a Causal Perspective
Hongyu Kuang, Jingjing Zhang 0005, Long Zhang 0004, Lin Yang 0031 |
KSEM (3) | 1 |
| 2024 | A Review on Binary Code Analysis Datasets
Shuguang Song, Hongyu Kuang |
WASA (3) | 4 |
| 2024 | MRC-VulLoc: Software source code vulnerability localization based on multi-choice reading comprehension
Gaigai Tang, Lin Yang 0031, Long Zhang 0004, Hongyu Kuang |
Comput. Secur. | 4 |
| 2024 | Distilling Quality Enhancing Comments From Code Reviews to Underpin Reviewer RecommendationabstractCode review is an important practice in software development. One of its main objectives is for the assurance of code quality. For this purpose, the efficacy of code review is subject to the credibility of reviewers, i.e., reviewers who have demonstrated strong evidence of previously making quality-enhancing comments are more credible than those who have not. Code reviewer recommendation (CRR) is designed to assist in recommending suitable reviewers for a specific objective and, in this context, assurance of code quality. Its performance is susceptible to the relevance of its training dataset to this objective, composed of all reviewers’ historical review comments, which, however, often contains a plethora of comments that are irrelevant to the enhancement of code quality. Furthermore, recommendation accuracy has been adopted as the sole metric to evaluate a recommender's performance, which is inadequate as it does not take reviewers’ relevant credibility into consideration. These two issues form the ground truth problem in CRR as they both originate from the relevance of dataset used to train and evaluate CRR algorithms. To tackle this problem, we first propose the concept of Quality-Enhancing Review Comments (QERC), which includes three types of comments - change-triggering inline comments, informative general comments, and approve-to-merge comments. We then devise a set of algorithms and procedures to obtain a distilled dataset by applyingQERCto the original dataset. We finally introduce a new metric – reviewer's credibility for quality enhancement (RCQE) – as a complementary metric to recommendation accuracy for evaluating the performance of recommenders. To validate the proposed QERC-based approach to CRR, we conduct empirical studies using real data from seven projects containing over 82K pull requests and 346K review comments. Results show that: (a)QERCcan effectively address the ground truth problem by distilling quality-enhancing comments from the dataset containing original code reviews, (b)QERCcan assist recommenders in finding highly credible reviewers at a slight cost of recommendation accuracy, and (c) even “wrong” recommendations using the distilled dataset are likely to be more credible than those using the original dataset. Guoping Rong, Yongda Yu, He Zhang 0001, Haifeng Shen, Dong Shao, Hongyu Kuang, Zhao Wei, Juhong Wang |
IEEE Trans. Software Eng. | 7 |
| 2023 | RL-Based CEP Operator Placement Method on Edge Networks Using Response Time Feedback
Yuyou Wang, Hao Hu 0001, Hongyu Kuang, Chenyou Fan, Liang Wang 0006, XianPing Tao |
WISA | 3 |
| 2023 | How Do Developers' Profiles and Experiences Influence their Logging Practices? An Empirical Study of Industrial PractitionersabstractLogs record the behavioral data of running programs and are typically generated by executing log statements. Software developers generally carry out logging practices with clear intentions and associated concerns (I&Cs). However, I&Cs may not be properly fulfilled in source code as log placement - specifically determination of a log statement's context and content - is often susceptible to an individual's profile and experience. Some industrial studies have been conducted to discern developers' main logging I&Cs and the way I&Cs are fulfilled. However, the findings are only based on the developers from a single company in each individual study and hence have limited generalizability. More importantly, there lacks a comprehensive and deep understanding of the relationships between developers' profiles and experiences and their logging practices from a wider perspective. To fill this significant gap, we conducted an empirical study using mixed methods comprising questionnaire surveys, semi-structured interviews, and code analyses with practitioners from a wide range of companies across a variety of industrial domains. Results reveal that while developers share common logging I&Cs and conduct logging practices mainly in the coding stage, their profiles and experiences profoundly influence their logging I&Cs and the way the I&Cs are fulfilled. These findings pave the way to facilitate the acceptance of important logging I&Cs and the adoption of good logging practices by developers Guoping Rong, Shenghui Gu, Haifeng Shen, He Zhang 0001, Hongyu Kuang |
ICSE | 5 |
| 2023 | Leveraging User-Defined Identifiers for Counterfactual Data Generation in Source Code Vulnerability DetectionabstractSoftware vulnerability detection is a critical aspect of ensuring the security and reliability of software systems. However, traditional vulnerability detection approaches often have limitations due to the scarcity and need for more diversity in labeled data. This research introduces a novel approach to overcome these challenges by utilizing user-defined identifiers in the source code to generate counterfactual training data. User-defined identffiers, such as variable and function names, contain essential information about the intentions and logic of the program. By perturbing these identifiers while maintaining the syntactic and semantic structure of the code, we create a diverse set of counterfactual examples that simulate potential vulnerabilities. When combined with existing labeled data, these counterfactual examples enrich the training process for vulnerability detection models. To evaluate the effectiveness of our approach, we conduct experiments on various datasets, achieving state-of-the-art performance on the VulDeePecker and Draper datasets. Our approach also outperforms models that utilize the same pre-trained language model in terms of accuracy. Hongyu Kuang, Long Zhang 0004, Gaigai Tang, Lin Yang 0031 |
SCAM | 1 |
| 2022 | The Influence of Sponsorship on Open-Source Software Developers' Activities on GitHubabstractStudies on the OSS communities have shown that financial supports are critical to OSS developers and projects to maintain their progress and sustainability. However, there were few developers being paid directly for maintaining OSS projects in the past. The GitHub Sponsors program that brings financial supports to the general OSS developers in GitHub-the world's largest OSS platform may make a difference on this situation in the future. In this paper, we present a data set on GitHub Sponsors and conduct a data-driven study to analyze the participants of the program and the impact of sponsorships to developers' activities and their projects' outcomes and qualities. The results of our survey suggest that most developers state they will contribute more with sponsorships and provide some privilege for their sponsors. And through quantitative study, we find that developers make more contributions on GitHub after they got/offered sponsorships. Moreover, gaining sponsorship also has a weakly positive impact on developers' collaborators that did not get sponsorship. And not only developers, but their own or contributed projects also can be motivate by sponsorships. Our findings are useful to the community by understanding the impact of sponsorships on users' activities and projects' progress and sustainability, and helping the managers to improve the current financial support mechanism. Liang Wang 0006, Hao Hu 0001, Jing Jiang 0005, Hongyu Kuang, XianPing Tao |
COMPSAC | 5 |
| 2022 | Modeling Review History for Reviewer Recommendation: A Hypergraph ApproachabstractModern code review is a critical and indispensable practice in a pull-request development paradigm that prevails in Open Source Software (OSS) development. Finding a suitable reviewer in projects with massive participants thus becomes an increasingly challenging task. Many reviewer recommendation approaches (recommenders) have been developed to support this task which apply a similar strategy, i.e. modeling the review history first then followed by predicting/recommending a reviewer based on the model. Apparently, the better the model reflects the reality in review history, the higher recommender's performance we may expect. However, one typical scenario in a pull-request development paradigm, i.e. one Pull-Request (PR) (such as a revision or addition submitted by a contributor) may have multiple reviewers and they may impact each other through publicly posted comments, has not been modeled well in existing recommenders. We adopted the hypergraph technique to model this high-order relationship (i.e. one PR with multiple reviewers herein) and developed a new recommender, namely HGRec, which is evaluated by 12 OSS projects with more than 87K PRs, 680K comments in terms of accuracy and recommendation distribution. The results indicate that HGRec outperforms the state-of-the-art recommenders on recommendation accuracy. Besides, among the top three accurate recommenders, HGRec is more likely to recommend a diversity of reviewers, which can help to relieve the core reviewers' workload congestion issue. Moreover, since HGRec is based on hypergraph, which is a natural and interpretable representation to model review history, it is easy to accommodate more types of entities and realistic relationships in modern code review scenarios. As the first attempt, this study reveals the potentials of hypergraph on advancing the pragmatic solutions for code reviewer recommendation. Guoping Rong, Lanxin Yang, Fuli Zhang, Hongyu Kuang, He Zhang 0001 |
ICSE | 5 |
| 2022 | Incorporating Pre-trained Transformer Models into TextCNN for Sentiment Analysis on Software Engineering TextsabstractSoftware information sites (e.g., Jira, Stack Overflow) are now wide-ly used in software development. These online platforms for collaborative development preserve a large amount of Software Engineering (SE) texts. These texts enable researchers to detect developers’ attitudes toward their daily development by analyzing the sentiments expressed in the texts. Unfortunately, recent works reported that neither off-the-shelf tools nor SE-specified tools for sentiment analysis on SE texts can provide satisfying and reliable results. In this paper, we propose to incorporate pre-trained transformer models into the sentence-classification oriented deep learning framework named TextCNN to better capture the unique expression of sentiments in SE texts. Specifically, we introduce an optimized BERT model named RoBERTa as the word embedding layer of TextCNN, along with additional residual connections between RoBERTa and TextCNN for better cooperation in our training framework. An empirical evaluation based on four datasets from different software information sites shows that our training framework can achieve overall better accuracy and generalizability than the four baselines. Xiaobo Shi, Hongyu Kuang, Xiaoxing Ma, Guoping Rong, Dong Shao, He Zhang 0001 |
Internetware | 4 |
| 2022 | Using Consensual Biterms from Text Structures of Requirements and Code to Improve IR-Based Traceability RecoveryabstractTraceability approves trace links among software artifacts based on whether two artifacts are related by system functionalities. The traces are valuable for software development, but are difficult to obtain manually. To cope with the costly and fallible manual recovery, automated approaches are proposed to recover traces through textual similarities among software artifacts, such as those based on Information Retrieval (IR). However, the low quality & quantity of artifact texts negatively impact the calculated IR values, thus greatly hindering the performance of IR-based approaches. In this study, we propose to extract co-occurred word pairs from the text structures of both requirements and code (i.e., consensual biterms) to improve IR-based traceability recovery. We first collect a set of biterms based on the part-of-speech of requirement texts, and then filter them through the code texts. We then use these consensual biterms to both enrich the input corpus for IR techniques and enhance the calculations of IR values. A nine-system-based evaluation shows that in general, when solely used to enhance IR techniques, our approach can outperform pure IR-based approaches and another baseline by 21.9% & 21.8% in AP, and 9.3% & 7.2% in MAP, respectively. Moreover, when used to collaborate with another enhancing strategy from different perspectives, it can outperform this baseline by 5.9% in AP and 4.8% in MAP. Hongyu Kuang, Xiaoxing Ma, Alexander Egyed, Patrick Mäder, Guoping Rong, Dong Shao, He Zhang 0001 |
ASE | 2 |
| 2022 | Automated Reliability Analysis of Redundancy Architectures Using Statistical Model Checking
Hongbin He, Hongyu Kuang, Lin Yang 0031, Qiang Wang 0020, Weipeng Cao |
KSEM (3) | 2 |
| 2022 | Semi-supervised pre-processing for learning-based traceability framework on real-world software projectsabstractThe traceability of software artifacts has been recognized as an important factor to support various activities in software development processes. However, traceability can be difficult and time-consuming to create and maintain manually, thereby automated approaches have gained much attention. Unfortunately, existing automated approaches for traceability suffer from practical issues. This paper aims to gain an understanding of the potential challenges for the underperforming of the state-of-the-art, ML-based trace link classifiers applied in real-world projects. By investigating different industrial datasets, we found that two critical (and classic) challenges, i.e. data imbalance and sparse problems, lie in real-world projects’ traceability automation. To overcome these challenges, we developed a framework called SPLINT to incorporate hybrid textual similarity measures and semi-supervised learning strategies as enhancements to the learning-based traceability approaches. We carried out experiments with six open-source platforms and ten industry datasets. The results confirm that SPLINT is able to operate at higher performance on two communities’ datasets. Specifically, the industrial datasets, which significantly suffer from data imbalance and sparsity problems, show an increase in F2-score over 14% and AUC over 8% on average. The adjusted class-balancing and self-training policies used in SPLINT (CBST-Adjust) also work effectively for the selection of pseudo-labels on minor classes from unlabeled trace sets, demonstrating SPLINT’s practicability. Liming Dong 0001, He Zhang 0001, Zhiluo Weng, Hongyu Kuang |
ESEC/SIGSOFT FSE | 5 |
| 2022 | Propagating frugal user feedback through closeness of code dependencies to improve IR-based traceability recovery
Hongyu Kuang, Xiaoxing Ma, Hao Hu 0001, Jian Lu 0001, Patrick Mäder, Alexander Egyed |
Empir. Softw. Eng. | 2 |
| 2022 | Automated detection on the security of the linked-list operations
Hongyu Kuang, Jian Wang 0020, Ruilin Li 0002, Chao Feng 0002, Yunfei Su |
Frontiers Comput. Sci. | 1 |
| 2021 | Exploiting the Unique Expression for Improved Sentiment Analysis in Software Engineering TextabstractSentiment analysis on software engineering (SE) texts has been widely used in the SE research, such as evaluating app reviews or analyzing developers' sentiments in commit messages. To better support the use of automated sentiment analysis for SE tasks, researchers built an SE-domain-specified sentiment dictionary to further improve the accuracy of the results. Unfortunately, recent work reported that current mainstream tools for sentiment analysis still cannot provide reliable results when analyzing the sentiments in SE texts. We suggest that the reason for this situation is because the way of expressing sentiments in SE texts is largely different from the way in social network or movie comments. In this paper, we propose to improve sentiment analysis in SE texts by using sentence structures, a different perspective from building a domain dictionary. Specifically, we use sentence structures to first identify whether the author is expressing her sentiment in a given clause of an SE text, and to further adjust the calculation of sentiments which are confirmed in the clause. An empirical evaluation based on four different datasets shows that our approach can outperform two dictionary-based baseline approaches, and is more generalizable compared to a learning-based baseline approach. Hongyu Kuang, Xiaoxing Ma, Guoping Rong, Dong Shao, He Zhang 0001 |
ICPC | 3 |
| 2019 | Using frugal user feedback with closeness analysis on code to improve IR-based traceability recoveryabstractTraceability recovery allows developers to extract and comprehend the trace links among software artifacts (e.g., requirements and code). These trace links can provide important support to software maintenance and evolution tasks. Information Retrieval (IR) is now widely accepted as the key technique of semi-automatic tools to recover candidate trace links based on textual similarities among artifacts. However, the vocabulary mismatch problem between different artifacts hinders the performance of these IR-based approaches. Thus, a growing body of enhancing strategies were proposed based on user feedback. They allow to adjust the textual similarities of candidate links after users accept or reject part of these links. Recently, several approaches successfully used this strategy to improve the performance of IR-based traceability recovery. However, these approaches require a large amount of user feedback, which is infeasible in practice. In this paper, we propose to improve IR-based traceability recovery by introducing only a small amount of user feedback into the closeness analysis on call and data dependencies in code. Specifically, our approach iteratively asks users to verify a chosen candidate link based on the quantified functional similarity for each code dependency (called closeness) and the generated IR values. The verified link is then used as the input to re-rank the unverified candidate links. An empirical evaluation based on five real-world systems shows that our approach can outperform four baseline approaches by using only a small amount of user feedback. Hongyu Kuang, Hao Hu 0001, Xiaoxing Ma, Jian Lu 0001, Patrick Mäder, Alexander Egyed |
ICPC | 1 |
| 2018 | Response Time Aware Operator Placement for Complex Event Processing in Edge Computing
Xinchen Cai, Hongyu Kuang, Hao Hu 0001, Wei Song 0003, Jian Lu 0001 |
ICSOC | 2 |
| 2017 | Parallelized Mobility-Aware Complex Event ProcessingabstractThe concept of complex event processing (CEP) and complex-event-aware service have been extensively studied to retrieve relevant information from massive amount of realtime streaming events. In mobile environment, the Mobilityaware CEP (MCEP) system was proposed to address the issue of synchronization problem between different query ranges and MCEP operators. We noticed that MCEP systems lack the ability to process event in parallel and scale out when system load is high. In this paper, we proposed a parallel architecture for MCEP. The architecture can handle the synchronization problem and guarantee the correctness of event processing result. We also proposed a scaling strategy that can automatically scale out operators while ensures semantic transparency. An empirical evaluation based on up to 10 ViMs demonstrated that our approach is able to achieve higher throughput while keeping the MCEP synchronization mechanism valid. Yuhao Gong, Hongyu Kuang, Xinchen Cai, Hao Hu 0001, Wei Song 0003, Jian Lu 0001 |
ICWS | 2 |
| 2017 | Analyzing closeness of code dependencies for improving IR-based Traceability RecoveryabstractInformation Retrieval (IR) identifies trace links based on textual similarities among software artifacts. However, the vocabulary mismatch problem between different artifacts hinders the performance of IR-based approaches. A growing body of work addresses this issue by combining IR techniques with code dependency analysis such as method calls. However, so far the performance of combined approaches is highly dependent to the correctness of IR techniques and does not take full advantage of the code dependency analysis. In this paper, we combine IR techniques with closeness analysis to improve IR-based traceability recovery. Specifically, we quantify and utilize the “closeness” for each call and data dependency between two classes to improve rankings of traceability candidate lists. An empirical evaluation based on three real-world systems suggests that our approach outperforms three baseline approaches. Hongyu Kuang, Jia Nie, Hao Hu 0001, Patrick Rempel, Jian Lu 0001, Alexander Egyed, Patrick Mäder |
SANER | 1 |
| 2015 | Can method data dependencies support the assessment of traceability between requirements and source code?abstractRequirements traceability benefits many software engineering activities, such as change impact analysis and risk assessment. However, these activities require complete and correct traceability links which is not trivial, making traceability assessment an important field of study. In recent years, requirements traceability research has focused on using call dependencies within source code to understand how code properties contribute to the implementation of a requirement and to assess whether traceability links are correct and complete. These approaches largely ignore the role of existing data dependencies within the source code. That is, methods may never call each other, but may still depend upon another by sharing data. We identified five research questions and validated them on five software systems, covering 4 to 72 KLOC. We found that data dependencies are as relevant as call dependencies for assessing requirements traceability. Even more interesting, our analyses show that data dependencies complement call dependencies in the assessment. These findings have strong implications on code understanding, including trace capture, maintenance, and validation techniques. Copyright © 2015 John Wiley & Sons, Ltd. Hongyu Kuang, Patrick Mäder, Hao Hu 0001, Achraf Ghabi, LiGuo Huang, Jian Lu 0001, Alexander Egyed |
J. Softw. Evol. Process. | 1 |
| 2012 | Do data dependencies in source code complement call dependencies for understanding requirements traceability?abstractIt is common practice for requirements traceability research to consider method call dependencies within the source code (e.g., fan-in/fan-out analyses). However, current approaches largely ignore the role of data. The question this paper investigates is whether data dependencies have similar relationships to requirements as do call dependencies. For example, if two methods do not call one another, but do have access to the same data then is this information relevant? We formulated several research questions and validated them on three large software systems, covering about 120 KLOC. Our findings are that data relationships are roughly equally relevant to understanding the relationship to requirements traces than calling dependencies. However, most interestingly, our analyses show that data dependencies complement call dependencies. These findings have strong implications on all forms of code understanding, including trace capture, maintenance, and validation techniques (e.g., information retrieval). Hongyu Kuang, Patrick Mäder, Hao Hu 0001, Achraf Ghabi, LiGuo Huang, Jian Lu 0001, Alexander Egyed |
ICSM | 1 |