Rong Chen 0003

dblp:22/6904-3 · DBLP profile ↗
← Back
57ranked-venue papers
3as first author
36since 2021 · last 2027
0000-0001-5848-6398ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 26 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2027 Towards discernible prototypes in heterogeneous federated learning via neural collapse
Rong Chen 0003, Shuai Hong, Shilong Jing, Hengyi Lv
Expert Syst. Appl.2
2026 Estimating Uncertainty in Line-Level Defect Prediction via Perceptual Borderline Oversampling
abstract
Software defect prediction aims to identify potentially defective software modules using various techniques, while fine-grained line-level defect prediction can pinpoint defective lines of code. This helps developers promptly discover and fix errors, thereby enhancing the efficiency of testing and code review. However, previous studies often overlook the impact of characterizing noise and the skewed distribution of defect knowledge in software projects, making it difficult for current methods to achieve satisfactory accuracy and cost-effectiveness in software defect prediction. To address these challenges, we propose a model named EU-LLDP, which effectively resolves the issue of low cost-effectiveness in line-level defect prediction models. Specifically, the EU-LLDP model consists of two main components: the defect mining component mines the most valuable defect knowledge from numerous software defects using prediction probability matrices, noise labels, and the borderline information of code vectorizations. The adaptive resampling component samples valuable defect knowledge through the density distribution of defect knowledge, thereby making full use of existing defect knowledge and improving the cost-effectiveness of line-level software defect prediction models. Seven comprehensive experiments were conducted on 32 defect datasets from 9 Java open source systems using file-level prediction models and line-level defect prediction models to evaluate the effectiveness of the EU-LLDP model. The EU-LLDP model improves the state-of-the-art file-level defect prediction model in terms of Balanced Accuracy by 9.87%, the MCC by 38.09%, and enhances the state-of-the-art line-level defect prediction method in terms of Recall@Top20%LOC by 44.16%, and Effort@Top20%Recall by 17.62%. These results fully demonstrate the effectiveness of EU-LLDP in improving the accuracy and cost-effectiveness of Software defect prediction.
Shikai Guo, Hui Li 0014, Rong Chen 0003
ACM Trans. Softw. Eng. Methodol.5
2026 Arbitrary-Scale Point Cloud Upsampling With Saliency-Aware Implicit Surface Guidance
abstract
Despite significant progress in point cloud upsampling, most existing methods rely heavily on supervised training with paired data, which are often difficult to acquire. Moreover, the inherent lack of explicit connectivity in point clouds makes it difficult to achieve both continuous and uniform densification while accurately recovering fine geometric structures. To address these problems, we propose an upsampling model that treats this task as saliency-aware implicit surface sampling, enabling self-supervised and fine-grained point densification. Central to our idea is correlating implicit surface reconstruction with salient point identification, and carrying out sampling on the saliency-aware surface representation. Motivated by this, we introduce a salient point detector along with a corresponding implicit surface-based interpolator, and a geometry filter, upon which we develop a three-stage architecture consisting of a pre-trained saliency guidance block, a saliency-aware enhancer, and an upsampler. The guidance block captures meaningful shape patterns to prevent detail loss and incomplete recovery, while the enhancer facilitates detail enhancement in complex salient regions. These components are integrated with the upsampler to generate dense results that accurately retain both global shape and meticulous structures. In comparison with state-of-the-art methods, our model significantly improves the upsampling quality. Extensive experiments conducted on various datasets, comprising both synthetic and real-world captured shapes, demonstrate the flexibility and availability of our method in processing the point clouds across different scales and distributions.
Yanzhe Liu, Rong Chen 0003, Yushi Li
IEEE Trans. Vis. Comput. Graph.2
2025 ShadowCraft-NeRF: Occlusion and Shadow Mitigation via SAM-Guided NeRF
Yushi Li, Yunyao Shen, Rong Chen 0003, Xiao-Bo Jin, Along Jin, Yu Han 0001
CASA4
2025 XGFu: Enhancing low-light visualization by feature and graph fusion of multiple artificial exposure images
Sihai Qiao, Ming An, Rong Chen 0003, Yushi Li
Expert Syst. Appl.4
2025 Unsupervised single-image dehazing via self-guided inverse-retinex GAN
Rong Chen 0003, Yushi Li
Multim. Syst.2
2025 Vehicle-Level Fairness-Oriented Constrained Multi-Agent Reinforcement Learning for Adaptive Traffic Signal Control
abstract
Multi-agent Reinforcement Learning (MARL) has shown considerable promise in enhancing the efficiency of adaptive traffic signal control (ATSC) systems. However, existing MARL approaches primarily focus on optimizing overall traffic flow, often overlooking the issue of fairness in vehicle waiting times. Considering that there is no need to strive for the ultimate fairness, this paper models the ATSC problem as a Constrained Partially Observable Markov Game (CPOMG), where fairness is modeled as a constraint on the maximum waiting time of vehicles on lanes of intersections instead of a reward term that pursues maximization. CPOMG aims to find a cooperative control policy with optimal traffic efficiency within the constrained solution space by multiple agents. On this basis, this paper proposes a new centralized training and decentralized execution cooperative MARL method, i.e., vehicle-level fairness multi-agent proximity policy optimization (VF-MAPPO). VF-MAPPO leverages a centralized trained global Critic Network to estimate the average vehicle traffic efficiency and vehicle maximum waiting time, and an Actor Network shared by all intersections for decentralized execution, which converts the optimization problem with constraints to an unconstrained optimization objective through the Lagrange multiplier method and adopts proximity policy optimization during training. Additionally, VF-MAPPO incorporates spatial-temporal graph attention in the Critic network to efficiently extract state representations in multi-intersection environments. We qualitatively analyzed the monotonic improvement guarantee of VF-MAPPO. Extensive experimental validation across two real-world and one synthetic scenarios substantiates that VF-MAPPO enhances vehicle-level fairness and maintains average traffic efficiency, surpassing state-of-the-art methods.
Wanting Liu, Chengwei Zhang 0001, Wanqing Fang, Kailing Zhou, Furui Zhan, Qi Wang 0044, Wanli Xue, Rong Chen 0003
IEEE Trans. Intell. Transp. Syst.9
2025 Sequential Decision MARL for Adaptive Traffic Signal Control With Different Intersections Priorities
abstract
Existing multi-agent reinforcement learning (MARL) in adaptive traffic signal control (ATSC) typically models cooperative control of multiple intersections as a cooperative Markov game, optimizing the average traffic efficiency of intersections with the same emphasis. However, it is insufficient to meet the requirements in real ATSC scenarios when all intersections are treated equally. To this end, this work proposes the Captain-Member Markov Game (CM-MG) that considers the different priorities between intersections. CM-MG categorizes intersections into special and ordinary intersections, controlled by captain agents and member agents to optimize the traffic efficiency of local and overall road networks, respectively. The cooperative requirements of CM-MG are achieved through a sequential decision-making principle. Captains have priority in choosing actions, and members make decisions sequentially, following the breadth-first traversal order in the road network after obtaining their precursors’ intentions. Then, a cooperative MARL algorithm, i.e, Sequential Decision Deep Graph Network (GNSD-Light), is proposed to learn the optimal joint policy that meets the learning goals of both captains and members. To be unrestricted by the scales of intersections, GNSD-Light adopts an autoregressive framework where all agents make decisions sequentially in a predetermined order by sharing the same decision model. In addition, to obtain sufficient state representation, two relative position encoding-based spatiotemporal representation modules are designed for GNSD-Light based on the characteristics of ATSC scenarios. Finally, through adequate experiments and qualitative analysis, we have confirmed that our method effectively balances traffic efficiency among both the overall road network and special intersections.
Wanting Liu, Chengwei Zhang 0001, Kailing Zhou, Furui Zhan, Wanli Xue, Rong Chen 0003
IEEE Trans. Intell. Transp. Syst.7
2025 DSANet: Dynamic and Structure-Aware GCN for Sparse and Incomplete Point Cloud Learning
abstract
Learning 3-D structures from incomplete point clouds with extreme sparsity and random distributions is a challenge since it is difficult to infer topological connectivity and structural details from fragmentary representations. Missing large portions of informative structures further aggravates this problem. To overcome this, a novel graph convolutional network (GCN) called dynamic and structure-aware NETwork (DSANet) is presented in this article. This framework is formulated based on a pyramidic auto-encoder (AE) architecture to address accurate structure reconstruction on the sparse and incomplete point clouds. A PointNet-like neural network is applied as the encoder to efficiently aggregate the global representations of coarse point clouds. On the decoder side, we design a dynamic graph learning module with a structure-aware attention (SAA) to take advantage of the topology relationships maintained in the dynamic latent graph. Relying on gradually unfolding the extracted representation into a sequence of graphs, DSANet is able to reconstruct complicated point clouds with rich and descriptive details. To associate analogous structure awareness with semantic estimation, we further propose a mechanism, called structure similarity assessment (SSA). This method allows our model to surmise semantic homogeneity in an unsupervised manner. Finally, we optimize the proposed model by minimizing a new distortion-aware objective end-to-end. Extensive qualitative and quantitative experiments demonstrate the impressive performance of our model in reconstructing unbroken 3-D shapes from deficient point clouds and preserving semantic relationships among different regional structures.
Yushi Li, George Baciu, Rong Chen 0003, Chenhui Li 0001, Hao Wang 0003, Yushan Pan, Weiping Ding 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Making Fault Localization in Online Service Systems More Actionable and Interpretable
abstract
Online service systems struggle with accurately and quickly pinpointing and resolving failures within their intricate systems, and it therefore emerges the solutions for fault localization in the code. However, the previous fault localization models suffer from low localization accuracy and poor interpretability due to the complex dependencies among fault characteristics in industrial practice. To address this issue, challenges brought by the long-distance dependencies among fault features and the unbalanced distribution of fault knowledge, and to improve the interpretability of the model, we present a fault localization model in online service systems more actionable and interpretable, named FL-AIer. Specifically, FL-AIer consists of two components: the feature encoding component and the fault localization component. The feature encoding component utilizes graph attention networks to capture the complex spatio-temporal dependencies within fault features. Then, the fault localization component adopts a three-stage approach, leveraging a multi-attention mechanism to identify and prioritize the most relevant fault features for precise localization. Additionally, the Fault Knowledge Balancing module it contains introduces a weighted Kullback-Leibler divergence loss function to ensure that the model pays adequate attention to all fault features, addressing the issue of imbalanced fault knowledge distribution and enhancing localization performance. We conducted extensive experiments on four datasets, and the results demonstrated that FL-AIer effectively addressing the challenges of fault localization in online system environments, and consistently outperforms the state-of-the-art methods across various evaluation metrics such as A@1, A@2, A@3, A@5, and MAR. For instance, FL-AIer achieves significant improvements of 5.82%, 10.77%, 4.20%, and 15.56% on the A@1 metric, respectively. These results fully demonstrate the excellent effectiveness of FL-AIer in effectively addressing the challenges of fault localization in online system environments, surpassing the performance of existing state-of-the-art methods.
Ke Xv, Shikai Guo, Hui Li 0014, Rong Chen 0003, He Jiang 0001
ACM Trans. Softw. Eng. Methodol.5
2025 Context-based Transfer Learning for Structuring Fault Localization and Program Repair Automation
abstract
Automated software debugging plays a crucial role in aiding software developers to swiftly identify and attempt to rectify faults, thereby significantly reducing developers’ workload. Previous researches have predominantly relied on simplistic semantic deep learning or statistical analysis methods to locate faulty statements in diverse projects. However, code repositories often consist of lengthy sequences with long-distance dependencies, posing challenges for accurately modeling fault localization using these methods. In addition, the lack of joint reasoning among various faults prevents existing models from deeply capturing fault information. To address these challenges, we propose a method named CodeHealer to achieve accurate fault localization and program repair. CodeHealer comprises three components: a Deep Semantic Information Extraction Component that effectively extracts deep semantic features from suspicious code statements using classifiers based on Joint-attention mechanisms; a Suspicious Statement Ranking Component that combines various fault localization features and employs multilayer perceptrons to derive multidimensional vectors of suspicion values; and a Fault Repair Component that, based on ranked suspicious statements generated by fault localization, adopts a top-down approach using multiple classifiers based on Co-teaching mechanisms to select repair templates and generate patches. The experimental results indicate that when applied to fault localization, CodeHealer outperforms the best baseline method with improvements of 11.4%, 2.7%, and 1.6% on Top-1/3/5 metrics, respectively. It also reduces the MFR and MAR by 9.8% and 2.1%, where lower values denote better fault localization effectiveness. Additionally, in automated software debugging, CodeHealer fixes an additional 6 faults compared to the current best method, totaling 53 faults repaired.
Lehuan Zhang, Shikai Guo, Hui Li 0014, Yu Chai, Rong Chen 0003, He Jiang 0001
ACM Trans. Softw. Eng. Methodol.6
2025 Line-Level Defect Prediction by Capturing Code Contexts With Graph Convolutional Networks
abstract
Software defect prediction refers to the systematic analysis and review of software using various approaches and tools to identify potential defects or errors. Software defect prediction aids developers in swiftly identifying defects and optimizing development resource allocation, thus enhancing software quality and reliability. Previous defect prediction approaches still face two main limitations: 1) lacking of contextual semantic information and 2) Ignoring the joint reasoning between different granularities of defect predictions. In response to these challenges, we propose LineDef, a line-level defect prediction approach by capturing code contexts with graph convolutional networks. Specifically, LineDef comprises three components: the token embedding component, the graph extraction component, and the multi-granularity defect prediction component. The token embedding component maps each token to a vector to obtain a high-dimensional semantic feature representation of the token. Subsequently, the graph extraction component utilizes a sliding window to extract line-level and token-level graphs, addressing the challenge of capturing contextual semantic relationships in the code. Finally, the multi-granularity defect prediction component leverages graph convolutional layers and attention mechanisms to acquire prediction labels and risk scores, thereby achieving file-level and line-level defect prediction. Experimental studies on 32 datasets across 9 different software projects show that LineDef exhibits significantly enhanced balanced accuracy, ranging from 15.61% to 45.20%, compared to state-of-the-art file-level defect prediction approaches, and a remarkable cost-effectiveness improvement ranging from 15.32% to 278%, compared to state-of-the-art line-level defect prediction approaches. These results demonstrate that LineDef approach can extract more comprehensive information from lines of code for defect prediction.
Shouyu Yin, Shikai Guo, Hui Li 0014, Rong Chen 0003, He Jiang 0001
IEEE Trans. Software Eng.5
2025 Point cloud upsampling via a coarse-to-fine network with transformer-encoder
Yixi Li, Yanzhe Liu, Rong Chen 0003, Hui Li 0014
Vis. Comput.3
2024 MGE-Net: Task-oriented Point Cloud Sampling based on Multi-scale Geometry Estimation
abstract
A large number of collaborative manufacturing tasks are directly performed on point clouds. With the growing size of point clouds, the computational demands of these tasks also increase. One possible solution is to sample the point clouds. The most commonly used sampling method is farthest point sampling, but it does not consider downstream tasks, often leading to sampling non-informative points for the tasks. With the development of neural networks, various methods have been proposed to sample point clouds in a task-oriented learning manner. However, most methods are based on generation rather than selecting a subset of point clouds. In this work, we propose a novel adaptive keypoint sampling method, called MGE-Net, that combines neural network-based learning with direct point selection based on multi-scale geometry estimation. In addition, we design a feature extraction module based on multi-scale attention graph convolution to provide accurate information for subsequent keypoint detection. Relying on the contribution of point clouds to the task, our framework aims to sample a subset of point clouds specifically optimized for downstream tasks. Both qualitative and quantitative experimental results demonstrate that our sampling method exhibits superior performance in common point cloud classification and segmentation tasks.
Weiliang Zeng, Yushi Li, Rong Chen 0003, Rong Xiang, Jinghang Gu
CSCWD3
2024 SPU-PMD: Self-Supervised Point Cloud Upsampling via Progressive Mesh Deformation
abstract
Despite the success of recent upsampling approaches, generating high-resolution point sets with uniform distribution and meticulous structures is still challenging. Unlike existing methods that only take spatial information of the raw data into account, we regard point cloud upsampling as generating dense point clouds from deformable topology. Motivated by this, we present SPU-PMD, a self-supervised topological mesh deformation network, for 3D densification. As a cascaded framework, our architecture is formu-lated by a series of coarse mesh interpolator and mesh de-formers. At each stage, the mesh interpolator first produces the initial dense point clouds via mesh interpolation, which allows the model to perceive the primitive topology better. Meanwhile, the deformer infers the morphing by estimating the movements of mesh nodes and reconstructs the de-scriptive topology structure. By associating mesh deformation with feature expansion, this module progressively re-fines point clouds' surface uniformity and structural details. To demonstrate the effectiveness of the proposed method, extensive quantitative and qualitative experiments are con-ducted on synthetic and real-scanned 3D data. Also, we compare it with state-of-the-art techniques to further illus-trate the superiority of our network. The project page is: https://github.com/lyz21/spU-PMd.
Yanzhe Liu, Rong Chen 0003, Yushi Li, Yixi Li, Xuehou Tan
CVPR2
2024 VF-Detector: Making Multi-Granularity Code Changes on Vulnerability Fix Detector Robust to Mislabeled Changes
Zhenkan Fu, Shikai Guo, Hui Li 0014, Rong Chen 0003, He Jiang 0001
IJCAI4
2024 WalkFormer: Point Cloud Completion via Guided Walks
abstract
Point clouds are often sparse and incomplete in real-world scenarios. The prevailing methods for point cloud completion typically rely on encoding the partial points and then decoding complete points from a global feature vector, which might lose the existing patterns and elaborate structures. To address these issues, we propose WalkFormer, a novel approach to predict complete point clouds through a partial deformation process. Concretely, our method samples locally dominant points based on feature similarity and moves the points to form the missing part. Since these points maintain representative information of the surrounding structures, they are appropriately selected as the starting points for multiple guided walks. Furthermore, we design a Route Transformer module to exploit and aggregate the walk information with topological relations. These guided walks facilitate the learning of long-range dependencies for predicting shape deformation. Qualitative and quantitative evaluations demonstrate that our proposed approach achieves superior performance compared to state-of-the-art methods in the 3D point cloud completion task.
Mohang Zhang, Yushi Li, Rong Chen 0003, Yushan Pan, Rong Xiang
WACV3
2024 Estimating Uncertainty in Labeled Changes by SZZ Tools on Just-In-Time Defect Prediction
abstract
The aim of Just-In-Time (JIT) defect prediction is to predict software changes that are prone to defects in a project in a timely manner, thereby improving the efficiency of software development and ensuring software quality. Identifying changes that introduce bugs is a critical task in just-in-time defect prediction, and researchers have introduced the SZZ approach and its variants to label these changes. However, it has been shown that different SZZ algorithms introduce noise to the dataset to a certain extent, which may reduce the predictive performance of the model. To address this limitation, we propose the Confident Learning Imbalance (CLI) model. The model identifies and excludes samples whose labels may be corrupted by estimating the joint distribution of noisy labels and true labels, and mitigates the impact of noisy data on the performance of the prediction model. The CLI consists of two components: identifying noisy data (Confident Learning Component) and generating a predicted probability matrix for imbalanced data (Imbalanced Data Probabilistic Prediction Component). The IDPP component generates precise predicted probabilities for each instance in the training set, while the CL component uses the generated predicted probability matrix and noise labels to clean up the noise and build a classification model. We evaluate the performance of our model through extensive experiments on a total of 126,526 changes from ten Apache open source projects, and the results show that our model outperforms the baseline methods.
Shikai Guo, Sijia Lv, Rong Chen 0003, Hui Li 0014, He Jiang 0001
ACM Trans. Softw. Eng. Methodol.5
2024 Analyzing and Detecting Information Types of Developer Live Chat Threads
abstract
Online chatrooms serve as vital platforms for information exchange among software developers. With multiple developers engaged in rapid communication and diverse conversation topics, the resulting chat messages often manifest complexity and lack structure. To enhance the efficiency of extracting information from chat threads , automatic mining techniques are introduced for thread classification. However, previous approaches still grapple with unsatisfactory classification accuracy due to two primary challenges that they struggle to adequately capture long-distance dependencies within chat threads and address the issue of category imbalance in labeled datasets. To surmount these challenges, we present a topic classification approach for chat information types named EAEChat. Specifically, EAEChat comprises three core components: the text feature encoding component captures contextual text features using a multi-head self-attention mechanism-based text feature encoder, and a siamese network is employed to mitigate overfitting caused by limited data; the data augmentation component expands a small number of categories in the training dataset using a technique tailored to developer chat messages, effectively tackling the challenge of imbalanced category distribution; the non-text feature encoding component employs a feature fusion model to integrate deep text features with manually extracted non-text features. Evaluation across three real-world projects demonstrates that EAEChat, respectively, achieves an average precision, recall, and F1-score of 0.653, 0.651, and 0.644, and it marks a significant 7.60% improvement over the state-of-the-art approaches. These findings confirm the effectiveness of our method in proficiently classifying developer chat messages in online chatrooms.
Xiuwei Shang, Shikai Guo, Yulong Li 0001, Rong Chen 0003, Hui Li 0014, He Jiang 0001
ACM Trans. Softw. Eng. Methodol.6
2024 Code Comment Inconsistency Detection Based on Confidence Learning
abstract
Code comments are a crucial source of software documentation that captures various aspects of the code. Such comments play a vital role in understanding the source code and facilitating communication between developers. However, with the iterative release of software, software projects become larger and more complex, leading to a corresponding increase in issues such as mismatched, incomplete, or outdated code comments. These inconsistencies in code comments can misguide developers and result in potential bugs, and there has been a steady rise in reports of such inconsistencies over time. Despite numerous methods being proposed for detecting code comment inconsistencies, their learning effect remains limited due to a lack of consideration for issues such as characterization noise and labeling errors in datasets. To overcome these limitations, we propose a novel approach called MCCL that first removes noise from the dataset and then detects inconsistent code comments in a timely manner, thereby enhancing the model's learning ability. Our proposed model facilitates better matching between code and comments, leading to improved development of software engineering projects. MCCL comprises two components, namely method comment detection and confidence learning denoising. The method comment detection component captures the intricate relationships between code and comments by learning their syntactic and semantic structures. It correlates the code and comments through an attention mechanism to identify how changes in the code affect the comments. Furthermore, confidence learning denoising component of MCCL identifies and removes characterization noises and labeling errors to enhance the quality of the datasets. This is achieved by implementing principles such as pruning noisy data, counting with probabilistic thresholds to estimate noise, and ranking examples to train with confidence. By effectively eliminating noise from the dataset, our model is able to more accurately learn inconsistencies between comments and source code. Our experiments on 1,518 open-source projects demonstrate that MCCL can accurately detect inconsistencies, achieving an averageF1-scoreof 82.6%. This result outperforms state-of-the-art methods by 2.4% to 28.0%. Therefore, MCCL is more effective in identifying inconsistent comments based on code changes compared to existing approaches.
Zhengkang Xu, Shikai Guo, Rong Chen 0003, Hui Li 0014, He Jiang 0001
IEEE Trans. Software Eng.4
2024 Ni-DehazeNet: representation learning via bilevel optimized architecture search for nighttime dehazing
Rong Chen 0003
Vis. Comput.3
2023 Constructing meaningful code changes via graph transformer
abstract
Abstract The rapid development of Open‐Source Software (OSS) has resulted in a significant demand for code changes to maintain OSS. Symptoms of poor design and implementation choices in code changes often occur, thus heavily hindering code reviewers to verify correctness and soundness of code changes. Researchers have investigated how to learn meaningful code changes to assist developers in anticipating changes that code reviewers may suggest for the submitted code. However, there are two main limitations to be addressed, including the limitation of long‐range dependencies of the source code and the missing syntactic structural information of the source code. To solve these limitations, a novel method is proposed, named Graph Transformer for learning meaningful Code Transformations (GTCT), to provide developers with preliminary and quick feedback when developers submit code changes, which can improve the quality of code changes and improve the efficiency of code review. GTCT comprises two components: code graph embedding and code transformation learning. To address the missing syntactic structural information of the source code limitation, the code graph embedding component captures the types and patterns of code changes by encoding the source code into a code graph structure from the lexical and syntactic representations of the source code. Subsequently, the code transformation learning component uses the multi‐head attention mechanism and positional encoding mechanism to address the long‐range dependencies limitation. Extensive experiments are conducted to evaluate the performance of GTCT by both quantitative and qualitative analyses. For the quantitative analysis, GTCT relatively outperforms the baseline on six datasets by 210%, 342.86%, 135%, 29.41%, 109.09%, and 91.67% in terms of perfect prediction. Meanwhile, the qualitative analysis shows that each type of code change by GTCT outperforms that of the baseline method in terms of bug fixed, refactoring code and others' taxonomy of code changes.
Shikai Guo, Mengxuan Li 0005, Hui Li 0014, Rong Chen 0003
IET Softw.5
2023 Optimization of Web Service Testing Task Assignment in Crowdtesting Environment
Wen-Jun Tang, Rong Chen 0003, Sheng-Jie Zheng, Shikai Guo
J. Comput. Sci. Technol.2
2023 Code samples summarization for knowledge exchange in developer community
abstract
Abstract A question title's function is to generate readable titles and describe a problem encountered by the code. Previous studies often used an end‐to‐end sequence‐to‐sequence system to generate question title's from source code. However, long‐term dependencies are often difficult to capture, and this may result in an incomplete source code representation. To address this issue, we propose a Transformer for Generating Code Title (hereinafter referred to as TGCT) model. Specifically, the TGCT model uses the position coding mechanism to model paired relationships between source terms by applying relative position representations. Multiple self‐attention mechanism components are also used to capture long‐term dependencies of the code. Comprehensive experiments on datasets from five coding languages, namely Python, Java, JavaScript, C#, and SQL, are conducted, and the results show that TGCT outperforms state‐of‐the‐art models based on the measurements of BLEU and ROUGE in general. In addition, a cross‐sectional comparison experiment was conducted to verify the effects of different model parameters, different data set sizes, position coding mechanism, and self‐attention mechanism on model results.
Shikai Guo, Zhongyan Liu, Zixuan Song, Hui Li 0014, Rong Chen 0003
Softw. Pract. Exp.5
2023 Feature transfer learning by reinforcement learning for detecting software defect
abstract
Abstract Software defects, produced inevitably in software projects, seriously affect the efficiency of software testing and maintenance. An appealing solution is the software defect prediction (SDP) that has achieved good performance in many software projects. However, the difference between features and the difference of the same feature between training data and test data may degrade defect prediction performance if such differences violate the model's assumption. To address this issue, we propose a SDP method based on feature transfer learning (FTL), which performs a transformation sequence for each feature in order to map the original features to another feature space. Specifically, FTL first uses the reinforcement learning scheme that automatically learns a strategy for transferring the potential feature knowledge from the training data. Then, we use the learned feature knowledge to inspire the transformation of the test data. The classifier is trained by the transformed training data and predicts defects for transformed test data. We evaluate the validity of FTL on 43 projects from PROMISE and NASA MDP using three classifiers, logistic regression, random forest, and Naive Bayes (NB). Experimental results indicate that FTL is better than the original classifiers and has the best performance on the NB classifier. For PROMISE, after using FTL, the average results of F1‐score, AUC, MCC are 0.601, 0.757, and 0.350 respectively, which are 24.9%, 2.6%, and 16.7% higher than the original NB classifier results. The number of projects with improved performance accounts for 83.87%, 83.87%, and 64.52%. Similarly, FTL performs well on NASA MDP. Besides, compared with four feature engineering (FE) methods, FTL achieves an excellent improvement on most projects and the average performance is also better than or close to the FE methods.
Shikai Guo, Hui Li 0014, Rong Chen 0003
Softw. Pract. Exp.6
2023 An Easy Data Augmentation Approach for Application Reviews Event Inference
abstract
Application review event inference aims to assess the effectiveness of application problems in response to user actions, which enables application developers to promptly discover and address potential issues in various applications, thereby improving their development and maintenance efficiency. Despite the development of event inference models for app reviews, which extract them as user action and app problem events and establish a relationship model between events and inference labels, the accuracy of these models is constrained due to limitations in labeling and characterizing noise and the lack of robustness and generalization. To address this challenge, we propose a model called Easy Data Augmentation for Application Reviews Event Inference (short for EDA-AREI), which comprises a denoising component, data augmentation component, and event inference prediction component. Specifically, the denoising component identifies labels and characterizes noisy data to enhance dataset quality, the data augmentation component replaces non-stop words with synonyms to increase textual diversity, and the event inference and prediction component reconstructs the classifier using denoised and augmented data. Experimental results on six datasets of one-star app reviews in the Apple App Store demonstrate that the EDA-AREI method achieves anAccuracyof 71.19%, 79.14%, 69.05%, 69.02%, 68.24% and 68.48%, respectively, representing an improvement of 0.83%–2.09% compared to state-of-the-art models. Regarding theF1-score, EDA-AREI achieves values of 71.30%, 69.93%, and 68.76% on the threshold_0.5, k-means_2, and random datasets, respectively, outperforming state-of-the-art models by 1.89%–4.02%. Furthermore, EDA-AREI achievesAUCvalues of 75.66% and 73.37% on the threshold_0.5 and k-means_2 datasets, respectively. As a result, EDA-AREI demonstrates substantial improvements inAccuracy, as well as enhancedF1-scoreandAUCacross most datasets, thereby enhancing the model's accuracy and robustness in identifying related action-problem pairs.
Shikai Guo, Haorui Lin, Jiaoru Zhao, Hui Li 0014, Rong Chen 0003, He Jiang 0001
IEEE Trans. Software Eng.5
2023 DupHunter: Detecting Duplicate Pull Requests in Fork-Based Development
abstract
The emergence of numerous fork-based development platforms facilitates the development of Open-Source Software (OSS) projects. Developers across the world can fork software projects and submit their Pull Requests (PRs) to the projects. However, as the number of forks increases, numerous duplicate PRs might be submitted. These duplicate PRs may cause extra code review workload and frustrate developers working on the projects. To detect duplicate PRs, many approaches have been proposed, which analyze the similarity of different elements in PRs. However, previous approaches still suffer from unsatisfied detection accuracy due to two challenges. That is, they ignore the syntactic structural information of text elements in PRs and lack the joint reasoning between different elements of two PRs. In this study, we propose an automated duplicate PRs detector namedDupHunter(Duplicate PRsHunter), which includes a graph embedding component and a duplicate PRs detection component to address the above challenges. The graph embedding component uses a feature graph to represent a PR. It encodes the syntactic structure and semantics of text elements (e.g., the title and the description), as well as the knowledge of non-text elements (e.g., the submission time), to address the syntactic structural information challenge. The duplicate PRs detection component tackles the joint reasoning challenge using a graph matching network, which enables the information exchange and matching across different elements of two feature graphs with an attention coefficient mechanism. Experiments on 26 open-source projects show that DupHunter achieves an averageF1-score@1value of 0.650, significantly outperforming the state-of-the-art approaches by 3.2% to 48.1%. DupHunter can accurately detect duplicate PRs, with an averagePrecision@1value of 0.922 and an averageRecall@1value of 0.502.
He Jiang 0001, Yulong Li 0001, Shikai Guo, Tao Zhang 0001, Hui Li 0014, Rong Chen 0003
IEEE Trans. Software Eng.7
2023 Single-image dehazing via depth-guided deep retinex decomposition
Rong Chen 0003
Vis. Comput.2
2022 Detecting Simulink compiler bugs via controllable zombie blocks mutation
abstract
As a popular Cyber-Physical System (CPS) development tool chain, MathWorks Simulink is widely used to prototype CPS models in safety-critical applications, e.g., aerospace and healthcare. It is crucial to ensure the correctness and reliability of Simulink compiler (i.e., the compiler module of Simulink) in practice since all CPS models depend on compilation. However, Simulink compiler testing is challenging due to millions of lines of source code and the lack of the complete formal language specification. Although several methods have been proposed to automatically test Simulink compiler, there still remains two challenges to be tackled, namely the limited variant space and the insufficient mutation diversity. To address these challenges, we propose COMBAT, a new differential testing method for Simulink compiler testing. COMBAT includes an EMI (Equivalence Modulo Input) mutation component and a diverse variant generation component. The EMI mutation component inserts assertion statements (e.g., If /While blocks) at arbitrary points of the seed CPS model. These statements break each insertion point into true and false branches. Then, COMBAT feeds all the data passed through the insertion point into the true branch to preserve the equivalence of CPS variants. In such a way, the body of the false branch could be viewed as a new variant space, thus addressing the first challenge. The diverse variant generation component uses Markov chain Monte Carlo optimization to sample the seed CPS model and generate complex mutations of long sequences of blocks in the variant space, thus addressing the second challenge. Experiments demonstrate that COMBAT significantly outperforms the state-of-the-art approaches in Simulink compiler testing. Within five months, COMBAT has reported 16 valid bugs for Simulink R2021b, of which 11 bugs have been confirmed as new bugs by MathWorks Support.
Shikai Guo, He Jiang 0001, Zhilei Ren, Zhide Zhou, Rong Chen 0003
ESEC/SIGSOFT FSE7
2022 Advertising Impression Resource Allocation Strategy with Multi-Level Budget Constraint DQN in Real-Time Bidding
Chengwei Zhang 0001, Kangjie Zheng, Wanli Xue, Tianpei Yang, Dou An, Yongqi Pi, Rong Chen 0003
Neurocomputing8
2022 A multi-objective optimization approach to package delivery by the crowd of occupied taxis
Zhifeng Zhou, Rong Chen 0003, Jian Gao 0007, Hu Xing
Knowl. Inf. Syst.2
2022 Self-admitted technical debt detection by learning its comprehensive semantics via graph neural networks
abstract
Abstract The goal of software development is to deliver software products with high quality and free from defects, but resource and time constraints often cause the developers to submit incomplete or temporary patches of codes and further bear the additional burden. Therefore, the investigations on identifying self‐admitted technical debt (SATD) to improve code quality have been conducted in recent years. However, missing syntactic structure information and the imbalance distribution bias shorten the SATD identification performance. Addressing to this issue, we present a graph neural network based SATD identification model (GNNSI) to improve the performance. Specifically, we obtain the structure information of the missing SATD in a compositional way to obtain different feature maps for different comments, and use focal loss to handle the imbalance between SATD and non‐SATD classes in the comments. Then extensive experiments on 10 open source projects are conducted, and the results show that GNNSI outperforms the baselines and can help developers to better predict SATDs.
Hui Li 0014, Rong Chen 0003, Jun Ai, Shikai Guo
Softw. Pract. Exp.4
2022 Neighborhood Cooperative Multiagent Reinforcement Learning for Adaptive Traffic Signal Control in Epidemic Regions
abstract
Nowadays, multiagent reinforcement learning (MARL) have shared significant advances in the adaptive traffic signal control (ATSC) problems. For most of the researches, agents are all isomorphic, which disregards the situation in which isomerous intersections cooperative together in a real ATSC scenario, especially in epidemic regions where different intersections have quite different levels of importance. To this end, this paper models the ATSC problem as a networked Markov game (NMG), in which agents take into account information, including traffic conditions of it and its connected neighbors. A cooperative MARL framework named neighborhood cooperative hysteretic DQN (NC-HDQN) is proposed. Specifically, for each NC-HDQN agent in the NMG, first, the framework analyses correlation degrees with their connected neighbors and weighs observations and rewards by these correlations. Second, NC-HDQN agents independently optimize their strategies on the weighted information using hysteretic DQN (HDQN), which is designed to learn optimal joint strategies in cooperative multiagent games. Third, a rule-based NC-HDQN method and a Pearson correlation coefficient based NC-HDQN method, i.e., empirical NC-HDQN (ENC-HDQN) and Pearson NC-HDQN (PNC-HDQN), respectively, are designed. The first method maps the correlation degree between two connected agents according to vehicle numbers on roads between the two agents. In contrast, the second method uses the Pearson correlation coefficient to calculate the correlation degree adaptively. Our methods are empirically evaluated in both a synthetic scenario and two real-world traffic scenarios and give better performances in almost every standard test metric for ATSC.
Chengwei Zhang 0001, Wanli Xue, Xiaofei Xie, Tianpei Yang, Rong Chen 0003
IEEE Trans. Intell. Transp. Syst.8
2021 MARL for Traffic Signal Control in Scenarios with Different Intersection Importance
Liguang Luan, Wanqing Fang, Chengwei Zhang 0001, Wanli Xue, Rong Chen 0003, Chen Sang
DAI6
2021 Toward more accurate developer recommendation via inference of development activities from interaction with bug repair process
abstract
Abstract Software projects usually receive a large number of submitted bug reports every day. Manually triaging the bug reports is often time‐consuming and error‐prone; thus, it is necessary to automatically assign the bug reports to the suitable developers for bug repair, with the help of bug tracking systems. Aiming to reducing the time consumption and mismatch of bug report assignments, we present a developer recommendation model for bug repair based on weighted recurrent neural network, namely, DTPM, which contains two parts: One obtains multisource semantic information of bug reports and fuses them into high‐dimensional semantic feature vectors, and the other combines a penalty matrix into a single hidden layer neural network to obtain more reasonable developer recommendations. We conduct experiments on five datasets of open bug repositories (NetBeans, OpenOffice, GCC, Mozilla, and Eclipse), and the experimental results show that DTPM can achieve better performance than state‐of‐the‐art models LDA_KL, LDA_KL, LDA_SVM, DERTOM, DREX, and DeepTriage.
Linhui Wang, Rong Chen 0003, Shikai Guo
J. Softw. Evol. Process.4
2021 Underwater image enhancement with image colorfulness measure
Xi Yang 0009, Hui Li 0014, Rong Chen 0003
Signal Process. Image Commun.3
2020 Solving quantified constraint satisfaction problems with value selection rules
Jian Gao 0007, Kuixian Wu, Rong Chen 0003
Frontiers Comput. Sci.4
2020 Developer Activity Motivated Bug Triaging: Via Convolutional Neural Network
Shikai Guo, Xi Yang 0009, Rong Chen 0003, Chen Guo 0001, Hui Li 0014
Neural Process. Lett.4
2019 CrowDIY: How to Design and Adapt Collaborative Crowdsourcing Workflows Under Budget Constraints
Rong Chen 0003, Bo Li 0120, Hu Xing
ICWE1
2019 Identify Severity Bug Report with Distribution Imbalance by CR-SMOTE and ELM
abstract
Manually inspecting bugs to determine their severity is often an enormous but essential software development task, especially when many participants generate a large number of bug reports in a crowdsourced software testing context. Therefore, boosting the capabilities of methods of predicting bug report severity is critically important for determining the priority of fixing bugs. However, typical classification techniques may be adversely affected when the severity distribution of the bug reports is imbalanced, leading to performance degradation in a crowdsourcing environment. In this study, we propose an enhanced oversampling approach called CR-SMOTE to enhance the classification of bug reports with a realistically imbalanced severity distribution. The main idea is to interpolate new instances into the minority category that are near the center of existing samples in that category. Then, we use an extreme learning machine (ELM) — a feedforward neural network with a single layer of hidden nodes — to predict the bug severity. Several experiments were conducted on three datasets from real bug repositories, and the results statistically indicate that the presented approach is robust against real data imbalance when predicting the severity of bug reports. The average accuracies achieved by the ELM in predicting the severity of Eclipse, Mozilla, and GNOME bug reports were 0.780, 0.871, and 0.861, which are higher than those of classifiers by 4.36%, 6.73%, and 2.71%, respectively.
Shikai Guo, Rong Chen 0003, Hui Li 0014, Tianlun Zhang
Int. J. Softw. Eng. Knowl. Eng.2
2019 The Influence Ranking for Testers in Bug Tracking Systems
abstract
At present, bug tracking systems are used to collect and manage bug reports in many software projects. As participants, the testers not only submit bug reports to the system, but also comment on bug reports in the system. The tester’s behaviors of submitting and commenting reflect his/her influence in bug tracking systems. However, with the rapid increase of the bug reports in software projects, evaluating the testers’ influence in the projects accurately becomes more and more difficult. Aiming at solving this problem, the submission and comment on bug report can be regarded as social behaviors of the testers, and thus the method of Influence Ranking for Testers (IRfT) in bug tracking systems is presented and used for measuring the influence of the testers in this paper. The case study of the Eclipse project in Bugzilla shows that the result produced by IRfT is consistent with the actual performance of the testers in this project. The ranking results can keep stable in the cases of link adding or removing and tester removing in tester networks, and the results are also proved to be valid in the future. The further investigation on the speed of network break-down by node removal demonstrates that the top-ranking testers are important in the organization of tester networks. Additionally, the results also show that the ranking of the testers is related to the existence time in bug tracking system. Therefore, IRfT is proved to be an effective measurement for evaluating the influence of the testers in bug tracking system, and it can further demonstrate the testers’ contributions in software testing, such as bug validations, bug fixes, etc.
Hui Li 0014, Guofeng Gao, Rong Chen 0003, Shikai Guo
Int. J. Softw. Eng. Knowl. Eng.3
2019 A Novel Approach to Publishing Tasks for Collaboratively Crowdsourcing Workflows
abstract
In recent years, crowdsourcing has gradually become a promising way of using netizens to accomplish tiny tasks on, or even complex works through crowdsourcing workflows that decompose them into tiny ones to publish sequentially on the crowdsourcing platforms. One of the significant challenges in this process is how to determine the parameters for task publishing. Still some technique applied constraint solving to select the optimal tasks parameters so that the total cost of completing all tasks is minimized. However, experimental results show that computational complexity makes these tools unsuitable for solving large-scale problems because of its excessive execution time. Taking into account the real-time requirements of crowdsourcing, this study uses a heuristic algorithm with four heuristic strategies to solve the problem in order to reduce execution time. The experiment results also show that the proposed heuristic strategies produce good quality approximate solutions in an acceptable timeframe.
Rong Chen 0003, Shikai Guo
Int. J. Softw. Eng. Knowl. Eng.2
2019 Improved SMOTE Algorithm to Deal with Imbalanced Activity Classes in Smart Homes
Shikai Guo, Rong Chen 0003, Xiao Sun 0003, Xiangxin Wang
Neural Process. Lett.3
2019 Fusion of Multi-RSMOTE With Fuzzy Integral to Classify Bug Reports With an Imbalanced Distribution
abstract
With the help of automated classification, severe bugs can be rapidly identified so that the latent damage to software projects can be minimized. However, bug report datasets commonly suffer from disproportionate number of category samples. When presented with the situation of class imbalance, most standard classification learning approaches fail to properly learn the distributive characteristics of the samples and tend to result in unfavorable performance to predict class label. In this case, imbalanced learning becomes critical to advance classification algorithms. In this paper, we propose an improved synthetic minority oversampling technique to avoid the degraded performance caused by class imbalance in bug report datasets. Moreover, to lessen the chance of occasionalities in random sampling process, we propose a repeated sampling technique to train different, but related classifiers. Finally, an ensemble algorithm based on Choquet fuzzy integral is employed to combine the wisdom of crowds and make better decisions. We conduct comprehensive experiments on several bug report datasets from real-world bug repositories. The results demonstrate that the proposed method boosts the classification performance across the classes of the data. Specifically, compared with various ensemble learning techniques, the Choquet fuzzy integral achieves outstanding results on integrating multiple random oversampling techniques.
Rong Chen 0003, Shikai Guo, Xizhao Wang, Tianlun Zhang
IEEE Trans. Fuzzy Syst.1
2019 Single Image Haze Removal via Region Detection Network
abstract
Haze removal typically works on a physical model to estimate how light is transmitted and lost due to absorption and scattering through the atmosphere. In this paper, a region detection network is proposed to learn the relationship between the hazy image and the medium transmission map in a patchwise manner; the transmission map is then used to remove haze via an atmospheric scattering model and enhance the detail of de-hazed images. To this end, we design a simple yet powerful deep convolutional neural network, which mainly consists of two types of network units and can be trained in an end-to-end manner. One network unit is a module with the residual structure that facilitates the learning process of the deep network. The other is a novel module with a cascaded cross channel pool, which fuses multi-level haze-relevant features and boosts the abstraction ability of the model on a nonlinear manifold. Moreover, an evolutionary-based enhancement method is developed to improve the level of detail of over-smoothed results. Several comparative experiments have been conducted on synthetic and real images, through which we conclude that the proposed method achieves state-of-the-art haze removal results, qualitatively and quantitatively. Supplementary experiments further indicate that our method works better against other adverse effects on vision quality (e.g., the mist formed by heavy rain and the veil met underwater). Moreover, we present a lightweight version of the proposed network, which achieves an impressive haze removal performance even on low-power devices.
Xi Yang 0009, Hui Li 0014, Yu-Long Fan, Rong Chen 0003
IEEE Trans. Multim.4
2018 An Exact Algorithm for Maximum k-Plexes in Massive Graphs
abstract
The maximum k-plex, a generalization of maximum clique, is used to cope with a great number of real-world problems. The aim of this paper is to propose a novel exact k-plex algorithm that can deal with large-scaled graphs with millions of vertices and edges. Specifically, we first propose several new graph reduction methods through a careful analyzing of structures of induced subgraphs. Afterwards, we present a preprocessing method to simplify initial graphs. Additionally, we present a branch-and-bound algorithm integrating the reduction methods as well as a new dynamic vertex selection mechanism. We perform intensive experiments to evaluate our algorithm, and show that the proposed strategies are effective and our algorithm outperforms state-of-the-art algorithms, especially for real-world massive graphs.
Jian Gao 0007, Jiejiang Chen, Minghao Yin, Rong Chen 0003, Yiyuan Wang 0002
IJCAI4
2018 Enhancing Bug Report Assignment with an Optimized Reduction of Training Set
Miaomiao Wei, Shikai Guo, Rong Chen 0003, Jian Gao 0007
KSEM (2)3
2018 Weighted Data Set Reduction for Automatic Bug Triaging (P)
abstract
Despite the great potential to save the labor cost of developers, automated bug triaging as a text classification problem has not been thoroughly investigated on long descriptions, which are informative but often noisy.In this paper an effective bug triage technique is proposed to build a high quality set of bug data by removing the noisy and noninformative bug reports while assigning new bugs to an appropriate developer.The proposed techniqueweighted data set reductionis built upon three feature selection algorithms and four instances selection algorithms with intention to recommend the bug and to automatically assign it more accurately even with noisy bug descriptions.Several experiments are conducted and the experimental results show that the reduced training sets by the proposed approach can achieve better accuracy in several cases, about 2-3% on average better than the original ones.
Miaomiao Wei, Shikai Guo, Rong Chen 0003
SEKE3
2018 Capability Matching and Heuristic Search for Job Assignment in Crowdsourced Web Application Testing
abstract
Web based commercial systems are increasingly becoming feature rich, interactive and functional as locally installed applications. Testing web applications is unique, as many factors affect the system performance and user experience. Crowdsourcing is an appealing and economic solution to web application testing due to the ability to reach a larger international audience. However, less is known about the quality control of crowdsourced testing to harness the collective efforts of individuals. In our study, the collaborative testing problem in a crowdsourcing environment is defined as a job assignment problem and is formulated as an integer linear programming (ILP) problem. The objective of this paper is to validate a greedy job assignment approach as a tool for the effective use of crowdsourced testing. We carried out a case study on Xturk, a prototype crowdsourced testing system, to understand the crowdsourced testers behaviour that the trustworthiness, the execution time of test cases and accuracy of feedback. Several experiments indicate that this approach is comparatively effective with regards to the feasibility verdict, efficiency and accuracy.
Shikai Guo, Rong Chen 0003, Hui Li 0014
SMC2
2018 Crowdsourced Web Application Testing Under Real-Time Constraints
abstract
Crowdsourcing carried out by cyber citizens instead of hired consultants and professionals has become increasingly an appealing solution to test the feature rich and interactive web. Despite having various online crowdsourcing testing services, the benefits of exposure to a wider audience and harnessing the collective efforts of individuals remain uncertain, especially when the quality control is problematic in an open environment. The objective of this paper is to propose a real-time collaborative testing approach (RCTA) to create a productive crowdsourced testing on a dynamic Internet. We implemented a prototype crowdsourcing system XTurk, and carried out a case study, to understand the crowdsourced testers behavior, the trustworthiness, the execution time of test cases and accuracy of feedback. Several experiments are carried out and experimental results validate the quality, efficiency and reliability of the present approach and the positive testing feedback is are shown to outperform the previous methods.
Shikai Guo, Rong Chen 0003, Hui Li 0014, Jian Gao 0007
Int. J. Softw. Eng. Knowl. Eng.2
2016 Probabilistic reasoning in diagnosing causes of program failures
abstract
Summary Fault localization is sensitive to program runs, and the pattern of fault propagation and manifestation in real software is extremely complex and uncertain. To accommodate the complexity and uncertainty, this paper presents a novel probabilistic graph model – the probabilistic cause–effect graph (PCEG) is built upon dynamic dependencies generated from running the faulty program against failed test cases and performs probabilistic inference with coverage information from the whole test suite. PCEG is an extension of the traditional probabilistic graph both in structural and inferential terms and is different from earlier probabilistic approaches to software diagnosis by introducing two forms of evidences (i.e. apparent faults and real faults). The proposed probabilistic reasoning algorithm works on the PCEG converted from a dynamic program dependency graph and diagnoses the causes with both top‐down and bottom‐up inference. The experimental results have shown the improvements on diagnostic effectiveness and accuracy in both single‐fault and multiple‐fault context, even when a program yields similar program runs through loop statements. Copyright © 2015 John Wiley & Sons, Ltd.
Rong Chen 0003, Zhenjun Du
Softw. Test. Verification Reliab.2
2014 SFDCloud: top-k service faults diagnosis in cloud computing
Zhichun Jia, Rong Chen 0003, Xing Xing, Yiwu Xie
Autom. Softw. Eng.2
2014 Extraction and Analysis of Crucial Fraction in Software Networks
abstract
Many complex systems, such as software systems, are full of complexity arising from interactions among basic units (such as classes, interfaces and struts in object-oriented software systems). One of the most successful approaches to capture the underlying structural features of large-scale software systems is the investigation of hierarchical organization. However, the hierarchy of software networks has not been thoroughly investigated. In this paper, the crucial fraction (CF) in software networks has been extracted and analyzed in a set of real-world software systems. First, the classes and the relationships between them have been extracted into software networks. Then software networks have been divided into different layers, and CF of software networks has been extracted by k-core. The empirical studies in this paper reveal that software networks represent flat hierarchical structure. Finally, CF has been measured by the relevant complex network parameters respectively, and the relations between CF and overall network have been analyzed by the case studies of software networks. The results show that CF represents characteristics of scale-free, small-world, strong connectivity, and the units in CF are frequently reused and dominate the overall system.
Hui Li 0014, Rong Chen 0003, Hai Zhao 0002
Int. J. Softw. Eng. Knowl. Eng.2
2014 A Fast Approach to Querying Multiple Ontology Versions Based on Concept Lattice
Rong Chen 0003, Wu Deng 0001
J. Web Eng.2
2013 Isolating and Understanding Program Errors Using Probabilistic Dispute Model
abstract
Automated software debugging can have a signifi-cant impact on the cost and quality of software development and maintenance. In recent years, researchers have invested a considerable amount of effort in developing automated techniques, and have demonstrated their effectiveness in helping developers in certain debugging tasks by pinpointing faulty statements. But there is still a gap between examining a faulty statement and understanding root causes of the cor-responding bug. As a step in this direction, we believe good developers have defensive programming in minds and software debugging is a process in search of arguments about why a statement is faulty. Therefore, a fault localization problem is rephrased as a dispute game between statements involved in successful runs and failing runs. A statement is OK if it can always provide arguments against other's blames, whereas a less defensive statement is thought to be faulty. In doing so, we propose a probabilistic dispute graph which is built upon dynamic dependencies between statements and statistics of program runs. Using such a graph, we put executed statements in dispute, compute acceptable statements, and thus figure out faulty statements if they have not strong arguments about their correctness. For empirical purpose, we carry out experiments on the well-known Siemens benchmark, and conclude that our approach not only casts new light on the causes of bugs in various cases, but also is statistically more effective in fault localization than competitors like Tarantula, SOBER, CT and PPDG.
Rong Chen 0003, Zhichun Jia, Jian Gao 0007
COMPSAC1
2012 A novel two-stage hybrid swarm intelligence optimization algorithm and application
Wu Deng 0001, Rong Chen 0003, Lifeng Yin, Jinghuan Guo
Soft Comput.2
2009 An Approach to Checking SOC Timing Safety
abstract
High-performance SOCs (system on a chip) are being widely used in multimedia signal processing. Timing behavior is a critical issue in high-performance design, and one of the significant and challenging problems in the SOC design is to ensure the safety of timing behavior. A model that combines Allen-Givone algebra in multi-valued logic and the waveform polynomial is presented in this paper. Functional and timing behavior can be precisely described simultaneously by one unified representation, allowing inputs and outputs of a multi-valued logic function to be described by multi-valued waveforms, which are consistent with commonly used intuitive waveforms. This method is applicable in precisely analyzing and checking the correctness and safety of SOC timing attributes.
Zhenjun Du, Rong Chen 0003
IAS2