Jie Chen 0060

dblp:92/6289-60 · DBLP profile ↗
← Back
29ranked-venue papers
5as first author
22since 2021 · last 2027
0000-0002-6701-084XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 15 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Security and privacy · 1
YearPublicationVenuePosition
2027 Method-level technical debt detection based on source code and comments
Dongjin Yu, Wangliang Yan, Bin Hu 0034, Jie Chen 0060
Sci. Comput. Program.6
2026 Code clone classification based on multi-dimension feature entropy
Bin Hu 0034, Lizhi Zheng, Dongjin Yu, Yijian Wu, Jie Chen 0060
Sci. Comput. Program.5
2025 Breaking the Curse of High-Dimensionality: A Transformer-Based Approach to Software Configuration Performance Prediction
abstract
Modern software systems are increasingly configurable, offering extensive options to meet diverse requirements. However, this flexibility results in exponentially growing configuration spaces, making exhaustive performance testing infeasible. Existing performance prediction models face significant challenges due to the curse of dimensionality, which limits their ability to generalize across high-dimensional spaces with sparse data. To address these challenges, this paper introduces a novel performance prediction framework comprising two components: Feature Transformer AutoEncoder (FTA) and Distance-Aware Greedy Sampling (DAGS). FTA leverages self-supervised learning and attention mechanisms to encode high-dimensional configuration data into compact, information-rich representations, effectively capturing long-range dependencies and mitigating data sparsity. DAGS, on the other hand, employs intelligent sampling strategies to maximize coverage within configuration spaces while minimizing measurement overhead. Extensive experimental evaluations across six real-world configurable software systems demonstrate that our approach significantly outperforms state-of-the-art methods, achieving up to a 79% reduction in Mean Relative Error (MRE). The results highlight the robustness and scalability of our framework in overcoming the traditional performance degradation associated with high-dimensional configuration spaces.
Zhengyue Pan, Mingjie Tu, Jie Chen 0060
APSEC5
2025 From Code Changes to Performance Variations: A Predictive Approach for Microbenchmark Prioritization
abstract
Microbenchmark-based performance regression detection is critical for software quality assurance but faces limitations in continuous integration (CI) environments. Existing test case prioritization methods primarily rely on code coverage and historical stability, overlooking the complex relationships between code changes and performance impacts. To address this, we presents JMH-TCP, an intelligent prioritization framework that leverages predictive modeling to establish connections between code changes and performance variations. Our method introduces a multi-dimensional analysis that integrates novel features, including semantic code distance metrics and change impact assessment powered by large language models, while incorporating traditional metrics such as historical performance stability and code coverage. Experimental evaluations across 10 open-source Java projects show that JMH-TCP achieving a 10.8% improvement in the Average Position of Faults in the Defect Prioritization (APFD-P) metric and an 18.2% enhancement in ranking effectiveness (NDCG) compared to state-of-the-art methods. Furthermore, our detailed ablation and correlation study reveals the synergistic effects of different features on prioritization effectiveness, offering valuable insights for optimizing performance testing strategies.
Jie Chen 0060, Dongjin Yu, Yingling Li, Hengyu Chen
IJCNN2
2025 Efficient feature envy detection and refactoring based on graph neural network
Dongjin Yu, Lehui Weng, Jie Chen 0060, Xin Chen 0032, Quanxin Yang
Autom. Softw. Eng.4
2025 Enhancing structural knowledge in code smell identification: A fusion learning framework combining AST-based metrics with semantic embeddings
Quanxin Yang, Dongjin Yu, Sixuan Wang, Xin Chen 0032, Jie Chen 0060, Bin Hu 0034
Expert Syst. Appl.6
2024 What makes a real change in software performance? An empirical study on analyzing the factors that affect the triagement of performance change points
Jie Chen 0060, Dongjin Yu
Sci. Comput. Program.1
2024 Actionable code smell identification with fusion learning of metrics and semantics
Dongjin Yu, Quanxin Yang, Xin Chen 0032, Jie Chen 0060, Sixuan Wang
Sci. Comput. Program.4
2024 Occlusion-robust workflow recognition with context-aware compositional ConvNet
Zhongjin Li, Jie Chen 0060
Soft Comput.4
2023 An Empirical Study on GitHub Sponsor Mechanism
abstract
From May 2019, GitHub launched sponsor mechanism indicating that GitHub is moving towards deeper integration of open source development and economic support. It will bring more comprehensive and diversified support to the open source community. However, the number of developers profiting from the sponsor mechanism follows a long tail distribution. Our study found that only 31% of developers who started the sponsor mechanism received rewards, and 39.3% of them only received a reward of one dollar. Our work focuses on identifying what factors affect the availability of sponsorship for developers in open source community. We start by defining 45 features to characterize the developers in four dimensions i.e. Personality, Advertisement, Repository and Behavior. The results of statistical analysis indicate that most of the proposed features differ significantly between the ones who received rewards (short for MTs_Yes) from those that are not. After that, we build machine learning model based on the proposed features to predict MTs_Yes. Compared with the existing work, results show that our method outperforms baselines by 30% for AUC (Area Under the Curve). In addition, we investigated the relative contribution of features in detecting MTs_Yes and analyzed the important features by using an interpretable model SHAP. Finally, based on the experimental results, we put forward corresponding and practical suggestions for developers who want to receive rewards so as to make the community of open source projects develop more harmonious.
Yiqian Yang, Haolan He, Jie Chen 0060
Int. J. Softw. Eng. Knowl. Eng.4
2023 Graph-based code semantics learning for efficient semantic code clone detection
Dongjin Yu, Quanxin Yang, Xin Chen 0032, Jie Chen 0060
Inf. Softw. Technol.4
2023 Pattern-based circular reference detection in Python
Jie Chen 0060, Dongjin Yu
Sci. Comput. Program.1
2023 Towards an understanding of memory leak patterns: an empirical study in Python
Jie Chen 0060, Dongjin Yu
Softw. Qual. J.1
2023 Action detection with two-stream enhanced detector
Zhongjin Li, Jie Chen 0060
Vis. Comput.4
2022 Detecting and Refactoring Feature Envy Based on Graph Neural Network
abstract
As one of the most common code smells, feature envy reduces the cohesion of classes and increases the coupling between classes, thus leading to difficulty of software maintainability. Though many studies have made good achievements on feature envy detection, they often despise or even ignore the inherent calling relationships between methods, causing unimpressive detection efficiency. To address this problem, we propose a Graph Neural Network (GNN) based approach towards feature envy detection. We first collect code metrics and calling relationships, and then convert them to the form of a graph, where the node represents the code metrics of a method and the edge represents the calling relationship between methods. Particularly, considering the unbalance of positive and negative samples, we introduce a graph augmenter to obtain an enhanced graph. Finally, we feed the enhanced graph into a GNN model for training and predicting. We conducted extensive experiments on a dataset containing five open-source software projects. The result shows that our approach achieves 78.90% in terms of average F1-score, which is 37.98% absolutely higher than the best comparison approach. Besides, we propose a refactoring recommendation approach based on calling strength. It achieves 61.44% of accuracy, which is 5.13% absolutely higher than the best competitive. Our code and datasets are available at https://github.com/HduDBSI/Feature-Envy-Detection.
Dongjin Yu, Lehui Weng, Jie Chen 0060, Xin Chen 0032, Quanxin Yang
ISSRE4
2022 Natural Language-Based Automatic Programming for Industrial Robots
Jie Chen 0060, Zhongjin Li, LiGuo Huang
J. Grid Comput.2
2022 Characterizing and Detecting Methods to be Benchmarked under Performance Unit Test
abstract
Continuous integration is a growing trend in the software engineering community and industry. Performance testing is becoming more important in this context. To support precise and fine-grained monitoring, performance unit tests are applied for small software components. However, the benchmarks for performance unit testing are still insufficient, which means that benchmark coverage is low and there is a room for improvement. Therefore, focusing on the most important parts of the software, such as methods, and ensuring that their performance is monitored closely with performance unit tests can greatly reduce the amount of work that needs to be done for testing and to prepare benchmarks. This paper aims to provide an assisting approach for detecting methods that need to be benchmarked in performance unit tests. We start by defining 30 features to characterize the methods in the projects and show that they can be used to tell the benchmarked methods (short for BDMs) from those that are not. Then, using the proposed features, we build machine learning-based models to detect BDMs. We perform an experiment with 10 open source projects from GitHub to see how well our approach works. First, we use seven binary classification techniques to evaluate the prediction performance of our machine learning models. We find that Random Forest makes the best predictions where AUC and MCC are between 0.77 and 0.89 and 0.5 and 0.75, respectively. In terms of cost effectiveness, the experiment reveals that by inspecting only 5% of the candidate methods detected by our model, 43% of the total real BDMs can be retrieved. Second, we conduct feature importance evaluations for individual features and feature categories. We find that eight features related to Scope, History, and Complexity are individually important for good predictions and that the combination of all features in the Scope category is paramount for our model, while the combination of features in the Control Flow category is less important. Third, we investigate the performance of our detection approach with different feature selection strategies and data sources. Our results show that we can make good predictions about whether a method needs to be benchmarked by using machine learning models. Practitioners can use our method and the results of the study to deal with BDMs detection effectively.
Jie Chen 0060, Dongjin Yu
Int. J. Softw. Eng. Knowl. Eng.1
2022 Proposal-Based Graph Attention Networks for Workflow Detection
Zhongjin Li, Jie Chen 0060
Neural Process. Lett.4
2022 Multiclass Classification for Self-Admitted Technical Debt Based on XGBoost
abstract
In software development, due to the demands from users or the limitations of time and resources, developers tend to adopt suboptimal solutions to achieve quick software development. In such a way, the released software usually involves not-quite-right code that is called technical debt, which will significantly decrease the quality of software and increase the maintenance cost. Recently, the concept of self-admitted technical debt (SATD) is proposed and refers to technical debt that is self-admitted by developers in code comments. Existing studies mainly focus on detecting technical debt by classifying code comments into either “SATD” or “non-SATD.” However, different types of SATD has different impacts on software maintenance and needs to be handled by different developers. Therefore, the detected SATD should be further classified so that developers can understand and remove technical debt better. In this article, we propose a new method based on eXtreme Gradient Boosting (XGBoost) to classify SATD into multiple classes. In our approach, we first preprocess the original code comments and adopt the easy data augmentation strategy to overcome the class unbalance problem. Then, chi-square is leveraged to select representative features from the textual feature set. Finally, we apply XGBoost to train a classifier and use the trained classifier to partition each comment into the corresponding class. We experimentally investigate the effectiveness of our approach on a public dataset, including 62 566 code comments from 10 open-source projects. Experimental results show that our approach achieves 56.66% in terms of macroaveraged precision, 59.07% in terms of macroaveraged recall, and 55.77% in terms of macroaveraged F-measure on average, and outperforms the natural language processing based method by 4.98%, 5.32%, and 3.17%, respectively. In addition, the experimental results also demonstrate that the data augmentation strategy is effective in improving the effectiveness of our approach.
Xin Chen 0032, Dongjin Yu, Xulin Fan, Jie Chen 0060
IEEE Trans. Reliab.5
2021 Using BiLSTM with attention mechanism to automatically detect self-admitted technical debt
Dongjin Yu, Xin Chen 0032, Jie Chen 0060
Frontiers Comput. Sci.4
2021 Attention-based encoder-decoder networks for workflow recognition
Zhongjin Li, Jie Chen 0060
Multim. Tools Appl.4
2021 Modeling and Analysis of Cyber-Physical System Based on Object-Oriente Generalized Stochastic Petri Net
abstract
Cyber–physical system (CPS) is a complex system that contains multiple components working cooperatively. According to its characteristics, we propose an object-oriented generalized stochastic Petri net (OGSPN), in which the CPS is abstracted into several types of objects and its logical structure and working process is visually described. Moreover, we model and measure the time consumed by each activity in CPS for quantitative analysis. To simplify the process of performance analysis on this model, in this article we propose a compression algorithm to convert OGSPN into a generalized stochastic Petri net (GSPN). Considering the uncertainty in CPS, we use a fuzzy mathematics based method to process the compressed model of GSPN for improving the accuracy of the performance analysis. We apply our method to a real-world thick metal plate production line in a manufacturing company, and the availability of our method is verified by extensive experiments.
Zhongjin Li, Jie Chen 0060, Hua Hu 0001
IEEE Trans. Reliab.4
2020 Security and performance-aware resource allocation for enterprise multimedia in mobile edge computing
Zhongjin Li, Binbin Huang 0006, Jie Chen 0060, Chuanyi Li, Hua Hu 0001, LiGuo Huang
Multim. Tools Appl.4
2020 Workflow recognition with structured two-stream convolutional networks
Kaiming Cheng, Zhongjin Li, Jie Chen 0060, Hua Hu 0001
Pattern Recognit. Lett.4
2019 DeepTLE: Learning Code-Level Features to Predict Code Performance before It Runs
abstract
With the continuous expansion of the software market and the updating of the maturity of the software development process, the performance requirements of software users are becoming increasingly prominent. Performance issues are essentially related to the source code. For solving the same problem, different programmers may write completely different "correct" code with the same functionality but have different performance. Most online judge system on programming make use of automated grading systems, usually rely on test results to quantify the correctness and performance for the submitted source code. However, traditional dynamic testing takes a lot of time, and the discovery of performance problems is usually after the fact even for those small scale programs. Therefore, we proposed DeepTLE which is used to effectively predict the performance of submitted source code before it runs. DeepTLE can automatically learn the semantic and structural features of the source code. In order to verify the effect of our approach, we applied it to the source code collected from the program competition website to predict if the source code would be time limit exceed or not without running its test cases. Experiment results show that our method can save 96% of the time cost compared to the dynamic testing, and the accuracy of the prediction reaches 82%.
Meiling Zhou, Jie Chen 0060, JiaCheng Yu, Zhongjin Li, Hua Hu 0001
APSEC2
2019 Analyzing performance-aware code changes in software development process
abstract
With the continuous expansion of software market and the updating of the maturity of the software development process, the performance requirements of software users have gradually become prominent. Performance issues are closely related to the source code. Thus, with the increasing complex of the software product, its performance changed during the evolution of the software product. Performance optimization related work has always been an important goal for developers who usually coding at a low-level. However, performance problems are well studied on architecture level. All too often, some developers are ignorant of the way their code modifications affect performance and simply to wait until performance drops to a point that is unacceptable to the business side. As software developers did a lot of daily work at code level, we think code level performance awareness can help developers in sight of the performance of the code that they are working with. To deal with this, we firstly build performance-aware code change model to identify the performance changes and its related code changes at the granularity of function between each two reversions of a program. Then, we analyzed the evolution history of the code performance and mined the frequent code change patterns that used to improve performance. We have build related tool to implement the proposed approach and applied it to 8 open source projects.
Jie Chen 0060, Dongjin Yu, Zhongjin Li, Hua Hu 0001
ICPC1
2018 Fault-Tolerant Scheduling for Scientific Workflow with Task Replication Method in Cloud
abstract
Cloud computing has become a revolutionary paradigm by provisioning on-demand and low cost computing resources for customers. As a result, scientific workflow, which is the big data application, is increasingly prone to adopt cloud computing resources. However, internal failure (host fault) is inevitable in such large distributed computing environment. It is also well studied that cloud data center will experience malicious attacks frequently. Hence, external failure (failure by malicious attack) should also be considered when executing scientific workflows in cloud. In this paper, a fault-tolerant scheduling (FTS) algorithm is proposed for scientific workflow in cloud computing environment, the aim of which is to minimize the workflow cost with the deadline constraint even in the presence of internal and external failures. The FTS algorithm, based on tasks replication method, is one of the widely used fault tolerant mechanisms. The experimental results in terms of real-world scientific workflow applications demonstrate the effectiveness and practicality of our proposed algorithm.
Zhongjin Li, JiaCheng Yu, Jie Chen 0060, Hua Hu 0001, Jidong Ge, Victor Chang 0001
IoTBDS4
2018 Multi-objective scheduling for scientific workflow in multicloud environment
Zhongjin Li, Hua Hu 0001, Jie Chen 0060, Jidong Ge, Chuanyi Li, Victor Chang 0001
J. Netw. Comput. Appl.4
2018 Efficiently detecting structural design pattern instances based on ordered sequences
Dongjin Yu, Jiazha Yang, Zhenli Chen, Chengfei Liu, Jie Chen 0060
J. Syst. Softw.6