Can Cheng

dblp:120/8679 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Code Reviewer Recommendation Based on a Hypergraph with Multiplex Relationships
abstract
Code review is an essential component of software development, playing a vital role in ensuring a comprehensive check of code changes. However, the continuous influx of pull requests and the limited pool of available reviewer candidates pose a significant challenge to the review process, making the task of assigning suitable reviewers to each review request increasingly difficult. To tackle this issue, we present MIRRec, a novel code reviewer recommendation method that leverages a hypergraph with multiplex relationships. MIRRec encodes high-order correlations that go beyond traditional pairwise connections using degree-free hyperedges among pull requests and developers. This way, it can capture high-order implicit connectivity and identify potential reviewers. To validate the effectiveness of MIRRec, we conducted experiments using a dataset comprising 48,374 pull requests from ten popular open-source software projects hosted on GitHub. The experiment results demonstrate that MIRRec, especially without PR-Review Commenters relationship, outperforms existing state-of-the-art code reviewer recommendation methods in terms of ACC and MRR, highlighting its significance in improving the code review process.
Yu Qiao 0001, Jian Wang 0018, Can Cheng, Wei Tang 0018, Peng Liang 0001, Yuqi Zhao 0001, Bing Li 0010
SANER3
2023 OFIDS : Online Learning-Enabled and Fingerprint-Based Intrusion Detection System in Controller Area Networks
abstract
As a widely used industrial field bus, the controller area network (CAN) lacks security mechanisms (e.g., encryption and authentication) and is vulnerable to security attacks (e.g., masquerade). A fingerprint-based intrusion detection system (IDS) in CAN networks can detect masquerade attacks by scanning the unique clock signals of CAN devices. However, most state-of-the-art fingerprint-based IDSs commonly use an analog-to-digital converter module with a low frequency of 60 MHz to sample CAN signals, lowering the detection accuracy of fingerprint-based IDSs. In addition, almost all fingerprint-based IDSs are trained offline and then detected online, ignoring that system clock signals of hardware change over time, resulting in degraded detection performance. This paper proposes an online learning-enabled and fingerprint-based IDS (OFIDS) in CAN networks to increase the sampling frequency, shorten the detection response time, and increase the detection accuracy. OFIDS uses a high-speed comparator (i.e., TLV3501) and FPGA (i.e., Xilinx ZYNQ-7010) to sample the CAN_High signal, achieving a low sampling delay time of 4.5 ns and a high sampling frequency of 1 GHz. The self-adaptability of the backpropagation neural network is taken advantage of and used to train the OFIDS model with a detection accuracy of 99.9992%. OFIDS is deployed to a CAN network prototype with five CAN devices (i.e., two Arduino UNO boards and three STM32 microcontrollers) and a real vehicle. Experimental results show that OFIDS can achieve at least 99.99% detection accuracy within 0.18μs in a CAN network prototype and can achieve 98% detection accuracy in a real vehicle.
Yehua Wei, Can Cheng, Guoqi Xie
IEEE Trans. Dependable Secur. Comput.2
2022 Improving generality and accuracy of existing public development project selection methods: a study on GitHub ecosystem
Can Cheng, Bing Li 0010, Zengyang Li, Peng Liang 0001
Autom. Softw. Eng.1
2022 An in-depth study of the effects of methods on the dataset selection of public development projects
abstract
Abstract Public development projects (PDPs) and documented public development projects (DPDPs) are two types of projects that can provide valuable information on how developers and users participate in OSS projects. However, it is hard for researchers to effectively select PDPs and DPDPs due to the lack of specific project selection methods for these two types of projects. To address this problem, a standard dataset was labelled and the base line methods (i.e. selecting projects according to a single feature like star number) under 60 configurations and the machine learning methods under 18 configurations were tested to identify the best configurations in precision and F‐measure for selecting PDPs and DPDPs. The results show that (1) to select PDPs or DPDPs with a high precision, the base line method is the best with precision of 0.877 (PDPs) and 0.831 (DPDPs); (2) to select PDPs or DPDPs with a high F‐measure, the machine learning methods are the best, with F‐measure of 0.817 (PDPs) and 0.789 (DPDPs); (3) existing sample selection strategies can be combined with the machine learning methods, and the precision of selecting PDPs can be increased by 6.39%–41.33% and the precision of selecting DPDPs can be can be increased by 35.50%–269.02%.
Can Cheng, Bing Li 0010, Zengyang Li, Peng Liang 0001
IET Softw.1
2018 Automatic Detection of Public Development Projects in Large Open Source Ecosystems: An Exploratory Study on GitHub
abstract
 -Hosting over 10 million of software projects, GitHub is one of the most important data sources to study behavior of developers and software projects.However, with the increase of the size of open source datasets, the potential threats to mining these datasets have also grown.As the dataset grows, it becomes gradually unrealistic for human to confirm quality of all samples.Some studies have investigated this problem and provided solutions to avoid threats in sample selection, but some of these solutions (e.g., finding development projects) require human intervention.When the amount of data to be processed increases, these semi-automatic solutions become less useful since the effort in need for human intervention is far beyond affordable.To solve this problem, we investigated the GHTorrent dataset and proposed a method to detect public development projects.The results show that our method can effectively improve the sample selection process in two ways: (1) We provide a simple model to automatically select samples (with 0.827 precision and 0.947 recall); (2) We also offer a complex model to help researchers carefully screen samples (with 63.2% less effort than manually confirming all samples, and can achieve 0.926 precision and 0.959 recall).
Can Cheng, Bing Li 0010, Zengyang Li, Peng Liang 0001
SEKE1
2017 Developer Role Evolution in Open Source Software Ecosystem: An Explanatory Study on GNOME
Can Cheng, Bing Li 0010, Zengyang Li, Yuqi Zhao 0001, Feng-Ling Liao
J. Comput. Sci. Technol.1
2014 A Petri Net-Based Approach to Service Composition and Monitoring in the IOT
abstract
Recently, there are many improvements in Internet of Things (IOT). Through the recombination and optimization, the real-world devices can provide their functionality as web services in IOT. However, it is very difficult to cost-effectively access to the Internet of Things due to the environmental changes. In this paper, firstly, a Petri net-based model for service composition in IOT is proposed, which uses a comprehensive performance function rtc (involving reliability, response time and cost) to evaluate the cost-effectiveness. Then, the Find to Optimal algorithm is addressed to find a cost-effective composition path. Furthermore, when environments change dynamically, the FBased Monitor algorithm can well solve the composition. Finally, the experiments prove the soundness and correctness of our model and algorithms.
Rong Yang 0007, Bing Li 0010, Can Cheng
APSCC3
2012 Supervised Isomap Based on Pairwise Constraints
Jian Cheng 0004, Can Cheng, Yinan Guo 0001
ICONIP (1)2