Jinfu Chen 0001

dblp:54/4698-1 · also Jin-Fu Chen 0001 · DBLP profile ↗
← Back
116ranked-venue papers
39as first author
79since 2021 · last 2026
0000-0002-3124-5452ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 65 · 26 first-author · 38 since 2021Security and privacy · 21 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 6 first-author · 8 since 2021Computer networks · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CPSA-VD: Contrastive Cross-Procedural Semantic Alignment for Vulnerability Detection
Jinfu Chen 0001, Saihua Cai, Zhangpei Huang
COMPSAC2
2026 CrossRPL: Propagation-Aware Detection and Function-Level Localisation of Cross-Contract Vulnerabilities in Ethereum Smart Contracts
Rexford Nii Ayitey Sosu, Jinfu Chen 0001, Edward Kwadwo Boahen, Saihua Cai, Wenjie Gu
COMPSAC2
2026 TIPSO-GAN: Malicious Network Traffic Detection Using a Novel Optimized Generative Adversarial Network
Ernest Akpaku, Jinfu Chen 0001, Joshua Ofoeda
NDSS2
2026 SENTRY: an adversarial robust anomaly detection approach in system log based on pattern unit extraction and time-step masking
Bo Geng, Jinfu Chen 0001, Saihua Cai, Yisong Liu
Autom. Softw. Eng.2
2026 DP-S3: software defect prediction through feature fusion with syntax trees, program slices and standard features
abstract
Abstract Software defect prediction (SDP) is crucial for enhancing software quality and reducing development costs. Prevailing SDP methods often depend on traditional code metrics, which inadequately capture vital semantic information from source code, thereby limiting defect identification accuracy. This paper introduces DP-S3, a novel SDP model that integrates features from abstract syntax trees (ASTs), program slices, and standard metrics. DP-S3 extracts ASTs and program slices, transforming them into vector representations. A hierarchical long short-term memory network then learns semantic features from these vectors, which are combined with standard metrics from the PROMISE repository. A key innovation is our feature fusion strategy employing a channel self-attention mechanism to dynamically weight the three feature sets. We evaluated DP-S3 on seven open-source Java projects from the Apache repository against several state-of-the-art methods. The results demonstrate DP-S3’s superior performance, achieving average improvements of up to 3.8% in area under the receiver operating characteristic curve, 4.5% in F1, and 7.5% in Matthews correlation coefficient over baselines, showcasing its effectiveness. Key limitations include its current focus on Java projects and within-project defect prediction. Nevertheless, this work concludes that a synergistic fusion of syntactic, semantic (slice-based), and traditional metric features, guided by attention mechanisms, significantly enhances SDP capabilities and offers a promising direction for future research.
Jinfu Chen 0001, Jiaping Xu, Saihua Cai, Rexford Nii Ayitey Sosu
Comput. J.1
2026 Optimized mutation scheduling for fuzzing based on estimation of distribution algorithm
abstract
Abstract In the discovery of modern software vulnerability, mutation-based fuzzing techniques are widely applied, with their performance highly dependent on the effectiveness of mutation scheduling strategies. Most current research primarily focuses on optimizing seed scheduling. However, mutation scheduling plays an equally critical role in fuzzing, as it determines how mutation operators are selected and applied to generate new test cases. Existing mutation scheduling schemes are faced with several issues, such as the need to input manual parameters from users and improper overhead management. To address these challenges, this paper proposes an innovative mutation scheduling model, HavocEDA, utilizing Estimation of Distribution Algorithm to optimize the selection of mutation operators in fuzzing, thereby enhancing the efficiency of vulnerability detection. The HavocEDA model can dynamically adjust the probability distribution of mutation operators based on fuzzing feedback, enabling a more efficient exploration of potential vulnerabilities within the software. We implemented a prototype based on the popular general-purpose fuzzer AFL. We evaluated HavocEDA on 9 open-source Linux programs in FuzzBench. Experimental results indicate that HavocEDA shows a performance improvement in edge coverage and crash discovery compared with the state-of-the-art fuzzers AFL, DARWIN, MOpt, and HavocMAB.
Guofan Lv, Jinfu Chen 0001, Haibo Chen 0005, Saihua Cai
Cybersecur.2
2026 An efficient framework for malicious network traffic detection using optimized deep learning techniques
Mukhtar Ahmed, Jinfu Chen 0001, Ernest Akpaku, Ajmal Latif
Eng. Appl. Artif. Intell.2
2026 A novel android malware classification approach based on multi-scale feature fusion for encrypted traffic
Jinfu Chen 0001, Saihua Cai, Yisong Liu, Shengran Wang
Eng. Appl. Artif. Intell.2
2026 MD-CGM: Malicious traffic detection model based on CycleGAN and multi-head self-attetion mechanism
Saihua Cai, Yige Zhao, Jinfu Chen 0001, Shengran Wang, Bingbing Gu
Future Gener. Comput. Syst.4
2026 SiftFuzz: Boosting structural diversity via efficient seed fusion
Jinfu Chen 0001, Saihua Cai, Shengran Wang, Xingquan Mao
Inf. Softw. Technol.2
2026 SAME: A Similarity Analysis Method for Evaluating Metamorphic Relations in Testing AI systems
Jinfu Chen 0001, Tsong Yueh Chen, Saihua Cai
Inf. Softw. Technol.2
2026 A novel android malware detection method based on CWInFs and MPTACF optimization
Shengran Wang, Jinfu Chen 0001, Saihua Cai, Ernest Akpaku, Xingquan Mao
J. Inf. Secur. Appl.2
2026 GS-HF: An anomaly detection method for network traffic based on heterogeneous features and GraphSAGE
abstract
With the rapid growth of network traffic, data has become increasingly complex and voluminous, posing significant challenges for accurate anomaly detection. Traditional deep learning approaches often fail to capture the rich interdependencies and heterogeneous nature of traffic features, limiting their effectiveness in identifying subtle or evolving abnormal patterns. To address these challenges, this paper proposes GS-HF (GraphSAGE with heterogeneous features), an anomaly detection framework. The method integrates both statistical and image-like features extracted from raw traffic data to capture multi-dimensional characteristics, and then constructs a graph based on the fused heterogeneous features to better represent relationships among traffic flows. An improved GraphSAGE model is applied for detection, incorporating Focal Loss to handle class imbalance in real-world anomaly datasets. Extensive experiments on multiple network traffic datasets demonstrate that GS-HF achieves superior detection performance compared to existing methods, highlighting its robustness and effectiveness in handling diverse and complex traffic patterns.
Bo Geng, Jinfu Chen 0001, Saihua Cai, Haodi Xie, Yisong Liu
J. Comput. Secur.2
2026 MAGNN: Multi-scale adaptive graph neural networks with contrastive learning for malicious network traffic detection
Mukhtar Ahmed, Jinfu Chen 0001, Ernest Akpaku, Ali Bux
J. Parallel Distributed Comput.2
2026 DEzzer: Efficient Fuzzing Mutation Scheduling Based on Differential Evolution
Jinfu Chen 0001, Wenjun Feng, Saihua Cai, Xingquan Mao, Yisong Liu
J. Syst. Softw.1
2026 A novel seed scheduling scheme using Thompson sampling for coverage-guided greybox fuzzing
Jinfu Chen 0001, Saihua Cai, Yisong Liu, Haotong Ding
J. Syst. Softw.2
2026 DATVD: A novel vulnerability detection method based on dynamic attention and hybrid convolutional pooling
Jinfu Chen 0001, Jinyu Mu, Saihua Cai, Jiapeng Zhou, Xinping Shi
Sci. Comput. Program.1
2026 CL-ViME: Contrastive Learning and Vision Mixture of Experts for Encrypted Traffic Classification
abstract
Network traffic classification is essential for application identification and malicious behavior detection. However, the widespread use of encryption protocols hides payloads and reduces the availability of high-quality labeled data, both of which constrain the effectiveness of current models. To address these challenges, we propose CL-ViME, a self-supervised encrypted traffic classification framework that integrates Contrastive Learning and Vision Mixture of Experts. First, we design a packet-temporal matrix that preserves fine-grained packet headers and flow-level temporal structure. Second, we introduce a Vertical Vision Transformer-Mixture of Experts model to extract dual-view features through vertical patching and dynamic expert routing. Third, we develop a dual-granularity contrastive learning framework that aligns packet-level and flow-level representations via an MoE projector, followed by lightweight classifier-head fine-tuning. Experiments on three public datasets show that CL-ViME significantly outperforms state-of-the-art self-supervised and supervised baselines across accuracy, macro-precision, macro-recall, and macro-F1. It also demonstrates strong generalization and stability.
Saihua Cai, Lizhou Chen, Jinfu Chen 0001, Shengran Wang, Guofeng Zhang 0015
IEEE Trans. Netw. Serv. Manag.3
2025 IDBFuzz: Web Storage DataBase Fuzzing with Controllable Semantics
abstract
Despite great progress in fuzzing browser APIs, systematic approaches for testing web storage techniques remain absent. IndexedDB, the most popular NoSql database in modern browsers, brings unique challenges for fuzzing its API due to its asynchronous event-driven feature and strict phase separation. Current browser fuzzing techniques frequently struggle to generate nested event flows and invocations, which significantly impacts semantic correctness. Moreover, they often rely heavily on the try-catch block to suppress exceptions, which introduces substantial performance overhead. We propose IDBFuzz, the first fuzzing approach tailored for the IndexedDB API, which effectively tackles the challenge of capturing the execution context and event semantics inherent to IndexedDB, as well as handling large persistent objects. We design a seed generator based on intermediate representation (IR) that decouples layered IR skeletons from input object generation. With the aid of a global database snapshot, IDBFuzz can generate semantically controllable seeds, enabling the efficient production of high-quality test cases that significantly improve coverage.
Jinfu Chen 0001, Saihua Cai, Shengran Wang
ASE2
2025 DiFuzzNMT: A Differential Fuzzing Framework for Neural Machine Translation
abstract
Neural machine translation (NMT) systems have been widely deployed in real-world applications. However, despite the great translation performance, these systems inevitably face robustness issues and sometimes produce erroneous outputs, particularly in complex and ambiguous scenarios. In this work, we propose a Fuzzing framework with Differential Testing for NMT systems, namely, DiFuzzNMT. DiFuzzNMT employs a heuristic strategy to continuously search inputs that produce greater output differences across various NMT systems for error detection. Specifically, DiFuzzNMT establishes token mappings between outputs of different NMT systems for the same input through word alignment. To guide the test input generation process of fuzzing, we design specific testing guidance that takes into account the differences between outputs, including word alignment differences and token semantic differences. All the test inputs are generated based on “seed” inputs (inputs to generate new inputs) by applying a mutation operator. Test inputs exhibiting higher testing guidance values are selected as new seeds, while the others are discarded. A potential translation error is reported when the same test input exhibits significant differences across different NMT systems. By iteratively retaining seeds and generating test inputs, DiFuzzNMT can effectively detect translation errors. To evaluate the effectiveness of DiFuzzNMT, we conduct experiments on two widely used NMT APIs (Baidu Translate and Tencent Translate), using a publicly available dataset of 800 original sentences across 8 thematic categories. The experimental results show that DiFuzzNMT detects more translation errors than baselines and exhibits greater diversity. Furthermore, the results show that the proposed testing guidance improves the method's ability to detect translation errors.
Haibo Chen 0005, Jinfu Chen 0001, Saihua Cai, Shengran Wang
QRS2
2025 Detecting encrypted malicious traffic with HEAT: a header-focused deep learning approach
abstract
Abstract The widespread adoption of encryption in network traffic significantly challenges traditional detection methods that rely on payload analysis. Existing approaches often convert traffic into images or sequences for deep learning models, producing redundant features and struggling with multi-protocol environments. In this study, we propose HEAT (Header-Embedded Attention for Traffic Detection), a novel model that leverages packet header fields to develop a robust characteristic representation for encrypted traffic analysis. HEAT introduces a hierarchical attention mechanism combined with a novel contextual embedding technique that enhances the semantic representation of header field values. Additionally, HEAT integrates an adapted Kolmogorov–Arnold Network classifier with B-spline activations and L1 weight regularization, optimizing the model for efficient real-time processing. Extensive evaluations on CICIDS-2018, Stratosphere, and ISCX2012 datasets demonstrate HEAT’s superior performance, achieving 98.95% accuracy and 98.28% F1-score on CICIDS-2018, 99.5% accuracy and 98.54% F1-score on Stratosphere, and 99.75% accuracy with 99.25% F1-score on ISCX2012. HEAT significantly outperforms CNN, LSTM, and BiGRU baselines. Moreover, it maintains detection accuracy above 98.95% during incremental learning, with only a 0.9% F1-score drop, compared with 6.55% in conventional models. These results highlight HEAT’s novelty, stability, and adaptability, making it a scalable and robust solution for encrypted malicious traffic detection.
Ernest Akpaku, Jinfu Chen 0001, Mukhtar Ahmed, William Leslie Brown-Acquaye, Francis Kwadzo Agbenyegah, Rexford Nii Ayitey Sosu
Comput. J.2
2025 BiRNN-SA: Context-aware malicious network traffic detection using self-attentive bidirectional RNNs
Mukhtar Ahmed, Jinfu Chen 0001, Ernest Akpaku, Ajmal Latif
Comput. Networks2
2025 MTCR-AE: A Multiscale Temporal Convolutional Recurrent Autoencoder for unsupervised malicious network traffic detection
Mukhtar Ahmed, Jinfu Chen 0001, Ernest Akpaku, Rexford Nii Ayitey Sosu
Comput. Networks2
2025 RAGN: Detecting unknown malicious network traffic using a robust adaptive graph neural network
Ernest Akpaku, Jinfu Chen 0001, Mukhtar Ahmed, Francis Kwadzo Agbenyegah, William Leslie Brown-Acquaye
Comput. Networks2
2025 APT-ATT: An efficient APT attribution model based on heterogeneous threat intelligence representation and CTGAN
Saihua Cai, Jinfu Chen 0001, Shengran Wang
Comput. Networks3
2025 CDDA-MD: An efficient malicious traffic detection method based on concept drift detection and adaptation technique
Saihua Cai, Jinfu Chen 0001, Yikai Hu, Wuhao Guo
Comput. Secur.3
2025 eBiTCN: Efficient bidirectional temporal convolution network for encrypted malicious network traffic detection
abstract
The growing prevalence of encrypted malicious network traffic poses significant challenges for cybersecurity, as it conceals the content from traditional detection methods. Temporal convolutional networks (TCNs) present promising capabilities for extracting complex temporal features and patterns from the dynamic traffic flow data. However, the unidirectional nature of traditional TCNs limits their effectiveness in capturing the full context of network traffic, which often exhibits bidirectional temporal dependencies. Consequently, a few studies have proposed bidirectional TCN (BiTCN) architectures to address the limitations. However, these methods present complex architectures that require a significant amount of parameters to be learned, which imposes high memory requirements on the computational resources for training such models. In this study, we introduce the efficient bidirectional TCN (eBiTCN) model, an efficient BiTCN that requires fewer parameters yet not at the expense of computational cost and effective detection. The eBiTCN framework combines a bidirectional processor, a lightweight gating mechanism, temporal attention, dropout, a novel loss function, and dense layers. Extensive experiments show that eBiTCN outperforms eight state-of-the-art competing models in terms of detection efficacy, speed, and scalability. The eBiTCN model showcased robust performance in detecting evolving attacks and excelled across various real-world datasets. Its efficiency in training speed and reduced memory usage translates to lower infrastructure costs, making it an accessible and effective choice for deployment. These findings highlight eBiTCN’s practicality and dependability in addressing contemporary network security needs.
Ernest Akpaku, Jinfu Chen 0001, Mukhtar Ahmed, Rexford Nii Ayitey Sosu, Francis Kwadzo Agbenyegah, Dominic Kofi Louis
J. Comput. Secur.2
2025 MTD-FRD: Malicious traffic detection method based on feature representation and conditional diffusion model
Saihua Cai, Jinfu Chen 0001, Yige Zhao, Shengran Wang
J. Netw. Comput. Appl.3
2025 CT-SSSA: Malicious traffic augmentation based on classifier transGAN and spatial-channel synergistic self-attention
Saihua Cai, Jinfu Chen 0001, Yige Zhao, Lizhou Chen
Knowl. Based Syst.3
2025 A Novel Vulnerability-Detection Method Based on the Semantic Features of Source Code and the LLVM Intermediate Representation
abstract
ABSTRACT With the increasingly frequent attacks on software systems, software security is an issue that must be addressed. Within software security, automated detection of software vulnerabilities is an important subject. Most existing vulnerability detectors rely on the features of a single code type (e.g., source code or intermediate representation [IR]), which may lead to both the global features of the code slices and the memory operation information not being captured or considered. In particular, vulnerability detection based on source‐code features cannot usually include some macro or type definition content. In this paper, we propose a vulnerability‐detection method that combines the semantic features of source code and the low level virtual machine (LLVM) IR. Our proposed approach starts by slicing (C/C++) source files using improved slicing techniques to cover more comprehensive code information. It then extracts semantic information from the LLVM IR based on the executable source code. This can enrich the features fed to the artificial neural network (ANN) model for learning. We conducted an experimental evaluation using a publicly‐available dataset of 11,381 C/C++ programs. The experimental results show the vulnerability‐detection accuracy of our proposed method to reach over 96% for code slices generated according to four different slicing criteria. This outperforms most other compared detection methods.
Jinfu Chen 0001, Jiapeng Zhou, Dave Towey, Saihua Cai, Haibo Chen 0005, Yemin Yin
J. Softw. Evol. Process.1
2025 Predicting Vulnerabilities in Computer Source Code Using Non-Investigated Software Metrics
Francis Kwadzo Agbenyegah, Jinfu Chen 0001, Micheal Asante, Ernest Akpaku
Softw. Qual. J.2
2025 DialTest-EA: An Enhanced Fuzzing Approach With Energy Adjustment for Dialogue Systems via Metamorphic Testing
abstract
ABSTRACT Deep neural networks (DNNs) possess potent feature learning capability, enabling them to comprehend natural language, which strongly support developing dialogue systems. However, dialogue systems usually perform incorrect behaviours in some corner cases, which may cause misunderstanding or economic loss. To test and debug dialogue systems, a popular fuzzing framework by metamorphic testing with Gini impurity guidance is proposed, namely, DialTest. However, DialTest treats all seeds (the initial test inputs to generate the mutated test inputs) equally during the fuzzing process and does not differentiate seeds, resulting in a certain limitation to its incorrect behaviour detection capability. In this paper, we propose to enhance the DialTest by applying a lightweight energy adjustment strategy called DialTest with Energy Adjustment (DialTest‐EA). DialTest‐EA employs the ant colony optimization algorithm (ACO) to adjust the mutation energy of each seed adaptively, ensuring that potential seeds have more opportunities to generate subsequent test inputs. To evaluate the effectiveness of the proposed DialTest‐EA, we conduct a series of comparisons with the original DialTest and random mutation strategy. The experimental results show that the proposed DialTest‐EA outperforms the compared methods both in the intent detection and slot filling tasks. Compared with the original DialTest, the intent detection accuracy of generated test cases by the proposed method is reduced by more than 14%, and the slot filling accuracy is reduced by more than 8%.
Haibo Chen 0005, Jinfu Chen 0001, Saihua Cai, Rubing Huang, Shengran Wang, Chi Zhang 0046
Softw. Test. Verification Reliab.2
2025 MGAN: A Multi-view Graph Adaptive Network for Robust Malicious Traffic Detection
abstract
Detecting malicious network traffic in large-scale, dynamic environments presents a significant challenge due to the complexity of network relationships and the evolving nature of cyber threats. Existing graph-based and sequence-based models often fail to capture both spatial dependencies and temporal patterns effectively, resulting in suboptimal detection. This study introduces the Multi-view Graph Adaptive Network (MGAN), a novel framework that integrates multi-hop graph neural network (GNN) aggregation with transformer-based sequence modeling to address these challenges. MGAN captures long-range spatial dependencies and temporal dynamics in network traffic, enabling the detection of complex attack patterns. It incorporates Dirichlet sampling for robust neighbor selection in sparse and noisy data environments and mutual information maximization to align multi-view representations for consistency. Additionally, a multi-view attention mechanism aggregates information across different hops, balancing local and global network context. Extensive experiments on four real-world datasets demonstrate MGAN’s superiority over 7 baseline models, achieving an average F1-Score above 97%, surpassing the best baseline by 2.35%. MGAN maintains detection accuracy above 97% and remains robust under data sparsity, achieving F1-Scores over 95% even when 40% of connectivity information is removed. Under noisy conditions, MGAN retains accuracy above 93%, outperforming baselines by over 4.5%. In zero-day attack scenarios, it achieves detection rates exceeding 96% for previously unseen attack categories. MGAN also exhibits exceptional computational efficiency, processing 2,034 samples per second with a detection time of 3.00 milliseconds per sample, outperforming all competing models in both accuracy and speed.
Ernest Akpaku, Jinfu Chen 0001, Mukhtar Ahmed, Francis Kwadzo Agbenyegah, Joshua Ofoeda
ACM Trans. Priv. Secur.2
2025 GSA-DT: A Malicious Traffic Detection Model Based on Graph Self-Attention Network and Decision Tree
abstract
Malicious attack has shown a rapid growth in recent years, it is very important to accurately detect malicious traffic to defend against malicious attacks. Compared with machine learning and deep learning technologies,graphconvolutional neuralnetwork (GCN) achieves better detection results of malicious traffic due to additional consideration of the correlation between network traffic features. However, existing GCN-based detection models suffer from fixed weight assignment, only focusing on local features, lack the ability to model graph structure and relationships as well as having gradient disappearance. To solve these problems, this paper proposes the GSA-DT model based ongraphself-attention network anddecisiontree. GSA-DT first preprocesses the original network traffic to obtain better traffic features and labels, and then uses GCN to extract the topological structure of network traffic as well as capture the correlation relationships among traffic features, where the ReLU activation function is replaced by LeakyReLU to overcome the problems of neuron “death” and gradient disappearance during the training process; It also introduces the self-attention mechanism into GCN to assign larger weights to the key features to reduce the interference of redundant features. Finally, GSA-DT uses decision tree to perform the detection of malicious traffic. Experimental results on four network traffic datasets show that GSA-DT model improves the detection accuracy over 1% on average than seven advanced malicious traffic detection models, and it also performs better in F1-measure, TPR, FPR as well as stability.
Saihua Cai, Jinfu Chen 0001, Tianxiang Lv, Chunlei Huang
IEEE Trans. Netw. Serv. Manag.3
2024 FMUZZ: A Novel Greybox Fuzzing Approach based on Mutation Strategy Optimization with Byte Scheduling
abstract
Mutation-based greybox fuzzing is an efficient and widely used software testing technique, and its performance heavily depends on the mutation strategy. Existing solutions guide the seed mutation by using program-adaptive mutation strategies or constraint solving techniques. However, they disregard the characteristic that the execution information of seeds with similar behavior contains general strategies for solving specific constraints. In this paper, we propose the FMUZZ, a lightweight fuzzing approach based on mutation strategy optimization. FMUZZ first clusters the seeds based on their execution information into different seed groups and then learns the byte mutation scheduling strategies applicable to different program paths to improve efficiency in generating seeds that satisfy specific branch constraints. Meanwhile, FMUZZ removes the redundant seeds during the learning process by using the customized multi-objective optimization algorithm, thereby improving the efficiency of learning byte mutation scheduling strategies for different program paths. We test the effectiveness of FMUZZ on 9 real-world programs with the comparison of 3 state-of-the-art mutation-based fuzzers. Extensive experimental results show that compared to the benchmark fuzzers, FMUZZ achieves 8.9% higher branch coverage and outperforms 35.3% in discovering unique crashes on average.
Jinfu Chen 0001, Saihua Cai, Shengran Wang
QRS1
2024 DA-CPVD: Vulnerability Detection Method based on Dual Attention Composite Pooling
abstract
Source code vulnerability detection is of great significance in securing software as well as addressing novel threats, and neural network-based methods have made significant progress in the field of vulnerability detection. However, the widely used neural network-based vulnerability detection methods suffer from the loss of complex structural and semantic information in the source code. To solve this problem, this paper proposes a dual-attention composite pooling-based vulnerability detection method called DA-CPVD for source code, it enhances the feature representation through fully considering the overall contextual information and complex dependencies of source code. The key of DA-CPVD is the use of a dual-attention composite pooling approach to emphasizes the key features more flexibly, thereby forming a comprehensive pooled feature representation. In specific, DA-CPVD utilizes a self-attention mechanism to adaptively assign the weights to features, as well as uses a composite pooling based on attention mechanism aggregation to dynamically adjust the weights of pooling results. The DA-CPVD method is evaluated on three publicly available and widely used datasets, and the experimental result shows that DA-CPVD improves Accuracy, Precision and F1-measure by an average of 11.52%, 28.7%, and 10.45%, and reduces False Positive Rate by an average of 19.15% compared to other existing methods.
Mengxuan Shi, Jinfu Chen 0001, Saihua Cai, Jiapeng Zhou
TrustCom2
2024 FD-WF: A Multi-tab Website Fingerprinting Attack Based on Fixed Dimensions for Tor Network
abstract
Website Fingerprinting Attack (WFA) is an effective method of network monitoring, which analyzes network traffic to identify the specific website or web page that a user is browsing. The performance of previous WFA, which assume singletab scenarios, deteriorates significantly in real-world multi-tab environments. While this issue has been acknowledged and some research has been conducted on multi-tab WFA, limitations in the accuracy and efficiency of multi-tab classification still persist. Specifically, extended research on mixed-tab scenarios, where the number of tabs is unknown, still lacks sufficient attention. In this paper, we propose FD-WF, a multi-tab WFA model based on fixed dimensions on the Tor network. This model mitigates the issue of blurred features in single-tab images caused by the expansion of multi-tab mixed traffic sequences. It enhances the ability to accurately identify and classify multiple web pages users are browsing. We propose a minimum padding optimization function to improve the performance of fixed-dimension image generation and introduce an enhanced ResNet-18 model to better address the fingerprint classification challenge for multi-tab scenarios. The experimental results verified the feasibility and effectiveness of the proposed method. While ensuring the performance of singletab WFA, our model achieved a F1 Score of 96% for multi-tab web pages. Even in the more complex mixed-tab scenario, we still achieved an F1 Score of 78%.
Shangnan Yin, Jinfu Chen 0001
TrustCom3
2024 GCN-MHSA: A novel malicious traffic detection method based on graph convolutional neural network and multi-head self-attention mechanism
Jinfu Chen 0001, Haodi Xie, Saihua Cai, Luo Song, Bo Geng, Wuhao Guo
Comput. Secur.1
2024 Exploiting DBSCAN and Combination Strategy to Prioritize the Test Suite in Regression Testing
abstract
Test case prioritization techniques improve the fault detection rate by adjusting the execution sequence of test cases. For static black‐box test case prioritization techniques, existing methods generally improve the fault detection rate by increasing the early diversity of execution sequences based on string distance differences. However, such methods have a high time overhead and are less stable. This paper proposes a novel test case prioritization method (DC‐TCP) based on density‐based spatial clustering of applications with noise (DBSCAN) and combination policies. By introducing a combination strategy to model the inputs to generate a mapping model, the test inputs are mapped to consistent types to improve generality. The DBSCAN method is then used to refine the classification of test cases further, and finally, the Firefly search strategy is introduced to improve the effectiveness of sequence merging. Extensive experimental results demonstrate that the proposed DC‐TCP method outperforms other methods in terms of the average percentage of faults detected and exhibits advantages in terms of time efficiency when compared to several existing static black‐box sorting methods.
Zikang Zhang, Jinfu Chen 0001, Yuechao Gu, Rexford Nii Ayitey Sosu
IET Softw.2
2024 Hybrid semantics-based vulnerability detection incorporating a Temporal Convolutional Network and Self-attention Mechanism
Jinfu Chen 0001, Bo Liu 0048, Saihua Cai, Dave Towey, Shengran Wang
Inf. Softw. Technol.1
2024 DCM-GIFT: An Android malware dynamic classification method based on gray-scale image and feature-selection tree
Jinfu Chen 0001, Zian Zhao, Saihua Cai, Xiao Chen 0003, Luo Song
Inf. Softw. Technol.1
2024 L′OP-ART: A linear-time adaptive random testing algorithm for object-oriented programs
Jinfu Chen 0001, Lili Zhu, Chengying Mao, Qihao Bao, Rubing Huang
J. Syst. Softw.1
2024 iGnnVD: A novel software vulnerability detection model based on integrated graph neural networks
Jinfu Chen 0001, Yemin Yin, Saihua Cai, Shengran Wang
Sci. Comput. Program.1
2024 TR-Fuzz: A syntax valid tool for fuzzing C compilers
Chi Zhang 0046, Jinfu Chen 0001, Saihua Cai, Rexford Nii Ayitey Sosu, Haibo Chen 0005
Sci. Comput. Program.2
2024 A novel test case prioritization approach for black-box testing based on K-medoids clustering
abstract
Abstract Regression testing is an essential and expensive process in software testing. However, there may be insufficient resources for the execution of all test cases during regression testing. Test case prioritization (TCP) techniques improve the efficiency of regression testing by adjusting the test case execution sequence. Traditional TCP techniques usually rely on the historical execution information of the software under test for more efficient results. String distance‐based TCP (SD‐TCP) avoids these limitations; it uses only the textual difference information of the test cases themselves for prioritization. However, the time overhead on the sorting process of this method is not ideal, and the extreme test case inputs have an impact on the stability of the method. To address these problems, we propose a novel test case prioritization strategy, it first classifies the test cases more finely using the K‐medoids algorithm and then transforms the set into subsequences and improves the early diversity by greedy sorting within clusters. Finally, the test cases are selected through a polling strategy to compose the execution sequence. Extensive experimental results demonstrate that the proposed approach outperforms SD‐TCP in better time efficiency on test case prioritization; it also has a higher average percentage of fault detected (APFD) value than random prioritization (RP) and SD‐TCP.
Jinfu Chen 0001, Yuechao Gu, Saihua Cai, Haibo Chen 0005
J. Softw. Evol. Process.1
2024 A novel defect prediction method based on semantic feature enhancement
abstract
Summary Although cross‐project defect prediction (CPDP) techniques that use traditional manual features to build defect prediction model have been well‐developed, they usually ignore the semantic and structural information inside the program and fail to capture the hidden features that are critical for program category prediction, resulting in poor defect prediction results. Researchers have proposed using deep learning to automatically extract the semantic features of programs and fuse them with traditional features as training data. However, in practice, it is important to explore the effective representation of the semantic features in the programs and how the fusion of a reasonable ratio between the two types of features can maximize the effectiveness of the model. In this paper, we propose a semantic feature enhancement‐based defect prediction framework (SFE‐DP), which augments the semantic feature set extracted from the program code with data. We also introduce a layer of self‐attentive mechanism and a matching layer to filter low‐efficiency and non‐critical semantic features in the model structure. Finally, we combine the idea of hybrid loss function to iteratively optimize the model parameters. Extensive experiments validate that SFE‐DP can outperform the baseline approaches on 90 pairs of CPDP tasks formed by 10 open‐source projects.
Chi Zhang 0046, Jinfu Chen 0001, Saihua Cai, Rexford Nii Ayitey Sosu
J. Softw. Evol. Process.3
2024 MUT Model: a metric for characterizing metamorphic relations diversity
Jinfu Chen 0001, Patrick Kwaku Kudjo
Softw. Qual. J.3
2024 DELM: Deep Ensemble Learning Model for Anomaly Detection in Malicious Network Traffic-based Adaptive Feature Aggregation and Network Optimization
abstract
With the rapid advancements in internet technology, the complexity and sophistication of network traffic attacks are increasing, making it challenging for traditional anomaly detection systems to analyze and detect malicious network attacks. The increasing advancedness of cyber threats calls for innovative approaches to identify malicious patterns within network traffic precisely. The primary issue lies in the fact that these approaches do not focus on the essential adaptive features of network traffic. We proposed an effective anomaly detection system for malicious network traffic attacks called the Deep Ensemble Learning Model (DELM). We leverage the structure of the Feedforward Deep Neural Network (FDNN), and Deep Belief Network (DBN), incorporating multiple hidden layers with non-linear activation functions. Integrating Adaptive Feature Aggregation (AFA) with the FDNN algorithm dynamically adjusts the feature aggregation process based on incoming traffic characteristics to improve adaptability. The Conditional Generative Network was employed to enhance DELM for generating data for minority classes. To improve the model’s accuracy, we applied batch normalization and data augmentation techniques for preprocessing, utilized n-gram, one-hot encoding, and feature aggregation methods for effective feature extraction. This study significantly contributes to network security by enhancing systems for detecting malicious network traffic. With its interpretability and adaptability, our proposed model shows promise in addressing the evolving cyber threat and fortifying critical network infrastructure. The experimental results demonstrate that our model performs with higher stability than the existing state-of-the-art detection approaches, as reflected by its higher accuracy, precision, recall, F1-score, and AUC-ROC.
Mukhtar Ahmed, Jinfu Chen 0001, Ernest Akpaku, Rexford Nii Ayitey Sosu, Ajmal Latif
ACM Trans. Priv. Secur.2
2024 Software Defect Prediction Approach Based on a Diversity Ensemble Combined With Neural Network
abstract
There is a severe class imbalance problem in defect datasets, with nondefective data dominating the distribution, making it easy to generate inaccurate software defect prediction models. Ensemble learning has been proven to be one of the best methods to solve class imbalance problem. Traditional ensemble prediction models usually ensemble the results of several base classifiers simply, and most of them only ensemble once, rarely consider the diversity of ensemble or the combination of ensemble learning and neural network. In order to explore whether the secondary ensemble of classifiers based on a diversity ensemble combined with neural network can improve the performance of defect prediction model, in this article, we propose a novel dual ensemble software defect prediction (DE-SDP) approach based on a diversity ensemble combined with neural network. In the first ensemble, we use cross-validation to build different subclassifiers, then, these subclassifiers are used to establish base ensemble classifiers with weighted average method. Through seven classification algorithms, seven base ensemble classifiers can be established. In the second ensemble, a neural network model and stacking are used to ensemble the base ensemble classifiers again. We have evaluated DE-SDP against other ensemble defect prediction methods on eight datasets of NASA MDP. The results show that our approach is superior to other ensemble approaches and effectively improves the performance of defect prediction model.
Jinfu Chen 0001, Jiaping Xu, Saihua Cai, Haibo Chen 0005
IEEE Trans. Reliab.1
2023 CGSA-RNN: Abnormal Network Traffic Detection Model Based on CycleGAN and Self-Attention Mechanism
abstract
Malicious attack is a major factor to endanger the cyberspace security. The accurate detection of abnormal network traffic generated by malicious attacks can effectively detect potential malicious attacks and thus protecting the network security. However, the scale of abnormal network traffic is relatively small (i.e., there is a data imbalance phenomenon), which causes a significant decrease of detection accuracy. The introduction of CycleGAN model can effectively deal with the data imbalance phenomenon, but it suffers from semantic inconsistency, image distortion and lack of diversity. This paper proposes an abnormal network traffic detection model called CGSA-RNN that incorporates CycleGAN, self-attention mechanism and RNN to overcome the disadvantages of CycleGAN model, thereby accurately detecting abnormal network traffic. CGSA-RNN model first takes the advantage of style migration of CycleGAN to perform data augmentation for the small-scale abnormal network traffic, and then replaces the ReLU activation function with LeakyReLU in the CycleGAN generator to reduce the effects of artifacts and distortion in the generated images. In addition, a self-attention mechanism is introduced into the CycleGAN to help it to better capture important features, thereby further improving the data augmentation capability. Finally, CGSA-RNN uses the RNN model to detect abnormal network traffic. Extensive experimental results on two publicly available network traffic datasets show that compared with four advanced detection models based on data augmentation, the average precision, recall and F1-measure of CGSA-RNN model are improved by more than 2%.
Saihua Cai, Jinfu Chen 0001, Wuhao Guo
QRS4
2023 VDABSys: A Novel Security-Testing Framework for Blockchain Systems Based on Vulnerability detection
Jinfu Chen 0001, Qiaowei Feng, Saihua Cai, Dengzhou Shi, Dave Towey
SecureComm (1)1
2023 EcoDialTest: Adaptive Mutation Schedule for Automated Dialogue Systems Testing
abstract
With the rapid growth of Artificial Intelligence, dialogue systems have become increasingly powerful. Though Recurrent Neural Network power the dialogue systems, it also bring challenges to the systems’ testing. In order to ensure the safety of these systems, which we must pay attention to, DialTest showed up. DialTest broke the traditional test methods, it made innovation at many levels. We have to acknowledge this great contribution. However, DialTest has a smattering of shortcomings. It treats all seeds as equal, implying that it cannot adjust the energy assignment quickly, resulting in energy waste. Moreover, DialTest’s mutant sentences generated by a few original seed sentences in the late stage of variation. This paper presents an improved DialTest with an adaptive mutation schedule, we called it EcoDialTest. EcoDialTest divides all the seed into three states, different states have different energy distribution strategies. We devise a new mutation strategy to improve the effectiveness and dependability of the seeds in the transformed seed set. All of these were implemented based on DialTest, we still adopt DeepGini impurity as the main guidance to guide the test generation process and utilize the three mutation operators as it does. Through ATIS, Snips and Facebook datasets, EcoDialTest was evaluated by two state-of-the-art models in the experiment. According to the result, we found that EcoDialTest attained lower values in both intent accuracy and slot accuracy than DialTest.
Xiangchen Shen, Haibo Chen 0005, Jinfu Chen 0001, Shuhui Wang
SANER3
2023 Minimal Rare Pattern-Based Outlier Detection Approach For Uncertain Data Streams Under Monotonic Constraints
abstract
Abstract Existing association-based outlier detection approaches were proposed to seek for potential outliers from huge full set of uncertain data streams ($UDS$), but could not effectively process the small scale of $UDS$ that satisfies preset constraints; thus, they were time consuming. To solve this problem, this paper proposes a novel minimal rare pattern-based outlier detection approach, namely Constrained Minimal Rare Pattern-based Outlier Detection (CMRP-OD), to discover outliers from small sets of $UDS$ that satisfy the user-preset succinct or convertible monotonic constraints. First, two concepts of ‘maximal probability’ and ‘support cap’ are proposed to compress the scale of extensible patterns, and then the matrix is designed to store the information of each valid pattern to reduce the scanning times of $UDS$, thus decreasing the time consumption. Second, more factors that can influence the determination of outlier are considered in the design of deviation indices, thus increasing the detection accuracy. Extensive experiments show that compared with the state-of-the-art approaches, CMRP-OD approach has at least 10% improvement on detection accuracy, and its time cost is also almost reduced half.
Saihua Cai, Jinfu Chen 0001, Haibo Chen 0005, Chi Zhang 0046, Qian Li 0042, Dengzhou Shi
Comput. J.2
2023 An optimized feature extraction algorithm for abnormal network traffic detection
Jinfu Chen 0001, Saihua Cai, Shang Yin, Lingling Zhao, Zikang Zhang
Future Gener. Comput. Syst.1
2023 A novel detection model for abnormal network traffic based on bidirectional temporal convolutional network
Jinfu Chen 0001, Tianxiang Lv, Saihua Cai, Luo Song, Shang Yin
Inf. Softw. Technol.1
2023 A memory-related vulnerability detection approach based on vulnerability model with Petri Net
Jinfu Chen 0001, Chi Zhang 0046, Saihua Cai
J. Log. Algebraic Methods Program.1
2023 BiTCN_DRSN: An effective software vulnerability detection model based on an improved temporal convolutional network
Jinfu Chen 0001, Saihua Cai, Yemin Yin, Haibo Chen 0005, Dave Towey
J. Syst. Softw.1
2023 A novel combinatorial testing approach with fuzzing strategy
abstract
Summary Combinatorial testing (CT) is considered as a practical approach to detect software faults, which has arisen from the interaction between factors affecting the software behavior. However, most of the traditional algorithms on CT generation did not take advantage of the execution results of the earlier test cases, as well as neglect the impact of the nonequilibrium input parameter model (NE‐IPM) effect on redundant test cases, which bring a deleterious effect to the detection accuracy of the software faults. To solve these problems, we propose a novel CT approach with fuzzing strategy called CTAF. Based on the idea that fuzzing is performed during execution, CTAF exploits the execution results of earlier tests to provide guidance for subsequent test generation thereby reducing the redundant test cases without compromising the diversity of test cases. And then, we designed three experiments on real subjects of six open source software systems, and the experimental results show that the proposed CTAF approach can effectively improve the NE‐IPM effect and enhance the detection accuracy of software faults.
Jinfu Chen 0001, Saihua Cai, Haibo Chen 0005, Chi Zhang 0046
J. Softw. Evol. Process.1
2023 Bug detection in Java code: An extensive evaluation of static analysis tools using Juliet Test Suites
abstract
Abstract Previous studies have demonstrated the usefulness of employing automated static analysis tools (ASAT) and techniques to detect security bugs in software systems. However, these studies are usually focused on analyzing the effectiveness of the tools using open‐source tools based on C/C++ source code. The choice for making an appropriate decision on the most suitable tool for bug detection in Java code software remains a relatively unexplored domain. To address this deficiency, this study empirically evaluates eight widely used ASATs, namely, Findbug, PMD, YASCA, LAPSE+, JLint, Bandera, ESC/Java, and Java Pathfinder using the Juliet Test Suite (Test Suite v1.2). Additionally, we assessed the performance of the detection capabilities for the aforementioned bug detection tools using robust performance measures such as precision, recall, Youden index, and the OWASP web benchmark evaluation (WBE). The experimental results show that the tools obtain precision values ranging from 83% to 90.7% based on the studied datasets. Specifically, the Java Pathfinder achieves the best precision score of 90.7%, followed by YASCA and Bandera with a precision score of 88.7% and 83%, respectively. Similarly, Bandera, ESC/Java, and Java Pathfinder obtain a Youden index of 0.8, which indicates the effectiveness of the tools in detecting security bugs in Java source code.
Richard Amankwah, Jinfu Chen 0001, Heping Song, Patrick Kwaku Kudjo
Softw. Pract. Exp.2
2023 TLS-MHSA: An Efficient Detection Model for Encrypted Malicious Traffic based on Multi-Head Self-Attention Mechanism
abstract
In recent years, the use of TLS (Transport Layer Security) protocol to protect communication information has become increasingly popular as users are more aware of network security. However, hackers have also exploited the salient features of the TLS protocol to carry out covert malicious attacks, which threaten the security of network space. Currently, the commonly used traffic detection methods are not always reliable when applied to the problem of encrypted malicious traffic detection due to their limitations. The most significant problem is that these methods do not focus on the key features of encrypted traffic. To address this problem, this study proposes an efficient detection model for encrypted malicious traffic based on transport layer security protocol and a multi-head self-attention mechanism called TLS-MHSA. Firstly, we extract the features of TLS traffic during pre-processing and perform traffic statistics to filter redundant features. Then, we use a multi-head self-attention mechanism to focus on learning key features as well as generate the most important combined features to construct the detection model, thereby detecting the encrypted malicious traffic. Finally, we use a public dataset to verify the effectiveness and efficiency of the TLS-MHSA model, and the experimental results show that the proposed TLS-MHSA model has high precision, recall, F1-measure, AUC-ROC as well as higher stability than seven state-of-the-art detection models.
Jinfu Chen 0001, Luo Song, Saihua Cai, Haodi Xie, Shang Yin
ACM Trans. Priv. Secur.1
2022 MPC: Multi-node Payment Channel for Off-chain Transactions
abstract
Payment channel (PC) greatly improves blockchain scalability by allowing an unlimited number of off-chain transactions instead of committing every transaction to the blockchain. However, one PC just confines to two nodes. Thus, for a node, like Cafe, who frequently receives money from multiple nodes, massive PCs need to be created, which leads to massively and repeatedly information addition of channels to the blockchain and costs much fees. In this paper, we introduce a multi-node payment channel (MPC) method to solve the problem above. MPC suits the scenario that one node frequently receives money from multiple nodes, such as Cafe, shop, etc. Compared to PC, MPC can contain nearly unlimited number of nodes, thus avoiding repetitive channel information addition to the blockchain and meanwhile reducing the fee. MPC supports the same transaction logic of PC and cooperates well with existing PC. The security, economy and efficiency of MPC are proved. We implement MPC’s smart contract in Ethereum and experimental results show that MPC outperforms the existing PC.
Longxia Huang, Liangmin Wang 0001, Jinfu Chen 0001
ICC4
2022 A novel classification approach for Android malware based on feature fusion and natural language processing
abstract
The growing use of Android software has made mobile devices the main platform for information services such as mobile social media and financial services. Mobile software provides great convenience but also brings challenges to the software community. For example, mobile malware, a malicious software specifically designed to target mobile devices, creates security concerns for the business network and the data stored on it. Therefore, it is becoming more and more important to effectively identify and classify malware. Most of the current malware-classification methods rely on the specific (static/dynamic) behaviour information from Android software for improved malware-detection capability. Nevertheless, these methods cannot detect new types of fraud software due to the limited generalisability. To address these issues, this paper proposes the AMC-FN, i.e. an Android-based malware classification method using feature fusion and natural language processing technologies. The proposed AMC-FN aims to improve the dimension and performance of classification and also some specific functions of natural language processing, i.e. mutual information method, n-gram word segmentation and feature mapping. The AMC-FN framework improves the classification dimensions by leveraging the information from Android APK permission, API calls and realistic network traffic. Moreover, the framework also contains a novel multi-level feature fusion algorithm (MFFA) designed to improve the weighted feature fusion. To obtain better fine granularity and generalisability, the fusion features are used by the optimized SVM (Support Vector Machine) classifier for training. Our experimental measurements and comparisons show the improved performance based on the proposed AMC-FN framework.
Jinfu Chen 0001, Zian Zhao, Xiao Chen 0003, Saihua Cai, Shang Yin, Luo Song
Internetware1
2022 An adaptive search optimization algorithm for improving the detection capability of software vulnerability
abstract
Deep learning-based vulnerability detection frees human experts from the tedious task of defining features and allows for better detection capabilities. The common practice is to convert program code into vector representation for neural network model training. Since the length of the vector representation varies across program code, finding the optimal vector length is critical to ensuring detection accuracy. This paper proposes an adaptive search optimization algorithm for finding the optimal vector length. It sorts all the vector lengths obtained by word2vec and takes the vector length corresponding to the point where the trend changes from slow to fast as the output. We evaluate our algorithm on three publicly available datasets against state-of-the-art algorithms. The results show that, without significantly increasing the time overhead, our algorithm can more accurately choose an appropriate vector length instead of setting a value empirically or arbitrarily. Furthermore, it shows that while a larger vector length can usually produces a higher detection accuracy, the extra time overhead incurred often does not suffice to compensate for the corresponding accuracy improvement.
Bo Liu 0048, Jinfu Chen 0001, Saihua Cai, Qiaowei Feng
Internetware2
2022 A Novel Coverage-guided Greybox Fuzzing based on Power Schedule Optimization with Time Complexity
abstract
Coverage-guided Greybox fuzzing is regarded as a practical approach to detect software vulnerabilities, which targets to expand code coverage as much as possible. A common implementation is to assign more energy to such seeds which find new edges with less execution time. However, solely considering new edges may be less effective because some hard-to-find branches often exist in the complex code of program. Code complexity is one of the key indicators to measure the code security. Compared to the code with simple structure, the program with higher code complexity is more likely to find more branches and cause security problems. In this paper, we propose a novel fuzzing method which further uses code complexity to optimize power schedule process in AFL (American Fuzzy Lop) and AFLFAST (American Fuzzy Lop Fast). The goal of our method is to generate inputs which are more biased toward the code with higher complexity of the program under test. In addition, we conduct a preliminary empirical study under three widely used real-world programs, and the experimental results show that the proposed approach can trigger more crashes as well as improve the coverage discovery.
Jinfu Chen 0001, Shengran Wang, Saihua Cai, Chi Zhang 0046, Haibo Chen 0005
ASE1
2022 Coverage-based Greybox Fuzzing with Pointer Monitoring for C Programs
abstract
C has been regarded as a dominant programming language for system software implementation. Meanwhile, it often suffers from various memory vulnerabilities due to its low-level memory control. Quite massive approaches are proposed to enhance memory security, among which Coverage-based Greybox Fuzzing (CGF) is very popular because of its practicality and satisfactory effectiveness. However, CGF identifies vulnerabilities based on the catched crashes, thus cannot detect vulnerabilities with non-crash. In this paper, we consider to trace pointer metadata (status, bounds and referents) to detect more various vulnerabilities. Additionally, since pointers in C are often directly related to memory operations, we design two standards to further use pointer metadata as the guidance of CGF, making fuzzing process target to the vulnerable part of programs.
Haibo Chen 0005, Jinfu Chen 0001
ASE2
2022 A formalization-based vulnerability detection method for cross-subject network components
abstract
With the rapid development of computer technology, the cross-subject network components (CSNC) is widely used in software. However, the existing of vulnerabilities in CSNC may seriously affect the security of software, which attracts the attention of software tester. This paper proposes a formal-based vulnerability detection method called FVDM for CSNC to detect the security vulnerabilities and defects in the logic of components. The proposed FVDM firstly selects the singleton as the medium of abstract computation as well as uses the formal description language to construct a vulnerability propagation model; And then, the FVDM classifies the vulnerabilities into explicit and implicit vulnerabilities through analyzing the types of vulnerabilities, thereby designing the vulnerability detection algorithm for explicit vulnerabilities and implicit vulnerabilities respectively. The experimental results on several COM (Component Object Model) components show that the proposed FVDM can detect the buffer overflow as well as illegal access vulnerabilities in the components.
Jinfu Chen 0001, Haodi Xie, Saihua Cai, Ye Geng, Yemin Yin, Zikang Zhang
TrustCom1
2022 Malware recognition approach based on self-similarity and an improved clustering algorithm
abstract
Abstract The recognition of malware in network traffic is an important research problem. However, existing solutions addressing this problem rely heavily on the source code and misrecognise vulnerabilities (i.e. incur a high false positive rate (FPR)) in some cases. In this paper, we initially use the K‐means clustering algorithm to extract malware patterns under user to root attacks in network traffic. Since the traditional K‐means algorithm needs to determine the number of clusters in advance and it is easily affected by the initial cluster centres, we propose an improved K‐means clustering algorithm (NIKClustering algorithm) for cluster analysis. Furthermore, we propose the use of self‐similarity and our improved clustering algorithm to recognise buffer overflow vulnerabilities for malware in network traffic. This motivates us to design and implement a recognition approach for buffer overflow vulnerabilities based on self‐similarity and our improved clustering algorithm, called Reliable Self‐Similarity with Improved K‐means Clustering (RSS‐IKClustering). Extensive experiments conducted on two different datasets demonstrate that the RSS‐IKClustering can achieve much fewer false positives than other notable approaches while increasing accuracy. We further apply our RSS‐IKClustering approach on a public dataset (Center for Applied Internet Data Analysis), which also exhibited a high accuracy and low FPR of 96% and 1.5%, respectively.
Jinfu Chen 0001, Chi Zhang 0046, Saihua Cai, Zufa Zhang, Longxia Huang
IET Softw.1
2022 MWFP-outlier: Maximal weighted frequent-pattern-based approach for detecting outliers from uncertain weighted data streams
Saihua Cai, Li Li 0059, Jinfu Chen 0001, Kaiyi Zhao, Ruizhi Sun, Rexford Nii Ayitey Sosu, Longxia Huang
Inf. Sci.3
2022 A software defect prediction method with metric compensation based on feature selection and transfer learning
abstract
Cross-project software defect prediction solves the problem of insufficient training data for traditional defect prediction, and overcomes the challenge of applying models learned from multiple different source projects to target project. At the same time, two new problems emerge: (1) too many irrelevant and redundant features in the model training process will affect the training efficiency and thus decrease the prediction accuracy of the model; (2) the distribution of metric values will vary greatly from project to project due to the development environment and other factors, resulting in lower prediction accuracy when the model achieves cross-project prediction. In the proposed method, the Pearson feature selection method is introduced to address data redundancy, and the metric compensation based transfer learning technique is used to address the problem of large differences in data distribution between the source project and target project. In this paper, we propose a software defect prediction method with metric compensation based on feature selection and transfer learning. The experimental results show that the model constructed with this method achieves better results on area under the receiver operating characteristic curve (AUC) value and F1-measure metric.
Jinfu Chen 0001, Saihua Cai, Jiaping Xu, Haibo Chen 0005
Frontiers Inf. Technol. Electron. Eng.1
2021 Fuzzing Methods Recommendation Based on Feature Vectors
abstract
Fuzzing is a technique that aims to detect vulnerabilities or exceptions through unexpected input and has found tremendous recent interest in both academia and industry. Although these fuzzing methods have great advantages in the field of vulnerability detection, they also have their own disadvantages in the face of different target programs. It is obviously impractical for a fuzzing test method to adapt to all the target programs. Therefore, we study how to select the appropriate fuzzing methods for different target programs. Specifically, we first analyze the program, and then extract the feature vectors of the target program to get the information of the program, such as syntax, context and so on. Next, we build a matching model to match the similarity of target program and the fuzzing algorithm to select the fuzzing algorithm with higher matching degree. Through our matching model, we get a more suitable fuzzing algorithm to improve the detection efficiency, precision, recall, F-measure, and other statistical measures.
Chi Zhang 0046, Jinfu Chen 0001
ASE2
2021 L-KPCA: an efficient feature extraction method for network intrusion detection
abstract
Network intrusion detection identifies malicious activity in the network by analyzing the behavior of network traffic. As an important part of network intrusion detection, feature extraction plays a crucial role in improving the performance of intrusion detection. This research proposes a novel secondary feature extraction method called L-KPCA based on the Liner Discriminant Analysis (LDA) and Kernel Principal Component Analysis (KPCA), to provide efficient features for network intrusion detection. While maintaining the effectiveness of processing nonlinear data in network traffic, the use of LDA effectively compensates for the problem that KPCA only focuses on the analysis of features in terms of variance and ignores the performance of features in terms of mean. Extensive experimental results verify that the use of the proposed, L-KPCA can make the intrusion detection classification model perform better in terms of recognition accuracy and recall.
Jinfu Chen 0001, Shang Yin, Saihua Cai, Lingling Zhao, Shengran Wang
MSN1
2021 MMFC-ART: a Fixed-size-Candidate-set Adaptive Random Testing approach based on the modified Metric-Memory tree
abstract
Adaptive random testing (ART) improves the failure-detection effectiveness of Random testing (RT) by making test cases more evenly distributed in the input domain. The Fixed-size-Candidate-set ART (FSCS-ART) is one of the most classical algorithms, which selects the candidate test case furthest from the previously executed test case as the next test case. However, when the number of executed test cases is large, the computational overhead will be very high. In this paper, we propose an enhanced version of FSCS-ART based on a modified Metric-Memory tree (MM-tree), namely Fixed-size-Candidate-set ART based on the modified MM-tree (MMFC-ART). Simulations and empirical studies are conducted to verify the effectiveness and efficiency of MMFC-ART. The experimental results indicate that MMFC-ART significantly reduces the computational overhead while ensuring comparable or better failure-detection effectiveness than FSCS-ART. Meanwhile, compared with KD-tree-enhanced Fixed-size-Candidate-set ART (KDFC-ART), MMFC-ART has better performance in high dimensions in terms of efficiency. In terms of effectiveness, MMFC-ART has better failure-detection effectiveness in some scenarios. Overall, MMFC-ART is cost-effective compared to FSCS-ART and KDFC-ART.
Jinfu Chen 0001, Yiming Wu 0012, Chengying Mao, Tsong Yueh Chen, Haibo Chen 0005
QRS1
2021 An Efficient Network Intrusion Detection Model Based on Temporal Convolutional Networks
abstract
Network intrusion detection plays an important role in the network security, but the increasingly complex network environment brings a serious challenge to intrusion detection. Although the existing efficient Convolutional Neural Network (CNN)-based network traffic intrusion detection models do not require manual design of the traffic features, but they do not make full use of the structured information of network traffic. In this paper, we propose a novel network intrusion detection model based on Temporal Convolutional Networks (TCN), it extracts the key features in the dataset through exploiting the characteristics of byte sequence in the network traffic packets. Compared with traditional recurrent neural networks (RNN), TCN shows a better performance in sequence modeling tasks and it can process the sequences in parallel for faster training. To solve the problem of poor detection accuracy caused by the “death” of some neurons on ReLU during the training stage, we use the ELU activation function in the TCN instead of ReLU. Finally, we compare our proposed TCN-based intrusion detection model with the state-of-the-art methods on the CTU public dataset, and the experimental results show that the use of TCN can obtain higher performance within less time consumption, in terms of higher average accuracy, higher average recall and higher average F1-measure.
Jinfu Chen 0001, Shang Yin, Saihua Cai, Chi Zhang 0046, Yemin Yin
QRS1
2021 AIdetectorX: A Vulnerability Detector Based on TCN and Self-attention Mechanism
Jinfu Chen 0001, Bo Liu 0048, Saihua Cai, Shengran Wang
SETTA1
2021 Covering Array Constructors: An Experimental Analysis of Their Interaction Coverage and Fault Detection
abstract
Abstract Combinatorial interaction testing (CIT) aims at constructing a covering array (CA) of all value combinations at a specific interaction strength, to detect faults that are caused by the interaction of parameters. CIT has been widely used in different applications, with many algorithms and tools having been proposed to support CA construction. To date, however, there appears to have been no studies comparing different CA constructors when only some of the CA test cases are executed. In this paper, we present an investigation of five popular CA constructors: ACTS, Jenny, PICT, CASA and TCA. We conducted empirical studies examining the five programs, focusing on interaction coverage and fault detection. The experimental results show that when there is no preference or special justification for using other CA constructors, then Jenny is recommended—because it achieves better interaction coverage and fault detection than the other four constructors in many cases. Our results also show that when using ACTS or CASA, their CAs must be prioritized before testing. The main reason for this is that these CAs can result in considerable interaction coverage or fault detection capabilities when executing a large number of test cases; however, they may also produce the lowest rates of fault detection and interaction coverage.
Rubing Huang, Haibo Chen 0005, Yunan Zhou, Tsong Yueh Chen, Dave Towey, Man Fai Lau, Sebastian Ng, Robert G. Merkel, Jinfu Chen 0001
Comput. J.9
2021 An efficient anomaly detection method for uncertain data based on minimal rare patterns with the consideration of anti-monotonic constraints
Saihua Cai, Jinfu Chen 0001, Haibo Chen 0005, Chi Zhang 0046, Qian Li 0042, Rexford Nii Ayitey Sosu, Shang Yin
Inf. Sci.2
2021 An efficient outlier detection method for data streams based on closed frequent patterns by considering anti-monotonic constraints
Saihua Cai, Rubing Huang, Jinfu Chen 0001, Chi Zhang 0046, Bo Liu 0048, Shang Yin, Ye Geng
Inf. Sci.3
2021 A Detection Approach for Vulnerability Exploiter Based on the Features of the Exploiter
abstract
With the wide application of software system, software vulnerability has become a major risk in computer security. The on-time detection and proper repair for possible software vulnerabilities are of great importance in maintaining system security and decreasing system crashes. The Control Flow Integrity (CFI) can be used to detect the exploit by some researchers. In this paper, we propose an improved Control Flow Graph with Jump (JCFG) based on CFI and develop a novel Vulnerability Exploit Detection Method based on JCFG (JCFG-VEDM). The detection method of the exploit program is realized based on the analysis results of the exploit program. Then the JCFG is addressed through combining the features of the exploit program and the jump instruction. Finally, we implement JCFG-VEDM and conduct the experiments to verify the effectiveness of the proposed method. The experimental results show that the proposed detection method (JCFG-VEDM) is feasible and effective.
Jinchang Hu, Jinfu Chen 0001, Sher Ali, Bo Liu 0048, Chi Zhang 0046
Secur. Commun. Networks2
2021 An Approach Based on the Improved SVM Algorithm for Identifying Malware in Network Traffic
abstract
Due to the growth and popularity of the internet, cyber security remains, and will continue, to be an important issue. There are many network traffic classification methods or malware identification approaches that have been proposed to solve this problem. However, the existing methods are not well suited to help security experts effectively solve this challenge due to their low accuracy and high false positive rate. To this end, we employ a machine learning-based classification approach to identify malware. The approach extracts features from network traffic and reduces the dimensionality of the features, which can effectively improve the accuracy of identification. Furthermore, we propose an improved SVM algorithm for classifying the network traffic dubbed Optimized Facile Support Vector Machine (OFSVM). The OFSVM algorithm solves the problem that the original SVM algorithm is not satisfactory for classification from two aspects, i.e., parameter optimization and kernel function selection. Therefore, in this paper, we present an approach for identifying malware in network traffic, called Network Traffic Malware Identification (NTMI). To evaluate the effectiveness of the NTMI approach proposed in this paper, we collect four real network traffic datasets and use a publicly available dataset CAIDA for our experiments. Evaluation results suggest that the NTMI approach can lead to higher accuracy while achieving a lower false positive rate compared with other identification methods. On average, the NTMI approach achieves an accuracy of 92.5% and a false positive rate of 5.527%.
Bo Liu 0048, Jinfu Chen 0001, Songling Qin, Zufa Zhang, Yisong Liu, Lingling Zhao
Secur. Commun. Networks2
2020 Minimal Rare-Pattern-Based Outlier Detection Method for Data Streams by Considering Anti-monotonic Constraints
Saihua Cai, Jinfu Chen 0001, Bo Liu 0048
ISC2
2020 An Approach to Determine the Optimal k-Value of K-means Clustering in Adaptive Random Testing
abstract
Adaptive Random Testing (ART) aims at improving detection effectiveness by evenly distributing test cases over the whole input domain. Many ART algorithms introducing clustering techniques (such as k-means Clustering) have been proposed to achieve an even spread of test cases. Though previous studies have demonstrated that ART with k-means clustering could achieve a good enhancement in testing effectiveness, k-means clustering is limited by the value of k, which will have a great impact on the test effectiveness. To improve the testing effectiveness of these techniques for object-oriented software, in this paper, we propose an approach named Determination Method of Optimal k-value based on the Experimental Process (DMOVk-EP) to determine the optimal k-value of k-means clustering and make the ART algorithms using k-means clustering technique achieve the best fault detection capability. The proposed method consists of two parts, one is a solution model for k based on the experimental process, and the other is an optimal k-value algorithm based on the presented model. We integrate this method with k-means clustering in ART and apply it to a set of open-source programs, with the experimental results showing that our approach obtains much more appropriate k, and also achieves much better testing effectiveness than other related methods.
Jinfu Chen 0001, Lingling Zhao, Minmin Zhou, Yisong Liu, Songling Qin
QRS1
2020 An Automatic Vulnerability Scanner for Web Applications
abstract
With the progressive development of web applications and the urgent requirement of web security, vulnerability scanner has been particularly emphasized, which is regarded as a fundamental component for web security assurance. Various scanners are developed with the intention of that discovering the possible vulnerabilities in advance to avoid malicious attacks. However, most of them only focus on the vulnerability detection with single target, which fail in satisfying the efficiency demand of users. In this paper, an effective web vulnerability scanner that integrates the information collection with the vulnerability detection is proposed to verify whether the target web application is vulnerable or not. The experimental results show that, by guiding the detection process with the useful collected information, our tool achieves great web vulnerability detection capability with a large scanning scope.
Haibo Chen 0005, Junzuo Chen, Jinfu Chen 0001, Shang Yin, Yiming Wu 0012, Jiaping Xu
TrustCom3
2020 An Automatic Vulnerability Classification System for IoT Softwares
abstract
Internet of Things(IoT) have been widely implemented in diverse domains of real-life, and become one of the most popular applications of the internet. Nevertheless, the development of IoT has suffered from its security issues so far. Various IoT vulnerabilities bring serious risks to the privacy and the property security of users. To study the security vulnerabilities in depth, the classification for IoT vulnerabilities becomes a basic requirement. However, manual classification relies on human experience and is very laborious. In this paper, an IoT vulnerabilities classification system based on Support Vector Machines(SVM) and Particle Swarm optimization(PSO) is developed to identify and classify IoT vulnerabilities automatically. The experimental results prove that our system presents great potential of vulnerability classification.
Haibo Chen 0005, Dalin Zhang 0004, Jinfu Chen 0001, Dengzhou Shi, Zian Zhao
TrustCom3
2020 iTES: Integrated Testing and Evaluation System for Software Vulnerability Detection Methods
abstract
To find software vulnerabilities using software vulnerability detection technology is an important way to ensure the system security. Existing software vulnerability detection methods have some limitations as they can only play a certain role in some specific situations. To accurately analyze and evaluate the existing vulnerability detection methods, an integrated testing and evaluation system (iTES) is designed and implemented in this paper. The main functions of the iTES are:(1) Vulnerability cases with source codes covering common vulnerability types are collected automatically to form a vulnerability cases library; (2) Fourteen methods including static and dynamic vulnerability detection are evaluated in iTES, involving the Windows and Linux platforms; (3) Furthermore, a set of evaluation metrics is designed, including accuracy, false positive rate, utilization efficiency, time cost and resource cost. The final evaluation and test results of iTES have a good guiding significance for the selection of appropriate software vulnerability detection methods or tools according to the actual situation in practice.
Chi Zhang 0046, Jinfu Chen 0001, Saihua Cai, Bo Liu 0048, Yiming Wu 0012, Ye Geng
TrustCom2
2020 Adaptive random testing based on flexible partitioning
abstract
Adaptive random testing (ART) achieves better failure‐detection effectiveness than random testing due to its even spreading of test cases. ART by random partitioning (RP‐ART) is a lightweight method, but its advantage over random testing is relatively low. Although iterative partition testing (IPT) method has good performance for detecting failures in a block pattern, it loses randomness during the test case generation. To overcome the shortcomings of the above two algorithms, a new algorithm named ART by flexible partitioning (FP‐ART) is proposed. In the FP‐ART, a set of random candidates is used to select an appropriate test case by considering their boundary distance. Accordingly, the corresponding sub‐domain is also partitioned by the new test case. Based on this kind of flexible partitioning, the randomness of test case selection can be guaranteed and the spatial distribution of test cases is even more diverse. According to the results in simulation and empirical experiments, FP‐ART demonstrates better failure‐detection effectiveness than RP‐ART and is more suitable to detect the failures in strip patterns than the IPT method. Meanwhile, its failure‐detection ability is much stronger than that of fixed‐size‐candidate‐set ART in the cases of a relatively high failure rate.
Chengying Mao, Xuzheng Zhan, Jinfu Chen 0001, Jifu Chen 0001, Rubing Huang
IET Softw.3
2020 A Proactive Approach to Test Case Selection - An Efficient Implementation of Adaptive Random Testing
abstract
Fixed Sized Candidate Set (FSCS) is the first of a series of methods proposed to enhance the effectiveness of random testing (RT) referred to as Adaptive Random Testing methods or ARTs. Since its inception, test case generation overheads have been a major drawback to the success of ART. In FSCS, the bulk of this cost is embedded in distance computations between a set of randomly generated candidate test cases and previously executed but unsuccessful test cases. Consequently, FSCS is caught in a logical trap of probing the distances between every candidate and all executed test cases before the best candidate is determined. Using data mining, however, we discovered that about 50% of all valid test cases are encountered much earlier in the distance computations process but without any benefit of a hindsight, FSCS is unable to validate them; a wild goose chase. This paper then uses this information to propose a new strategy that predictively and proactively selects valid candidates anywhere during the distance computation process without vetting every candidate. Theoretical analysis, simulations and experimental studies conducted led to a similar conclusion: 25% of the distance computations are wasteful and can be discarded without any repercussion on effectiveness.
Michael Omari, Jinfu Chen 0001, Robert French-Baidoo, Yunting Sun
Int. J. Softw. Eng. Knowl. Eng.2
2020 An automatic software vulnerability classification framework using term frequency-inverse gravity moment and feature selection
Jinfu Chen 0001, Patrick Kwaku Kudjo, Solomon Mensah, Selasie Brown Aformaley, George Akorfu
J. Syst. Softw.1
2020 Regression test case prioritization by code combinations coverage
Rubing Huang, Quanjun Zhang, Dave Towey, Weifeng Sun 0004, Jinfu Chen 0001
J. Syst. Softw.5
2020 An automated framework for evaluating open-source web scanner vulnerability severity
Richard Amankwah, Jinfu Chen 0001, Patrick Kwaku Kudjo, Beatrice Korkor Agyemang, Alfred Adutwum Amponsah
Serv. Oriented Comput. Appl.2
2020 An empirical comparison of commercial and open-source web vulnerability scanners
abstract
Summary Web vulnerability scanners (WVSs) are tools that can detect security vulnerabilities in web services. Although both commercial and open‐source WVSs exist, their vulnerability detection capability and performance vary. In this article, we report on a comparative study to determine the vulnerability detection capabilities of eight WVSs (both open and commercial) using two vulnerable web applications: WebGoat and Damn vulnerable web application. The eight WVSs studied were: Acunetix; HP WebInspect; IBM AppScan; OWASP ZAP; Skipfish; Arachni; Vega; and Iron WASP. The performance was evaluated using multiple evaluation metrics: precision; recall; Youden index; OWASP web benchmark evaluation; and the web application security scanner evaluation criteria. The experimental results show that, while the commercial scanners are effective in detecting security vulnerabilities, some open‐source scanners (such as ZAP and Skipfish) can also be effective. In summary, this study recommends improving the vulnerability detection capabilities of both the open‐source and commercial scanners to enhance code coverage and the detection rate, and to reduce the number of false‐positives.
Richard Amankwah, Jinfu Chen 0001, Patrick Kwaku Kudjo, Dave Towey
Softw. Pract. Exp.2
2020 The effect of Bellwether analysis on software vulnerability severity prediction models
Patrick Kwaku Kudjo, Jinfu Chen 0001, Solomon Mensah, Richard Amankwah, Christopher Kudjo
Softw. Qual. J.2
2020 Abstract Test Case Prioritization Using Repeated Small-Strength Level-Combination Coverage
abstract
Abstract test cases (ATCs) have been widely used in practice, including in combinatorial testing and in software product line testing. When constructing a set of ATCs, due to limited testing resources in practice (e.g., in regression testing), test case prioritization (TCP) has been proposed to improve the testing quality, aiming at ordering test cases to increase the speed with which faults are detected. One intuitive and extensively studied TCP technique for ATCs is λ-wise Level-combination Coverage based Prioritization (λLCP), a static, black-box prioritization technique that only uses the ATC information to guide the prioritization process. A challenge facing λLCP, however, is the necessity for the selection of the fixed prioritization strength λ before testing-testers need to choose an appropriate λ value before testing begins. Choosing higher λ values may improve the testing effectiveness of λLCP (e.g., by finding faults faster), but may reduce the testing efficiency (by incurring additional prioritization costs). Conversely, choosing lower λ values may improve the efficiency, but may also reduce the effectiveness. In this paper, we propose a new family of λLCP techniques, Repeated Small-strength Level-combination Coverage-based Prioritization (RSLCP), that repeatedly achieves the full combination coverage at lower strengths. RSLCP maintains λLCP's advantages of being static and black box, but avoids the challenge of prioritization strength selection. We have performed an empirical study involving five different versions of each of five C programs. Compared with λLCP, and Incremental-strength LCP (ILCP), our results show that RSLCP could provide a good tradeoff between testing effectiveness and efficiency. Our results also show that RSLCP is more effective and efficient than two popular techniques of Similarity-based Prioritization (SP). In addition, the results of empirical studies also show that RSLCP can remain robust over multiple system releases.
Rubing Huang, Weifeng Sun 0004, Tsong Yueh Chen, Dave Towey, Jinfu Chen 0001, Weiwen Zong, Yunan Zhou
IEEE Trans. Reliab.5
2019 A cost-effective strategy for software vulnerability prediction based on bellwether analysis
abstract
Vulnerability Prediction Models (VPMs) aims to identify vulnerable and non-vulnerable components in large software systems. Consequently, VPMs presents three major drawbacks (i) finding an effective method to identify a representative set of features from which to construct an effective model. (ii) the way the features are utilized in the machine learning setup (iii) making an implicit assumption that parameter optimization would not change the outcome of VPMs. To address these limitations, we investigate the significant effect of the Bellwether analysis on VPMs. Specifically, we first develop a Bellwether algorithm to identify and select an exemplary subset of data to be considered as the Bellwether to yield improved prediction accuracy against the growing portfolio benchmark. Next, we build a machine learning approach with different parameter settings to show the improvement of performance of VPMs. The prediction results of the suggested models were assessed in terms of precision, recall, F-measure, and other statistical measures. The preliminary result shows the Bellwether approach outperforms the benchmark technique across the applications studied with F-measure values ranging from 51.1%-98.5%.
Patrick Kwaku Kudjo, Jinfu Chen 0001
ISSTA2
2019 Improving the Accuracy of Vulnerability Report Classification Using Term Frequency-Inverse Gravity Moment
abstract
Software vulnerability analysis is one of the critical issues in the software industry, and vulnerability classification plays a major role in this analysis. A typical vulnerability classification model usually involves a stage of term selection, in which the relevant terms are identified via feature selection. It also involves a stage of term weighting, in which document weights for the selected terms are computed, and a stage for classifier learning. Generally, the term frequency-inverse document frequency (TF-IDF) is the most widely used term-weighting method. However, empirical evidence shows that the TF-IDF is plagued with issues pertaining to its effectiveness. This paper introduces a new approach for vulnerability classification, which is based on term frequency and inverse gravity moment (TF-IGM). The proposed method is validated by empirical experiments using three machine learning algorithms on ten publicly available vulnerability datasets. The result shows that TF-IGM outperforms the benchmark method across the applications studied.
Patrick Kwaku Kudjo, Jinfu Chen 0001, Minmin Zhou, Solomon Mensah, Rubing Huang
QRS2
2019 Random Border Mirror Transform: A Diversity Based Approach to an Effective and Efficient Mirror Adaptive Random Testing
abstract
Mirror Adaptive random testing (MART) is an overhead reduction strategy for adaptive random testing methods. Theoretically speaking, MART's advantage over ordinary ARTs is determined by the mirroring scheme selected. Incidentally, an inherent problem with MART relates to the difficulty in the choice of a scheme for any testing task. This is because a higher scheme (larger mirror domains) does not necessarily guarantee efficient utilization of testing resources due to lack of diversity of mirror generated test cases. The culprit has been identified as the mapping functions used as substitutes to complex ART methods. In this paper, we present a new method for generating diversified mirror test cases by randomly displacing the mirror partitions upon which the mapping functions of MART operates. The result of simulations and experiments conducted shows remarkable improvement over MART's effectiveness and efficiency across MART schemes, especially where program failures are unrelated to one or more input parameters.
Michael Omari, Jinfu Chen 0001, Patrick Kwaku Kudjo, Hilary Ackah-Arthur, Rubing Huang
QRS2
2019 Toward a K-means clustering approach to adaptive random testing for object-oriented software
Jinfu Chen 0001, Minmin Zhou, T. H. Tse, Tsong Yueh Chen, Yuchi Guo, Rubing Huang, Chengying Mao
Sci. China Inf. Sci.1
2019 Prioritising abstract test cases: an empirical study
abstract
Test‐case prioritisation (TCP) attempts to schedule the order of test‐case execution such that faults can be detected as quickly as possible. TCP has been widely applied in many testing scenarios such as regression testing and fault localisation. Abstract test cases (ATCs) are derived from models of the system under test and have been applied to many testing environments such as model‐based testing and combinatorial interaction testing. Although various empirical and analytical comparisons for some ATC prioritisation (ATCP) techniques have been conducted, to the best of the authors’ knowledge, no comparative study focusing on the most current techniques has yet been reported. In this study, they investigated 18 ATCP techniques, categorised into four classes. They conducted a comprehensive empirical study to compare 16 of the 18 ATCP techniques in terms of their testing effectiveness and efficiency. They found that different ATCP techniques could be cost‐effective in different testing scenarios, allowing us to present recommendations and guidelines for which techniques to use under what conditions.
Rubing Huang, Weiwen Zong, Tsong Yueh Chen, Dave Towey, Yunan Zhou, Jinfu Chen 0001
IET Softw.6
2019 A Modified Similarity Metric for Unit Testing of Object-Oriented Software Based on Adaptive Random Testing
abstract
Finding an effective method for testing object-oriented software (OOS) has proven elusive in the software community due to the rapid development of object-oriented programming (OOP) technology. Although significant progress has been made by previous studies, challenges still exist in relation to the object distance measurement of OOS using Adaptive Random Testing (ART). This is partly due to the unique features of OOS such as encapsulation, inheritance and polymorphism. In a previous work, we proposed a new similarity metric called the Object and Method Invocation Sequence Similarity (OMISS) metric to facilitate multi-class level testing using ART. In this paper, we broaden the set of models in the metric (OMISS) by considering the method parameter and adding the weight in the metric to develop a new distance metric to improve unit testing of OOS. We used the new distance metric to calculate the distance between the set of objects and the distance between the method sequences of the test cases. Additionally, we integrate the new metric in unit testing with ART and applied it to six open source subject programs. The experimental result shows that the proposed method with method parameter considered in this study is better than previous methods without the method parameter in the case of the single method. Our finding further shows that the proposed unit testing approach is a promising direction for assisting software engineers who seek to improve the failure-detection effectiveness of OOS testing.
Jinfu Chen 0001, Patrick Kwaku Kudjo, Zufa Zhang, Chenfei Su, Yuchi Guo, Rubing Huang, Heping Song
Int. J. Softw. Eng. Knowl. Eng.1
2019 One-Domain-One-Input: Adaptive Random Testing by Orthogonal Recursive Bisection With Restriction
abstract
One goal of software testing may be the identification or generation of a series of test cases that can detect a fault with as few test executions as possible. Motivated by insights from research into failure-causing regions of input domains, the even-spreading (even distribution) of tests across the input domain has been identified as a useful heuristic to more quickly find failures. This finding has encouraged a shift in focus from traditional random testing (RT) to its enhancement, adaptive random testing (ART), which retains the randomness of test input selection, but also attempts to maintain a more evenly distributed spread of test inputs across the input domain. Given that there are different ways to achieve the even distribution, several different ART methods and approaches have been proposed. This paper presents a new ART method, called ART by orthogonal recursive bisection (ART-ORB), which explores the advantages of repeated geometric bisection of the input domain, combined with restriction regions, to evenly spread test inputs. Experimental results show a better performance in terms of fewer test executions than RT to find failures. Compared with other ART methods, ART-ORB has comparable performance (in terms of required test executions), but incurs lower test input selection overheads, especially in higher dimensional input space. It is recommended that ART-ORB can be used in testing situations involving expensive test input execution.
Hilary Ackah-Arthur, Jinfu Chen 0001, Dave Towey, Michael Omari, Jiaxiang Xi, Rubing Huang
IEEE Trans. Reliab.2
2018 On the Selection of Strength for Fixed-Strength Interaction Coverage Based Prioritization
abstract
Abstract test cases are derived by modeling the system under test, and have been widely applied in practice, such as for software product line testing and combinatorial testing. Abstract test case prioritization (ATCP) is used to prioritize abstract test cases and aims at achieving higher rates of fault detection. Many ATCP algorithms have been proposed, using different prioritization criteria and information. One ATCP approach makes use of fixed-strength level-combinations information covered by abstract test cases, and is called fixed-strength interaction coverage based prioritization (FICBP). Before using FICBP, the prioritization strength λ needs to be decided. Previous studies have generally focused on λ values ranging between 1 and 6. However, no study has investigated the appropriateness of such a range, nor how to assign the prioritization strength for FICBP. To answer these questions, this paper reports on an empirical study involving four real-life programs (each of which with six versions). The experimental results indicate that λ should be set approximately equal to a value corresponding to half of the number of parameters, when testing resources are sufficient. Our results also show that when testing resources are limited or insufficient, either small or large λ values are suggested for FICBP.
Rubing Huang, Weiwen Zong, Tsong Yueh Chen, Dave Towey, Jinfu Chen 0001, Yunan Zhou, Weifeng Sun 0004
COMPSAC (1)5
2018 A cost-effective adaptive random testing approach by dynamic restriction
abstract
A key objective of software testing is to find program errors that cause failure in software, at less cost. One basic testing technique is random testing (RT), but many researchers have criticised its failure‐detection effectiveness. Several researchers have proposed that an enhancement of the failure‐detection effectiveness of RT is achieved if test cases are evenly spread within the input domain. Adaptive RT (ART) describes a family of algorithms that employ various strategies to evenly and randomly spread test cases. Fixed sized candidate set ART (FSCS‐ART) is an ART algorithm that has gained many research studies far and wide; however, the high distance computations make its algorithm computationally expensive. The authors propose a new ART method that restricts distance computations to only test cases inside an exclusion zone. The experimental results show that the new ART method not only improves RT but also provides failure‐detection effectiveness similar to FSCS‐ART, while significantly minimising computation overhead.
Hilary Ackah-Arthur, Jinfu Chen 0001, Jiaxiang Xi, Michael Omari, Heping Song, Rubing Huang
IET Softw.2
2018 Test case prioritization for object-oriented software: An adaptive random sequence approach based on clustering
Jinfu Chen 0001, Lili Zhu, Tsong Yueh Chen, Dave Towey, Fei-Ching Kuo, Rubing Huang, Yuchi Guo
J. Syst. Softw.1
2017 An Empirical Comparison of Similarity Measures for Abstract Test Case Prioritization
abstract
Test case prioritization (TCP) attempts to order test cases such that those which are more important, according to some criterion or measurement, are executed earlier. TCP has been applied in many testing situations, including, for example, regression testing. An abstract test case (also called a model input) is an important type of test case, and has been widely used in practice, such as in configurable systems and software product lines. Similarity-based test case prioritization (STCP) has been proven to be cost-effective for abstract test cases (ATCs), but because there are many similarity measures which could be used to evaluate ATCs and to support STCP, we face the following question: How can we choose the similarity measure(s) for prioritizing ATCs that will deliver the most effective results? To address this, we studied fourteen measures and two popular STCP algorithms - local STCP (LSTCP), and global STCP (GSTCP). We also conducted an empirical study of five realworld programs, and investigated the efficacy of each similarity measure, according to the interaction coverage rate and fault detection rate. The results of these studies show that GSTCP outperforms LSTCP - in 61% to 84% of the cases, in terms of interaction coverage rates; and in 76% to 78% of the cases with respect to fault detection rates. Our studies also show that Overlap, the simplest similarity measure examined in this study, could obtain the overall best performance for LSTCP; and that Goodall3 has the best performance for GSTCP.
Rubing Huang, Yunan Zhou, Weiwen Zong, Dave Towey, Jinfu Chen 0001
COMPSAC (1)5
2017 Detecting Implicit Security Exceptions Using an Improved Variable-Length Sequential Pattern Mining Method
abstract
The process of component security testing can produce massive amounts of monitor logs. Current approaches to detect implicit security exceptions (those which cannot be identified by visual inspection alone) compare correct execution sequences with fixed patterns mined from the execution of sequential patterns in the monitor logs. However, this is not efficient and is not suitable for mining large monitor logs. To enable effective mining of implicit security exceptions from large monitor logs, this paper proposes a method based on improved variable-length sequential pattern mining. The proposed method first mines the variable-length sequential patterns from correct execution sequences and from actual execution sequences, thus reducing the number of patterns. The sequential patterns are then detected using the Sunday string-searching algorithm. We conducted an experimental study based on this method, the results of which show that the proposed method can efficiently detect the implicit security exceptions of components.
Jinfu Chen 0001, Saihua Cai, Dave Towey, Lili Zhu, Rubing Huang, Hilary Ackah-Arthur, Michael Omari
Int. J. Softw. Eng. Knowl. Eng.1
2017 A Similarity Metric for the Inputs of OO Programs and Its Application in Adaptive Random Testing
abstract
Random testing (RT) has been identified as one of the most popular testing techniques, due to its simplicity and ease of automation. Adaptive random testing (ART) has been proposed as an enhancement to RT, improving its fault-detection effectiveness by evenly spreading random test inputs across the input domain. To achieve the even spreading, ART makes use of distance measurements between consecutive inputs. However, due to the nature of object-oriented software (OOS), its distance measurement can be particularly challenging: Each input may involve multiple classes, and interaction of objects through method invocations. Two previous studies have reported on how to test OOS at a single-class level using ART. In this study, we propose a new similarity metric to enable multiclass level testing using ART. When generating test inputs (for multiple classes, a series of objects, and a sequence of method invocations), we use the similarity metric to calculate the distance between two series of objects, and between two sequences of method invocations. We integrate this metric with ART and apply it to a set of open-source OO programs, with the empirical results showing that our approach outperforms other RT and ART approaches in OOS testing.
Jinfu Chen 0001, Fei-Ching Kuo, Tsong Yueh Chen, Dave Towey, Chenfei Su, Rubing Huang
IEEE Trans. Reliab.1
2016 Prioritizing Interaction Test Suites Using Repeated Base Choice Coverage
abstract
Combinatorial interaction testing is a well-studied testing strategy that aims at constructing an effective interaction test suite (ITS) of a specific generation strength to identify interaction faults caused by the interactions among factors. Due to limited testing resources in practice, for example in combinatorial interaction regression testing, interaction test suite prioritization (ITSP) has been proposed to improve the efficiency of testing. An intuitive ITSP strategy that has been widely used in practice is fixed-strength interaction coverage based prioritization (FICBP). FICBP makes use of a property of the ITS: interaction coverage at a fixed prioritization strength. However, a challenge facing FICBP is that, when the ITS is large, the prioritization cost can be very high. In this paper, we propose a new FICBP method that, by repeatedly using base choice coverage (i.e., one-wise coverage) during the prioritization process, improves testing efficiency while maintaining testing effectiveness. The empirical studies show that our method has fault detection capability comparable to current FICBP methods, but obtains more stable results in many cases. Additionally, our method requires considerably less prioritization time than other FICBP methods at different prioritization strengths.
Rubing Huang, Weiwen Zong, Jinfu Chen 0001, Dave Towey, Yunan Zhou, Deng Chen
COMPSAC3
2016 An approach of security testing for third-party component based on state mutation
abstract
ABSTRACT It is essential to study an effective approach of security testing for third‐party component. In this paper, to effectively trigger implicit vulnerabilities of third‐party components, an approach of security testing for third‐party component is proposed based on state mutation. To start with, executable method sequences of components are transformed into extended finite state machine. Then, according to characteristics of condition conflict and behavior conflict, two test case generation algorithms are addressed, that is, Operations Conflict Sequences Generation Algorithm and Conditions Conflict Sequences Generation Algorithm, which are designed to generate inaccessible sequences of behavior and condition conflicts. These conflict sequences are run. Furthermore, the security detecting algorithms are addressed to detect implicit vulnerabilities of third‐party components, and then, testing report of component security is obtained. In the end, some experiments are conducted on the basis of the proposed approach, and the experimental results show that the proposed approach can effectively detect security exceptions of third‐party components. Copyright © 2015 John Wiley & Sons, Ltd.
Jinfu Chen 0001, Jiamei Chen, Rubing Huang, Yuchi Guo
Secur. Commun. Networks1
2015 Search-based QoS ranking prediction for web services in cloud environments
Chengying Mao, Jifu Chen 0001, Dave Towey, Jinfu Chen 0001, Xiaoyuan Xie
Future Gener. Comput. Syst.4
2015 Enhancing mirror adaptive random testing through dynamic partitioning
Rubing Huang, Huai Liu, Jinfu Chen 0001
Inf. Softw. Technol.4
2015 Aggregate-strength interaction test suite prioritization
Rubing Huang, Jinfu Chen 0001, Dave Towey, Alvin Chan Toong Shoon, Yansheng Lu
J. Syst. Softw.2
2014 How to Do Tie-breaking in Prioritization of Interaction Test Suites?
Rubing Huang, Jinfu Chen 0001, Rongcun Wang, Deng Chen
SEKE2
2014 A Web services vulnerability testing approach based on combinatorial mutation and SOAP message mutation
Jinfu Chen 0001, Chengying Mao, Dave Towey
Serv. Oriented Comput. Appl.1
2013 Describing Component Behavior Using Improved Chemical Abstract Machine
abstract
This paper proposes an improved chemical abstract machine to accurately describe the behavior characteristics of components based on chemical computation model. The chemical abstract machine is analyzed from the perspective of software field, thereupon the formal description of chemical abstract machine is given based on γccomputation model. Firstly, we analyze the dynamic characteristics of components and chemical computation model. Secondly, the definition of chemical abstract machine and relevant rules are extended to more accurately describe the dynamic behavior of some components. Then, the γccomputation model of the component is given for accurately describing the component behavior. Finally, an actual case of component is described by using the improved component chemical abstract machine model. The case shows that the improved component chemical abstract machine model can provide a good theoretical foundation for generating effective test cases in component testing.
Jinfu Chen 0001, Rubing Huang
COMPSAC1
2013 Prioritizing Variable-Strength Covering Array
abstract
Combinatorial interaction testing is a well-studied testing strategy, and has been widely applied in practice. Combinatorial interaction test suite, such as fixed-strength and variable-strength interaction test suite, is widely used for combinatorial interaction testing. Due to constrained testing resources in some applications, for example in combinatorial interaction regression testing, prioritization of combinatorial interaction test suite has been proposed to improve the efficiency of testing. However, nearly all prioritization techniques may only support fixed-strength interaction test suite rather than variable-strength interaction test suite. In this paper, we propose two heuristic methods in order to prioritize variable-strength interaction test suite by taking advantage of its special characteristics. The experimental results show that our methods are more effective for variable-strength interaction test suite by comparing with the technique of prioritizing combinatorial interaction test suites according to test case generation order, the random test prioritization technique, and the fixed-strength interaction test suite prioritization technique. Besides, our methods have additional advantages compared with the prioritization techniques for fixed-strength interaction test suite.
Rubing Huang, Jinfu Chen 0001, Tao Zhang 0001, Rongcun Wang, Yansheng Lu
COMPSAC2
2013 Prioritization of Combinatorial Test Cases by Incremental Interaction Coverage
abstract
Combinatorial interaction testing is a well-recognized testing method, and has been widely applied in practice, often with the assumption that all test cases in a combinatorial test suite have the same fault detection capability. However, when testing resources are limited, an alternative assumption may be that some test cases are more likely to reveal failure, thus making the order of executing the test cases critical. To improve testing cost-effectiveness, prioritization of combinatorial test cases is employed. The most popular approach is based on interaction coverage, which prioritizes combinatorial test cases by repeatedly choosing an unexecuted test case that covers the largest number of uncovered parameter value combinations of a given strength (level of interaction among parameters). However, this approach suffers from some drawbacks. Based on previous observations that the majority of faults in practical systems can usually be triggered with parameter interactions of small strengths, we propose a new strategy of prioritizing combinatorial test cases by incrementally adjusting the strength values. Experimental results show that our method performs better than the random prioritization technique and the technique of prioritizing combinatorial test suites according to test case generation order, and has better performance than the interaction-coverage-based test prioritization technique in most cases.
Rubing Huang, Dave Towey, Tsong Yueh Chen, Yansheng Lu, Jinfu Chen 0001
Int. J. Softw. Eng. Knowl. Eng.6
2012 Component Security Testing Approach Based on Extended Chemical Abstract Machine
abstract
Unreliable component security hinders the development of component technology. Component security testing is rarely researched with comprehensive focus; several approaches or technologies for detecting vulnerabilities in component security have been proposed, but most are infeasible. A testing approach for component security, which is based on the chemical abstract machine, is proposed for detecting explicit and implicit component security vulnerabilities. We develop an extended chemical abstract machine model, called eCHAM, and generate a state transfer tree and testing sequence for components based on the proposed model. The model can help test the explicit security exceptions of components according to the testing approach of interface fault injection. Condition and state mutation algorithms for identifying implicit security exceptions are also proposed. Vulnerability testing reports are obtained according to the test results. Experiments were conducted in an integration testing platform to verify the applicability of the proposed approach. Results show that the approach is effective and practicable. The proposed approach can detect explicit and implicit security exceptions of components.
Jinfu Chen 0001, Yansheng Lu
Int. J. Softw. Eng. Knowl. Eng.1