Yue Yan 0001

dblp:148/1648-1 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0001-7657-4789ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 CAST: Contrastive Analysis of Spatial and Temporal Features for QIM-Based VoIP Steganalysis
abstract
QIM(quantization index modulation)-based VoIP steganography is an information-hiding technology that malicious users could misuse to engage in illegal activities. Its countermeasure, commonly known as the QIM-based VoIP steganalysis, has been one of the research hotspots over the past decades. VoIP speech data is sequential, so most previous studies have focused on the temporal features extracted from VoIP encoding codewords to improve detection accuracy. As a result, spatial features are often ignored or fully investigated. In practice, VoIP speech data has unique spatial characteristics. Spatial features could capture the properties of nearby codewords and frames. Inspired by the success of CLIP (contrastive language–image pre-training), we propose a novel model that could efficiently incorporate spatial and temporal features named CAST (contrastive analysis of spatial and temporal features). CLIP introduces the concept of contrastive language–image learning, which has demonstrated exceptional efficiency and effectiveness in aligning textual and visual representations. However, it is not directly applicable to the field of QIM-based VoIP steganalysis. In CAST, we introduce a way to align spatial and temporal features through contrastive analysis. Experimental results demonstrate that this alignment is resource-efficient and could enhance detection accuracy. Meanwhile, CAST outperforms other state-of-the-art models in most scenarios.
Cheng Zhang 0042, Yue Yan 0001, Shujuan Jiang
IEEE Signal Process. Lett.2
2025 Efficient Detection of QIM-Based VoIP Steganography Using Adjacent Frame Integration and Multi-Codeword Priority Attention
abstract
With the growing volume of VoIP traffic, many steganography algorithms exploit VoIP speech as a carrier, posing a threat to cybersecurity. Among them, quantization index modulation (QIM)-based VoIP steganography has demonstrated excellent stealth, making detection difficult. In recent years, more studies have focused on developing feasible QIM-based VoIP steganalysis methods for detecting QIM-based VoIP steganography. Previous studies have mostly focused on improving detection performance while neglecting efficiency, resulting in insufficient research on lightweight models. In online detection scenarios, detection efficiency is crucial. On the one hand, the long inference time of large models can delay warnings. On the other hand, the high computational requirements of these models make them difficult to deploy on remote devices, which reduces their practical value. In this letter, we propose a simple yet efficient model named EQVS (efficient QIM-based VoIP steganalysis network) for detecting QIM-based VoIP steganography. In EQVS, the fold and unfold operations are redesigned based on the characteristics of VoIP speech samples and the requirements of the QIM-based VoIP steganalysis task, to avoid disrupting correlation features. Then, multi-codeword priority attention mechanism, inspired by the multi-query attention and retention mechanisms, redefines the calculation procedure for the query, key, and value matrices, as well as the normalization and softmax operations, to further reduce computational resource consumption in a single attention head. Experimental results demonstrate that EQVS outperforms other state-of-the-art models in both detection performance and efficiency.
Cheng Zhang 0042, Yue Yan 0001, Shujuan Jiang
IEEE Signal Process. Lett.2
2024 Combining Error Guessing and Logical Reasoning for Software Fault Localization via Deep Learning
abstract
Automated fault localization has been extensively studied to improve the effectiveness of software debugging. Existing automated fault localization methods neglect the guidance of the simple and easily available debugging information on fault localization. To bridge manual fault localization with automated fault localization, we propose a fault localization approach combining error guessing and logical reasoning via deep learning. The proposed approach simulates the actual debugging process. Specifically, developers’ debugging experience and context dependencies between methods are mapped into two different types of coverage matrices. The constructed matrices are fed to a convolutional neural network (CNN) to predict whether a method is buggy or not. To validate the effectiveness of the proposed approach, we designed and constructed the empirical study on the widely used Defect4J datasets. With respect to the top-n ([Formula: see text]) metric, our approach outperforms the state-of-the-art DeepFL and other five methods including Ochai, Muse, MULTRIC, TraPT and FLUCSS. Particularly, compared with the above methods, our approach has an improvement of 5–182% for top-1. In terms of MFR and MAR, the proposed approach is slightly lower than the best DeepFL but better than the other five methods. The approach we presented achieving the unification of manual and automatic debugging can aid in the improvement of fault localization accuracy.
Rongcun Wang, Mingmei Fan, Yue Yan 0001, Shujuan Jiang
Int. J. Softw. Eng. Knowl. Eng.3
2024 Improving fault localization via weighted execution graph and graph attention network
abstract
Abstract Software fault localization is commonly recognized as arduous and time consuming. Spectrum‐based fault localization (SBFL) has been widely used due to its lightness. However, the effectiveness of SBFL is limited since it only considers simple statistics on the coverage information, ignoring the tie problem that the spectrum matrixes of some statements are the same. Most existing deep learning‐based fault localization (DLFL) techniques convert the coverage information into a vector, which utilizes the spectrum in a simplified manner and still has limitations in practice. To solve the above problem, we propose an approach via the weighted execution graph and graph attention network (WEGAT). We use a graph structure to represent the coverage information between test cases and program elements. Then, we generate a weighted execution graph by applying the predicate execution sequence. Furthermore, we combine the weighted execution graph with the AST as an integrated graph, which is the input of the GAT for fault localization. We evaluate WEGAT in within‐project and cross‐project prediction scenarios on the Defects4J benchmark. Experimental results show that our approach outperforms traditional SBFL (Ochiai, DStar and Tarantula) and DLFL (TraPT, CNN‐FL, Grace, and AGFL) methods, effectively improving the accuracy of fault localization.
Yue Yan 0001, Shujuan Jiang, Cheng Zhang 0042
J. Softw. Evol. Process.1
2023 A Hybrid Multiple Models Transfer Approach for Cross-Project Software Defect Prediction
abstract
For a new project, it is impossible to get a reliable prediction model because of the lack of sufficient training data. To solve the problem, researchers proposed cross-project defect prediction (CPDP). For CPDP, most researchers focus on how to reduce the distribution difference between training data and test data, and ignore the impact of class imbalance on prediction performance. This paper proposes a hybrid multiple models transfer approach (HMMTA) for cross-project software defect prediction. First, several instances that are most similar to each target project instance are selected from all source projects to form the training data. Second, the same number of instances as that of the defected class are randomly selected from all the non-defect class in each iteration. Next, instances selected from the non-defect classes and all defected class instances are combined to form the training data. Third, the transfer learning method called ETrAdaBoost is used to iteratively construct multiple prediction models. Finally, the prediction models obtained from multiple iterations are integrated by the ensemble learning method to obtain the final prediction model. We evaluate our approach on 53 projects from AEEEM, PROMISE, SOFTLAB and ReLink four defect repositories, and compare it with 10 baseline CPDP approaches. The experimental results show that the prediction performance of our approach significantly outperforms the state-of-the-art CPDP methods. Besides, we also find that our approach has the comparable prediction performance as within-project defect prediction (WPDP) approaches. These experimental results demonstrate the effectiveness of HMMTA approach for CPDP.
Shenggang Zhang, Shujuan Jiang, Yue Yan 0001
Int. J. Softw. Eng. Knowl. Eng.3
2023 A Hierarchical Feature Ensemble Deep Learning Approach for Software Defect Prediction
abstract
Software defect prediction can detect modules that may have defects in advance and optimize resource allocation to improve test efficiency and reduce development costs. Traditional features cannot capture deep semantic and grammatical information, which limits the further development of software defect prediction. Therefore, it has gradually become a trend to use deep learning technology to automatically learn valuable deep features from source code or relevant data. However, most software defect prediction methods based on deep learning extraction features from a single information source or only use a single deep learning model, which leads to the fact that the extracted features are not comprehensive enough to affect the final prediction performance. In view of this, this paper proposes a Hierarchical Feature Ensemble Deep Learning (HFEDL) Approach for software defect prediction. Firstly, the HFEDL approach needs to obtain three types of information sources: abstract syntax tree (AST), class dependency network (CDN) and traditional features. Then, the Convolutional Neural Network (CNN) and the Bidirectional Long Short-Term Memory based on Attention mechanism (BiLSTM+Attention) are used to extract different valuable features from the three information sources and multiple prediction sub-models are constructed. Next, all the extracted features are fused by a filter mechanism to obtain more comprehensive features and construct a fusion prediction sub-model. Finally, all the sub-models are integrated by an ensemble learning method to obtain the final prediction model. We use 11 projects in the PROMISE defect repository and evaluate our approach in both non-effort-aware and effort-aware scenarios. The experimental results show that the prediction performance of our approach is superior to state-of-the-art methods in both scenarios.
Shenggang Zhang, Shujuan Jiang, Yue Yan 0001
Int. J. Softw. Eng. Knowl. Eng.3
2023 A fault localization approach based on fault propagation context
Yue Yan 0001, Shujuan Jiang, Shenggang Zhang, Cheng Zhang 0042
Inf. Softw. Technol.1
2023 An effective fault localization approach based on PageRank and mutation analysis
Yue Yan 0001, Shujuan Jiang, Cheng Zhang 0042
J. Syst. Softw.1
2022 A Fault Localization Approach Based on BiRNN and Multi-Dimensional Features
abstract
Software fault localization is notoriously tedious and time-consuming. Developed rapidly, machine learning techniques have been adopted for fault localization by researchers. Most existing approaches use the test coverage information as feature input to the learning model, ignoring the limited ability of the single-dimensional features. The effectiveness of fault localization is not greatly improved. To overcome the limitation, we propose a fault localization approach based on Bidirectional Recurrent Neural Networks (BiRNNs) and multi-dimensional features. Our approach collects suspiciousness-based, text similarity-based and fault-proneness-based features from the traditional fault localization areas and software metrics. To evaluate our approach, the experiments have been studied on the real-fault benchmark Defects4J and seeded fault program NanoXML. The experimental results show that our approach effectively improves fault localization accuracy.
Yue Yan 0001, Shujuan Jiang, Rongcun Wang, Cheng Zhang 0042, Shengang Zhang
Int. J. Softw. Eng. Knowl. Eng.1
2021 CSFL: Fault Localization on Real Software Bugs Based on the Combination of Context and Spectrum
Yue Yan 0001, Shujuan Jiang, Shenggang Zhang
SETTA1