VLDB 2026 Research / reviewers in the wild / expert
Qingfeng Du
dblp:27/319
· DBLP profile ↗
28ranked-venue papers
6as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FedEPA: Enhancing Personalization and Modality Alignment in Multimodal Federated Learning
Qingfeng Du |
ICIC (10) | 2 |
| 2025 | Real-World Code Vulnerability Detection Framework: From Data Preprocessing to Multi-Feature Fusion DetectionabstractCode vulnerability detection (CVD) is a critical approach to ensuring the security, stability, and reliability of software. When exploited by malicious actors or hackers, code vulnerabilities can lead to a series of severe consequences and cause significant losses. However, the effectiveness of real-world CVD is currently unsatisfactory, with several issues that need to be optimized and resolved. These issues include the poor quality of real-world CVD datasets, the high resource consumption of code intermediate structures, the imbalance between positive and negative samples, and insufficient feature modeling of code vulnerabilities. To address these challenges, we propose a comprehensive and efficient framework for real-world CVD called MARCOVul. It offers a complete process from data preprocessing to final vulnerability detection, optimizing the entire real-world CVD pipeline. Our approach begins with a data derivation technique which seeks to improve the overall quality of the dataset. Next, we propose a code-specific data augmentation method to tackle the issue of sample imbalance in the dataset. We then propose a code intermediate structure simplification method to reduce computational complexity and resource consumption while fully leveraging the power of language models. Finally, we propose a real-world CVD method based on multi-feature fusion to identify potential security vulnerabilities in the code. Experiments on a large-scale real-world CVD dataset demonstrate the effectiveness of MARCOVul in detecting real-world code vulnerabilities, achieving up to 12.75% BF1 and 6.98% MCC improvements over the best unweighted baselines, and 1.55% BF1 and 1.42% MCC gains over the best weighted ones. Jixian Zhang 0001, Qingfeng Du, Zhongda Lu |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | KEWS: A KPIs-Based Evaluation Framework of Workload Simulation On Microservice SystemabstractSimulating the workload is an essential procedure in microservice systems as it helps augment realistic workloads whilst safeguarding user privacy. The efficacy of such simulation depends on its dynamic assessment. The straightforward and most efficient approach to this is comparing the original workload with the simulated one using Key Performance Indicators (KPIs), which capture the state of the system. Nonetheless, due to the extensive volume and complexity of KPIs, fully evaluating them is not feasible, and measuring their similarity poses a significant challenge. This paper introduces a similarity metric algorithm for KPIs, the Extended Shape-Based Distance (ESBD), which gauges similarity in both shape and intensity. Additionally, we propose a KPI-based Evaluation Framework for Workload Simulations (KEWS), comprising three modules: preprocessing, compression, and evaluation. These methodologies effectively counteract the adverse effects of KPIs’ characteristics and offer a holistic evaluation. Experimental results substantiate the effectiveness of both ESBD and KEWS. Pengsheng Li, Qingfeng Du, Shengjie Zhao 0001 |
CSCWD | 2 |
| 2024 | Advancing Root Cause Analysis in Cloud-native System with Knowledge Graph Path Embedding TranslationabstractCloud computing technologies, including cloud-native and containerization, have gained prominence in recent years, attributed to their exceptional scalability, enhanced resource utilization, and expedited deployment capabilities. However, their inherent complexity and the intricate interplay of internal components heighten the risk of sporadic and unforeseen anomalies. To address these challenges, Root Cause Analysis (RCA) is employed to accurately identify problematic services (pods) and mine the precise faults behind observed anomalies. Tailored to the limitations of conventional RCA algorithms, we propose a novel approach that jointly models operation entities and their relationships as learnable embeddings. Additionally, this method integrates fault propagation information to further improve RCA accuracy. Our evaluation involves developing a prototype within the Kubernetes cloud-native system. Extensive experimental results validate the efficacy of our approach. Pengsheng Li, Qingfeng Du, Shengjie Zhao 0001, Pei Fang |
CSCWD | 2 |
| 2024 | Semi-Supervised Metrics-Based Self-Training Root Cause Analysis for Cloud-Native Systems with Class-Imbalanced DataabstractRoot cause analysis is crucial for cloud-native systems. However, existing supervised approaches ignore the potential of unlabeled data, which is frequent in the cloud-native root cause analysis scenarios. Moreover, the class-imbalanced distribution of faults presents obstacles to applying semi-supervised learning. To overcome these limitations, we propose STRCA, a metrics-based semi-supervised self-training approach for root cause analysis. Furthermore, STRCA employs minority priority self-training, which selects pseudo-labels of high quality during generations. Additionally, the stepwise distribution alignment is introduced to rebalance the predicted distribution with gradually decreasing strength. These two strategies mitigate the class-imbalance of data in semi-supervised learning. Experiments on the public dataset show the effectiveness of STRCA with limited labels. Qingfeng Du, Yongqi Han 0001, Fulong Tian |
ICASSP | 2 |
| 2024 | The Potential of One-Shot Failure Root Cause Analysis: Collaboration of the Large Language Model and Small ClassifierabstractFailure root cause analysis (RCA), which systematically identifies underlying faults, is essential for ensuring the reliability of widely adopted microservice-based applications and cloud-native systems. However, manual analysis by simple rules faces significant burdens due to the heterogeneous nature of resource entities and the massive amount of observability data. Furthermore, existing approaches for automating RCA struggle to perform in-depth fault analysis without extensive fault labels. To address the scarcity of fault labels, we examine an extreme RCA scenario where each fault type has only one example (one-shot). We propose LasRCA, a framework for one-hot RCA in cloud-native systems that leverages the collaboration of the large language model (LLM) and the small classifier. In the training stage, LasRCA initially trains a small classifier based on one-shot fault examples. The small classifier then iteratively selects high-confusion samples and receives feedback on their fault types from LLM-driven fault labeling. These samples are applied to retrain the small classifier. In the inference stage, LasRCA performs a joint RCA through the collaboration of the LLM and small classifier, achieving a trade-off between effectiveness and cost. Experiment results on public datasets with heterogeneous nature and prevalent fault types show the effectiveness of LasRCA in one-shot RCA. Yongqi Han 0001, Qingfeng Du, Fulong Tian |
ASE | 2 |
| 2024 | Holistic Root Cause Analysis for Failures in Cloud-Native Systems Through Observability DataabstractMicroservices are widely adopted in large IT enterprises, leveraging the scalability, resiliency, and elasticity of the cloud-native architecture. Effective root cause analysis is crucial for ensuring the reliability of such cloud-native systems. Many efforts have focused on using the three modalities of observability data–traces, metrics, and logs. However, existing approaches are limited by inconsistent problem definitions and cloud-native heterogeneity. To address these challenges, we proposeHolisticRCA, a root cause analysis framework in cloud-native systems from a holistic perspective.HolisticRCAformally defines root cause analysis through three dimensions. ThenHolisticRCAuses an “assembling building blocks” strategy to address the cloud-native heterogeneity. It maps each observability feature into a shared vector space and concatenates the vector embeddings associated with each resource entity for standardized resource entity vector embeddings. Then it applies Graph Attention Network to capture intertwined resource entity relations and incorporates mask embeddings to enable holistic analysis. The evaluation results on three public datasets show thatHolisticRCAoutperforms existing approaches in holistic root cause analysis of cloud-native systems. Yongqi Han 0001, Qingfeng Du, Pengsheng Li, Xiaonan Shi, Pei Fang, Fulong Tian |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | LogFold: Enhancing Log Anomaly Detection Through Sequence Folding and ReconstructionabstractModern large-scale systems and networks necessitate automated anomaly detection to support the high availability and quality of services. Since logs are an essential data source that can accurately reflect the state of a system, log anomaly detection has attracted a lot of attention from researchers in both academia and industry. As the technology of artificial intelligence advances, plenty of work has adopted deep learning to detect log anomalies and achieved promising results. Nevertheless, it usually suffers from a lack of labels, excessive log sequence length, and low throughput problems when deploying to real-world systems. To address these challenges, we propose Log-Fold, an unsupervised Transformer-based log anomaly detection approach. In LogFold, we propose fold embedding, which can compress long log sequences to enhance the efficiency of anomaly detection. And we design a sequence reconstruction technique to enhance the effectiveness of anomaly detection. Our evaluation shows LogFold achieves 90.55% and 99.90% Fl-score on HDFS and BGL datasets, respectively, outperforming state-of-the-art methods. Besides, the fold embedding layer achieves compression rates of 36.55% and 64.86% on HDFS and BGL datasets, respectively, which helps to improve the throughput of LogFold. Xiaonan Shi, Qingfeng Du, Fulong Tian |
APSEC | 3 |
| 2023 | Trace-Based Anomaly Detection with Contextual Sequential Invocations
Qingfeng Du, Fulong Tian, Yongqi Han 0001 |
DEXA (2) | 1 |
| 2023 | LWS: A framework for log-based workload simulation in session-based SUT
Yongqi Han 0001, Qingfeng Du, Jincheng Xu, Shengjie Zhao 0001, Zhekang Chen, Kanglin Yin, Dan Pei |
J. Syst. Softw. | 2 |
| 2021 | A Requirement-based Regression Test Selection Technique in Behavior-Driven DevelopmentabstractRegression testing is an essential software maintenance activity before the release of a new version implementing a bug fix or a new feature. A regression test selection (RTS) technique chooses a subset of existing test cases to ensure that the system will not be adversely affected by the latest modifications. With the rise of DevOps, behavior-driven development (BDD) is growing in popularity as it is in close alignment with agile practices, for example, continuous integration. Hence, it is necessary to propose a novel and effective RTS technique for BDD specifically to accelerate the development process while ensuring software quality. Since most existing techniques for RTS are code-based and thus subject to some limitations, we present a requirement-based technique which uses the requirements in BDD to select test cases in both high-level (acceptance testing) and low-level (unit testing). Our technique firstly illustrates the new requirement with a scenario, and subsequently computes the semantic similarity of the new scenario and all existing scenarios with the vector space model. According to the results, the modification-traversing regression test cases can be selected in a semi-automated way. We also conduct an experimental study to evaluate our technique in terms of inclusiveness, precision, efficiency and generality. The study shows that our technique is applicable for BDD and effective in practice. Jincheng Xu, Qingfeng Du |
COMPSAC | 2 |
| 2021 | Log-Based Anomaly Detection with Multi-Head Scaled Dot-Product Attention Mechanism
Qingfeng Du, Jincheng Xu, Yongqi Han 0001, Shuangli Zhang |
DEXA (1) | 1 |
| 2021 | Model-Agnostic Local Explanations with Genetic Algorithms for Text ClassificationabstractThe interpretability of black-box text classification models has been receiving widespread attention in recent years accompanying the growing popularity of artificial intelligence.To garner user trust on the model's decision-making process, it is imperative to provide faithful instance-wise justifications and rationalize the prediction in a human-readable way.In this paper, we address this challenge by introducing Locally Universal Rules (LURs) as model-agnostic local explanations.LURs are a subset of input words sufficient for the model to arrive at a particular prediction, even if the rest of words are perturbed slightly.We show the identification of the optimal LUR is NP-complete.Consequently, we propose a population-based algorithm LUR-Locator to perform the constrained optimization efficiently.We conduct extensive experiments to evaluate our algorithm on a cross product of well-established text classification datasets and models.The empirical results demonstrate that LURLocator can efficiently generate high-quality local explanations, as compared to existing explanatory methods. Qingfeng Du, Jincheng Xu |
SEKE | 1 |
| 2021 | Towards a Better Understanding of Gradient-Based Explanatory Methods in NLPabstractTo grasp what makes the deep learning models arrive at a particular prediction, gradient-based explanatory methods have been widely used in Natural Language Processing (NLP) recently.While the saliency maps of images can be computed directly in the pixel-level input space, the continuous gradient vector for words has to be reduced to a single value to indicate the word-level importance, and existing methods such as Sensitivity Analysis (SA) and Gradient × Input (GI) are either tricky or short of a deep investigation.In this paper, we review the family of gradient-based explanatory methods and discuss their practical implications.Specially, we propose the signed version of GI, namely SignedGI, while some previous work may have misunderstandings on its signedness.We also show the weakness of SA-based methods.We conduct extensive experiments to evaluate these explanatory methods both qualitatively and quantitatively. Qingfeng Du, Jincheng Xu |
SEKE | 1 |
| 2021 | Predicting Cutterhead Torque for TBM based on Different Characteristics and AGA-Optimized LSTM-MLPabstractAdaptive adjustment of excavation parameters makes a significant role in the process of tunneling by tunnel boring machine (TBM), which ensures the tunneling carried out safely and efficiently. Though substantial effort has been devoted to this area, there is still a lack of a comprehensive method for TBM data analysis. In this paper, we analyzed the TBM data from different perspectives. The data source is from the Songhua River Water Conveyance Project. In order to facilitate the processing and analysis of the data, we proposed the concepts of rising characteristic interval (RCI) and stable characteristic interval (SCI), which are the first 30 seconds of the rising stage and one sixth of the center part of the stable stage respectively. As a key parameter, the cutterhead torque (T), which reflects the interaction between the cutter and the soil, is selected as our prediction target. In order to forecast the value of T in the SCIs, the time series characteristic and the non time series (mean and variance) characteristic of the important excavation parameters in the RCIs are analyzed. A sequential combination of long short-term memory (LSTM) and multi-layer perceptrons (MLP), LSTM-MLP for short, is used to make a comprehensive analysis of the two characteristics. Notably, adaptive genetic algorithm (AGA) was employed to optimize the topology structure and the hyper parameters of our neural network, which ensures the convergence of the basic genetic algorithm and maintains the diversity of the population at the same time. The experimental results indicate that, LSTM-MLP performs better in comparison with LSTM network and backpropagation neural network (BPNN, a kind of MLP). Our work provides a reference for the control and optimization of TBM’s excavation parameters. To make our results fully reproducible, all the relevant source codes and the preprocessed dataset are publicly available at https://github.com/Dandelionslove/LSTM MLP for TBM. Shuangli Zhang, Qingfeng Du, Sicheng Zhao |
SMC | 2 |
| 2021 | On Representing Resilience Requirements of Microservice Architecture SystemsabstractTogether with the spread of DevOps practices and container technologies, Microservice Architecture has become a mainstream architecture style in recent years. Resilience is a key characteristic in Microservice Architecture (MSA) Systems, and it shows the ability to cope with various kinds of system disturbances which cause degradations of services. However, due to lack of consensus definition of resilience in the software field, although a lot of work has been done on resilience for MSA Systems, developers still do not have a clear idea on how resilient an MSA System should be, and what resilience mechanisms are needed. In this paper, by referring to existing systematic studies on resilience in other scientific areas, the definition of microservice resilience is provided and a Microservice Resilience Measurement Model is proposed to measure service resilience. And a requirement model to represent resilience requirements of MSA Systems is given. The requirement model uses elements in KAOS to represent notions in the measurement model, and decompose service resilience goals into system behaviors that can be executed by system components. As a proof of concept, a case study is conducted on an MSA System to illustrate how the proposed models are applied. Kanglin Yin, Qingfeng Du |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2020 | Software Defect Prediction and Localization with Attention-Based Models and Ensemble LearningabstractSoftware defect prediction (SDP) utilizes a trained prediction model to predict the defect proneness of code modules in a software system by mining the inherent characteristics of historical defect data. An effective model can optimize the allocation of testing resources, thus improving the quality of software products. Most previous studies use handcrafted features to represent code snippets, but the main problem is that it is difficult to capture the semantic and structural information of the code context, which is often crucial for software defect prediction. Meanwhile, most of the existing software defect prediction models cannot make predictions at the code line level, which makes it extremely arduous to provide developers with more detailed reference information. To address these issues, in this paper, we propose a model based on ensemble learning techniques and attention mechanisms to offer more comprehensive prediction information to developers by locating suspect lines of code when making method-level defect predictions. This model leverages abstract syntax trees (ASTs) as the intermediate representation of code snippets. Since the historical defect data has a striking characteristic of class-imbalance, an approach based on Self-organizing Map (SOM) clustering is employed to handle noisy data. Experimental results show that, on average, the proposed model improves the F-measure by 17.7% and AUC by 37.8%, compared with the other four machine learning algorithms. Tianhang Zhang, Qingfeng Du, Jincheng Xu, Jiechu Li |
APSEC | 2 |
| 2020 | On the Interpretation of Convolutional Neural Networks for Text Classification
Jincheng Xu, Qingfeng Du |
ECAI | 2 |
| 2020 | Document-Improved Hierarchical Modular Attention for Event Detection
Yiwei Ni, Qingfeng Du, Jincheng Xu |
KSEM (2) | 2 |
| 2020 | TextTricker: Loss-based and gradient-based adversarial attacks on text classification models
Jincheng Xu, Qingfeng Du |
Eng. Appl. Artif. Intell. | 2 |
| 2020 | Adversarial attacks on text classification models using layer-wise relevance propagationabstractDue to the nested nonlinear structure inside neural networks, most existing deep learning models are treated as black boxes, and they are highly vulnerable to adversarial attacks. On the one hand, adversarial examples shed light on the decision-making process of these opaque models to interrogate the interpretability. On the other hand, interpretability can be used as a powerful tool to assist in the generation of adversarial examples by affording transparency on the relative contribution of each input feature to the final prediction. Recently, a post-hoc explanatory method, layer-wise relevance propagation (LRP), shows significant value in instance-wise explanations. In this paper, we attempt to optimize the recently proposed explanation-based attack algorithms (EAAs) on text classification models with LRP. We empirically show that LRP provides good explanations and benefits existing EAAs notably. Apart from that, we propose a LRP-based simple but effective EAA, LRPTricker. LRPTricker uses LRP to identify important words and subsequently performs typo-based perturbations on these words to generate the adversarial texts. The extensive experiments show that LRPTricker is able to reduce the performance of text classification models significantly with infinitesimal perturbations as well as lead to high scalability. Jincheng Xu, Qingfeng Du |
Int. J. Intell. Syst. | 2 |
| 2020 | Learning neural networks for text classification by exploiting label relations
Jincheng Xu, Qingfeng Du |
Multim. Tools Appl. | 2 |
| 2020 | Learning transferable features in meta-learning for few-shot text classification
Jincheng Xu, Qingfeng Du |
Pattern Recognit. Lett. | 2 |
| 2019 | Short-Term Performance Metrics Forecasting for Virtual Machine to Support Anomaly Detection Using Hybrid ARIMA-WNN ModelabstractAnomaly detection is a significant functionality in most cloud monitoring applications. Time-series forecasting model could be easily used for predicting the values of the performance metrics which could be used for representing the performance status of the cloud environment. The proposed hybrid model combines both Autoregressive Integrated Moving Average (ARIMA) and Wavelet Neural Network (WNN) models. Firstly, ARIMA model is employed to firstly predict the linear component and then WNN model is used for the nonlinear residual component prediction. Finally, the results of the two parts are combined into the final prediction value of the performance metric. Finally the experimental results show that the hybrid model could produce more accurate short-term prediction than other models. Qingfeng Du, Kanglin Yin |
COMPSAC (2) | 2 |
| 2018 | An Approach of Collecting Performance Anomaly Dataset for NFV Infrastructure
Qingfeng Du, Tiandi Xie, Kanglin Yin |
ICA3PP (3) | 1 |
| 2018 | Anomaly Detection and Diagnosis for Container-Based Microservices with Performance Monitoring
Qingfeng Du, Tiandi Xie |
ICA3PP (4) | 1 |
| 2018 | Performance Anomaly Detection Models of Virtual Machines for Network Function Virtualization Infrastructure with Machine Learning
Qingfeng Du, YiQun Lin, Jiaye Zhu, Kanglin Yin |
ICANN (2) | 2 |
| 2018 | Helpful or Not? An investigation on the feasibility of identifier splitting via CNN-BiLSTM-CRF
Jiechu Li, Qingfeng Du, Jincheng Xu |
SEKE | 2 |