VLDB 2026 Research / reviewers in the wild / expert
Qifan He
dblp:232/9873
· DBLP profile ↗
9ranked-venue papers
2as first author
8since 2021 · last 2025
0009-0003-7976-1928ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TD4ITG: A Test Data Generation Method for Issue Title Generation ModelsabstractIn open-source software platforms, users utilize issues to report software bugs or request new features. To improve the quality of issues, researchers have proposed several methods for issue title generation. It is widely recognized that deep learning models often suffer from robustness limitations, as minor input perturbations can lead to incorrect or significantly altered outputs. In this paper, we investigate the robustness of issue title generation models and propose a corresponding test data generation method, TD4ITG. This method leverages large language models in combination with chain-of-thought prompting to automatically generate test data to evaluate robustness. Experimental results demonstrate that both iTAPE and iTiger, two issue title generation models, exhibit robustness problems. Specifically, the test data generated by TD4ITG leads to a performance degradation of 21.73% for iTAPE, reducing its score to 75.00%, and a degradation of 17.34% for iTiger, reducing its score to 45.98%. Compared to MATS, a recently proposed testing method for text summarization models, TD4ITG is more effective in revealing the robustness limitations of the models. Qifan He, Zhanqi Cui |
SMC | 3 |
| 2025 | MR-OT: a Metamorphic Testing Method for Object Tracking ModelsabstractTo ensure the reliability of DNN models, researchers have proposed various testing methods. However, most existing methods focus on static tasks such as image recognition, neglecting the challenges posed by temporal continuity of input data and environmental changes in dynamic scenarios like object tracking and behavior detection. In this paper, we propose MR-OT, a metamorphic testing method for object tracking models in dynamic scenarios. We design five metamorphic relations to evaluate the robustness of the model under three common scenarios, which include natural weather variations, speed changes, and environmental disturbances. Additionally, leveraging generative adversarial networks (GANs) and other techniques, we generate large-scale, realistic, and temporally consistent test data based on original test data. Experimental results demonstrate that MR-OT achieves greater metamorphic relation violation rate, outperforming DeepTest by 0.06 to 0.49 and MT4MOT by 0.04 to 0.57, validating its effectiveness in uncovering robustness issues in dynamic tracking tasks. Zhanqi Cui, Qifan He |
SMC | 2 |
| 2025 | Testing Autonomous Driving System via Object-Level ReplacementabstractWith the rapid advancement of autonomous driving technology, ensuring the robustness and reliability of decision-making modules has become a critical challenge for the safety of autonomous driving systems (ADSs). In this paper, we propose a novel method, SOGR (Semantic-Guided Object Replacement), to evaluate the decision consistency of ADSs by constructing highly misleading test images that preserve the semantic integrity of the original driving scenes. SOGR identifies important objects using Grad-CAM and YOLOv8, and replaces them with semantically equivalent objects generated by Stable Diffusion. Experiments conducted on the BDD100K dataset demonstrate that SOGR outperforms the pixel-level perturbation baseline DeepIA, achieving higher misleading rates while maintaining lower perceptual similarity (LPIPS) scores. These results indicate that SOGR can effectively expose model vulnerabilities while maintaining high visual realism, offering a practical and semantically grounded approach for robustness testing in real-world autonomous driving scenarios. Songcheng Xie, Qifan He, Zhanqi Cui |
SMC | 2 |
| 2025 | IATT: Interpretation Analysis-based Transferable Test Generation for Convolutional Neural NetworksabstractConvolutional Neural Networks (CNNs) have been widely used in various fields. However, it is essential to perform sufficient testing to detect internal defects before deploying CNNs, especially in security-sensitive scenarios. Generating error-inducing inputs to trigger erroneous behavior is the primary way to detect CNN model defects. However, in practice, when the model under test is a black-box CNN model without accessible internal information, in some scenarios it is still necessary to generate high-quality test inputs within a limited testing budget. In such a new scenario, a potential approach is to generate transferable test inputs by analyzing the internal knowledge of other white-box CNN models similar to the model under test, and then use transferable test inputs to test the black-box CNN model. The main challenge in generating transferable test inputs is how to improve their error-inducing capability for different CNN models without changing the test oracle. We found that different CNN models make predictions based on features of similar important regions in images. Adding targeted perturbations to important regions will generate transferable test inputs with high realism. Therefore, we propose the Interpretable Analysis-based Transferable Test (IATT) Generation method for CNNs, which employs interpretation methods of CNN models to explain and localize important regions in test inputs, using backpropagation optimizer and perturbation mask process to add targeted perturbations to these important regions, thereby generating transferable test inputs. This process is repeated to iteratively optimize the transferability and realism of the test inputs. To verify the effectiveness of IATT, we perform experimental studies on nine deep learning models, including ResNet-50 and Vit-B/16, and commercial computer vision system Google Cloud Vision , and compared our method with four state-of-the-art baseline methods. Experimental results show that transferable test inputs generated by IATT can effectively cause black-box target models to output incorrect results. Compared to existing testing and adversarial attack methods, the average Error-inducing Success Rate (ESR) in different testing scenarios is 18.1%–52.7% greater than the baseline methods. Additionally, the test inputs generated by IATT achieve high ESR while maintaining high realism. Ruilin Xie, Xiang Chen 0005, Qifan He, Bixin Li, Zhanqi Cui |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2024 | Issue Title Generation: How Far Can Large Language Models Go?abstractIn open-source software and platforms, developers utilize issues to record software failures or propose new features. The title of an issue, which is a mandatory field, should accurately describe the core content in a concise way. However, developers often face challenges in crafting high-quality issue titles due to insufficient experience or limited proficiency. As a result, researchers have proposed several methods for automatically generating issue titles, but typically relying on constructing large datasets to train models. Recently, Large Language Models (LLMs) have exhibited exceptional performance across a variety of general tasks, suggesting significant potential for issue title generation. Initial experiments indicate that the direct application of LLMs fails to yield satisfactory results. Therefore, we propose a method named LBITG (LLMs-Based Issue Title Generation). LBITG enhances the effectiveness of LLMs by providing contextual information through four types of prompts, which include example prompt and label prompt. These prompts serve as guidance for LLMs, thereby further improving their performance. Experimental results demonstrate that LBITG can significantly enhance the quality of issue titles generated by LLMs without any training. In the within-project scenario, LBITG achieves a minimum improvement of 111.29% in ROUGE, 104.54% in BLEU, and 188.48% in METEOR compared to iTAPE, and achieves performance comparable to that of the SOTA method iTiger. Moreover, LBITG demonstrates superior performance in the cross-project scenario, which outperforms iTiger by 25.33%, 30.14%, and 27.29% in terms of ROUGE-1, BLEU-1, and METEOR, respectively. Shifan Liu, Qifan He, Songcheng Xie, Zhanqi Cui |
SMC | 3 |
| 2023 | DeepIA: An Interpretability Analysis based Test Data Generation Method for DNNabstractRecently, deep neural networks (DNN) have been widely applied in various fields, such as image classification, even replace humans to make decisions in some specific tasks. However, like traditional software, DNNs inevitably contain defects. If defective DNN models are applied in safety-critical fields, such as autonomous driving and medical diagnosis, it may cause disastrous consequences. Therefore, effective testing methods are urgently needed to improve the reliability of DNNs. The existing DNN testing methods typically generate test data by either globally modifying the original data or taking adversarial approaches. The generated test data typically struggle to simultaneously achieve good performance in both the degree of difference from the original data and the Error-inducing Success Rate (ESR) with respect to the target DNN model. Moreover, the perturbation-based methods are difficult to be understood by humans. To address the above issue, this paper proposes DeepIA, an interpretability analysis based test data generation method for DNN. DeepIA analyzes the interpretability of decision-making behaviors for DNN. According to the interpretability analysis results, the original training data is split into different regions to evaluate their influences on decision-making results of the DNN. After that, the most significant regions of the original test data are transformed to generate new test data. Experimental results show that the interpretability method effectively enhances the misleading ability of DeepIA for the DNN model under test. Compared with DeepTest and DeepSearch, DeepIA can generate test data with minor permutations and greater ESR. Qifan He, Ruilin Xie, Li Li 0114, Zhanqi Cui |
QRS | 1 |
| 2023 | MOBTAG: Multi-Objective Optimization Based Textual Adversarial Example GenerationabstractNatural language processing (NLP) models are vulnerable to adversarial examples. Generating high-quality adversarial examples, which expose the vulnerability of NLP models and can be used to evaluate and improve their robustness, deserves further research. Existing techniques of generating adversarial examples in the NLP field are typically based on greedy synonym replacements, which may result in out-of-context and unnatural perturbations, and are easily identifiable by humans. In this paper, we present MOB-TAG, a Multi-objective Optimization based Textual Adversarial Example Generation method, which includes three types of perturbations, and utilizes pre-trained models such as BERT and RoBERTa to generate high-quality adversarial examples. MOBTAG generates fluent and grammatical output through a mask-then-infill procedure, with introducing multi-objective optimization and genetic algorithm to pursue a high attack success rate while maintaining a high level of similarity and readability. Experimental results show that compared with methods such as TextFooler, BERTAttack, and CLARE, MOBTAG improves the attack success rate and the textual similarity by at least 11.8% and 0.09 on average, respectively. Yuanxin Qiao, Ruilin Xie, Li Li 0114, Qifan He, Zhanqi Cui |
SMC | 4 |
| 2022 | Immune infiltration and clinical significance analyses of the coagulation-related genes in hepatocellular carcinomaabstractHepatocellular carcinoma (HCC) is one of the most common types of cancers and a global health challenge with a low early diagnosis rate and high mortality. The coagulation cascade plays an important role in the tumor immune microenvironment (TME) of HCC. In this study, based on the coagulation pathways collected from the KEGG database, two coagulation-related subtypes were distinguished in HCC patients. We demonstrated the distinct differences in immune characteristics and prognostic stratification between two coagulation-related subtypes. A coagulation-related risk score prognostic model was developed in the Cancer Genome Atlas (TCGA) cohort for risk stratification and prognosis prediction. The predictive values of the coagulation-related risk score in prognosis and immunotherapy were also verified in the TCGA and International Cancer Genome Consortium cohorts. A nomogram was also established to facilitate the clinical use of this risk score and verified its effectiveness using different approaches. Based on these results, we can conclude that there is an obvious correlation between the coagulation and the TME in HCC, and the risk score could serve as a robust prognostic biomarker, provide therapeutic benefits for chemotherapy and immunotherapy and may be helpful for clinical decision making in HCC patients. Qifan He, Yonghai Jin |
Briefings Bioinform. | 1 |
| 2018 | Light-Weight Object Detection and Decision Making via Approximate Computing in Resource-Constrained Mobile RobotsabstractMost of the current solutions for autonomous flights in indoor environments rely on purely geometric maps (e.g., point clouds). There has been, however, a growing interest in supplementing such maps with semantic information (e.g., object detections) using computer vision algorithms. Unfortunately, there is a disconnect between the relatively heavy computational requirements of these computer vision solutions, and the limited computation capacity available on mobile autonomous platforms. In this paper, we propose to bridge this gap with a novel Markov Decision Process framework that adapts the parameters of the vision algorithms to the incoming video data rather than fixing them a priori. As a concrete example, we test our framework on a object detection and tracking task, showing significant benefits in terms of energy consumption without considerable loss in accuracy, using a combination of publicly available and novel datasets. Parul Pandey, Qifan He, Dario Pompili, Roberto Tron |
IROS | 2 |