VLDB 2026 Research / reviewers in the wild / expert
Tiantian Wang 0001
dblp:66/4971-1
· DBLP profile ↗
26ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-2958-8066ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 16 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 3Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Func: reducing the impact of Android framework evolution on malware detection
Tiantian Wang 0001, Lwin Khin Shar, Hanmeng Li, David Lo 0001 |
Empir. Softw. Eng. | 2 |
| 2026 | Shield Broken: Black-Box Adversarial Attacks on LLM-Based Vulnerability DetectorsabstractVulnerability detection is critical for ensuring software security. Although deep learning (DL) methods, particularly those employing large language models (LLMs), have shown strong performance in automating vulnerability identification, they remain susceptible to adversarial examples, which are carefully crafted inputs with subtle perturbations designed to evade detection. Existing adversarial attack methods often require access to model architectures or confidence scores, making them impractical for real-world black-box systems. In this paper, we propose SVulAttack, a novel label-only adversarial attack framework targeting LLM-based vulnerability detectors. Our key innovation lies in a similarity-based strategy that estimates statement importance and model confidence, thereby enabling more effective selection of semantic-preserving code perturbations. SVulAttack combines this strategy with a transformation component and a search component, based on either greedy or genetic algorithms, to effectively identify and apply optimal combinations of transformations. We evaluate SVulAttack on open-source models (LineVul, StagedVulBERT, Code Llama, Deepseek-Coder) and closed-source models (GPT-5 nano, GPT-4o, GPT-4o-mini, Claude Sonnet 4). Results show that SVulAttack significantly outperforms existing label-only black-box attack methods. For example, against LineVul, our method with genetic algorithm achieves an attack success rate of 49.0%, improving over DIP and CODA by 150.0% and 240.3%, respectively. Christoph Treude, Xiaohong Su, Tiantian Wang 0001 |
IEEE Trans. Software Eng. | 5 |
| 2025 | Effective Code Membership Inference for Code Completion Models via Adversarial PromptsabstractMembership inference attacks (MIAs) on code completion models offer an effective way to assess privacy risks by inferring whether a given code snippet was part of the training data. Existing black- and gray-box MIAs rely on expensive surrogate models or manually crafted heuristic rules, which limit their ability to capture the nuanced memorization patterns exhibited by over-parameterized code language models. To address these challenges, we propose AdvPrompt-MIA, a method specifically designed for code completion models, combining code-specific adversarial perturbations with deep learning. The core novelty of our method lies in designing a series of adversarial prompts that induce variations in the victim code model’s output. By comparing these outputs with the ground-truth completion, we construct feature vectors to train a classifier that automatically distinguishes member from non-member samples. This design allows our method to capture richer memorization patterns and accurately infer training set membership. We conduct comprehensive evaluations on widely adopted models, such as Code Llama 7B, over the APPS and HumanEval benchmarks. The results show that our approach consistently outperforms state-of-the-art baselines, with AUC gains of up to 102%. In addition, our method exhibits strong transferability across different models and datasets, underscoring its practical utility and generalizability. Christoph Treude, Xiaohong Su, Tiantian Wang 0001 |
ASE | 6 |
| 2025 | Visual Modeling and Simulation of AUTOSAR Application Layer Models Using ModelicaabstractAs automotive electronic architectures grow increasingly complex and software development costs escalate, AUTOSAR plays a critical role in standardizing and enhancing the reusability of automotive controllers. However, existing AUTOSAR application layer modeling tools, such as the Simulink AUTOSAR Blockset, primarily adopt causal modeling paradigms, which constrain flexibility in capturing intricate system interactions. Additionally, their proprietary nature limits model accessibility, hindering cross-platform collaboration and multi-domain integration. Modelica, an object-oriented, equation-based modeling language, is particularly well-suited for multi-domain simulations due to its acausal modeling capabilities and strong support for component reuse. This paper proposes a Modelica-based visual modeling approach for AUTOSAR application layer models. Specifically, it establishes encapsulation rules for representing AUTOSAR constructs in Modelica and develops an open-source AUTOSAR model library, facilitating industry collaboration and accelerating rapid prototyping. A formal mathematical representation of AUTOSAR models is introduced to enhance both expressiveness and verifiability. Furthermore, a structured visual modeling methodology is presented to lower the development barrier for AUTOSAR application layer modeling. Comparative analysis with Simulink's AUTOSAR Blockset demonstrates that the proposed approach successfully integrates Modelica's multi-domain modeling capabilities into the AUTOSAR workflow while ensuring simulation consistency with Simulink. To the best of our knowledge, this work represents the first integration of AUTOSAR within Modelica's multi-domain simulation framework. Compared to Simulink, Modelica's acausal modeling paradigm enables more flexible system representations, its open ecosystem supports cross-platform collaboration, and its multi-domain integration enhances interoperability. Beyond the automotive domain, the proposed approach can also be applied to controller design in other industries, further demonstrating its potential for cross-disciplinary adoption. Peihao Yang, Tiantian Wang 0001, Xiaohong Su |
MODELS | 2 |
| 2025 | Cross-Level Requirements Tracing Based on Large Language ModelsabstractCross-level requirements traceability, linkinghigh-level requirements(HLRs) andlow-level requirements(LLRs), is essential for maintaining relationships and consistency in software development. However, the manual creation of requirements links necessitates a profound understanding of the project and entails a complex and laborious process. Existing machine learning and deep learning methods often fail to fully understand semantic information, leading to low accuracy and unstable performance. This paper presents the first approach for cross-level requirements tracing based on large language models (LLMs) and introduces a data augmentation strategy (such as synonym replacement, machine translation, and noise introduction) to enhance model robustness. We compare three fine-tuning strategies—LoRA, P-Tuning, and Prompt-Tuning—on different scales of LLaMA models (1.1B, 7B, and 13B). The fine-tuned LLMs exhibit superior performance across various datasets, including six single-project datasets, three cross-project datasets within the same domain, and one cross-domain dataset. Experimental results show that fine-tuned LLMs outperform traditional information retrieval, machine learning, and deep learning methods on various datasets. Furthermore, we compare the performance of GPT and DeepSeek LLMs under different prompt templates, revealing their high sensitivity to prompt design and relatively poor result stability. Our approach achieves superior performance, outperforming GPT-4o and DeepSeek-r1 by 16.27% and 16.8% in F1 score on cross-domain datasets. Compared to the baseline method that relies on prompt engineering, it achieves a maximum improvement of 13.8%. Chuyan Ge, Tiantian Wang 0001, Christoph Treude |
IEEE Trans. Software Eng. | 2 |
| 2025 | Enhancing Fine-Grained Vulnerability Detection With Reinforcement LearningabstractThe rapid growth of vulnerabilities has significantly accelerated the development of automated vulnerability detection methods, especially those based on data-driven models. However, most of them primarily focus on extracting accurate code representations while overlooking the complex vulnerability patterns among vulnerable statements, thereby leaving room for improvement. To overcome this limitation, we present a novel reinforcement learning framework (RLFD) for detecting vulnerabilities at a fine-grained level.RLFDredefines the detection task as a sequential decision-making process and then employs reinforcement learning to automatically learn vulnerability-relevant structures from code snippets. Moreover, by designing reward functions aligned with fine-grained evaluation metrics,RLFDfocuses on the co-existence relations among statements from a global perspective, enabling the model to capture complex interactions that lead to vulnerabilities. Additionally, the framework utilizes CodeBERT-HLS for code representation, ensuring consistency with the state-of-the-art method while highlighting the improvements brought by the proposed reinforcement learning-based approach. Comprehensive experiments show that our method achieves a locating precision (IoU) of 69.7% and a Top-5% Acc of 67.7% on thebig_vuldataset, outperforming the state-of-the-art method by an overall 3.4% improvement in IoU. Notably, our method achieves up to a 19.7% increase in IoU for specific categories, e.g., CWE-416 (use-after-free). Zhichen Qu, Christoph Treude, Xiaohong Su, Tiantian Wang 0001 |
IEEE Trans. Software Eng. | 5 |
| 2024 | StagedVulBERT: Multigranular Vulnerability Detection With a Novel Pretrained Code ModelabstractThe emergence of pre-trained model-based vulnerability detection methods has significantly advanced the field of automated vulnerability detection. However, these methods still face several challenges, such as difficulty in learning effective feature representations of statements for fine-grained predictions and struggling to process overly long code sequences. To address these issues, this study introduces StagedVulBERT, a novel vulnerability detection framework that leverages a pre-trained code language model and employs a coarse-to-fine strategy. The key innovation and contribution of our research lies in the development of the CodeBERT-HLS component within our framework, specialized in hierarchical, layered, and semantic encoding. This component is designed to capture semantics at both the token and statement levels simultaneously, which is crucial for achieving more accurate multi-granular vulnerability detection. Additionally, CodeBERT-HLS efficiently processes longer code token sequences, making it more suited to real-world vulnerability detection. Comprehensive experiments demonstrate that our method enhances the performance of vulnerability detection at both coarse- and fine-grained levels. Specifically, in coarse-grained vulnerability detection, StagedVulBERT achieves an F1 score of 92.26%, marking a 6.58% improvement over the best-performing methods. At the fine-grained level, our method achieves a Top-5% accuracy of 65.69%, which outperforms the state-of-the-art methods by up to 75.17%. Yujian Zhang, Xiaohong Su, Christoph Treude, Tiantian Wang 0001 |
IEEE Trans. Software Eng. | 5 |
| 2023 | A Graph Neural Network-Based Smart Contract Vulnerability Detection Method with Artificial Rule
Ziyue Wei, Weining Zheng, Xiaohong Su, Wenxin Tao, Tiantian Wang 0001 |
ICANN (4) | 5 |
| 2023 | Does Deep Learning improve the performance of duplicate bug report detection? An empirical study
Xiaohong Su, Christoph Treude, Tiantian Wang 0001 |
J. Syst. Softw. | 5 |
| 2022 | Fault localization based on wide & deep learning model by mining software behavior
Tiantian Wang 0001, HaiLong Yu, Kechao Wang, Xiaohong Su |
Future Gener. Comput. Syst. | 1 |
| 2022 | Golden Mutator Recommendation Based on Mutation Pattern MiningabstractMutation testing is widely used in the research of evaluation and optimization of test set quality, and has been paid attention to the study of bug localization and fixing. But one inherent problem of mutation testing is huge computation cost. Selective mutation is an important method to reduce mutation testing cost. However, the existing selective mutation researches indicate that there is no universal selection strategy. This paper proposes a method which integrates with historical software data mining to recommend suitable mutator for program under test. The basic idea of the method is in the software version control system, any pair of nonfixed and post-fixed programs can be original program and mutant of each other. The edit performed during bug fixing contains mutators, and such mutator can guide buggy program to correct program, hence it is called Golden Mutator. The basic method is to compare the buggy files and the corresponding fixed files in the version control system, obtain historical faulty statements and their fix editing operations, thereby accumulating the bug-fix instance base, and then mining mutation patterns from it. Before mutating the target statement, we first find out faulty instances similar to the target statement, and use their fixing edits to match with the mutation pattern, so as to get the golden mutator applicable to the target statement. This paper uses Defects4j dataset in the test, verifies the accuracy of the proposed recommendation method, and further uses the method in bug localization. Compared to fixedly selected mutators, when applying the golden mutator recommended by using the proposed method, the average accuracy of bug localization is higher. Dan Gong, Tiantian Wang 0001, Xiaohong Su |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2022 | Equivalent Mutants Detection Based on Weighted Software Behavior GraphabstractThe equivalent mutants problem is one of the crucial problems in mutation testing. In consequence of its existence, the effectiveness of mutation testing is underestimated. In addition, it will produce a certain amount of useless overhead. Equivalent mutants cannot be detected by any test input. The existing works mostly focus on static analysis to detect, or avoid generating, the equivalent mutants. The essence of these methods is to use prior knowledge to establish some rules of program equivalence. However, (1) it needs a lot of professional labor to sort out the equivalence rules, and (2) only a small part of the rules can be determined in advance, because of the diversity of mutation operators and mutation targets. Consequently, the best result reported so far is 50% of the equivalent mutants can be detected. Since it is generally believed that manual judgment of program equivalence is the most reliable, this paper proposes a novel method to automatically detect equivalent mutants by tracing program behavior like the professionals. The weighted software behavior graph is utilized in the detection of equivalent mutants for the first time. This method can not only figure out different execution paths, but also be sensitive to execution frequency. By comparing the weighted software behavior graphs of an alive mutant and its original program, we are able to examine more precisely whether the alive mutant is the same as the original program, in terms of the state of infection and/or the propagation. Evaluation results on an open dataset of manually evaluated equivalent mutants show that our approach can detect 77.5% of all the equivalent mutants, which is much higher than the existing static methods. Dan Gong, Tiantian Wang 0001, Xiaohong Su, Yanhang Zhang |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2022 | Hierarchical semantic-aware neural code representation
Xiaohong Su, Christoph Treude, Tiantian Wang 0001 |
J. Syst. Softw. | 4 |
| 2020 | LTRWES: A new framework for security bug report detection
Pengcheng Lu, Xiaohong Su, Tiantian Wang 0001 |
Inf. Softw. Technol. | 4 |
| 2019 | Invariant based fault localization by analyzing error propagation
Tiantian Wang 0001, Kechao Wang, Xiaohong Su, Lei Zhang 0036 |
Future Gener. Comput. Syst. | 1 |
| 2019 | Automatic debugging of operator errors based on efficient mutation analysis
Tiantian Wang 0001, Jiahuan Xu, Xiaohong Su, ChenShi Li, Yang Chi |
Multim. Tools Appl. | 1 |
| 2018 | An Empirical Study on Software Defect Prediction Using Over-Sampling by SMOTEabstractSoftware defect prediction suffers from the class-imbalance. Solving the class-imbalance is more important for improving the prediction performance. SMOTE is a useful over-sampling method which solves the class-imbalance. In this paper, we study about some problems that faced in software defect prediction using SMOTE algorithm. We perform experiments for investigating how they, the percentage of appended minority class and the number of nearest neighbors, influence the prediction performance, and compare the performance of classifiers. We use paired t-test to test the statistical significance of results. Also, we introduce the effectiveness and ineffectiveness of over-sampling, and evaluation criteria for evaluating if an over-sampling is effective or not. We use those concepts to evaluate the results in accordance with the evaluation criteria for the effectiveness of over-sampling. The results show that they, the percentage of appended minority class and the number of nearest neighbors, influence the prediction performance, and show that the over-sampling by SMOTE is effective in several classifiers. CholMyong Pak, Tiantian Wang 0001, Xiaohong Su |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2017 | Greening Software Requirements Change Management Strategy Based on Nash EquilibriumabstractRecently, green computing has become more and more important in software engineering (SE), which can be achieved by effectively recycling the software system and utilizing the computing resources. However, the requirement change may lead to unnecessary labor and time cost. Moreover, it may also result in the waste of hardware and computing resources once unreasonable requirements are realized. Thus, to perform green computing in SE, it is necessary to propose effective strategies to manage the requirement change. For this decision-making problem, game theoretical methods can be feasible solutions. In this paper, we propose a novel requirement change management approach based on game theory. Specifically, we model the problem as a game between the stakeholders and the developer and devise the payoff matrix between different strategies of the players. We then propose a Nash equilibrium-based game theoretical algorithm to manage requirement change. The evaluation results show that, compared to the exhaustive algorithm, our method not only can achieve almost the same optimal results but also can significantly reduce the computational time complexity. Thus, our method is feasible for a lot of requirement changes and can facilitate the green computing targets from the perspective of software engineering. Zhixiang Tong, Xiaohong Su, Longzhu Cen, Tiantian Wang 0001 |
Wirel. Commun. Mob. Comput. | 4 |
| 2016 | Optimization and improvements of a Moodle-Based online learning system for C programmingabstractIt is important for students to solve problems with specific requirements in the programming teaching. Our teaching system is a Moodle-based interactive teaching platform for C programming. Its online judging system can grade students code automatically. It plays an extremely important role in programming language teaching. This paper is devoted to optimizing and improving the system. We firstly analyze the five problems in the system according to the feedback from the teachers and students: 1) logical errors in students programs cannot be located; 2) a cheating that directly outputs answers cannot be detected; 3) evaluation results lack statistics and visualization; 4) the code submitting procedure is complicated; 5) the feedback of incorrect answers is not detailed. In order to solve these problems, we employ a fault localization algorithm, revise the evaluation logic, introduce third-party visualization plug-ins and refactor the system, respectively. Detailed and exact solutions are given also. After optimizing and improving the system, the user experience is significantly improved. It is convenient for the student to find and correct errors in their programs. Also, it is easier for teachers to acquire valuable feedback and master students' learning situation. Xiaohong Su, Jing Qiu 0003, Tiantian Wang 0001, Lingling Zhao |
FIE | 3 |
| 2015 | Motivating students with new mechanisms of online assignments and examination to meet the MOOC challenges for programmingabstractThe advent of massive open online courses (MOOC) poses challenges for teaching and learning programming. This paper has analyzed these challenges and thereby proposed a self-motivating learning platform for students in the introductory programming course. Novel mechanisms of online assignments and examination have been introduced. Our platform provides functions for self-motivating learning and practicing in MOOC, which makes it distinguish from the others. For example, self-paced timetable with supervision, self-motivated exercise contents, exercise market, and relative ranking. The automatic grading approach is also a highlight. Programs even with syntactic or semantic errors can be automatically graded. Our platform gains popularity among both students and teachers. The platform has been used together with a programming MOOC. This course is ranked as the third most popular courses among over 500 courses. The platform has also been used by more than 100 other universities. The application of the platform in both MOOC and the traditional classroom courses has shown that students' self-motivation in learning programming has been greatly promoted, and their practical skills have also been significantly improved. Xiaohong Su, Tiantian Wang 0001, Jing Qiu 0003, Lingling Zhao |
FIE | 2 |
| 2015 | Interest-driven and innovation-oriented practice for programming courseabstractIn order to maximize the motivation of students in the programming practice, this paper offers an analysis on the core factors of practice case motivating students put in effort in programming practice, namely, "interest", "usability", and "hierarchy". Furthermore, we present typical practice cases which are carefully designed according to the motivating factors and give a description on the implementation and experience of our programming practice course at Harbin Institute of Technology. The designed programming practice can not only train the students' practical programming skills but also enhance their self-regulated learning skills, creativity and self-efficacy. Lingling Zhao, Xiaohong Su, Tiantian Wang 0001, Yongfeng Yuan |
FIE | 3 |
| 2015 | State dependency probabilistic model for fault localization
Dandan Gong, Xiaohong Su, Tiantian Wang 0001, Peijun Ma |
Inf. Softw. Technol. | 3 |
| 2014 | Detection of semantically similar code
Tiantian Wang 0001, Kechao Wang, Xiaohong Su, Peijun Ma |
Frontiers Comput. Sci. | 1 |
| 2013 | Searching for better configurations: a rigorous approach to clone evaluationabstractClone detection finds application in many software engineering activities such as comprehension and refactoring. However, the confounding configuration choice problem poses a widely-acknowledged threat to the validity of previous empirical analyses. We introduce desktop and parallelised cloud-deployed versions of a search based solution that finds suitable configurations for empirical studies. We evaluate our approach on 6 widely used clone detection tools applied to the Bellon suite of 8 subject systems. Our evaluation reports the results of 9.3 million total executions of a clone tool; the largest study yet reported. Our approach finds significantly better configurations (p < 0.05) than those currently used, providing evidence that our approach can ameliorate the confounding configuration choice problem. Tiantian Wang 0001, Mark Harman, Yue Jia 0001, Jens Krinke |
ESEC/SIGSOFT FSE | 1 |
| 2013 | A test-suite reduction approach to improving fault-localization effectiveness
Dandan Gong, Tiantian Wang 0001, Xiaohong Su, Peijun Ma |
Comput. Lang. Syst. Struct. | 2 |
| 2007 | Semantic similarity-based grading of student programs
Tiantian Wang 0001, Xiaohong Su, Peijun Ma |
Inf. Softw. Technol. | 1 |