Yanqi Su

dblp:16/4518 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 4 first-author · 6 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2025 A Safe-Critical and Efficient Self-Merging Strategy for CAVs in Mixed Traffic Scenarios
abstract
Connected and autonomous vehicles (CAVs) are emerging as a potential solution to merging safety problems. However, in mixed traffic scenarios where CAVs coexist with human-driven vehicles (HVs), challenges arise due to the lack of proactive cooperation and the limited length of the acceleration lane, complicating the merging processes of CAVs. In these cases, CAVs should actively seize the transit opportunity and perform safe and efficient merging. Failure to do so can lead to decreased traffic efficiency, increased fuel consumption and emissions, compromised self-merging capacities, and heightened crash risk. Therefore, this article employs the roadside unit and proposes a two-level hierarchical self-merging strategy for CAVs to increase the merging efficiency while ensuring high safety. Since the surrogate safety measures (SSMs) can formulate reliable safety assessment and identify the merging conflict risk by setting appropriate threshold, the upper level uses a novel SSM-based method (i.e., the minimum acceleration rate, MIAR) to determine the merging sequence (MS). A theoretical model for merging safety (TMMS) is developed to estimate the MIAR value, and the MIAR threshold is determined using signal detection theory (SDT). At the lower level, the strategy recommends the optimized merging maneuvers, pregenerated using sequential quadratic programming-model predictive control (SQP-MPC), based on the determined MS and the CAV’s velocity. A case study at a real-world freeway merging area demonstrates the effectiveness of the MIAR in measuring merging conflict risk. Besides, numerous simulations are conducted, and the results demonstrate that the proposed strategy significantly improves merging success rate and overall traffic efficiency.
Siyang Jiang, Menglu Gu, Yanqi Su, Chang Wang 0002, Wenhui Wei
IEEE Internet Things J.3
2025 Automating TODO-missed Methods Detection and Patching
abstract
TODO comments are widely used by developers to remind themselves or others about incomplete tasks. In other words, TODO comments are usually associated with temporary or suboptimal solutions. In practice, all the equivalent suboptimal implementations should be updated (e.g., adding TODOs) simultaneously. However, due to various reasons (e.g., time constraints or carelessness), developers may forget or even are unaware of adding TODO comments to all necessary places, which results in the TODO-missed methods . These “hidden” suboptimal implementations in TODO-missed methods may hurt the software quality and maintainability in the long-term. Therefore, in this article, we propose the novel task of TODO-missed methods detection and patching and develop a novel model, namely T O D O-comment Patcher ( TDPatcher ), to automatically patch TODO comments to the TODO-missed methods in software projects. Our model has two main stages: offline learning and online inference. During the offline learning stage, TDPatcher employs the GraphCodeBERT and contrastive learning for encoding the TODO comment (natural language) and its suboptimal implementation (code fragment) into vector representations. For the online inference stage, we can identify the TODO-missed methods and further determine their patching position by leveraging the offline trained model. We built our dataset by collecting TODO-introduced methods from the top-10,000 Python GitHub repositories and evaluated TDPatcher on them. Extensive experimental results show the promising performance of our model over a set of benchmarks. We further conduct an in-the-wild evaluation that successfully detects 26 TODO-missed methods from 50 GitHub repositories.
Zhipeng Gao 0002, Yanqi Su, Xing Hu 0008, Xin Xia 0001
ACM Trans. Softw. Eng. Methodol.2
2024 Enhancing Exploratory Testing by Large Language Model and Knowledge Graph
abstract
Exploratory testing leverages the tester's knowledge and creativity to design test cases for effectively uncovering system-level bugs from the end user's perspective. Researchers have worked on test scenario generation to support exploratory testing based on a system knowledge graph, enriched with scenario and oracle knowledge from bug reports. Nevertheless, the adoption of this approach is hindered by difficulties in handling bug reports of inconsistent quality and varied expression styles, along with the infeasibility of the generated test scenarios. To overcome these limitations, we utilize the superior natural language understanding (NLU) capabilities of Large Language Models (LLMs) to construct a System KG of User Tasks and Failures (SysKG-UTF). Leveraging the system and bug knowledge from the KG, along with the logical reasoning capabilities of LLMs, we generate test scenarios with high feasibility and coherence. Particularly, we design chain-of-thought (CoT) reasoning to extract human-like knowledge and logical reasoning from LLMs, simulating a developer's process of validating test scenario feasibility. Our evaluation shows that our approach significantly enhances the KG construction, particularly for bug reports with low quality. Furthermore, our approach generates test scenarios with high feasibility and coherence. The user study further proves the effectiveness of our generated test scenarios in supporting exploratory testing. Specifically, 8 participants find 36 bugs from 8 seed bugs in two hours using our test scenarios, a significant improvement over the 21 bugs found by the state-of-the-art baseline.
Yanqi Su, Dianshu Liao, Zhenchang Xing, Mulong Xie, Qinghua Lu 0001, Xiwei Xu 0001
ICSE1
2023 Still Confusing for Bug-Component Triaging? Deep Feature Learning and Ensemble Setting to Rescue
abstract
To speed up the bug-fixing process, it is essential to triage bugs into the right components as soon as possible. Given the large number of bugs filed everyday, a reliable and effective bug-component triaging tool is needed to assist this task. LR-BKG is the state-of-the-art toolkit for doing this. However, the suboptimal performance for recommending the right component at the first position (low Top-1 accuracy) limits its usage in practice. We thoroughly investigate the limitations of LR-BKG and find out the gap between the manual feature design of LR-BKG and the characteristics of bug reports causes such suboptimal performance. Therefore, we propose an approach, DEEPTRIAG, which uses the large scale pre-trained models to extract deep features automatically from bug reports (including bug summary and description), to fill this gap. DEEPTRIAG transforms bug-component triaging into a multi-classification task (CodeBERT-Classifier) and a generation task (CodeT5-Generator). Then, we ensemble the prediction results from them to improve the performance of bug-component triaging further. Extensive experimental results demonstrate the superior performance of DEEPTRIAG on bug-component triaging over LR-BKG. In particular, the overall Top-1 accuracy is improved from 56.2% to 68.3% on Mozilla dataset and from 51.3% to 64.1% on Eclipse dataset, which verifies the effectiveness and generalization of our approach on improving the practical usage for bug-component triaging.
Yanqi Su, Zheming Han, Zhipeng Gao 0002, Zhenchang Xing, Qinghua Lu 0001, Xiwei Xu 0001
ICPC1
2023 Comparing the Effects of Visual Distraction in a High-Fidelity Driving Simulator and on a Real Highway
abstract
Driving simulators have been widely used in driving behavior analysis and intelligent driving algorithm development. However, the validity of driving behavior data derived from driving simulators remains unclear. In this study, 30 Chinese drivers were recruited to participate in two experiments: on-road and simulator experiments. An instrumented vehicle and a high-fidelity simulator were used in the on-road and simulator experiments, respectively, to investigate the effects of high speed (60, 80, and 100 km/h) and a visual distraction task on the lateral driving performance, including lane deviation (LD), standard deviation of the lane position (SDLP) rate, standard deviation of the steering wheel angle (SDSWA) rate, and steering wheel reversal rates (SRRs) (at levels of 1.3° and 2.5°). It was found that the visual distraction task impaired the drivers’ lane-keeping ability. Furthermore, the driving task had similar effects on the LD, SDLP rate, and SRRs (2.5°) in the on-road and simulator experiments. The effects of the driving speed on the LD, SDLP rate, and SDSWA rate were comparable in both driving environments. However, the results confirmed that even a high-fidelity driving simulator could not achieve perfect absolute validity. The results provided preliminary evidence that the high-fidelity driving simulator used in this study might be an effective tool for investigating the effects of visual distractions task on lateral driving behavior.
Qinyu Sun, Yingshi Guo, Chang Wang 0002, Menglu Gu, Yanqi Su
IEEE Trans. Intell. Transp. Syst.6
2022 Constructing a System Knowledge Graph of User Tasks and Failures from Bug Reports to Support Soap Opera Testing
abstract
Exploratory testing is an effective testing approach which leverages the tester’s knowledge and creativity to design test cases to provoke and recognize failures at the system level from the end user’s perspective. Although some principles and guidelines have been proposed to guide exploratory testing, there are no effective tools for automatic generation of exploratory test scenarios (a.k.a soap opera tests). Existing test generation techniques rely on specifications, program differences and fuzzing, which are not suitable for exploratory test generation. In this paper, we propose to leverage the scenario and oracle knowledge in bug reports to generate soap opera test scenarios. We develop open information extraction methods to construct a system knowledge graph (KG) of user tasks and failures from the steps to reproduce, expected results and observed results in bug reports. We construct a proof-of-concept KG from 25,939 bugs of the Firefox browser. Our evaluation shows the constructed KG is of high quality. Based on the KG, we create soap opera test scenarios by combining the scenarios of relevant bugs, and develop a web tool to present the created test scenarios and support exploratory testing. In our user study, 5 users find 18 bugs from 5 seed bugs in 2 hours using our tool, while the control group finds only 5 bugs based on the recommended similar bugs.
Yanqi Su, Zheming Han, Zhenchang Xing, Xin Xia 0001, Xiwei Xu 0001, Liming Zhu 0001, Qinghua Lu 0001
ASE1
2021 Reducing Bug Triaging Confusion by Learning from Mistakes with a Bug Tossing Knowledge Graph
abstract
Assigning bugs to the right components is the prerequisite to get the bugs analyzed and fixed. Classification-based techniques have been used in practice for assisting bug component assignments, for example, the BugBug tool developed by Mozilla. However, our study on 124,477 bugs in Mozilla products reveals that erroneous bug component assignments occur frequently and widely. Most errors are repeated errors and some errors are even misled by the BugBug tool. Our study reveals that complex component designs and misleading component names and bug report keywords confuse bug component assignment not only for bug reporters but also developers and even bug triaging tools. In this work, we propose a learning to rank framework that learns to assign components to bugs from correct, erroneous and irrelevant bug-component assignments in the history. To inform the learning, we construct a bug tossing knowledge graph which incorporates not only goal-oriented component tossing relationships but also rich information about component tossing community, component descriptions, and historical closed and tossed bugs, from which three categories and seven types of features for bug, component and bug-component relation can be derived. We evaluate our approach on a dataset of 98,587 closed bugs (including 29,100 tossed bugs) of 186 components in six Mozilla products. Our results show that our approach significantly improves bug component assignments for both tossed and non-tossed bugs over the BugBug tool and the BugBug tool enhanced with component tossing relationships, with >20% Top-k accuracies and >30% NDCG@k (k=1,3,5,10).
Yanqi Su, Zhenchang Xing, Xin Peng 0001, Xin Xia 0001, Chong Wang 0013, Xiwei Xu 0001, Liming Zhu 0001
ASE1
2021 User Review-Based Change File Localization for Mobile Applications
abstract
In the current mobile app development, novel and emerging DevOps practices (e.g., Continuous Delivery, Integration, and user feedback analysis) and tools are becoming more widespread. For instance, the integration of user feedback (provided in the form of user reviews) in the software release cycle represents a valuable asset for the maintenance and evolution of mobile apps. To fully make use of these assets, it is highly desirable for developers to establish semantic links between the user reviews and the software artefacts to be changed (e.g., source code and documentation), and thus to localize the potential files to change for addressing the user feedback. In this paper, we proposeRISING(ReviewIntegration via claSsification, clusterIng, and linkiNG), an automated approach to support the continuous integration of user feedback via classification, clustering, and linking of user reviews.RISINGleverages domain-specific constraint information and semi-supervised learning to group user reviews into multiple fine-grained clusters concerning similar users’ requests. Then, by combining the textual information from both commit messages and source code, it automatically localizes potential change files to accommodate the users’ requests. Our empirical studies demonstrate that the proposed approach outperforms the state-of-the-art baseline work in terms of clustering and localization accuracy, and thus produces more reliable results.
Yu Zhou 0010, Yanqi Su, Taolue Chen 0001, Harald C. Gall, Sebastiano Panichella
IEEE Trans. Software Eng.2
2019 Mining and Comparing User Reviews across Similar Mobile Apps
abstract
With the rapid development of the market for mobile apps, there are a number of apps with similar functions. To gain an advantage in such a competitive environment, developers need to understand not only the strengths and weaknesses of their app but also competitive apps. User reviews contain valuable information for comparing similar apps from user preference. In this paper, we propose UISAT (User-review mining via topic Identification, Sentiment Analysis and Topic matching across apps), an automated approach to compare user reviews from similar apps with the goal of mining user feedback from competitive apps by (i) extracting the hidden topics from large volumes of user reviews using topic modeling, (ii) combining a rule-based model, user rating and user-helpful for sentiment analysis of topics and (iii) matching relevant topics across apps. Empirical studies demonstrate that UISAT is effective and promising for developers to build and maintain a more competitive app.
Yanqi Su, Yongchao Wang 0003, Wenhua Yang 0001
MSN1
2006 SynView: a GBrowse-compatible approach to visualizing comparative genome data
abstract
UNLABELLED: We present SynView, a simple and generic approach to dynamically visualize multi-species comparative genome data. It is a light-weight application based on the popular and configurable web-based GBrowse framework. It can be used with a variety of databases and provides the user with a high degree of interactivity. The tool is written in Perl and runs on top of the GBrowse framework. It is in use in the PlasmoDB (http://www.PlasmoDB.org) and the CryptoDB (http://www.CryptoDB.org) projects and can be easily integrated into other cross-species comparative genome projects. AVAILABILITY: The program and instructions are freely available at http://www.ApiDB.org/apps/SynView/ CONTACT: [email protected].
Yanqi Su, Aaron J. Mackey, Eileen T. Kraemer, Jessica C. Kissinger
Bioinform.2