VLDB 2026 Research / reviewers in the wild / expert
Yiming Tang 0002
dblp:98/5588-2
· DBLP profile ↗
17ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0003-2378-8972ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 17 · 2 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Empirical Study of Privacy Leakage Vulnerability in Third-Party Android Logs Libraries
Yixi Zhao, Kundi Yao, Yiming Tang 0002, Weiyi Shang |
SANER | 3 |
| 2025 | Batch Execution of Microbenchmarks for Efficient Performance TestingabstractPerformance microbenchmarking is essential for ensuring software quality by providing granular insights into code efficiency. While automated performance microbenchmark generation tools (e.g., ju2jmh) are proposed to alleviate practitioners from manually curating microbenchmarks, the high volume of generated benchmarks can lead to protracted benchmarking execution time, as many of the generated benchmarks are too short in nature to be valuable for evaluating performance. In this paper, we present a novel approach that optimizes microbenchmark execution through a batching strategy, i.e., grouping benchmarks with similar code coverage and treating them as a single unit to 1) reduce execution overhead and 2) reduce the bias from microbenchmarks that are too short. We evaluate the effectiveness of this enhancement across various Java projects, comparing the execution times of clustered and individual micro benchmarks. Our findings demonstrate substantial improvements in execution efficiency, reducing execution time by up to 89.81% while preserving high microbenchmark stability. Mostafa Jangali, Kundi Yao, Yiming Tang 0002, Diego Costa 0001, Weiyi Shang |
ICST | 3 |
| 2025 | An Empirical Study of Logging Practice in CUDA-Based Deep Learning SystemsabstractAlthough logging practices have been extensively explored in conventional software systems, there remains a lack of understanding of how logging is applied in CUDAbased deep learning (DL) systems, despite their growing adoption in practice. In this paper, we conduct an empirical study to examine the characteristics and rationales of logging practices in these systems. We analyze logging statements from 33 CUDA-based open-source DL projects, covering both general-purpose logging libraries and DL-specific logging frameworks. For each type, we identify the development or execution phases in which the logs are used and investigate the reasoning behind their usage. Our quantitative analysis reveals that the majority of logging statements occur during the model training phase, with significant usage also in the model loading phase and model evaluation/validation phase. Furthermore, we observe that logging is predominantly used for monitoring purposes and tracking model-related information. Our findings not only shed light on current logging practices in CUDAbased DL development but also provide practical guidance on when to use DL-specific versus general-purpose logging, helping practitioners make more informed decisions and guiding the evolution of DL-focused logging tools to better support developer needs. Kundi Yao, Haonan Zhang 0006, Yiming Tang 0002, Weiyi Shang |
QRS | 4 |
| 2024 | From Logging to Leakage: A Study of Privacy Leakage in Android App LogsabstractAndroid phones are among the most popular mobile devices today, providing users with a wide array of convenient services through various apps. These apps generate software logs during their runtime, which record their behavior, status, and error information. However, these logs can also inadvertently capture sensitive information and user privacy data, often without the developer's awareness. In this study, we constructed a dataset comprising 67,702 log records from 83 Android apps. Our analysis of this dataset identified 610 instances of privacy leakage, which indicates the prevalence of such issues in Android app logs. Additionally, our analysis identified characteristics of Android app logs with exposed sensitive information and revealed a gap between developers' awareness of privacy protection and privacy leakage in real-world scenarios. Soham Sanjay Deo, Poorna Chander Reddy Puttaparthi, Yiming Tang 0002, Xueling Zhang, Weiyi Shang |
ASE | 4 |
| 2024 | LoGenText-Plus: Improving Neural Machine Translation Based Logging Texts Generation with Syntactic TemplatesabstractDevelopers insert logging statements in the source code to collect important runtime information about software systems. The textual descriptions in logging statements (i.e., logging texts) are printed during system executions and exposed to multiple stakeholders including developers, operators, users, and regulatory authorities. Writing proper logging texts is an important but often challenging task for developers. Prior studies find that developers spend significant efforts modifying their logging texts. However, despite extensive research on automated logging suggestions, research on suggesting logging texts rarely exists. To fill this knowledge gap, we first propose LoGenText (initially reported in our conference paper), an automated approach that uses neural machine translation (NMT) models to generate logging texts by translating the related source code into short textual descriptions. LoGenText takes the preceding source code of a logging text as the input and considers other context information, such as the location of the logging statement, to automatically generate the logging text. LoGenText ’s evaluation on 10 open source projects indicates that the approach is promising for automatic logging text generation and significantly outperforms the state-of-the-art approach. Furthermore, we extend LoGenText to LoGenText-Plus by incorporating the syntactic templates of the logging texts. Different from LoGenText , LoGenText-Plus decomposes the logging text generation process into two stages. LoGenText-Plus first adopts an NMT model to generate the syntactic template of the target logging text. Then LoGenText-Plus feeds the source code and the generated template as the input to another NMT model for logging text generation. We also evaluate LoGenText-Plus on the same 10 projects and observe that it outperforms LoGenText on 9 of them. According to a human evaluation from developers’ perspectives, the logging texts generated by LoGenText-Plus have a higher quality than those generated by LoGenText and the prior baseline approach. By manually examining the generated logging texts, we then identify five aspects that can serve as guidance for writing or generating good logging texts. Our work is an important step toward the automated generation of logging statements, which can potentially save developers’ efforts and improve the quality of software logging. Our findings shed light on research opportunities that leverage advances in NMT techniques for automated generation and suggestion of logging statements. Zishuo Ding, Yiming Tang 0002, Heng Li 0007, Weiyi Shang |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | PILAR: Studying and Mitigating the Influence of Configurations on Log ParsingabstractThe significance of logs has been widely acknowledged with the adoption of various log analysis techniques that assist in software engineering tasks. Many log analysis techniques require structured logs as input while raw logs are typically unstructured. Automated log parsing is proposed to convert unstructured raw logs into structured log templates. Some log parsers achieve promising accuracy, yet they rely on significant efforts from the users to tune the parameters to achieve optimal results. In this paper, we first conduct an empirical study to understand the influence of the configurable parameters of six state-of-the-art log parsers on their parsing results on three aspects: 1) varying the parameters while using the same dataset, 2) keeping the same parameters while using different datasets, and 3) using different samples from the same dataset. Our results show that all these parsers are sensitive to the parameters, posing challenges to their adoption in practice. To mitigate such challenges, we propose PILAR (Parameter Insensitive Log Parser), an entropy-based log parsing approach. We compare PILAR with the existing log parsers on the same three aspects and find that PILAR is the most parameter-insensitive one. In addition, PILAR achieves the second highest parsing accuracy and efficiency among all the state-of-the-art log parsers. This paper paves the road for easing the adoption of log analysis in software engineer practices. Hetong Dai, Yiming Tang 0002, Heng Li 0007, Weiyi Shang |
ICSE | 2 |
| 2023 | On the Temporal Relations between Logging and CodeabstractPrior work shows that misleading logging texts (i.e., the textual descriptions in logging statements) can be counterproductive for developers during their use of logs. One of the most important types of information provided by logs is the temporal information of the recorded system behavior. For example, a logging text may use a perfective aspect to describe a fact that an important system event has finished. Although prior work has performed extensive studies on automated logging suggestions, few of these studies investigate the temporal relations between logging and code. In this work, we make the first attempt to comprehensively study the temporal relations between logging and its corresponding source code. In particular, we focus on two types of temporal relations: (1) logical temporal relations, which can be inferred from the execution order between the logging statement and the corresponding source code; and (2) semantic temporal relations, which can be inferred based on the semantic meaning of the logging text. We first perform qualitative analyses to study these two types of logging-code temporal relations and the inconsistency between them. As a result, we derive rules to detect these two types of temporal relations and their inconsistencies. Based on these rules, we propose a tool named TempoLo to automatically detect the issues of temporal inconsistencies between logging and code. Through an evaluation of four projects, we find that TempoLo can effectively detect temporal inconsistencies with a small number of false positives. To gather developers' feedback on whether such inconsistencies are worth fixing, we report 15 detected instances from these projects to developers. 13 instances from three projects are confirmed and fixed, while two instances of the remaining project are pending at the time of this writing. Our work lays the foundation for describing temporal relations between logging and code and demonstrates the potential for a deeper understanding of the relationship between logging and code. Zishuo Ding, Yiming Tang 0002, Heng Li 0007, Weiyi Shang |
ICSE | 2 |
| 2023 | IoPV: On Inconsistent Option Performance VariationsabstractMaintaining a good performance of a software system is a primordial task when evolving a software system. The performance regression issues are among the dominant problems that large software systems face. In addition, these large systems tend to be highly configurable, which allows users to change the behaviour of these systems by simply altering the values of certain configuration options. However, such flexibility comes with a cost. Such software systems suffer throughout their evolution from what we refer to as “Inconsistent Option Performance Variation” (IoPV ). An IoPV indicates, for a given commit, that the performance regression or improvement of different values of the same configuration option is inconsistent compared to the prior commit. For instance, a new change might not suffer from any performance regression under the default configuration (i.e., when all the options are set to their default values), while altering one option’s value manifests a regression, which we refer to as a hidden regression as it is not manifested under the default configuration. Similarly, when developers improve the performance of their systems, performance regression might be manifested under a subset of the existing configurations. Unfortunately, such hidden regressions are harmful as they can go unseen to the production environment. In this paper, we first quantify how prevalent (in)consistent performance regression or improvement is among the values of an option. In particular, we study over 803 Hadoop and 502 Cassandra commits, for which we execute a total of 4,902 and 4,197 tests, respectively, amounting to 12,536 machine hours of testing. We observe that IoPV is a common problem that is difficult to manually predict. 69% and 93% of the Hadoop and Cassandra commits have at least one configuration that hides a performance regression. Worse, most of the commits have different options or tests leading to IoPV and hiding performance regressions. Therefore, we propose a prediction model that identifies whether a given combination of commit, test, and option (CTO) manifests an IoPV. Our evaluation for different models shows that random forest is the best performing classifier, with a median AUC of 0.91 and 0.82 for Hadoop and Cassandra, respectively. Our paper defines and provides scientific evidence about the IoPV problem and its prevalence, which can be explored by future work. In addition, we provide an initial machine learning model for predicting IoPV. Jinfu Chen 0002, Zishuo Ding, Yiming Tang 0002, Mohammed Sayagh, Heng Li 0007, Bram Adams, Weiyi Shang |
ESEC/SIGSOFT FSE | 3 |
| 2023 | Studying and Complementing the Use of Identifiers in LogsabstractLogs contain a large amount of curated run-time information about the process of a software. Modern software systems have become more complex and larger in scale. They are typically executed in parallel or distributively, resulting in interleaved software logs and making log analysis challenging. Despite extensive research on automated logging analysis, none to our knowledge focuses on the use of logs, and they rarely augment logs to help with simpler analysis. Software log IDs are unique identifiers that developers can use to group and filter log entries. However, we found that, on average, only 21% of logging statements produce IDs, which can lead to loss of information in the log file. We propose LTID, a static analysis approach on log IDs, to remediate the aforementioned issue by extracting a dependency relation between log statements from source code. We build a dependency graph using static analysis and compute the dominance relations of each logging statement. We then propagate IDs to logs that do not contain them based on the dependency graph. We studied 21 well-known Java open-source software subjects and were able to inject IDs on average into 12% of logs without IDs. Through an open coding process, we also establish a categorization, which has a Cohen’s Kappa agreement coefficient of 0.74, of the information gained to better understand the relations recovered by the ID propagation process. Jianchen Zhao, Yiming Tang 0002, Sneha Sunil, Weiyi Shang |
SANER | 2 |
| 2023 | Automated Generation and Evaluation of JMH Microbenchmark Suites From Unit TestsabstractPerformance is a crucial non-functional requirement of many software systems. Despite the widespread use of performance testing, developers still struggle to construct and evaluate the quality of performance tests. To address these two major challenges, we implement a framework, dubbedju2jmh, to automatically generate performance microbenchmarks from JUnit tests and use mutation testing to study the quality of generated microbenchmarks. Specifically, we compare ourju2jmhgenerated benchmarks to manually written JMH benchmarks and to automatically generated JMH benchmarks using the AutoJMH framework, as well as directly measuring system performance with JUnit tests. For this purpose, we have conducted a study on three subjects (Rxjava,Eclipse-collections, andZipkin) with$\sim$454Ksource lines of code(SLOC), 2,417 JMH benchmarks (including manually written and generated AutoJMH benchmarks) and 35,084 JUnit tests. Our results show that theju2jmhgenerated JMH benchmarks consistently outperform using the execution time and throughput of JUnit tests as a proxy of performance and JMH benchmarks automatically generated using the AutoJMH framework while being comparable to JMH benchmarks manually written by developers in terms of tests’ stability and ability to detect performance bugs. Nevertheless,ju2jmhbenchmarks are able to cover more of the software applications than manually written JMH benchmarks during the microbenchmark execution. Furthermore,ju2jmhbenchmarks are generated automatically, while manually written JMH benchmarks require many hours of hard work and attention; therefore our study can reduce developers’ effort to construct microbenchmarks. In addition, we identify three factors (too low test workload, unstable tests and limited mutant coverage) that affect a benchmark's ability to detect performance bugs. To the best of our knowledge, this is the first study aimed at assisting developers in fully automated microbenchmark creation and assessing microbenchmark quality for performance testing. Mostafa Jangali, Yiming Tang 0002, Niclas Alexandersson, Philipp Leitner 0001, Jinqiu Yang 0001, Weiyi Shang |
IEEE Trans. Software Eng. | 2 |
| 2022 | Studying logging practice in test code
Haonan Zhang 0006, Yiming Tang 0002, Maxime Lamothe, Heng Li 0007, Weiyi Shang |
Empir. Softw. Eng. | 2 |
| 2022 | Automated evolution of feature logging statement levels using Git histories and degree of interest
Yiming Tang 0002, Allan Spektor, Raffi Khatchadourian, Mehdi Bagherzadeh 0001 |
Sci. Comput. Program. | 1 |
| 2021 | An Empirical Study of Refactorings and Technical Debt in Machine Learning SystemsabstractMachine Learning (ML), including Deep Learning (DL), systems, i.e., those with ML capabilities, are pervasive in today's data-driven society. Such systems are complex; they are comprised of ML models and many subsystems that support learning processes. As with other complex systems, ML systems are prone to classic technical debt issues, especially when such systems are long-lived, but they also exhibit debt specific to these systems. Unfortunately, there is a gap of knowledge in how ML systems actually evolve and are maintained. In this paper, we fill this gap by studying refactorings, i.e., source-to-source semantics-preserving program transformations, performed in real-world, open-source software, and the technical debt issues they alleviate. We analyzed 26 projects, consisting of 4.2 MLOC, along with 327 manually examined code patches. The results indicate that developers refactor these systems for a variety of reasons, both specific and tangential to ML, some refactorings correspond to established technical debt categories, while others do not, and code duplication is a major cross-cutting theme that particularly involved ML configuration and model code, which was also the most refactored. We also introduce 14 and 7 new ML-specific refactorings and technical debt categories, respectively, and put forth several recommendations, best practices, and anti-patterns. The results can potentially assist practitioners, tool developers, and educators in facilitating long-term ML system usefulness. Yiming Tang 0002, Raffi Khatchadourian, Mehdi Bagherzadeh 0001, Rhia Singh, Ajani Stewart, Anita Raja |
ICSE | 1 |
| 2020 | An Empirical Study on the Use and Misuse of Java 8 StreamsabstractStreaming APIs allow for big data processing of native data structures by providing MapReduce-like operations over these structures. However, unlike traditional big data systems, these data structures typically reside in shared memory accessed by multiple cores. Although popular, this emerging hybrid paradigm opens the door to possibly detrimental behavior, such as thread contention and bugs related to non-execution and non-determinism. This study explores the use and misuse of a popular streaming API, namely, Java 8 Streams. The focus is on how developers decide whether or not to run these operations sequentially or in parallel and bugs both specific and tangential to this paradigm. Our study involved analyzing 34 Java projects and 5:53 million lines of code, along with 719 manually examined code patches. Various automated, including interprocedural static analysis, and manual methodologies were employed. The results indicate that streams are pervasive, parallelization is not widely used, and performance is a crosscutting concern that accounted for the majority of fixes. We also present coincidences that both confirm and contradict the results of related studies. The study advances our understanding of streams, as well as benefits practitioners, programming language and API designers, tool developers, and educators alike. Raffi Khatchadourian, Yiming Tang 0002, Mehdi Bagherzadeh 0001, Baishakhi Ray |
FASE | 2 |
| 2020 | Safe automated refactoring for intelligent parallelization of Java 8 streams
Raffi Khatchadourian, Yiming Tang 0002, Mehdi Bagherzadeh 0001 |
Sci. Comput. Program. | 2 |
| 2019 | Safe automated refactoring for intelligent parallelization of Java 8 streamsabstractStreaming APIs are becoming more pervasive in mainstream Object-Oriented programming languages. For example, the Stream API introduced in Java 8 allows for functional-like, MapReduce-style operations in processing both finite and infinite data structures. However, using this API efficiently involves subtle considerations like determining when it is best for stream operations to run in parallel, when running operations in parallel can be less efficient, and when it is safe to run in parallel due to possible lambda expression side-effects. In this paper, we present an automated refactoring approach that assists developers in writing efficient stream code in a semantics-preserving fashion. The approach, based on a novel data ordering and typestate analysis, consists of preconditions for automatically determining when it is safe and possibly advantageous to convert sequential streams to parallel and unorder or de-parallelize already parallel streams. The approach was implemented as a plug-in to the Eclipse IDE, uses the WALA and SAFE analysis frameworks, and was evaluated on 11 Java projects consisting of ?642K lines of code. We found that 57 of 157 candidate streams (36.31%) were refactorable, and an average speedup of 3.49 on performance tests was observed. The results indicate that the approach is useful in optimizing stream code to their full potential. Raffi Khatchadourian, Yiming Tang 0002, Mehdi Bagherzadeh 0001, Syed Ahmed |
ICSE | 2 |
| 2018 | [Engineering Paper] A Tool for Optimizing Java 8 Stream Software via Automated RefactoringabstractStreaming APIs are pervasive in mainstream Object-Oriented languages and platforms. For example, the Java 8 Stream API allows for functional-like, MapReduce-style operations in processing both finite, e.g., collections, and infinite data structures. However, using this API efficiently involves subtle considerations like determining when it is best for stream operations to run in parallel, when running operations in parallel can be less efficient, and when it is safe to run in parallel due to possible lambda expression side-effects. In this paper, we describe the engineering aspects of an open source automated refactoring tool called Optimize Streams that assists developers in writing optimal stream software in a semantics-preserving fashion. Based on a novel ordering and typestate analysis, the tool is implemented as a plug-in to the popular Eclipse IDE, using both the WALA and SAFE frameworks. The tool was evaluated on 11 Java projects consisting of ~642 thousand lines of code, where we found that 36.31% of candidate streams were refactorable, and an average speedup of 1.55 on a performance suite was observed. We also describe experiences gained from integrating three very different static analysis frameworks to provide developers with an easy-to-use interface for optimizing their stream code to its full potential. Raffi Khatchadourian, Yiming Tang 0002, Mehdi Bagherzadeh 0001, Syed Ahmed |
SCAM | 2 |