EDBT 2026 Demo / reviewers in the wild / expert
Dayi Lin
dblp:187/9420
· DBLP profile ↗
20ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0002-4034-6650ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 16 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort EstimationabstractHigh-quality labeled datasets are crucial for training and evaluating foundation models in software engineering, but creating them is often prohibitively expensive and labor-intensive. We introduce SPICE, a scalable, automated pipeline for labeling SWE-bench-style datasets with annotations for issue clarity, test coverage, and effort estimation. SPICE combines context-aware code navigation, rationale-driven prompting, and multi-pass consensus to produce labels that closely approximate expert annotations. SPICE’s design was informed by our own experience and frustration in labeling more than 800 instances from SWE-Gym. SPICE achieves strong agreement with human-labeled SWE-bench Verified data while reducing the cost of labeling 1,000 instances from around $100,000 (manual annotation) to only $5.10. These results demonstrate SPICE’s potential to enable cost-effective, large-scale dataset creation for SE-focused FMs. To support the community, we release both SPICE tool and SPICE Bench, a new dataset of 6,802 SPICE-labeled instances curated from 291 open-source projects in SWE-Gym (over 13x larger than SWE-bench Verified). Gustavo Ansaldi Oliva, Gopi Krishnan Rajbahadur, Aaditya Bhatia, Haoxiang Zhang 0001, Zhilong Chen, Arthur Leung, Dayi Lin, Boyuan Chen 0002, Ahmed E. Hassan |
ASE | 8 |
| 2025 | Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered AgentsabstractLarge language models (LLMs) are increasingly integrated into autonomous systems, giving rise to a new class of software known as Agentware, where LLM-powered agents perform complex, open-ended tasks in domains such as software engineering, customer service, and data analysis. However, their high autonomy and opaque reasoning processes pose significant challenges for traditional software observability methods. To address this, we introduce the concept of cognitive observability—the ability to recover and inspect the implicit reasoning behind agent decisions. We present Watson, a general-purpose framework for observing the reasoning processes of fast-thinking LLM agents without altering their behavior. Watson retroactively infers reasoning traces using prompt attribution techniques. We evaluate Watson in both manual debugging and automated correction scenarios across the MMLU benchmark and the AutoCodeRover and OpenHands agents on the SWE-bench-lite dataset. In both static and dynamic settings, Watson surfaces actionable reasoning insights and supports targeted interventions, demonstrating its practical utility for improving transparency and reliability in Agentware systems. Benjamin Rombaut 0002, Sogol Masoumzadeh, Kirill Vasilevski, Dayi Lin, Ahmed E. Hassan |
ASE | 4 |
| 2025 | The Hitchhikers Guide to Production-ready Trustworthy Foundation Model Powered Software (FMware)abstractFoundation Models (FMs) such as Large Language Models (LLMs) are reshaping the software industry by enabling FMware, systems that integrate these FMs as core components.In this KDD 2025 tutorial, we present a comprehensive exploration of FMware that combines a curated catalogue of challenges with real-world production concerns.We first discuss the state of research and practice in building FMware.We further examine the difficulties in selecting suitable models, aligning high-quality domain-specific data, engineering robust prompts, and orchestrating autonomous agents.We then address the complex journey from impressive demos to production-ready systems by outlining issues in system testing, optimization, deployment, and integration with legacy software.Drawing on our industrial experience and recent research in the area, we provide actionable insights and a technology roadmap for overcoming these challenges.Attendees will gain practical strategies to enable the creation of trustworthy FMware in the evolving technology landscape. Kirill Vasilevski, Gopi Krishnan Rajbahadur, Gustavo Ansaldi Oliva, Benjamin Rombaut 0002, Keheliya Gallaba, Filipe Roseiro Côgo, Jiahuei Lin, Dayi Lin, Haoxiang Zhang 0001, Bouyan Chen, Kishanthan Thangarajah, Ahmed E. Hassan, Zhen Ming (Jack) Jiang |
KDD (2) | 8 |
| 2025 | A Framework and Taxonomy for Characterizing the Applicability of Software Architecture Recovery Approaches: A Tertiary-Mapping StudyabstractSummary Software architecture assists developers in addressing non‐functional requirements and in maintaining, debugging, and upgrading their software systems. Consequently, consistency between the designed architecture and the implemented software system itself is important; without this consistency the non‐functional requirements targeted may not be addressed and architectural documentation may mis‐direct maintenance efforts that target the associated code‐base. But often, when software is initially implemented or subsequently evolved, the designed architecture and software architecture become inconsistent, with the implemented structure degraded due to issues like developer time‐pressures, or ambiguous communication of the designed architecture. In such cases, Software Architecture Recovery (SAR) or consistency approaches can be applied to reconstruct the architecture of the software system and possibly to compare it to/re‐align it with the designed architecture. Many SAR approaches have been proposed in the research. However, choosing an appropriate architecture recovery approach for software systems is still an open issue. Consequently, this research aims to conduct a tertiary‐mapping study based on available secondary studies of architecture recovery approaches, to uncover important characteristics, towards the selection of appropriate SAR approaches. This research has aggregated 13 secondary studies and 10 primary studies beyond 2020 from 5 databases and, in doing so, identified 111 architecture recovery approaches. Based on these approaches, a taxonomy, containing nine main SAR‐selection categories is proposed and a framework (in the form of a supporting tool to help developers select an appropriate SAR approach) has been developed. Finally, this research identifies six potential open research gaps related to the underlying research that could be helpful for guiding research in the future. Abdul Qayum, Simon Colreavy-Donnelly, Muslim Chochlov, Jim Buckley, Dayi Lin, Ashish Rajendra Sai |
Softw. Pract. Exp. | 6 |
| 2025 | SimClone: Detecting Tabular Data Clones Using Value SimilarityabstractData clones are defined as multiple copies of the same data among datasets. The presence of data clones between datasets can cause issues such as difficulties in managing data assets and data license violations when using datasets with clones to build AI software. However, detecting data clones is not trivial. The majority of the prior studies in this area rely on structural information to detect data clones (e.g., font size, column header). However, tabular datasets used to build AI software are typically stored without any structural information. In this article, we propose a novel method called SimClone for data clone detection in tabular datasets without relying on structural information. SimClone method utilizes value similarities for data clone detection. We also propose a visualization approach as a part of our SimClone method to help locate the exact position of the cloned data between a dataset pair. Our results show that our SimClone outperforms the current state-of-the-art method by at least 20% in terms of both F1-score and AUC. In addition, SimClone’s visualization component helps identify the exact location of the data clone in a dataset with a Precision@10 value of 0.80 in the top 20 true positive predictions. Gopi Krishnan Rajbahadur, Dayi Lin, Shaowei Wang 0002, Zhen Ming (Jack) Jiang |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | MetaSel: A Test Selection Approach for Fine-Tuned DNN ModelsabstractDeep Neural Networks (DNNs) face challenges during deployment due to covariate shift, i.e., data distribution shifts between development and deployment contexts. Fine-tuning adapts pre-trained models to new contexts requiring smaller labeled sets. However, testing fine-tuned models under constrained labeling budgets remains a critical challenge. This paper introduces MetaSel, a new approach tailored for DNN models that have been fine-tuned to address covariate shift, to select tests from unlabeled inputs. MetaSel assumes that fine-tuned and pre-trained models share related data distributions and exhibit similar behaviors for many inputs. However, their behaviors diverge within the input subspace where fine-tuning alters decision boundaries, making those inputs more prone to misclassification. Unlike general approaches that rely solely on the DNN model and its input set, MetaSel leverages information from both the fine-tuned and pre-trained models and their behavioral differences to estimate misclassification probability for unlabeled test inputs, enabling more effective test selection. Our extensive empirical evaluation, comparing MetaSel against 11 state-of-the-art approaches and involving 68 fine-tuned models across weak, medium, and strong distribution shifts, demonstrates that MetaSel consistently delivers significant improvements in Test Relative Coverage (TRC) over existing baselines, particularly under highly constrained labeling budgets. MetaSel shows average TRC improvements of 28.46% to 56.18% over the most frequent second-best baselines while maintaining a high TRC median and low variability. Our results confirm MetaSel’s practicality, robustness, and cost-effectiveness for test selection in the context of fine-tuned models. Amin Abbasishahkoo, Mahboubeh Dadkhah, Lionel C. Briand, Dayi Lin |
IEEE Trans. Software Eng. | 4 |
| 2024 | TEASMA: A Practical Methodology for Test Adequacy Assessment of Deep Neural NetworksabstractSuccessful deployment of Deep Neural Networks (DNNs), particularly in safety-critical systems, requires their validation with an adequate test set to ensure a sufficient degree of confidence in test outcomes. Although well-established test adequacy assessment techniques from traditional software, such as mutation analysis and coverage criteria, have been adapted to DNNs in recent years, we still need to investigate their application within a comprehensive methodology for accurately predicting the fault detection ability of test sets and thus assessing their adequacy. In this paper, we propose and evaluateTEASMA, a comprehensive and practical methodology designed to accurately assess the adequacy of test sets for DNNs. In practice,TEASMAallows engineers to decide whether they can trust high-accuracy test results and thus validate the DNN before its deployment. Based on a DNN model's training set,TEASMAprovides a procedure to build accurate DNN-specific prediction models of the Fault Detection Rate (FDR) of a test set using an existing adequacy metric, thus enabling its assessment. We evaluatedTEASMAwith four state-of-the-art test adequacy metrics: Distance-based Surprise Coverage (DSC), Likelihood-based Surprise Coverage (LSC), Input Distribution Coverage (IDC), and Mutation Score (MS). We calculated MS based on mutation operators that directly modify the trained DNN model (i.e., post-training operators) due to their significant computational advantage compared to the operators that modify the DNN's training set or program (i.e., pre-training operators). Our extensive empirical evaluation, conducted across multiple DNN models and input sets, including large input sets such as ImageNet, reveals a strong linear correlation between the predicted and actual FDR values derived from MS, DSC, and IDC, with minimum$R^{2}$values of 0.94 for MS and 0.90 for DSC and IDC. Furthermore, a low average Root Mean Square Error (RMSE) of 9% between actual and predicted FDR values across all subjects, when relying on regression analysis and MS, demonstrates the latter's superior accuracy when compared to DSC and IDC, with RMSE values of 0.17 and 0.18, respectively. Overall, these results suggest thatTEASMAprovides a reliable basis for confidently deciding whether to trust test results for DNN models. Amin Abbasishahkoo, Mahboubeh Dadkhah, Lionel C. Briand, Dayi Lin |
IEEE Trans. Software Eng. | 4 |
| 2023 | Analyzing Gamer Complaints in Reviews of Cross-Platform Video Games on SteamabstractVideo gaming now represents the largest category in the entertainment industry in terms of revenue. To expand their market share, game developers are creating more cross-platform games, which are compatible with various platforms, including PCs, consoles, and smartphones. However, creating such games poses challenges as developers encounter platform-specific issues that may only surface on one of the target platforms. Consequently, many ported games fail due to careless adaptation from one exclusive platform to another. This paper presents the first empirical study on cross-platform issues by analyzing game users’ reviews for video games on both PC and game console(s). Our findings reveal that platform-related issues occur more frequently on the PC side, particularly for games that are ported from consoles. To address this challenge, we develop machine learning-based approaches to automatically identify and categorize reviews discussing platform-related issues, achieving a reasonable classification performance with 79.73% to 90.06% accuracy. Our approach would help cross-platform game developers save considerable time when analyzing user reviews. Hanwen Hu, Yuan Tian 0008, Safwat Hassan, Dayi Lin |
CoG | 4 |
| 2022 | Towards Training Reproducible Deep Learning ModelsabstractReproducibility is an increasing concern in Artificial Intelligence (AI), particularly in the area of Deep Learning (DL). Being able to reproduce DL models is crucial for AI-based systems, as it is closely tied to various tasks like training, testing, debugging, and auditing. However, DL models are challenging to be reproduced due to issues like randomness in the software (e.g., DL algorithms) and non-determinism in the hardware (e.g., GPU). There are various practices to mitigate some of the aforementioned issues. However, many of them are either too intrusive or can only work for a specific usage context. In this paper, we propose a systematic approach to training reproducible DL models. Our approach includes three main parts: (1) a set of general criteria to thoroughly evaluate the reproducibility of DL models for two different domains, (2) a unified framework which leverages a record-and-replay technique to mitigate software-related randomness and a profile-and-patch technique to control hardware-related non-determinism, and (3) a reproducibility guideline which explains the rationales and the mitigation strategies on conducting a reproducible training process for DL models. Case study results show our approach can successfully reproduce six open source and one commercial DL models. Boyuan Chen 0002, Mingzhi Wen, Yong Shi 0010, Dayi Lin, Gopi Krishnan Rajbahadur, Zhen Ming (Jack) Jiang |
ICSE | 4 |
| 2022 | What Causes Wrong Sentiment Classifications of Game Reviews?abstractSentiment analysis is a popular technique to identify the sentiment of a piece of text. Several different domains have been targeted by sentiment analysis research, such as Twitter, movie reviews, and mobile app reviews. Although several techniques have been proposed, the performance of current sentiment analysis techniques is still far from acceptable, mainly when applied in domains on which they were not trained. In addition, the causes of wrong classifications are not clear. In this article, we study how sentiment analysis performs on game reviews. We first report the results of a large-scale empirical study on the performance of widely used sentiment classifiers on game reviews. Then, we investigate the root causes for the wrong classifications and quantify the impact of each cause on the overall performance. We study three existing classifiers:Stanford CoreNLP,NLTK, andSentiStrength. Our results show that most classifiers do not perform well on game reviews, with the best one beingNLTK(with an AUC of 0.70). We also identified four main causes for wrong classifications, such as reviews that point out advantages and disadvantages of the game, which might confuse the classifier. The identified causes are not trivial to be resolved and we call upon sentiment analysis and game researchers and developers to prioritize a research agenda that investigates how the performance of sentiment analysis of game reviews can be improved, for instance by developing techniques that can automatically deal with specific game-related issues of reviews (e.g., reviews with advantages and disadvantages). Finally, we show that training sentiment classifiers on reviews that are stratified by the game genre is effective. Markos Viggiato, Dayi Lin, Abram Hindle, Cor-Paul Bezemer |
IEEE Trans. Games | 2 |
| 2022 | Towards a Consistent Interpretation of AIOps ModelsabstractArtificial Intelligence for IT Operations (AIOps) has been adopted in organizations in various tasks, including interpreting models to identify indicators of service failures. To avoid misleading practitioners, AIOps model interpretations should be consistent (i.e., different AIOps models on the same task agree with one another on feature importance). However, many AIOps studies violate established practices in the machine learning community when deriving interpretations, such as interpreting models with suboptimal performance, though the impact of such violations on the interpretation consistency has not been studied. In this article, we investigate the consistency of AIOps model interpretation along three dimensions: internal consistency, external consistency, and time consistency. We conduct a case study on two AIOps tasks: predicting Google cluster job failures and Backblaze hard drive failures. We find that the randomness from learners, hyperparameter tuning, and data sampling should be controlled to generate consistent interpretations. AIOps models with AUCs greater than 0.75 yield more consistent interpretation compared to low-performing models. Finally, AIOps models that are constructed with the Sliding Window or Full History approaches have the most consistent interpretation with the trends presented in the entire datasets. Our study provides valuable guidelines for practitioners to derive consistent AIOps model interpretation. Yingzhe Lyu, Gopi Krishnan Rajbahadur, Dayi Lin, Boyuan Chen 0002, Zhen Ming (Jack) Jiang |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2022 | The Impact of Data Merging on the Interpretation of Cross-Project Just-In-Time Defect ModelsabstractJust-In-Time (JIT) defect models are classification models that identify the code commits that are likely to introduce defects. Cross-project JIT models have been introduced to address the suboptimal performance of JIT models when historical data is limited. However, many studies built cross-project JIT models using a pool of mixed data from multiple projects (i.e., data merging)—assuming that the properties of defect-introducing commits of a project are similar to that of the other projects, which is likely not true. In this paper, we set out to investigate the interpretation of JIT defect models that are built from individual project data and a pool of mixed project data with and without consideration of project-level variances. Through a case study of 20 datasets of open source projects, we found that (1) the interpretation of JIT models that are built from individual projects varies among projects; and (2) the project-level variances cannot be captured by a JIT model that is trained from a pool of mixed data from multiple projects without considering project-level variances (i.e., a global JIT model). On the other hand, a mixed-effect JIT model that considers project-level variances represents the different interpretations better, without sacrificing performance, especially when the contexts of projects are considered. The results hold for different mixed-effect learning algorithms. When the goal is to derive sound interpretation of cross-project JIT models, we suggest that practitioners and researchers should opt to use a mixed-effect modelling approach that considers individual projects and contexts. Dayi Lin, Chakkrit Tantithamthavorn, Ahmed E. Hassan |
IEEE Trans. Software Eng. | 1 |
| 2021 | An Empirical Study of Trends of Popular Virtual Reality Games and Their ComplaintsabstractThe market for virtual reality (VR) games is growing rapidly and is expected to grow from 3.3 billion in 2018 to 13.7 billion in 2022. Due to the immersive nature of such games and the use of VR headsets, players may have complaints about VR games, which are distinct from those about traditional computer games, and an understanding of those complaints could enable developers to better take advantage of the growing VR market. We conduct an empirical study of 750 popular VR games and 17 635 user reviews on Steam in order to understand trends in VR games and their complaints. We find that the VR games market is maturing. Fewer VR games are released each month, but their quality appears to be improving over time. Most games support multiple headsets and play areas, and support for smaller scale play areas is increasing. Complaints of cybersickness are rare and declining, indicating that players are generally more concerned with other issues. Recently, complaints about game-specific issues have become the most frequent type of complaint, and VR game developers can now focus on these issues and worry less about VR-comfort issues such as cybersickness. Rain Epp, Dayi Lin, Cor-Paul Bezemer |
IEEE Trans. Games | 2 |
| 2020 | Building the perfect game - an empirical study of game modifications
Dayi Lin, Cor-Paul Bezemer, Ahmed E. Hassan |
Empir. Softw. Eng. | 2 |
| 2020 | An empirical study of the characteristics of popular Minecraft mods
Gopi Krishnan Rajbahadur, Dayi Lin, Mohammed Sayagh, Cor-Paul Bezemer, Ahmed E. Hassan |
Empir. Softw. Eng. | 3 |
| 2019 | Identifying gameplay videos that exhibit bugs in computer games
Dayi Lin, Cor-Paul Bezemer, Ahmed E. Hassan |
Empir. Softw. Eng. | 1 |
| 2019 | An empirical study of game reviews on the Steam platform
Dayi Lin, Cor-Paul Bezemer, Ying Zou 0001, Ahmed E. Hassan |
Empir. Softw. Eng. | 1 |
| 2018 | An empirical study of early access games on the steam platformabstract"Early access" is a release strategy for software that allows consumers to purchase an unfinished version of the software. In turn, consumers can influence the software development process by giving developers early feedback. This early access model has become increasingly popular through digital distribution platforms, such as Steam which is the most popular distribution platform for games. The plethora of options offered by Steam to communicate between developers and game players contribute to the popularity of the early access model. Dayi Lin, Cor-Paul Bezemer, Ahmed E. Hassan |
ICSE | 1 |
| 2018 | An empirical study of early access games on the Steam platform
Dayi Lin, Cor-Paul Bezemer, Ahmed E. Hassan |
Empir. Softw. Eng. | 1 |
| 2017 | Studying the urgent updates of popular games on the Steam platform
Dayi Lin, Cor-Paul Bezemer, Ahmed E. Hassan |
Empir. Softw. Eng. | 1 |