VLDB 2026 Research / reviewers in the wild / expert
Ze Shi Li
dblp:247/5207
· DBLP profile ↗
14ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-2888-1025ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Integrating AI and Entrepreneurship: Preparing CS Undergraduates for Startups and InnovationabstractEntrepreneurship has become a defining force in the technology sector, with startups and unicorns largely fueled by computer science skills. However, most undergraduate computer science programs focus on technical competencies while outsourcing entrepreneurial education to business schools, where curricula are often generic and disconnected from the realities of CS-driven ventures. In addition, existing capstone and software engineering courses emphasize the technical details of development rather than entrepreneurial processes. This paper presents a curriculum that explicitly integrates artificial intelligence (AI) tools into entrepreneurship learning for computer science students. The program is designed to help students move beyond the notion that entrepreneurship is accidental or sudden and instead develop sustained exposure to ideation, customer discovery, prototyping, and pitching. Our project-based, one-year program equips students with entrepreneurial skills while leveraging AI for brainstorming, market validation, prototyping, and storytelling. We argue that, particularly in tight labor markets where evidence suggests graduates increasingly turn to entrepreneurship, there is an urgent need to embed CS-centered entrepreneurial education into undergraduate programs. Ze Shi Li, Sridhar Radhakrishnan |
SIGCSE (2) | 1 |
| 2025 | Objectifying the Subjective: Cognitive Biases in Topic InterpretationsabstractAbstract Interpretation of topics is crucial for their downstream applications. State-of-the-art evaluation measures of topic quality such as coherence and word intrusion do not measure how much a topic facilitates the exploration of a corpus. To design evaluation measures grounded on a task, and a population of users, we do user studies to understand how users interpret topics. We propose constructs of topic quality and ask users to assess them in the context of a topic and provide rationale behind evaluations. We use reflexive thematic analysis to identify themes of topic interpretations from rationales. Users interpret topics based on availability and representativeness heuristics rather than probability. We propose a theory of topic interpretation based on the anchoring-and-adjustment heuristic: users anchor on salient words and make semantic adjustments to arrive at an interpretation. Topic interpretation can be viewed as making a judgment under uncertainty by an ecologically rational user, and hence cognitive biases aware user models and evaluation frameworks are needed. Swapnil Hingmire, Ze Shi Li, Shiyu Zeng, Ahmed Musa Awon, Luiz Pedro Franciscatto Guerra, Neil A. Ernst |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | A Region-Aware Dual Latent State Mining Framework for Service Recommendation in Large-Scale Service Networks
Xiaohong Zhang 0002, Ze Shi Li, Meng Yan 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Feature Noise Resilient for QoS Prediction With Probabilistic Deep SupervisionabstractAccurate Quality of Service (QoS) prediction is essential for enhancing user satisfaction in web recommendation systems, yet existing prediction models often overlook feature noise, focusing predominantly on label noise. In this paper, we present the Probabilistic Deep Supervision Network (PDS-Net), a robust framework designed to effectively identify and mitigate feature noise, thereby improving QoS prediction accuracy. PDS-Net operates with adual-branch architecture: the main branch utilizes a decoder network to learn a Gaussian-based prior distribution from known features, while the second branch derives a posterior distribution based on true labels. A key innovation of PDS-Net is its condition-based noise recognition loss function, which enables precise identification of noisy features in objects (users or services). Once noisy features are identified, PDS-Net refines the feature's prior distribution, aligning it with the posterior distribution, and propagates this adjusted distribution to intermediate layers, effectively reducing noise interference. Extensive experiments conducted on two real-world QoS datasets demonstrate that PDS-Net consistently outperforms existing models, achieving an average improvement of 8.91% in MAE on Dataset D1 and 8.32% on Dataset D2 compared to the state-of-the-art. These results highlight PDS-Net's ability to accurately capture complex user-service relationships and handle feature noise, underscoring its robustness and versatility across diverse QoS prediction environments. Xiaohong Zhang 0002, Ze Shi Li, Sheng Huang 0001, Meng Yan 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | QoSBERT: An Uncertainty-Aware Approach Based on Pretrained Language Models for Service Quality PredictionabstractAccurate prediction of Quality of Service (QoS) metrics is fundamental for selecting and managing cloud-based services. Traditional QoS models rely on manual feature engineering and yield only point estimates, offering no insight into the confidence of their predictions. In this paper, we propose QoSBERT, the framework that reformulates QoS prediction as a semantic regression task based on pre-trained language models. Unlike previous approaches relying on sparse numerical features, QoSBERT automatically encodes user-service metadata into natural language descriptions, enabling deep semantic understanding. Furthermore, we integrate a Monte Carlo Dropout–based uncertainty estimation module, allowing for trustworthy and risk-aware service quality prediction, which is crucial yet underexplored in existing QoS models. QoSBERT encodes user-service metadata as natural language and leverages a pre-trained model to capture contextual semantics. It applies attentive pooling over the encoded embeddings and employs a lightweight regressor optimized to minimize prediction error. To quantify predictive confidence, Monte Carlo Dropout is applied at inference time. The resulting uncertainty estimates further support high-confidence sample selection, enhancing robustness in low-resource scenarios. On standard QoS benchmark datasets, QoSBERT achieves an average reduction of 11.7% in MAE and 6.7% in RMSE for response time prediction, and 6.9% in MAE for throughput prediction compared to the strongest baselines, while providing well-calibrated confidence intervals for robust and trustworthy service quality estimation. Our approach not only advances the accuracy of service quality prediction but also delivers reliable uncertainty quantification, paving the way for more trustworthy, datadriven service selection and optimization. Xiaohong Zhang 0002, Ze Shi Li, Meng Yan 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2024 | Unveiling the Life Cycle of User Feedback: Best Practices from Software PractitionersabstractUser feedback has grown in importance for organizations to improve software products. Prior studies focused primarily on feedback collection and reported a high-level overview of the processes, often overlooking how practitioners reason about, and act upon this feedback through a structured set of activities. In this work, we conducted an exploratory interview study with 40 practitioners from 32 organizations of various sizes and in several domains such as e-commerce, analytics, and gaming. Our findings indicate that organizations leverage many different user feedback sources. Social media emerged as a key category of feedback that is increasingly critical for many organizations. We found that organizations actively engage in a number of non-trivial activities to curate and act on user feedback, depending on its source. We synthesize these activities into a life cycle of managing user feedback. We also report on the best practices for managing user feedback that we distilled from responses of practitioners who felt that their organization effectively understood and addressed their users' feedback. We present actionable empirical results that organizations can leverage to increase their understanding of user perception and behavior for better products thus reducing user attrition. Ze Shi Li, Nowshin Nawar Arony, Kezia Devathasan, Manish Sihag, Neil A. Ernst, Daniela E. Damian |
ICSE | 1 |
| 2024 | "Do you have Time for a Quick Call?": Exploring Remote and Hybrid Requirements Engineering Practices and Challenges in IndustryabstractWith the onset of the COVID-19 pandemic and the ensuing shift away from co-located work arrangements towards working from home, practitioners encountered many collabora-tive and coordination challenges. While companies are slowly moving back towards in-person arrangements, many employers have permanently adopted fully remote and hybrid work modes. However, the work modes are more blurred, as hybrid work makes coordinating the requirements engineering practices that rely on rich interactions more challenging. Therefore, gaining more understanding of RE challenges and practices from prac-titioners transitioning to the new modes of work is imperative to identify insights that can be useful for organizations shifting to hybrid and remote work. In this paper, we use a mixed-methods approach to gain insights into remote and hybrid requirements engineering practices and challenges in the industry. Through interviews with 12 industry practitioners and a survey with 49 practitioners, we report on 7 adopted practices and 7 challenges encountered in these work arrangements. We found challenges such as organizing co-located tasks, lack of interpersonal con-nections, keeping everyone in the loop, and engagement barriers, which fall under coordination, communication and collaboration. To offset such challenges, we provide 20 recommendations based on our findings, such as proactive planning and using newer tools that support comprehensive tracking of important knowl-edge for requirements documentation. Our findings suggest that practitioners are facing challenges in remote and hybrid work arrangements, which they are mitigating with various strategies. Nonetheless, there remains a need for further research, as not all challenges are equally addressed across different work contexts. Ze Shi Li, Delina Ly, Lukas Nagel, Nowshin Nawar Arony, Daniela E. Damian |
RE | 1 |
| 2023 | Leveraging User Feedback for Requirements Through Trend and Narrative AnalysisabstractWith the rapid rise of new mediums and surge of user discussion regarding software products, it is increasingly important to consider these user concerns to fulfill user needs, otherwise users may opt for alternatives. Traditional requirements elicitation approaches relied on interviews and surveys with stakeholders to elicit key and important requirements. Despite the advances from recent studies to leverage the “crowd” via increased involvement of crowd based discussions, we still lack empirical structured guidance and approaches to handle feedback from multiple sources and synthesize the themes that emerge from these sources. As providing feedback becomes more accessible for users, managing the volume of feedback is correspondingly challenging. Since development resources are often limited, organizations need to make trade-offs between different user concerns. Gaining more insights on the themes and the trends of these themes from feedback should help organizations conduct these requirement trade-offs. To better explain the emergent trends of user feedback, I use the concept of narratives from economics to explain the phenomenon of what and when users change in their user discussions. Narrative analysis can help explain the causes of feedback trends, which can support organizations determining the priority and validity of various themes from feedback. In this work, I describe my preliminary findings which indicate the profound role that trends and narratives in user feedback have on user perception and concerns. My work has shown that feedback sources, like social media, can provide an avenue to identify requirements. Finally, I outline my plan towards developing a framework to identify trends and narratives, and a tool to support automated identification of requirements in the form of user stories. Ze Shi Li |
RE | 1 |
| 2023 | A Data-Driven Approach for Finding Requirements Relevant Feedback from TikTok and YouTubeabstractThe increasing importance of videos as a medium for engagement, communication, and content creation makes them critical for organizations to consider for user feedback. However, sifting through vast amounts of video content on social media platforms to extract requirements-relevant feedback is challenging. This study delves into the use of TikTok and YouTube, two widely used social media platforms that focus on video content, in identifying relevant user feedback that may be further refined into requirements using subsequent requirement generation steps. We demonstrate an approach of using videos as a source of user feedback by analyzing audio and visual text, and metadata (i.e., description/title) from 6276 videos of 20 popular products across various industries. We employed state-of-the-art deep learning transformer-based models, and classified 3097 videos consisting of requirements relevant information. We then clustered relevant videos and found multiple requirements relevant feedback themes for each of the 20 products. This feedback can later be refined into requirements artifacts. We found that product ratings (feature, design, performance), bug reports, and usage tutorial are persistent themes from the videos. Video-based social media such as TikTok and YouTube can provide valuable user insights, making them a powerful and novel resource for companies to improve customer-centric development. Manish Sihag, Ze Shi Li, Amanda Dash, Nowshin Nawar Arony, Kezia Devathasan, Neil A. Ernst, Alexandra Branzan Albu, Daniela E. Damian |
RE | 2 |
| 2022 | Narratives: the Unforeseen Influencer of Privacy ConcernsabstractPrivacy requirements are increasingly growing in importance as new privacy regulations are enacted. To adequately manage privacy requirements, organizations not only need to comply with privacy regulations, but also consider user privacy concerns. In this exploratory study, we used Reddit as a source to understand users’ privacy concerns regarding software applications. We collected 4.5 million posts from Reddit and classified 129075 privacy related posts, which is a non-negligible number of privacy discussions. Next, we clustered these posts and identified 9 main areas of privacy concerns. We use the concept of narratives from economics (i.e., posts that can go viral) to explain the phenomenon of what and when users change in their discussion of privacy. We further found that privacy discussions change over time and privacy regulatory events have a short term impact on such discussions. However, narratives have a notable impact on what and when users discussed about privacy. Considering narratives could guide software organizations in eliciting the relevant privacy concerns before developing them as privacy requirements. Ze Shi Li, Manish Sihag, Nowshin Nawar Arony, Joao Bezerra Junior, Thanh Phan, Neil A. Ernst, Daniela E. Damian |
RE | 1 |
| 2022 | Towards privacy compliance: A design science study in a small organization
Ze Shi Li, Colin M. Werner, Neil A. Ernst, Daniela E. Damian |
Inf. Softw. Technol. | 1 |
| 2022 | Uncovering the Benefits and Challenges of Continuous Integration PracticesabstractIn 2006, Fowler and Foemmel defined ten core Continuous Integration (CI) practices that could increase the speed of software development feedback cycles and improve software quality. Since then, these practices have been widely adopted by industry and subsequent research has shown they improve software quality. However, there is poor understanding ofhoworganizations implement these practices, of thebenefitsdevelopers perceive they bring, and of thechallengesdevelopers and organizations experience in implementing them. In this article, we discuss a multiple-case study of three small- to medium-sized companies using the recommended suite of ten CI practices. Using interviews and activity log mining, we learned that these practices are broadly implemented buthowthey are implemented varies depending on their perceived benefits, the context of the project, and the CI tools used by the organization. We also discovered that CI practices can create new constraints on the software process that hurt feedback cycle time. For researchers, we show that how CI is implemented varies, and thus studying CI (for example, using data mining) requires understanding these differences as important context for research studies. For practitioners, our findings reveal in-depth insights on the possible benefits and challenges from using the ten practices, and how project context matters. Omar Elazhary, Colin M. Werner, Ze Shi Li, Derek Lowlind, Neil A. Ernst, Margaret-Anne D. Storey |
IEEE Trans. Software Eng. | 3 |
| 2022 | Continuously Managing NFRs: Opportunities and Challenges in PracticeabstractNon-functional requirements (NFR), which include performance, availability, and maintainability, are vitally important to overall software quality. However, research has shown NFRs are, in practice, poorly defined and difficult to verify. Continuous software engineering practices, which extend agile practices, emphasize fast paced, automated, and rapid release of software that poses additional challenges to handling NFRs. In this multi-case study we empirically investigated how three organizations, for which NFRs are paramount to their business survival, manage NFRs in their continuous practices. We describe four practices these companies use to manage NFRs, such as offloading NFRs to cloud providers or the use of metrics and continuous monitoring, both of which enable almost real-time feedback on managing the NFRs. However, managing NFRs comes at a cost—as we also identified a number of challenges these organizations face while managing NFRs in their continuous software engineering practices. For example, the organizations in our study were able to realize an NFR by strategically and heavily investing in configuration management and infrastructure as code, in order to offload the responsibility of NFRs; however, this offloading implied potential loss of control. Our discussion and key research implications show the opportunities, trade-offs, and importance of the unique give-and-take relationship between continuous software engineering and NFRs. Research artifacts may be found athttps://doi.org/10.5281/zenodo.3376342. Colin M. Werner, Ze Shi Li, Derek Lowlind, Omar Elazhary, Neil A. Ernst, Daniela E. Damian |
IEEE Trans. Software Eng. | 2 |
| 2020 | The Lack of Shared Understanding of Non-Functional Requirements in Continuous Software Engineering: Accidental or Essential?abstractBuilding shared understanding of requirements is key to ensuring downstream software activities are efficient and effective. However, in continuous software engineering (CSE) some lack of shared understanding is an expected, and essential, part of a rapid feedback learning cycle. At the same time, there is a key trade-off with avoidable costs, such as rework, that come from accidental gaps in shared understanding. This tradeoff is even more challenging for non-functional requirements (NFRs), which have significant implications for product success. Comprehending and managing NFRs is especially difficult in small, agile organizations. How such organizations manage shared understanding of NFRs in CSE is understudied. We conducted a case study of three small organizations scaling up CSE to further understand and identify factors that contribute to lack of shared understanding of NFRs, and its relationship to rework. Our in-depth analysis identified 41 NFR-related software tasks as rework due to a lack of shared understanding of NFRs. Of these 41 tasks 78% were due to avoidable (accidental) lack of shared understanding of NFRs. Using a mixed-methods approach we identify factors that contribute to lack of shared understanding of NFRs, such as the lack of domain knowledge, rapid pace of change, and cross-organizational communication problems. We also identify recommended strategies to mitigate lack of shared understanding through more effective management of requirements knowledge in such organizations. We conclude by discussing the complex relationship between shared understanding of requirements, rework and, CSE. Colin M. Werner, Ze Shi Li, Neil A. Ernst, Daniela E. Damian |
RE | 2 |