Yuan Tian 0008

dblp:39/5423-8 · DBLP profile ↗
← Back
5ranked-venue papers in the field
1as first author
4since 2021 · last 2025
0000-0002-2208-3893ORCID · conflict

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 4 (1 first)Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2025 Understanding Abandonment and Slowdown Dynamics in the Maven Ecosystem
abstract
The sustainability of libraries is critical for modern software development, yet many libraries face abandonment, posing significant risks to dependent projects. This study explores the prevalence and patterns of library abandonment in the Maven ecosystem. We investigate abandonment trends over the past decade, revealing that approximately one in four libraries fail to survive beyond their creation year. We also analyze the release activities of libraries, focusing on their lifespan and release speed, and analyze the evolution of these metrics within the lifespan of libraries. We find that while slow release speed and relatively long periods of inactivity are often precursors to abandonment, some abandoned libraries exhibit bursts of high frequent release activity late in their life cycle. Our findings contribute to a new understanding of library abandonment dynamics and offer insights for practitioners to identify and mitigate risks in software ecosystems.
Kazi Amit Hasan, Jerin Yasmin, Huizi Hao, Yuan Tian 0008, Safwat Hassan, Steven H. H. Ding
MSR4
2024 PeaTMOSS: A Dataset and Initial Analysis of Pre-Trained Models in Open-Source Software
abstract
The development and training of deep learning models have become increasingly costly and complex. Consequently, software engineers are adopting pre-trained models (PTMs) for their downstream applications. The dynamics of the PTM supply chain remain largely unexplored, signaling a clear need for structured datasets that document not only the metadata but also the subsequent applications of these models. Without such data, the MSR community cannot comprehensively understand the impact of PTM adoption and reuse.
Wenxin Jiang 0001, Jerin Yasmin, Jason Jones, Nicholas Synovic, Jiashen Kuo, Nathaniel Bielanski, Yuan Tian 0008, George K. Thiruvathukal, James C. Davis 0001
MSR7
2023 Understanding the Time to First Response in GitHub Pull Requests
abstract
The pull-based development is widely adopted in modern open-source software (OSS) projects, where developers propose changes to the codebase by submitting a pull request (PR). However, due to many reasons, PRs in OSS projects frequently experience delays across their lifespan, including prolonged waiting times for the first response. Such delays may significantly impact the efficiency and productivity of the development process, as well as the retention of new contributors as long-term contributors.In this paper, we conduct an exploratory study on the time-to-first-response for PRs by analyzing 111,094 closed PRs from ten popular OSS projects on GitHub. We find that bots frequently generate the first response in a PR, and significant differences exist in the timing of bot-generated versus human-generated first responses. We then perform an empirical study to examine the characteristics of bot- and human-generated first responses, including their relationship with the PR’s lifetime. Our results suggest that the presence of bots is an important factor contributing to the time-to-first-response in the pull-based development paradigm, and hence should be separately analyzed from human responses. We also report the characteristics of PRs that are more likely to experience long waiting for the first human-generated response. Our findings have practical implications for newcomers to understand the factors contributing to delays in their PRs.
Kazi Amit Hasan, Marcos Macedo, Yuan Tian 0008, Bram Adams, Steven H. H. Ding
MSR3
2021 Stock Price Prediction Leveraging Reddit: The Role of Trust Filter and Sliding Window
abstract
The increasing attraction of high-risk investments for retail investors has fuelled speculation across social media platforms on stocks with a perceived likelihood for disproportionate returns. This paper investigates the capacity of one such forum’s (Reddit’s r/WallStreetBets) speculation and the prevailing literature on how value should be maximally extracted from a community’s insights. With comparative research focused on aggregating the sentiment of a community – we took an individualized approach – and sought to understand the role of forum member posts in stock price predictions. Utilizing a sentiment sliding window and trust filter enabled a 40% reduction in errors compared to not applying these techniques with the financial and Reddit data.
Dennis Huynh, Garrett Audet, Nikolay Alabi, Yuan Tian 0008
IEEE BigData4
2012 What does software engineering community microblog about?
abstract
Microblogging is a new trend to communicate and to disseminate information. One microblog post could potentially reach millions of users. Millions of microblogs are generated on a daily basis on popular sites such as Twitter. The popularity of microblogging among programmers, software engineers, and software users has also led to their use of microblogs to communicate software engineering issues apart from using emails and other traditional communication channels. Understanding how millions of users use microblogs in software engineering related activities would shed light on ways we could leverage the fast evolving microblogging content to aid software development efforts. In this work, we perform a preliminary study on what the software engineering community microblogs about. We analyze the content of microblogs from Twitter and categorize the types of microblogs that are posted. We investigate the relative popularity of each category of microblogs. We also investigate what kinds of microblogs are diffused more widely in the Twitter network via the “retweet” feature. Our experiments show that microblogs commonly contain job openings, news, questions and answers, or links to download new tools and code. We find that microblogs concerning real-world events are more widely diffused in the Twitter network.
Yuan Tian 0008, Palakorn Achananuparp, Nelman Lubis Ibrahim, David Lo 0001, Ee-Peng Lim
MSR1