Xianhao Jin

dblp:223/4525 · DBLP profile ↗
← Back
9ranked-venue papers
9as first author
6since 2021 · last 2025
0000-0001-8536-1523ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 9 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
YearPublicationVenuePosition
2025 How Can Infrastructure as Code Accelerate Data Center Bring-ups? A Case Study at ByteDance
abstract
Software companies are establishing new data centers to enhance software performance with lower response times as well as meet security requirements for storing local user data. ByteDance also has a strong need to deploy its services to new data centers worldwide quickly and with minimal error. Unfortunately, this process can also be time-consuming and error-prone, e.g., services often have dependencies on one another, requiring a strict deployment execution order. Moreover, the high similarity between resources in the new and the old data centers enables minimal modifications and maximizes the reuse of existing infrastructure configurations while providing the ability of global resource management. Additionally, manual migration across data centers often requires multiple confirmation check-points, which can significantly slow down the process. Therefore, to accelerate the new data center bring-ups, we adopt the idea of Infrastructure as code (IaC) which is a practice to automatically configure system dependencies and to provision local and remote instances [41]. In this work, we propose ByteRollout, an automatic intent-based software resource deployment system that is able to take customized infrastructure configurations as input and actuate deployments accordingly. We also assess ByteRollout in the following events of new data centers creations driven by site reliability engineering (SRE) teams. The evaluation results demonstrate that ByteRollout significantly accelerates the data center bring-up process, saving months of human effort and reducing costs by millions of USD simultaneously. The results also highlight that the Infrastructure as Code practice can be leveraged in the context of data center setups, offering benefits such as reduced process time and minimized errors throughout the deployment.
Xianhao Jin, Yongning Hu, Luchuan Guo
ASE1
2024 PIPELINEASCODE: A CI/CD Workflow Management System through Configuration Files at ByteDance
abstract
Continuous Integration (CI) and Continuous Deployment (CD) are widely used practices in modern software engineering. Unfortunately, it is also an expensive and complicated practice - setting up a CI/CD pipeline can be time-consuming and error-prone. CI/CD pipelines at ByteDance can be set up and modified through the user interface by dragging each atom onto the screen to maximize its customizability. However, these operations increase the difficulty of reuse and version tracking. In this work, we design PIPELINEASCODE as a CI/CD workflow management system using configuration files to provide high portability and traceability. PIPELINEASCODE is seamlessly integrated with the CI/CD system at ByteDance which is named as “ByteCycle”. We also assess PIPELINEASCODE in two dimensions to determine its strengths by addressing two key questions: what types of pipelines developers prefer to manage using PIPELINEASCODE, and how pipelines evolve after the adoption of PIPELINEASCODE. The evaluation results show that pipelines managed by PIPELINEASCODE generally have a higher build frequency, fewer steps, a higher change frequency, a lower success rate and a longer build duration. After the adoption of PIPELINEASCODE, pipelines tend to have fewer steps, a higher build frequency, a longer build duration, a higher success rate and a higher change frequency. The results show that PIPELINEASCODE is capable of improving the reliability and flexibility of CI/CD pipelines and thus encourages users to build and deploy more frequently.
Xianhao Jin, Yongning Hu, Luchuan Guo
SANER1
2023 HybridCISave: A Combined Build and Test Selection Approach in Continuous Integration
abstract
Continuous Integration (CI) is a popular practice in modern software engineering. Unfortunately, it is also a high-cost practice—Google and Mozilla estimate their CI systems in millions of dollars. To reduce the computational cost in CI, researchers developed approaches to selectively execute builds or tests that are likely to fail (and skip those likely to pass). In this article, we present a novel hybrid technique ( HybridCISave ) to improve on the limitations of existing techniques: to provide higher cost savings and higher safety. To provide higher cost savings, HybridCISave combines techniques to predict and skip executions of both full builds that are predicted to pass and partial ones (only the tests in them predicted to pass). To provide higher safety, HybridCISave combines the predictions of multiple techniques to obtain stronger certainty before it decides to skip a build or test. We evaluated HybridCISave by comparing its effectiveness with the existing build selection techniques over 100 projects and found that it provided higher cost savings at the highest safety. We also evaluated each design decision in HybridCISave and found that skipping both full and partial builds increased its cost savings and that combining multiple test selection techniques made it safer.
Xianhao Jin, Francisco Servant
ACM Trans. Softw. Eng. Methodol.1
2022 Which builds are really safe to skip? Maximizing failure observation for build selection in continuous integration
abstract
Continuous integration (CI) is a widely used practice in modern software engineering.Unfortunately, it is also an expensive practice.Google and Mozilla estimate their expenses for their CI systems in millions of dollars.To reduce the cost of CI, researchers developed multiple approaches to reduce its computational workload requirements.However, these approaches sometimes make mispredictions and skip failing builds which are not desirable to be skipped.Thus, in this paper, we aim to save computational cost in CI, while also maximizing the observation of failing builds, i.e., to skip builds more safely.First, we perform empirical studies to understand which builds are safe to skip, starting from CI-Skip rules [1] that characterize builds that developers decide to skip.We observe that CI-Skip rules are not so safe as expected.We then develop a collection of CI-Run rules that can complement these rules.Based on our findings, we propose PreciseBuildSkip, a novel approach that maximizes build failure observation and reduces the cost of CI through the strategy of build selection.We evaluate our approach and results show that our approach saved more cost (5.5%)than the safest existing technique but reduced the falsely skipped failing builds from 4.1% to 0% (median value).
Xianhao Jin, Francisco Servant
J. Syst. Softw.1
2021 What helped, and what did not? An Evaluation of the Strategies to Improve Continuous Integration
abstract
Continuous integration (CI) is a widely used practice in modern software engineering. Unfortunately, it is also an expensive practice - Google and Mozilla estimate their CI systems in millions of dollars. There are a number of techniques and tools designed to or having the potential to save the cost of CI or expand its benefit - reducing time to feedback. However, their benefits in some dimensions may also result in drawbacks in others. They may also be beneficial in other scenarios where they are not designed to help. In this paper, we perform the first exhaustive comparison of techniques to improve CI, evaluating 14 variants of 10 techniques using selection and prioritization strategies on build and test granularity. We evaluate their strengths and weaknesses with 10 different cost and time-tofeedback saving metrics on 100 real-world projects. We analyze the results of all techniques to understand the design decisions that helped different dimensions of benefit. We also synthesized those results to lay out a series of recommendations for the development of future research techniques to advance this area.
Xianhao Jin, Francisco Servant
ICSE1
2021 Reducing cost in continuous integration with a collection of build selection approaches
abstract
Continuous integration (CI) is a widely used practice in modern software engineering. Unfortunately, it is also an expensive practice — Google and Mozilla estimate their CI systems in millions of dollars. To reduce CI computation cost, I propose the strategy of build selection to selectively execute those builds whose outcomes are failing and skip those passing builds for cost-saving. In my research, I firstly designed SmartBuildSkip as my first build selection approach that can skip unfruitful builds in CI automatically. Next, I evaluated SmartBuildSkip with all CI-improving approaches for understanding the strength and weakness of existing approaches to recommend future technique design. Then I proposed PreciseBuildSkip as a build selection approach to maximize the safety of skipping builds in CI. I also combined existing approaches both within and across granularity to be applied as a new build selection approach — HybridBuildSkip to save builds in a hybrid way. Finally, I plan to propose a human study to understand how to increase developers' trust on build selection approaches.
Xianhao Jin
ESEC/SIGSOFT FSE1
2020 A cost-efficient approach to building in continuous integration
abstract
Continuous integration (CI) is a widely used practice in modern software engineering. Unfortunately, it is also an expensive practice --- Google and Mozilla estimate their CI systems in millions of dollars. In this paper, we propose a novel approach for reducing the cost of CI. The cost of CI lies in the computing power to run builds and its value mostly lies on letting developers find bugs early --- when their size is still small. Thus, we target reducing the number of builds that CI executes by still executing as many failing builds as early as possible. To achieve this goal, we propose SmartBuildSkip, a technique which predicts the first builds in a sequence of build failures and the remaining build failures separately. SmartBuildSkip is customizable, allowing developers to select different preferred trade-offs of saving many builds vs. observing build failures early. We evaluate the motivating hypothesis of SmartBuildSkip, its prediction power, and its cost savings in a realistic scenario. In its most conservative configuration, SmartBuildSkip saved a median 30% of builds by only incurring a median delay of 1 build in a median of 15% failing builds.
Xianhao Jin, Francisco Servant
ICSE1
2019 What edits are done on the highly answered questions in stack overflow?: an empirical study
abstract
Stack Overflow is the most-widely-used online question-and-answer platform for software developers to solve problems and communicate experience. Stack Overflow believes in the power of community editing, which means that one is able to edit questions without the changes going through peer review. Stack Overflow users may make edits to questions for a variety of reasons, among others, to improve the question and try to obtain more answers. However, to date the relationship between edit actions on questions and the number of answers that they collect is unknown. In this paper, we perform an empirical study on Stack Overflow to understand the relationship between edit actions and number of answers obtained in different dimensions from different attributes of the edited questions. We find that questions are more commonly edited by question owners, on bodies with relatively big changes before obtaining an accepted answer. However, edited questions that obtained more answers in a shorter time, were edited by other users rather than question owners, and their edits tended to be small, focused on titles and in adding addendums.
Xianhao Jin, Francisco Servant
MSR1
2018 The hidden cost of code completion: understanding the impact of the recommendation-list length on its efficiency
abstract
Automatic code completion is a useful and popular technique that software developers use to write code more effectively and efficiently. However, while the benefits of code completion are clear, its cost is yet not well understood. We hypothesize the existence of a hidden cost of code completion, which mostly impacts developers when code completion techniques produce long recommendations. We study this hidden cost of code completion by evaluating how the length of the recommendation list affects other factors that may cause inefficiencies in the process. We study how common long recommendations are, whether they often provide low-ranked correct items, whether they incur longer time to be assessed, and whether they were more prevalent when developers did not select any item in the list. In our study, we observe evidence for all these factors, confirming the existence of a hidden cost of code completion.
Xianhao Jin, Francisco Servant
MSR1