VLDB 2026 Research / reviewers in the wild / expert
Yi Liu 0014
dblp:97/4626-14
· DBLP profile ↗
13ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-1766-2444ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hallucination detection in LLM code generation: A sampling-based consensus verification approach
Taicheng Huang, Zhanhui Ren, Yuan Huang 0002, Xiangping Chen, Yi Liu 0014, Zibin Zheng |
Autom. Softw. Eng. | 5 |
| 2026 | LLM-Based Misconfiguration Detection for AWS Serverless ComputingabstractServerless computing is a popular cloud computing paradigm that enables developers to build applications at the function level, known as serverless applications. The Serverless Application Model (AWS SAM) is the most widely adopted configuration schema. However, misconfigurations pose a significant challenge due to the complexity of serverless configurations and the limitations of traditional data-driven techniques. Recent advancements in Large Language Models (LLMs), pre-trained on large-scale public data, offer promising potential for identifying and explaining misconfigurations. In this article, we present SlsDetector , the first framework that harnesses the capabilities of LLMs to perform static misconfiguration detection in serverless applications. SlsDetector utilizes effective prompt engineering with zero-shot prompting to identify configuration issues. It designs multi-dimensional constraints aligned with serverless configuration characteristics and leverages the Chain of Thought technique to enhance LLM inferences, alongside generating structured responses. We evaluate SlsDetector on a curated dataset of 110 configuration files, which includes correct configurations, real-world misconfigurations, and intentionally injected errors. Our results show that SlsDetector , based on ChatGPT-4o (one of the most representative LLMs), achieves a precision of 72.88%, recall of 88.18%, and F1-score of 79.75%, outperforming state-of-the-art data-driven methods by 53.82, 17.40, and 49.72 percentage points, respectively. We further investigate the generalization capability of SlsDetector across recent LLMs, including Llama 3.1 (405B) Instruct Turbo, Gemini 1.5 Pro, and DeepSeek V3, with consistently high effectiveness. Jinfeng Wen, Zhenpeng Chen 0001, Zixi Zhu, Federica Sarro, Yi Liu 0014, Haodi Ping, Shangguang Wang |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2025 | ServerlessIE: A Reusable Information Extraction System with Serverless FunctionabstractInformation extraction (IE) plays a pivotal role for data-driven applications. However, existing IE systems face challenges in poor generalization and re-usability. They perform well within specific domains or tasks but struggle to generalize to new domains or tasks. Meanwhile, it is difficult for users to reuse or replicate existing information extraction algorithms. In this paper, we present ServerlessIE, a reusable IE framework supporting multi-modal data (text, images, videos) via serverless functions. It unifies heterogeneous extraction algorithms and simplifies their replication across tasks. We also develop a deep learning-based function recommendation model to help users find appropriate functions for their tasks more conveniently. Experiments demonstrate the effectiveness of our approach in cross-modal extraction and adaptive function recommendation. Jingru Yang, Lingran Bu, Jinfeng Wen, Yi Liu 0014 |
ICWS | 5 |
| 2024 | MeDiC: Metasearch Service on Distributed Confidential DataabstractTraditional search engines aggregate vast amount of data on the Internet to provide keyword search services. However, in some privacy-sensitive fields like healthcare and e-government, the personal data often contains a substantial amount of private or sensitive information across different stakeholders’ data repositories. Due to the presence of private or sensitive information within the raw data or its metadata, as well as the lack of unified local search models, employing the approach with global data aggregation and indexing is not feasible. To solve this problem, we propose MeDiC, a metasearch service for discovering globally distributed confidential data. By distributing search requests through a unified interface, MeDiC offers a metasearch service to users without aggregating distributed confidential data. Experiments demonstrate that the MeDiC exhibits strong extensibility and user friendliness to enable developers to select different models, algorithms and search parameters based on specific search scenarios. Wenchun Jing, Jingru Yang, Yun Ma 0002, Yi Liu 0014, Chaoran Luo |
ICWS | 4 |
| 2023 | LegoDroid: flexible Android app decomposition and instant installation
Yi Liu 0014, Yun Ma 0002, Xusheng Xiao, Tao Xie 0001, Xuanzhe Liu |
Sci. China Inf. Sci. | 1 |
| 2023 | Characterizing commodity serverless computing platformsabstractAbstract Serverless computing has become a new trending paradigm in cloud computing, allowing developers to focus on the development of core application logic and rapidly construct the prototype via the composition of independent functions. With the development and prosperity of serverless computing, major cloud vendors have successively rolled out their commodity serverless computing platforms. However, the characteristics of these platforms have not been systematically studied. Measuring these characteristics can help developers to select the most adequate serverless computing platform and develop their serverless‐based applications in the right way. To fill this knowledge gap, we present a comprehensive study on characterizing mainstream commodity serverless computing platforms, including AWS Lambda, Google Cloud Functions, Azure Functions, and Alibaba Cloud Function Compute. Specifically, we conduct both qualitative analysis and quantitative analysis. In qualitative analysis, we compare these platforms from three aspects (i.e., development, deployment, and runtime) based on their official documentation to construct a taxonomy of characteristics. In quantitative analysis, we analyze the runtime performance of these platforms from multiple dimensions with well‐designed benchmarks. First, we analyze three key factors that can influence the startup latency of serverless‐based applications. Second, we compare the resource efficiency of different platforms with 16 representative benchmarks. Finally, we measure their performance difference when dealing with different concurrent requests and explore the potential causes in a black‐box fashion. Based on the results of both qualitative and quantitative analysis, we derive a series of findings and provide insightful implications for both developers and cloud vendors. Jinfeng Wen, Yi Liu 0014, Zhenpeng Chen 0001, Junkai Chen, Yun Ma 0002 |
J. Softw. Evol. Process. | 2 |
| 2023 | FaaSLight: General Application-level Cold-start Latency Optimization for Function-as-a-Service in Serverless ComputingabstractServerless computing is a popular cloud computing paradigm that frees developers from server management. Function-as-a-Service (FaaS) is the most popular implementation of serverless computing, representing applications as event-driven and stateless functions. However, existing studies report that functions of FaaS applications severely suffer from cold-start latency. In this article, we propose an approach, namely, FaaSLight , to accelerating the cold start for FaaS applications through application-level optimization. We first conduct a measurement study to investigate the possible root cause of the cold-start problem of FaaS. The result shows that application code loading latency is a significant overhead. Therefore, loading only indispensable code from FaaS applications can be an adequate solution. Based on this insight, we identify code related to application functionalities by constructing the function-level call graph and separate other code (i.e., optional code) from FaaS applications. The separated optional code can be loaded on demand to avoid the inaccurate identification of indispensable code causing application failure. In particular, a key principle guiding the design of FaaSLight is inherently general, i.e., platform - and language-agnostic . In practice, FaaSLight can be effectively applied to FaaS applications developed in different programming languages (Python and JavaScript), and can be seamlessly deployed on popular serverless platforms such as AWS Lambda and Google Cloud Functions, without having to modify the underlying OSes or hypervisors, nor introducing any additional manual engineering efforts to developers. The evaluation results on real-world FaaS applications show that FaaSLight can significantly reduce the code loading latency (up to 78.95%, 28.78% on average), thereby reducing the cold-start latency. As a result, the total response latency of functions can be decreased by up to 42.05% (19.21% on average). Compared with the state-of-the-art, FaaSLight achieves a 21.25× improvement in reducing the average total response latency. Xuanzhe Liu, Jinfeng Wen, Zhenpeng Chen 0001, Ding Li 0001, Junkai Chen, Yi Liu 0014, Haoyu Wang 0001, Xin Jin 0008 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2022 | Fission: Autonomous, Scalable Sharding for IoT BlockchainabstractIoT blockchain suffers heavy performance issues because of the massive transactions generated by various IoT nodes. By dividing nodes into different shards, sharding can produce blocks in parallel and hence improve the throughput of the blockchain system. Unlike the traditional blockchain system, IoT blockchain mainly consists of smart devices and the transactions are usually generated from the real world, such as the sensor data, photos taken by cameras, and so on. In IoT blockchain, closer nodes usually share a lower network latency and the transactions they generate are more related. Therefore, location-based sharding is an effective approach to improve the performance of IoT blockchain. Traditionally, IoT nodes are di-vided into different shards based on geographical locations or the connected edge server. However, the key challenge of sharding in IoT blockchain is how to guarantee the equality of shards division as to the unpredictable distribution and the dynamic behavior of IoT nodes. On one hand, shards can not be pre-divided because we can not predict the number or the distribution of the IoT nodes. On the other hand, nodes continuously joining or quitting shards will also break the equality of the shards division. In this paper, we propose Fission, a sharding mechanism designed for IoT blockchain. Fission divide shards based on the Voronoi diagram without any preknowledge about the nodes distribution, and support dynamic, autonomous sharding adjustment based on distributed Delaunay Triangulation. In addition, Fission uses a new diffusion-based consensus algorithm to achieve the linear scalability of throughput. The experimental results show that Fission can construct and adjust shards at a very low cost and can execute in a decentralized manner. The throughput can reach 1900tps in 500 nodes with only 5M bps bandwidth, and can scale linearly as the nodes increase. Chaoran Luo, Yueyang Hu, Ying Zhang 0012, Yi Liu 0014, Xingchun Diao, Gang Huang 0001 |
COMPSAC | 5 |
| 2021 | A Measurement Study on Serverless Workflow ServicesabstractMajor cloud providers increasingly roll out their serverless workflow services to orchestrate serverless functions, making it possible to construct complex applications effectively. A comprehensive study is necessary to help developers understand the pros and cons, and make better choices among these serverless workflow services. However, the characteristics of these serverless workflow services have not been systematically analyzed. To fill the knowledge gap, we conduct a comprehensive measurement study on four mainstream serverless workflow services, focusing on both features and the performance. First, we review their official documentation and extract their features from six dimensions, including programming model, state management, etc. Then, we compare their performance (i.e., the execution time of functions, execution time of workflows, orchestration overhead time of workflows) under various settings considering activity complexity and data-flow complexity of workflows, as well as function complexity of serverless functions. Our findings and implications could help developers and cloud providers improve their development efficiency and user experience. Jinfeng Wen, Yi Liu 0014 |
ICWS | 2 |
| 2021 | An empirical study on challenges of application development in serverless computingabstractServerless computing is an emerging paradigm for cloud computing, gaining traction in a wide range of applications such as video processing and machine learning. This new paradigm allows developers to focus on the development of the logic of serverless computing based applications (abbreviated as serverless-based applications) in the granularity of function, thereby freeing developers from tedious and error-prone infrastructure management. Meanwhile, it also introduces new challenges on the design, implementation, and deployment of serverless-based applications, and current serverless computing platforms are far away from satisfactory. However, to the best of our knowledge, these challenges have not been well studied. To fill this knowledge gap, this paper presents the first comprehensive study on understanding the challenges in developing serverless-based applications from the developers’ perspective. We mine and analyze 22,731 relevant questions from Stack Overflow (a popular Q&A website for developers), and show the increasing popularity trend and the high difficulty level of serverless computing for developers. Through manual inspection of 619 sampled questions, we construct a taxonomy of challenges that developers encounter, and report a series of findings and actionable implications. Stakeholders including application developers, researchers, and cloud providers can leverage these findings and implications to better understand and further explore the serverless computing paradigm. Jinfeng Wen, Zhenpeng Chen 0001, Yi Liu 0014, Yiling Lou, Yun Ma 0002, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
ESEC/SIGSOFT FSE | 3 |
| 2019 | A First Look at Instant Service Consumption with Quick Apps on Mobile DevicesabstractMobile app ecosystem has gained giant success in providing services on mobile devices to facilitate almost all aspects in our daily life. However, the whole-package installation and dramatically increasing package size are now preventing users from trying more apps. To address the issue, many lightweight frameworks have emerged, enabling to provide the experience of instant service consumption where apps are of small size and no installation is needed to consuming services provided by the apps. In this paper, we conduct the first empirical study on instant service consumption on mobile devices. We focus on one of the most popular frameworks, quick apps, which are proposed and supported by nine mainstream mobile phone manufacturers in China. Quick apps are implemented with Web-based technologies, and run as native apps without the need of installation. We find that quick apps have much smaller size and only provide a limited set of services compared to their corresponding native apps. Then, we characterize the performance differences between quick apps and native apps in terms of launching time, data drain, and network connections, when the two kinds of apps provide the same services. Our observations reveal that quick apps perform better than native apps thanks to its much smaller size and less functionalities in a single page. Finally, we propose a machine learning based approach to helping developers construct the quick app from an existing native app. Yi Liu 0014, Enze Xu, Yun Ma 0002, Xuanzhe Liu |
ICWS | 1 |
| 2018 | A Tale of Two Fashions: An Empirical Study on the Performance of Native Apps and Web Apps on Androidabstractprevalent smartphones have become the major entrance to accessing services on the Internet. On smartphones, users can have two options as the clients, i.e., native apps and Web apps. There have been several debates about native apps and Web apps. However, major service providers such as Google, Amazon, and Facebook provide both native apps and Web apps to end-users. Essentially, the performance differences between these two types of apps haven't been addressed. Indeed, the performance differences make non-trivial impacts on apps development, deployment, and distribution. In this article, we conduct a measurement study on the performance of native apps and Web apps on Android smartphones. Specifically, we want to explore given the same functionalities, do Web apps always perform poorly compared to native apps. We select 328 services from some popular providers, covering various domains such as e-commerce, map, social networking, and entertainment. With HTTP-level trace analysis, we demystify the workflows on how native apps and Web apps deliver services on mobile devices, respectively. Then, we characterize the performance differences between native apps and Web apps with the metrics including the number of requests, response time, data drain, and energy consumption. We find that the performance of Web apps is better than native apps in more than 31 percent cases. Our derived knowledge can suggest some recommendations to improve the performance for mobile apps. Yun Ma 0002, Xuanzhe Liu, Yi Liu 0014, Yunxin Liu 0001, Gang Huang 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2015 | Characterizing RESTful Web Services Usage on Smartphones: A Tale of Native Apps and Web AppsabstractThe burst of Web-based Restful services brings us a number of facilities in our life and work. We are used to take smartphones to access these Web services, like location-based services, weather search, mapping, social networking, et al. On smartphones, we have two options of service consumers, a.k.a, Native apps and Web apps. Despite the platform-independence, Web apps are claimed to provide the same features and comparable user experiences with native apps. However, one fact is that more and more people prefer native apps rather than Web apps. In this paper, we make an empirical study on characterizing the performance disparity of native apps and Web apps. Given the same functionalities provided by the same service providers, we explore the Restful Web services that are used by native apps and Web apps. With HTTP-level trace analysis, we demystify the workflows on how native apps and Web apps use Web services and summarize different service usage patterns from architectural style perspective. Then we characterize the performance differences between native apps and Web apps on realizing Restful Web services including GET, DELETE, PUT & POST, in terms of number of network connections, response time, and data drain, given the same functional features. Our observations reveal that Web apps do not always perform worse than native apps using Restful Web services under the same context. We further propose some implications to improve both native apps and Web apps on smartphones. Yi Liu 0014, Xuanzhe Liu, Yun Ma 0002, Yunxin Liu 0001, Zibin Zheng, Gang Huang 0001, M. Brian Blake |
ICWS | 1 |