VLDB 2026 Research / reviewers in the wild / expert
Xiaoling Li 0002
dblp:74/4356-2
· DBLP profile ↗
16ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0002-9479-2541ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SoapFL: A Standard Operating Procedure for LLM-Based Method-Level Fault LocalizationabstractFault Localization (FL) is an essential step during the debugging process. With the strong capabilities of code comprehension, the recent Large Language Models (LLMs) have demonstrated promising performance in diagnosing bugs in the code. Nevertheless, due to LLMs’ limited performance in handling long contexts, existing LLM-based fault localization remains on localizing bugs within asmall code scope(i.e., a method or a class), which struggles to diagnose bugs for alarge code scope(i.e., an entire software system). To address the limitation, this paper presents SoapFL, which builds an LLM-driven standard operating procedure (SOP) to automatically localize buggy methods from the entire software. By simulating the behavior of a human developer, SoapFL models the FL task as a three-step process, which involves comprehension, navigation, and confirmation. Within specific steps, SoapFL provides useful test behavior or coverage information to LLM through program analysis. Particularly, we adopt a series of auxiliary strategies such as Test Behavior Tracking, Document-Guided Search, and Multi-Round Dialogue to overcome the challenges in each step. The evaluation on the widely used Defects4J-V1.2.0 benchmark shows that SoapFL can localize 175 out of 395 bugs within Top-1, which outperforms the other LLM-based approaches and exhibits complementarity to the state-of-the-art learning-based techniques. Additionally, we confirm the indispensability of the components in SoapFL with the ablation study and demonstrate the usability of SoapFL through a user study. Finally, the cost analysis shows that SoapFL spends an average of only 0.081 dollars and 92 seconds for a single bug. Yihao Qin, Shangwen Wang, Yiling Lou, Jinhao Dong, Xiaoling Li 0002, Xiaoguang Mao |
IEEE Trans. Software Eng. | 6 |
| 2024 | A Knowledge Graph Based Technology of Operating System Software Repository Evolution AnalysisabstractThe operating system software repository is a collection of software package resources that are used to build operating system distributions. It also serves as a platform for users to install and update the system software. Owing to the extensive array of software packages within the software repository, intricate dependency relationships among these packages, and the asynchronous nature of updates for different software packages, the evolution and upgrade of both the software repository and operating system version are challenging to predict. This presents significant hidden risks for version updates and ecological compatibility governance. To tackle this issue, the paper designs a knowledge graph for the operating system software repository. Additionally, it suggests a method for identifying changes and ensuring coherence within the software repository by utilizing the knowledge graph. We select typical open source operating system software repositories for evolutionary analysis. Results demonstrate that the proposed method, as compared to traditional dependency detection methods such as Boolean expressions, examines package dependencies and overall self-consistency of the repository from a macro perspective of the operating system. This approach effectively identifies and guides the improvement of deep inconsistent relationships. It offers a new tool and perspective for constructing, managing, and maintaining operating system software repositories. Jun Ma 0015, Xiaoling Li 0002, Xinran Hong, Jie Yu 0008, Shasha Li 0001 |
ISPA | 4 |
| 2024 | How to Pet a Two-Headed Snake? Solving Cross-Repository Compatibility Issues with HeraabstractMany programming languages and operating system communities maintain software repositories to build their own ecosystems. The repositories often provide management tools to help users using the packages. The tools are often, if not all the times, well-designed to handle intra-repository dependencies without considering inter-repository dependencies. The users, however, often need packages from different repositories, and thus may suffer from compatibility issues. We refer to these issues as Cross-repository Compatibility (CC) issues. Existing works typically focus on a single software repository and are insufficient to detect CC issues. Zhouyang Jia, Shanshan Li 0001, Ying Wang 0038, Jun Ma 0015, Xiaoling Li 0002, Xiangke Liao |
ASE | 6 |
| 2024 | Exploring nonintrusive measurements of spatio-temporal portrait of microservicesabstractAbstract As cloud native technology advances, the scale and complexity of applications built on microservice architecture continue to expand, leading to increasingly intricate differences between software within the same application. Microservice applications, offering high flexibility, are deployed in data centers as black boxes from the users' perspective, leaving them with no insight into the orchestration of cloud service providers. Consequently, users face challenges in promptly recognizing performance imbalances within their deployed applications. Meanwhile, cloud service providers may cut costs by offering a mix of qualified and unqualified services, potentially deceiving users. To enhance the understanding of microservice application organization, we propose a non‐intrusive measurement framework, termed NMPI. NMPI facilitates rapid identification of microservice application defects, offering insights into cloud services and detecting fraudulent behavior in microservice‐based applications. We model microservice applications using a queue analysis‐based approach and filter the dominant frequency components of average response time signals by employing k‐means on the fast fourier transform (FFT). Our model constructs a library of performance portraits for various software, with these portraits resembling human fingerprints that carry and mark the software's internal information. Utilizing a two‐tier microservices‐based application incorporating a database as a case study allows us to demonstrate the effectiveness of NMPI. Our experimental results show that NMPI can produce differentiable profiles of data service performance portraits across a diverse and extensive range of workloads, enabling the identification of software types and the analysis of performance conditions. Zichen Xu 0001, Dan Wu 0010, Xiaoling Li 0002, Biyong Liu, Haichuan Hu, Shuang Tan, Yusong Tan, Chenren Xu, Christopher Stewart, Qihe Zhou |
Softw. Pract. Exp. | 4 |
| 2019 | Multi-resource workload mapping with minimum cost in cloud environmentabstractSummary Workload mapping in cloud environment refers to map multiple workloads provided by the cloud users/tenants to the substrate network provided by the cloud providers, which is a NP‐hard problem. The workload is a service demand made to the cloud, which is modeled as a logical network consists of virtual nodes and virtual links. Substrate network is a physical network consists of physical nodes that are inter‐connected via communication links. Devising heuristic methods has become the mainstream of workload mapping problem, which can obtain a feasible solution, but the quality of the solution is not guaranteed. Pointing to this issue, this paper takes the mapping cost of the workloads as the solving objective, and models the workload mapping as a constraint optimization problem. Based on the constraint optimization model, we devise two algorithms to solve the problem. These algorithms can not only find the feasible solution, but also ensure the solution is optimal. Lastly, we have demonstrated the optimality of the proposed algorithms through theoretical proof and evaluated the performance of them through simulation experiment. Xiaoling Li 0002, Xiaoyong Li 0002, Yusong Tan, Shuang Tan |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | MicRun: A framework for scale-free graph algorithms on SIMD architecture of the Xeon PhiabstractGraph algorithms currently play increasingly important roles, especially in social networks and language modeling scenarios. Recently, accelerating graph algorithms by heterogeneous high performance computers with the integrated cores and expanded SIMD lanes has been becoming the mainstream. However, the existing methods, restricted by the low-efficiency grouping strategy and the non-optimized selection mechanism of tile size of a graph, are far below our expectations in many ways. Moreover, there are few convenient integrated tools provided for deploying the graph algorithms on MIC architecture. In this paper, we propose a high-efficiency framework MicRun, which is flexible to be used for graph algorithms on SIMD architecture of the Xeon Phi. There are two key components in MicRun, the Bucket Grouping module and Auto-tuning module. In the Grouping module, an optimization algorithm is designed for splitting graph tiles into conflict-free groups, which can be directly processed on SIMD parallelism. In the Auto-tuning module, a novel strategy is proposed for optimizing the tile size to boost execution efficiency of the graph computation. MicRun currently supports Bellman-Ford and PageRank algorithms, we also conduct extensive validation experiments on MicRun. Experimental results show that MicRun outperforms existing mechanisms in terms of storage and time overhead. As a consequence, both graph algorithms achieve an average speedup of 1.1× by MicRun, compared with the state-of-the-art. Qingbo Wu 0003, Yusong Tan, Jie Yu 0008, Qi Zhang 0028, Xiaoling Li 0002, Lei Luo 0002 |
ASAP | 6 |
| 2017 | Efficient skyline computation over distributed interval dataabstractSummary The increasing volume of uncertain data has resulted in a dire need for supporting efficient uncertain data management. The skyline query as an important aspect of data management has received considerable attention in recent years, because of its importance in making intelligent decisions over complex data. Moreover, data collection and storage have become increasingly distributed, which makes the central assembly of data for storage and query infeasible and inefficient. Although many research efforts have been conducted to address the skyline query problem in various distributed scenarios, we still lack algorithms to address the queries over interval data, which is a special kind of attribute‐level uncertain data that widely exists in many applications. In this paper, we extensively study the skyline query over distributed interval data. We model the skyline query problem and define the distributed skyline query over interval data. Particularly, 2 efficient algorithms are proposed to retrieve the skylines progressively from distributed local sites with a highly optimized feedback framework. Moreover, we exploit 2 strategies for further improving the queries. Extensive experiments on synthetic and real datasets with real deployment are conducted to validate the effectiveness and efficiency of our proposals. Xiaoyong Li 0002, Kaijun Ren, Xiaoling Li 0002, Jie Yu 0006 |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | A novel optimization scheme for caching in locality-aware P2P networksabstractDeploying cache has been generally adopted by Internet service providers (ISPs) to mitigate P2P traffic in recent years. Most traditional caching algorithms are designed for locality-unaware P2P networks, which mainly consider the requested frequency of contents as the principle of caching policies. However, in more prevalent locality-aware conditions with biased neighbor-selection policies, the existing caching schemes can hardly optimize the situation. In this paper we show that, what need to be cached in locality-aware conditions are the contents that can not be well provided by local neighbors, rather than the contents which are requested most frequently. Therefore, states of local neighbors should be taken into consideration in caching policies. We first present a new model in which P2P cache and locality-aware neighbor selection work together. We focus on inter-ISP traffic and available bandwidth of users in order to benefit both ISPs and users. Based on the mathematical model, a novel caching algorithm is proposed which considers replacement and allocation policies together. According to trace-driven simulations, the proposed algorithm outperforms other two representative caching algorithms in various scenarios. Shaoduo Gan, Jiexin Zhang 0001, Jie Yu 0008, Xiaoling Li 0002, Jun Ma 0015, Lei Luo 0002, Qingbo Wu 0003 |
ISCC | 4 |
| 2016 | ERPC: An Edge-Resources Based Framework to Reduce Bandwidth Cost in the Personal Cloud
Shaoduo Gan, Jie Yu 0008, Xiaoling Li 0002, Jun Ma 0015, Lei Luo 0002, Qingbo Wu 0003, Shasha Li 0001 |
WAIM (2) | 3 |
| 2015 | BLOR: An efficient bandwidth and latency sensitive overlay routing approach for flash data disseminationabstractSummary The problem of flash data dissemination refers to transmitting time‐critical data to a large group of distributed receivers in a timely manner, which widely exists in many mission‐critical applications and Web services. However, existing approaches for flash data dissemination fail to ensure the timely and efficient transmission, because of the unpredictability of the dissemination process. Overlay routing has been widely used as an efficient routing primitive for providing better end‐to‐end routing quality by detouring inefficient routing paths in the real networks. To improve the predictability of the flash data dissemination process, we propose a bandwidth and latency sensitive overlay routing approach named BLOR, by optimizing the overlay routing and avoiding inefficient paths in flash data dissemination. BLOR tries to select optimal routing paths in terms of network latency, bandwidth capacity, and available bandwidth in nature, which has never been studied before. Additionally, a location‐aware unstructured overlay topology construction algorithm, an unbiased top‐kdominance model, and an efficient semi‐distributed information management strategy are proposed to assist the routing optimization of BLOR. Extensive experiments have been conducted to verify the effectiveness and efficiency of the proposals with real‐world data sets. Copyright © 2014 John Wiley & Sons, Ltd. Xiaoyong Li 0002, Yijie Wang 0001, Yongquan Fu, Xiaoling Li 0002 |
Concurr. Comput. Pract. Exp. | 4 |
| 2014 | MABP: an optimal resource allocation approach in data center networks
Xiaoling Li 0002, Huaimin Wang 0001, Bo Ding 0001, Xiaoyong Li 0002 |
Sci. China Inf. Sci. | 1 |
| 2014 | Resource allocation with multi-factor node ranking in data center networks
Xiaoling Li 0002, Huaimin Wang 0001, Bo Ding 0001, Xiaoyong Li 0002 |
Future Gener. Comput. Syst. | 1 |
| 2014 | Parallelizing skyline queries over uncertain data streams with sliding window partitioning and grid index
Xiaoyong Li 0002, Yijie Wang 0001, Xiaoling Li 0002 |
Knowl. Inf. Syst. | 3 |
| 2013 | Parallelizing Probabilistic Streaming Skyline Operator in Cloud Computing EnvironmentsabstractThe skyline query processing over uncertain data streams has received considerable attention, due to its importance in helping users make intelligent decisions over complex data. Nevertheless, existing studies only focus on retrieving the skylines over data streams in a centralized environment typically with one processor, which limits the scalability of algorithms and cannot meet the requirement for massive data analysis. The emerging cloud computing environment provides much more reliable and stable environments than the traditional distributed environments, which can be well adapted to the massive data management and complex queries. Unfortunately, existing parallel frameworks in clouds such as MapReduce and its variants are not suitable for the skyline queries over uncertain data streams. In this paper, we propose a general framework for parallelizing the probabilistic streaming skyline operator with the sliding window partitioning. Particularly, we propose four items mapping strategies CMS, AMS, DMS and APS to optimize the queries based on the proposed parallel framework. Extensive experiments with real deployment are conducted to demonstrate the effectiveness and efficiency of the proposals. Xiaoyong Li 0002, Yijie Wang 0001, Xiaoling Li 0002, Rubing Huang |
COMPSAC | 3 |
| 2013 | A survey of queries over uncertain data
Yijie Wang 0001, Xiaoyong Li 0002, Xiaoling Li 0002 |
Knowl. Inf. Syst. | 3 |
| 2012 | Topology awareness algorithm for virtual network mappingabstractNetwork virtualization is recognized as an effective way to overcome the ossification of the Internet. However, the virtual network mapping problem (VNMP) is a critical challenge, focusing on how to map the virtual networks to the substrate network with efficient utilization of infrastructure resources. The problem can be divided into two phases: node mapping phase and link mapping phase. In the node mapping phase, the existing algorithms usually map those virtual nodes with a complete greedy strategy, without considering the topology among these virtual nodes, resulting in too long substrate paths (with multiple hops). Addressing this problem, we propose a topology awareness mapping algorithm, which considers the topology among these virtual nodes. In the link mapping phase, the new algorithm adopts the k -shortest path algorithm. Simulation results show that the new algorithm greatly increases the long-term average revenue, the acceptance ratio, and the long-term revenue-to-cost ratio ( R/C ). Xiaoling Li 0002, Huaimin Wang 0001, Changguo Guo, Bo Ding 0001, Xiaoyong Li 0002, Wen-qi Bi, Shuang Tan |
J. Zhejiang Univ. Sci. C | 1 |