VLDB 2026 Research / reviewers in the wild / expert
Tao Ma 0006
dblp:60/6411-6
· DBLP profile ↗
19ranked-venue papers
0as first author
14since 2021 · last 2025
0009-0008-8629-6274ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 since 2021Software engineering, systems software and programming languages · 7 · 7 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RPG: Linux Kernel Fuzzing Guided by Distribution-Specific Runtime Parameter InterfacesabstractThe Linux distribution kernel differs significantly from the mainline kernel, incorporating additional features and vendor-specific extensions. Among these additions, many runtime parameter interfaces are unique to distribution kernels, which expands the attack surface and increases the risk of potential vulnerabilities. Fuzzing has been used to assess Linux distributions, but existing tools cannot systematically test these distribution-specific interfaces due to two main challenges: (1) generating test cases for these runtime parameter interfaces, and (2) concentrating test resources on the distribution-specific interface code. To address these challenges, we propose RPG, a distribution-specific runtime parameter-guided kernel fuzzer. RPG operates in three phases: First, RPG extracts distribution-specific runtime parameter interfaces. Then, RPG uses LLM and tuning software databases to model each parameter range to generate meaningful interface test cases. Third, RPG utilizes the distribution kernel’s function control flow graph to guide the fuzzer to generate generic test cases that are more closely related to the distribution-specific interface code. We evaluated RPG on four Linux distribution kernels: Ubuntu 22.04, Fedora 42, OpenAnolis 8.8, and OpenAnolis 23.1. RPG detected 22 previously unknown bugs (13 distribution-specific), of which 15 were confirmed and 10 fixed by kernel maintainers. RPG also achieved 20.4% and 21.2% higher branch coverage than Syzkaller and Healer, respectively. Yuheng Shen, Guoyu Yin, Runzhe Wang, Tao Ma 0006, Xiaohai Shi, Heyuan Shi |
ASE | 6 |
| 2025 | LLM-assisted Industrial-Scale Differential Testing of Package Incompatibilities in Linux DistributionsabstractAn open source Linux distribution often undergoes version upgrades and migrations, which is prone to incompatibility issues especially when it comes to large-scale software changes. Although differential testing has been widely used in software testing, it is still challenging to apply it for detecting such incompatibilities in the context of industrial settings. In this paper, we report our experience in leveraging LLMs to address the challenges faced by the Linux distribution community. Specifically, we develop an LLM-based differential testing method called Versify to assist maintainers of Linux distributions in locating incompatibilities during version upgrades and migrations. Its trial operation period within the Linux distribution community shows that it uncovered 8,489 instances of differing behavior, of which 644 were prioritized for attention by developers. After deduplication and filtering, 39 unique compatibility reports were identified. Feedback from Linux distributions developers indicates that our reports have provided valuable recommendations for package selection in future OS releases. Chijin Zhou, Runzhe Wang, Weibo Zhang, Yuheng Shen, Xiaohai Shi, Tao Ma 0006, Zhe Wang 0015, Heyuan Shi |
ASE | 7 |
| 2025 | LogLabeler: Towards Effective Acquisition of Log Labels in Industrial Log-Based AnalysisabstractLog-based AIOps is a widely researched topic aiming at reducing the developer burden in system maintenance. Since industrial developers prefer lightweight supervised solutions for log-based AIOps, the strong dependence of these solutions on labeled data creates significant challenges for teams new to building log-based AIOps capabilities, such as high labeling costs, inconsistent annotations, and manual management issues. Log-labeling faces challenges in integrating existing artifacts to reduce labeling costs and manage labels effectively. To the best of our knowledge, no prior research addresses of assisting log-labeling problem. In this article, we propose a new approach called LogLabeler to assist developers in annotating and managing log labels. LogLabeler leverages existing artifacts for initial-label-acquisition, minimizes labeling costs by automatically generating all log labels, and shields developers from manual label management through a human-in-the-loop refinement approach. Evaluations on real-world datasets from Alibaba and open-source datasets show that LogLabeler can effectively supplement log labels, achieving comparable accuracy to existing baselines while operating more efficiently. Furthermore, we demonstrate LogLabeler's practical effectiveness at Alibaba through a case study, highlighting its benefits to developers. Zongyang Li, Qinglong Wang 0003, Shangming Cai, Zheng Liu 0022, Tao Ma 0006, Wei Yang 0013, Ying Li 0012, Tao Xie 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | TrEnv: Transparently Share Serverless Execution Environments Across Different Functions and NodesabstractServerless computing is renowned for its computation elasticity, yet its full potential is often constrained by the requirement for functions to operate within local and dedicated background environments, resulting in limited memory elasticity. To address this limitation, this paper introduces TrEnv, a co-designed integration of the serverless platform with the operating system and CXL/RDMA-based remote memory pools in two key areas. Firstly, TrEnv introduces repurposable sandboxes, which can be shared across different functions and hence, substantially decrease the overhead associated with creating isolation sandboxes. Secondly, it augments the OS with "memory templates" that enable rapid restoration of function states stored on remote memory. These innovations allow TrEnv to facilitate rapid transitions between instances of different functions and enable memory sharing across multiple nodes. Our evaluations using a variety of representative and real-world workloads demonstrate that TrEnv can initiate a container within 10 milliseconds, achieving up to a 7× speedup in P99 end-to-end latency and reducing memory usage by 48% on average compared to state-of-the-art on-demand restoring systems. Teng Ma 0006, Zheng Liu 0022, Sixing Lin, Kang Chen 0001, Jinlei Jiang, Xia Liao, Yingdi Shan, Mengting Lu, Tao Ma 0006, Haifeng Gong, Yongwei Wu 0001 |
SOSP | 12 |
| 2024 | Diagnosing Application-network Anomalies for Millions of IPs in Production Clouds
Zhe Wang 0015, Huanwu Hu, Linghe Kong, Xinlei Kang, Qiao Xiang, Peihao Yang, Jiejian Wu, Yong Yang 0013, Tao Ma 0006, Zheng Liu 0022, Xianlong Zeng, Dennis Cai, Guihai Chen |
USENIX ATC | 12 |
| 2024 | HydraRPC: RPC in the CXL Era
Teng Ma 0006, Zheng Liu 0022, Chengkun Wei, Youwei Zhuo, Yijin Guan, Dimin Niu, Tao Ma 0006 |
USENIX ATC | 11 |
| 2024 | Zero+: Monitoring Large-Scale Cloud-Native Infrastructure Using One-Sided RDMAabstractCloud services have shifted from monolithic designs to microservices running on cloud-native infrastructure with monitoring systems to ensure service level agreements (SLAs). However, traditional monitoring systems no longer meet the demands of cloud-native monitoring. In Alibaba’s “double eleven” shopping festival, it is observed that the monitor occupies resources of the monitored infrastructure and even disrupts services. In this paper, we propose a novel monitoring system named for cloud-native monitoring. achieves zero overhead in collecting raw metrics using one-sided remote direct memory access (RDMA) and remedies network congestion by adopting a receiver-driven flow control scheme. also features a priority queue mechanism to meet different quality of service requirements and an efficient batch processing design to relieve CPU occupation. has been deployed and evaluated in four different clusters with heterogeneous RDMA NIC devices and architectures in Alibaba Cloud. Results show that achieves no CPU occupation at the monitored host and supports$1\sim10k$hosts with$0.1\sim1s$sampling interval using a single thread for network I/O. significantly relieves the incast issue and maintains$80\sim95\%$of bandwidth utilization in several clusters when monitoring$1k$hosts. also ensures services with high priority accomplish collecting metrics earlier than low priority ones by at least$400 \mu s$when monitoring$1k$hosts. Jiejian Wu, Teng Ma 0006, Zhe Wang 0015, Linghe Kong, Zhenzao Wen, Yong Yang 0013, Tao Ma 0006, Zheng Liu 0022, Guihai Chen |
IEEE/ACM Trans. Netw. | 10 |
| 2023 | KeenTune: Automated Tuning Tool for Cloud Application Performance Testing and OptimizationabstractThe performance testing and optimization of cloud applications is challenging, because manual tuning of cloud computing stacks is tedious and automated tuning tools are rare used for cloud services. To address this issue, we introduce KeenTune, an automated tuning tool designed to optimize application performance and facilitate performance testing. KeenTune is a lightweight and flexible tool that can be deployed with to-be-tuned applications with negligible impact on their performance. Specifically, KeenTune uses a surrogate model that can be implemented with machine learning models to filter out less relevant parameters for efficient tuning. Our empirical evaluation shows that KeenTune significantly enhances the throughput performance of Nginx web servers, resulting in performance improvements of up to 90.43% and 117.23% in certain cases. This study highlights the benefits of using KeenTune for achieving efficient and effective performance testing of cloud applications. The video and source code for KeenTune are provided as supplementary materials. Qinglong Wang 0003, Runzhe Wang, Xiaohai Shi, Zheng Liu 0022, Tao Ma 0006, Houbing Song, Heyuan Shi |
ISSTA | 6 |
| 2023 | PVM: Efficient Shadow Paging for Deploying Secure Containers in Cloud-native EnvironmentabstractIn cloud-native environments, containers are often deployed within lightweight virtual machines (VMs) to ensure strong security isolation and privacy protection. With the growing demand for customized cloud services, third-party vendors are turning to infrastructure-as-a-service (IaaS) cloud providers to build their own cloud-native platforms, necessitating the need to run a VM or a guest that hosts containers inside another VM instance leased from an IaaS cloud. State-of-the-art nested virtualization in the x86 architecture relies heavily on the host hypervisor to expose hardware virtualization support to the guest hypervisor, not only complicating cloud management but also raising concerns about an increased attack surface at the host hypervisor. Hang Huang, Jiangshan Lai, Jia Rao, Hui Lu 0001, Wenlong Hou, Zhengyu He, Weidong Han 0003, Tao Ma 0006, Song Wu 0001 |
SOSP | 14 |
| 2023 | Partial Failure Resilient Memory Management System for (CXL-based) Distributed Shared MemoryabstractThe efficiency of distributed shared memory (DSM) has been greatly improved by recent hardware technologies. But, the difficulty of distributed memory management can still be a major obstacle to the democratization of DSM, especially when a partial failure of the participating clients (e.g., due to crashed processes or machines) should be tolerated. Teng Ma 0006, Jinqi Hua, Zheng Liu 0022, Kang Chen 0001, Fan Du, Jinlei Jiang, Tao Ma 0006, Yongwei Wu 0001 |
SOSP | 9 |
| 2023 | Kronos: towards bus contention-aware job scheduling in warehouse scale computers
Shang Zhao 0003, Quan Chen 0002, Shanpei Chen, Tao Ma 0006, Yong Yang 0013, Wenli Zheng, Minyi Guo |
Frontiers Comput. Sci. | 6 |
| 2023 | Async-fork: Mitigating Query Latency Spikes Incurred by the Fork-based Snapshot Mechanism from the OS LevelabstractIn-memory key-value stores (IMKVSes) serve many online applications. They generally adopt the fork-based snapshot mechanism to support data backup. However, this method can result in query latency spikes because the engine is out-of-service for queries during the snapshot. In contrast to existing research optimizing snapshot algorithms, we address the problem from the operating system (OS) level, while keeping the data persistent mechanism in IMKVSes unchanged. Specifically, we first study the impact of the fork operation on query latency. Based on findings in the study, we propose Async-fork, which performs the fork operation asynchronously to reduce the out-of-service time of the engine. Async-fork is implemented in the Linux kernel and deployed into the online Redis database in public clouds. Our experiment results show that Async-fork can significantly reduce the tail latency of queries during the snapshot. Pu Pang, Kaihao Bai, Quan Chen 0002, Shixuan Sun, Bo Liu 0122, Hongbo Yao, Zhengheng Wang, Zheng Liu 0022, Yong Yang 0013, Tao Ma 0006, Minyi Guo |
Proc. VLDB Endow. | 14 |
| 2022 | Help Rather Than Recycle: Alleviating Cold Startup in Serverless Computing Through Inter-Function Container Sharing
Zijun Li 0001, Linsong Guo, Quan Chen 0002, Jiagan Cheng, Chuhao Xu, Deze Zeng, Tao Ma 0006, Yong Yang 0013, Chao Li 0009, Minyi Guo |
USENIX ATC | 8 |
| 2021 | Enable simultaneous DNN services based on deterministic operator overlap and precise latency predictionabstractWhile user-facing services experience diurnal load patterns, co-locating services improve hardware utilization. Prior work on co-locating services on GPUs run queries sequentially, as the latencies of the queries are neither stable nor predictable when running simultaneously. The input sensitiveness and the non-deterministic operator overlap are two primary factors of the latency unpredictability. Hence, We propose Abacus, a runtime system that runs multiple services simultaneously. Abacus enables deterministic operator overlap to enforce latency predictability. Abacus composes of an overlap-aware latency predictor, a headroom-based query controller, and segmental model executors. The predictor predicts the latencies of the deterministic operator overlap. The controller determines the appropriate operator overlap for the QoS guarantee of all the services. The executors run the operators as needed to support the deterministic operator overlap. Our evaluation shows that Abacus reduces 51.3% of the QoS violation and improves the throughput by 29.8% on average compared with state-of-the-art solutions. Weihao Cui, Han Zhao 0005, Quan Chen 0002, Ningxin Zheng, Jingwen Leng, Jieru Zhao, Tao Ma 0006, Yong Yang 0013, Chao Li 0009, Minyi Guo |
SC | 8 |
| 2020 | URSA: Precise Capacity Planning and Fair Scheduling based on Low-level Statistics for Public CloudsabstractDatabase platform-as-a-service (dbPaaS) is developing rapidly and a large number of databases have been migrated to run on the Clouds for the low cost and flexibility. Emerging Clouds rely on the tenants to provide the resource specification for their database workloads. However, they tend to over-estimate the resource requirement of their databases, resulting in the unnecessarily high cost and low Cloud utilization. A methodology that automatically suggests the “just-enough” resource specification that fulfills the performance requirement of every database workload is profitable. Wei Zhang 0149, Ningxin Zheng, Quan Chen 0002, Yong Yang 0013, Tao Ma 0006, Jingwen Leng, Minyi Guo |
ICPP | 6 |
| 2020 | Amoeba: QoS-Awareness and Reduced Resource Usage of Microservices with Serverless ComputingabstractWhile microservices that have stringent Quality-of-Service constraints are deployed in the Clouds, the long-term rented infrastructures that host the microservices are under-utilized except peak hours due to the diurnal load pattern. It is resource efficient for Cloud vendors and cost efficient for service maintainers to deploy the microservices in the long-term infrastructure at high load and in the serverless computing platform at low load. However, prior work fails to take advantage of the opportunity, because the contention between microservices on the serverless platform seriously affects their response latencies.Our investigation shows that the load of a microservice, the shared resource contentions on the serverless platform, and its sensitivities to the contention together affect the response latency of the microservice on the platform. To this end, we propose Amoeba, a runtime system that dynamically switches the deployment of a microservice. Amoeba is comprised of a contention-aware deployment controller, a hybrid execution engine, and a multi-resource contention monitor. The deployment controller predicts the tail latency of a microservice based on its load and the contention on the serverless platform, and determines the appropriate deployment of the microservice. The hybrid execution engine enables the quick switch of the two deploy modes. The contention monitor periodically quantifies the contention on multiple types of shared resources. Experimental results show that Amoeba is able to significantly reduce up to 72.9% of CPU usage and up to 84.9% of memory usage compared with the traditional pure IaaS-based deployment, while ensuring the required latency target. Zijun Li 0001, Quan Chen 0002, Tao Ma 0006, Yong Yang 0013, Minyi Guo |
IPDPS | 4 |
| 2020 | Alita: comprehensive performance isolation through bias resource management for public cloudsabstractThe tenants of public cloud platforms share hard-ware resources on the same node, resulting in the potential for performance interference (or malicious attacks). A tenant is able to degrade the performance of its neighbors on the same node significantly through overuse of the shared memory bus, last level cache (LLC)/memory bandwidth, and power. To eliminate such unfairness we propose Alita, a runtime system consisting of an online interference identifier and adaptive interference eliminator. The interference identifier monitors hardware and system-level event statistics to identify resource polluters. The eliminator improves the performance of normal applications by throttling only the resource usage of polluters. Specifically, Alita adopts bus lock sparsification, bias LLC/bandwidth isolation, and selective power throttling to throttle the resource usage of polluters. Results for an experimental platform and in-production cloud platform with 30,000 nodes demonstrate that Alita significantly improves the performance of co-located virtual machines in the presence of resource polluters based on system-level knowledge. Quan Chen 0002, Shang Zhao 0003, Shanpei Chen, Tao Ma 0006, Yong Yang 0013, Minyi Guo |
SC | 8 |
| 2020 | Spool: Reliable Virtualized NVMe Storage Pool in Public Cloud Infrastructure
Shang Zhao 0003, Quan Chen 0002, Zheng Liu 0022, Tao Ma 0006, Yong Yang 0013, Yanbo Zhou, Keqiang Niu, Sijie Sun, Minyi Guo |
USENIX ATC | 8 |
| 2019 | X-RDMA: Effective RDMA Middleware in Large-scale Production EnvironmentsabstractX-RDMA is a communication middleware deployed and heavily used in Alibaba's large-scale cluster hosting cloud storage and database systems. Unlike recent research projects which purely focus on squeezing out the raw hardware performance, it puts emphasis on robustness, scalability and maintainability of large-scale production clusters. X-RDMA integrates necessary features, not available in current RDMA ecosystem, to release the developers from complex and imperfect details. X-RDMA simplifies the programming model, extends RDMA protocols for application awareness, and proposes mechanisms for resource management with thousands of connections per machine. It also reduces the work for administration and performance tuning with built-in tracing, tuning and monitoring tools. X-RDMA has been deployed in several large-scale clusters with over 4000 servers in Alibaba cloud since 2016. It can save at least 70% development and maintenance time over RDMA, effectively improve performance and reduce network jitter especially when production servers are under pressure. It also helped locate over 30 issues in different layers of productions with over 5000 connections for each server on average. Teng Ma 0006, Tao Ma 0006, Huaixin Chang, Kang Chen 0001, Hai Jiang 0003, Yongwei Wu 0001 |
CLUSTER | 2 |