EDBT 2026 Demo / reviewers in the wild / expert
Aidi Pi
dblp:206/4519
· DBLP profile ↗
14ranked-venue papers
4as first author
6since 2021 · last 2022
0000-0001-9468-2185ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Holmes: SMT Interference Diagnosis and CPU Scheduling for Job Co-locationabstractCo-location of latency-critical services with best-effort batch jobs is commonly adopted in production systems to increase resource utilization. Although memory and CPU isolation have been extensively studied, we find Simultaneous Multi-Threading (SMT) technology imposes non-trivial interference on memory access which jeopardizes efficient co-location and performance assurance of latency-critical services. However, there is not an existing metric to quantitatively measure and lacks a deterministic approach to tackle SMT interference on memory access. Aidi Pi, Xiaobo Zhou 0002, Cheng-Zhong Xu 0001 |
HPDC | 1 |
| 2022 | Improving Concurrent GC for Latency Critical Services in Multi-tenant SystemsabstractFor resource utilization efficiency, latency critical (LC) services are commonly co-located with best-effort batch jobs in datacenter servers. Many LC services, such as Cassandra and HBase, run in Java Virtual Machine (JVM). We find that LC services often experience heavy-tailed latency due to performance interference of the concurrent garbage collection (GC) as well as multi-tenancy. The root cause is a semantic gap of resource allocation between JVM and the underlying Linux OS in multi-tenant systems. That is, the OS is unaware of the characteristics of different kinds of threads in JVM (i.e., GC threads and LC worker threads), which may lead to GC threads competing for CPUs; JVM is unaware of the resource utilization in the OS, which may trigger CPU-intensive GC operations when CPUs are busy. Furthermore, we find that co-located batch jobs can interfere with LC services due to Simultaneous Multi-Threading (SMT). Junxian Zhao, Aidi Pi, Xiaobo Zhou 0002, Sang-Yoon Chang, Cheng-Zhong Xu 0001 |
Middleware | 2 |
| 2022 | Elastic Parameter Server: Accelerating ML Training With Scalable Resource SchedulingabstractParameter server (PS) based on worker-server communication is designed for distributed machine learning (ML) training in clusters. In feedback-driven exploration of ML model training, users exploit early feedback from each job to decide whether to kill the job or keep it running so as to find the optimal model configuration. However, PS does not support adjusting the number of workers and servers of a job at runtime. It becomes the bottleneck of scalable distributed ML training because the cluster resources cannot be dynamically allocated or deallocated to jobs, resulting in significant early feedback latency and resource under-utilization. This article rethinks the principle of PS architecture. We present Elastic Parameter Server (EPS), a lightweight and user-transparent PS that accelerates feedback-driven exploration for distributed ML training. EPS allows to remove a subset of workers and servers from running jobs and allocate the released resources to an incoming job at runtime so as to reduce its early feedback latency. It can also use the released resources from a killed job to add workers and servers to running jobs to improve resource utilization and the training speed. We develop a heuristic scheduler that leverages EPS and offers scalable resource scheduling for multiple ML jobs. We implement EPS in Tencent Angel and the scheduler in Apache Yarn, and conduct evaluations with various ML models. Experimental results show that EPS achieves up to 1.5x improvement on the ML training speed compared to PS. Aidi Pi, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | FlashByte: Improving Memory Efficiency with Lightweight Native StorageabstractIn-memory caching of intermediate data is effective in reducing re-computation and I/O cost in distributed data-analytics frameworks, but it also generates a large amount of data in Java heap which increases the overhead of garbage collection (GC). An alternative off-heap approach caches data in native storage by transmitting the data from heap to native storage so as to reduce GC overhead. However, it incurs severe serialization and de-serialization overheads. Serialization also generates non-trivial metadata of the cached data in native storage. We propose and develop FlashByte, a lightweight native storage that efficiently caches intermediate data. FlashByte improves memory efficiency by achieving low GC overhead, low data transmission overhead, and low memory consumption. Specifically, the cached data are divided into two parts: metadata stored in Java heap and raw data stored in native storage. The metadata is generated based on the profile of workloads. Its size is trivial because it only contains a concise format of raw data, which achieves low memory consumption in the heap as well as low GC overhead. Native storage only stores the raw data to reduce its memory consumption. According to the metadata, the raw data are efficiently transmitted between the heap and native storage without serialization and de-serialization. We implement FlashByte in Spark and conduct evaluation with benchmark workloads. Experimental results show that, compared with the in-heap approach of Vanilla Spark, FlashByte achieves up to 4x speedup of the job execution time, reduces GC time by up to 96%, and reduces the memory consumption in heap by up to 36%. Compared with the alternative off-heap approach, FlashByte achieves up to 2.3x speedup of the job execution time, reduces the data transmission time by up to 84%, and reduces the cache size in native storage by up to 34%. Junxian Zhao, Aidi Pi, Xiaobo Zhou 0002 |
CCGRID | 2 |
| 2021 | Memory at your service: fast memory allocation for latency-critical servicesabstractCo-location and memory sharing between latency-critical services, such as key-value store and web search, and best-effort batch jobs is an appealing approach to improving memory utilization in multi-tenant datacenter systems. However, we find that the very diverse goals of job co-location and the GNU/Linux system stack can lead to severe performance degradation of latency-critical services under memory pressure in a multi-tenant system. Aidi Pi, Junxian Zhao, Xiaobo Zhou 0002 |
Middleware | 1 |
| 2021 | Overlapping Communication With Computation in Parameter Server for Scalable DL TrainingabstractScalability of distributed deep learning (DL) training with parameter server (PS) architecture is often communication constrained in large clusters. There are recent efforts that use a layer by layer strategy to overlap gradient communication with backward computation so as to reduce the impact of communication constraint on the scalability. However, the approaches could bring significant overhead in gradient communication. Meanwhile, they cannot be effectively applied to the overlap between parameter communication and forward computation. In this article, we propose and develop iPart, a novel approach that partitions communication and computation in various partition sizes to overlap gradient communication with backward computation and parameter communication with forward computation. iPart formulates the partitioning decision as an optimization problem and solves it based on a greedy algorithm to derive communication and computation partitions. We implement iPart in the open-source DL framework BigDL and perform evaluations with various DL workloads. Experimental results show that iPart improves the scalability of a cluster of 72 nodes by up to 94 percent over the default PS and 52 percent over the layer by layer strategy. Aidi Pi, Xiaobo Zhou 0002, Jun Wang 0001, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Scalable Distributed DL Training: Batching Communication and ComputationabstractScalability of distributed deep learning (DL) training with parameter server architecture is often communication constrained in large clusters. There are recent efforts that use a layer by layer strategy to overlap gradient communication with backward computation so as to reduce the impact of communication constraint on the scalability. However, the approaches cannot be effectively applied to the overlap between parameter communication and forward computation. In this paper, we propose and design iBatch, a novel communication approach that batches parameter communication and forward computation to overlap them with each other. We formulate the batching decision as an optimization problem and solve it based on greedy algorithm to derive communication and computation batches. We implement iBatch in the open-source DL framework BigDL and perform evaluations with various DL workloads. Experimental results show that iBatch improves the scalability of a cluster of 72 nodes by up to 73% over the default PS and 41% over the layer by layer strategy. Aidi Pi, Xiaobo Zhou 0002 |
AAAI | 2 |
| 2019 | Pufferfish: Container-driven Elastic Memory Management for Data-intensive ApplicationsabstractData-intensive applications often suffer from significant memory pressure, resulting in excessive garbage collection (GC) and out-of-memory (OOM) errors, harming system performance and reliability. In this paper, we demonstrate how lightweight virtualization via OS containers opens up opportunities to address memory pressure and realize memory elasticity: 1) tasks running in a container can be set to a large heap size to avoid OutOfMemory (OOM) errors, and 2) tasks that are under memory pressure and incur significant swapping activities can be temporarily "suspended" by depriving resources from the hosting containers, and be "resumed" when resources are available. We propose and develop Pufferfish, an elastic memory manager, that leverages containers to flexibly allocate memory for tasks. Memory elasticity achieved by Pufferfish can be exploited by a cluster scheduler to improve cluster utilization and task parallelism. We implement Pufferfish on the cluster scheduler Apache Yarn. Experiments with Spark and MapReduce on real-world traces show Pufferfish is able to avoid OOM errors, improve cluster memory utilization by 2.7x and the median job runtime by 5.5x compared to a memory over-provisioning solution. Wei Chen 0038, Aidi Pi, Xiaobo Zhou 0002 |
SoCC | 2 |
| 2019 | Semantic-aware Workflow Construction and Analysis for Distributed Data Analytics SystemsabstractLogging is a universal approach to recording important events in system workflows of distributed systems. Current log analysis tools ignore the semantic knowledge that is key to workflow construction and analysis. In addition, they focus on infrastructure-level distributed systems. Because of fundamental differences in log features, they are ineffective in distributed data analytics systems. This paper proposes IntelLog, a semantic-aware non-intrusive workflow reconstruction tool for distributed data analytics systems. It is capable of building hierarchical relationships between components and events from logs generated by the targeted systems with little or even no domain knowledge. Leveraging natural language processing, IntelLog automatically extracts and formats semantic information in each log message, including system events, identifiers, locality information, and metrics values. It builds a graph to represent the hierarchical relationship of components in the targeted system via nomenclature conventions. We implement IntelLog for Hadoop MapReduce, Spark and Tez. Evaluation results show that IntelLog provides a fine-grained view of the system workflows with semantics. It outperforms existing tools in automatically detecting anomalies caused by real-world problems, misconfigurations and system bugs. Users can query the formatted semantic knowledge to understand and further troubleshoot the systems. Aidi Pi, Wei Chen 0038, Xiaobo Zhou 0002 |
HPDC | 1 |
| 2019 | OS-Augmented Oversubscription of Opportunistic Memory with a User-Assisted OOM KillerabstractExploiting opportunistic memory by oversubscription is an appealing approach to improving cluster utilization and throughput. In this paper, we find the efficacy of memory oversubscription depends on whether or not the oversubscribed tasks can be killed by an OutOf Memory (OOM) killer in a timely manner to avoid significant memory thrashing upon memory pressure. However, current approaches in modern cluster schedulers are actually unable to unleash the power of opportunistic memory because their user space OOM killers are unable to timely deliver a task killing signal to terminate the oversubscribed tasks. Our experiments observe that a user space OOM killer fails to do that because of lacking the memory pressure knowledge from OS while the kernel space Linux OOM killer is too conservative to relieve memory pressure. Wei Chen 0038, Aidi Pi, Xiaobo Zhou 0002 |
Middleware | 2 |
| 2018 | Profiling distributed systems in lightweight virtualized environments with logs and resource metricsabstractUnderstanding and troubleshooting distributed systems in the cloud is considered a very difficult problem because the execution of a single user request is distributed to multiple machines. Further, the multi-tenancy nature of cloud environments further introduces interference that causes performance issues. Most existing troubleshooting tools either focus on log analysis or intrusive tracing methods, leaving resource usage monitoring unexplored. Aidi Pi, Wei Chen 0038, Xiaobo Zhou 0002, Mike Ji |
HPDC | 1 |
| 2018 | Characterizing Scheduling Delay for Low-Latency Data Analytics WorkloadsabstractData analytics workloads are shifting to shorter task execution time, higher degree of parallelism, and execution on faster hardware. As a result, job scheduling is becoming a bottleneck, which needs to offer extreme low-latency, massive throughput, and high scalability. However, few efforts have been focused on systematically understanding the scheduling delay. In this paper, we propose a method and develop a tool, SD-checker, that decomposes the job scheduling delay into multiple components and characterizes each by extensive experiments. SDchecker extracts event messages through mining both cluster scheduler logs and application logs, and constructs a scheduling order graph for the ease of analysis. We use SDchecker to evaluate Spark-SQL on a popular cluster scheduler Yarn. Results show that the scheduling delay may account for 60% of the job runtime of small data analytics workloads. After decomposing the total scheduling delay, we find Spark itself contributes 70% of the delay. Through the evaluation and analysis, we conclude that (1) The causes of scheduling delay are determined by many factors, and (2) The job scheduling is not well optimized yet, and far from ideal for low-latency data analytics workloads. Wei Chen 0038, Aidi Pi, Xiaobo Zhou 0002 |
IPDPS | 2 |
| 2018 | Aggressive Synchronization with Partial Processing for Iterative ML Jobs on ClustersabstractExecuting distributed machine learning (ML) jobs on Spark follows Bulk Synchronous Parallel (BSP) model, where parallel tasks execute the same iteration at the same time and the generated updates must be synchronized on parameters when all tasks are finished. However, the parallel tasks rarely have the same execution time due to sparse data so that the synchronization has to wait for tasks finished late. Moreover, running Spark on heterogeneous clusters makes it even worse because of stragglers, where the synchronization is significantly delayed by the slowest task. Wei Chen 0038, Aidi Pi, Xiaobo Zhou 0002 |
Middleware | 3 |
| 2017 | mBalloon: enabling elastic memory management for big data processingabstractBig Data processing often suffers from significant memory pressure, resulting in excessive garbage collection (GC) and out-of-memory (OOM) errors, harming system performance and reliability. Therefore, users tend to give an excessive heap size to applications to avoid job failure, causing low cluster utilization. Wei Chen 0038, Aidi Pi, Jia Rao, Xiaobo Zhou 0002 |
SoCC | 2 |