EDBT 2026 Demo / reviewers in the wild / expert
Danlin Jia
dblp:218/6176
· DBLP profile ↗
10ranked-venue papers
8as first author
6since 2021 · last 2024
0000-0003-2858-0505ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Computer networks · 3 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Data-Loader Tunable Knob to Shorten GPU Idleness for Distributed Deep LearningabstractDeep Neural Networks (DNNs) have been applied as an effective machine learning algorithm to tackle problems in different domains. However, the endeavor to train sophisticated DNN models can stretch from days into weeks, presenting substantial obstacles in the realm of research focused on large-scale DNN architectures. Distributed Deep Learning (DDL) contributes to accelerating DNN training by distributing training workloads across multiple computation accelerators, for example, graphics processing units (GPUs). Despite the considerable amount of research directed toward enhancing DDL training, the influence of data loading on GPU utilization and overall training efficacy remains relatively overlooked. It is non-trivial to optimize data-loading in DDL applications that need intensive central processing unit (CPU) and input/output (I/O) resources to process enormous training data. When multiple DDL applications are deployed on a system (e.g., Cloud and High-Performance Computing (HPC) system), the lack of a practical and efficient technique for data-loader allocation incurs GPU idleness and degrades the training throughput. Therefore, our work first focuses on investigating the impact of data-loading on the global training throughput. We then propose a throughput prediction model to predict the maximum throughput for an individual DDL training application. By leveraging the predicted results, A-Dloader is designed to dynamically allocate CPU and I/O resources to concurrently running DDL applications and use the data-loader allocation as a knob to reduce GPU idle intervals and thus improve the overall training throughput. We implement and evaluate A-Dloader in a DDL framework for a series of DDL applications arriving and completing across the runtime. Our experimental results show that A-Dloader can achieve a 28.9% throughput improvement and a 10% makespan improvement compared with allocating resources evenly across applications. Danlin Jia, Geng Yuan, Xue Lin 0001, Ningfang Mi |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | Learning-Based Dynamic Memory Allocation Schemes for Apache Spark Data ProcessingabstractApache Spark is an in-memory analytic framework that has been adopted in the industry and research fields. Two memory managers, Static and Unified, are available in Spark to allocate memory for caching Resilient Distributed Datasets (RDDs) and executing tasks. However, we find that the static memory manager (SMM) lacks flexibility, while the unified memory manager (UMM) puts heavy pressure on the garbage collection of the JVM on which Spark resides. To address these issues, we design a learning-based bidirectional usage-bounded memory allocation scheme to support dynamic memory allocation with the consideration of both memory demands and latency introduced by garbage collection. We first develop an auto-tuning memory manager (ATuMm) that adopts an intuitive feedback-based learning solution. However, ATuMm is a slow learner that can only alter the states of Java Virtual Memory (JVM) Heap in a limited range. That is, ATuMm decides to increase or decrease the boundary between the execution and storage memory pools by a fixed portion of JVM Heap size. To overcome this shortcoming, we further develop a new reinforcement learning-based memory manager (Q-ATuMm) that uses a Q-learning intelligent agent to dynamically learn and tune the partition of JVM Heap. We implement our new memory managers in Spark 2.4.0 and evaluate them by conducting experiments in a real Spark cluster. Our experimental results show that our memory manager can reduce the total garbage collection time and thus further improve Spark applications’ performance (i.e., reduced latency) compared to the existing Spark memory management solutions. By integrating our machine learning-driven memory manager into Spark, we can further obtain around 1.3x times reduction in the latency. Danlin Jia, Natalia Valencia, Janki Bhimani, Bo Sheng, Ningfang Mi |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | MoKE: Modular Key-value Emulator for Realistic Studies on Emerging Storage DevicesabstractKey-value stores are widely used as building blocks in today's IT infrastructure for managing and storing large amounts of data. Storage technologies are undergoing continuous innovations to accelerate KV workloads. However, designing high-performance KV or object storage devices is challenging and still needs more research to address the performance bottlenecks of the existing designs. There is a void for an inexpensive and extendable research platform that enables in-depth exploration of the index management components within the KV devices. To fill this void, we design Modular Key-value Emulator (MoKE). MoKE is a software emulator for fostering future full-stack software/hardware KV and object storage device research. MoKE is cheap (software-based emulator), usable with SNIA KV API (supports popular host-device interfaces), extendable (supports internal KV device research), and adaptable (QEMU-based). Manoj Pravakar Saha, Danlin Jia, Janki Bhimani, Ningfang Mi |
CLOUD | 2 |
| 2023 | SRC: Mitigate I/O Throughput Degradation in Network Congestion Control of Disaggregated Storage SystemsabstractThe industry has adopted disaggregated storage systems to provide high-quality services for hyper-scale architectures. This infrastructure enables organizations to access storage resources that can be independently managed, configured, and scaled. It is supported by the recent advances of all-flash arrays and NVMe-over-Fabric protocol, enabling remote access to NVMe devices over different network fabrics. A surge of research has been proposed to mitigate network congestion in traditional remote direct memory access protocol (RDMA). However, NVMe-oF raises new challenges in congestion control for disaggregated storage systems.In this work, we investigate the performance degradation of the read throughput on storage nodes caused by traditional network congestion control mechanisms. We design a storage-side rate control (SRC) to relieve network congestion while avoiding performance degradation on storage nodes. First, we design an I/O throughput control mechanism in the NVMe driver layer to enable throughput control on storage nodes. Second, we construct a throughput prediction model to learn a mapping function between workload characteristics and I/O throughput. Third, we deploy SRC on storage nodes to cooperate with traditional network congestion control on an NVMe-over-RDMA architecture. Finally, we evaluate SRC with varying workloads, SSD configurations, and network topologies. The experimental results show that SRC achieves significant performance improvement. Danlin Jia, Xuebin Yao, Mahsa Bayati, Pradeep Subedi, Bo Sheng, Ningfang Mi |
IPDPS | 1 |
| 2022 | A Data-Loader Tunable Knob to Shorten GPU Idleness for Distributed Deep LearningabstractDeep Neural Network (DNN) has been applied as an effective machine learning algorithm to tackle problems in different domains. However, training a sophisticated DNN model takes days to weeks and becomes a challenge in constructing research on large-scale DNN models. Distributed Deep Learning (DDL) contributes to accelerating DNN training by distributing training workloads across multiple computation accelerators (e.g., GPUs). Although a surge of research works has been devoted to optimizing DDL training, the impact of data-loading on GPU usage and training performance has been relatively under-explored. It is non-trivial to optimize data-loading in DDL applications that need intensive CPU and I/O resources to process enormous training data. When multiple DDL applications are deployed on a system (e.g., Cloud and HPC), the lack of a practical and efficient technique for data-loader allocation incurs GPU idleness and degrades the training throughput. Therefore, our work first focuses on investigating the impact of data-loading on the global training throughput. We then propose a throughput prediction model to predict the maximum throughput for an individual DDL training application. By leveraging the predicted results, A-Dloader is designed to dynamically allocate CPU and I/O resources to concurrently running DDL applications and use the data-loader allocation as a knob to reduce GPU idle intervals and thus improve the overall training throughput. We implement and evaluate A-Dloader in a DDL framework for a series of DDL applications arriving and completing across the runtime. Our experimental results show that A-Dloader can achieve a 23.5% throughput improvement and a 10% makespan improvement, compared to allocating resources evenly across applications. Danlin Jia, Geng Yuan, Xue Lin 0001, Ningfang Mi |
CLOUD | 1 |
| 2021 | SNIS: Storage-Network Iterative Simulation for Disaggregated Storage SystemsabstractIn recent years, designs and optimizations on disaggregated storage systems supported by cutting-edge storage and network techniques emerge dramatically. However, conducting experiments in a disaggregated architecture is often expensive. A comprehensive modeling system of disaggregated storage systems is indispensable for researchers to construct fast and reliable experiments. Modeling and evaluating the performance of disaggregated storage systems is a challenge for the following reasons. First, the performance of a disaggregated storage system depends on network protocols and storage solutions jointly. Second, the available trace datasets for generating the workload may not suffice the need for the simulation, as they are collected without considering the integration of network delay and storage processing time. This work proposes a storage-network iterative simulation (SNIS) for disaggregated storage systems by considering the issues above. Our simulation methodology integrates storage and network simulations to model the end-to-end performance of disaggregated storage systems and conducts multiple rounds of simulations to update arrival times of read/write requests. The evaluation results show that SNIS can converge to a relatively stable state after a certain number of iterations. Danlin Jia, Tengpeng Li, Mahsa Bayati, Ron Lee, Bo Sheng, Ningfang Mi |
IPCCC | 1 |
| 2020 | RITA: Efficient Memory Allocation Scheme for Containerized Parallel Systems to Improve Data Processing LatencyabstractThe emerging resource-sharing container-based virtualization is prevalent in IT, as it is a much lighter deployment in the cloud environment compared to VM-based virtualization. Distributed data-processing workloads executing in parallel take advantages of resource sharing, fast delivery, and excellent portability of containerization, but also suffer from resource competition and performance interference. Especially for memory virtualization, data-processing frameworks allocate physical memory (i.e., RAM) and swap to applications specified by users, without considering cache-characteristics and parallelism of applications, which induces performance degradation and significantly protracted latency which is worse given over-provisioning. We design an efficient memory allocation scheme (RITA) for containerized parallel systems to improve data processing latency. RITA monitors memory usage and cache characteristics of applications, and dynamically re-allocates memory resources. We implement RITA in a real-world system, which can easily migrate to other container-based virtualization environments. Our experimental results show that RITA provides remarkable latency improvement for memory intensive distributed data-processing workloads. Danlin Jia, Mahsa Bayati, Ron Lee, Ningfang Mi |
CLOUD | 1 |
| 2020 | Performance and Consistency Analysis for Distributed Deep Learning ApplicationsabstractAccelerating the training of Deep Neural Network (DNN) models is very important for successfully using deep learning techniques in fields like computer vision and speech recognition. Distributed frameworks help to speed up the training process for large DNN models and datasets. Plenty of works have been done to improve model accuracy and training efficiency, based on mathematical analysis of computations in the Con-volutional Neural Networks (CNN). However, to run distributed deep learning applications in the real world, users and developers need to consider the impacts of system resource distribution. In this work, we deploy a real distributed deep learning cluster with multiple virtual machines. We conduct an in-depth analysis to understand the impacts of system configurations, distribution typologies, and application parameters, on the latency and correctness of the distributed deep learning applications. We analyze the performance diversity under different model consistency and data parallelism by profiling run-time system utilization and tracking application activities. Based on our observations and analysis, we develop design guidelines for accelerating distributed deep-learning training on virtualized environments. Danlin Jia, Manoj Pravakar Saha, Janki Bhimani, Ningfang Mi |
IPCCC | 1 |
| 2019 | ATuMm: Auto-tuning Memory Manager in Apache SparkabstractApache Spark is an in-memory analytic framework that has been adopted in the industry and research fields. Two memory managers, Static and Unified, are available in Spark to allocate memory for caching Resilient Distributed Datasets (RDDs) and executing tasks. However, we found that the static memory manager (SMM) lacks flexibility, while the unified memory manager (UMM) puts heavy pressure on the garbage collection of JVM on which Spark resides. To address these issues, we design an auto-tuning memory manager (ATuMm) to support dynamic memory allocation with the consideration of both memory demands and latency introduced by garbage collection. We implement our new memory manager in Spark 2.2.0 and evaluate it by conducting experiments in a real Spark cluster. Our experimental results show that our auto-tuning memory manager can reduce the total garbage collection time and thus further improve the performance (i.e., reduced latency) of Spark applications, compared to the existing Spark memory management solutions. Danlin Jia, Janki Bhimani, Son Nam Nguyen, Bo Sheng, Ningfang Mi |
IPCCC | 1 |
| 2018 | Intermediate Data Caching Optimization for Multi-Stage and Parallel Big Data FrameworksabstractIn the era of big data and cloud computing, large amounts of data are generated from user applications and need to be processed in the datacenter. Data-parallel computing frameworks, such as Apache Spark, are widely used to perform such data processing at scale. Specifically, Spark leverages distributed memory to cache the intermediate results, represented as Resilient Distributed Datasets (RDDs). This gives Spark an advantage over other parallel frameworks for implementations of iterative machine learning and data mining algorithms, by avoiding repeated computation or hard disk accesses to retrieve RDDs. By default, caching decisions are left at the programmer's discretion, and the LRU policy is used for evicting RDDs when the cache is full. However, when the objective is to minimize total work, LRU is woefully inadequate, leading to arbitrarily suboptimal caching decisions. In this paper, we design an algorithm for multi-stage big data processing platforms to adaptively determine and cache the most valuable intermediate datasets that can be reused in the future. Our solution automates the decision of which RDDs to cache: this amounts to identifying nodes in a direct acyclic graph (DAG) representing computations whose outputs should persist in the memory. Our experiment results show that our proposed cache optimization solution can improve the performance of machine learning applications on Spark decreasing the total work to recompute RDDs by 12%. Zhengyu Yang 0001, Danlin Jia, Stratis Ioannidis, Ningfang Mi, Bo Sheng |
IEEE CLOUD | 2 |