Zan Zong

dblp:205/9119 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-7828-9030ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse
abstract
Recent advances in AI-generated content have driven widespread adoption of Diffusion Transformers (DiTs) for high-resolution, long-duration content generation. While parallelization techniques accelerate diffusion inference, they face significant scalability challenges due to excessive communication overhead in multi-node environments. We observe that sequence partitions in Context Parallelism (CP) exhibit distinct heterogeneity: spatially proximate partitions contribute more significantly to attention computation results. By mapping this heterogeneous pattern to hierarchical communication topology, we can access high-contribution partitions with reduced communication cost. This insight motivates our novel selective attention state mechanism that strategically balances partial attention computation and historical result reuse across denoising steps. We present DiTango, an efficient parallel framework for DiT generation. DiTango features an anchor-guided state selection planner that optimizes computation-reuse decisions for each partition, complemented by a runtime that orchestrates efficient state-centric operations. This design achieves superior system efficiency while preserving generation quality. Experimental evaluation on popular diffusion models demonstrates that DiTango achieves up to 1.9x end-to-end and 3.2x attention speedup with near-linear scaling in multi-node settings, while maintaining generation quality comparable to state-of-the-art approaches.
Runxin Zhong, Zan Zong, Hengjie Li, Yuyang Jin 0001, Jidong Zhai
HPDC3
2025 TraceFlow: Efficient Trace Analysis for Large-Scale Parallel Applications via Interaction Pattern-Aware Trace Distribution
abstract
Trace analysis of large-scale parallel applications is crucial for understanding and optimizing performance. It primarily focuses on the interaction behaviors between different parallel processes, such as synchronization waits and asynchronous overlaps. The trace size explodes as the parallel scale of applications, thus current methods analyze traces in parallel to ensure analysis speed. However, due to the interaction pattern-agnostic trace distribution, they often introduce inter-process communications to fetch non-local event data during interaction analysis, leading to excessively long trace analysis time.
Yuyang Jin 0001, Xirui Shui, Mingshu Zhai, Zan Zong, Feng Zhang 0007, Felix Wolf 0001, Jidong Zhai
SC4
2025 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling
abstract
Long-context comprehension is critical for large language models. Context parallelism and irregular block-sparse attention are keyss to accelerating long-context training and inference. Existing context parallelism suffers from poor scalability due to the striped-like partition pattern, which causes high communication traffic, and the ring-based communication pattern, which limits kernel granularity, reduces device utilization, and incurs redundant communication.
Zan Zong, Yuyang Jin 0001, Kinman Lei, Jiaao He, Qigang Yang, Jidong Zhai
SC2
2025 Multi-View Trace Clustering Based on Graph Convolutional Networks in Process Mining
abstract
Process mining techniques can extract process models from event logs produced by information systems. However, in flexible environments, simply using existing methods often leads to complex process models that are hard to understand, due to the less structured and greater complexity of processes in real life. Trace clustering is a pre-processing technique that enhances the effectiveness of model mining by partitioning similar behaviors in logs. In this paper, we present a Multi-view Trace Clustering method, named MTC, that improves the homogeneity of trace subclusters. Our method consists of three parts: (1) We use trace profiles to depict the traces from different views, and each profile would be transformed into a graph based on the k-nearest neighbor algorithm; (2) A fusion graph is designed to capture the information among these graphs based on an attention coefficient matrix, and then the graph convolutional networks are used to encode all graphs for obtaining the common representation; (3) We also enhance the characterization of the common representation with an inner decoder. Finally, we adopt k-means to cluster the traces in the log based on the common representation. Extensive experiments using multiple datasets illustrate that MTC significantly surpasses state-of-the-art trace clustering methods
Leilei Lin, Yunuo Cao, Zan Zong, Chen Qian 0003, Lijie Wen 0001
IEEE Trans. Serv. Comput.4
2023 SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online Parallelization
Mingshu Zhai, Jiaao He, Zixuan Ma, Zan Zong, Runqing Zhang, Jidong Zhai
USENIX ATC4
2023 STR: Hybrid Tensor Re-Generation to Break Memory Wall for DNN Training
abstract
With the growth of the depth of neural networks and the scale of data, the difficulty of network training also increases. When the GPU memory is insufficient, it is challenging to train deeper models. Recent research uses tensor swapping and recomputation techniques in a combined manner to optimize memory usage. However, complex dependencies and enormous scales of the DNN graph limit the improvement of single GPU memory optimization. Improper swap and recomputation decisions even bring negative effects on training performance. In this article, we propose a novel hybrid tensor re-generation strategy, called STR, which combines swap and recomputation techniques to find the optimal execution plan for the DNN training when the memory is limited. We formalize our memory optimization problem with constraints that describe the dependency of the operator calculation and the bandwidth usage of the swap. Ahost checkpointmechanism is designed to make full use of the swapped tensors, which reduces the cost of the recomputation. We also present arecursive source tracingalgorithm to improve the optimization efficiency by constraint relaxation with a performance bound. To optimize large models, we further introduce an approximation method based on a weighted graph coarsening. We implement a prototype of STR as a plugin on TensorFlow and evaluated based on 5 popular DNN models. The experimental result shows that the approximate solution of STR improves the training throughput of ResNet series of models by up to 28.1% compared to the state-of-the-art hybrid optimization strategy.
Zan Zong, Li Lin 0011, Leilei Lin, Lijie Wen 0001, Yu Sun 0027
IEEE Trans. Parallel Distributed Syst.1
2022 A Swap Dominated Tensor Re-Generation Strategy for Training Deep Learning Models
abstract
With the growing of the depth of neural networks and the scale of data, the difficulty of network training also increases. When the GPU memory is insufficient, it is challenging to train deeper models. Recent research uses tensor swapping and recomputation techniques in a combined manner to optimize the memory usage. However, complex dependencies of the DNN graph limit the improvement of the single GPU memory optimization. Improper swap decisions even brings negative effects because the source of the recomputation may have been swapped out. In this paper, we propose a novel swap dominated tensor re-generation strategy, called STR, which combines swap and recomputation techniques to find the optimal execution plan for the DNN training when the memory is limited. We formalize our memory optimization problem with constraints which describe the dependency of the operator calculation and the bandwidth usage of swap. A host checkpoint mechanism is designed to make full use of the swapped tensors, which reduces the cost of the recomputation. We also present an approximation method based on a recursive source tracing procedure to improve the optimization efficiency. We implement a prototype of STR as a plugin on TensorFlow. The experimental result shows that STR improves up to 21.3% throughput compared with the state-of-the-art hybrid optimization strategy.
Lijie Wen 0001, Zan Zong, Li Lin 0011, Leilei Lin
IPDPS2
2022 MespaConfig: Memory-Sparing Configuration Auto-Tuning for Co-Located In-Memory Cluster Computing Jobs
abstract
Distributed in-memory computing frameworks usually have lots of parameters (e.g., the buffer size of shuffle) to form a configuration for each execution. A well-tuned configuration can bring large improvements of performance. However, to improve resource utilization, jobs are often share the same cluster, which causes dynamic cluster load conditions. According to our observation, the variation of cluster load reduces effectiveness of configuration tuning. Besides, as a common problem of cluster computing jobs, overestimation of resources also occurs during configuration tuning. It is challenging to efficiently find the optimal configuration in a shared cluster with the consideration of memory-sparing. In this article, we introduce MespaConfig, a job-level configuration optimizer for distributed in-memory computing jobs. Advancements of MespaConfig over previous work are features including memory-sparing and load-sensitive. We evaluate MespaConfig by 6 typical Spark programs under different load conditions. The evaluation results show that MespaConfig improves the performance of six typical programs by up to 12× compared with default configurations. MespaConfig also achieves at most 41 percent reduction of configuration memory usage and reduces the optimization time overhead by 10.8× compared with the state-of-the-art approach.
Zan Zong, Lijie Wen 0001, Xuming Hu, Rui Han 0001, Chen Qian 0003, Li Lin 0011
IEEE Trans. Serv. Comput.1
2021 MM-CPred: A Multi-task Predictive Model for Continuous-Time Event Sequences with Mixture Learning Losses
Li Lin 0011, Zan Zong, Lijie Wen 0001, Chen Qian 0003, Shuang Li 0015, Jianmin Wang 0001
DASFAA (1)2
2020 An Approach for Process Model Extraction by Multi-grained Text Classification
Chen Qian 0003, Lijie Wen 0001, Akhil Kumar 0001, Leilei Lin, Li Lin 0011, Zan Zong, Shuang Li 0015, Jianmin Wang 0001
CAiSE6
2019 Workload-Adaptive Configuration Tuning for Hierarchical Cloud Schedulers
abstract
Cluster schedulers provide flexible resource sharing mechanism for best-effort cloud jobs, which occupy a majority in modern datacenters. Properly tuning a scheduler's configurations is the key to these jobs' performance because it decides how to allocate resources among them. Today's cloud scheduling systems usually rely on cluster operators to set the configuration and thus overlook the potential performance improvement through optimally configuring the scheduler according to the heterogeneous and dynamic cloud workloads. In this paper, we introduce AdaptiveConfig, a run-time configurator for cluster schedulers that automatically adapts to the changing workload and resource status in two steps. First, a comparison approach estimates jobs' performances under different configurations and diverse scheduling scenarios. The key idea here is to transform a scheduler's resource allocation mechanism and their variable influence factors (configurations, scheduling constraints, available resources, and workload status) into business rules and facts in a rule engine, thereby reasoning about these correlated factors in job performance comparison. Second, a workload-adaptive optimizer transforms the cluster-level searching of huge configuration space into an equivalent dynamic programming problem that can be efficiently solved at scale. We implement AdaptiveConfig on the popular YARN Capacity and Fair schedulers and demonstrate its effectiveness using real-world Facebook and Google workloads, i.e., successfully finding best configurations for most of scheduling scenarios and considerably reducing latencies by a factor of two with low optimization time.
Rui Han 0001, Chi Harold Liu, Zan Zong, Lydia Y. Chen, Wending Liu, Jianfeng Zhan
IEEE Trans. Parallel Distributed Syst.3
2018 AdaptiveConfig: Run-Time Configuration of Cluster Schedulers for Cloud Short-Running Jobs
abstract
Cluster schedulers provide flexible resource sharing mechanism for short-running jobs, which occupy a majority of cloud jobs. A scheduler's configuration decides how to allocate resources among jobs and hence it is crucial to their performances. Today's cloud platforms usually rely on cluster administrators to set this configuration, thus it is difficult to optimally configure the scheduler so as to minimize the latencies of heterogeneous and dynamically changing jobs in the cloud. In this paper, we introduce AdaptiveConfig, a run-time configurator for cluster schedulers that automatically adapts to the changing workload and resource status. This includes: (1) an estimator to calculate jobs' performances under different configurations and various scheduling scenarios. The key idea here is to transform a scheduler's resource allocation mechanisms and their variable influence factors (configuration parameters, scheduling constraints, available resources, and workload status) into business rules and facts in a rule engine, thereby reasoning about these correlated factors in job performance estimation. (2) A run-time optimizer that efficiently searches the configuration space to find the optimal configuration for the current workload. We implemented AdaptiveConfig on the popular YARN Capacity and Fair schedulers and demonstrate its effectiveness using workloads of Facebook jobs, i.e. considerably reducing latencies by 2.22 times (and up to 4.50 times) with low optimization overheads.
Rui Han 0001, Zan Zong, Lydia Y. Chen, Jianfeng Zhan
ICDCS2
2017 CloudMix: Generating Diverse and Reducible Workloads for Cloud Systems
abstract
The prosperity of cloud computing offers common infrastructures to a wide range of applications. Understanding these applications' workload behaviors is the premise of designing, managing, and optimizing cloud systems. Considering the heterogeneity and diversity of cloud workloads, for the sake of fairness, cloud benchmarks must be able to accurately replicate their behaviors in cloud systems, including both the usages of cloud resources and the micro-architectural behaviors beyond the virtualization layer. Furthermore, workloads spanning long durations are usually required to achieve representativeness in evaluation. Hence the more challenging issue is to significantly reduce the evaluation duration while still preserving their workload characteristics. This paper presents our efforts towards generating cloud workloads of diverse behaviors and reducible durations. Our benchmark tool, CloudMix, employs a repository of reducible workload blocks (RWBs) as the high level abstraction of workload behaviors, including usages of the two most important cloud resources (CPU and memory) and their pairing micro-architectural operations. CloudMix further introduces an efficient methodology to combine RWBs to synthesize and replicate diverse cloud workloads in real-world traces. The effectiveness of CloudMix is demonstrated by generating a variety of reducible workloads according to a Google cluster trace and by applying these workloads in job scheduling optimization on Hadoop YARN. The evaluation results show: (i) when the workload durations are reduced by 100 times, the replication errors of workload behaviors are smaller than 2.08%; (ii) when providing fast evaluations (workload durations are reduced by 10 to 100 times) to recommend the optimal setting in YARN job scheduling, the performance degradation in the recommended setting is just 0.69% compared to that of the actual optimal setting. CloudMix is publicly available from the project home page http://prof.ict.ac.cn/BigDataBench/multi tenancy/.
Rui Han 0001, Zan Zong, Fan Zhang 0047, José Luis Vázquez-Poletti, Zhen Jia 0001, Lei Wang 0004
CLOUD2