EDBT 2026 Demo / reviewers in the wild / expert
Haojie Wang 0004
dblp:121/3324-4
· DBLP profile ↗
34ranked-venue papers
3as first author
30since 2021 · last 2026
0000-0003-4605-148XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 2 first-author · 25 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy ServersabstractMixture-of-Experts (MoE) models face memory and PCIe latency bottlenecks when deployed on commodity hardware. Offloading expert weights to CPU memory results in PCIe transfer latency that exceeds GPU computation by several folds. We present PreScope, a prediction-driven expert scheduling system that addresses three key challenges: inaccurate activation prediction, PCIe bandwidth competition, and cross-device scheduling complexity. Our solution includes: 1) Learnable Layer-Aware Predictor (LLaPor) that captures layer-specific expert activation patterns; 2) Prefetch-Aware Cross-Layer Scheduling (PreSched) that generates globally optimal plans balancing prefetching costs and loading overhead; 3) Asynchronous I/O Optimizer (AsyncIO) that decouples I/O from computation, eliminating waiting bubbles. PreScope achieves 141% higher throughput and 74.6% lower latency than state-of-the-art solutions. Enda Yu, Dezun Dong, Zhaoning Zhang 0001, Zhe Bai, Weiling Yang, Haojie Wang 0004, Dongsheng Li 0001, Yongwei Wu 0001, Xiangke Liao |
ICS | 6 |
| 2026 | Taming Dynamic Diffusion LLM Inference through Virtual Static Execution
Jianian Zhu, Haojie Wang 0004, Ruixuan Li 0001, Jidong Zhai |
ICS | 4 |
| 2026 | ChituDiffusion: A Data-Characteristic-Aware Serving System for Diffusion ModelsabstractDiffusion models have become the dominant approach for generative tasks in images, videos, and other domains. However, diverse data properties in generation requests, which are critical for efficient serving, remain underexploited. To address this issue, we propose a diffusion model serving system ChituDiffusion. ChituDiffusion leverages the locality of data properties to recompose a diffusion pipeline into dGraphs with shared optimization opportunities, enabling thorough compile-time and runtime co-optimizations. During compilation, ChituDiffusion compiles each dGraph into multiple execution engines optimized for specific data properties. At runtime, heterogeneous requests are elaborately reorganized into fine-grained batching tasks with similar properties and then efficiently executed by matched engines. Evaluation on five diffusion applications shows that ChituDiffusion improves the throughput by up to 2.13× (1.58× on average) on A100 and 2.19× (1.51× on average) on H100 compared with existing frameworks. The code for ChituDiffusion and the production traces have been made open-source at https://github.com/thu-pacman/chitu/tree/Diffusion. Chengzhang Wu, Liyan Zheng 0001, Haojie Wang 0004, Kezhao Huang, Zixuan Ma, Dong Dong 0001, Jidong Zhai |
PPoPP | 3 |
| 2026 | UniOrch: A Unified Mixed Framework for High-Efficiency LLM Training on Heterogeneous AI ChipsabstractEfficient coordination of heterogeneous AI chips (GPU/NPU/DCU) in data centers is crucial for Large Language Model (LLM) training, but this process is hindered by architectural mismatches, protocol fragmentation, and network partitioning. Existing solutions fail to achieve unified resource management across different chips, resulting in severe resource fragmentation and reduced allreduce efficiency. To overcome these limitations, this paper proposes the unified coordination framework UniOrch, which not only integrates three core functionalities, including hardware abstraction, software standardization, and communication coordination, but also enables the training and inference of large models across heterogeneous AI chips. UniOrch's hardware-agnostic bare-metal cloud eliminates virtualization overhead through Border Gateway Protocol Ethernet Virtual Private Network (BGP EVPN) overlay networks and gateway-based chip integration; its PyTorch-based adaptation layer masks hardware differences and reduces migration costs; the Transformer Collective Communication Library (TCCL) unifies NCCL, HCCL, and OpenMPIprotocols to support seamless hybrid parallel training. Furthermore, the framework ' score scheduling mechanism is the Heterogeneous Hybrid Estimation Model (HHEM), which employs a two stage cost model combining static analysis with dynamic runtime feedback to dynamically allocate computing power based on Transformer task loads, ensuring cross-chip task synchronization, resource pooling, and dynamic allocation. Deployment verification in real-world production environments (e.g., China Construction Bank) shows that UniOrch achieves significant improvements: resource utilization of heterogeneous AI infrastructure is increased by 35%, cross-chip latency is reduced by 42%, accuracy loss in heterogeneous environments is <0.8%. Jia Wang 0038, Yang Zhai, Haojie Wang 0004, Wanxin Song, Wei Li 0008, Zhaofeng He 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | IntelliGen: Instruction-Level Auto-tuning for Tensor Program with Monotonic Memory OptimizationabstractTensor compilers play a critical role in optimizing deep neural networks (DNNs), with memory performance emerging as a key bottleneck in code generation for DNN models. Existing tensor compilers are constrained by inefficient auto-tuning algorithms. They either must deploy coarse-grained descriptions, thus miss potential optimization, or struggle with vast search spaces, rendering auto-tuning inapplicable. Tensor compilers require a more holistic optimization of memory performance to overcome these constraints. To address this issue, we focus our optimization objective on memory performance, which allows us to design monotonic optimization methods, significantly enhancing the efficiency of auto-tuning and thus enabling auto-tuning on a fine-granularity description. Based on these observations, we propose IntelliGen, a tensor compiler with instruction-level auto-tuning and monotonic memory optimization. We design an instruction-level graph description, and a monotonic optimization method for optimization on . Benefiting from auto-tuning techniques with fine-grained description, IntelliGen demonstrates significant speedup of up to 3.13×, 3.55×, and 16.9× (averaging 1.46×, 1.85×, and 2.30×, respectively) on NVIDIA GPUs, AMD GPUs, and Cambricon MLUs over the most efficient existing frameworks. Zixuan Ma, Haojie Wang 0004, Jingze Xing, Shuhong Huang, Liyan Zheng 0001, Chen Zhang 0001, Huanqi Cao, Kezhao Huang, Mingshu Zhai, Shizhi Tang, Penghan Wang, Jidong Zhai |
CGO | 2 |
| 2025 | Leveraging Graph Analysis to Pinpoint Root Causes of Scalability Issues for Parallel ApplicationsabstractIt is challenging to scale parallel applications to modern supercomputers because of load imbalance, resource contention, and communications between processes. Profiling and tracing are two main performance analysis approaches for detecting these scalability bottlenecks. Profiling is low-cost but lacks detailed dependence for identifying root causes. Tracing records plentiful information but incurs significant overheads. To address these issues, we presentScalAna, which employs static analysis techniques to combine the benefits of profiling and tracing - it enables tracing's analyzability with overhead similar to profiling.ScalAnauses static analysis to capture program structures and data dependence of parallel applications, and leverages lightweight profiling approaches to record performance data during runtime. Then a parallel performance graph is generated with both static and dynamic data. Based on this graph, we design a backtracking detection approach to automatically pinpoint the root causes of scaling issues. We evaluate the efficacy and efficiency ofScalAnausing several real applications with up to 704K lines of code and demonstrate that our approach can effectively pinpoint the root causes of scaling loss with an average overhead of 5.65% for up to 16,384 processes. By fixing the root causes detected by our tool, it achieves up to 33.01% performance improvement. Yuyang Jin 0001, Haojie Wang 0004, Xiongchao Tang, Zhenhua Guo 0003, Yaqian Zhao, Torsten Hoefler, Tao Liu 0029, Xu Liu 0001, Jidong Zhai |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Optimal Kernel Orchestration for Tensor Programs with KorchabstractKernel orchestration is the task of mapping the computation defined in different operators of a deep neural network (DNN) to the execution of GPU kernels on modern hardware platforms. Prior approaches optimize kernel orchestration by greedily applying operator fusion, which fuses the computation of multiple operators into a single kernel, and miss a variety of optimization opportunities in kernel orchestration. Muyan Hu, Ashwin Venkatram, Shreyashri Biswas, Balamurugan Marimuthu, Bohan Hou, Gabriele Oliaro, Haojie Wang 0004, Liyan Zheng 0001, Xupeng Miao, Jidong Zhai |
ASPLOS (3) | 7 |
| 2024 | AdaPipe: Optimizing Pipeline Parallelism with Adaptive Recomputation and PartitioningabstractLarge language models (LLMs) have demonstrated powerful capabilities, requiring huge memory with their increasing sizes and sequence lengths, thus demanding larger parallel systems. The broadly adopted pipeline parallelism introduces even heavier and unbalanced memory consumption. Recomputation is a widely employed technique to mitigate the problem but introduces extra computation overhead. Zhenbo Sun, Huanqi Cao, Yuanwei Wang, Guanyu Feng, Shengqi Chen 0001, Haojie Wang 0004 |
ASPLOS (3) | 6 |
| 2024 | WiseGraph: Optimizing GNN with Joint Workload Partition of Graph and OperationsabstractGraph Neural Network (GNN) has emerged as an important workload for learning on graphs. With the size of graph data and the complexity of GNN model architectures increasing, developing an efficient GNN system grows more important. As GNN has heavy neural computation workloads on a large graph, it is crucial to partition the entire workload into smaller parts for parallel execution and optimization. However, existing approaches separately partition graph data and GNN operations, resulting in inefficiency and large data movement overhead. Kezhao Huang, Jidong Zhai, Liyan Zheng 0001, Haojie Wang 0004, Yuyang Jin 0001, Qihao Zhang, Runqing Zhang, Zhen Zheng, Youngmin Yi, Xipeng Shen |
EuroSys | 4 |
| 2024 | MAGPY: Compiling Eager Mode DNN Programs by Monitoring Execution States
Chen Zhang 0001, Rongchao Dong, Haojie Wang 0004, Runxin Zhong, Jike Chen, Jidong Zhai |
USENIX ATC | 3 |
| 2024 | Graph-Centric Performance Analysis for Large-Scale Parallel ApplicationsabstractPerformance analysis is essential for understanding the performance behaviors of parallel programs and detecting performance bottlenecks. Whereas, complex interconnections across several types of performance bugs, as well as inter-process communications and data dependence, make efficient performance analysis even more difficult. Despite the fact that many performance tools have been developed, accurately identifying underlying performance bottlenecks for such complex scenarios requires specific in-depth analysis. Significant human efforts and analysis knowledge are often required to implement each specific analytic task. To alleviate the complexity of developing specific performance analytic tasks, we present a programmable performance analysis tool, calledPerFlow. InPerFlow, a step-by-step performance analysis process is represented as an Analysis Flow Diagram, which is constructed with several performance analysis sub-tasks, namely passes, that can be defined by developers or provided byPerFlow's built-in analysis pass library. Furthermore, we define a Performance Abstraction Graph to describe the performance behavior of a parallel program, where the edges indicate the interactions between parallel units, therefore the analytic sub-tasks are converted to graph analysis tasks.PerFlowprovides plentiful Python APIs for developing analytic tasks. Several case studies of real-world applications with up to 700 K lines of code are used to demonstrate the effectiveness ofPerFlow. The results indicate thatPerFlowmakes it much easier to implement specific performance analytic tasks, and these tasks are performed automatically and efficiently to detect underlying performance bottlenecks. Yuyang Jin 0001, Haojie Wang 0004, Runxin Zhong, Chen Zhang 0001, Xia Liao, Feng Zhang 0007, Jidong Zhai |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | EINNET: Optimizing Tensor Programs with Derivation-Based Transformations
Liyan Zheng 0001, Haojie Wang 0004, Jidong Zhai, Muyan Hu, Zixuan Ma, Tuowei Wang, Shuhong Huang, Xupeng Miao, Shizhi Tang, Kezhao Huang |
OSDI | 2 |
| 2023 | GraphSet: High Performance Graph Mining through Equivalent Set TransformationsabstractGraph mining is of critical use in a number of fields such as social networks, knowledge graphs, and fraud detection. As an NP-complete problem, accelerating computation performance is the main target for current optimizations. Due to excellent performance, state-of-the-art graph mining systems mainly rely on pattern-aware algorithms. Despite previous efforts, complex control flows introduced by pattern-aware algorithms bring significant overhead and also impede further acceleration on heterogeneous hardware. Tianhui Shi, Jidong Zhai, Haojie Wang 0004, Qiqian Chen, Mingshu Zhai, Zixu Hao |
SC | 3 |
| 2023 | Unified Programming Models for Heterogeneous High-Performance Computers
Zixuan Ma, Yuyang Jin 0001, Shizhi Tang, Haojie Wang 0004, Wei-Cheng Xue, Jidong Zhai |
J. Comput. Sci. Technol. | 4 |
| 2023 | Optimizing DNNs With Partially Equivalent Transformations and Automated CorrectionsabstractDeep neural network (DNN) applications are typically represented by tensor programs. To boost the performance of DNN computations, existing works adopt fully equivalent transformations for tensor program optimization by guaranteeing the equivalence on each element of tensors. However, as there are thousands of elements in a tensor, such optimization misses the opportunities that allow the in-equivalence of minority elements. In this work, we proposePet, the first work that introduces partially equivalent transformations to optimize tensor programs. To maintain the functional equivalence of tensor programs,Petautomatically finds and corrects the in-equivalent positions by leveraging the multi-linearity of DNN computations.Petfurther uses a mutation manager to improve search efficiency. Evaluation results show thatPetcan achieve up to 1.98$\times$and 2.20$\times$speedups on NVIDIA Tesla A100 and V100 respectively compared with existing DNN frameworks by introducing new optimization opportunities of partially equivalent transformations. Haojie Wang 0004, Jidong Zhai, Mingyu Gao 0001, Feng Zhang 0007, Tuowei Wang, Zixuan Ma, Shizhi Tang, Liyan Zheng 0001, Kaiyuan Rong, Yuanyong Chen |
IEEE Trans. Computers | 1 |
| 2022 | An Efficient Sparse CNNs Accelerator on FPGAabstractConvolutional Neural Networks (CNNs) have achieved remarkable performance at a huge computational cost. By improving the model sparsity, it can effectively reduce the complexity. However, with deepening of sparsity, the problems of unbalanced workloads, computing fragmentation and mapping access conflict caused by irregular sparsity have become more and more remarkable. These problems pose great challenges for efficient computation of sparse CNN s. In order to make full use of two side of sparsity introduced by activations and weights, and overcome the above problems, this paper proposes an efficient sparse CNN s accelerator on FPGA to achieve the inference acceleration. We designed and implemented the accelerator on the Zynq UltraScale+ MPSoC ZCU102 evaluation board. By running AlexNet, VGG16 and ResNet50 networks on the accelerator to evaluated the peeformance. Experimental results show that the method proposed in this paper can achieve more than 97% reduction in collision rate and 2.35x improvement in computing performance and 9.37x improvement in energy efficiency. Haojie Wang 0004, Dong Dong 0001, Yongxiang Cao |
CLUSTER | 4 |
| 2022 | Efficiently emulating high-bitwidth computation with low-bitwidth hardwareabstractDomain-Specific Accelerators (DSAs) are being rapidly developed to support high-performance domain-specific computation. Although DSAs provide massive computation capability, they often only support limited native data types. To mitigate this problem, previous works have explored software emulation for certain data types, which provides some compensation for hardware limitations. However, how to efficiently design more emulated data types and choose a high-performance one without hurting correctness or precision for a given application still remains an open problem. Zixuan Ma, Haojie Wang 0004, Guanyu Feng, Chen Zhang 0001, Jiaao He, Shengqi Chen 0001, Jidong Zhai |
ICS | 2 |
| 2022 | FreeTensor: a free-form DSL with holistic optimizations for irregular tensor programsabstractTensor programs are of critical use in many domains. Existing frameworks, such as PyTorch, TensorFlow, and JAX, adopt operator-based programming to ease programming, increase performance, and perform automatic differentiation. However, as the rapid development of tensor programs, operator-based programming shows significant limitations for irregular patterns since a large amount of redundant computation or memory access is introduced. Shizhi Tang, Jidong Zhai, Haojie Wang 0004, Liyan Zheng 0001, Zhenhao Yuan, Chen Zhang 0001 |
PLDI | 3 |
| 2022 | Scaling graph traversal to 281 trillion edges with 40 million coresabstractGraph processing, especially high-performance graph traversal, plays a more and more important role in data analytics. The successor of Sunway TaihuLight, New Sunway, is equipped with nearly 10 PB memory and over 40 million cores, which brings the opportunity to process hundreds of trillions of edges graphs. However, the graph with an unprecedented scale also brings severe performance challenges, including load imbalance, poor locality, and irregular access of graph traversal workload. Huanqi Cao, Yuanwei Wang, Haojie Wang 0004, Heng Lin, Zixuan Ma, Wanwang Yin |
PPoPP | 3 |
| 2022 | FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained modelsabstractThe current trend in deep learning is to scale models to extremely large sizes with the objective of increasing their accuracy. Mixture-of-Expert (MoE) is the most popular pre-trained model that makes feasible the training of models with parameters beyond trillion-scale. Thanks to the dynamic activation of experts, i.e., shallow layers specialized in certain domains, it allows for sparse training of bigger models, removing the linearity between model size and computation. However, different from traditional deep learning models, it draws huge challenges to the efficiency of these training systems, including dynamic load imbalance, inefficient synchronous execution mode, and congested all-to-all communication. Jiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang 0004, Fuwen Luo, Shangfeng Shi |
PPoPP | 4 |
| 2022 | PerFlow: a domain specific framework for automatic performance analysis of parallel applicationsabstractPerformance analysis is widely used to identify performance issues of parallel applications. However, complex communications and data dependence, as well as the interactions between different kinds of performance issues make high-efficiency performance analysis even harder. Although a large number of performance tools have been designed, accurately pinpointing root causes for such complex performance issues still needs specific in-depth analysis. To implement each such analysis, significant human efforts and domain knowledge are normally required. Yuyang Jin 0001, Haojie Wang 0004, Runxin Zhong, Chen Zhang 0001, Jidong Zhai |
PPoPP | 2 |
| 2022 | BaGuaLu: targeting brain scale pretrained models with over 37 million coresabstractLarge-scale pretrained AI models have shown state-of-the-art accuracy in a series of important applications. As the size of pretrained AI models grows dramatically each year in an effort to achieve higher accuracy, training such models requires massive computing and memory capabilities, which accelerates the convergence of AI and HPC. However, there are still gaps in deploying AI applications on HPC systems, which need application and system co-design based on specific hardware features. Zixuan Ma, Jiaao He, Jiezhong Qiu, Huanqi Cao, Yuanwei Wang, Zhenbo Sun, Liyan Zheng 0001, Haojie Wang 0004, Shizhi Tang, Tianyu Zheng, Junyang Lin, Guanyu Feng, Zeqiang Huang, Aohan Zeng, Jianwei Zhang 0012, Runxin Zhong, Tianhui Shi, Jie Tang 0001, Hongxia Yang, Xin Liu 0086, Jidong Zhai |
PPoPP | 8 |
| 2022 | Vapro: performance variance detection and diagnosis for production-run parallel applicationsabstractPerformance variance is a serious problem for parallel applications, which can cause performance degradation and make applications' behavior hard to understand. Therefore, detecting and diagnosing performance variance are of crucial importance for users and application developers. However, previous detection approaches either bring too large overhead and hurt applications' performance, or rely on nontrivial source code analysis that is impractical for production-run parallel applications. Liyan Zheng 0001, Jidong Zhai, Xiongchao Tang, Haojie Wang 0004, Yuyang Jin 0001, Shuaiwen Song |
PPoPP | 4 |
| 2022 | UniQ: A Unified Programming Model for Efficient Quantum Circuit SimulationabstractQuantum circuit simulation is critical for verifying quantum computers. Given exponential complexity in the simulation, existing simulators use different architectures to accelerate the simulation. However, due to the variety of both simulation methods and modern architectures, it is challenging to design a high-performance yet portable simulator. In this work, we propose UniQ, a unified programming model for multiple simulation methods on various hardware architectures. We provide a unified application abstraction to describe different applications, and a unified hierarchical hardware abstraction upon different hardware. Based on these abstractions, UniQ can perform various circuit transformations without being aware of either concrete application or architecture detail, and generate high-performance execution schedules on different platforms without much human effort. Evaluations on CPU, GPU, and Sunway platforms show that UniQ can accelerate quantum circuit simulation by up to 28.59× (4.47× on average) over state-of-the-art frameworks, and successfully scale to 399,360 cores on 1,024 nodes. Chen Zhang 0001, Haojie Wang 0004, Zixuan Ma, Jidong Zhai |
SC | 2 |
| 2022 | $TC-Stream$TC-Stream: Large-Scale Graph Triangle Counting on a Single Machine Using GPUsabstractIn this paper, we build a TC-Stream, a high-performance graph processing system specific for a triangle counting algorithm on graph data with up to tens of billions of edges, which significantly exceeds the device memory capacity of Graphics Processing Units (GPUs). The triangle counting problem is a broad research topic in data mining and social network analysis in the graph processing field. As the scale of the graph data grows, a portion of the graph data must be loaded iteratively. To solve the above problem, we propose TC-Stream. It focuses on three issues: 1) For power-law graphs, because the amount of tasks of each vertex or edge is inconsistent, it is bound to cause different demands of computing and memory resources for different task types. We propose a parallel vertex approach and the reordering of vertices for graph data that can be placed in the GPU device memory to ensure the maximum workload balancing; 2) A binary-search-based set intersection method is designed to achieve the maximum parallelism in GPU; 3) For the graph data that exceeds the GPU device memory capacity, we develop a novel vertical partition algorithm to guarantee the independent computing on each partition so that the three computation processes, i.e., the computation on GPU, the data transmission between main memory of CPU and SSD, and the communication between the CPU and the GPU can be perfectly overlapped. Extensive experiments conducted on large-scale datasets showed that the TC-stream running on a single Tesla V100 GPU performs 2.4 6 and 1.8 4.4 faster than the state-of-the-art single-machine in-memory triangle counting system and GPU-based triangle counting system, respectively, and achieves 2.4faster than the state-of-the-art out-of-core distributed system PDTL running on an 8-node cluster when processing the graph data with 42.5 billion edges. Jianqiang Huang 0002, Haojie Wang 0004, Xiaoying Wang 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Detecting Performance Variance for Parallel Applications Without Source CodeabstractFor parallel applications, performance variance is a critical issue that can degrade performance and make applications’ behavior difficult to explain. Therefore, users and application developers should be able to detect and diagnose performance variance. Previous detection methods either introduce too much overhead and slow down applications, or rely on nontrivial source code analysis, which is impractical for production-run parallel systems. In this article, we proposeVapro, a framework for detecting and diagnosing performance variance in production-run parallel systems. Our method is based on an observation that most parallel programs contain code snippets that are executed repeatedly with a fixed workload and can be utilized to detect performance variance. We present State Transition Graph (STG) to track program execution and then do light-weight workload analysis on STG to locate performance variance.Vaprois able to successfully identify these snippets at runtime even without program source code. To diagnose the discovered variation,Vaprouses a progressive diagnosis method based on a hybrid model combining variance breakdown and statistical analysis. According to evaluating results,Vapro's performance overhead is only 1.38% on average.Vaprocan identify performance variance in real applications caused by hardware issues, such as memory and IO. The standard deviation of the execution time is decreased by up to 73.5% when the identified variance is fixed.Vaproachieves 30.0% larger detection coverage than the state-of-the-art variance detection approach based on source code analysis. Jidong Zhai, Liyan Zheng 0001, Feng Zhang 0007, Xiongchao Tang, Haojie Wang 0004, Yuyang Jin 0001, Shuaiwen Song |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | Sparker: Efficient Reduction for More Scalable Machine Learning with SparkabstractMachine learning applications on Spark suffers from poor scalability. In this paper, we reveal that the key reasons is the non-scalable reduction, which is restricted by the non-splittable object programming interface in Spark. This insight guides us to propose Sparker, Spark with Efficient Reduction. By providing a split aggregation interface, Sparker is able to perform split aggregation with scalable reduction while being backward compatible with existing applications. We implemented Sparker in 2,534 lines of code. Sparker can improve the aggregation performance by up to 6.47 × and can improve the end-to-end performance of MLlib model training by up to 3.69 × with a geometric mean of 1.81 × . Bowen Yu 0003, Huanqi Cao, Tianyi Shan, Haojie Wang 0004, Xiongchao Tang |
ICPP | 4 |
| 2021 | HyQuas: hybrid partitioner based quantum circuit simulation system on GPUabstractQuantum computing has shown its strong potential in solving certain important problems. Due to the intrinsic limitations of current real quantum computers, quantum circuit simulation still plays an important role in both research and development of quantum computing. GPU-based quantum circuit simulation has been explored due to GPU's high computation capability. Despite previous efforts, existing quantum circuit simulation systems usually rely on a single method to improve poor data locality caused by complex quantum entanglement. However, we observe that existing simulation methods show significantly different performance for different circuit patterns. The optimal performance cannot be obtained only with any single method. Chen Zhang 0001, Haojie Wang 0004, Kaiyuan Rong, Jidong Zhai |
ICS | 3 |
| 2021 | PET: Optimizing Tensor Programs with Partially Equivalent Transformations and Automated Corrections
Haojie Wang 0004, Jidong Zhai, Mingyu Gao 0001, Zixuan Ma, Shizhi Tang, Liyan Zheng 0001, Yuanzhi Li, Kaiyuan Rong, Yuanyong Chen |
OSDI | 1 |
| 2021 | Chukonu: A Fully-Featured Big Data Processing System by Efficiently Integrating a Native Compute Engine into SparkabstractApache Spark is a widely deployed big data analytics framework that offers such attractive features as resiliency, load-balancing, and a rich ecosystem. However, there is still plenty of room for improvement in its performance. Although a data-parallel system in a native programming language significantly improves performance, it may require re-implementing many functionalities of Spark to become a full-featured system. It is desirable for native big data systems to just write a compute engine in native languages to ensure high efficiency, and reuse other mature features provided by Spark rather than re-implement everything. But the interaction between the JVM and the native world risks becoming a bottleneck. This paper proposes Chukonu, a native big data framework that re-uses critical big data features provided by Spark. Owing to our novel DAG-splitting approach, the potential Spark integration overhead is alleviated, and its even outperforms existing pure native big data frameworks. Chukonu splits DAG programs into run-time parts and compile-time parts: The run-time parts are delegated to Spark to offload the complexities due to feature implementations. The compile-time parts are natively compiled. We propose a series of optimization techniques to be applied to the compile-time parts, such as operator fusion, vectorization, and compaction, to significantly reduce the Spark integration overhead. The results of evaluation show that Chukonu has a speedup of up to 71.58X (geometric mean 6.09X) over Apache Spark, and up to 7.20X (geometric mean 2.30X) over pure-native frameworks on six commonly-used big data applications. By translating the physical plan produced by SparkSQL into Chukonu programs, Chukonu accelerates Spark-SQL's TPC-DS performance by 2.29X. Bowen Yu 0003, Guanyu Feng, Huanqi Cao, Zhenbo Sun, Haojie Wang 0004, Xiaowei Zhu 0001 |
Proc. VLDB Endow. | 6 |
| 2020 | Identifying scalability bottlenecks for large-scale parallel programs with graph analysisabstractScaling a parallel program to modern supercomputers is challenging due to inter-process communication, code serialization, and resource contention. Performance analysis tools for finding such scaling bottlenecks either base on profiling or tracing. Profiling incurs lower overheads but does not capture detailed dependencies needed for root-cause analyses. Tracing collects all information at prohibitive overheads. In this work, we develop ScalAna that uses static analysis techniques to achieve the best of both worlds---it enables the analyzability of traces at a cost similar to profiling. We leverage compiler and runtime lightweight techniques to generate performance graph and perform graph analysis algorithm to detect the root cause of scaling issues. We evaluate ScalAna with real applications on the Tianhe-2 supercomputer. Results show that our approach can effectively locate the root cause of scalability bottlenecks for real applications and incur less than 6.38% overhead (1.89% on average) for up to 2,048 processes. Yuyang Jin 0001, Haojie Wang 0004, Xiongchao Tang, Torsten Hoefler, Xu Liu 0001, Jidong Zhai |
PPoPP | 2 |
| 2020 | ScalAna: automating scaling loss detection with graph analysisabstractScaling a parallel program to modern supercomputers is challenging due to inter-process communication, Amdahl’s law, and resource contention. Performance analysis tools for finding such scaling bottlenecks either base on profiling or tracing. Profiling incurs low overheads but does not capture detailed dependencies needed for root-cause analysis. Tracing collects all information at prohibitive overheads. In this work, we design SCALANA that uses static analysis techniques to achieve the best of both worlds - it enables the analyzability of traces at a cost similar to profiling. SCALANA first leverages static compiler techniques to build a Program Structure Graph, which records the main computation and communication patterns as well as the program’s control structures. At runtime, we adopt lightweight techniques to collect performance data according to the graph structure and generate a Program Performance Graph. With this graph, we propose a novel approach, called backtracking root cause detection, which can automatically and efficiently detect the root cause of scaling loss. We evaluate SCALANA with real applications. Results show that our approach can effectively locate the root cause of scaling loss for real applications and incurs 1.73parcent overhead on average for up to 2,048 processes. We achieve up to 11.11parcent performance improvement by fixing the root causes detected by SCALANA on 2,048 processes. Yuyang Jin 0001, Haojie Wang 0004, Xiongchao Tang, Torsten Hoefler, Xu Liu 0001, Jidong Zhai |
SC | 2 |
| 2019 | Spread-n-share: improving application performance and cluster throughput with resource-aware job placementabstractTraditional batch job schedulers adopt the Compact-n-Exclusive (CE) strategy, packing processes of a parallel job into as few compute nodes as possible. While CE minimizes inter-node network communication, it often brings self-contention among tasks of a resource-intensive application. Recent studies have used virtual containers to balance CPU utilization and memory capacity across physical nodes, but the imbalance in cache and memory bandwidth usage is still under-investigated. Xiongchao Tang, Haojie Wang 0004, Xiaosong Ma, Nosayba El-Sayed, Jidong Zhai, Ashraf Aboulnaga |
SC | 2 |
| 2018 | Spindle: Informed Memory Access Monitoring
Haojie Wang 0004, Jidong Zhai, Xiongchao Tang, Bowen Yu 0003, Xiaosong Ma |
USENIX ATC | 1 |