EDBT 2026 Demo / reviewers in the wild / expert
Zhen Jia 0001
dblp:06/180-1
· DBLP profile ↗
20ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-3543-2324ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorArtificial intelligence and machine learning · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context ParallelismabstractContext parallelism has emerged as a key technique to support long-context training, a growing trend in generative AI for modern large models. However, existing context parallel methods rely on static parallelization configurations that overlook the dynamic nature of training data, specifically, the variability in sequence lengths and token relationships (i.e., attention patterns) across samples. As a result, these methods often suffer from unnecessary communication overhead and imbalanced computation. In this paper, we present DCP, a dynamic context parallel training framework that introduces fine-grained blockwise partitioning of both data and computation. By enabling flexible mapping of data and computation blocks to devices, DCP can adapt to varying sequence characteristics, effectively reducing communication and improving memory and computation balance. Micro-benchmarks demonstrate that DCP accelerates attention by 1.19x~2.45x under causal masks and 2.15x~3.77x under sparse attention patterns. Additionally, we observe up to 0.94x~1.16x end-to-end training speed-up for causal masks, and 1.00x~1.46x for sparse masks. Chenyu Jiang 0002, Zhenkun Cai, Zhen Jia 0001, Yida Wang 0003, Chuan Wu 0001 |
SOSP | 4 |
| 2024 | DynaPipe: Optimizing Multi-task Training through Dynamic PipelinesabstractMulti-task model training has been adopted to enable a single deep neural network model (often a large language model) to handle multiple tasks (e.g., question answering and text summarization). Multi-task training commonly receives input sequences of highly different lengths due to the diverse contexts of different tasks. Padding (to the same sequence length) or packing (short examples into long sequences of the same length) is usually adopted to prepare input samples for model training, which is nonetheless not space or computation efficient. This paper proposes a dynamic micro-batching approach to tackle sequence length variation and enable efficient multi-task model training. We advocate pipelineparallel training of the large model with variable-length micro-batches, each of which potentially comprises a different number of samples. We optimize micro-batch construction using a dynamic programming-based approach, and handle micro-batch execution time variation through dynamic pipeline and communication scheduling, enabling highly efficient pipeline training. Extensive evaluation on the FLANv2 dataset demonstrates up to 4.39x higher training throughput when training T5, and 3.25x when training GPT, as compared with packing-based baselines. DynaPipe's source code is publicly available at https://github.com/awslabs/optimizing-multitask-training-through-dynamic-pipelines. Chenyu Jiang 0002, Zhen Jia 0001, Shuai Zheng 0004, Yida Wang 0003, Chuan Wu 0001 |
EuroSys | 2 |
| 2024 | Fast Convolution Meets Low Precision: Exploring Efficient Quantized Winograd Convolution on Modern CPUsabstractLow-precision computation has emerged as one of the most effective techniques for accelerating convolutional neural networks and has garnered widespread support on modern hardware. Despite its effectiveness in accelerating convolutional neural networks, low-precision computation has not been commonly applied to fast convolutions, such as the Winograd algorithm, due to numerical issues. In this article, we propose an effective quantized Winograd convolution, named LoWino, which employs an in-side quantization method in the Winograd domain to reduce the precision loss caused by transformations. Meanwhile, we present an efficient implementation that integrates well-designed optimization techniques, allowing us to fully exploit the capabilities of low-precision computation on modern CPUs. We evaluate LoWino on two Intel Xeon Scalable Processor platforms with representative convolutional layers and neural network models. The experimental results demonstrate that our approach can achieve an average of 1.84× and 1.91× operator speedups over state-of-the-art implementations in the vendor library while preserving accuracy loss at a reasonable level. Xueying Wang 0003, Guangli Li, Zhen Jia 0001, Xiaobing Feng 0002, Yida Wang 0003 |
ACM Trans. Archit. Code Optim. | 3 |
| 2023 | Performance Characterization of Large Language Models on High-Speed InterconnectsabstractLarge Language Models (LLMs) have recently gained significant popularity due to their ability to generate human-like text and perform a wide range of natural language processing tasks. Training these models usually requires a large amount of computational resources and is often done in a distributed manner. The use of high-speed interconnects can significantly influence the efficiency of distributed training. Therefore, there poses a need for systematic studies to explore the distributed training characteristics of these models on high-speed interconnects. This paper presents a comprehensive performance characterization of representative large language models: GPT, BERT, and T5. We evaluate their training performance in terms of iteration time, interconnect utilization, and scalability, over different high-speed interconnects and communication protocols, including TCP/IP, IPoIB, and RDMA. We observe that interconnects play a vital role in LLM training. Specifically, RDMA-100 Gbps outperforms IPoIB-100 Gbps and TCP/IP-10 Gbps by an average of 2.51x and 4.79x regarding training iteration time, and scores the highest interconnect utilization (up to 60 Gbps) in both strong and weak scaling, compared to IPoIB with up to 20 Gbps and TCP/IP with up to 9 Gbps, leading to the shortest training time. We also observe that larger models tend to have higher requirements for communication bandwidth, especially for AllReduce during backward propagation, which can take up to 91.12% of training time. Through our evaluation, we envision opportunities to improve the communication time for better training performance of LLMs. We extensively explore and summarize the role communication plays in distributed LLM training. Hao Qi 0008, Liuyao Dai, Weicong Chen 0002, Zhen Jia 0001, Xiaoyi Lu 0001 |
HOTI | 4 |
| 2023 | GEMINI: Fast Failure Recovery in Distributed Training with In-Memory CheckpointsabstractLarge deep learning models have recently garnered substantial attention from both academia and industry. Nonetheless, frequent failures are observed during large model training due to large-scale resources involved and extended training time. Existing solutions have significant failure recovery costs due to the severe restriction imposed by the bandwidth of remote storage in which they store checkpoints. Zhen Jia 0001, Shuai Zheng 0004, Zhen Zhang 0063, Xinwei Fu, T. S. Eugene Ng, Yida Wang 0003 |
SOSP | 2 |
| 2021 | LoWino: Towards Efficient Low-Precision Winograd Convolutions on Modern CPUsabstractLow-precision computation, which has been widely supported in contemporary hardware, is considered as one of the most effective methods to accelerate convolutional neural networks. However, low-precision computation is not widely used to speed up Winograd, an algorithm for fast convolution computation, due to the numerical error introduced by combining Winograd transformation and quantization. In this paper, we propose a low-precision Winograd convolution approach, LoWino, based on post-training quantization, which employs a linear quantization method in the Winograd domain to reduce the precision loss caused by transformations. Moreover, we present an efficient implementation that integrates well-designed optimization techniques, thereby adequately exploiting the capability of low-precision computation on modern CPUs. We evaluate our approach on Intel Xeon Scalable Processors by leveraging representative convolutional layers in prevailing deep neural networks. Experimental results show that LoWino achieves up to 2.04 × speedup over state-of-the-art implementations in the vendor library while maintaining the accuracy at a reasonable level. Guangli Li, Zhen Jia 0001, Xiaobing Feng 0002, Yida Wang 0003 |
ICPP | 2 |
| 2019 | The anatomy of efficient FFT and winograd convolutions on modern CPUsabstractWinograd-based convolution has quickly gained traction as a preferred approach to implement convolutional neural networks (ConvNet) on various hardware platforms because it could require fewer floating point operations than FFT-based or direct convolutions. Aleksandar Zlateski, Zhen Jia 0001, Kai Li 0001, Frédo Durand |
ICS | 2 |
| 2019 | RAGuard: An Efficient and User-Transparent Hardware Mechanism against ROP AttacksabstractControl-flow integrity (CFI) is a general method for preventing code-reuse attacks, which utilize benign code sequences to achieve arbitrary code execution. CFI ensures that the execution of a program follows the edges of its predefined static Control-Flow Graph: any deviation that constitutes a CFI violation terminates the application. Despite decades of research effort, there are still several implementation challenges in efficiently protecting the control flow of function returns (Return-Oriented Programming attacks). The set of valid return addresses of frequently called functions can be large and thus an attacker could bend the backward-edge CFI by modifying an indirect branch target to another within the valid return set. This article proposes RAGuard, an efficient and user-transparent hardware-based approach to prevent Return-Oreiented Programming attacks. RAGuard binds a message authentication code (MAC) to each return address to protect its integrity. To guarantee the security of the MAC and reduce runtime overhead: RAGuard (1) computes the MAC by encrypting the signature of a return address with AES-128, (2) develops a key management module based on a Physical Unclonable Function (PUF) and a True Random Number Generator (TRNG), and (3) uses a dedicated register to reduce MACs’ load and store operations of leaf functions. We have evaluated our mechanism based on the open-source LEON3 processor and the results show that RAGuard incurs acceptable performance overhead and occupies reasonable area. Rui Hou 0001, Wei Song 0002, Sally A. McKee, Zhen Jia 0001, Chen Zheng 0001, Mingyu Chen 0001, Lixin Zhang 0002, Dan Meng 0002 |
ACM Trans. Archit. Code Optim. | 5 |
| 2019 | Understanding Processors Design Decisions for Data Analytics in Homogeneous Data CentersabstractOur global economy increasingly depends on our ability to gather, analyze, link, and compare very large data sets. Keeping up with such big data poses challenges in terms of both computational performance and energy efficiency, and motivates different approaches to explore data center systems and architectures. To better understand the processor design decisions in context of data analytics in data centers, we conduct comprehensive evaluations using representative data analaytics workloads on representative conventional multi-core and many-core processors. After a comprehensive analysis of performance, power, energy efficiency and performance-cost efficiency, we have the following observations: contrasted with the conventional wisdom that uses wimpy many-core processors to improve energy-efficiency, the brawny multi-core processors with SMT (simultaneous multithreading) and dynamic overclocking technologies outperform the counterparts in terms of not only execution time, but also energy-efficiency for most of data analytics workloads in our experiments. Zhen Jia 0001, Wanling Gao, Yingjie Shi, Sally A. McKee, Zhenyan Ji, Jianfeng Zhan, Lei Wang 0004, Lixin Zhang 0002 |
IEEE Trans. Big Data | 1 |
| 2018 | CVR: efficient vectorization of SpMV on x86 processorsabstractSparse Matrix-vector Multiplication (SpMV) is an important computation kernel widely used in HPC and data centers. The irregularity of SpMV is a well-known challenge that limits SpMV’s parallelism with vectorization operations. Existing work achieves limited locality and vectorization efficiency with large preprocessing overheads. To address this issue, we present the Compressed Vectorization-oriented sparse Row (CVR), a novel SpMV representation targeting efficient vectorization. The CVR simultaneously processes multiple rows within the input matrix to increase cache efficiency and separates them into multiple SIMD lanes so as to take the advantage of vector processing units in modern processors. Our method is insensitive to the sparsity and irregularity of SpMV, and thus able to deal with various scale-free and HPC matrices. We implement and evaluate CVR on an Intel Knights Landing processor and compare it with five state-of-the-art approaches through using 58 scale-free and HPC sparse matrices. Experimental results show that CVR can achieve a speedup up to 1.70 × (1.33× on average) and a speedup up to 1.57× (1.10× on average) over the best existing approaches for scale-free and HPC sparse matrices, respectively. Moreover, CVR typically incurs the lowest preprocessing overhead compared with state-of-the-art approaches. Biwei Xie, Jianfeng Zhan, Xu Liu 0001, Wanling Gao, Zhen Jia 0001, Xiwen He, Lixin Zhang 0002 |
CGO | 5 |
| 2018 | Optimizing N-dimensional, winograd-based convolution for manycore CPUsabstractRecent work on Winograd-based convolution allows for a great reduction of computational complexity, but existing implementations are limited to 2D data and a single kernel size of 3 by 3. They can achieve only slightly better, and often worse performance than better optimized, direct convolution implementations. We propose and implement an algorithm for N-dimensional Winograd-based convolution that allows arbitrary kernel sizes and is optimized for manycore CPUs. Our algorithm achieves high hardware utilization through a series of optimizations. Our experiments show that on modern ConvNets, our optimized implementation, is on average more than 3 x, and sometimes 8 x faster than other state-of-the-art CPU implementations on an Intel Xeon Phi manycore processors. Moreover, our implementation on the Xeon Phi achieves competitive performance for 2D ConvNets and superior performance for 3D ConvNets, compared with the best GPU implementations. Zhen Jia 0001, Aleksandar Zlateski, Frédo Durand, Kai Li 0001 |
PPoPP | 1 |
| 2017 | CloudMix: Generating Diverse and Reducible Workloads for Cloud SystemsabstractThe prosperity of cloud computing offers common infrastructures to a wide range of applications. Understanding these applications' workload behaviors is the premise of designing, managing, and optimizing cloud systems. Considering the heterogeneity and diversity of cloud workloads, for the sake of fairness, cloud benchmarks must be able to accurately replicate their behaviors in cloud systems, including both the usages of cloud resources and the micro-architectural behaviors beyond the virtualization layer. Furthermore, workloads spanning long durations are usually required to achieve representativeness in evaluation. Hence the more challenging issue is to significantly reduce the evaluation duration while still preserving their workload characteristics. This paper presents our efforts towards generating cloud workloads of diverse behaviors and reducible durations. Our benchmark tool, CloudMix, employs a repository of reducible workload blocks (RWBs) as the high level abstraction of workload behaviors, including usages of the two most important cloud resources (CPU and memory) and their pairing micro-architectural operations. CloudMix further introduces an efficient methodology to combine RWBs to synthesize and replicate diverse cloud workloads in real-world traces. The effectiveness of CloudMix is demonstrated by generating a variety of reducible workloads according to a Google cluster trace and by applying these workloads in job scheduling optimization on Hadoop YARN. The evaluation results show: (i) when the workload durations are reduced by 100 times, the replication errors of workload behaviors are smaller than 2.08%; (ii) when providing fast evaluations (workload durations are reduced by 10 to 100 times) to recommend the optimal setting in YARN job scheduling, the performance degradation in the recommended setting is just 0.69% compared to that of the actual optimal setting. CloudMix is publicly available from the project home page http://prof.ict.ac.cn/BigDataBench/multi tenancy/. Rui Han 0001, Zan Zong, Fan Zhang 0047, José Luis Vázquez-Poletti, Zhen Jia 0001, Lei Wang 0004 |
CLOUD | 5 |
| 2017 | Understanding Big Data Analytics Workloads on Modern ProcessorsabstractBig data analytics workloads are very significant ones in modern data centers, and it is more and more important to characterize their representative workloads and understand their behaviors so as to improve the performance of data center computer systems. In this paper, we embark on a comprehensive study to understand the impacts and performance implications of the big data analytics workloads on the systems equipped with modern superscalar out-of-order processors. After investigating three most important application domains in Internet services in terms of page views and daily visitors, we choose 11 representative data analytics workloads and characterize their micro-architectural behaviors by using hardware performance counters. Our study reveals that the big data analytics workloads share many inherent characteristics, which place them in a different class from the traditional workloads and the scale-out services. To further understand the characteristics of big data analytics workloads, we perform correlation analysis to identify the most key factors that affect cycles per instruction (CPI). Also, we reveal that the increasing complexity of the big data software stacks will put higher pressures on the modern processor pipelines. Zhen Jia 0001, Jianfeng Zhan, Lei Wang 0004, Chunjie Luo, Wanling Gao, Rui Han 0001, Lixin Zhang 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Auto-tuning Spark Big Data Workloads on POWER8: Prediction-Based Dynamic SMT ThreadingabstractMuch research work devotes to tuning big data analytics in modern data centers, since %the truth that even a small percentage of performance improvement immediately translates to huge cost savings because of the large scale. Simultaneous multithreading (SMT) receives great interest from data center communities, as it has the potential to boost performance of big data analytics by increasing the processor resources utilization. For example, the emerging processor architectures like POWER8 support up to 8-way multithreading. However, as different big data workloads have disparate architectural characteristics, how to identify the most efficient SMT configuration to achieve the best performance is challenging in terms of both complex application behaviors and processor architectures. In this paper, we specifically focus on auto-tuning SMT configuration for Spark-based big data workloads on POWE-R8. However, our methodology could be generalized and extended to other programming software stacks and other architectures. Zhen Jia 0001, Guancheng Chen, Jianfeng Zhan, Lixin Zhang 0002, Yonghua Lin, H. Peter Hofstee |
PACT | 1 |
| 2016 | BDTUne: Hierarchical correlation-based performance analysis and rule-based diagnosis for big data systemsabstractAlthough big data systems are in widespread use and there have much research efforts for improving big data systems performance, efficiently analysing and diagnosing performance bottlenecks over these massively distributed systems remain a major challenge. In this paper, we propose a hierarchical correlation-based analysis and rule-based diagnostic approach for big data systems. The key approaches lie in identifying performance bottlenecks, classifying root causes, analyzing performance according to multi-level performance metrics, and setting diagnostic rules for performance tuning. Based on this approach, we have implemented BDTune - a lightweight, extensible and transparent tool that can provide valuable insights into performance of big data applications with a very low overhead. We also report our experience on how to use BDTune to conduct performance analysis and performance bottlenecks diagnosis, and demonstrate BDTune can help users find the performance bottlenecks and provide optimization recommendations. Zhen Jia 0001, Lei Wang 0004, Jianfeng Zhan, Tianxu Yi |
IEEE BigData | 2 |
| 2016 | Characterization and architectural implications of big data workloadsabstractThe previous major efforts on big data benchmark either propose a large amount of workloads (e.g. a recent comprehensive big data benchmark suite—BigDataBench [4]), which impose cognitive difficulty on workload characterization and serious benchmarking cost; or only select a few workloads according to so-called popularity[1], which lead to partial or biased observations. Lei Wang 0004, Jianfeng Zhan, Zhen Jia 0001 |
ISPASS | 4 |
| 2014 | BigDataBench: A big data benchmark suite from internet servicesabstractAs architecture, systems, and data management communities pay greater attention to innovative big data systems and architecture, the pressure of benchmarking and evaluating these systems rises. However, the complexity, diversity, frequently changed workloads, and rapid evolution of big data systems raise great challenges in big data benchmarking. Considering the broad use of big data systems, for the sake of fairness, big data benchmarks must include diversity of data and workloads, which is the prerequisite for evaluating big data systems and architecture. Most of the state-of-the-art big data benchmarking efforts target evaluating specific types of applications or system software stacks, and hence they are not qualified for serving the purposes mentioned above. This paper presents our joint research efforts on this issue with several industrial partners. Our big data benchmark suite-BigDataBench not only covers broad application scenarios, but also includes diverse and representative data sets. Currently, we choose 19 big data benchmarks from dimensions of application scenarios, operations/ algorithms, data types, data sources, software stacks, and application types, and they are comprehensive for fairly measuring and evaluating big data systems and architecture. BigDataBench is publicly available from the project home page http://prof.ict.ac.cn/BigDataBench. Also, we comprehensively characterize 19 big data workloads included in BigDataBench with varying data inputs. On a typical state-of-practice processor, Intel Xeon E5645, we have the following observations: First, in comparison with the traditional benchmarks: including PARSEC, HPCC, and SPECCPU, big data applications have very low operation intensity, which measures the ratio of the total number of instructions divided by the total byte number of memory accesses; Second, the volume of data input has non-negligible impact on micro-architecture characteristics, which may impose challenges for simulation-based big data architecture research; Last but not least, corroborating the observations in CloudSuite and DCBench (which use smaller data inputs), we find that the numbers of L1 instruction cache (L1I) misses per 1000 instructions (in short, MPKI) of the big data applications are higher than in the traditional benchmarks; also, we find that L3 caches are effective for the big data applications, corroborating the observation in DCBench. Lei Wang 0004, Jianfeng Zhan, Chunjie Luo, Yuqing Zhu 0001, Qiang Yang 0012, Yongqiang He, Wanling Gao, Zhen Jia 0001, Yingjie Shi, Chen Zheng 0001, Kent Zhan, Bizhu Qiu |
HPCA | 8 |
| 2012 | LogMaster: Mining Event Correlations in Logs of Large-Scale Cluster SystemsabstractThis paper presents a set of innovative algorithms and a system, named Log Master, for mining correlations of events that have multiple attributions, i.e., node ID, application ID, event type, and event severity, in logs of large-scale cloud and HPC systems. Different from traditional transactional data, e.g., supermarket purchases, system logs have their unique characteristics, and hence we propose several innovative approaches to mining their correlations. We parse logs into an n-ary sequence where each event is identified by an informative nine-tuple. We propose a set of enhanced apriori-like algorithms for improving sequence mining efficiency, we propose an innovative abstraction-event correlation graphs (ECGs) to represent event correlations, and present an ECGs-based algorithm for fast predicting events. The experimental results on three logs of production cloud and HPC systems, varying from 433490 entries to 4747963 entries, show that our method can predict failures with a high precision and an acceptable recall rates. Xiaoyu Fu, Jianfeng Zhan, Wei Zhou 0019, Zhen Jia 0001 |
SRDS | 5 |
| 2012 | CloudRank-D: benchmarking and ranking cloud computing systems for data processing applications
Chunjie Luo, Jianfeng Zhan, Zhen Jia 0001, Lei Wang 0004, Lixin Zhang 0002, Cheng-Zhong Xu 0001, Ninghui Sun |
Frontiers Comput. Sci. | 3 |
| 2012 | Precise, Scalable, and Online Request Tracing for Multitier Services of Black BoxesabstractAs more and more multitier services are developed from commercial off-the-shelf components or heterogeneous middleware without source code available, both developers and administrators need a request tracing tool to (1) exactly know how a user request of interest travels through services of black boxes and (2) obtain macrolevel user request behaviors of services without manually analyzing massive logs. This need is further exacerbated by IT system “agility,” which mandates the tracing tool to provide online performance data since offline approaches cannot reflect system changes in real time. Moreover, considering the large scale of deployed services, a pragmatic tracing approach should be scalable in terms of the cost in collecting and analyzing logs. In this paper, we introduce a precise, scalable, and online request tracing tool for multitier services of black boxes. Our contributions are threefold. First, we propose a precise request tracing algorithm for multitier services of black boxes, which only uses application-independent knowledge. Second, we present a microlevel abstraction, component activity graph, to represent causal paths of each request. On the basis of this abstraction, we use dominated causal path patterns to represent repeatedly executed causal paths that account for significant fractions, and we further present a derived performance metric of causal path patterns, latency percentages of components, to enable debugging performance-in-the-large. Third, we develop two mechanisms, tracing on demand and sampling, to significantly increase the system scalability. We implement a prototype of the proposed system, called PreciseTracer, and release it as open source code. In comparison with WAP5-a black-box tracing approach, PreciseTracer achieves higher tracing accuracy and faster response time. Our experimental results also show that PreciseTracer has low overhead, and still achieves high tracing accuracy even if an aggressive sampling policy is adopted, indicating that PreciseTracer is a promising tracing tool for large-scale production systems. Bo Sang, Jianfeng Zhan, Haining Wang 0001, Dongyan Xu, Lei Wang 0004, Zhen Jia 0001 |
IEEE Trans. Parallel Distributed Syst. | 8 |