Lijuan Jiang

dblp:33/6295 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 4 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2024 A Holistic Functionalization Approach to Optimizing Imperative Tensor Programs in Deep Learning
abstract
As deep learning empowers various fields, many domain-specific non-neural network operators have been proposed to improve the accuracy of deep learning models. Researchers often use the imperative programming diagram (PyTorch) to express these new operators, leaving the fusion optimization of these operators to deep learning compilers. Unfortunately, the inherent side effects introduced by imperative tensor programs, especially tensor-level mutations, often make optimization extremely difficult. Previous works either fail to eliminate the side effects of tensor-level mutations or require programmers to manually analyze and transform them. In this paper, we present a holistic functionalization approach (TensorSSA) to optimizing imperative tensor programs beyond control flow boundaries. We first introduce TensorSSA intermediate representation for removing tensor-level mutation and expanding the scope and ability of operator fusion. Based on TensorSSA IR, we propose a TensorSSA conversion algorithm that performs functionalization crossing the boundary of control flow. TensorSSA achieves a 1.79X (1.34X on average) speedup in representative deep learning tasks than state-of-the-art works.
Xingcheng Zhang, Shengen Yan, Yuting Chen 0001, Yueqian Zhang, Minxi Jin, Lijuan Jiang, Yun Liang 0001, Chao Yang 0002, Dahua Lin
DAC9
2023 xMath2.0: a high-performance extended math library for SW26010-Pro many-core processor
Fangfang Liu 0004, Wenjing Ma, Daokun Chen, Qinglin Lu, Wanwang Yin, Xinhui Yuan, Lijuan Jiang, Hongsen Wang, Chao Yang 0002
CCF Trans. High Perform. Comput.9
2023 Publisher Correction: xMath2.0: a high-performance extended math library for SW26010-Pro many-core processor
Fangfang Liu 0004, Wenjing Ma, Daokun Chen, Qinglin Lu, Wanwang Yin, Xinhui Yuan, Lijuan Jiang, Hongsen Wang, Chao Yang 0002
CCF Trans. High Perform. Comput.9
2022 EasyView: Enabling and Scheduling Tensor Views in Deep Learning Compilers
abstract
In recent years, memory-intensive operations are becoming dominant in efficiency of running novel neural networks. Just-in-time operator fusion on accelerating devices like GPU proves an effective method for optimizing memory-intensive operations, and suits the numerous varying model structures. In particular, we find memory-intensive operations on tensor views are ubiquitous in neural network implementations. Tensors are the de facto representation for numerical data in deep learning areas, while tensor views cover a bunch of sophisticated syntax, which allow various interpretations on the underlying tensor data without memory copy. The support of views in deep learning compilers could greatly enlarge operator fusion scope, and appeal to optimizing novel neural networks. Nevertheless, mainstream solutions in state-of-the-art deep learning compilers exhibit imperfections either in view syntax representations or operator fusion. In this article, we propose EasyView, which enables and schedules tensor views in an end-to-end workflow from neural networks onto devices. Aiming at maximizing memory utilization and reducing data movement, we categorize various view contexts in high-level language, and lower views in accordance with different scenarios. Reference-semantic in terms of views are kept in the lowering from native high-level language features to intermediate representations. Based on the reserved reference-semantics, memory activities related to data dependence of read and write are tracked for further compute and memory optimization. Besides, ample operator fusion is applied to memory-intensive operations with views. In our tests, the proposed work could get average 5.63X, 2.44X, and 4.67X speedup compared with the XLA, JAX, and TorchScript, respectively for hotspot Python functions. In addition, operation fusion with views could bring 8.02% performance improvement in end-to-end neural networks.
Lijuan Jiang, Qianchao Zhu, Shengen Yan, Xingcheng Zhang, Dahua Lin, Wenjing Ma, Zhouyang Li, Minxi Jin, Chao Yang 0002
ICPP1
2020 Enabling Highly Efficient Batched Matrix Multiplications on SW26010 Many-core Processor
abstract
We present a systematic methodology for optimizing batched matrix multiplications on SW26010 many-core processor of the Sunway TaihuLight supercomputer. Five surrogate algorithms and a machine learning–based algorithm selector are proposed to fully exploit the computing capability of SW26010 and cope with the sophisticated algorithm characteristics of batched matrix multiplications. Experiment results show that the algorithm selector is able to adaptively choose the appropriate algorithm for various matrix shapes and batch sizes with low overhead and high accuracy. In particular, the optimized batched matrix multiplications can substantially outperform the non-batched version and reach around 84.8% of the performance upper bound.
Lijuan Jiang, Chao Yang 0002, Wenjing Ma
ACM Trans. Archit. Code Optim.1
2019 Unilateral left-tail Anderson Darling test-based spectrum sensing with Laplacian noise
abstract
This study focuses on spectrum sensing under Laplacian noise. To mitigate the negative effects caused by the heavy‐tailed behaviour of Laplacian noise, the fractional lower order moments (FLOM) technology is employed to pre‐process the received samples before spectrum sensing. Through exploiting the asymmetrical difference between the distribution for the FLOM of received samples in the absence and presence of primary users, the authors formulate the spectrum sensing problem under Laplacian noise as a unilateral goodness‐of‐fit (GoF) test problem. Based on this test problem, they propose a new GoF‐based detector, which is called a unilateral left‐tail Anderson Darling (ULAD) detector. The analytical expressions for the theoretical performance, in terms of false‐alarm and detection probabilities, of the ULAD are derived. Moreover, a closed‐form expression for the optimal detection threshold is also derived to minimise the total error rate. Simulation results are provided to validate the theoretical analyses and to demonstrate the superior performance of the proposed detector than others.
Lijuan Jiang, Yongzhao Li, Yinghui Ye, Yunfei Chen 0001, Hailin Zhang 0001
IET Commun.1
2018 Performance Optimization of the HPCG Benchmark on the Sunway TaihuLight Supercomputer
abstract
In this article, we present some key techniques for optimizing HPCG on Sunway TaihuLight and demonstrate how to achieve high performance in memory-bound applications by exploiting specific characteristics of the hardware architecture. In particular, we utilize a block multicoloring approach for parallelization and propose methods such as requirement-based data mapping and customized gather collective to enhance the effective memory bandwidth. Experiments indicate that the optimized HPCG code can sustain 77% of the theoretical memory bandwidth and scale to the full system of more than 10 million cores, with an aggregated performance of 480.8 Tflop/s and a weak scaling efficiency of 87.3%.
Yulong Ao, Chao Yang 0002, Fangfang Liu 0004, Wanwang Yin, Lijuan Jiang, Qiao Sun 0005
ACM Trans. Archit. Code Optim.5
2017 Towards Highly Efficient DGEMM on the Emerging SW26010 Many-Core Processor
abstract
The matrix-matrix multiplication is an essential building block that can be found in various scientific and engineering applications. High-performance implementations of the matrix-matrix multiplication on state-of-the-art processors may be of great importance for both the vendors and the users. In this paper, we present a detailed methodology of implementing and optimizing the double-precision general format matrix-matrix multiplication (DGEMM) kernel on the emerging SW26010 processor, which is used to build the Sunway TaihuLight supercomputer. We propose a three level blocking algorithm to orchestrate data on the memory hierarchy and expose parallelism on different hardware levels, and design a collective data sharing scheme by using the register communication mechanism to exchange data efficiently among different cores. On top of those, further optimizations are done based on a data-thread mapping method for efficient data distribution, a double buffering scheme for asynchronous DMA data transfer, and an instruction scheduling method for maximizing the pipeline usage. Experiment results show that the proposed DGEMM implementation can fully exploit the unique hardware features provided by SW26010 and can sustain up to 95% of the peak performance.
Lijuan Jiang, Chao Yang 0002, Yulong Ao, Wanwang Yin, Wenjing Ma, Qiao Sun 0005, Fangfang Liu 0004, Rongfen Lin
ICPP1