Renwei Zhang

dblp:44/10920 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
7since 2021 · last 2024
0000-0002-9744-5676ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 4 since 2021Systems, architecture and hardware · 5 · 4 since 2021Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2024 PolyTOPS: Reconfigurable and Flexible Polyhedral Scheduler
abstract
Polyhedral techniques have been widely used for automatic code optimization in low-level compilers and higher-level processes. Loop optimization is central to this technique, and several polyhedral schedulers like Feautrier, Pluto, isl and Tensor Scheduler have been proposed, each of them targeting a different architecture, parallelism model, or application scenario. The need for scenario-specific optimization is growing due to the heterogeneity of architectures. One of the most critical cases is represented by NPUs (Neural Processing Units) used for AI, which may require loop optimization with different objectives. Another factor to be considered is the framework or compiler in which polyhedral optimization takes place. Different scenarios, depending on the target architecture, compilation environment, and application domain, may require different kinds of optimization to best exploit the architecture feature set. We introduce a new configurable polyhedral scheduler, PolyTOPS, that can be adjusted to various scenarios with straightforward, high-level configurations. This scheduler allows the creation of diverse scheduling strategies that can be both scenario-specific (like state-of-the-art schedulers) and kernel-specific, breaking the concept of a one-size-fits-all scheduler approach. PolyTOPS has been used with isl and CLooG as code generators and has been integrated in MindSpore AKG deep learning compiler. Experimental results in different scenarios show good performance: a geomean speedup of 7.66x on MindSpore (for the NPU Ascend architecture) hybrid custom operators over isl scheduling, a geomean speedup up to 1.80× on PolyBench on different multicore architectures over Pluto scheduling. Finally, some comparisons with different state-of-the-art tools are presented in the PolyMage scenario.
Gianpietro Consolaro, Harenome Razanajato, Nelson Lossing, Nassim Tchoulak, Adilla Susungi, Artur Cesar Araujo Alves, Renwei Zhang, Denis Barthou, Corinne Ancourt, Cédric Bastoul
CGO8
2024 Enabling Tensor Language Model to Assist in Generating High-Performance Tensor Programs for Deep Learning
Yi Zhai 0005, Keyu Pan, Renwei Zhang, Shuo Liu 0019, Zichun Ye, Jianmin Ji, Jie Zhao 0002, Yu Zhang 0086, Yanyong Zhang
OSDI4
2024 TSCompiler: efficient compilation framework for dynamic-shape models
Chenbo Geng, Yanzhi Yi, Renwei Zhang, Gianpietro Consolaro, Fan Yang 0001, Tun Lu, Ning Gu 0001
Sci. China Inf. Sci.6
2023 Modeling the Interplay between Loop Tiling and Fusion in Optimizing Compilers Using Affine Relations
abstract
Loop tiling and fusion are two essential transformations in optimizing compilers to enhance the data locality of programs. Existing heuristics either perform loop tiling and fusion in a particular order, missing some of their profitable compositions, or execute ad-hoc implementations for domain-specific applications, calling for a generalized and systematic solution in optimizing compilers. In this article, we present a so-called basteln (an abbreviation for backward slicing of tiled loop nests) strategy in polyhedral compilation to better model the interplay between loop tiling and fusion. The basteln strategy first groups loop nests by preserving their parallelism/tilability and next performs rectangular/parallelogram tiling to the output groups that produce data consumed outside the considered program fragment. The memory footprints required by each tile are then computed, from which the upward exposed data are extracted to determine the tile shapes of the remaining fusion groups. Such a tiling mechanism can construct complex tile shapes imposed by the dependences between these groups, which are further merged by a post-tiling fusion algorithm for enhancing data locality without losing the parallelism/tilability of the output groups. The basteln strategy also takes into account the amount of redundant computations and the fusion of independent groups, exhibiting a general applicability. We integrate the basteln strategy into two optimizing compilers, with one a general-purpose optimizer and the other a domain-specific compiler for deploying deep learning models. The experiments are conducted on CPU, GPU, and a deep learning accelerator to demonstrate the effectiveness of the approach for a wide class of application domains, including deep learning, image processing, sparse matrix computation, and linear algebra. In particular, the basteln strategy achieves a mean speedup of 1.8× over cuBLAS/cuDNN and 1.1× over TVM on GPU when used to optimize deep learning models; it also outperforms PPCG and TVM by 11% and 20%, respectively, when generating code for the deep learning accelerator.
Jie Zhao 0002, Jinchen Xu, Peng Di, Wang Nie, Yanzhi Yi, Zhen Geng, Renwei Zhang, Bojie Li, Zhiliang Gan, Xuefeng Jin 0004
ACM Trans. Comput. Syst.9
2022 Parallelizing Neural Network Models Effectively on GPU by Implementing Reductions Atomically
abstract
Due to the missing of a good orchestration of loop transformations, existing optimizing compilers for deploying neural networks on GPU either parallelize reductions ineffectively or miss the fusion opportunities with other operators. Neural network models thus exhibit sub-optimal performance on GPU. We present a practical approach called Panamera for the effective parallelization of reductions in neural networks on GPU. Panamera first leverages loop coalescing to flatten the loop dimensions of reductions, converting all reduction operators into canonical forms eligible for the polyhedral model. Next, Panamera uses polyhedral transformations to reduce the data movements caused by unfused reductions and perform multi-block hardware binding not considered by many compilers. Finally, Panamera embeds a highly optimized routine implemented using GPU atomic instructions, further improving the performance of neural network models while guaranteeing the correctness of parallel reductions. The experimental results demonstrate the effectiveness of our approach: for single operators our code obtains a mean speedup of 33.7×, 3.5×, 5.4× and 9.6× over cuDNN, CUB, TVM and Ansor, for sub-graphs our approach outperforms cuDNN, TVM and Ansor by 9.5×, 2.6× and 2.7×, and for end-to-end workloads, a tensor compiler integrated with our approach outperforms them by 122.5%, 19.3% and 15.2%.
Jie Zhao 0002, Cédric Bastoul, Yanzhi Yi, Wang Nie, Renwei Zhang, Zhen Geng, Chong Li 0003, Thibaut Tachon, Zhiliang Gan
PACT6
2022 Optimizing GPU Deep Learning Operators with Polyhedral Scheduling Constraint Injection
abstract
Automatic parallel code generation from high-level abstractions such as those manipulated by artificial intelligence and deep learning (AI/DL) frameworks heavily rely on compiler techniques for automatic parallelization and optimization. Many recent advances rely on the polyhedral framework for this task because of its ability to model and to apply a wide range of loop transformations. However, modeling the complexity of the target architecture and of efficient cost models to decide about the best transformation is in general out of reach for a framework based on linear/affine constraints. In this work, we propose to decouple the polyhedral framework into linear and non-linear components. We introduce the constraint tree abstraction which may be generated by a non-linear optimizer and injected to the polyhedral optimization process to build better solutions. We present how to benefit from such a mechanism to generate efficient codes for GPU in the context of AI/DL operators. Our constraint injection allows to drive the polyhedral scheduler towards efficient solutions for load/store vectorization relying both on memory coalescing and vector types. We implemented our scheduler supporting constraint injection and our constraint construction system within a production AI/DL framework. Experiments on well known neural networks show the efficiency of this approach with respect to state-of-the-art polyhedral scheduling for GPU.
Cédric Bastoul, Harenome Razanajato, Nelson Lossing, Adilla Susungi, Javier de Juan, Etienne Filhol, Baptiste Jarry, Gianpietro Consolaro, Renwei Zhang
CGO10
2021 AKG: automatic kernel generation for neural processing units using polyhedral transformations
abstract
Existing tensor compilers have proven their effectiveness in deploying deep neural networks on general-purpose hardware like CPU and GPU, but optimizing for neural processing units (NPUs) is still challenging due to the heterogeneous compute units and complicated memory hierarchy.
Jie Zhao 0002, Bojie Li, Wang Nie, Zhen Geng, Renwei Zhang, Xiong Gao, Zheng Li 0035, Peng Di, Xuefeng Jin 0004
PLDI5
2020 A lightweight and aggregated system for indoor/outdoor detection using smart devices
Zheng Qin 0003, Houbing Song, Chengxiang Si, Renwei Zhang
Future Gener. Comput. Syst.7
2019 Statically-Directed Assertion Recommendation for C Programs
abstract
Assertions are helpful in program analysis, such as software testing and verification. The oracles encoded in the assertions help detect potential flaws and release engineers from the manually check of the reported weaknesses, which is error-prone, burdensome and time-consuming. While in practice, few engineers would write assertions during programming, and it is challenging to generate assert statements, and insert them into proper locations automatically. In this paper, we propose a statically directed assertion recommendation approach for C programs. It combines static analysis, dynamic testing, and program verification to automatically recommend and validate weakness-oriented assertions, which is defined as an assert statement used to detect program weaknesses. Firstly, we integrate a static analysis tool such as FlawFinder and some learned patterns about CWE (Common Weakness Enumeration) to report potential program flaws. Secondly, we insert the corresponding assertions into the suspicious locations of those flaws. Then, we validate the program inserted with the assertions through two methods, the first is to execute the code with some test cases generated by automatic test-case generators such as Klee and Dart, and the second is to verify the program with some automatic verifier such as CPAchecker and Smack. Finally, we report on whether those flaws could be a real weakness. Experimental results show that our approach helps to find 125 real weaknesses in open source software from Github. Furthermore, our performance in detecting static analysis' true positives can reach 81.42%.
Cong Wang 0020, Renwei Zhang, Weiliang Ying
COMPSAC (1)3
2018 Fuzz testing in practice: Obstacles and solutions
abstract
Fuzz testing has helped security researchers and organizations discover a large number of vulnerabilities. Although it is efficient and widely used in industry, hardly any empirical studies and experience exist on the customization of fuzzers to real industrial projects. In this paper, collaborating with the engineers from Huawei, we present the practice of adapting fuzz testing to a proprietary message middleware named libmsg, which is responsible for the message transfer of the entire distributed system department. We present the main obstacles coming across in applying an efficient fuzzer to libmsg, including system configuration inconsistency, system build complexity, fuzzing driver absence. The solutions for those typical obstacles are also provided. For example, for the most difficult and expensive obstacle of writing fuzzing drivers, we present a low-cost approach by converting existing sample code snippets into fuzzing drivers. After overcoming those obstacles, we can effectively identify software bugs, and report 9 previously unknown vulnerabilities, including flaws that lead to denial of service or system crash.
Jie Liang 0006, Yuanliang Chen, Yu Jiang 0001, Renwei Zhang
SANER5
2018 A method for measuring the thermal geometric parameters of large hot rectangular forgings based on projection feature lines
Jinghao Yang, Wei Liu 0037, Renwei Zhang, Zhenyuan Jia, Fuji Wang
Mach. Vis. Appl.3
2016 Composite-based conflict resolution in merging versions of UML models
abstract
Model-driven engineering is now playing an essential role in software development. Adequate model versioning systems are critical to enable efficient team-based development of models. The state-of-art model versioning systems are able to detect and help resolving basic conflicts which arise during the merging of different model versions. However, conflict resolution is typically conducted at the primitive operation level in operation-based system and user interaction is required to choose from the conflicting operations. In this study, we present an approach to resolve conflicts automatically at composite level in model versioning systems for Unified Modeling Language (UML). This approach has two main stages. During the merging stage, a temporary merged model is generated, which represent the central intention of model developers. And during the conflict resolution stage, our approach automatically finds and presents to the model developers all solutions for resolving all inconsistencies in the merged model. The approach was empirically evaluated on a range of test models and proved to be scalable to models of large size.
Hao Chong, Renwei Zhang, Zheng Qin 0003
SNPD2