VLDB 2026 Research / reviewers in the wild / expert
Zhen Geng
dblp:136/7681
· DBLP profile ↗
8ranked-venue papers
3as first author
3since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSystems, architecture and hardware · 2 · 2 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
loop transformation |
0.7 | 1 | 2023 | Modeling the Interplay between Loop Tiling and Fusion in Optimizing Compilers Using Affine Relations · ACM Trans. Comput. Syst. 2023 |
Compilers and program optimization › loop transformation
polyhedral compilation |
0.7 | 1 | 2023 | Modeling the Interplay between Loop Tiling and Fusion in Optimizing Compilers Using Affine Relations · ACM Trans. Comput. Syst. 2023 |
Compilers and program optimization › domain-specific compilation
tensor algebra compilation |
0.5 | 1 | 2021 | AKG: automatic kernel generation for neural processing units using polyhedral transformations · PLDI 2021 |
Compilers and program optimization › memory optimization
data locality optimization |
0.2 | 1 | 2023 | Modeling the Interplay between Loop Tiling and Fusion in Optimizing Compilers Using Affine Relations · ACM Trans. Comput. Syst. 2023 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit |
0.1 | 1 | 2021 | AKG: automatic kernel generation for neural processing units using polyhedral transformations · PLDI 2021 |
Methods — techniques the papers use, named apart from their topics
polyhedral compilation · 1.7post-tiling fusion · 0.7backward slicing · 0.7affine relations · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Modeling the Interplay between Loop Tiling and Fusion in Optimizing Compilers Using Affine RelationsabstractLoop tiling and fusion are two essential transformations in optimizing compilers to enhance the data locality of programs. Existing heuristics either perform loop tiling and fusion in a particular order, missing some of their profitable compositions, or execute ad-hoc implementations for domain-specific applications, calling for a generalized and systematic solution in optimizing compilers. In this article, we present a so-called basteln (an abbreviation for backward slicing of tiled loop nests) strategy in polyhedral compilation to better model the interplay between loop tiling and fusion. The basteln strategy first groups loop nests by preserving their parallelism/tilability and next performs rectangular/parallelogram tiling to the output groups that produce data consumed outside the considered program fragment. The memory footprints required by each tile are then computed, from which the upward exposed data are extracted to determine the tile shapes of the remaining fusion groups. Such a tiling mechanism can construct complex tile shapes imposed by the dependences between these groups, which are further merged by a post-tiling fusion algorithm for enhancing data locality without losing the parallelism/tilability of the output groups. The basteln strategy also takes into account the amount of redundant computations and the fusion of independent groups, exhibiting a general applicability. We integrate the basteln strategy into two optimizing compilers, with one a general-purpose optimizer and the other a domain-specific compiler for deploying deep learning models. The experiments are conducted on CPU, GPU, and a deep learning accelerator to demonstrate the effectiveness of the approach for a wide class of application domains, including deep learning, image processing, sparse matrix computation, and linear algebra. In particular, the basteln strategy achieves a mean speedup of 1.8× over cuBLAS/cuDNN and 1.1× over TVM on GPU when used to optimize deep learning models; it also outperforms PPCG and TVM by 11% and 20%, respectively, when generating code for the deep learning accelerator. Jie Zhao 0002, Jinchen Xu, Peng Di, Wang Nie, Yanzhi Yi, Zhen Geng, Renwei Zhang, Bojie Li, Zhiliang Gan, Xuefeng Jin 0004 |
ACM Trans. Comput. Syst. | 8 |
| 2022 | Parallelizing Neural Network Models Effectively on GPU by Implementing Reductions AtomicallyabstractDue to the missing of a good orchestration of loop transformations, existing optimizing compilers for deploying neural networks on GPU either parallelize reductions ineffectively or miss the fusion opportunities with other operators. Neural network models thus exhibit sub-optimal performance on GPU. We present a practical approach called Panamera for the effective parallelization of reductions in neural networks on GPU. Panamera first leverages loop coalescing to flatten the loop dimensions of reductions, converting all reduction operators into canonical forms eligible for the polyhedral model. Next, Panamera uses polyhedral transformations to reduce the data movements caused by unfused reductions and perform multi-block hardware binding not considered by many compilers. Finally, Panamera embeds a highly optimized routine implemented using GPU atomic instructions, further improving the performance of neural network models while guaranteeing the correctness of parallel reductions. The experimental results demonstrate the effectiveness of our approach: for single operators our code obtains a mean speedup of 33.7×, 3.5×, 5.4× and 9.6× over cuDNN, CUB, TVM and Ansor, for sub-graphs our approach outperforms cuDNN, TVM and Ansor by 9.5×, 2.6× and 2.7×, and for end-to-end workloads, a tensor compiler integrated with our approach outperforms them by 122.5%, 19.3% and 15.2%. Jie Zhao 0002, Cédric Bastoul, Yanzhi Yi, Wang Nie, Renwei Zhang, Zhen Geng, Chong Li 0003, Thibaut Tachon, Zhiliang Gan |
PACT | 7 |
| 2021 | AKG: automatic kernel generation for neural processing units using polyhedral transformationsabstractExisting tensor compilers have proven their effectiveness in deploying deep neural networks on general-purpose hardware like CPU and GPU, but optimizing for neural processing units (NPUs) is still challenging due to the heterogeneous compute units and complicated memory hierarchy. Jie Zhao 0002, Bojie Li, Wang Nie, Zhen Geng, Renwei Zhang, Xiong Gao, Zheng Li 0035, Peng Di, Xuefeng Jin 0004 |
PLDI | 4 |
| 2016 | ORB feature based web pornographic image recognition
Li Zhuo 0001, Zhen Geng, Jing Zhang 0023 |
Neurocomputing | 2 |
| 2015 | Fast Level-Set-Based Inverse Lithography Algorithm for Process Robustness Improvement and Its Application
Zhen Geng, Zheng Shi 0002, Xiaolang Yan, Kai-sheng Luo |
J. Comput. Sci. Technol. | 1 |
| 2014 | SVM based layout retargeting for fast and regularized inverse lithographyabstractInverse lithography technology (ILT), also known as pixel-based optical proximity correction (PB-OPC), has shown promising capability in pushing the current 193 nm lithography to its limit. By treating the mask optimization process as an inverse problem in lithography, ILT provides a more complete exploration of the solution space and better pattern fidelity than the traditional edge-based OPC. However, the existing methods of ILT are extremely time-consuming due to the slow convergence of the optimization process. To address this issue, in this paper we propose a support vector machine (SVM) based layout retargeting method for ILT, which is designed to generate a good initial input mask for the optimization process and promote the convergence speed. Supervised by optimized masks of training layouts generated by conventional ILT, SVM models are learned and used to predict the initial pixel values in the ‘undefined areas’ of the new layout. By this process, an initial input mask close to the final optimized mask of the new layout is generated, which reduces iterations needed in the following optimization process. Manufacturability is another critical issue in ILT; however, the mask generated by our layout retargeting method is quite irregular due to the prediction inaccuracy of the SVM models. To compensate for this drawback, a spatial filter is employed to regularize the retargeted mask for complexity reduction. We implemented our layout retargeting method with a regularized level-set based ILT (LSB-ILT) algorithm under partially coherent illumination conditions. Experimental results show that with an initial input mask generated by our layout retargeting method, the number of iterations needed in the optimization process and runtime of the whole process in ILT are reduced by 70.8% and 69.0%, respectively. Kai-sheng Luo, Zheng Shi 0002, Xiaolang Yan, Zhen Geng |
J. Zhejiang Univ. Sci. C | 4 |
| 2013 | A New Level-Set-Based Inverse Lithography Algorithm for Process Robustness Improvement with Attenuated Phase Shift MaskabstractInverse lithography technology (ILT) is one of the promising resolution enhancement techniques (RET), as the advanced integrated circuits (IC) technology nodes still use the 193nm light source. Among all the algorithms for ILT, the level-set-based ILT (LSB-ILT) is a feasible choice with good production result in practice. However, existing ILT algorithms optimize mask at nominal process condition without giving sufficient attention to the process variations, and thus the optimized masks show poor performance with focus and dose variations. In this paper, we put forward a new LSB-ILT algorithm for process robustness improvement with attenuated Phase Shift Mask (att-PSM) which is extensively used in the semiconductor foundries. In order to account for the process variations in the optimization, we adopt a new form of the cost function by adding the objective function of process variation band (PV band) to the nominal cost. The test patterns are from the M1 layer of a 28nm layout. Experimental results show that our new algorithm has a larger process window (PW) and reduces the process manufacturability index (PMI) by 41.37% compared with the LSB-ILT algorithm without PV band consideration. Zhen Geng, Zheng Shi 0002, Xiaolang Yan, Kai-sheng Luo |
CAD/Graphics | 1 |
| 2013 | Regularized level-set-based inverse lithography algorithm for IC mask synthesisabstractInverse lithography technology (ILT) is one of the promising resolution enhancement techniques, as the advanced IC technology nodes still use the 193 nm light source. In ILT, optical proximity correction (OPC) is treated as an inverse imaging problem to find the optimal solution using a set of mathematical approaches. Among all the algorithms for ILT, the level-set-based ILT (LSB-ILT) is a feasible choice with good production in practice. However, the manufacturability of the optimized mask is one of the critical issues in ILT; that is, the topology of its result is usually too complicated to manufacture. We put forward a new algorithm with high pattern fidelity called regularized LSB-ILT implemented in partially coherent illumination (PCI), which has the advantage of reducing mask complexity by suppressing the isolated irregular holes and protrusions in the edges generated in the optimization process. A new regularization term named the Laplacian term is also proposed in the regularized LSB-ILT optimization process to further reduce mask complexity in contrast with the total variation (TV) term. Experimental results show that the new algorithm with the Laplacian term can reduce the complexity of mask by over 40% compared with the ordinary LSB-ILT. Zhen Geng, Zheng Shi 0002, Xiaolang Yan, Kai-sheng Luo |
J. Zhejiang Univ. Sci. C | 1 |