Chengrui Zhang

dblp:88/5845 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dolphium: Co-Optimizing Quantization Dataflow and Paradigms on Poly-Hierarchical NPUs
abstract
Poly-hierarchical NPUs integrate distributed memory modules with heterogeneous computation units, posing significant challenges for mapping quantized operators. The difficulty arises from the need to coordinate data transfers across memory hierarchies and to assign diverse operations to suitable computation units. In this work, we systematically construct the mapping space from quantization to dataflows by addressing three key aspects: generation of NPU-friendly computation flows, integrated operation–data co-mapping, and determination of transfer granularity and frequency. Building on this foundation, we further exploit quantization dataflows to guide the selection of quantization paradigms. Compared with the state-of-the-art quantization compiler, our mapping achieves a 1.67–2.03× speedup. Moreover, the selected quantization paradigms deliver an average 2.18× efficiency improvement on NPUs without accuracy loss.
Xiuping Cui, Chengrui Zhang, Yun Liang 0001
DATE2
2026 LATIAS: A General Architecture-Operator Model for Spatial Accelerators with Complex Topology and Memory Hierarchy
abstract
Spatial accelerators are widely deployed for deep neural networks, but their architectural diversity—from hierarchical to dataflow designs—makes accurate architecture–operator modeling difficult, limiting operator optimization and hardware utilization. Existing models abstract hardware as hierarchical chains and operators as loop trees, which cannot capture essential features of modern dataflow accelerators, including heterogeneous processing elements (PEs), uni-directional interconnects, and cross-PE memory hierarchies, leading to inaccurate latency prediction. We propose LATIAS, a unified framework that introduces (1) an architecture graph with uni-directional edges to represent arbitrary topologies, and (2) a dataflow-aware tile-centric notation that augments loop trees with transfer nodes to model diverse dataflows. Building on these, LATIAS further provides a graph-guided tree analysis that accurately resolves tensor residency and latency under hardware constraints. Experiments on representative operators (GEMM, vector, fused vector) and operator shapes extracted from DNNs (BERT, ViT, T5) on Huawei Ascend 910B3 show that LATIAS achieves over 0.99 correlation with runtime measurements—substantially outperforming prior models—and provides actionable insights for architectural design.
Chengrui Zhang, Liancheng Jia, Renze Chen, Xiuping Cui, Size Zheng 0001, Shengen Yan, Yu Wang 0002, Yun Liang 0001
DATE1
2026 Embodied dataset enhanced reachability-aware 6-DoF grasp detection
Tongjia Zhang, Tianliang Hu, Chengrui Zhang, Shuai Ji
Expert Syst. Appl.3
2026 FSTL-Net: A few-shot welding defect detection method based on temporal modeling and metric learning
Shuai Ji, Yisheng Yin, Chengrui Zhang, Tieshuang Zhu
Knowl. Based Syst.4
2025 MReg: A Novel Regression Model with MoE-Based Video Feature Mining for Mitral Regurgitation Diagnosis
Yuhao Huang 0001, Chengrui Zhang, Haotian Lin 0009, Tong Han, Ruiyue Chen, Dong Ni 0001, Zhongshan Gou, Xin Yang 0009
MICCAI (9)4
2024 MAGIS: Memory Optimization via Coordinated Graph Transformation and Scheduling for DNN
abstract
Recently, memory consumption of Deep Neural Network (DNN) rapidly increases, mainly due to long lifetimes and large shapes of tensors. Graph scheduling has emerged as an effective memory optimization technique, which determines the optimal execution, re-computation, swap-out, and swap-in timings for each operator/tensor. However, it often hurts performance significantly and can only manipulate tensors' lifetimes but not shapes, limiting the optimization space. We find that graph transformation, which can change the tensor shapes and graph structure, creates a new trade-off space between memory and performance. Nevertheless, graph transformation are applied separately so far, with primary focus on optimizing performance and not memory.
Renze Chen, Zijian Ding, Size Zheng 0001, Chengrui Zhang, Jingwen Leng, Xuanzhe Liu, Yun Liang 0001
ASPLOS (3)4
2024 Exploiting Substitution Box for Cryptanalyzing Image Encryption Schemes With DNA Coding and Nonlinear Dynamics
abstract
In recent years, a number of image encryption schemes based on DNA coding and nonlinear dynamics have been proposed. Generally, these DNA-based schemes first encode plaintext images into DNA sequences and then encrypt them with pseudorandom elements produced by chaotic systems or other nonlinear dynamics. Although ciphertexts can pass some security tests, many image encryption schemes are being shown to have intrinsic flaws and that they cannot guarantee a high level of security. In this article, we cryptanalyze a family of image encryption schemes for which the encryption kernel is DNA coding or its variant. The complex DNA operation can be simplified as a substitution box (S-box). The whole cryptosystem's security level is thus significantly decreased and is vulnerable to the chosen-plaintext attack. Applications of this concept to break five ciphers are theoretically presented and experimentally verified. In addition, some suggestions for resisting similar attacks are also given in this article.
Chengrui Zhang, Junxin Chen 0001, Dongming Chen, Wei Wang 0077, Yushu Zhang 0001, Yicong Zhou
IEEE Trans. Multim.1
2023 Robot suction region prediction method from knowledge to learning in disordered manufacturing scenarios
Tongjia Zhang, Chengrui Zhang, Shuai Ji, Tianliang Hu
Eng. Appl. Artif. Intell.2
2022 Thermal-Aware Layout Optimization and Mapping Methods for Resistive Neuromorphic Engines
abstract
Resistive neuromorphic engines can accelerate spiking neural network tasks with memristor crossbars. However, the stored weight is influenced by the temperature, which leads to accuracy and endurance degradation. The higher the temperature is, the larger the influence is. In this work, we propose a cross-array mapping method and a layout optimization method to reduce the thermal effect with the consideration of input distribution, weight value and layout of memristor crossbars. Experimental results show that our method reduces the peak temperature up to 10.4K and improves the endurance up to 1.72×.
Chengrui Zhang, Pingqiang Zhou
ASP-DAC1
2006 Fuzzy Logic Thermal Error Compensation for Computer Numerical Control Noncircular Turnning System
abstract
With the new emerging technologies of high performance machining and the increasing demand for improved machining accuracy in recent years, the problem of thermal deformation of machine tool structures is becoming more critical than ever. The computer numerical control (CNC) turning system for noncircular section pistons is designed with giant magnetostrictive actuator (GMA) as the turning module. When the temperature of the cooler system of the GMA varies for about 6degC, the dimension error of the surface contour varies for about 20 micron, and cannot meet the precision requirement of the piston's contour dimension. In this paper, a method using fuzzy logic control for the compensation of the thermally induced error is developed. The fuzzy rule is used to compensate directly for the nonlinearity and uncertainty of the cooler system. The rule development strategy for the compensation system is to change the feed quantity of GMA to control the tool the pistons dimension. The fuzzy logic control developed here is a two-input single-output controller. The two inputs are the temperature deviation from setpoint error, and error rate. The output is the compensated value of the feed system. The triangular membership functions are used to define the input linguistic variables and the output linguistic variables. The fuzzy logic approach incorporates many advantages of using fuzzy logic such as the incorporation of heuristic knowledge, ease of implementation and the lack of a need for an accurate mathematical model. The experimental results are presented and the effectiveness of the fuzzy thermal error compensation control technique is discussed. Accuracy is greatly improved using the developed error compensation system
Hong'en Wu, Guili Li, Daguang Shi, Chengrui Zhang
ICARCV4