Zhongcheng Zhang

dblp:84/248 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0003-4835-7650ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-authorSystems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 39% Processor architecture and microarchitecture · 30% Parallel and multicore computing · 30%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.712023
Portable and Scalable All-Electron Quantum Perturbation Simulations on Exascale Supercomputers · SC 2023
Processor architecture and microarchitecture
SIMD
0.712023
Occamy: Elastically Sharing a SIMD Co-processor across Multiple CPU Cores · ASPLOS (3) 2023
Compilers and program optimization
vectorization
0.212023
Occamy: Elastically Sharing a SIMD Co-processor across Multiple CPU Cores · ASPLOS (3) 2023
High-performance computing › supercomputing
exascale computing
0.212023
Portable and Scalable All-Electron Quantum Perturbation Simulations on Exascale Supercomputers · SC 2023

Methods — techniques the papers use, named apart from their topics

task mapping · 1.3phase behavior analysis · 1.3hierarchical collective communication · 1.3dynamic lane partitioning · 1.3OpenCL · 1.3
YearPublicationVenuePosition
2025 Qiwu: Exploiting Ciphertext-Level SIMD Parallelism in Homomorphic Encryption Programs
abstract
Fully Homomorphic Encryption (FHE), particularly the CKKS scheme, enables computation on encrypted data, facilitating secure task offloading to untrusted servers. CKKS allows packing multiple complex values into a single ciphertext, crucial for fixed-point arithmetic in machine learning, while leveraging SIMD parallelism at the plaintext level. However, operations such as reductions can degrade performance by creating a large number of bubbles (or gaps) in intermediate ciphertexts, leading to wasted computational resources. We introduce Qiwu, a ciphertext-level vectorization approach that enhances performance by fusing multiple ciphertexts containing bubbles. Qiwu uses a DSL to specify zero bubbles in input ciphertexts and nonzero bubbles in output ciphertexts, employs data-flow analysis to track them, and formulates a fusion plan guided by a cost-benefit assessment. Implemented in an existing FHE compiler, Qiwu was evaluated on four applications (including three machine learning tasks) and three kernels. It achieves speedups of up to 18.0× on CPUs, averaging 3.4× (geometric mean), compared to the state-of-the-art compiler that exploit only plaintext-level parallelism.
Zhongcheng Zhang, Ying Liu 0055, Zhenchuan Chen, Xiaobing Feng 0002, Huimin Cui, Jingling Xue
CGO1
2023 Occamy: Elastically Sharing a SIMD Co-processor across Multiple CPU Cores
abstract
SIMD extensions are widely adopted in multi-core processors to exploit data-level parallelism. However, when co-running workloads on different cores, compute-intensive workloads cannot take advantage of the underutilized SIMD lanes allocated to memoryintensive workloads, reducing the overall performance. This paper proposes Occamy, a SIMD co-processor that can be shared by multiple CPU cores, so that their co-running workloads can spatially share its SIMD lanes. The key idea is to enable elastic spatial sharing by dynamically partitioning all the SIMD lanes across different workloads based on their phase behaviors, so that each workload may execute in variable-length SIMD mode. We also introduce an Occamy compiler to support such variable-length vectorization by analyzing such phase behaviors and generating the vectorized code that works with varying vector lengths. We demonstrate that Occamy can improve SIMD utilization, and consequently, performance over three representative SIMD architectures, with negligible chip area cost.
Zhongcheng Zhang, Yan Ou, Ying Liu 0055, Chenxi Wang 0005, Yongbin Zhou, Yucheng Ouyang, Jiahao Shan, Ying Wang 0001, Jingling Xue, Huimin Cui, Xiaobing Feng 0002
ASPLOS (3)1
2023 Portable and Scalable All-Electron Quantum Perturbation Simulations on Exascale Supercomputers
abstract
Quantum perturbation theory is pivotal in determining the critical physical properties of materials. The first-principles computations of these properties have yielded profound and quantitative insights in diverse domains of chemistry and physics. In this work, we propose a portable and scalable OpenCL implementation for quantum perturbation theory, which can be generalized across various high-performance computing (HPC) systems. Optimal portability is realized through the utilization of a cross-platform unified interface and a collection of performance-portable heterogeneous optimizations. Exceptional scalability is attained by addressing major constraints on memory and communication, employing a locality-enhancing task mapping strategy and a packed hierarchical collective communication scheme. Experiments on two advanced supercomputers demonstrate that our implementation exhibits remarkably performance on various material systems, scaling the system to 200,000 atoms with all-electron precision. This research enables all-electron quantum perturbation simulations on substantially larger molecular scales, with a potentially significant impact on progress in material sciences.
Zhikun Wu, Yangjun Wu, Ying Liu 0055, Honghui Shang, Yingxiang Gao, Zhongcheng Zhang, Yingchi Long, Xiaobing Feng 0002, Huimin Cui
SC6
2009 Allocation Method of Total Permitted Pollution Discharge Capacity Based on Uniform Price Auction
Congjun Rao, Zhongcheng Zhang, June Liu
ISNN (1)2
2009 Fuzzy Group Decision Making Method and Its Application
Zhongcheng Zhang, Congjun Rao
ISNN (1)2
2009 Alternating Iterative Projection Algorithm of Multivariate Time Series Mixed Models
Zhongcheng Zhang
ISNN (1)1