Chao-Yang Lu

dblp:213/7822 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0002-8227-9177ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Emerging computing paradigms · 61% High-performance computing · 30% GPUs and heterogeneous computing · 5%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
performance optimization at scale
0.812024
Surpassing Sycamore: Achieving Energetic Superiority Through System-Level Circuit Simulation · SC 2024
Emerging computing paradigms › quantum computer architecture
quantum circuit simulation
0.812024
Surpassing Sycamore: Achieving Energetic Superiority Through System-Level Circuit Simulation · SC 2024
Emerging computing paradigms
quantum computer architecture
0.812024
Surpassing Sycamore: Achieving Energetic Superiority Through System-Level Circuit Simulation · SC 2024
Emerging computing paradigms
quantum computing
0.612022
Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight · IEEE Trans. Parallel Distributed Syst. 2022
Emerging computing paradigms › quantum computing
quantum simulation
0.612022
Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight · IEEE Trans. Parallel Distributed Syst. 2022
High-performance computing
supercomputing
0.612022
Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight · IEEE Trans. Parallel Distributed Syst. 2022
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU scaling
0.212024
Surpassing Sycamore: Achieving Energetic Superiority Through System-Level Circuit Simulation · SC 2024
Performance modeling and evaluation
benchmarking
0.212022
Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight · IEEE Trans. Parallel Distributed Syst. 2022

Methods — techniques the papers use, named apart from their topics

tensor network optimization · 0.8system-level circuit simulation · 0.8parallel framework · 0.6multiple-precision fixed-point arithmetic · 0.6instruction scheduling · 0.6
YearPublicationVenuePosition
2024 Surpassing Sycamore: Achieving Energetic Superiority Through System-Level Circuit Simulation
abstract
In this paper, we present a groundbreaking largescale system technology that leverages optimization on global, node, and device levels to achieve unprecedented scalability for tensor networks. Our techniques enable accommodating largescale tensor networks with up to tens of terabytes of memory, reaching up to 2304 GPUs with a peak computing power of 561 PFLOPS. Notably, we have achieved a time-to-solution of 14.22 seconds with an energy consumption of 2.39 kWh which achieved a fidelity of 0.002. Our most remarkable result is a time-to-solution of 17.18 seconds, with energy consumption of only 0.29 kWh which achieved a XEB of 0.002 after post-processing. The experiments conducted demonstrate that our research outperforms Google’s quantum processor Sycamore in both speed and energy efficiency, which recorded 600 seconds and 4.3 kWh, respectively. The code is available at https://github.com/DeepLinkorg/OpenTenNet.
Zhongling Su, Han-Sen Zhong, Xiti Zhao, Jianyang Zhang, Xianhe Zhao, Ming-Cheng Chen, Chao-Yang Lu, Jian-Wei Pan, Zhilin Pei, Xingcheng Zhang, Wanli Ouyang
SC10
2022 Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight
abstract
Boson sampling is expected to be an important milestone that will demonstrate quantum computational advantage (or quantum supremacy). This work establishes the benchmarking of Gaussian boson sampling (GBS) with threshold detection based on the Sunway TaihuLight supercomputer. To achieve the best performance and provide a competitive scenario for future quantum computing studies, the selected simulation algorithm is fully optimized based on a set of innovative approaches, including a parallel framework with almost perfect load balance and an instruction-level optimizing scheme based on a shortest-path-based instruction scheduling. In addition, data precision is carefully processed by an integer-instruction-based and multiple-precision fixed-point implementation, including 128- and 256-bit precison mode, which can be appropriately selected based on an adaptive precision optimizing scheme. Based on these methods, a highly efficient parallel quantum sampling algorithm is designed. The largest run enables us to obtain one Torontonian function of a$100\times 100$submatrix from 50-photon GBS within 20 hours in 128-bit precision and 2 days in 256-bit precision. To our knowledge, this was the largest quantum computing simulation based on Boson Sampling by using modern supercomputers.
Lin Gan 0001, Mingcheng Chen, Yaojian Chen, Haitian Lu, Chao-Yang Lu, Jian-Wei Pan, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.6