Longshan Xu

dblp:364/3492 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0001-7759-7683ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Emerging computing paradigms · 72% Distributed systems · 28%
Theoretical computer science
1 paper
Quantum computing and quantum information · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
communication optimization
1.012026
Minimizing Communications of Quantum Circuit Simulations on Distributed Systems · IEEE Trans. Parallel Distributed Syst. 2026
Quantum computing and quantum information
quantum circuit simulation
1.012026
Minimizing Communications of Quantum Circuit Simulations on Distributed Systems · IEEE Trans. Parallel Distributed Syst. 2026
Emerging computing paradigms › quantum computer architecture
distributed quantum computing
0.912025
Optimizing Quantum Circuit Mapping to Reduce Inter-Module Communications in Distributed Architectures · SC 2025
Emerging computing paradigms
quantum computer architecture
0.912025
Optimizing Quantum Circuit Mapping to Reduce Inter-Module Communications in Distributed Architectures · SC 2025
Emerging computing paradigms › quantum computer architecture
qubit mapping
0.912025
Optimizing Quantum Circuit Mapping to Reduce Inter-Module Communications in Distributed Architectures · SC 2025
Quantum computing and quantum information
quantum circuit
0.312026
Minimizing Communications of Quantum Circuit Simulations on Distributed Systems · IEEE Trans. Parallel Distributed Syst. 2026

Methods — techniques the papers use, named apart from their topics

hybrid simulation · 2.0circuit slicing algorithm · 2.0hierarchical mapping · 0.9gate teleportation · 0.9circuit segmentation · 0.9
YearPublicationVenuePosition
2026 Minimizing Communications of Quantum Circuit Simulations on Distributed Systems
abstract
Efficient full-state quantum circuit simulations are useful tools for the design of quantum algorithms. Multi-node distributed systems are commonly employed as such simulations require a large amount of computation power and memory space. In distributed systems, communication overhead can be the performance bottleneck. This paper presents a distributed simulation framework called QuanTrans. A quantum circuit is composed of many levels of quantum gates. The simulation is conducted level by level. For circuits with particular structures, it employs a hybrid simulation approach to replace intermediate multi-level communications with one level of final merge operation, whose communication volume is comparable to that of one level of simulation in previous work. A circuit without such structures is sliced to find applicable sub-circuits with a single or multiple consecutive level(s). One level of communication is required for each sub-circuit, so we further propose a polynomial-time optimal circuit slicing algorithm. It can transform any circuit such that the number of sliced sub-circuits is the minimum after transformation. Experimental results show that QuanTrans can effectively reduce communication time and simulation time.
Longshan Xu, Edwin H.-M. Sha, Yuhong Song, Yunfan Chi, Qingfeng Zhuge
IEEE Trans. Parallel Distributed Syst.1
2025 Optimizing Quantum Circuit Mapping to Reduce Inter-Module Communications in Distributed Architectures
abstract
Modular quantum architectures have emerged as a promising solution for scalable quantum computing systems. Executing circuits in such distributed systems necessitates non-local operations between modules, incurring significant communication overhead. In this work, an optimized quantum circuit mapping technique called DQTetris is proposed to reduce inter-module communications. DQTetris employs a hierarchical framework that first seeks a global communication-free qubit mapping assignment under module capacity constraints. If infeasible, it searches for subcircuits with local communication-free qubit assignments via layer-wise gate pruning. Executing adjacent subcircuits with different qubit assignments incurs inter-module data teleportation. DQTetris minimizes these overheads by reducing qubit reassignment events through optimal circuit segmentation, qubit assignment selection, and adaptive gate teleportation. Experiments show that compared with existing methods, DQTetris can achieve average reductions in communication costs ranging from 28% to 75% across various benchmarks.
Longshan Xu, Edwin H.-M. Sha, Xiulin Cui, Qingfeng Zhuge
SC1
2025 MuDP: multi-granularity data placement for uniform loops on SPM-DRAM architectures to minimize latency
Edwin H.-M. Sha, Yuhong Song, Yibo Guo, Longshan Xu, Qingfeng Zhuge
Frontiers Comput. Sci.5
2024 Mera: Memory Reduction and Acceleration for Quantum Circuit Simulation via Redundancy Exploration
abstract
With the development of quantum computing, quantum processor demonstrates the potential supremacy in specific applications, such as Grover's database search and popular quantum neural networks (QNNs). For better calibrating the quantum algorithms and machines, quantum circuit simulation on classical computers becomes crucial. However, as the number of quantum bits (qubits) increases, the memory requirement grows exponentially. In order to reduce memory usage and accelerate simulation, we propose a multi-level optimization, namely Mera, by exploring memory and computation redundancy. First, for a large number of sparse quantum gates, we propose two compressed structures for low-level full-state simulation. The corresponding gate operations are designed for practical implementations, which are relieved from the longtime compression and decompression. Second, for the dense Hadamard gate, which is definitely used to construct the superposition, we design a customized structure for significant memory saving as a regularity-oriented simulation. Meanwhile, an ondemand amplitude updating process is optimized for execution acceleration. Experiments show that our compressed structures increase the number of qubits from 17 to 35, and achieve up to$6.9 \times$acceleration for QNN.
Yuhong Song, Edwin H.-M. Sha, Longshan Xu, Qingfeng Zhuge, Zili Shao
ICCD3