EDBT 2026 Demo / reviewers in the wild / expert
Qiao Sun 0005
dblp:10/6242-5
· DBLP profile ↗
12ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0002-8770-7847ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DH_Aligner: A fast short-read aligner on multicore platforms with AVX vectorization
Qiao Sun 0005, Leisheng Li, Huiyuan Li 0002 |
J. Parallel Distributed Comput. | 1 |
| 2024 | A novel HPL-AI approach for FP16-only accelerator and its instantiation on Kunpeng+Ascend AI-specific platform
Zijian Cao 0006, Qiao Sun 0005, Changcheng Song, Huiyuan Li 0002 |
J. Parallel Distributed Comput. | 2 |
| 2023 | Evolving the HPL benchmark towards multi-GPGPU clusters
Qiao Sun 0005, Wenjing Ma, Jiachang Sun, Huiyuan Li 0002 |
CCF Trans. High Perform. Comput. | 1 |
| 2020 | A Spatiotemporal Causality Based Governance Framework for Noisy Urban Sensory Data
Biying Yan, Chao Yang 0002, Qiao Sun 0005, Feng Chen 0009 |
J. Comput. Sci. Technol. | 4 |
| 2019 | Enabling Highly Efficient k-Means Computations on the SW26010 Many-Core Processor of Sunway TaihuLight
Chao Yang 0002, Qiao Sun 0005, Wenjing Ma, Wenlong Cao, Yulong Ao |
J. Comput. Sci. Technol. | 3 |
| 2018 | Distributed Parallel Simulation of Primary Sample Space Metropolis Light Transport
Changmao Wu, Changyou Zhang, Qiao Sun 0005 |
ICA3PP (1) | 3 |
| 2018 | Bandwidth Reduced Parallel SpMV on the SW26010 Many-Core PlatformabstractSpMV (Sparse Matrix-Vector multiplication), in its simplest form y = Ax, multiplies a sparse matrix with a dense vector and is a widely used computing primitive in the domain of HPC. On the newly SW26010 many-core platform, we propose a highly efficient CSR (Compressed Storage Row) based implementation of parallel SpMV, referred to as SWCSR-SpMV in the sequel. SpMV in the CSR format can be trivially parallelized but its performance is majorly impeded by memory access efficiency, and therefore to leverage high-throughput memory access mechanism while avoiding redundant bandwidth usage becomes the major goal of designing high performance SpMV on the target platform. The original problem is sequentially partitioned into row-slices, each of which can reside in the fast scratchpad memory, so that the loaded x'es can be reused; meanwhile, a dynamic look-ahead scheme is applied to avoid redundant memory access; we split the many-core mesh into smaller communication scope to facilitate the sharing of the common data across the working threads via the high speed on-mesh data bus. Beyond the above, to leverage massive parallelism balanced workload is ensured by both static and dynamic means. Performance evaluation is done on a benchmark of 36 frequently used sparse matrices in the fields of graph computing, data mining, computational fluid dynamics, etc.. While the performance upper-bound is defined by the ratio between the minimal data access volume required against the practically optimal bandwidth, ignoring the computing overhead, SWCSR-SpMV can achieve an efficiency of nearly 87%, maintaining over 75% for 1/3 of the testing matrices. SWCSR-SpMV is further applied in a PETSc based application, a 1.75x-2.6x speedup is sustained in a multi-process environment on the Sunway TaiHuLight supercomputer. Qiao Sun 0005, Changyou Zhang, Changmao Wu, Leisheng Li |
ICPP | 1 |
| 2018 | Performance Optimization of the HPCG Benchmark on the Sunway TaihuLight SupercomputerabstractIn this article, we present some key techniques for optimizing HPCG on Sunway TaihuLight and demonstrate how to achieve high performance in memory-bound applications by exploiting specific characteristics of the hardware architecture. In particular, we utilize a block multicoloring approach for parallelization and propose methods such as requirement-based data mapping and customized gather collective to enhance the effective memory bandwidth. Experiments indicate that the optimized HPCG code can sustain 77% of the theoretical memory bandwidth and scale to the full system of more than 10 million cores, with an aggregated performance of 480.8 Tflop/s and a weak scaling efficiency of 87.3%. Yulong Ao, Chao Yang 0002, Fangfang Liu 0004, Wanwang Yin, Lijuan Jiang, Qiao Sun 0005 |
ACM Trans. Archit. Code Optim. | 6 |
| 2018 | PEPS++: Towards Extreme-Scale Simulations of Strongly Correlated Quantum Many-Particle Models on Sunway TaihuLightabstractThe study of strongly frustrated magnetic systems has drawn great attentions from both theoretical and experimental physics. Efficient simulations of these models are essential for understanding their exotic properties. Here we present PEPS++, a novel computational paradigm for simulating frustrated magnetic systems and other strongly correlated quantum many-body systems. PEPS++ can accurately solve these models at the extreme scale with low cost and high scalability on modern heterogeneous supercomputers. We implement PEPS++ on Sunway TaihuLight based on a carefully designed tensor computation library for manipulating high-rank tensors and optimize it by invoking various high-performance matrix and tensor operations. By solving a 2D strongly frustrated$J_1$-$J_2$model with over ten million cores, PEPS++ demonstrates the capability of simulating strongly correlated quantum many-body problems at unprecedented scales with accuracy and time-to-solution far beyond the previous state of the art. Lixin He, Hong An, Chao Yang 0002, Junshi Chen 0003, Weihao Liang, Shao-Jun Dong, Qiao Sun 0005, Wenting Han, Yongjian Han, Wenjun Yao |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2017 | Towards Highly Efficient DGEMM on the Emerging SW26010 Many-Core ProcessorabstractThe matrix-matrix multiplication is an essential building block that can be found in various scientific and engineering applications. High-performance implementations of the matrix-matrix multiplication on state-of-the-art processors may be of great importance for both the vendors and the users. In this paper, we present a detailed methodology of implementing and optimizing the double-precision general format matrix-matrix multiplication (DGEMM) kernel on the emerging SW26010 processor, which is used to build the Sunway TaihuLight supercomputer. We propose a three level blocking algorithm to orchestrate data on the memory hierarchy and expose parallelism on different hardware levels, and design a collective data sharing scheme by using the register communication mechanism to exchange data efficiently among different cores. On top of those, further optimizations are done based on a data-thread mapping method for efficient data distribution, a double buffering scheme for asynchronous DMA data transfer, and an instruction scheduling method for maximizing the pipeline usage. Experiment results show that the proposed DGEMM implementation can fully exploit the unique hardware features provided by SW26010 and can sustain up to 95% of the peak performance. Lijuan Jiang, Chao Yang 0002, Yulong Ao, Wanwang Yin, Wenjing Ma, Qiao Sun 0005, Fangfang Liu 0004, Rongfen Lin |
ICPP | 6 |
| 2016 | Fast Parallel Stream Compaction for IA-Based Multi/many-core ProcessorsabstractStream compaction, frequently found in a large variety of applications, serves as a general primitive that reduces an input stream to a subset containing only the wanted elements so that the follow-on computation can be done efficiently. In this paper, we propose a fast parallel stream compaction for IA-based multi-/many-core processors. Unlike the previously studied algorithms that depend heavily on a black-box parallel scan, we open the black-box in the proposed algorithm and manually tailor it so that both the workload and the memory footprint is significantly reduced. By further eliminating the conditional statements and applying automatic code generation/optimization for performance-critical kernels, the proposed parallel stream compaction achieves high performance in different cases and for various data types across different IA-based multi/manycore platforms. Experimental results on three typical IA-based processors, including a quad-core Core-i7 CPU, a dual-socket 8-core Xeon CPU, and a 61-core Xeon Phi accelerator show that the proposed implementation outperforms the referenced parallel counterpart in the state-of-art library Thrust. On top of the above, we apply it in the random forest based data classifier to show its potential to boost the performance of real-world applications. Qiao Sun 0005, Chao Yang 0002, Changmao Wu, Leisheng Li, Fangfang Liu 0004 |
CCGrid | 1 |
| 2014 | Optimization of scan algorithms on multi- and many-core processorsabstractScan is a basic building block widely utilized in many applications. With the emergence of multi-core and many-core processors, the study of highly scalable parallel scan algorithms becomes increasingly important. In this paper, we first propose a novel parallel scan algorithm based on the fine grain dynamic task scheduling in QUARK, and then derive a cache-friendly framework for any parallel scan kernel. The QUARK-scan is superior to the fastest available counterpart proposed by Zhang in 2012 and many other parallel scans in several aspects, including the greatly improved load balance and the substantially reduced number of global barriers. On the other hand, the cache-friendly framework helps in improving the cache line usage and is flexible to apply to any parallel scan kernel. A variety of optimization techniques such as SIMD vectorization, loop unrolling, adjacent synchronization and thread affinity are exploited in QUARKscan and the cache-friendly versions of both QUARK-scan and Zhang's scan. Experiments done on three typical multi- and many-core platforms indicate that the proposed QUARK-scan and the cache-friendly Zhang's scan are superior in different scenarios. Qiao Sun 0005, Chao Yang 0002 |
HiPC | 1 |