EDBT 2026 Demo / reviewers in the wild / expert
Dexun Chen
dblp:184/6268
· DBLP profile ↗
20ranked-venue papers
0as first author
16since 2021 · last 2024
0000-0002-4524-4186ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 15 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Heterogeneous many-core optimization for Monte Carlo path-tracing on new generation Sunway HPC systemabstractAbstract We present swRender, a new parallel rendering pipeline based on the new Sunway many-core architecture (SW26010P) for the Monte Carlo path-tracing algorithm. Previous parallel rendering schemes are unsuitable for our task due to issues such as vast differences in hardware architectures and bottlenecks in I/O communication efficiency. To that end, we create a new two-level parallel tile rendering framework to fully utilize the Sunway computing resources, a practical tile-grouping load-balancing method to maintain the framework’s stability, and a novel many-core acceleration optimization to improve the rendering performance at the pixel level. Our method achieves (1) an average speedup of 16x in multiple benchmarks when compared to the baseline path-tracing model on the Sunway architecture, and (2) an average speedup of 2x when compared to state-of-the-art CPU, co-processor, and GPU-based parallel rendering approaches. Moreover, we scale swRender to run on 15 million cores and obtain high scalable parallel efficiency of 92%. Xinjie Wang 0003, Guanghao Ma, Jiaying Song, Mingyao Geng, Wenhui Hu, Xi Duan, Xiaogang Jin 0001, Dexun Chen, Maoxue Yu |
CCF Trans. High Perform. Comput. | 11 |
| 2023 | HadaFS: A File System Bridging the Local and Shared Burst Buffer for Exascale Supercomputers
Xiaobin He, Bin Yang 0043, Shupeng Shi, Dexun Chen, Wei Xue 0003, Zuoning Chen |
FAST | 7 |
| 2023 | Lifetime-Based Optimization for Simulating Quantum Circuits on a New Sunway SupercomputerabstractHigh-performance classical simulator for quantum circuits, in particular the tensor network contraction algorithm, has become an important tool for the validation of noisy quantum computing. In order to address the memory limitations, the slicing technique is used to reduce the tensor dimensions, but it could also lead to additional computation overhead that greatly slows down the overall performance. This paper proposes novel lifetime-based methods to reduce the slicing overhead and improve the computing efficiency, including, an interpretation method to deal with slicing overhead, an inplace slicing strategy to find the smallest slicing set and an adaptive tensor network contraction path refiner customized for Sunway architecture. Experiments show that in most cases the slicing overhead with our inplace slicing strategy would be less than the Cotengra, which is the most used graph path optimization software at present. Finally, the resulting simulation time is reduced to 96.1s for the Sycamore quantum processor RQC, with a sustainable single-precision performance of 308.6Pflops using over 41M cores to generate 1M correlated samples, which is more than 5 times performance improvement compared to 60.4 Pflops in 2021 Gordon Bell Prize work. Yaojian Chen, Xinmin Shi, Jiawei Song, Xin Liu 0081, Lin Gan 0001, Chu Guo, Haohuan Fu, Dexun Chen, Guangwen Yang 0002 |
PPoPP | 10 |
| 2023 | Enabling Real World Scale Structural Superlubricity All-Atom Simulation on the Next-Generation Sunway SupercomputerabstractMolecular dynamics (MD) simulation can provide an affordable way for inspecting microscopic phenomena, which is a powerful complement to real-world experiments. But the spatial scale of MD simulations is usually magnitudes smaller than experiment systems. In this paper, we present our work, redesigning the widely used inter-layer potential in structural superlubricity. By carrying out a specialized neighbor list for inter-layer potential computation, the total memory access amount is reduced significantly. Besides, a simple but efficient vectorization strategy is implemented based on the new neighbor list. In the extreme case, our work can scale to 38 million cores to achieve a sustainable performance of 61 PFLOPS, enabling a simulation of a superlubricity system of 32 μm2 with 7.2 billion atoms at 4.75 ns/day, which is 11,834 times of reported largest scale simulation in superlubricity systems in contact area and almost ten times faster in time-to-solution. Furthermore, we have done a simulation at 9 μm2 which results in consistency with real-world experiments and verified some theoretical predictions in the mesoscopic scale. Xiaohui Duan, Ping Gao 0005, Ming Ma 0012, Lin Gan 0001, Xin Liu 0081, Haohuan Fu, Wei Xue 0003, Dexun Chen, Guangwen Yang 0002 |
SC | 9 |
| 2023 | Establishing a Modeling System in 3-km Horizontal Resolution for Global Atmospheric Circulation triggered by Submarine Volcanic Eruptions with 400 Billion Smoothed Particle HydrodynamicsabstractPeople are increasingly concerned about how tectonic processes affect climate and vice versa. We establish a cross-sphere modeling system for volcanic eruptions and atmosphere circulation on a new Sunway supercomputer with a spatial resolution from 10m locally to 3km globally, using an improved multimedium and multiphase smoothed particle hydrodynamics (SPH) combined with a fully coupled meteorology-chemistry global atmospheric modeling scheme. We achieve 400 billion particles and 80% parallel efficiency using 39,000,000 processor cores. The simulation captures the whole dynamic process of the Tonga eruption from shock waves, earthquakes, tsunamis, mushroom clouds to the following 6--7 days of transport and diffusion of ash and water vapor, and preliminarily obtains the influence effect of full coupling of volcano, earthquake, ocean and atmosphere. This work is of great significance for deeply understanding the interaction between tectonic processes and climate change, and establishing an early warning simulation system for similar global hazard events. Shenghong Huang, Junshi Chen 0003, Ziyu Zhang 0003, Hong An, Yan Hu 0004, Zhanming Wang, Longkui Chen, Jineng Yao, Yang Zhao 0040, Dongning Jia, Changming Song, Xisheng Luo, Xiaobin He, Dexun Chen |
SC | 21 |
| 2023 | Scalability and efficiency challenges for the exascale supercomputing system: practice of a parallel supporting environment on the Sunway exascale prototype systemabstractWith the continuous improvement of supercomputer performance and the integration of artificial intelligence with traditional scientific computing, the scale of applications is gradually increasing, from millions to tens of millions of computing cores, which raises great challenges to achieve high scalability and efficiency of parallel applications on super-large-scale systems. Taking the Sunway exascale prototype system as an example, in this paper we first analyze the challenges of high scalability and high efficiency for parallel applications in the exascale era. To overcome these challenges, the optimization technologies used in the parallel supporting environment software on the Sunway exascale prototype system are highlighted, including the parallel operating system, input/output (I/O) optimization technology, ultra-large-scale parallel debugging technology, 10-million-core parallel algorithm, and mixed-precision method. Parallel operating systems and I/O optimization technology mainly support large-scale system scaling, while the ultra-large-scale parallel debugging technology, 10-million-core parallel algorithm, and mixed-precision method mainly enhance the efficiency of large-scale applications. Finally, the contributions to various applications running on the Sunway exascale prototype system are introduced, verifying the effectiveness of the parallel supporting environment design. Xiaobin He, Xin Chen 0023, Xin Liu 0081, Dexun Chen, Yuling Yang, Yunlong Feng, Longde Chen, Xiaona Diao, Zuoning Chen |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2023 | Coupled Incomplete Cholesky and Jacobi Preconditioned Conjugate Gradient on the New Generation of Sunway Many-Core ArchitectureabstractThe Preconditioned Conjugate Gradient method is one of the most important solvers in linear algebra system, and is widely used in scientific and engineering computing applications. Based on the Sunway heterogeneous many-core architecture, we propose a Coupled Incomplete Cholesky and Jacobi preconditioner (CICJ). The preconditioner applies a block Jacobi method to the matrix inversion in preconditioning process, localizes matrix inversions and completely eliminates the data correlation on slave cores. It strikes a better trade-off between convergence and parallelism than other preconditioners on the Sunway heterogeneous many-core architecture. Besides, a two-level software-controlled cache is designed for sparse matrix-vector multiplication operations, which makes full use of the Sunway heterogeneous many-core architecture. We apply our CICJ method on Intel, GPU, and Sunway, and the results show great generality on all three architectures. We also conduct experiments on the underwater submarine models using the open-source framework OpenFOAM. The results show that when the matrix column size is 0.82 billion and the number of non-zero values is 59 billion, our method accelerates the whole algorithm by 8.42 times compared with the diagonal incomplete Cholesky preconditioner (DIC) and 6.5 times compared with the geometric algebraic multi-grid preconditioner (GAMG) using 133,120 processors on the Sunway system. Yuejin Ye, Heng Guo 0006, Bingzhuo Wang, Pengxiao Wang, Dexun Chen |
IEEE Trans. Computers | 5 |
| 2022 | SeqDLM: A Sequencer-Based Distributed Lock Manager for Efficient Shared File Access in a Parallel File SystemabstractDistributed locks are used to guarantee the distributed client-cache coherence in parallel file systems. However, they lead to poor performance in the case of parallel writes under high-contention workloads. We analyze the distributed lock manager and find out that lock conflict resolution is the root cause of the poor performance, which involves frequent lock revocations and slow data flushing from client caches to data servers. We design a distributed lock manager named SeqDLM by exploiting the sequencer mechanism. SeqDLM mitigates the lock conflict resolution overhead using early grant and early revocation while keeping the same semantics as traditional distributed locks. To evaluate SeqDLM, we have implemented a parallel file system called ccPFS using both SeqDLM and traditional distributed locks. Evaluations on 96 nodes show SeqDLM outperforms the traditional distributed locks by up to$\boldsymbol{10.3}\times$for high-contention parallel writes on a shared file with multiple stripes. Shaonan Ma, Kang Chen 0001, Teng Ma 0006, Xin Liu 0081, Dexun Chen, Yongwei Wu 0001, Zuoning Chen |
SC | 6 |
| 2022 | Increasing the Efficiency of Massively Parallel Sparse Matrix-Matrix Multiplication in First-Principles Calculation on the New-Generation Sunway SupercomputerabstractThe first-principles approach based on density-functional theory (DFT)/density-functional perturbation theory (DFPT) is widely used in calculations of the systems’ ground state energy, response properties (e.g., polarizability, phonon dispersions) and is playing an increasingly important role in chemistry, physics and materials science. For the large-scale calculations, the computation of the density matrix/response density matrix in DFT/DFPT has become the main performance bottleneck. One of the solutions is using the linear scaling method to get the density matrix and response density matrix. Here a massively parallel medium sparse matrix-matrix multiplication algorithm is designed for first-principle calculations and implemented on the new-generation Sunway supercomputer. Experiments show that the proposed method has obvious performance advantages compared to the original parallel version under moderate sparsity. The computing cores scale to 3,900,000 with strong scalability of 77.3$\%$. Xin Chen 0023, Yingxiang Gao, Honghui Shang, Zhiqian Xu 0005, Xin Liu 0081, Dexun Chen |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2022 | Jdebug: A Fast, Non-Intrusive and Scalable Fault Locating Tool for Ten-Million-Scale Parallel ApplicationsabstractThis article presents Jdebug, a fast, non-intrusive and scalable fault locating tool for extreme-scale parallel applications. Large-scale debugging has drawn more attention with the increasing scale of supercomputers and applications. To eliminate program intrusion caused by traditional instrumentation or interception during debugging information acquisition, we introduce the out-of-band management into large-scale debugging. We propose a rapid information gathering scheme that separates user and debugging traffic to solve scalability problem and to eliminate program interference during merging data. Observations of Program Counters (PC) and performance characteristics in suspended applications find abnormalities and help locate abnormal threads caused by software errors or hardware failures effectively. Evaluation shows that Jdebug collects PCs of over 20 million cores on the new Sunway supercomputer within 1.97 seconds, and can locate the abnormal threads in 1.4 seconds with an accuracy of 92.5%. In the running test of three fundamental benchmarks (HPL, HPCG, Graph500) and seventeen real-world applications, Jdebug quickly and accurately locates abnormal threads to help find scalability errors and hardware failures including memory access failures, communication failures, and execution component failures, which validates its effectiveness. Dajia Peng, Yunlong Feng, Xin Liu 0081, Wei Xue 0003, Dexun Chen, Jiawei Song, Zuoning Chen |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | LMFF: efficient and scalable layered materials force field on heterogeneous many-core processorsabstractLAMMPS is one of the most popular Molecular Dynamic (MD) packages and is widely used in the field of physics, chemistry and materials simulation. Layered Materials Force Field (LMFF) is our expansion of the LAMMPS potential function based on the Tersoff potential and inter-layer potential (ILP) in LAMMPS. LMFF is designed to study layered materials such as graphene and boron hexanitride. It is universal and does not depend on any platform. We have also carried out a series of optimizations on LMFF and the optimization work is carried out on the new generation of Sunway supercomputer, called SWLMFF. Experiments show that our implementation is efficient, scalable and portable. When generic LMFF is ported to Intel Xeon Gold 6278C, 2X performance improvement is achieved. For the optimized SWLMFF, the overall performance improvement is nearly 200--330X compared to the original ILP and Tersoff potentials. And SWLMFF has good parallel efficiency of 95%-100% under weak scaling with 2.7 million atoms on a single process. The maximum atomic system simulated by SWLMFF is close to 231 atoms. And nanosecond simulations in one day can be realized. Ping Gao 0005, Xiaohui Duan, Jiaxu Guo, Zhenya Song, Li-Zhen Cui 0001, Xiangxu Meng, Xin Liu 0081, Wusheng Zhang, Ming Ma 0012, Dexun Chen, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002 |
SC | 12 |
| 2021 | SW_Qsim: a minimize-memory quantum simulator with high-performance on a new Sunway supercomputerabstractClassical simulation of quantum computation plays a critical role in numerical studies of quantum algorithms and the validation of quantum devices. Here, we introduce SW_Qsim, a tensor-network-based quantum simulator, which is designed with a two-level parallel structure for efficient implementation on the many-core New Sunway Supercomputer. We propose a minimize-memory contraction path algorithm for rectangular quantum grids to reduce the memory overhead, and provide the memory-limited simulation capacity of SW26010pro. Moreover, tensor operations are carefully optimized on the SW processor to achieve high performance. We design a fault tolerance mechanism to improve the extreme-scale parallel stability. We benchmark SW_Qsim's simulation of RQCs up to 400-qubits, achieving near-linear strong and weak scaling with up to 28.75 million cores, far beyond the previous state of the art. Our work sheds light on the development of efficient quantum algorithms for use in the physical, chemical, and engineering science fields. Xin Liu 0081, Pengpeng Zhao 0006, Yuling Yang, Honghui Shang, Weizhe Sun, Enming Dong, Dexun Chen |
SC | 10 |
| 2021 | Closing the "quantum supremacy" gap: achieving real-time simulation of a random quantum circuit using a new Sunway supercomputerabstractWe develop a high-performance tensor-based simulator for random quantum circuits(RQCs) on the new Sunway supercomputer. Our major innovations include: (1) a near-optimal slicing scheme, and a path-optimization strategy that considers both complexity and compute density; (2) a three-level parallelization scheme that scales to about 42 million cores; (3) a fused permutation and multiplication design that improves the compute efficiency for a wide range of tensor contraction scenarios; and (4) a mixed-precision scheme to further improve the performance. Our simulator effectively expands the scope of simulatable RQCs to include the 10X10(qubits)X(1+40+1)(depth) circuit, with a sustained performance of 1.2 Eflops (single-precision), or 4.4 Eflops (mixed-precision)as a new milestone for classical simulation of quantum circuits; and reduces the simulation sampling time of Google Sycamore to 304 seconds, from the previously claimed 10,000 years. Yong (Alexander) Liu, Xin (Lucy) Liu, Fang (Nancy) Li, Haohuan Fu, Yuling Yang, Jiawei Song, Pengpeng Zhao 0006, Dajia Peng, Huarong Chen, Chu Guo, Heliang Huang, Wenzhao Wu, Dexun Chen |
SC | 14 |
| 2021 | Accelerating all-electron ab initio simulation of raman spectra for biological systemsabstractRaman spectroscopy provides chemical and compositional information that can serve as a structural fingerprint for various materials. Therefore, simulations of Raman spectra, including both quantum perturbation analyses and ground-state calculations are of significant interest. However, highly accurate full quantum mechanical (QM) simulations of Raman spectra have previously been confined to small systems. For large systems such as biological materials, the computational cost of full QM simulations is extremely high, and their extension to such systems remains challenging. In the work described here, by employing robust new algorithms and advances in implementation for the many-core architectures, we are able to perform fast, accurate, and massively parallel full ab initio simulations of the Raman spectra of biological systems with excellent strong and weak scaling, thereby providing a starting point for applying QM approaches to structural studies of such systems. Honghui Shang, Yunquan Zhang, Ying Liu 0055, Mingchuan Wu, Yangjun Wu, Di Wei, Huimin Cui, Xin Liu 0081, Fei Wang 0096, Yuxi Ye, Yingxiang Gao, Shuang Ni, Xin Chen 0023, Dexun Chen |
SC | 16 |
| 2021 | Extreme-scale ab initio quantum raman spectra simulations on the leadership HPC system in ChinaabstractRaman spectroscopy provides chemical and compositional information that can serve as a structural fingerprint for various materials. Therefore, simulations of Raman spectra, including both quantum perturbation analyses and ground-state calculations, are of significant interest. However, highly accurate full quantum mechanical (QM) simulations of Raman spectra have previously been confined to small systems. For large systems such as biological materials, full QM simulations have an extremely high computational cost and remain challenging. In this work, robust new algorithms and advanced implementations on many-core architectures are employed to enable fast, accurate, and massively parallel full ab initio simulations of the Raman spectra of realistic biological systems containing up to 3006 atoms, with excellent strong and weak scaling. Up to a performance of 468.5 PFLOP/s in double-precision and 813.7 PLOPS/s in mixed-half precision is achieved on the new-generation Sunway high-performance computing system, suggesting the potential for new applications of the QM approach to biological systems. Honghui Shang, Yunquan Zhang, You Fu, Yingxiang Gao, Yangjun Wu, Xiaohui Duan, Rongfen Lin, Xin Liu 0081, Ying Liu 0055, Dexun Chen |
SC | 12 |
| 2021 | Symplectic structure-preserving particle-in-cell whole-volume simulation of tokamak plasmas to 111.3 trillion particles and 25.7 billion gridsabstractWe employ our recently developed explicit 2nd-order charge-conservative symplectic electromagnetic particle-in-cell (PIC) scheme in the cylindrical mesh to simulate the whole-volume magnetic confinement toroidal plasmas on the new Sunway supercomputer. From a large-scale simulation of magneticized toroidal plasma with 111.3 trillion particles and 25.7 billion grids, we have obtained a sustained performance exceeding 201.1 PFLOP/s (double precision) with the fastest iteration step achieving 298.2 PFLOP/s (double precision). For the first time, unprecedented high resolution evolution of 6D electromagnetic fully kinetic plasmas based on 2D equilibrium profiles from Experimental Advanced Superconducting Tokamak (EAST) and designed operation state of China Fusion Engineering Test Reactor (CFETR) are presented, and edge micro-instabilities can be investigated directly. This shows the possibility to study crucial problems and phenomena in the magnetic confinement toroidal plasma directly using the symplectic electromagnetic fully kinetic PIC method on world's leading supercomputers. Jianyuan Xiao, Junshi Chen 0003, Jiangshan Zheng, Hong An, Shenghong Huang, Chao Yang 0001, Ziyu Zhang 0003, Yeqi Huang, Wenting Han, Xin Liu 0081, Dexun Chen, Ge Zhuang, Qiang Chen 0005 |
SC | 12 |
| 2019 | OpenKMC: a KMC design for hundred-billion-atom simulation using millions of cores on Sunway TaihulightabstractWith more attention attached to nuclear energy, the formation mechanism of the solute clusters precipitation within complex alloys becomes intriguing research in the embrittlement of nuclear reactor pressure vessel (RPV) steels. Such phenomenon can be simulated with atomic kinetic Monte Carlo (AKMC) software, which evaluates the interactions of solute atoms with point defects in metal alloys. In this paper, we propose OpenKMC to accelerate large-scale KMC simulations on Sunway many-core architecture. To overcome the constraints caused by complex many-core architecture, we employ six levels of optimization in OpenKMC: (1) a new efficient potential computation model; (2) a group reaction strategy for fast event selection; (3) a software cache strategy; (4) combined communication optimizations; (5) a Transcription-Translation-Transmission algorithm for many-core optimization; (6) vectorization acceleration. Experiments illustrate that our OpenKMC has high accuracy and good scalability of applying hundred-billion-atom simulation over 5.2 million cores with a performance of over 80.1% parallel efficiency. Kun Li 0016, Honghui Shang, Yunquan Zhang, Shigang Li 0002, Baodong Wu, Dexun Chen, Zhiqiang Wei 0004 |
SC | 9 |
| 2018 | Redesigning LAMMPS for peta-scale and hundred-billion-atom simulation on Sunway TaihuLight
Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Wusheng Zhang, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Dexun Chen, Xiangxu Meng, Guangwen Yang 0002 |
SC | 10 |
| 2016 | Refactoring and optimizing the community atmosphere model (CAM) on the sunway taihulight supercomputerabstractThis paper reports our efforts on refactoring and optimizing the Community Atmosphere Model (CAM) on the Sunway TaihuLight supercomputer, which uses a many-core processor that consists of management processing elements (MPEs) and clusters of computing processing elements (CPEs). To map the large code base of CAM to the millions of cores on the Sunway system, we take OpenACC-based refactoring as the major approach, and apply source-to-source translator tools to exploit the most suitable parallelism for the CPE cluster, and to fit the intermediate variable into the limited on-chip fast buffer. For individual kernels, when comparing the original ported version using only MPEs and the refactored version using both the MPE and CPE clusters, we achieve up to 22× speedup for the compute-intensive kernels. For the 25km resolution CAM global model, we manage to scale to 24,000 MPEs, and 1,536,000 CPEs, and achieve a simulation speed of 2.81 model years per day. Haohuan Fu, Junfeng Liao, Wei Xue 0003, Lanning Wang, Dexun Chen, Long Gu, Jinxiu Xu 0001, Nan Ding 0006, Conghui He, Shizhen Xu, Yishuang Liang, Jiarui Fang, Yuanchao Xu 0001, Weijie Zheng 0001, Jingheng Xu, Zhen Zheng, Wanjing Wei, Bingwei Chen, Xiaomeng Huang, Guangwen Yang 0002 |
SC | 5 |
| 2016 | Extreme-scale phase field simulations of coarsening dynamics on the sunway taihulight supercomputerabstractMany important properties of materials such as strength, ductility, hardness and conductivity are determined by the microstructures of the material. During the formation of these microstructures, grain coarsening plays an important role. The Cahn-Hilliard equation has been applied extensively to simulate the coarsening kinetics of a two-phase microstructure. It is well accepted that the limited capabilities in conducting large scale, long time simulations constitute bottlenecks in predicting microstructure evolution based on the phase field approach. We present here a scalable time integration algorithm with large stepsizes and its efficient implementation on the Sunway TaihuLight supercomputer. The highly nonlinear and severely stiff Cahn-Hilliard equations with degenerate mobility for microstructure evolution are solved at extreme scale, demonstrating that the latest advent of high performance computing platform and the new advances in algorithm design are now offering us the possibility to simulate the coarsening dynamics accurately at unprecedented spatial and time scales. Jian Zhang 0070, Chunbao Zhou, Yangang Wang 0002, Lili Ju, Qiang Du 0001, Xuebin Chi, Dexun Chen |
SC | 8 |