EDBT 2026 Demo / reviewers in the wild / expert
Xin Liu 0081
dblp:76/1820-81
· DBLP profile ↗
26ranked-venue papers
1as first author
23since 2021 · last 2025
0000-0002-7870-6535ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 1 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Behaviour-diverse automatic penetration testing: a coverage-based deep reinforcement learning approach
Yizhou Yang, Longde Chen, Lanning Wang, Haohuan Fu, Xin Liu 0081, Zuoning Chen |
Frontiers Comput. Sci. | 6 |
| 2025 | Minimizing transformer inference overhead using controlling element on Shenwei AI acceleratorabstractTransformer models have become a cornerstone of various natural language processing (NLP) tasks. However, the substantial computational overhead during the inference remains a significant challenge, limiting their deployment in practical applications. In this study, we address this challenge by minimizing the inference overhead in transformer models using the controlling element on artificial intelligence (AI) accelerators. Our work is anchored by four key contributions. First, we conduct a comprehensive analysis of the overhead composition within the transformer inference process, identifying the primary bottlenecks. Second, we leverage the management processing element (MPE) of the Shenwei AI (SWAI) accelerator, implementing a three-tier scheduling framework that significantly reduces the number of host-device launches to approximately 1/10 000 of the original PyTorch-GPU setup. Third, we introduce a zero-copy memory management technique using segment-page fusion, which significantly reduces memory access latency and improves overall inference efficiency. Finally, we develop a fast model loading method that eliminates redundant computations during model verification and initialization, reducing the total loading time for large models from 22 128.31 ms to 1041.72 ms. Our contributions significantly enhance the optimization of transformer models, enabling more efficient and expedited inference processes on AI accelerators. Chunzhi Wu, Lufei Zhang, Yaguang Zhang, Wenyuan Shen, Hankang Fang, Xin Liu 0081 |
Frontiers Inf. Technol. Electron. Eng. | 10 |
| 2024 | Multilevel Load Balancing Algorithm for Domestic Heterogeneous Manycore ArchitectureabstractLoad imbalance often occurs in particle-in-cell simulations on parallel computing, which seriously affects the efficiency of applications. Due to the characteristics of multilevel parallelism and communication asymmetry of compute nodes in domestic heterogeneous manycore architecture, the impact of load imbalance is more prominent. The paper proposes a multilevel load-balancing algorithm for domestic heterogeneous manycore architecture. Inside the supernode, computing tasks are redivided based on manycore acceleration. Between the supernodes, a greedy-based communication mode is designed to minimize communication across supernodes. The experimental results show that the proposed algorithm achieves almost ideal dynamic load balance, and improves the performance of the evaporation module in two-phase flow simulation by 10.9-19.7 times for the 50 million-sized grid. Xin Chen 0023, Xin Liu 0081 |
ISPA | 5 |
| 2024 | SWattention: designing fast and memory-efficient attention for a new Sunway SupercomputerabstractAbstract In the past few years, Transformer-based large language models (LLM) have become the dominant technology in a series of applications. To scale up the sequence length of the Transformer, FlashAttention is proposed to compute exact attention with reduced memory requirements and faster execution. However, implementing the FlashAttention algorithm on the new generation Sunway Supercomputer faces many constraints such as the unique heterogeneous architecture and the limited memory bandwidth. This work proposes SWattention, a highly efficient method for computing the exact attention on the SW26010pro processor. To fully utilize the 6 core groups (CG) and 64 cores per CG on the processor, we design a two-level parallel task partition strategy. Asynchronous memory access is employed to ensure that memory access overlaps with computation. Additionally, a tiling strategy is introduced to determine optimal SRAM block sizes. Compared with the standard attention, SWattention achieves around 2.0x speedup for FP32 training and 2.5x speedup for mixed-precision training. The sequence lengths range from 1k to 8k and scale up to 16k without being out of memory. As for the end-to-end performance, SWattention achieves up to 1.26x speedup for training GPT-style models, which demonstrates that SWattention enables longer sequence length for LLM training. Ruohan Wu, Xianyu Zhu, Junshi Chen 0003, Tianyu Zheng, Xin Liu 0081, Hong An |
J. Supercomput. | 6 |
| 2023 | SetTron: Towards Better Generalisation in Penetration Testing with Reinforcement LearningabstractIntelligent penetration testing (pen-testing), utilising Deep Reinforcement Learning (DRL) has gained attention due to its potential for improving testing efficiency and cost-effectiveness in evaluating network system security. Nonetheless, current approaches which rely on simplistic neural network architectures suffer limitations in transferability and their ability to generalise to new tasks, thus impeding their practical application. This paper aims to address these issues by formalising the pen-testing decision process as a Host-Centric Markov decision process (HC- MDP), as well as establishing a structural representation of the relationships among the hosts within a network system. Further, we propose a flexible policy architecture, the “SetTron”, that leverages this structural representation to augment architectural inductive bias in a DRL agent and then practically evaluate our approach on pen-testing simulator platforms. The findings show SetTron to demonstrate superior performance, in terms of learning efficiency and policy convergence, compared to state-of-the-art methods and baselines with shorter penetration sequences and enhanced rewards. Besides, SetTron exhibits remarkable zero-shot generalisation capabilities, enabling perfect transfer to new tasks with randomly placed target hosts, achieving a 100 % success rate, and outperforming baselines by a factor of 6 when comparing normalised scores. Yizhou Yang, Mengxuan Chen, Haohuan Fu, Xin Liu 0081 |
GLOBECOM | 4 |
| 2023 | SW-LCM: A Scalable and Weakly-supervised Land Cover Mapping Method on a New Sunway SupercomputerabstractHigh-resolution land cover mapping (LCM) is an important application for studying and understanding the change of the earth surface. While deep learning (DL) methods demonstrate great potential in analyzing satellite images, they largely depend on massive high-quality labels. This paper proposes SW-LCM, a Scalable and Weakly-supervised two-stage Land Cover Mapping method on a new Sunway Supercomputer. Our method consists of a k-means clustering module as a first stage, and an iterative deep learning module as a second stage. With the k-means module providing a good enough starting point (taking inaccurate results as noisy labels), the deep learning module improves the classification results in an iterative way, without any labelling efforts required for processing large scenarios. To achieve efficiency for country-level land cover mapping, we design a customized data partition scheme and an on-the-fly assembly for k-means. Through careful parallelization and optimization, our k-means module scales to 98,304 computing nodes (over 38 million cores), and provides a sustained performance of 437.56 PFLOPS, in a real LCM task of the entire region of China; the iterative updating part scales to 24,576 nodes, with a performance of 11 PFLOPS. We produce a 10-m resolution land cover map of China, with an accuracy of 83.5% (10-class) or 73.2% (25-class), 7% to 8% higher than best existing products, paving ways for finer land surveys to support sustainability-related applications. Yi Zhao 0024, Juepeng Zheng, Haohuan Fu, Wenzhao Wu, Mengxuan Chen, Jinxiao Zhang, Lixian Zhang 0002, Runmin Dong, Zhenrong Du, Xin Liu 0081, Shaoqing Zhang, Le Yu 0001 |
IPDPS | 12 |
| 2023 | Lifetime-Based Optimization for Simulating Quantum Circuits on a New Sunway SupercomputerabstractHigh-performance classical simulator for quantum circuits, in particular the tensor network contraction algorithm, has become an important tool for the validation of noisy quantum computing. In order to address the memory limitations, the slicing technique is used to reduce the tensor dimensions, but it could also lead to additional computation overhead that greatly slows down the overall performance. This paper proposes novel lifetime-based methods to reduce the slicing overhead and improve the computing efficiency, including, an interpretation method to deal with slicing overhead, an inplace slicing strategy to find the smallest slicing set and an adaptive tensor network contraction path refiner customized for Sunway architecture. Experiments show that in most cases the slicing overhead with our inplace slicing strategy would be less than the Cotengra, which is the most used graph path optimization software at present. Finally, the resulting simulation time is reduced to 96.1s for the Sycamore quantum processor RQC, with a sustainable single-precision performance of 308.6Pflops using over 41M cores to generate 1M correlated samples, which is more than 5 times performance improvement compared to 60.4 Pflops in 2021 Gordon Bell Prize work. Yaojian Chen, Xinmin Shi, Jiawei Song, Xin Liu 0081, Lin Gan 0001, Chu Guo, Haohuan Fu, Dexun Chen, Guangwen Yang 0002 |
PPoPP | 5 |
| 2023 | Enabling Real World Scale Structural Superlubricity All-Atom Simulation on the Next-Generation Sunway SupercomputerabstractMolecular dynamics (MD) simulation can provide an affordable way for inspecting microscopic phenomena, which is a powerful complement to real-world experiments. But the spatial scale of MD simulations is usually magnitudes smaller than experiment systems. In this paper, we present our work, redesigning the widely used inter-layer potential in structural superlubricity. By carrying out a specialized neighbor list for inter-layer potential computation, the total memory access amount is reduced significantly. Besides, a simple but efficient vectorization strategy is implemented based on the new neighbor list. In the extreme case, our work can scale to 38 million cores to achieve a sustainable performance of 61 PFLOPS, enabling a simulation of a superlubricity system of 32 μm2 with 7.2 billion atoms at 4.75 ns/day, which is 11,834 times of reported largest scale simulation in superlubricity systems in contact area and almost ten times faster in time-to-solution. Furthermore, we have done a simulation at 9 μm2 which results in consistency with real-world experiments and verified some theoretical predictions in the mesoscopic scale. Xiaohui Duan, Ping Gao 0005, Ming Ma 0012, Lin Gan 0001, Xin Liu 0081, Haohuan Fu, Wei Xue 0003, Dexun Chen, Guangwen Yang 0002 |
SC | 6 |
| 2023 | Rapid simulations of atmospheric data assimilation of hourly-scale phenomena with modern neural networksabstractAtmospheric data assimilation is essential for numerical weather prediction. Ensemble data assimilation connects multiple instances of an atmospheric model through a Kalman filter-based algorithm, which is regarded as a challenging computing task today. In this work, we build a fast, low-cost, and scalable atmospheric data assimilation prototype, DIDA, for the new-generation Sunway supercomputer, including: (1) a framework that enables flexible deployment of components, and manages and optimizes data communication among modules, achieving maximum resource efficiency; (2) an accurate, robust, UNet-based surrogate model for atmospheric dynamic simulation to generate the background ensemble; (3) a batch-LETKF algorithm with high-performance eigenvalue decomposition, which is up to 7.37 times faster than existing numerical libraries while exhibiting almost linear scalability. Experimental evaluations show that our AI-integrated ensemble data assimilation prototype can complete hour-cycle assimilation in minutes, maintain linear scalability, and save an order of magnitude of computing resources, compared with the traditional method. Yiyuan Li, Xiting Ju, Qilong Jia, Yongxiao Zhou, Simeng Qian, Rongfen Lin, Bin Yang 0043, Shupeng Shi, Xin Liu 0081, Jian Tan 0005, Zhengding Hu, Limin Yan, Wei Xue 0003 |
SC | 10 |
| 2023 | 69.7-PFlops Extreme Scale Earthquake Simulation with Crossing Multi-faults and Topography on SunwayabstractA high-scalable and fully optimized earthquake model is presented based on the latest Sunway supercomputer. Contributions include: 1) the curvilinear grid finite-difference method (CGFDM) and flexible model applying perfectly matched layer (PML) and enabling more accurate and realistic terrain descriptions; 2) a hybrid and non-uniform domain decomposition scheme that efficiently maps the model across different levels of the computing system; and 3) sophisticated optimizations that largely alleviate or even eliminate bottlenecks in memory, communication, etc., obtaining a speedup of over 140×. Combining all innovations, the design fully exploits the hardware potential of all aspects and enables us to perform the largest CGFDM-based earthquake simulation ever reported (69.7 PFlops using over 39 million cores). Based on our design, the Turkey earthquakes (February 6, 2023), and the Ridgecrest earthquake (July 4, 2019), are successfully simulated with a maximum resolution of 12-m. Precise hazard evaluations for the hazardous reduction of earthquake-stricken areas are also conducted. Wubing Wan, Lin Gan 0001, Zekun Yin, Haodong Tian, Mengyuan Hua, Shengye Xiang, Zhongqiu He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002, Yaojian Chen, Xin Liu 0081, Wei Zhang 0321 |
SC | 22 |
| 2023 | Scalability and efficiency challenges for the exascale supercomputing system: practice of a parallel supporting environment on the Sunway exascale prototype systemabstractWith the continuous improvement of supercomputer performance and the integration of artificial intelligence with traditional scientific computing, the scale of applications is gradually increasing, from millions to tens of millions of computing cores, which raises great challenges to achieve high scalability and efficiency of parallel applications on super-large-scale systems. Taking the Sunway exascale prototype system as an example, in this paper we first analyze the challenges of high scalability and high efficiency for parallel applications in the exascale era. To overcome these challenges, the optimization technologies used in the parallel supporting environment software on the Sunway exascale prototype system are highlighted, including the parallel operating system, input/output (I/O) optimization technology, ultra-large-scale parallel debugging technology, 10-million-core parallel algorithm, and mixed-precision method. Parallel operating systems and I/O optimization technology mainly support large-scale system scaling, while the ultra-large-scale parallel debugging technology, 10-million-core parallel algorithm, and mixed-precision method mainly enhance the efficiency of large-scale applications. Finally, the contributions to various applications running on the Sunway exascale prototype system are introduced, verifying the effectiveness of the parallel supporting environment design. Xiaobin He, Xin Chen 0023, Xin Liu 0081, Dexun Chen, Yuling Yang, Yunlong Feng, Longde Chen, Xiaona Diao, Zuoning Chen |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2022 | SeqDLM: A Sequencer-Based Distributed Lock Manager for Efficient Shared File Access in a Parallel File SystemabstractDistributed locks are used to guarantee the distributed client-cache coherence in parallel file systems. However, they lead to poor performance in the case of parallel writes under high-contention workloads. We analyze the distributed lock manager and find out that lock conflict resolution is the root cause of the poor performance, which involves frequent lock revocations and slow data flushing from client caches to data servers. We design a distributed lock manager named SeqDLM by exploiting the sequencer mechanism. SeqDLM mitigates the lock conflict resolution overhead using early grant and early revocation while keeping the same semantics as traditional distributed locks. To evaluate SeqDLM, we have implemented a parallel file system called ccPFS using both SeqDLM and traditional distributed locks. Evaluations on 96 nodes show SeqDLM outperforms the traditional distributed locks by up to$\boldsymbol{10.3}\times$for high-contention parallel writes on a shared file with multiple stripes. Shaonan Ma, Kang Chen 0001, Teng Ma 0006, Xin Liu 0081, Dexun Chen, Yongwei Wu 0001, Zuoning Chen |
SC | 5 |
| 2022 | Increasing the Efficiency of Massively Parallel Sparse Matrix-Matrix Multiplication in First-Principles Calculation on the New-Generation Sunway SupercomputerabstractThe first-principles approach based on density-functional theory (DFT)/density-functional perturbation theory (DFPT) is widely used in calculations of the systems’ ground state energy, response properties (e.g., polarizability, phonon dispersions) and is playing an increasingly important role in chemistry, physics and materials science. For the large-scale calculations, the computation of the density matrix/response density matrix in DFT/DFPT has become the main performance bottleneck. One of the solutions is using the linear scaling method to get the density matrix and response density matrix. Here a massively parallel medium sparse matrix-matrix multiplication algorithm is designed for first-principle calculations and implemented on the new-generation Sunway supercomputer. Experiments show that the proposed method has obvious performance advantages compared to the original parallel version under moderate sparsity. The computing cores scale to 3,900,000 with strong scalability of 77.3$\%$. Xin Chen 0023, Yingxiang Gao, Honghui Shang, Zhiqian Xu 0005, Xin Liu 0081, Dexun Chen |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | Customer Adaptive Resource Provisioning for Long-Term Cloud Profit Maximization under Constrained BudgetabstractAs an efficient commercial information technology, cloud computing has attracted more and more users and enterprises to use it. Faced with such a large number and variety of customers, it is necessary for cloud providers (CPs) with limited budget to provide satisfactory customized pricing services, profitable customer and system investments, and flexible system resource provisioning strategies to improve both customer experience and long-term profit. Existing profit optimization research rarely considers customer diversity and dynamics, which may have a negative impact on long-term profit growth due to poor management of customer relations. In this article, we implement customer relationship management by considering both customer diversity and dynamics, and propose a customer adaptive resource provisioning scheme to maximize long-term profit under constrained budget. We consider four customer types (i.e., loyal, old, new, and lost) that can transition to each other during the customer's lifetime of interaction with the CP. The CP builds multiple cloud service sub-platforms, each of which contains multiple multiserver systems and serves the same type of customers. For the cloud service platform, we first analyze single multiserver system using an analytical method to obtain its optimal profit, invested funding, and system configuration. In particular, for systems serving new and lost customers, we develop a novel customer lifetime value (CLV)-based customer investment scheme that selects valuable customers for investment under limited marketing budget. Based on the above analysis, we then present a customer retention rate (CRR)-driven three-stage heuristic scheme that prioritizes investment in multiserver systems with endangered customers under limited infrastructure budget for reducing customer churn and promoting long-term profit growth. We conduct extensive simulation experiments to validate the effectiveness of our method. Simulation results show that compared with the benchmark algorithms, our method can improve the long-term profit and CRR by up to 3.4x and 7.8x, respectively. Peijin Cong, Junlong Zhou, Xin Liu 0081, Yao Liu 0017, Tongquan Wei |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Jdebug: A Fast, Non-Intrusive and Scalable Fault Locating Tool for Ten-Million-Scale Parallel ApplicationsabstractThis article presents Jdebug, a fast, non-intrusive and scalable fault locating tool for extreme-scale parallel applications. Large-scale debugging has drawn more attention with the increasing scale of supercomputers and applications. To eliminate program intrusion caused by traditional instrumentation or interception during debugging information acquisition, we introduce the out-of-band management into large-scale debugging. We propose a rapid information gathering scheme that separates user and debugging traffic to solve scalability problem and to eliminate program interference during merging data. Observations of Program Counters (PC) and performance characteristics in suspended applications find abnormalities and help locate abnormal threads caused by software errors or hardware failures effectively. Evaluation shows that Jdebug collects PCs of over 20 million cores on the new Sunway supercomputer within 1.97 seconds, and can locate the abnormal threads in 1.4 seconds with an accuracy of 92.5%. In the running test of three fundamental benchmarks (HPL, HPCG, Graph500) and seventeen real-world applications, Jdebug quickly and accurately locates abnormal threads to help find scalability errors and hardware failures including memory access failures, communication failures, and execution component failures, which validates its effectiveness. Dajia Peng, Yunlong Feng, Xin Liu 0081, Wei Xue 0003, Dexun Chen, Jiawei Song, Zuoning Chen |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | LMFF: efficient and scalable layered materials force field on heterogeneous many-core processorsabstractLAMMPS is one of the most popular Molecular Dynamic (MD) packages and is widely used in the field of physics, chemistry and materials simulation. Layered Materials Force Field (LMFF) is our expansion of the LAMMPS potential function based on the Tersoff potential and inter-layer potential (ILP) in LAMMPS. LMFF is designed to study layered materials such as graphene and boron hexanitride. It is universal and does not depend on any platform. We have also carried out a series of optimizations on LMFF and the optimization work is carried out on the new generation of Sunway supercomputer, called SWLMFF. Experiments show that our implementation is efficient, scalable and portable. When generic LMFF is ported to Intel Xeon Gold 6278C, 2X performance improvement is achieved. For the optimized SWLMFF, the overall performance improvement is nearly 200--330X compared to the original ILP and Tersoff potentials. And SWLMFF has good parallel efficiency of 95%-100% under weak scaling with 2.7 million atoms on a single process. The maximum atomic system simulated by SWLMFF is close to 231 atoms. And nanosecond simulations in one day can be realized. Ping Gao 0005, Xiaohui Duan, Jiaxu Guo, Zhenya Song, Li-Zhen Cui 0001, Xiangxu Meng, Xin Liu 0081, Wusheng Zhang, Ming Ma 0012, Dexun Chen, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002 |
SC | 8 |
| 2021 | SW_Qsim: a minimize-memory quantum simulator with high-performance on a new Sunway supercomputerabstractClassical simulation of quantum computation plays a critical role in numerical studies of quantum algorithms and the validation of quantum devices. Here, we introduce SW_Qsim, a tensor-network-based quantum simulator, which is designed with a two-level parallel structure for efficient implementation on the many-core New Sunway Supercomputer. We propose a minimize-memory contraction path algorithm for rectangular quantum grids to reduce the memory overhead, and provide the memory-limited simulation capacity of SW26010pro. Moreover, tensor operations are carefully optimized on the SW processor to achieve high performance. We design a fault tolerance mechanism to improve the extreme-scale parallel stability. We benchmark SW_Qsim's simulation of RQCs up to 400-qubits, achieving near-linear strong and weak scaling with up to 28.75 million cores, far beyond the previous state of the art. Our work sheds light on the development of efficient quantum algorithms for use in the physical, chemical, and engineering science fields. Xin Liu 0081, Pengpeng Zhao 0006, Yuling Yang, Honghui Shang, Weizhe Sun, Enming Dong, Dexun Chen |
SC | 2 |
| 2021 | Accelerating all-electron ab initio simulation of raman spectra for biological systemsabstractRaman spectroscopy provides chemical and compositional information that can serve as a structural fingerprint for various materials. Therefore, simulations of Raman spectra, including both quantum perturbation analyses and ground-state calculations are of significant interest. However, highly accurate full quantum mechanical (QM) simulations of Raman spectra have previously been confined to small systems. For large systems such as biological materials, the computational cost of full QM simulations is extremely high, and their extension to such systems remains challenging. In the work described here, by employing robust new algorithms and advances in implementation for the many-core architectures, we are able to perform fast, accurate, and massively parallel full ab initio simulations of the Raman spectra of biological systems with excellent strong and weak scaling, thereby providing a starting point for applying QM approaches to structural studies of such systems. Honghui Shang, Yunquan Zhang, Ying Liu 0055, Mingchuan Wu, Yangjun Wu, Di Wei, Huimin Cui, Xin Liu 0081, Fei Wang 0096, Yuxi Ye, Yingxiang Gao, Shuang Ni, Xin Chen 0023, Dexun Chen |
SC | 10 |
| 2021 | Extreme-scale ab initio quantum raman spectra simulations on the leadership HPC system in ChinaabstractRaman spectroscopy provides chemical and compositional information that can serve as a structural fingerprint for various materials. Therefore, simulations of Raman spectra, including both quantum perturbation analyses and ground-state calculations, are of significant interest. However, highly accurate full quantum mechanical (QM) simulations of Raman spectra have previously been confined to small systems. For large systems such as biological materials, full QM simulations have an extremely high computational cost and remain challenging. In this work, robust new algorithms and advanced implementations on many-core architectures are employed to enable fast, accurate, and massively parallel full ab initio simulations of the Raman spectra of realistic biological systems containing up to 3006 atoms, with excellent strong and weak scaling. Up to a performance of 468.5 PFLOP/s in double-precision and 813.7 PLOPS/s in mixed-half precision is achieved on the new-generation Sunway high-performance computing system, suggesting the potential for new applications of the QM approach to biological systems. Honghui Shang, Yunquan Zhang, You Fu, Yingxiang Gao, Yangjun Wu, Xiaohui Duan, Rongfen Lin, Xin Liu 0081, Ying Liu 0055, Dexun Chen |
SC | 10 |
| 2021 | Symplectic structure-preserving particle-in-cell whole-volume simulation of tokamak plasmas to 111.3 trillion particles and 25.7 billion gridsabstractWe employ our recently developed explicit 2nd-order charge-conservative symplectic electromagnetic particle-in-cell (PIC) scheme in the cylindrical mesh to simulate the whole-volume magnetic confinement toroidal plasmas on the new Sunway supercomputer. From a large-scale simulation of magneticized toroidal plasma with 111.3 trillion particles and 25.7 billion grids, we have obtained a sustained performance exceeding 201.1 PFLOP/s (double precision) with the fastest iteration step achieving 298.2 PFLOP/s (double precision). For the first time, unprecedented high resolution evolution of 6D electromagnetic fully kinetic plasmas based on 2D equilibrium profiles from Experimental Advanced Superconducting Tokamak (EAST) and designed operation state of China Fusion Engineering Test Reactor (CFETR) are presented, and edge micro-instabilities can be investigated directly. This shows the possibility to study crucial problems and phenomena in the magnetic confinement toroidal plasma directly using the symplectic electromagnetic fully kinetic PIC method on world's leading supercomputers. Jianyuan Xiao, Junshi Chen 0003, Jiangshan Zheng, Hong An, Shenghong Huang, Chao Yang 0001, Ziyu Zhang 0003, Yeqi Huang, Wenting Han, Xin Liu 0081, Dexun Chen, Ge Zhuang, Qiang Chen 0005 |
SC | 11 |
| 2021 | Establishing high performance AI ecosystem on Sunway platform
Xin Liu 0081, Zeqiang Huang, Tianyu Zheng |
CCF Trans. High Perform. Comput. | 3 |
| 2021 | Towards Efficient Short-Range Pair Interaction on Sunway Many-Core Architecture
Junshi Chen 0003, Hong An, Wenting Han, Zeng Lin, Xin Liu 0081 |
J. Comput. Sci. Technol. | 5 |
| 2021 | Parallelization and Optimization of NSGA-II on Sunway TaihuLight SystemabstractSunway TaihuLight system is the first supercomputer offering a peak performance over 100 PFlops, which can be utilized to parallelize Non-dominated Sorting Genetic Algorithm II (NSGA-II), a standard approach to multi-objective optimization. However, insufficient off-chip memory bandwidth and limited scratchpad memory capacity of the supercomputer hinder the performance improvement of parallellizing NSGA-II. In this article, we propose an optimized parallel NSGA-II on Sunway TaihuLight system, called swNSGA-II, by utilizing process- and thread-level parallelism of the system based on an improved island/master-slave model. To overcome the hurdles of low memory bandwidth and capacity, we propose a data sharing scheme based on register-level communication that can efficiently parallelize non-dominated sorting and crowding-distance computation of NSGA-II. Several optimization techniques including vectorization, direct memory accessing, and double buffering are also adopted to further accelerate swNSGA-II. Experiment results show that the proposed swNSGA-II can achieve a speedup of 41284 on a use case of path planning, and a speedup of 62692 on ZDT1 as compared to conventional NSGA-II. Xin Liu 0081, Su Wang 0005, Yao Liu 0017, Tongquan Wei |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Towards Efficient SpMV on Sunway Manycore ArchitecturesabstractSparse Matrix-Vector Multiplication (SpMV) is an essential computation kernel for many data-analytic workloads running in both supercomputers and data centers. The intrinsic irregularity in SpMV is challenging to achieve high performance, especially when porting to new architectures. In this paper, we present our work on designing and implementing efficient SpMV algorithms on Sunway, a novel architecture with many unique features. To fully exploit the Sunway architecture, we have designed a dual-side multi-level partition mechanism on both sparse matrices and hardware resources to improve locality and parallelism. On one hand, we partition sparse matrices into blocks, tiles, and slices for different granularities. On the other hand, we partition cores in a Sunway processor into fleets, and further dedicate part of cores in a fleet as computation and I/O cores. Moreover, we have optimized the communication between partitions to further improve the performance. Our scheme is generally applicable to different SpMV formats and implementations. For evaluation, we have applied our techniques atop a popular SpMV format, CSR. Experimental results on 18 datasets show that our optimization yields up to 15.5x (12.3x on average) speedups. Changxi Liu, Biwei Xie, Xin Liu 0081, Wei Xue 0003, Hailong Yang 0002, Xu Liu 0001 |
ICS | 3 |
| 2018 | ShenTu: processing multi-trillion edge graphs on millions of cores in seconds
Heng Lin, Xiaowei Zhu 0001, Bowen Yu 0003, Xiongchao Tang, Wei Xue 0003, Lufei Zhang, Torsten Hoefler, Xiaosong Ma, Xin Liu 0081, Jingfang Xu |
SC | 10 |
| 2012 | Microwave radiation anomaly of Yushu earthquake and its mechanismabstractA violent earthquake (Ms 7.1) occurred in Yushu county of China on April 14, 2010, and the epicenter was (33.1°N, 96.7°E). It caused 2689 people death. We use the AMER-E satellite remote sensing data to analyze the variation of microwave radiation around the epicenter. It was found 45 days before the earthquake a high microwave radiation strip appeared in the southwest of the epicenter, and gradually became longer and wider. The largest increment of microwave brightness temperature was up to 12K. 9 days before earthquake another high microwave radiation strip appeared in the northwest of the epicenter and the epicenter was just in the intersection of the two the high microwave radiation strips. 4 days before the earthquake the strength of the microwave anomaly became weak, but one day before the earthquake the strength became increase again. After the main shock the anomaly still existed for about 20days and was finally disappeared. To explore the cause of microwave radiation anomaly the microwave polarization difference index (MPDI) and the air temperature variation was investigated. The final result indicated the microwave anomaly was caused by the temperature increase, i.e. the anomaly should belong to thermal anomaly. Shanjun Liu, Xin Liu 0081, Lixin Wu |
IGARSS | 2 |