Guangwen Yang 0002

dblp:67/3001-2 · DBLP profile ↗
← Back
141ranked-venue papers
4as first author
39since 2021 · last 2026
0000-0002-8673-8254ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 115 · 1 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 2 since 2021Computer networks · 4Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 HierCut: Enabling 16-bit Format Mixed Precision for Molecular Dynamics through Hierarchical Cutoff
abstract
Mixed-precision methods offer the potential to achieve better performance while maintaining accuracy comparable to that of high-precision formats. However, the adoption of mixed precision—particularly with 16-bit formats—in scientific computing remains limited due to precision truncation.
Lin Gan 0001, Xiaohui Duan, Zhengrui Li, Jiayu Fu, Guangzhao Li, Guangwen Yang 0002
PPoPP8
2026 Accelerating Molecular Dynamics Simulations on ARM Multi-Core Processors
abstract
LAMMPS is a widely used molecular dynamics (MD) software package in materials science, computational chemistry, and biophysics, supporting parallel computing from a single CPU core to large supercomputers. The Kunpeng processor features both high memory bandwidth and core density and is therefore an interesting candidate for accelerating compute-intensive workloads. In this paper, we target the Kunpeng multi-core architecture and focus on optimizing LAMMPS for modern ARM-based platforms by using the Lennard-Jones (L-J) and Tersoff potentials as representative case studies. We investigate both common and specific optimization challenges, and present a comprehensive performance analysis addressing four key aspects: neighbor list algorithm design, force computation optimization, efficient vectorization, and multi-thread parallelization. Experimental results show that the optimized potentials achieve speedups of approximately$2 \times$and$5 \times$, reaching$4.55 \times$and$7.04\times$the performance of the original Intel version for L-J and Tersoff, respectively. Both potentials outperform Intel's acceleration library, with a peak performance up to$2.9\times$-$3.5\times$. In terms of parallel efficiency, we evaluate scalability both within a single CPU (small-scale) and across multiple nodes (large-scale). Strong and weak scaling tests within a single CPU show that when the expansion factor is 32 times, parallel efficiency remains above$90\%$. Large-scale weak scaling across multiple nodes achieves up to$86\%$efficiency when the expansion factor is 32. Using 32 nodes (18,432 processes), our implementation enables billion-atom simulations with L-J and Tersoff potentials. This work achieves breakthrough performance and provides critical support for large-scale molecular dynamics in engineering applications.
Huihai An, Zhihua Sa, Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Yizhen Chen, Lin Gan 0001, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.10
2026 Exploiting the Performance Potential of Extreme-Scale Earthquake Simulation: Achieving 86.7 PFLOPS With Over 39 Million Cores
abstract
Leveraging the latest Sunway supercomputer, we developed a fully optimized earthquake simulation model that accurately captures topographic effects for realistic seismic analysis. Optimizing for the SW26010Pro architecture with DMA/RMA communication mechanisms, data compression schemes, and vectorization, we achieved a speedup exceeding 160×. Our pipeline-based computation and communication overlapping scheme, combined with performance prediction models further minimized computational costs. These optimizations enabled the largest-scale curvilinear grid finite-difference method (CGFDM) earthquake simulations to date, covering 197 trillion grid points and achieving 86.7 PFLOPS on 39 million cores with a weak scaling efficiency of 97.9%. These advancements enabled the successful simulation of the 2008 Wenchuan earthquake, providing high-resolution seismic insights and robust assessments for regional hazard mitigation and disaster preparedness.
Lin Gan 0008, Wubing Wan, Zekun Yin, Zhong He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.13
2026 SMEStencil: Optimizing High-Order Stencils on ARM Multicore Using SME Unit
abstract
Matrix-accelerated stencil computation is a hot research topic, yet its application to 3 dimensional (3D) high-order stencils and HPC remains underexplored. With the emergence of Scalable Matrix Extension(SME) on ARMv9-A CPU, we analyze SME-based accelerating strategies and tailor an optimal approach for 3D high-order stencils. We introduce algorithmic optimizations based on Scalable Vector Extension(SVE) and SME unit to address strided memory accesses, alignment conflicts, and redundant accesses. We propose memory optimizations to boost on-package memory efficiency, and a novel multi-thread parallelism paradigm to overcome data-sharing challenges caused by the absence of shared data caches. SMEStencil sustains consistently high hardware utilization across diverse stencil shapes and dimensions. Our DMA-based inter-NUMA communication further mitigates NUMA effects and MPI limitations in hybrid parallelism. Combining all the innovations, SMEStencil outperforms state-of-the-art libraries on Nividia A100 GPGPU by up to 2.1× . Moreover, the performance improvements enabled by our optimizations translate directly to real-world HPC applications and enable Reverse Time Migration(RTM) real-world applications to yield 1.8x speedup versus highly-optimized Nvidia A100 GPGPU version.
Tianqi Mao 0003, Lin Gan 0008, Wubing Wan, Jiayu Fu, Lanke He, Zekun Yin, Wei Xue 0003, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.11
2025 Auto-Stencil: Performance-Driven Stencil Optimization with Hardware Feedback for LLMs
abstract
Stencil computation is an important computing pattern from numerous scientific simulations, and optimizing stencil for modern GPU architectures demands specialized expertise in both parallel programming and hardware-specific optimizations. However, traditional domain-specific languages based auto-tuning tools offer limited flexibility; generalized auto-parallelization tools offer limited performance; and large language models produce code that often fails to compile or underperforms. This paper presents Auto-Stencil, a novel framework that bridges this gap by integrating LLMs with hardware-aware reinforcement learning. Our approach combines a comprehensive stencil optimization dataset with a dual-objective training methodology that systematically aligns model outputs with both functional correctness and performance requirements. By incorporating execution feedback through a performance-driven reward model, Auto-Stencil generates highly optimized CUDA implementations that not only pass unit tests but also deliver exceptional performance across diverse stencil patterns. Experimental results demonstrate the superiority of the framework over state-of-the-art alternatives: 100% compilation accuracy, 100% optimization rate, and an average of 171 × speedup across test cases compared to a single CPU core. The results show Auto-Stencil is particularly suitable for large-scale HPC workflows where automation of complex, architecture-specific optimizations can significantly reduce development effort while maintaining excellent performance.
Quan Deng 0001, Lin Gan 0001, Hongkun Yu 0002, Wenlai Zhao, Guangwen Yang 0002
ICPP5
2025 Trillion Ligands per Day: Performance-Portable Virtual Screening via Compound Database Optimization and Multi-Target Docking
abstract
Structure-based virtual screening confronts a grand challenge in scaling to trillion-ligand libraries for drug discovery. We present SWDOCKP2, a performance-portable virtual screening framework achieving 1.9 trillion ligand-receptor pairs daily across eight targets on the Sunway OceanLight supercomputer with 39-million cores — 10× faster than prior state-of-the-art. Key innovations combine (1) a ligand database optimizer with conformational sorting and merging, (2) multi-receptor grid alignment enabling parallel target screening and SIMD-accelerated trilinear interpolation, and (3) a Sunway architecture emulator for cross-platform efficiency. These advancements bridge computational scalability with novel drug discovery demands, offering a blueprint for next-generation supercomputing in structure-based drug design. Additionally, SWDOCKP2 will generate an unprecedented dataset of predicted protein-ligand interactions, creating a transformative resource for machine learning applications. By addressing experimental data scarcity, this dataset empowers accurate ligand prediction, generative chemistry, and AI-driven drug discovery.
Xiaohui Duan, Gaowei Chen, Yizhen Chen, Qixin Chang, Qiancheng Xia, Zekun Yin, Lin Gan 0001, Yibing Shan, Guangwen Yang 0002, Niu Huang
SC12
2025 T2-RELION: Task Parallelism, Tensor Core Accelerated RELION for Cryo-EM 3D Reconstruction
abstract
Cryo-electron microscopy (cryo-EM) is a key technique for structural biology, but its computational efficiency, particularly during 3D reconstruction, remains a bottleneck. We introduce T2-RELION, a highly optimized version of RELION for cryo-EM 3D reconstruction on CPU-GPU platforms. RELION is a widely used open-source package in the cryo-EM community. We identify and resolve key inefficiencies in RELION’s parallelization strategy and memory management by proposing task parallelism and a three-phase GPU memory management strategy. Furthermore, we leverage Tensor Cores to accelerate the hot-spot kernel for difference calculation, employing an advanced pipelining strategy to hide latency and enable thread-block-level data reuse. On a quad-A100 GPU machine, performance evaluations demonstrate that T2-RELION outperforms RELION 4.0. For the hot-spot kernel, our optimizations achieve 1.90-23.7 times speedup. For the whole application using CNG and Trpv1 datasets, we observe 3.86 times and 2.68 times speedups, respectively.
Jiayu Fu, Jingle Xu, Lin Gan 0001, Tianqi Mao 0003, Zirong Shen, Xiaohui Duan, Wei Xue 0003, Guangwen Yang 0002
SC10
2025 Leveraging the Hardware Resources to Accelerate cryo-EM Reconstruction of RELION on the New Sunway Supercomputer
abstract
The fast development of biomolecular structure determination has enabled the fine-grained study of objects in the micro-world, such as proteins and RNAs. The world is benefited. However, as the computational algorithms are constantly developed, the enrichment of features increases the algorithmic complexity and brings more computationally unfriendly modules. It calls for efficient solutions to leverage the rich and various hardware resources from the world’s most state-of-the-art supercomputing systems, and to fully accelerate the performance of the applications. In this article, we present our efforts on porting and optimizing the 3D reconstruction of RELION, one of the most popular cryo-EM software for biomolecular structure determinations, by leveraging different resources of the latest generation of Sunway heterogeneous supercomputer. Several novel approaches are proposed to resolve different challenges faced by the complex algorithm, including a multi-level parallel scheme and operator optimizations to smartly map and scale RELION, efficient strategies to largely address the memory bottlenecks and improve data locality, lock-free writing solutions to minimize write-write conflicts, and pipelining approaches to obtain excellent computation and communication overlap. Combining all proposed optimizations, the computation time is greatly reduced to under 2 hours, achieving 11.9× and 8.9× speedups on two different datasets. The overall design scales to 131,072 cores, increasing parallel efficiency from 33% to 61% and from 46% to 70%, respectively. To the best of our knowledge, this is the first work that fully optimized and scaled the 3D reconstruction of RELION using the latest Sunway system.
Jingle Xu, Jiayu Fu, Lin Gan 0001, Yaojian Chen, Zhaoqi Sun, Zhenchun Huang, Guangwen Yang 0002
ACM Trans. Archit. Code Optim.7
2025 Accelerating Half-Precision Seismic Simulation on Neural Processing Unit
abstract
Due to the superiority of handling irregular regions of interest, the curvilinear grid finite difference method (CGFDM) has become wildly used in seismic simulation for earthquake hazard evaluation and understanding of earthquake physics. This paper proposes a novel approach that optimizes a CGFDM solver on the Ascend, a cutting-edge Neural Processing(NPU) Unit using half-precision storage and mixed-precision arithmetic. The approach increases the data throughput and computing efficiency, enabling more effective seismic modeling. Furthermore, we propose an efficient matrix unit enabled 3D difference algorithm that employs matrix unit on NPU to accelerate the computation. By fully exploiting the capability of matrix unit and wide SIMD lane, our solver on Ascend achieves a speedup of 4.19 × over the performance of parallel solver on two AMD CPUs and has successfully simulated real-world Wenchuan earthquake. For the best of our knowledge, we are the first to conduct seismic simulations on NPU.
Wubing Wan, Lin Gan 0008, Ping Gao 0005, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.11
2024 Enabling High-Performance Physical Based Rendering on New Sunway Supercomputer
abstract
Physical based rendering is widely applied in diverse fields requiring realistic scene visualization. This paper outlines our efforts in implementing a high-performance and highly scalable physical based rendering framework on the next-generation Sunway supercomputer based on PBRT. To effectively tailor the rendering application to the hardware attributes of the state-of-the-art architecture, we primarily carried out four approaches, 1) a memory management strategy, 2) a solution for runtime polymorphism, 3) measures to mitigate instruction cache misses, and 4) a two level load balancing strategy. Our design can achieve at most 41.95x speedup relative to baseline implementation. By using 32 Sunway processors, we can achieve at most 27.97x speedup relative to a 40-core 5218R CPU and at most 6.90x speedup relative to a RTX 3090 GPU. Nearly linear scalabilities are obtained when scaling up to 2,048 Sunway processors.
Lin Gan 0001, Shengye Xiang, Xiaohui Duan, Guangwen Yang 0002
IPDPS6
2024 ESFLOW: Mapping Large-Scale Earthquake Simulation to Spatial Computing Systems
abstract
In the last ten years, the frequent earthquakes have pushed the experts to watch the earth’s movements more closely. Fortunately, recent enhancements in modern High-Performance Computing (HPC) power help researchers understand the internal earthquake mechanisms using the numerical simulation method. Considering the performance and energy requirements, specialized FPGA-based accelerators have become a promising solution for high-performance earthquake simulation. In this work, we propose a resource-aware decomposition framework of earthquake simulation based on an analytic resource model. Then, we demonstrate our efforts in computation design to ensure continuous streaming operation and prevent deadlocks. Compared with the floating-point implementation based on NVIDIA GPU A6000, our design is 1.6 times and 3.4 times better in performance and energy efficiency, respectively.
Qiang Liu 0011, Lin Gan 0001, Guangwen Yang 0002
ISCAS4
2024 Collaborative Metapath Enhanced Corporate Default Risk Assessment on Heterogeneous Graph
abstract
Default risk assessment for small companies is a tough problem in financial services. Recent efforts utilize advanced Heterogeneous Graph Neural Networks (HGNNs) with metapaths to exploit interactive features in corporate activities for risk analysis. However, few works are proposed for commercial banks. Given a real financial graph, how to detect corporate default risks? We identify two challenges for the task. (1) Massive noisy connections hinder HGNNs to achieve strong results. (2) Multiple semantic connections greatly increase transitive default risk, while existing aggregation schemes do not leverage such connection patterns. In this work, we propose a novel Heterogeneous Graph Co-Attention Network for corporate default risk assessment. Our model takes advantage of collaborative metapaths to distill risky features by a co-attentive aggregation mechanism. First, the local attention score models the importance of neighbors under each metapath by holistic metapath context. Second, the global attention score fuse local attention scores to filter valuable/noisy signals. Then, pairwise importance learning aims to enhance attention scores of multi-metapath neighbors for risky feature distillation. Extensive experiments on large-scale banking datasets demonstrate the effectiveness of our method.
Yingsheng Ji, Yushu Chen, Xi Zhang 0008, Guangwen Yang 0002
WWW6
2024 O2ath: an OpenMP offloading toolkit for the sunway heterogeneous manycore platform
Lifeng Yan, Qixin Chang, Haitian Lu, Chenlin Li, Quanjie He, Xiaohui Duan, Zekun Yin, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Guangwen Yang 0002
CCF Trans. High Perform. Comput.15
2024 Towards optimized tensor code generation for deep learning on sunway many-core processor
Mingzhen Li 0001, Changxi Liu, Jianjin Liao, Xuegui Zheng, Hailong Yang 0002, Rujun Sun, Lin Gan 0001, Guangwen Yang 0002, Zhongzhi Luan, Depei Qian 0001
Frontiers Comput. Sci.9
2024 A Joint Time-Frequency Domain Transformer for multivariate time series forecasting
Yushu Chen, Shengzhuo Liu, Jinzhe Yang, Wenlai Zhao, Guangwen Yang 0002
Neural Networks6
2024 Acceleration of Multi-Body Molecular Dynamics With Customized Parallel Dataflow
abstract
FPGAs are drawing increasing attention in resolving molecular dynamics (MD) problems, and have already been applied in problems such as two-body potentials, force fields composed of these potentials, etc. Competitive performance is obtained compared with traditional counterparts such as CPUs and GPUs. However, as far as we know, FPGA solutions for more complex and real-world MD problems, such as multi-body potentials, are seldom to be seen. This work explores the prospects of state-of-the-art FPGAs in accelerating multi-body potential. An FPGA-based accelerator with customized parallel dataflow that features multi-body potential computation, motion update, and internode communication is designed. Major contributions include: (1) parallelization applied at different levels of the accelerator; (2) an optimized dataflow mixing atom-level pipeline and cell-level pipeline to achieve high throughput; (3) a mixed-precision method using different precision at different stages of simulations; and (4) a communication-efficient method for internode communication. Experiments show that, our single-node accelerator is over 2.7× faster than an 8-core CPU design, performing 20.501 ns/day on a 55,296-atom system for theTersoffsimulation. Regarding power efficiency, our accelerator is 28.9× higher than I7-11700 and 4.8× higher than RTX 3090 when running the same test case.
Quan Deng 0001, Qiang Liu 0011, Xiaohui Duan, Lin Gan 0008, Jinzhe Yang, Wenlai Zhao, Zhenxiang Zhang, Guiming Wu, Wayne Luk, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.12
2024 SunwayLB: Enabling Extreme-Scale Lattice Boltzmann Method Based Computing Fluid Dynamics Simulations on Advanced Heterogeneous Supercomputers
abstract
The Lattice Boltzmann Method (LBM) is a class of Computational Fluid Dynamics methods which models the fluid as fictive particles. In this paper, we report our work on SunwayLB, which enables LBM based solutions aiming for industrial applications using advanced heterogeneous systems such as the Sunway supercomputers. We propose several techniques to boost the simulation speed and improve the scalability of SunwayLB, including a customized multi-level domain decomposition and data sharing scheme, a carefully orchestrated strategy to fuse kernels with different performance constraints for a more balanced workload, and optimization strategies for assembly code. Based on these optimization schemes, we manage to scale SunwayLB on three advanced supercomputers: Sunway TaihuLight, the new Sunway Supercomputer and a GPU cluster. On Sunway TaihuLight, our largest simulation involves up to 5.6 trillion lattice cells, achieving 11,245 billion cell updates per second (GLUPS), 77% memory bandwidth utilization and a sustained performance of 4.7 PFlops. We further improve the memory bandwidth utilization and computational efficiency using the unique features of a new generation of Sunway supercomputer. On the new Sunway Supercomputer, the largest simulation contains over 4.2 trillion lattice cells, resulting in 6,583 GLUPS, 81% memory bandwidth utilization and a sustained performance of 2.76 PFlops. To evaluate the portability of our code, we also adapt our code to a GPU cluster with tailored optimization techniques, resulting in 191x speedup and 83.8% memory bandwidth utilization. We demonstrate a series of computational experiments for extreme-large scale fluid flow, as examples of real-world applications, to check the validity and performance of our work. The results show that our implementation is competent to be a highly scalable and efficient solution for large-scale CFD problems on heterogeneous systems.
Xuesen Chu, Xiaojing Lv, Hongsong Meng, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.8
2024 Optimizing I/O Performance Through Effective vCPU Scheduling Interference Management
abstract
Virtual machines (VMs) heavily rely on virtual CPUs (vCPUs) scheduling to achieve efficient I/O performance. The vCPU scheduling interference can cause inconsistent scheduling latency and degraded I/O performance, potentially compromising the services provided by affected VMs. Existing solutions have limitations, such as inefficiency in diagnosing interference issues or imposing undesired side effects on cloud systems. To address these challenges, we present Otter, a holistic technique for optimizing I/O performance in the presence of vCPU scheduling interference. Otter employs innovative methods to enhance interference diagnosis efficiency. First, we propose lightweight methods to measure the dynamic changes in scheduling latencies for co-running vCPUs, ensuring both flexibility and accuracy. Second, we propose fine-grained quantification methods to timely determine the interference, with low false positive and false negative rates. Third, we identify interference patterns that aid in analyzing the root causes of interference and preventing similar issues from recurring. Otter has been operational for one year in the production cloud at the National Supercomputing Center (Wuxi). It diagnoses and helps fix more than 470 vCPU scheduling interference-related issues, resulting in a 19.6% improvement in cloud service I/O performance with negligible overhead in production.
Jinzhe Yang, Jidong Zhai, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.4
2023 Automatic Deep Learning Operator Fusion on Sunway SW26010 Many-Core Processor
abstract
Deep learning networks (DNNs) have been growing rapidly in recent years, with increasing demands on computing power. Therefore, accelerating the execution of DNN models has become a research hotspot. Operator fusion is a critical optimization strategy to enhance DNN performance in Deep Learning (DL) frameworks, such as TensorFlow, Pytorch, TVM and Halide. However, these frameworks are designed for general optimization and cannot fully harness the specific features of emerging hardware. Moreover, they primarily implement operator fusion at the operator level, missing out on many fusion opportunities and heavily relying on extensive manual optimizations for fused operators. Targeting the Sunway SW26010 Many-Core processor, the basic building block of Sunway TaihuLight supercomputer, we introduce swAutoFuser, an end-to-end automatic operator fusion and code generation framework. swAutoFuser proposes a set of low-level primitives to leverage hardware features and employs an autofuser to achieve primitive level fusion, which breaks operator boundaries and enables more fusion opportunities. In addition, swAutoFuser can automatically generate high-performance fused operator implementations based on a static cost model, significantly reducing the overhead of manually optimizing fused operators. Our experiments demonstrate that swAutoFuser can improve operator performance by 10% to 56%.
Wenxiang Zhang, Wenzhao Wu, Yanjie Zhen, Wenlai Zhao, Guangwen Yang 0002
ICPADS6
2023 Accelerating Large-Scale CFD Simulations with Lattice Boltzmann Method on a 40-Million-Core Sunway Supercomputer
abstract
The Lattice Boltzmann Method (LBM) has gained widespread popularity due to its applicability in fluid dynamics, chemical engineering, material science, and other domains. In this work, we present an optimized implementation of the LBM, with a specific focus on achieving superior performance and scalability on advanced heterogeneous systems such as the new Sunway supercomputer. To accomplish this, we employ several techniques, including kernel fusion to enhance temporal and spatial locality, a customized multi-level domain decomposition and data sharing scheme, and pipelining strategies that are tailored to the SW26010-Pro processor. As a result of these optimizations, we have successfully scaled our code to a total of 39,000,000 CPU cores. Our largest simulation, which encompassed over 42 trillion lattice cells, achieved an impressive 67,018 billion lattice cell updates per second (GLUPS), with 82.9% memory bandwidth utilization, and a sustained performance of 28 PFlops. In order to assess the portability of our implementation, we also adapted our code to run on a GPU cluster, utilizing a range of tailored optimization techniques. Our results demonstrated a 191x speedup, along with 83.8% memory bandwidth utilization. Our proposed approach marks a significant milestone in the field of LBM implementations, as it demonstrates unprecedented scalability by effectively utilizing over 39,000,000 cores while maintaining exceptional parallel efficiency and computational performance. This achievement establishes our method as a compelling solution for addressing large-scale computational fluid dynamics challenges on heterogeneous systems.
Xuesen Chu, Xiaojing Lv, Haohuan Fu, Guangwen Yang 0002
ICPP6
2023 Lifetime-Based Optimization for Simulating Quantum Circuits on a New Sunway Supercomputer
abstract
High-performance classical simulator for quantum circuits, in particular the tensor network contraction algorithm, has become an important tool for the validation of noisy quantum computing. In order to address the memory limitations, the slicing technique is used to reduce the tensor dimensions, but it could also lead to additional computation overhead that greatly slows down the overall performance. This paper proposes novel lifetime-based methods to reduce the slicing overhead and improve the computing efficiency, including, an interpretation method to deal with slicing overhead, an inplace slicing strategy to find the smallest slicing set and an adaptive tensor network contraction path refiner customized for Sunway architecture. Experiments show that in most cases the slicing overhead with our inplace slicing strategy would be less than the Cotengra, which is the most used graph path optimization software at present. Finally, the resulting simulation time is reduced to 96.1s for the Sycamore quantum processor RQC, with a sustainable single-precision performance of 308.6Pflops using over 41M cores to generate 1M correlated samples, which is more than 5 times performance improvement compared to 60.4 Pflops in 2021 Gordon Bell Prize work.
Yaojian Chen, Xinmin Shi, Jiawei Song, Xin Liu 0081, Lin Gan 0001, Chu Guo, Haohuan Fu, Dexun Chen, Guangwen Yang 0002
PPoPP11
2023 Enabling Real World Scale Structural Superlubricity All-Atom Simulation on the Next-Generation Sunway Supercomputer
abstract
Molecular dynamics (MD) simulation can provide an affordable way for inspecting microscopic phenomena, which is a powerful complement to real-world experiments. But the spatial scale of MD simulations is usually magnitudes smaller than experiment systems. In this paper, we present our work, redesigning the widely used inter-layer potential in structural superlubricity. By carrying out a specialized neighbor list for inter-layer potential computation, the total memory access amount is reduced significantly. Besides, a simple but efficient vectorization strategy is implemented based on the new neighbor list. In the extreme case, our work can scale to 38 million cores to achieve a sustainable performance of 61 PFLOPS, enabling a simulation of a superlubricity system of 32 μm2 with 7.2 billion atoms at 4.75 ns/day, which is 11,834 times of reported largest scale simulation in superlubricity systems in contact area and almost ten times faster in time-to-solution. Furthermore, we have done a simulation at 9 μm2 which results in consistency with real-world experiments and verified some theoretical predictions in the mesoscopic scale.
Xiaohui Duan, Ping Gao 0005, Ming Ma 0012, Lin Gan 0001, Xin Liu 0081, Haohuan Fu, Wei Xue 0003, Dexun Chen, Guangwen Yang 0002
SC10
2023 Toward Exascale Computation for Turbomachinery Flows
abstract
A state-of-the-art large eddy simulation code has been developed to solve compressible flows in turbomachinery. The code has been engineered with a high degree of scalability, enabling it to effectively leverage the many-core architecture of the new Sunway system. A consistent performance of 115.8 DP-PFLOPs has been achieved on a high-pressure turbine cascade consisting of over 1.69 billion mesh elements and 865 billion Degree of Freedoms (DOFs). By leveraging a high-order unstructured solver and its portability to large heterogeneous parallel systems, we have progressed towards solving the grand challenge problem outlined by NASA [1], which involves a time-dependent simulation of a complete engine, incorporating all the aerodynamic and heat transfer components.
Yuhang Fu, Weiqi Shen, Jiahuan Cui, Yao Zheng 0003, Guangwen Yang 0002, Jifa Zhang, Tingwei Ji, Fangfang Xie, Xiaojing Lv, Guocheng Tao, Paul Tucker, Steven A. E. Miller, Shirui Luo, Seid Koric
SC5
2023 69.7-PFlops Extreme Scale Earthquake Simulation with Crossing Multi-faults and Topography on Sunway
abstract
A high-scalable and fully optimized earthquake model is presented based on the latest Sunway supercomputer. Contributions include: 1) the curvilinear grid finite-difference method (CGFDM) and flexible model applying perfectly matched layer (PML) and enabling more accurate and realistic terrain descriptions; 2) a hybrid and non-uniform domain decomposition scheme that efficiently maps the model across different levels of the computing system; and 3) sophisticated optimizations that largely alleviate or even eliminate bottlenecks in memory, communication, etc., obtaining a speedup of over 140×. Combining all innovations, the design fully exploits the hardware potential of all aspects and enables us to perform the largest CGFDM-based earthquake simulation ever reported (69.7 PFlops using over 39 million cores). Based on our design, the Turkey earthquakes (February 6, 2023), and the Ridgecrest earthquake (July 4, 2019), are successfully simulated with a maximum resolution of 12-m. Precise hazard evaluations for the hazardous reduction of earthquake-stricken areas are also conducted.
Wubing Wan, Lin Gan 0001, Zekun Yin, Haodong Tian, Mengyuan Hua, Shengye Xiang, Zhongqiu He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002, Yaojian Chen, Xin Liu 0081, Wei Zhang 0321
SC18
2023 Bio-ESMD: A Data Centric Implementation for Large-Scale Biological System Simulation on Sunway TaihuLight Supercomputer
abstract
Molecular dynamics (MD) simulations of biological systems are playing an increasingly important role in the research of pathogens and drugs. Most MD methods for biological simulations rely on the listed bonds which interact among specific groups of atoms identified by atom tags (unique atom tags regardless the storage location). However, efficient mapping of tags to atom locations is often challenging on modern many-core processors because data locality can not always be guaranteed for large-scale systems. In this paper, we present Bio-ESMD, a new MD implementation supporting listed bonds. Bio-ESMD is designed and developed based on our previously designed ESMD framework for many-core processors. In Bio-ESMD, we have introduced a data-centric approach for refactoring MD algorithms by reorganizing the cell list data structure to adopt bond lists with guaranteed data locality. Our implementation achieves speedups of over two compared to SW_GROMACS on Sunway TaihuLight. Furthermore, Bio-ESMD can simulate a system of 308.8 million atoms at 1.33 ns/day or 14.44 million atoms at 17.28 ns/day with linear weak scaling efficiency.
Xiaohui Duan, Junben Weng, Bertil Schmidt, Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.10
2023 Redesign and Accelerate the AIREBO Bond-Order Potential on the New Sunway Supercomputer
abstract
Molecular dynamics (MD) is one of the most crucial computer simulation methods for understanding real-world processes at the atomic level. Reactive potentials based on the bond order concept have the ability to model dynamic bond breaking and formation with close to quantum mechanical (QM) precision without actually requiring expensive QM calculations. In this article, we focus on the adaptive intermolecular reactive empirical bond-order (AIREBO) potential in LAMMPS for the simulation of carbon and hydrocarbon systems on the new Sunway supercomputer. To achieve scalable performance, we propose a parallel two-level building scheme and periodic buffering strategy for the tailored data design to explore data locality and data reuse. Furthermore, we design two optimized nearest-neighbor access algorithms: the redistribution of accumulated coefficients algorithm and the double-end search connectivity algorithm. Finally, we implement parallel force computation with an AoS data layout and hardware/software co-cache. In addition, we have designed a low-overhead atomic operation-based load balancing method and vectorization. The overall performance of AIREBO achieves a speedup of nearly$20\times$on a single core group (CG), and more than$5\times$and$4\times$over an Intel Xeon E5 2680 v3 core and an Intel Xeon Gold 6138 core, respectively. Compared with the Intel accelerator package in LAMMPS, our performance further achieves$3.0\times$of an Intel Xeon E5 2680 v3 core and is better than that of an Intel Xeon Gold 6138 core. We complete the validation of the results in no more than 20.5 hours on a single node with 2,000,000 running steps (i.e., 1 ns). Our experiments show that the simulation of 2,139,095,040 atoms on 798,720 ((1MPE+64CPEs) × 12,288 processes) cores exhibits a parallel efficiency of 88% under weak scaling.
Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Wubing Wan, Jiaxu Guo, Wusheng Zhang, Lin Gan 0008, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.11
2022 Radio: Reconciling Disk I/O Interference in a Para-virtualized Cloud
abstract
As more virtual machines (VMs) are consolidated in the cloud system, interference among VMs sharing underlying resources may occur more frequently than ever. In particular, certain VMs’ disk I/O performance gets impacted, leading to related cloud services being seriously compromised. Existing interference analysis approaches cannot guarantee desired results due to 1) lack of effective techniques for characterizing disk I/O interference and 2) considerable runtime overhead for determining interference and related culprits. To overcome these barriers, we present Radio, an end-to-end analysis tool for disk I/O interference diagnostics in a para-virtualized cloud. Radio quantifies the dynamic changes in I/O strength across virtual CPUs (vCPUs), constructs the performance repository to efficiently identify VMs’ abnormal behaviors, and then exploits interference heat maps and non-constant correlation approaches to infer the culprits of interference. With Radio's deployment at the National Supercomputing Center in Wuxi for more than 10 months, we demonstrate its effectiveness in real-world use cases on the cloud system with more than 300 VMs deployed. Radio can effectively analyze the interference issues within 20 seconds, incurring only 0.2% extra CPU overhead on the host machine. With this achievement, Radio has successfully assisted system administrators in reducing the daily incidence of interference from more than 65% to less than 10% and improving the overall disk throughput of the cloud system by more than 27.5%.
Guangwen Yang 0002, Liana Wang, Wei Xue 0003
CLOUD1
2022 Detecting Cash-out Users via Dense Subgraphs
abstract
Cash-out fraud refers to the withdrawal of cash from a credit card by illegitimate payments with merchants. Conventional data-driven approaches for cash-out detection commonly construct a classifier with domain specific feature engineering. To further spot cash-out behaviors in complex scenarios, recent efforts adopt graph models to exploit the interaction relations rich in financial transactions. However, most existing graph-based methods are proposed for online payment activities in internet financial institutions. Moreover, these methods commonly rely on a large amount of online user data, which are not well suitable for the traditional credit card services in commercial banks. In this paper, we focus on discerning fraudulent cash-out users by taking advantage of only the personal credit card data from banks. To alleviate the scarcity of available labeled data, we formulate the cash-out detection problem as identifying dense blocks. First, we define a bipartite multigraph to hold transactions between users and merchants, where cash-out activities generate cyclically intensive and high-volume flows. Second, we give a formal definition of cash-out behaviors from four perspectives: time, capital, cyclicity, and topotaxy. Then, we develop ANTICO, with a class of metrics to capture suspicious signals of the activities and a greedy algorithm to spot suspicious blocks by optimizing the proposed metric. Theoretical analysis shows a provable upper bound of ANTICO on the effectiveness of detecting cash-out users. Experimental results show that ANTICO outperforms state-of-the-art methods in accurately detecting cash-out users on both synthetic and real-world banking data.
Yingsheng Ji, Xinlei Tang, Xi Zhang 0008, Guangwen Yang 0002
KDD6
2022 A fully-customized dataflow engine for 3D earthquake simulation with a complex topography
Bingwei Chen, Haohuan Fu, Wayne Luk, Guangwen Yang 0002
Sci. China Inf. Sci.4
2022 Enabling Large-Scale Simulation of CAM on the Sunway TaihuLight Supercomputer
abstract
The Community Atmosphere Model (CAM) has been ported, redesigned, and scaled to the full system of the Sunway TaihuLight, and provides peta-scale climate modeling performance. Based on a novel domain decomposition method, we have fully optimized the complete model code by using both OpenACC refactoring and more aggressive and finer-grained Athread approaches. The Athread approach enables us to achieve exceptional memory control and usage, efficient vectorization, and sophisticated utilization of the thread-level communication mechanism. We have also further refined the load-balance behaviors towards ultra-large-scale numerical simulation. By combining all these novelties, we achieved a simulation speed of 7.2 and 25.6 simulation-year-per-day (SYPD) for global 25-km and 100-km resolution, respectively (1.2- to 2.2-fold improvements over previous efforts), and a sustainable double-precision performance of 3.3 PFlops for a 750-m global simulation when using 10075000 cores.
Xiaohui Duan, Lin Gan 0001, Wubing Wan, Yuhu Chen, Jinzhe Yang, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Computers11
2022 Input-Aware Sparse Tensor Storage Format Selection for Optimizing MTTKRP
abstract
The major bottleneck of Canonical polyadic decomposition (CPD) is matricized tensor times Khatri-Rao product (MTTKRP). To optimize the performance of MTTKRP, various sparse tensor formats have been proposed such as CSF and HiCOO. However, due to the spatial complexity of the tensors, no single format fits all tensors. To address this problem, we propose SpTFS, a framework that automatically predicts the optimal storage format for an input sparse tensor. Specifically, SpTFS leverages a set of sampling methods to lower the sparse tensor to fix-sized matrices and sparsity features. In addition, SpTFS adopts both supervised learning based and unsupervised learning based methods to predict the optimal sparse tensor storage formats. For supervised learning, we propose TnsNet that combines convolution neural network (CNN) and the feature layer, which effectively captures the sparsity patterns of the input tensors. Whereas for unsupervised learning, we propose TnsClustering that consists of a feature encoder using convolutional layers and fully connected layers, and a K-means++ model to cluster sparse tensors for optimal tensor format prediction, without massively profiling on the hardware platform. The experimental results show that both TnsNet and TnsClustering can achieve higher prediction accuracy and performance speedup compared to the state-of-the-art works.
Qingxiao Sun, Yi Liu 0013, Hailong Yang 0002, Ming Dun, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001
IEEE Trans. Computers7
2022 Optimization of Reactive Force Field Simulation: Refactor, Parallelization, and Vectorization for Interactions
abstract
Molecular dynamics (MD) simulations are playing an increasingly important role in many areas ranging from chemical materials to biological molecules. With the continuing development of MD models, the potentials are getting larger and more complex. In this article, we focus on the reactive force field (ReaxFF) potential from LAMMPS to optimize the computation of interactions. We present our efforts on refactoring for neighbor list building, bond order computation, as well as valence angles and torsion angles computation. After redesigning these kernels, we develop a vectorized implementation for non-bonded interactions, which is nearly 100 × faster than the management processing element (MPE) on the Sunway TaihuLight supercomputer. Furthermore, we have implemented the three-body-list free torsion angles computation, and propose a line-locked software cache method to eliminate write conflicts in the torsion angle and valence angle interactions resulting in an order-of-magnitude speedup on a single Sunway TaihuLight node. In addition, we achieve a speedup of up to 3.5 compared to the KOKKOS package on an Intel Xeon Gold 6148 core. When executed on 1,024 processes, our implementation enables the simulation of 21,233,664 atoms on 66,560 cores with a performance of 0.032 ns/day and a weak scaling efficiency of 95.71 percent.
Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Wusheng Zhang, Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.9
2022 Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight
abstract
Boson sampling is expected to be an important milestone that will demonstrate quantum computational advantage (or quantum supremacy). This work establishes the benchmarking of Gaussian boson sampling (GBS) with threshold detection based on the Sunway TaihuLight supercomputer. To achieve the best performance and provide a competitive scenario for future quantum computing studies, the selected simulation algorithm is fully optimized based on a set of innovative approaches, including a parallel framework with almost perfect load balance and an instruction-level optimizing scheme based on a shortest-path-based instruction scheduling. In addition, data precision is carefully processed by an integer-instruction-based and multiple-precision fixed-point implementation, including 128- and 256-bit precison mode, which can be appropriately selected based on an adaptive precision optimizing scheme. Based on these methods, a highly efficient parallel quantum sampling algorithm is designed. The largest run enables us to obtain one Torontonian function of a$100\times 100$submatrix from 50-photon GBS within 20 hours in 128-bit precision and 2 days in 256-bit precision. To our knowledge, this was the largest quantum computing simulation based on Boson Sampling by using modern supercomputers.
Lin Gan 0001, Mingcheng Chen, Yaojian Chen, Haitian Lu, Chao-Yang Lu, Jian-Wei Pan, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.9
2022 Redesigning and Optimizing UCSF DOCK3.7 on Sunway TaihuLight
abstract
Molecular docking is the process of posing, scoring, and ranking small molecules at the binding sites of proteins to prioritize compounds for experimental testing. It is a widely-used computational method in the drug discovery process. However, it is a highly time-consuming procedure since a receptor may need to find favorable ligand orientations in billions of ligands. UCSF DOCK3.7 is one of the most widely used molecular docking applications. In this paper, we port and optimize UCSF DOCK3.7 on the Sunway TaihuLight supercomputer. To avoid the impact of load imbalance, we employ a producer-consumer strategy that can overlap I/O and computation in order to achieve high performance. Furthermore, we present a new binary file format to replace the mol2db2 file format for ligand storage and adopt xzip rather than gzip to compress ligand files. We show that our file format can reduce I/O time significantly while xzip saves significant storage. For the routines which determine the orientation of a ligand relative to the receptor, we present an improved algorithm to discard geometrically similar orientations. Furthermore, we fuse loops and compress memory usage to store data in fast Local Device Memory (LDM) in order to score ligand orientations with high efficiency. In addition, we propose a number of architecture-specific optimizations. Asynchronous data transfer and vectorization of computation are implemented to take full advantage of the SW26010 processor. Our experiments show that a speedup of 167 can be achieved by using the proposed strategies. Compared to a core of an Intel(R) Core(TM) i9-10900K CPU, our approach achieves speedups of 15 on a SW26010 core group. Furthermore, our implementation achieves strong scalability to hundreds of thousands of heterogeneous cores on the next-generation Sunway supercomputer.
Jinxiao Zhang, Xiaohui Duan, Xiaobo Wan, Niu Huang, Bertil Schmidt, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.8
2021 CUBIST: High-Quality 360-Degree Video Streaming Services via Tile-based Edge Caching and FoV-Adaptive Prefetching
abstract
360-degree video streaming, which is becoming more and more popular as the fast development of VR/AR applications nowadays due to the immersive viewing experience it can offer, poses enormous challenges to the current network infrastructure in terms of high bandwidth and low latency requirements. To address this problem and to ensure the QoE (quality of experience) of end-users, this paper presents CUBIST, a method and system for high-quality 360-degree video streaming in networks with cache nodes at the edge. To the best of our knowledge, it is the first tile-based edge caching solution that incorporates proactive tile prefetching and hierarchical cache organization into reactive caching to maximize the caching benefit while reducing the cost of 360-degree video streaming. Experimental results show that CUBIST can achieve a cache hit ratio of 87 % and improve the effective video bitrate by 12.9 % with most rate transitions being small when compared with the latest FoV-aware edge caching scheme.
Dongbiao He, Jinlei Jiang, Teng Ma 0006, Guangwen Yang 0002, Cédric Westphal, J. J. Garcia-Luna-Aceves, Shutao Xia
ICWS4
2021 LMFF: efficient and scalable layered materials force field on heterogeneous many-core processors
abstract
LAMMPS is one of the most popular Molecular Dynamic (MD) packages and is widely used in the field of physics, chemistry and materials simulation. Layered Materials Force Field (LMFF) is our expansion of the LAMMPS potential function based on the Tersoff potential and inter-layer potential (ILP) in LAMMPS. LMFF is designed to study layered materials such as graphene and boron hexanitride. It is universal and does not depend on any platform. We have also carried out a series of optimizations on LMFF and the optimization work is carried out on the new generation of Sunway supercomputer, called SWLMFF. Experiments show that our implementation is efficient, scalable and portable. When generic LMFF is ported to Intel Xeon Gold 6278C, 2X performance improvement is achieved. For the optimized SWLMFF, the overall performance improvement is nearly 200--330X compared to the original ILP and Tersoff potentials. And SWLMFF has good parallel efficiency of 95%-100% under weak scaling with 2.7 million atoms on a single process. The maximum atomic system simulated by SWLMFF is close to 231 atoms. And nanosecond simulations in one day can be realized.
Ping Gao 0005, Xiaohui Duan, Jiaxu Guo, Zhenya Song, Li-Zhen Cui 0001, Xiangxu Meng, Xin Liu 0081, Wusheng Zhang, Ming Ma 0012, Dexun Chen, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
SC16
2021 Towards efficient canonical polyadic decomposition on sunway many-core processor
Ming Dun, Yunchun Li, Qingxiao Sun, Hailong Yang 0002, Wei Li 0125, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001
Inf. Sci.8
2021 Towards efficient tile low-rank GEMM computation on sunway many-core processors
Qingchang Han, Hailong Yang 0002, Ming Dun, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001
J. Supercomput.6
2021 The Deep Learning Compiler: A Comprehensive Survey
abstract
The difficulty of deploying various deep learning (DL) models on diverse DL hardware has boosted the research and development of DL compilers in the community. Several DL compilers have been proposed from both industry and academia such as Tensorflow XLA and TVM. Similarly, the DL compilers take the DL models described in different DL frameworks as input, and then generate optimized codes for diverse DL hardware as output. However, none of the existing survey has analyzed the unique design architecture of the DL compilers comprehensively. In this paper, we perform a comprehensive survey of existing DL compilers by dissecting the commonly adopted design in details, with emphasis on the DL oriented multi-level IRs, and frontend/backend optimizations. Specifically, we provide a comprehensive comparison among existing DL compilers from various aspects. In addition, we present detailed analysis on the design of multi-level IRs and illustrate the commonly adopted optimization techniques. Finally, several insights are highlighted as the potential research directions of DL compiler. This is the first survey paper focusing on the design architecture of DL compilers, which we hope can pave the road for future research towards DL compiler.
Mingzhen Li 0001, Yi Liu 0013, Qingxiao Sun, Xin You 0001, Hailong Yang 0002, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001
IEEE Trans. Parallel Distributed Syst.9
2020 Efficient Edge Caching for High-Quality 360-Degree Video Delivery
Dongbiao He, Jinlei Jiang, Cédric Westphal, Guangwen Yang 0002
MMM (2)4
2020 Neighbor-list-free molecular dynamics on sunway TaihuLight supercomputer
abstract
Molecular dynamics (MD) simulations are playing an increasingly important role in many research areas. Pair-wise potentials are widely used in MD simulations of bio-molecules, polymers, and nano-scale materials. Due to a low compute-to-memory-access ratio, their calculation is often bounded by memory transfer speeds. Sunway TaihuLight is one of the fastest supercomputers featuring a custom SW26010 many-core processor. Since the SW26010 has some critical limitations regarding main memory bandwidth and scratchpad memory size, it is considered as a good platform to investigate the optimization of pair-wise potentials especially in terms of data reusage. MD algorithms often use a neighbor-list data structure to reduce the computational workload. In this paper, we show that a cell-list-based approach is more suitable for the SW26010 processor. We apply a number of novel optimization methods including self-adaptable replica-summation for conflict-free parallelization, parameter profiles for flexible vectorization, and particle-cell cutoff checking filters for reducing the computational workload. We also established an open source standalone framework featuring the techniques above, ESMD1, which is at least 50% faster than the latest existing LAMMPS port on a single TaihuLight node. Furthermore, EMSD achieves a weak scaling efficiency of 88% on 4,096 nodes.
Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Hongsong Meng, Bertil Schmidt, Haohuan Fu, Lin Gan 0001, Wei Xue 0003, Guangwen Yang 0002
PPoPP11
2020 Cell-list based molecular dynamics on many-core processors: a case study on sunway TaihuLight supercomputer
abstract
Molecular dynamics (MD) simulations are playing an increasingly important role in several research areas. The most frequently used potentials in MD simulations are pair-wise potentials. Due to the memory wall, computing pair-wise potentials on many-core processors are usually memory bounded. In this paper, we take the SW26010 processor as an exemplary platform to explore the possibility to break the memory bottleneck by improving data reusage via cell-list-based methods. We use cell-lists instead of neighbor-lists in the potential computation, and apply a number of novel optimization methods. Theses methods include: an adaptive replica arrangement strategy, a parameter profile data structure, and a particle-cell cutoff checking filter. An incremental cell-list building method is also realized to accelerate the construction of cell-lists. Furthermore, we have established an open source standalone framework, ESMD, featuring the techniques above. Experiments show that ESMD is 50~170% faster than previous ports on a single node, and can scale to 1,024 nodes with a weak scalibility of 95%.
Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Hongsong Meng, Bertil Schmidt, Haohuan Fu, Lin Gan 0001, Wei Xue 0003, Guangwen Yang 0002
SC12
2020 SpTFS: sparse tensor format selection for MTTKRP via deep learning
abstract
Canonical polyadic decomposition (CPD) is one of the most common tensor computations adopted in many scientific applications. The major bottleneck of CPD is matricized tensor times Khatri-Rao product (MTTKRP). To optimize the performance of MTTKRP, various sparse tensor formats have been proposed such as CSF and HiCOO. However, due to the spatial complexity of the tensors, no single format fits all tensors. To address this problem, we propose SpTFS, a framework that automatically predicts the optimal storage format for an input sparse tensor. Specifically, SpTFS leverages a set of sampling methods to lower the sparse tensor to fix-sized matrices and specific features. Then, TnsNet combines CNN and the feature layer to accurately predict the optimal format. The experimental results show that SpTFS achieves prediction accuracy of 92.7% and 96% on CPU and GPU respectively.
Qingxiao Sun, Yi Liu 0013, Ming Dun, Hailong Yang 0002, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001
SC7
2020 Tuning a general purpose software cache library for TaihuLight's SW26010 processor
Xiaohui Duan, Haohuan Fu, Lin Gan 0001, Wei Xue 0003, Guangwen Yang 0002
CCF Trans. High Perform. Comput.7
2020 High performance reconfigurable computing for numerical simulation and deep learning
Lin Gan 0001, Jinzhe Yang, Wenlai Zhao, Wayne Luk, Guangwen Yang 0002
CCF Trans. High Perform. Comput.6
2020 Efficient AES implementation on Sunway TaihuLight supercomputer: A systematic approach
Liandeng Li, Jiarui Fang, Jinlei Jiang, Lin Gan 0001, Weijie Zheng 0001, Haohuan Fu, Guangwen Yang 0002
J. Parallel Distributed Comput.7
2020 Millimeter-Scale and Billion-Atom Reactive Force Field Simulation on Sunway Taihulight
abstract
Large-scale molecular dynamics (MD) simulations on supercomputers play an increasingly important role in many research areas. With the capability of simulating charge equilibration (QEq), bonds and so on, Reactive force field (ReaxFF) enables the precise simulation of chemical reactions. Compared to the first principle molecular dynamics (FPMD), ReaxFF has far lower requirements on computational resources so that it can achieve higher efficiencies for large-scale simulations. In this article, we present our efforts on scaling ReaxFF on the Sunway TaihuLight Supercomputer (TaihuLight). We have carefully redesigned the force analysis and neighbor list building steps. By applying fine-grained optimizations we gain better single process performance. For the many-body interactions, we propose an isolated computation and update strategy and implement inverse trigonometric functions. For QEq, we implement a pipelined conjugate gradient (CG) approach to achieving better scalability. Furthermore, we reorganize the data layout and implement the update operation based on data locality in ReaxFF. Our experiments show that this approach can simulate chemical reactions with 1,358,954,496 atoms using 4,259,840 cores with a performance of 0.015 ns/day. To our best knowledge, this is the first realization of chemical reaction simulation with a millimeter-scale force field.
Ping Gao 0005, Xiaohui Duan, Tingjian Zhang, Bertil Schmidt, Wusheng Zhang, Lin Gan 0001, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.13
2020 Massively Scaling Seismic Processing on Sunway TaihuLight Supercomputer
abstract
Common Midpoint (CMP) and Common Reflection Surface (CRS) are widely used methods for improving the signal-to-noise ratio in the field of seismic processing. These methods are computationally intensive and require high-performance computing. This article optimizes these methods on the Sunway many-core architecture and implements large-scale seismic processing on the Sunway Taihulight supercomputer. We propose the following three optimization techniques: 1) we propose a software cache method to reduce the overhead of memory accesses, and share data among CPEs via the register communication; 2) we re-design the semblance calculation procedure to further reduce the overhead of memory accesses; 3) we propose a vectorization method to improve the performance when processing the small volume of data within short loops. The experimental results show that our implementations of CMP and CRS methods on Sunway achieve 3.50× and 3.01× speedup on average compared to the-state-of-the-art implementations on CPU. In addition, our implementation is capable to run on more than one million cores of Sunway TaihuLight with good scalability.
Yongmin Hu, Hailong Yang 0002, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001
IEEE Trans. Parallel Distributed Syst.5
2020 Accelerating Sparse Cholesky Factorization on Sunway Manycore Architecture
abstract
To improve the performance of sparse Cholesky factorization, existing research divides the adjacent columns of the sparse matrix with the same nonzero patterns into supernodes for parallelization. However, due to the various structures of sparse matrices, the computation of the generated supernodes varies significantly, and thus hard to optimize when computed by dense matrix kernels. Therefore, how to efficiently map sparse Choleksy factorization to the emerging architectures, such as Sunway many-core processor, remains an active research direction. In this article, we propose swCholesky, which is a highly optimized implementation of sparse Cholesky factorization on Sunway processor. Specifically, we design three kernel task queues and a dense matrix library to dynamically adapt to the kernel characteristics and architecture features. In addition, we propose an auto-tuning mechanism to search for the optimal settings of the important parameters in swCholesky. Our experiments show that swCholesky achieves better performance than state-of-the-art implementations.
Mingzhen Li 0001, Yi Liu 0013, Hailong Yang 0002, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001
IEEE Trans. Parallel Distributed Syst.6
2020 Quantum Supremacy Circuit Simulation on Sunway TaihuLight
abstract
With the rapid progress made by industry and academia, quantum computers with dozens of qubits or even larger size are being realized. However, the fidelity of existing quantum computers often sharply decreases as the circuit depth increases. Thus, an ideal quantum circuit simulator on classical computers, especially on high-performance computers, is needed for benchmarking and validation. We design a large-scale simulator of universal random quantum circuits, often called “quantum supremacy circuits”, and implement it on Sunway TaihuLight. The simulator can be used to accomplish the following two tasks: 1) Computing a complete output state-vector; 2) Calculating one or a few amplitudes. We target the simulation of 49-qubit circuits. For task 1), we successfully simulate such a circuit of depth 39, and for task 2) we reach the 55-depth level. To the best of our knowledge, both of the simulation results reach the largest depth for 49-qubit quantum supremacy circuits.
Riling Li, Bujiao Wu, Mingsheng Ying, Xiaoming Sun 0001, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.5
2020 Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core Supercomputer
abstract
This article presents an automatic k-means clustering solution targeting the Sunway TaihuLight supercomputer. We first introduce a multilevel parallel partition approach that not only partitions by dataflow and centroid, but also by dimension, which unlocks the potential of the hierarchical parallelism in the heterogeneous many-core processor and the system architecture of the supercomputer. The parallel design is able to process large-scale clustering problems with up to 196,608 dimensions and over 160,000 targeting centroids, while maintaining high performance and high scalability. Furthermore, we propose an automatic hyper-parameter determination process for k-means clustering, by automatically generating and executing the clustering tasks with a set of candidate hyper-parameter, and then determining the optimal hyper-parameter using a proposed evaluation method. The proposed autoclustering solution can not only achieve high performance and scalability for problems with massive high-dimensional data, but also support clustering without sufficient prior knowledge for the number of targeted clusters, which can potentially increase the scope of k-means algorithm to new application areas.
Wenlai Zhao, Pan Liu 0002, Vladimir Janjic, Xiaohan Yan, Shicai Wang, Haohuan Fu, Guangwen Yang 0002, John Thomson
IEEE Trans. Parallel Distributed Syst.8
2019 Large-scale Parallel Design for Cryo-EM Structure Determination on Heterogeneous Many-core Architectures
abstract
Cryo-EM structure determination is the most important research area in structural biology. With the development of cryo-electron microscopy, the resolution has been enhanced significantly, which leads to the huge computation to reconstruct the biomolecule in recent years. In this paper, we present a large-scale parallel design for Cryo-EM structure determination on heterogeneous many-core architectures. A novel task parallel strategy is proposed to reduce the redundant computation and improve the scalability on large-scale systems. Further, We distribute the reconstruction model to each node and rearrange the data layout to achieve high parallel efficiency and reduce the memory footprint. The proposed comprehensive parallel design shows highly parallel efficiency and scalability on large-scale heterogeneous architectures, which could significantly accelerate the whole period of Cryo-EM structure determination process.
Hongkun Yu 0002, Ruixin Sun, Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002
BIBM7
2019 Million-Core-Scalable Simulation of the Elastic Migration Algorithm on Sunway TaihuLight Supercomputer
abstract
Migration algorithm is one of the most essential methods in seismic application to image the underground geology, and to help scientists and researchers in geophysics exploration better understand the earth system. However, due to the desire in migration algorithm for covering lager region and acquiring better resolution, many tough challenges have to be tackled for current state-of-the-art computing systems. This work optimized and scaled the elastic migration algorithm onto the Sunway TaihuLight supercomputer, one of the most powerful systems of the world. Targeting at the major process, the reverse time migration (RTM) algorithm, a set of algorithmic, process-level, and thread-level optimizations is proposed, to significantly improve the performance (up to 163× speedup in time-to-solution) on Sunway CPU. Our design is successfully scaled to over two million cores (2,662,400 cores in total) on the Sunway TaihuLight supercomputer, with nearly ideal weak-scaling efficiency. The largest run is able to achieve a sustainable performance of processing over 859 billion cells per second.
Lin Gan 0001, Jingheng Xu, Xin Wang 0233, Sihai Wu, Xiaohui Duan, Haohuan Fu, Guangwen Yang 0002
CCGRID8
2019 swATOP: Automatically Optimizing Deep Learning Operators on SW26010 Many-Core Processor
abstract
Achieving an optimized mapping of Deep Learning (DL) operators to new hardware architectures is the key to building a scalable DL system. However, handcrafted optimization involves huge engineering efforts, due to the variety of DL operator implementations and complex programming skills. Targeting the innovative many-core processor SW26010 adopted by the 3rd fastest supercomputer Sunway TaihuLight, an end-to-end automated framework called swATOP is presented as a more practical solution for DL operator optimization. Arithmetic intensive DL operators are expressed into an auto-tuning-friendly form, which is based on tensorized primitives. By describing the algorithm of a DL operator using our domain specific language (DSL), swATOP is able to derive and produce an optimal implementation by separating hardware-dependent optimization and hardware-agnostic optimization. Hardware-dependent optimization is encapsulated in a set of tensorized primitives with sufficient utilization of the underlying hardware features. The hardware-agnostic optimization contains a scheduler, an intermediate representation (IR) optimizer, an auto-tuner, and a code generator. These modules cooperate to perform an automatic design space exploration, to apply a set of programming techniques, to discover a near-optimal solution, and to generate the executable code. Our experiments show that swATOP is able to bring significant performance improvement on DL operators in over 88% of cases, compared with the best-handcrafted optimization. Compared to a black-box autotuner, the tuning and code generation time can be reduced to minutes from days using swATOP.
Jiarui Fang, Wenlai Zhao, Jinzhe Yang, Long Wang 0014, Lin Gan 0001, Haohuan Fu, Guangwen Yang 0002
ICPP8
2019 Parallelizing cryo-EM 3D reconstruction on GPU cluster with a partitioned and streamed model
abstract
As a vital approach to determine the structure of biomacromolecules, high-resolution cryo-electron microscopy (cryo-EM) 3D reconstruction is extremely compute-intensive, and has gradually migrated to GPU accelerators in recent years. With certain kernels already achieving high speedup and efficiency on GPUs, the reconstruction part, which inherently requires accesses of a large 3D model in different orientations, brings tough challenges to GPU architectures and has no effective GPU-based options. To fill the above gap, in this paper, we propose Stream3D, a novel GPU-based parallel design for cryo-EM 3D reconstruction. Our major idea is to reorganize the related problem space as streams of key-value pairs, so that we can achieve both the flexibility and efficiency to compute and accumulate the contribution to the final 3D model from all different 2D image inputs. In addition, we design a hybrid communication mechanism to reduce intra-node communications and enable the solving process on a larger scale. With the addition of our GPU-based reconstruction design, we are able to improve the performance of the reconstruction part itself by 9.50 times, and the performance of the entire processing part (the reconstruction part and the other parts with mature GPU options) by 2.83 times. Moreover, Stream3D enables using the approach at a large scale, with 65.32-fold speedup when using up to 80 GPUs.
Shizhen Xu, Haohuan Fu, Hongkun Yu 0002, Wenlai Zhao, Guangwen Yang 0002
ICS6
2019 SunwayLB: Enabling Extreme-Scale Lattice Boltzmann Method Based Computing Fluid Dynamics Simulations on Sunway TaihuLight
abstract
The Lattice Boltzmann Method (LBM) is a relatively new class of Computational Fluid Dynamics methods. In this paper, we report our work on SunwayLB, which enables LBM based solutions aiming for industrial applications. We propose several techniques to boost the simulation speed and improve the scalability of SunwayLB, including a customized multi-level domain decomposition and data sharing scheme, a carefully orchestrated strategy to fuse kernels with different performance constraints for a more balanced workload, and optimization strategies for assembly code, which bring up to 137x speedup. Based on these optimization schemes, we manage to perform the largest direct numerical simulation which involves up to 5.6 trillion lattice cells, achieving 11,245 billion cell updates per second (GLUPS), 77% memory bandwidth utilization and a sustained performance of 4.7 PFlops. We also demonstrate a series of computational experiments for extreme-large scale fluid flow, as examples of real-world applications, to check the validity and performance of our work. The results show that SunwayLB is competent for a practical solution for industrial applications.
Xuesen Chu, Xiaojing Lv, Hongsong Meng, Shupeng Shi, Wenji Han, Jingheng Xu, Haohuan Fu, Guangwen Yang 0002
IPDPS9
2019 Pushing smart caching to the edge with BayCache
abstract
Caching contents in a small cell base station (SBS) is getting supported more and more widely today due to the Internet traffic growth and the requirement of low access latency. A primary concern and challenging issue with cache-enabled SBSs is how to better utilize network resources to achieve high overall performance. Though existing caching strategies can solve the problem to some extent, they are far from perfect --- pure popularity-based ones usually lead to sub-optimal caching performance whereas global coordination ones suffer from extra cost of many control messages. To deal with the issue, we present an adaptive caching scheme based on Bayesian inference, which 1) identifies the traffic features over time for each SBS; 2) synthesizes various features to rank the contents via a Bayesian ranking model; and 3) does cache placement online according to the ranking results. Unlike existing approaches that only highlight some specific factor, our scheme, due to the adoption of a Bayesian approach, can easily support additional features of high impact on caching performance and measure them in a decentralized way within a single SBS. We evaluate our scheme under various circumstances in terms of SBSs density, cache size, content popularity and skewness. The results show that our solution using multiple features exhibits improved performance --- it can reduce more than 30% the overall network latency in some cases when compared with solutions that only use a single feature.
Dongbiao He, Jinlei Jiang, Guangwen Yang 0002, Cédric Westphal
MobiQuitous3
2019 Towards Tile Based Distribution Simulation in Immersive Video Streaming
abstract
There has been increasing attention to virtual reality applications in recent years, especially to immersive or 360-degree videos that typically consume much more bandwidth than traditional ones. Though all produced data is transferred, only a small part (denoted as Field of View or viewport) is watched by users due to the nature of immersive videos. Obviously, this causes a large waste of network resources. Hence, it is important to define a viewport-dependent streaming transmission strategy by detecting where the user is gazing and the movement of the user's head. Unfortunately, there are few datasets providing this information. In this paper, we propose a tile-based simulation approach to generate the distribution of the user's behavior and to provide information that can be used to optimize future view-dependent streaming protocols. We first characterize the users' viewport pattern from datasets gathered from real users by decomposing the 360-degree stream into tiles and analyzing the frequency and time-interval distribution for each tile. Then, we devise a hierarchical Markov model that incorporates the beta distribution of each tile time interval to predict tile transition. The results show that the simulation tool characterizes the tile sequences of users accurately, performing close to the empirical results.
Dongbiao He, Cédric Westphal, Jinlei Jiang, Guangwen Yang 0002, J. J. Garcia-Luna-Aceves
Networking4
2019 GPU-based 3D cryo-EM reconstruction with key-value streams: poster
abstract
The 3D reconstruction of cryo-electron microscopy (cryo-EM) structural determination process is highly compute-intensive. It inherently requires accesses of a large 3D model in different and variable orientations, brings tough challenges to GPU architecture and has no effective solutions currently. To fill this gap, we propose a novel GPU-based parallel design for cryo-EM 3D reconstruction. The major idea is to reorganize the related problem space as streams of key-value pairs, so that we can achieve both the flexibility and efficiency to compute and accumulate the contribution to the final 3D model from all different 2D image inputs. In addition, we design a hybrid communication mechanism to reduce intra-node communications and enable the solving process on a larger scale.
Shizhen Xu, Hongkun Yu 0002, Haohuan Fu, Guangwen Yang 0002
PPoPP5
2019 SW_GROMACS: accelerate GROMACS on Sunway TaihuLight
abstract
GROMACS is one of the most popular Molecular Dynamic (MD) applications and is widely used in the field of chemical and bimolecular system study. Similar to other MD applications, it needs long run-time for large-scale simulations. Therefore, many high performance platforms have been employed to accelerate it, such as Knights Landing (KNL), Cell Processor, Graphics Processing Unit (GPU) and so on. As the third fastest supercomputer in the world, Sunway TaihuLight contains 40960 SW26010 processors and SW26010 is a typical many-core processor. To make full use of the superior computation ability of TaihuLight, we port GROMACS to SW26010 with following new strategies: (1) a new deferred update strategy; (2) a new update mark strategy; (3) a full pipeline acceleration. Furthermore, we redesign GROMACS to enable all possible vectorization. Experiments show that our implementation achieves better performance than both Intel KNL and Nvidia P100 GPU when using appropriate number of SW26010 processors for a fair comparison.
Tingjian Zhang, Ping Gao 0005, Mingshan Shao, Jinxiao Zhang, Xiaohui Duan, Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
SC14
2019 Extreme-scale earthquake simulations on Sunway TaihuLight
Haohuan Fu, Bingwei Chen, Wei Zhang 0321, Guangwen Yang 0002
CCF Trans. High Perform. Comput.6
2019 swTensor: accelerating tensor decomposition on Sunway architecture
Xiaogang Zhong, Hailong Yang 0002, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001
CCF Trans. High Perform. Comput.5
2019 RedSync: Reducing synchronization bandwidth for distributed deep learning training system
Jiarui Fang, Haohuan Fu, Guangwen Yang 0002, Cho-Jui Hsieh
J. Parallel Distributed Comput.3
2019 Performance Tuning and Analysis for Stencil-Based Applications on POWER8 Processor
abstract
This article demonstrates an approach for combining general tuning techniques with the POWER8 hardware architecture through optimizing three representative stencil benchmarks. Two typical real-world applications, with kernels similar to those of the winning programs of the Gordon Bell Prize 2016 and 2017, are employed to illustrate algorithm modifications and a combination of hardware-oriented tuning strategies with the application algorithms. This work fills the gap between hardware capability and software performance of the POWER8 processor, and provides useful guidance for optimizing stencil-based scientific applications on POWER systems.
Jingheng Xu, Haohuan Fu, Lin Gan 0001, Wayne Luk, Guangwen Yang 0002
ACM Trans. Archit. Code Optim.7
2019 Optimizing Finite Volume Method Solvers on Nvidia GPUs
abstract
As scientific applications are increasingly ported to GPUs to benefit from both the powerful computing capacity and high throughput, accelerating explicit solvers for GPU-based finite volume methods is gaining more and more attention. In this paper, based on the detailed analysis of the FVM algorithm, we present a set of novel optimization methods, including the explicit data cache mechanism, optimal global memory loading strategy, as well as the inner-thread rescheduling method, which derives a suitable mapping from the solver algorithm to the underlying GPU hardware architecture, so as to remarkably improve the solving performance of structured mesh based FVM. We demonstrate the impact of our tuning techniques on two widely-used atmospheric dynamic kernels (3-D Euler and 2-D SWE) on five kinds of mainstream GPU platforms, and make a detailed analysis of the different tuning methodologies so as to demonstrate how to select the proper tuning strategy to different applications on various GPU platforms. Specifically, 93.9x speedup is achieved for the 3D Euler solver on Nvidia V100 over one 12-core Intel E5-2697 (v2) CPU, which is a 77 percent improvement compared with the original speedup without adopting the tuning techniques presented in this work.
Jingheng Xu, Guangwen Yang 0002, Haohuan Fu, Wayne Luk, Lin Gan 0001, Wei Xue 0003, Chao Yang 0002, Yong Jiang 0001, Conghui He
IEEE Trans. Parallel Distributed Syst.2
2018 swCaffe: A Parallel Framework for Accelerating Deep Learning Applications on Sunway TaihuLight
abstract
This paper reports our efforts on swCaffe, a highly efficient parallel framework for accelerating deep neural networks (DNNs) training on Sunway TaihuLight, the current fastest supercomputer in the world that adopts a unique many-core heterogeneous architecture, with 40,960 SW26010 processors connected through a customized communication network.First, we point out some insightful principles to fully exploit the performance of the innovative many-core architecture.Second, we propose a set of optimization strategies for redesigning a variety of neural network layers based on Caffe.Third, we put forward a topology-aware parameter synchronization scheme to scale the synchronous Stochastic Gradient Descent (SGD) method to multiple processors efficiently.We evaluate our framework by training a variety of widely used neural networks with the ImageNet dataset.On a single node, swCaffe can achieve 23%˜119% overall performance compared with Caffe running on K40m GPU.As compared with the Caffe on CPU, swCaffe runs 3.04˜7.84xfaster on all the networks.Finally, we present the scalability of swCaffe for training of ResNet-50 and AlexNet on the scale of 1024 nodes.
Liandeng Li, Jiarui Fang, Haohuan Fu, Jinlei Jiang, Wenlai Zhao, Conghui He, Xin You 0001, Guangwen Yang 0002
CLUSTER8
2018 Working principles of binary differential evolution
abstract
We conduct a first fundamental analysis of the working principles of binary differential evolution (BDE), an optimization heuristic for binary decision variables that was derived by Gong and Tuson (2007) from the very successful classic differential evolution (DE) for continuous optimization. We show that unlike most other optimization paradigms, it is stable in the sense that neutral bit values are sampled with probability close to 1/2. This is generally a desirable property, however, it makes it harder to find the optima for decision variables with small influence on the objective function. This can result in an optimization time exponential in the dimension when optimizing simple symmetric functions like OneMax. On the positive side, BDE quickly detects and optimizes the most important decision variables. For example, dominant bits converge to the optimal value in time logarithmic in the population size. This leads to a very good performance in the situation where the decision variables have a differently strong influence on the result, in particular, when the target is not to find the optimal solution, but only a good one. Overall, our results indicate that BDE is an interesting optimization paradigm having characteristics significantly different from the classic evolutionary algorithms or EDAs.
Weijie Zheng 0001, Guangwen Yang 0002, Benjamin Doerr
GECCO2
2018 PLZMA: A Parallel Data Compression Method for Cloud Computing
Xin Wang 0233, Lin Gan 0001, Jingheng Xu, Jinzhe Yang, Maocai Xia, Haohuan Fu, Xiaomeng Huang, Guangwen Yang 0002
ICA3PP (3)8
2018 MCPC: Improving In-Network Caching with Network Partitions
abstract
In-network caching is considered to be an important solution to efficiently using network resources to achieve a high overall content delivery performance in both information-centric networks (ICNs) and 5G wireless networks. Content placement plays a key role in achieving this goal. Unfortunately, most content placement strategies today rely on opportunistic caching due to the problem complexity. We analyze the content placement problem in detail and present MCPC, a new content placement strategy for architectures that support in-network caching. Unlike existing content placement approaches that try to increase cache hit ratio, MCPC leverages the information recorded in each network node to reduce content access latency. MCPC proposes two new mechanisms: 1) a content load allocation estimation method based on local requests aggregation information; and 2) a content placement algorithm that avoids long-distance signaling messages by partitioning the network into smaller domains. We evaluate MCPC on a variety of network topologies, cache sizes and content popularity distributions. The experimental results show that MPCP can reduce by up to 56% the content access latency while providing comparable cache hit ratios as traditional benchmarks.
Dongbiao He, Jinlei Jiang, Guangwen Yang 0002, Cédric Westphal
ICPADS3
2018 A Fast Sparse Triangular Solver for Structured-grid Problems on Sunway Many-core Processor SW26010
abstract
The sparse triangular solver (SpTRSV) is one of the most essential kernels in many scientific and engineering applications. Efficiently parallelizing the SpTRSV on modern many-core architectures is considerably difficult due to inherent dependency of computation and discontinuous memory accesses. Achieving high performance of SpTRSV is even more challenging for SW26010, the new-generation customized heterogeneous many-core processor equipped in the top-rank Sunway TaihuLight supercomputer. Owing to regular sparse pattern, structured-grid triangular problems show much different computing characteristics with general ones as well as new opportunities to algorithm design on many-core architectures, which ever lacks attention. In this work, we focus on how to design and implement fast SpTRSV for structured-grid problems on SW26010. A generalized algorithm framework of parallel SpTRSV is proposed for best utilization of the features and flexibilities of SW26010 many-core architecture according to the fine-grained Producer-Consumer model. Moreover, a novel parallel structured-grid SpTRSV is presented by using direct data transfers across registers of the computing elements of SW26010. Experiments on four typical structured-grid triangular problems with different problem sizes demonstrate that our SpTRSV can achieve an average momory bandwidth utilization of 79.7% according to the stream benchmark, which leads to a speedup of 17.7 over serial version on SW26010. Furthermore, experiments with real world sparse linear problems show that our proposed SpTRSV can achieve superior preconditioning performance over the Intel Xeon E5-2670 v3 CPU and Intel Xeon Phi 7210 KNL over DDR4 memory.
Wei Xue 0003, Yulong Ao, Chao Yang 0002, Haohuan Fu, Lin Gan 0001, Guangwen Yang 0002
ICPP8
2018 CODA: Achieving Multipath Data Transmission in NDN
abstract
The exponential growth of data traffic raises a great challenge to content delivery in current TCP/IP networks. To answer this challenge, Information-Centric Networking (ICN) has been proposed with the purpose of bringing content caching and name-based content access to the network layer. Though great progress has been made, most existing ICN proposals lack support for parallel data transfer over multiple paths with low data redundancy. To deal with the issue, we present CODA, a fully distributed cooperative multipath data transmission solution that enhances content delivery further. Taking Named Data Networking (NDN) as a basis, CODA works in a distributed manner with the following contributions: 1) it extends the standard Interest model in NDN to support transmission of data over multiple paths so as to reduce the flow completion time; 2) it devises a traffic scheduling model to form parallel paths for transmitting data in a cooperative way; and 3) it proposes a transmission control scheme to select paths in an efficient and reliable manner. Extensive simulation comparisons with existing data transmission methods show that: 1) CODA speeds up the data rate twice as high as that of the best-route method; and 2) the amount of Interests required by CODA to build multiple data transmission paths in the network accounts for only 66% of that by MSRT, another multipath transmission proposal.
Dongbiao He, Jinlei Jiang, Guangwen Yang 0002, Cédric Westphal
IPCCC3
2018 Taming the "Monster": Overcoming Program Optimization Challenges on SW26010 Through Precise Performance Modeling
abstract
This paper presents an effort for overcoming the complexities of program optimizations on SW26010, the heterogeneous many-core processor that powers Sunway TaihuLight, the world top one supercomputer. The solution centers around a precise, static performance model for modern many-core processor. Through a careful design that leverages the special properties of SW26010 and an effective treatment to massive parallelism, the model achieves a high accuracy, showing less than 5% average errors in estimating program execution performance. The precise performance model opens many opportunities for analyzing and guiding code optimizations. The paper demonstrates the usefulness by revealing a series of insights on the effects of some important code optimizations on SW26010. Moreover, it demonstrates that with such a precise performance model, it is feasible to replace empirical auto-tuning with static auto-tuning for optimizing regular loops on heterogeneous many-core systems. Such a replacement speeds up the tuning process by as much as a factor of 43 while keeping the tuning quality loss below 6%.
Shizhen Xu, Yuanchao Xu 0001, Wei Xue 0003, Xipeng Shen, Fang Zheng 0015, Xiaomeng Huang, Guangwen Yang 0002
IPDPS7
2018 Simulating the Wenchuan earthquake with accurate surface topography on Sunway TaihuLight
Bingwei Chen, Haohuan Fu, Yanwen Wei, Conghui He, Wubin Wan, Lin Gan 0001, Wei Zhang 0321, Guangwen Yang 0002
SC12
2018 Redesigning LAMMPS for peta-scale and hundred-billion-atom simulation on Sunway TaihuLight
Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Wusheng Zhang, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Dexun Chen, Xiangxu Meng, Guangwen Yang 0002
SC12
2018 Large-scale hierarchical k-means for heterogeneous many-core supercomputers
Liandeng Li, Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002, John Thomson
SC7
2018 Application software beyond exascale: challenges and possible trends
abstract
With various exascale systems in different countries planned over the next three to five years, developing application software for such unprecedented computing capabilities and parallel scaling becomes a major challenge. In this study, we start our discussion with the current 125-Pflops Sunway TaihuLight system in China and its related application challenges and solutions. Based on our current experience with Sunway TaihuLight, we provide a projection into the next decade and discuss potential challenges and possible trends we would probably observe in future high performance computing software.
Guangwen Yang 0002, Haohuan Fu
Frontiers Inf. Technol. Electron. Eng.1
2018 Optimizing Convolutional Neural Networks on the Sunway TaihuLight Supercomputer
abstract
The Sunway TaihuLight supercomputer is powered by SW26010, a new 260-core processor designed with on-chip fusion of heterogeneous cores. In this article, we present our work on optimizing the training process of convolutional neural networks (CNNs) on the Sunway TaihuLight supercomputer. Specifically, a highly efficient library (swDNN) and a customized Caffe framework (swCaffe) are proposed. Architecture-oriented optimization methods targeting the many-core architecture of SW26010 are introduced and are able to achieve 48× speedup for the convolution routine in swDNN and 4× speedup for the complete training process of the VGG-16 network using swCaffe, compared to the unoptimized algorithm and framework. Compared to the cuDNN library and the Caffe framework based on the NVIDIA K40m GPU, the proposed swDNN library and swCaffe framework on SW26010 have nearly half the performance of K40m in single -precision and have 3.6× and 1.8× speedup over K40m in double precision, respectively.
Wenlai Zhao, Haohuan Fu, Jiarui Fang, Weijie Zheng 0001, Lin Gan 0001, Guangwen Yang 0002
ACM Trans. Archit. Code Optim.6
2018 Accelerating MapReduce on Commodity Clusters: An SSD-Empowered Approach
abstract
MapReduce, as a programming model and implementation for processing large data sets on clusters with hundreds or thousands of nodes, has gained wide adoption. In spite of the fact, we found that MapReduce on commodity clusters, which are usually equipped with limited memory and hard-disk drive (HDD) and have processors of multiple or many cores, does not scale as expected as the number of processor cores increases. The key reason for this is that the underlying low-speed HDD storage cannot meet the requirement of frequent IO operations. Though in-memory caching can improve IO, it is costly and sometimes cannot get the desired result either due to memory limitation. To deal with the problem and make MapReduce more scalable on commodity clusters, we present mpCache, a solution that utilizes solid-state drive (SSD) to cache input data and localized data of MapReduce tasks. In order to make a good trade-off between cost and performance, mpCache proposes ways to dynamically allocate the cache space between the input data and localized data and to do cache replacement. We have implemented mpCache in Hadoop and evaluated it on a 7-node commodity cluster by 13 benchmarks. The experimental results show that mpCache can gain an average speedup of 2.09× when compared with Hadoop, and can achieve an average speedup of 1.79× when compared with PACMan, the latest in-memory optimization of MapReduce.
Jinlei Jiang, Yongwei Wu 0001, Guangwen Yang 0002, Keqin Li 0001
IEEE Trans. Big Data4
2017 A Nanosecond-Level Hybrid Table Design for Financial Market Data Generators
abstract
This paper proposes a hybrid sorted table design for minimizing electronic trading latency, with three main contributions. First, a hierarchical sorted table with two levels, a fast cache table in reconfigurable hardware storing megabytes of data items and a master table in software storing gigabytes of data items. Second, a full set of operations, including insertion, deletion, selection and sorting, for the hybrid table with latency in a few cycles. Third, an on-demand synchronization scheme between the cache table and the master table. An implementation has been developed that targets an FPGA-based network card in the environment of the China Financial Futures Exchange (CFFEX) which sustains 1-10Gb/s bandwidth with latency of 400 to 700 nanoseconds, providing an 80- to 125-fold latency reduction compared to a fully optimized CPU-based solution, and a 2.2-fold reduction over an existing FPGA-based solution.
Haohuan Fu, Conghui He, Wayne Luk, Guangwen Yang 0002
FCCM5
2017 Accelerating Financial Market Server through Hybrid List Design (Abstract Only)
Haohuan Fu, Conghui He, Huabin Ruan, Itay Greenspon, Wayne Luk, Yongkang Zheng, Junfeng Liao, Guangwen Yang 0002
FPGA9
2017 swDNN: A Library for Accelerating Deep Learning Applications on Sunway TaihuLight
abstract
To explore the potential of training complex deep neural networks (DNNs) on other commercial chips rather than GPUs, we report our work on swDNN, which is a highly-efficient library for accelerating deep learning applications on the newly announced world-leading supercomputer, Sunway TaihuLight. Targeting SW26010 processor, we derive a performance model that guides us in the process of identifying the most suitable approach for mapping the convolutional neural networks (CNNs) onto the 260 cores within the chip. By performing a systematic optimization that explores major factors, such as organization of convolution loops, blocking techniques, register data communication schemes, as well as reordering strategies for the two pipelines of instructions, we manage to achieve a double-precision performance over 1.6 Tflops for the convolution kernel, achieving 54% of the theoretical peak. Compared with Tesla K40m with cuDNNv5, swDNN results in 1.91-9.75x performance speedup in an evaluation with over 100 parameter configurations.
Jiarui Fang, Haohuan Fu, Wenlai Zhao, Bingwei Chen, Weijie Zheng 0001, Guangwen Yang 0002
IPDPS6
2017 18.9-Pflops nonlinear earthquake simulation on Sunway TaihuLight: enabling depiction of 18-Hz and 8-meter scenarios
abstract
This paper reports our large-scale nonlinear earthquake simulation software on Sunway TaihuLight. Our innovations include: (1) a customized parallelization scheme that employs the 10 million cores efficiently at both the process and the thread levels; (2) an elaborate memory scheme that integrates on-chip halo exchange through register communcation, optimized blocking configuration guided by an analytic model, and coalesced DMA access with array fusion; (3) on-the-fly compression that doubles the maximum problem size and further improves the performance by 24%. With these innovations to remove the memory constraints of Sunway TaihuLight, our software achieves over 15% of the system's peak, better than the 11.8% efficiency achieved by a similar software running on Titan, whose byte to flop ratio is 5 times better than TaihuLight. The extreme cases demonstrate a sustained performance of over 18.9 Pflops, enabling the simulation of Tangshan earthquake as an 18-Hz scenario with an 8-meter resolution.
Haohuan Fu, Conghui He, Bingwei Chen, Zekun Yin, Tingjian Zhang, Wei Xue 0003, Wanwang Yin, Guangwen Yang 0002
SC11
2017 Redesigning CAM-SE for peta-scale climate modeling performance and ultra-high resolution on Sunway TaihuLight
abstract
The Community Atmosphere Model (CAM) is ported, redesigned, and scaled to the full system of the Sunway TaihuLight, and provides peta-scale climate modeling performance. We refactored and optimized the complete code using OpenACC directives at the first stage. A more aggressive and finer-grained redesign is then applied on the CAM, to achieve finer memory control and usage, more efficient vectorization and compute and communication overlapping. We further improve the CAM performance of a 260-core Sunway processor to the range of 28 to 184 Intel CPU cores, and achieve a sustainable double-precision performance of 3.3 PFlops for a 750 m global simulation when using 10,075,000 cores. CAM on Sunway achieves the simulation speed of 3.4 and 21.5 simulation-year-per-day (SYPD) for global 25-km and 100-km resolution respectively; and enables us to perform, to our knowledge, the first simulation of the complete lifecycle of hurricane Katrina, and achieve close-to-observation simulation results for both track and intensity.
Haohuan Fu, Junfeng Liao, Nan Ding 0006, Xiaohui Duan, Lin Gan 0001, Yishuang Liang, Jinzhe Yang, Lanning Wang, Guangwen Yang 0002
SC12
2017 Designing and implementing a heuristic cross-architecture combination for graph traversal
Yang You 0001, Haohuan Fu, David A. Bader, Guangwen Yang 0002
J. Parallel Distributed Comput.4
2017 An EnKF-based scheme to optimize hyper-parameters and features for SVM classifier
Yingsheng Ji, Yushu Chen, Haohuan Fu, Guangwen Yang 0002
Pattern Recognit.4
2017 A Fully-Pipelined Hardware Design for Gaussian Mixture Models
abstract
Gaussian Mixture Models (GMMs) are widely used in many applications such as data mining, signal processing and computer vision, for probability density modeling and soft clustering. However, the parameters of a GMM need to be estimated from data by, for example, the Expectation-Maximization algorithm for Gaussian Mixture Models (EM-GMM), which is computationally demanding. This paper presents a novel design for the EM-GMM algorithm targeting reconfigurable platforms, with five main contributions. First, a pipeline-friendly EM-GMM with diagonal covariance matrices that can easily be mapped to hardware architectures. Second, a function evaluation unit for Gaussian probability density based on fixed-point arithmetic. Third, our approach is extended to support a wide range of dimensions or/and components by fitting multiple pieces of smaller dimensions onto an FPGA chip. Fourth, we derive a cost and performance model that estimates logic resources. Fifth, our dataflow design targeting the Maxeler MPCX2000 with a Stratix-5SGSD8 FPGA can run over 200 times faster than a 6-core Xeon E5645 processor, and over 39 times faster than a Pascal TITAN-X GPU. Our design provides a practical solution to applications for training and explores better parameters for GMMs with hundreds of millions of high dimensional input instances, for low-latency and high-performance applications.
Conghui He, Haohuan Fu, Ce Guo 0002, Wayne Luk, Guangwen Yang 0002
IEEE Trans. Computers5
2016 Unleashing the performance potential of CPU-GPU platforms for the 3D atmospheric Euler solver
abstract
As a traditional application on various supercomputers, atmospheric modeling has long been suffering from the low performance efficiency. In this paper, we pick the 3D Euler equation solver (the most essential dynamic component for a non-hydrostatic atmospheric model) as the target application, and explore the maximum performance efficiency that can be achieved on CPU-GPU hybrid architectures. Besides presenting the suitable hybrid domain decomposition methodology and taking proper usage of tuning techniques for both the CPU and GPU parts, we further propose a novel GPU tuning technique, namely the customizable data caching mechanism with thread warp rescheduling scheme, which is specifically designed for the Euler solver. Combining all the optimizing approaches together, remarkable performance boost has been achieved on mainstream GPU architectures including Tesla Fermi C2050, K20×, K40 and K80. Especially, on the latest Tesla K80, we demonstrate a 31.64× speedup over the performance of 12-core E5-2697 CPU. In addition, based on a hybrid CPU-GPU node with two 12-core E5-2697 CPUs and two Tesla K80 GPUs, a sustained double-precision performance of 1.04 Tflops (16% of the peak) is achieved, which is remarkably higher than the efficiency of similar optimizing tasks based on heterogeneous platforms (strictly less than 10%, as demonstrated in the related work). In addition, a nearly linear weak scaling efficiency is achieved which demonstrate the effectiveness of our domain decomposition method.
Haohuan Fu, Jingheng Xu, Lin Gan 0001, Chao Yang 0002, Wei Xue 0003, Wenlai Zhao, Guangwen Yang 0002
ASAP9
2016 Performance optimization of Jacobi stencil algorithms based on POWER8 architecture
abstract
In this paper we choose the widely used Jacobi stencil algorithm as our target program to evaluate the effectiveness of tuning techniques based on the latest POWER8 processor, thus to provide optimization guidelines to similar stencil based algorithms.
Jingheng Xu, Haohuan Fu, Lin Gan 0001, Hongbo Peng, Guangwen Yang 0002
ASAP7
2016 F-CNN: An FPGA-based framework for training Convolutional Neural Networks
abstract
This paper presents a novel reconfigurable framework for training Convolutional Neural Networks (CNNs). The proposed framework is based on reconfiguring a streaming datapath at runtime to cover the training cycle for the various layers in a CNN. The streaming datapath can support various parameterized modules which can be customized to produce implementations with different trade-offs in performance and resource usage. The modules follow the same input and output data layout, simplifying configuration scheduling. For different layers, instances of the modules contain different computation kernels in parallel, which can be customized with different layer configurations and data precision. The associated models on performance, resource and bandwidth can be used in deriving parameters for the datapath to guide the analysis of design trade-offs to meet application requirements or platform constraints. They enable estimation of the implementation specifications given different layer configurations, to maximize performance under the constraints on bandwidth and hardware resources. Experimental results indicate that the proposed module design targeting Maxeler technology can achieve a performance of 62.06 GFLOPS for 32-bit floating-point arithmetic, outperforming existing accelerators. Further evaluation based on training LeNet-5 shows that the proposed framework achieves about 4 times faster than CPU implementation of Caffe and about 7.5 times more energy efficient than the GPU implementation of Caffe.
Wenlai Zhao, Haohuan Fu, Wayne Luk, Yuchun Ma, Guangwen Yang 0002
ASAP8
2016 Graph-Oriented Code Transformation Approach for Register-Limited Stencils on GPUs
abstract
Stencil kernels play an important role in many scientific and engineering disciplines. With the development of numerical algorithms and the increasing requirements of accuracy, register-limited stencils containing massive variables and operations are widely used. However, these register-limited stencils consume vast resources when executing on GPUs. The excessive use of registers reduces the number of active threads dramatically, and consequently leads to a serious performance decline. To improve the performance of these register-limited stencils, we propose a DDG (data-dependency-graph) oriented code transformation approach in this paper. By analyzing, deleting and transforming the original stencil program on GPUs, our graph-oriented code transformation approach explores for the best trade-off between the calculation amount and the parallelism degree, and further achieves better performance. The graph-oriented code transformation approach is evaluated using the Weighted Nearly Analytic Discrete stencil, and the experimental result shows that a speedup of 2.16X can be achieved when compared with the original fairly-optimized implementation. To the best of our knowledge, our study takes the first step towards balancing the calculation amount and parallelism degree of the extremely register-limited stencils on GPUs.
Mengyao Jin, Haohuan Fu, Zihong Lv, Guangwen Yang 0002
CCGrid4
2016 Generalized GPU Acceleration for Applications Employing Finite-Volume Methods
abstract
Scientific HPC applications are increasingly ported to GPUs to benefit from both the high throughput and the powerful computing capacity. Many of these applications, such as atmospheric modeling and hydraulic erosion simulation, are adopting the finite volume method (FVM) as the solver algorithm. However, the communication components inside these applications generally lead to a low flop-to-byte ratio and an inefficient utilization of GPU resources. This paper aims at optimizing FVM solver based on the structured mesh. Besides a high-level overview of the finite-volume method as well as its basic optimizations on modern GPU platforms, we further present two generalized tuning techniques including an explicit cache mechanism as well as an inner-thread rescheduling method that tries to achieve a suitable mapping between the algorithm feature and the platform architecture. To the end, we demonstrate the impact of our generalized optimization methods in two typical atmospheric dynamic kernels (Euler and SWE) based on four mainstream GPU platforms. According to the experimental results of Tesla K80, speedups of 24.4x for SWE and 31.5x for Euler could be achieved over a 12-core Intel E5-2697 CPU, which is a great promotion compared with its original speedup (18x and 15.47x) without applying these two methods.
Jingheng Xu, Haohuan Fu, Lin Gan 0001, Chao Yang 0002, Wei Xue 0003, Shizhen Xu, Wenlai Zhao, Bingwei Chen, Guangwen Yang 0002
CCGrid10
2016 Cache-Friendly Design for Complex Spatially-Variable Coefficient Stencils on Many-Core Architectures
abstract
Many-core architectures, such as the NVIDIA graphics processing unit and Intel Xeon Phi, which are characterized by high computation resources but limited on-chip memory capacity, have been used to significantly accelerate various computationally demanding tasks. Stencil operators are naturally suitable for such architectures because of their parallel calculation patterns. However, only simple stencils with points distributed along the axes and with constant coefficients have been fully investigated. This study first provides insights into optimization strategies for stencils with complex shapes, including off-axial points and spatially variable coefficients. Through our proposed stencil-decomposition schemes, we maintain read-only coefficients in on-chip caches to avoid unvectorized memory access. To alleviate the resulting severe cache-starvation situation, a generalized cache-friendly design for many-core architecture is proposed. It can reduce cache miss times and cache space consumption. The proposed methodology significantly improves the performance of stencil operations in a real seismic imaging application and introduces a new option to write highly efficient memory-bound stencil-like loops.
Jiarui Fang, Haohuan Fu, Guangwen Yang 0002
HiPC3
2016 TADE: Tight Adaptive Differential Evolution
Weijie Zheng 0001, Haohuan Fu, Guangwen Yang 0002
PPSN3
2016 Refactoring and optimizing the community atmosphere model (CAM) on the sunway taihulight supercomputer
abstract
This paper reports our efforts on refactoring and optimizing the Community Atmosphere Model (CAM) on the Sunway TaihuLight supercomputer, which uses a many-core processor that consists of management processing elements (MPEs) and clusters of computing processing elements (CPEs). To map the large code base of CAM to the millions of cores on the Sunway system, we take OpenACC-based refactoring as the major approach, and apply source-to-source translator tools to exploit the most suitable parallelism for the CPE cluster, and to fit the intermediate variable into the limited on-chip fast buffer. For individual kernels, when comparing the original ported version using only MPEs and the refactored version using both the MPE and CPE clusters, we achieve up to 22× speedup for the compute-intensive kernels. For the 25km resolution CAM global model, we manage to scale to 24,000 MPEs, and 1,536,000 CPEs, and achieve a simulation speed of 2.81 model years per day.
Haohuan Fu, Junfeng Liao, Wei Xue 0003, Lanning Wang, Dexun Chen, Long Gu, Jinxiu Xu 0001, Nan Ding 0006, Conghui He, Shizhen Xu, Yishuang Liang, Jiarui Fang, Yuanchao Xu 0001, Weijie Zheng 0001, Jingheng Xu, Zhen Zheng, Wanjing Wei, Bingwei Chen, Xiaomeng Huang, Guangwen Yang 0002
SC25
2016 10M-core scalable fully-implicit solver for nonhydrostatic atmospheric dynamics
abstract
An ultra-scalable fully-implicit solver is developed for stiff time-dependent problems arising from the hyperbolic conservation laws in nonhydrostatic atmospheric dynamics. In the solver, we propose a highly efficient hybrid domain-decomposed multigrid preconditioner that can greatly accelerate the convergence rate at the extreme scale. For solving the overlapped subdomain problems, a geometry-based pipelined incomplete LU factorization method is designed to further exploit the on-chip fine-grained concurrency. We perform systematic optimizations on different hardware levels to achieve best utilization of the heterogeneous computing units and substantial reduction of data movement cost. The fully-implicit solver successfully scales to the entire system of the Sunway TaihuLight supercomputer with over 10.5M heterogeneous cores, sustaining an aggregate performance of 7.95 PFLOPS in double-precision, and enables fast and accurate atmospheric simulations at the 488-m horizontal resolution (over 770 billion unknowns) with 0.07 simulated-years-per-day. This is, to our knowledge, the largest fully-implicit simulation to date.
Chao Yang 0002, Wei Xue 0003, Haohuan Fu, Hongtao You, Yulong Ao, Fangfang Liu 0004, Lin Gan 0001, Lanning Wang, Guangwen Yang 0002
SC11
2016 The Sunway TaihuLight supercomputer: system and applications
Haohuan Fu, Junfeng Liao, Jinzhe Yang, Lanning Wang, Zhenya Song, Xiaomeng Huang, Chao Yang 0002, Wei Xue 0003, Fangfang Liu 0004, Fangli Qiao, Xunqiang Yin, Chaofeng Hou, Jian Zhang 0070, Yangang Wang 0002, Chunbo Zhou, Guangwen Yang 0002
Sci. China Inf. Sci.19
2015 Optimizing Residue Number Reverse Converters through Bitwise Arithmetic on FPGAs
abstract
As a promising number representation method to provide inspiring operational performance, the Residue Number System (RNS) has been widely applied in many key applications for data pocessing. However, a highly-efficient and general-purpose reverse converter, which is the key component in an RNS system, is still less to be seen, due to the costly and complex operators that require large amounts of computing resources and a long latency to accomplish. In this paper, we are targeting at reverse converters that are highly efficient and can support general moduli sets. We first propose optimizing methods based on the bit wise arithmetic to improve the performance of general reverse converters such as CRT and New CRT. The methods are capable of replacing expensive operations such as additions and multiplications with bit wise operations. We also optimize the performance of specific reverse converter through condition reduction and pre-calculation methods. Furthermore, we develop a user controlled FPGA design generator that can produce optimized reverse converter designs for a number of different moduli sets. Compared with the existing optimized converter designs, our proposed methods can further reduce the latency and resource consumption by 54.2% to 84.6% and 65% to 88.5% respectively.
Bangtian Liu, Haohuan Fu, Lin Gan 0001, Wenlai Zhao, Guangwen Yang 0002
FCCM5
2015 CSAP: A Performance Predictor for Climate Simulation Applications on Intel CPUs
Guangwen Yang 0002
ICA3PP (4)4
2015 Performance Characterization and Optimization for Intel Xeon Phi Coprocessor
Guangwen Yang 0002
ICA3PP (1)4
2015 Optimizing Complex Spatially-Variant Coefficient Stencils for Seismic Modeling on GPU
abstract
The Explicit Time Evolution (ETE) method is an innovative Finite-Difference (FD) type method to simulate the wave propagation in acoustic media with higher spatial and temporal accuracy. However, different from FD, it is difficult to achieve an efficient GPU design because of the poor memory access patterns caused by the off-axis points and spatially-variant coefficients. In this paper, we present a set of new optimization strategies for ETE stencils according to the memory hierarchy of NVIDIA GPU. To handle the problem caused by the complexity of the stencil shapes, we design a one-to-multi updating scheme for shared memory usage. To alleviate the performance damage resulted from the poor memory access pattern of reading spatially-variant coefficients, we propose a stencil decomposition method to reduce un-coalesced global memory access. Based on the state-of-the-art GPU architecture, combining with existing spatial and temporal stencil blocking schemes, we manage to achieve 9.6x and 9.9x speedups compared with a well-tuned 12-core CPUs version for 37-point and 73-point ETE stencils, respectively. Compared with a well-tuned MIC version, the best speedups for the 2 type stencils are 3.7x and 4.7x. Our designs leads to an ETE method that is 31.2x faster than conventional CPU-FD method and make it a practical seismic imaging technology.
Jiarui Fang, Haohuan Fu, Nanxun Dai, Lin Gan 0001, Guangwen Yang 0002
ICPADS7
2015 Targeted Mutation: A Novel Mutation Strategy for Differential Evolution
abstract
Differential Evolution (DE) has been shown as an effective, efficient and robust evolutionary computing algorithm. The main force to generate promising offspring is the mutation operator. Usually, two randomly selected vectors are used to generate the differential vector, which maintains the large diversity of mutant directions and ensures the possibility to find global optima. However, strong randomness also leads to the ineffective searching and slow convergence speed. A proper degree of certainty in differential vector will help the population evolve efficiently. This paper proposes a novel mutation strategy called Targeted Mutation that takes the determined target vector as the starting point of the differential vector and maintains the randomness of the ending point, which makes a better trade-off between the certainty and randomness in the differential vector. Besides, Targeted Mutation adopts the best vector as the base vector. The extensive experiments of comparison with two popular mutation operators on 20 benchmark functions demonstrate the competitive performance of our proposed targeted mutation scheme. Our method achieves better or equivalent performance over 70% of total benchmarks against the other two methods. 17 out of 20 function results can get further improved when roughly tuning parameters on each function, showing the potential ability to get even better results. In addition, an integrated evaluation scoring scheme is designed to provide a more concrete demonstration of the overall performance of different approaches, and our method gains the highest score.
Weijie Zheng 0001, Haohuan Fu, Guangwen Yang 0002
ICTAI3
2015 Improving the scalability of the ocean barotropic solver in the community earth system model
abstract
High-resolution climate simulations are increasingly in demand and require tremendous computing resources. In the Community Earth SystemModel (CESM), the Parallel Ocean Model (POP) is computationally expensive for high-resolution grids (e.g., 0.1°) and is frequently the least scalable component of CESM for certain production simulations. In particular, the modified Preconditioned Conjugate Gradient (PCG), used to solve the elliptic system of equations in the barotropic mode, scales poorly at the high core counts, which is problematic for high-resolution simulations. In this work, we demonstrate that the communication costs in the barotropic solver occupy an increasing portion of the total POP execution time as core counts are increased. To mitigate this problem, we implement a preconditioned Chebyshev-type iterative method in POP (called P-CSI), which requires far fewer global reductions than PCG. We also develop an effective block preconditioner based on the Error Vector Propagation Method to attain a competitive convergence rate for P-CSI. We demonstrate that the improved scalability of P-CSI results in a 5.2x speedup of the barotropic mode in high-resolution POP on 16,875 cores, which yields a 1.7x speedup of the overall POP simulation. Further, we ensure that the new solver produces an ocean climate consistent with the original one via an ensemble-based statistical method.
Xiaomeng Huang, Allison H. Baker, Yu-heng Tseng, Frank O. Bryan, John M. Dennis, Guangwen Yang 0002
SC7
2015 Scaling Support Vector Machines on modern HPC platforms
Yang You 0001, Haohuan Fu, Shuaiwen Song, Amanda Randles, Darren J. Kerbyson, Andrés Márquez 0001, Guangwen Yang 0002, Adolfy Hoisie
J. Parallel Distributed Comput.7
2015 Solving the Global Atmospheric Equations through Heterogeneous Reconfigurable Platforms
abstract
One of the most essential and challenging components in climate modeling is the atmospheric model. To solve multiphysical atmospheric equations, developers have to face extremely complex stencil kernels that are costly in terms of both computing and memory resources. This article aims to accelerate the solution of global shallow water equations (SWEs), which is one of the most essential equation sets describing atmospheric dynamics. We first design a hybrid methodology that employs both the host CPU cores and the field-programmable gate array (FPGA) accelerators to work in parallel. Through a careful adjustment of the computational domains, we achieve a balanced resource utilization and a further improvement of the overall performance. By decomposing the resource-demanding SWE kernel, we manage to map the double-precision algorithm into three FPGAs. Moreover, by using fixed-point and reduced-precision floating point arithmetic, we manage to build a fully pipelined mixed-precision design on a single FPGA, which can perform 428 floating-point and 235 fixed-point operations per cycle. The mixed-precision design with four FPGAs running together can achieve a speedup of 20 over a fully optimized design on a CPU rack with two eight-core processorsand is 8 times faster than the fully optimized Kepler GPU design. As for power efficiency, the mixed-precision design with four FPGAs is 10 times more power efficient than a Tianhe-1A supercomputer node.
Lin Gan 0001, Haohuan Fu, Wayne Luk, Chao Yang 0002, Wei Xue 0003, Xiaomeng Huang, Youhui Zhang, Guangwen Yang 0002
ACM Trans. Reconfigurable Technol. Syst.8
2014 An approach of processor core customization for stencil computation
abstract
Architecture customization is believed as one of the most promising methods to meet ever-increasing computing needs and power density limitations. This paper presents an approach to enhance a preliminary customizable core with some common architecture features, to adapt to the specific applications while keeping the programming flexibility. Those features include several effective software/hardware co-optimizing strategies, such as loop tiling, pre-fetching, cache customization, customized Single Instruction Multiple Data (SIMD) and Direct Memory Access (DMA), as well as the necessary ISA extensions. Currently we select stencil computation as the research target. Detailed tests of power-efficiency to evaluate the effect of all these optimizations comprehensively shows impressive performance speedup and power efficiency, even compared to X86, GPU and FPGA platforms. All these proposed customizations here could be applied to other computing applications.
Youhui Zhang, Wayne Luk, Guangwen Yang 0002
ASAP5
2014 A Fully-Pipelined FPGA Design for Tree-Reweighted Message Passing Algorithm
Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002
FCCM3
2014 A highly-efficient and green data flow engine for solving euler atmospheric equations
abstract
Atmospheric modeling is an essential issue in the study of climate change. However, due to the complicated algorithmic and communication models, scientists and researchers are facing tough challenges in finding efficient solutions to solve the atmospheric equations. In this paper, we accelerate a solver for the three-dimensional Euler atmospheric equations through reconfigurable data flow engines. We first propose a hybrid design that achieves efficient resource allocation and data reuse. Furthermore, through algorithmic offsetting, fast memory table, and customizable-precision arithmetic, we map a complex Euler kernel into a single FPGA chip, which can perform 956 floating point operations per cycle. In a 1U-chassis, our CPU-DFE unit with 8 FPGA chips is 18.5 times faster and 8.3 times more power efficient than a multicore system based on two 12-core Intel E5-2697 (Ivy Bridge) CPUs, and is 6.2 times faster and 5.2 times more power efficient than a hybrid unit equipped with two 12-core Intel E5-2697 (Ivy Bridge) CPUs and three Intel Xeon Phi 5120d (MIC) cards.
Lin Gan 0001, Haohuan Fu, Chao Yang 0002, Wayne Luk, Wei Xue 0003, Oskar Mencer, Xiaomeng Huang, Guangwen Yang 0002
FPL8
2014 Patra: Parallel tree-reweighted message passing architecture
abstract
Maximum a posteriori probability inference algorithms for Markov Random Field are widely used in many applications, such as computer vision and machine learning. Sequential tree-reweighted message passing (TRW-S) is an inference algorithm which shows good quality in finding optimal solutions. However, the performance of TRW-S in software cannot meet the requirements of many real-time applications, due to the sequential scheme and the high memory, bandwidth and computational costs. This paper proposes Patra, a novel parallel tree-reweighted message passing architecture, which involves a fully pipelined design targeting FPGA technology. We build a hybrid CPU/FPGA system to test the performance of Patra for stereo matching. Experimental results show that Patra provides about 100 times faster than a software implementation of TRW-S, and 12 times faster than a GPU-based message passing algorithm. Compared with an existing design in four FPGAs, we can achieve 2 times speedup in a single FPGA. Moreover, Patra can work at video rate in many cases, such as a rate of 167 frame/sec for a standard stereo matching test case, which makes it promising for many real-time applications.
Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002, Wayne Luk
FPL3
2014 Porting the Princeton Ocean Model to GPUs
Shizhen Xu, Xiaomeng Huang, Haohuan Fu, Guangwen Yang 0002
ICA3PP (1)6
2014 Scaling and analyzing the stencil performance on multi-core and many-core architectures
abstract
Stencils are among the most important and time-consuming kernels in many applications. While stencil optimization has been a well-studied topic on CPU platforms, achieving higher performance and efficiency for the evolving numerical stencils on the more recent multi-core and many-core architectures is still an important issue. In this paper, we explore a number of different stencils, ranging from a basic 7-point Jacobi stencil to more complex high-order stencils used in finer numerical simulations. By optimizing and analyzing those stencils on the latest multi-core and many-core architectures (the Intel Sandy Bridge processor, the Intel Xeon Phi coprocessor, and the NVIDIA Fermi C2070 and Kepler K20x GPUs), we investigate the algorithmic and architectural factors that determine the performance and efficiency of the resulting designs. While multi-threading, vectorization, and optimization on cache and other fast buffers are still the most important techniques that provide performance, we observe that the different memory hierarchy and the different mechanism for issuing and executing parallel instructions lead to the different performance behaviors on CPU, MIC and GPU. With vector-like processing units becoming the major provider of computing power on almost all architectures, the compiler's inability to align all the computing and memory operations would become the major bottleneck from getting a high efficiency on current and future platforms. Our specific optimization of the complex WNAD stencil on GPU provides a good example of what the compiler could do to help.
Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Yangtong Xu, Chao Yang 0002, Zihong Lv, Yang You 0001, Guangwen Yang 0002, Kaijian Ou
ICPADS9
2014 MIC-SVM: Designing a Highly Efficient Support Vector Machine for Advanced Modern Multi-core and Many-Core Architectures
abstract
Support Vector Machine (SVM) has been widely used in data-mining and Big Data applications as modern commercial databases start to attach an increasing importance to the analytic capabilities. In recent years, SVM was adapted to the field of High Performance Computing for power/performance prediction, auto-tuning, and runtime scheduling. However, even at the risk of losing prediction accuracy due to insufficient runtime information, researchers can only afford to apply offline model training to avoid significant runtime training overhead. Advanced multi- and many-core architectures offer massive parallelism with complex memory hierarchies which can make runtime training possible, but form a barrier to efficient parallel SVM design. To address the challenges above, we designed and implemented MIC-SVM, a highly efficient parallel SVM for x86 based multi-core and many-core architectures, such as the Intel Ivy Bridge CPUs and Intel Xeon Phi co-processor (MIC). We propose various novel analysis methods and optimization techniques to fully utilize the multilevel parallelism provided by these architectures and serve as general optimization methods for other machine learning tools. MIC-SVM achieves 4.4-84x and 18-47x speedups against the popular LIBSVM, on MIC and Ivy Bridge CPUs respectively, for several real-world data-mining datasets. Even compared with GPUSVM, run on a top of the line NVIDIA k20x GPU, the performance of our MIC-SVM is competitive. We also conduct a cross-platform performance comparison analysis, focusing on Ivy Bridge CPUs, MIC and GPUs, and provide insights on how to select the most suitable advanced architectures for specific algorithms and input data patterns.
Yang You 0001, Shuaiwen Song, Haohuan Fu, Andrés Márquez 0001, Maryam Mehri Dehnavi, Kevin J. Barker, Kirk W. Cameron, Amanda Randles, Guangwen Yang 0002
IPDPS9
2014 A High Performance Compression Method for Climate Data
abstract
Climate modeling data are usually multidimensional arrays of floating-point numbers. These arrays typically have two or three spatial dimensions and one temporal dimension, describing the evolvement of climate variables in a time span. With the advances of high performance computing, the volume of climate data is expanding exponentially, bringing tough challenges for climate data archiving and sharing. In this paper, we propose a lossless compression algorithm for the time-spatial climate floating-point arrays. Our compression algorithm can eliminate more data redundancy efficiently through adaptive prediction, XOR-differencing, and multi-way compression. In addition, static regions, which are very common in climate data, can be identified and compressed more efficiently. Moreover, to utilize the multi-cores on modern computers, we proposed a method to parallelize our compression algorithm. Evaluations demonstrate that single thread version of our compression method can achieve the best balance in compression ratios, deflating throughputs and inflating throughputs. And the parallel version can achieve 800 MB/s deflating throughputs and over 2600 MB/s inflating throughputs on a 16-core server.
Songbin Liu, Xiaomeng Huang, Yufang Ni, Haohuan Fu, Guangwen Yang 0002
ISPA5
2014 mpCache: Accelerating MapReduce with Hybrid Storage System on Many-Core Clusters
Jinlei Jiang, Guangwen Yang 0002
NPC3
2014 CFIO2: Overlapping Communications and I/O with Computations Using RDMA Technology
Xiaomeng Huang, Shizhen Xu, Haohuan Fu, Guangwen Yang 0002
NPC6
2013 Understanding Data Characteristics and Access Patterns in a Cloud Storage System
abstract
Understanding the inherent system characteristics is crucial to the design and optimization of cloud storage system, and few studies have systematically investigated its data characteristics and access patterns. This paper presents an analysis of file system snapshot and five-month access trace of a campus cloud storage system that has been deployed on Tsinghua campus for three years. The system provides online storage and data sharing services for more than 19,000 students and 500 student groups. We report several data characteristics including file size and file type, as well as some access patterns, including read/write ratio, read-write dependency and daily traffic. We find that there are many differences between cloud storage system and traditional file systems: our cloud storage system has larger file sizes, lower read/write ratio, and smaller set of active files than those of a typical traditional file system. With a trace-driven simulation, we find that the cache efficiency can be improved by 5 times using the guidance from our observations.
Songbin Liu, Xiaomeng Huang, Haohuan Fu, Guangwen Yang 0002
CCGRID4
2013 A Scalable Barotropic Mode Solver for the Parallel Ocean Program
Xiaomeng Huang, Xiaoge Wang, Haohuan Fu, Shizhen Xu, Huabin Ruan, Wei Xue 0003, Guangwen Yang 0002
Euro-Par8
2013 Global Atmospheric Simulation on a Reconfigurable Platform
abstract
Summary form only given. As the only method to study long-term climate trend and to predict potential climate risk, climate modeling is becoming a key research topic among governments and research organizations. One of the most essential and challenging components in climate modeling is the atmospheric model. To cover high resolution in climate simulation scenarios, developers have to face the challenges from billions of mesh points and extremely complex algorithms. Shallow Water Equations (SWEs) are a set of conservation laws that perform most of the essential characteristics of the atmosphere. The study of SWEs can serve as the starting point for understanding the dynamic behavior of the global atmosphere. We choose cubed-sphere mesh as the computational mesh for its better load balance in pole regions over other meshes such as the latitude-longitude mesh. The cubed-sphere mesh is obtained by mapping a cube to the surface of the sphere. The computational domain is then the six patches, each of which is covered with N × N mesh points to be calculated. When written in local coordinates, SWEs have an identical expression on the six patches, that is ∂Q/∂t + 1/Λ ∂(ΛF1)/∂x1+ 1/Λ ∂(ΛF1)/∂z2+ S=0, (1) where (x1, x2) ∈ [-π/4, π/4] are the local coordinates, Q = (h, hu1, hu2)Tis the prognostic variable, Fi= uiQ (i = 1, 2) are the convective fluxes, S is the source term. Spatially discretized with a cell-centered finite volume method and integrated with a second-order accurate TVD Runge-Kutta method, SWE solvers are transferred to the computation of a 13-point upwind stencil that exhibits a diamond shape. To get the prognostic components (h, hu1and hu2) of the central point, its neighboring 12 points need to be accessed. The stencil kernel includes at least 434 ADD/SUB operations, 570 multiplications, 99 divisions. The high arithmetic density of the SWEs algorithm makes it difficult to implement one kernel into the resource-limited FPGA card. In this study, we first proposes a hybrid algorithm that utilizes both CPUs and FPGAs to simulate the global shallow water equations (SWEs). In each of the computational patch, most of the complicated communications happen in the two layers of the outer boundary, whose value need to be exchanged with other patches. Therefore, we decompose each of the six patches into an outer part that includes two layers of the outer boundary meshes, and an inner part that is the remaining part. We assign CPU to handle the communications and the stencil calculation of the outer part, while assign FPGA to process the inner-part stencil. In this way, FPGA and CPU will work simultaneously and the CPU time for stencil and communication can be hidden in the FPGA time for stencil. For the Virtex-6 SX475T that we use in our study, the original program in double-precision will require 299% of the on-board LUTs, 283% of the FFs and 189% of the DSPs, and cannot fit into one FPGA. In order to fit the SWE kernel into one FPGA chip, we apply two algorithmic optimizations to the original design. One is to replace certain computations by lookup tables, so as to reduce the usage of computation resources. The other one is to locate common factors in the algorithm and to remove redundant computations. These two optimizations reduce the resource usage by 20%. To further reduce the resource cost and to fit the extremely complex stencil kernel into one FPGA chip, we perform optimization in the space of customizable representations and precisions. For the variables with a relatively small range, we apply fixed-point number to replace the double-precisions. For the rest parts with a wide dynamic range, we use floating-point numbers with a mixed-precision. Through mixed-precision floating-point and fixed-point arithmetic, we build a complex upwind stencil kernel on a single FPGA. The design includes a highly-efficient pipeline that can perform hundreds of floating-point and fixed-point arithmetic operations concurrently. Compared with our previous work in [1], the solution based on one FPGA acceleration card provides 100 times speedup over a 6-core CPU, and 4 times speedup over a Tianhe-1A supercomputer node that consists of 12 CPU cores and one Fermi GPU.
Lin Gan 0001, Haohuan Fu, Wayne Luk, Chao Yang 0002, Wei Xue 0003, Guangwen Yang 0002
FCCM6
2013 An FPGA-Based Data Flow Engine for Gaussian Copula Model
abstract
The Gaussian Copula Model (GCM) plays an important role in the state-of-the-art financial analysis field for modeling the dependence of financial assets. However, the existing implementations of GCM are all computationallydemanding and time-consuming. In this paper, we propose a Dataflow Engine (DFE) design to accelerate the GCM computation. Specifically, a commonly used CPU-friendly GCM algorithm is converted into a fully-pipelined dataflow graph through four steps of optimization: recomposing the algorithm to be pipeline-friendly, removing unnecessary computation, sharing common computing results, and reducing the computing precision while maintaining the same level of accuracy for the computation results. The performance of the proposed DFE design is compared with three CPU-based implementations that are well-optimized. Experimental results show that our DFE solution not only generates fairly accurate result, but also achieves a maximum of 467x speedup over a single-thread CPU-based solution, 120x speedup over a multi-thread CPUbased solution, and 47x speedup over an MPI-based solution.
Huabin Ruan, Xiaomeng Huang, Haohuan Fu, Guangwen Yang 0002, Wayne Luk, Sébastien Racanière, Oliver Pell, Wenjing Han
FCCM4
2013 Accelerating solvers for global atmospheric equations through mixed-precision data flow engine
abstract
One of the most essential and challenging components in a climate system model is the atmospheric model. To solve the multi-physical atmospheric equations, developers have to face extremely complex stencil kernels. In this paper, we propose a hybrid CPU-FPGA algorithm that applies single and multiple FPGAs to compute the upwind stencil for the global shallow water equations. Through mixed-precision arithmetic, we manage to build a fully pipelined upwind stencil design on a single FPGA, which can perform 428 floating-point and 235 fixed-point operations per cycle. The CPU-FPGA algorithm using one Virtex-6 FPGA provides 100 times speedup over a 6-core CPU and 4 times speedup over a hybrid node with 12 CPU cores and a Fermi GPU card. The algorithm using four FPGAs provides 330 times speedup over a 6-core CPU; it is also 14 times faster and 9 times more power efficient than the hybrid CPU-GPU node.
Lin Gan 0001, Haohuan Fu, Wayne Luk, Chao Yang 0002, Wei Xue 0003, Xiaomeng Huang, Youhui Zhang, Guangwen Yang 0002
FPL8
2013 Optimize Multidimensional Arrays Queries with Heterogeneous Replica Method
abstract
Multidimensional arrays are commonly used in scientific and engineering applications. The disk layout for the multidimensional arrays will obviously affect the performance of data querying. Homogeneous Replica method are widely used to maintain the data reliability in most of the distributed storage systems and used to improve the data locality in some parallel processing systems. In this paper, we propose a novel method, that is heterogeneous replicas, to makes better use of the replica method to optimize the performance of multidimensional arrays querying. The experimental results shows that heterogeneous replicas method can significantly reduce the overhead of disk I/O for most of the queries. With three heterogeneous replicas, the performance of random generated range queries for multidimensional datasets can be improved for 30% on the average.
Xiaomeng Huang, Songbin Liu, Haohuan Fu, Qiming Fang, Guangwen Yang 0002
NAS6
2013 A peta-scalable CPU-GPU algorithm for global atmospheric simulations
abstract
Developing highly scalable algorithms for global atmospheric modeling is becoming increasingly important as scientists inquire to understand behaviors of the global atmosphere at extreme scales. Nowadays, heterogeneous architecture based on both processors and accelerators is becoming an important solution for large-scale computing. However, large-scale simulation of the global atmosphere brings a severe challenge to the development of highly scalable algorithms that fit well into state-of-the-art heterogeneous systems. Although successes have been made on GPU-accelerated computing in some top-level applications, studies on fully exploiting heterogeneous architectures in global atmospheric modeling are still very less to be seen, due in large part to both the computational difficulties of the mathematical models and the requirement of high accuracy for long term simulations.
Chao Yang 0002, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Yangtong Xu, Yutong Lu, Jiachang Sun, Guangwen Yang 0002
PPoPP9
2011 Location-Aware MapReduce in Virtual Cloud
abstract
MapReduce is an important programming model for processing and generating large data sets in parallel. It is commonly applied in applications such as web indexing, data mining, machine learning, etc. As an open-source implementation of MapReduce, Hadoop is now widely used in industry. Virtualization, which is easy to configure and economical to use, shows great potential for cloud computing. With the increasing core number in a CPU and involving of virtualization technique, one physical machine can hosts more and more virtual machines, but I/O devices normally do not increase so rapidly. As MapReduce system is often used to running I/O intensive applications, decreasing of data redundancy and load unbalance, which increase I/O interference in virtual cloud, come to be serious problems. This paper builds a model and defines metrics to analyze the data allocation problem in virtual environment theoretically. And we design a location-aware file block allocation strategy that retains compatibility with the native Hadoop. Our model simulation and experiment in real system shows our new strategy can achieve better data redundancy and load balance to reduce I/O interference. Execution time of applications such as RandomWriter, Text Sort and Word Count are reduced by up to 33% and 10% on average.
Yifeng Geng, Shimin Chen, Yongwei Wu 0001, Ryan Wu, Guangwen Yang 0002
ICPP5
2011 Optimizing write operation on replica in data grid
Pengzhi Xu, Yongwei Wu 0001, Xiaomeng Huang, Guangwen Yang 0002
Sci. China Inf. Sci.4
2011 Optimization of sub-query processing in distributed data integration systems
Gang Chen 0003, Yongwei Wu 0001, Jia Liu 0023, Guangwen Yang 0002
J. Netw. Comput. Appl.4
2011 Automatically constructing trusted cluster computing environment
Yongwei Wu 0001, Gang Chen 0003, Jia Liu 0023, Xiaomeng Huang, Guangwen Yang 0002
J. Supercomput.6
2010 PV-EASY: a strict fairness guaranteed and prediction enabled scheduler in parallel job scheduling
abstract
As the most widely used parallel job scheduling strategy in production schedulers, EASY has achieved great success, not only because it can balance fairness and performance, but also because it is universally applicable to most HPC systems. However, unfairness still exists in EASY. For real workloads used in this work, our simulation shows that a blocked job can be delayed by later jobs for more than 90 hours. In addition, EASY cannot directly employ parallel job runtime prediction techniques, because this would lead to a serious situation called reservation violation.
Yulai Yuan, Guangwen Yang 0002, Yongwei Wu 0001
HPDC2
2010 DABGPM: A Double Auction Bayesian Game-Based Pricing Model in Cloud Market
Shifeng Shang, Jinlei Jiang, Yongwei Wu 0001, Zhenchun Huang, Guangwen Yang 0002
NPC5
2010 Improving grid performance by dynamically deploying applications
abstract
Abstract Grid applications are normally deployed on computing nodes beforehand, which may cause the undesirable situation that some of these nodes (with hot applications deployed) are always busy whereas the others are consistently idle. Therefore, the overall performance (e.g. throughput and load balancing) of such a Grid system would be seriously degraded. In this paper, we present the idea of Hierarchical and Dynamic Deployment of Application (HDDA) in Grid to improve the system performance. With HDDA, an application can be dynamically deployed and undeployed when necessary. In order to reduce the overhead caused by HDDA, the Average Latency Ratio Minimum (ALR‐MIN) replacement strategy is also proposed. It deploys applications to nodes with minimum ALR of Node (NALR), and evicts applications with minimum increment of ALR. The results of the experiment we conducted on ChinaGrid show that HDDA can achieve 10 and 24% less average complete time (ACT) than the schemes of non‐HDDA and Static Deployment of Application (SDA), respectively. Additionally, throughput and load balancing of HDDA are also better than the other two schemas. Results of the simulation performed on a simulator particularly developed for this research show that our ALR‐MIN replacement strategy results in 17% less relative delay‐time of jobs than the well‐known Least Recently Used (LRU)‐based strategies in a typical setting. Copyright © 2010 John Wiley & Sons, Ltd.
Yongwei Wu 0001, Gang Chen 0003, Jia Liu 0023, Guangwen Yang 0002
Concurr. Comput. Pract. Exp.5
2010 VDB-MR: MapReduce-based distributed data integration using virtual database
Yulai Yuan, Yongwei Wu 0001, Guangwen Yang 0002
Future Gener. Comput. Syst.5
2010 Distributed bandwidth allocation based on alternating evolution algorithm
Xiaomeng Huang, Yongwei Wu 0001, Guangwen Yang 0002, Jinlei Jiang
J. Parallel Distributed Comput.3
2010 An adaptive task-level fault-tolerant approach to Grid
Yongwei Wu 0001, Yulai Yuan, Guangwen Yang 0002
J. Supercomput.3
2009 Integrating Cloud-Computing-Specific Model into Aircraft Design
Zhimin Tian, Guangwen Yang 0002
CloudCom3
2008 Adaptive Hybrid Model for Long Term Load Prediction in Computational Grid
abstract
Long term load prediction can assist task scheduling and load balancing greatly in distributed environment such as computational grid. Due to the dynamic property of grid environment, fixed-parameter prediction model can not exert its forecast capability completely. In this paper we first observe and analyze parameters' impact on prediction accuracy for our previous long term load prediction hybrid model (HModel) in detail. And then, a parameter-level adaptive method based on previous analysis is proposed in order to make HModel adapt to the time-varying characteristics of load in computational grid. The results of the experiments demonstrate that our adaptive hybrid model (AHModel) outperforms the widely used autoregressive (AR) model in long term load prediction significantly, and it also achieves obvious reduction in prediction mean square error comparing with HModel which uses fixed parameter value.
Yulai Yuan, Yongwei Wu 0001, Guangwen Yang 0002
CCGRID3
2008 End-to-End Congestion Control for High Speed Networks Based on Population Ecology Models
abstract
Since TCP congestion control is ill-suited for high speed networks, designing a replacement for TCP has become a challenge. To address this problem, we extend the population ecology theory to design a novel congestion control algorithm. We treat the network flows as the species in nature, the throughput of the flows as the population number, and the bottleneck bandwidth as the food resources. Then we use the key idea of constructing population ecology models to develop a novel congestion control model, and implement the corresponding end-to-end transport protocol through measurement, which called Population Ecology TCP (PE-TCP). The theoretical analysis and simulation results validate that PE-TCP achieves high utilization, fast convergence, fair bandwidth allocation, and near-zero packet drops. These qualities are desirable for high speed networks.
Xiaomeng Huang, Fengyuan Ren, Guangwen Yang 0002, Yongwei Wu 0001, W. Zhen, Chuang Lin 0002
ICDCS3
2007 Improving the Convergence and Stability of Congestion Control Algorithm
abstract
The traditional TCP congestion control is inefficient for high speed networks and it is a challenge to design a high speed replacement for TCP. By simulating some existing high speed protocols, we find that these high speed protocols have limitations in convergence and stability. To address these problems, we apply a population ecology model to design a novel congestion control algorithm-Coupling Logistic TCP(CLTCP). It is based on bandwidth pre-assignment that is similar to XCP and MaxNet. The pre-assignment rate factor is computed in the routers based on the information of the router capacity, the aggregate incoming traffic and the queue length. Then the senders adjust the sending rate according to the pre-assignment rate factor which carries by the packet to strengthen the convergence and stability of transport protocol. The theoretical analysis and simulation results show that CLTCP provides not only fast convergence and strong stability, but also high utilization and fair bandwidth allocation regardless of round trip time.
Xiaomeng Huang, Chuang Lin 0002, Fengyuan Ren, Guangwen Yang 0002, Peter D. Ungsunan, Yuanzhuo Wang
ICNP4
2006 A Resource-Autonomy Based Monitoring Architecture for Grids
Meizhi Hu, Guangwen Yang 0002
GPC2
2004 Efficiently Rationing Resources for Grid and P2P Computing
Ming Chen 0004, Yongwei Wu 0001, Guangwen Yang 0002, Xuezheng Liu
NPC3
2004 Paramecium: Assembling Raw Nodes into Composite Cells
Ming Chen 0004, Guangwen Yang 0002, Yongwei Wu 0001, Xuezheng Liu
NPC2
2004 Lookup-Ring: Building Efficient Lookups for High Dynamic Peer-to-Peer Overlays
Xuezheng Liu, Guangwen Yang 0002, Jinfeng Hu, Ming Chen 0004, Yongwei Wu 0001
NPC2
2004 Grid Computing in China
Guangwen Yang 0002, Hai Jin 0001, Minglu Li 0001, Wei Li 0008, Zhaohui Wu 0001, Yongwei Wu 0001, Feilong Tang 0001
J. Grid Comput.1
2003 DSI: Distributed Service Integration for Service Grid
Guangwen Yang 0002, Shuming Shi 0003, Dingxing Wang, Qifeng Huang, Xuezheng Liu
J. Comput. Sci. Technol.1