EDBT 2026 Demo / reviewers in the wild / expert
Hong An
dblp:73/5552
· DBLP profile ↗
68ranked-venue papers
1as first author
42since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 53 · 33 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAW3D: A Performance-Portable Solver for Compressible Turbulence at the 풪(1010)-Cell Scale on Heterogeneous Supercomputers
Siyue Chen, Junshi Chen 0003, Yunchi Deng, Hong An, Juchun Ding |
APPT | 6 |
| 2026 | CISim: ISA-Agnostic Custom Instruction Simulation for General-Purpose ProcessorabstractPre-RTL ISA-agnostic simulators have been established for designing heterogeneous systems, but few of them are suitable for evaluating a general-purpose processor (GPP) with custom instructions (CIs). MosaicSim [1], a state-of-the-art ISA-agnostic simulator, still has several limitations for CI design and simulation. First, it shows inaccuracy in simulating GPPs due to an oversimplified performance model. Second, as designed for kernel simulation, it lacks support for running complex real-world benchmarks. Third, it cannot evaluate fine-grained irregular CIs due to the lack of the ability to represent or define them in benchmarks. To this end, we propose CISim, a new ISA-agnostic simulation framework containing an offloader that generates and integrates CIs into benchmarks, along with a simulator capable of executing benchmarks with CIs. Evaluations show that CISim is accurate by validating against Gem5 [2] and achieves higher accuracy than MosaicSim. A case study evaluating CI exploration methods highlights the strength and flexibility of CISim. Jun Shi 0007, Junshi Chen 0003, Hong An |
DATE | 6 |
| 2026 | AtSpMV: Model-Guided Adaptive Tiling and Load Balancing for SpMV on GPUs
Junshi Chen 0003, Longsheng Song, Jun Shi 0007, Hong An |
Euro-Par (1) | 8 |
| 2026 | RVDeformer: Sparse Point Cloud-Guided Right Ventricle 3-D Reconstruction in Echocardiogramsabstract3D reconstruction of the Right Ventricle (RV) from echocardiograms is crucial for accurate clinical evaluation of cardiac function. However, existing methods are hindered by the complex RV anatomy and the incomplete spatial information inherent in 2D multi-view echocardiograms. Therefore, we propose RVDeformer, a sparse point cloud-guided framework for RV 3D reconstruction. RVDeformer reformulates the reconstruction task as a mesh deformation problem, learning to deform a predefined template mesh to match the target structure under the guidance of the sparse anatomical point cloud. Specifically, this framework employs the end-to-end neural network RVDeformNet to extract the features of the point cloud and template mesh for predicting the displacement of each mesh vertex. We design a point cloud-mesh fusion module that can effectively align and fuse features from the two modalities to enhance the representation ability of the model. We conduct extensive validation on a clinical dataset of 1,278 cases and demonstrate that RVDeformer outperforms existing state-of-the-art methods, achieving a Chamfer Distance (CD) of $2.24\pm 0.55$ mm, an F1-score of $0.74\pm 0.10$ at the 3 mm threshold, and a Volumetric Similarity (VS) of $91.53\pm 2.28$ %, with significant potential for clinical applications. The code is available at https://github.com/onezh95/RVDeformer. Zhaohui Wang 0002, Jun Shi 0007, Minfan Zhao, Hong An |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program TuningabstractTensor program tuning is essential for the efficient deployment of deep neural networks. Search-based approaches have demonstrated scalability and effectiveness in automatically finding high-performance programs for specific hardware. However, the search process is often inefficient, taking hours or even days to discover optimal programs due to the exploration mechanisms guided by an accurate but slow-learned cost model. Meanwhile, the learned cost model trained on one platform cannot seamlessly adapt online to another, which we call cross-platform online unawareness. In this work, we propose Pruner and MoA-Pruner. Pruner is a ''Draft-then-Verify'' exploration mechanism that accelerates the schedule search process. Instead of applying the complex learned cost model to all explored candidates, Pruner drafts small-scale potential candidates by introducing a naive Symbol-based Analyzer (draft model), then identifies the best candidates by the learned cost model. MoA-Pruner introduces a Momentum online Adaptation strategy to address the cross-platform online unawareness. Jun Shi 0007, Minfan Zhao, Junshi Chen 0003, Hong An, Xulong Tang, Honghui Yuan |
ASPLOS (2) | 9 |
| 2025 | NDFT: Accelerating Density Functional Theory Calculations via Hardware/Software Co-Design on Near-Data Computing SystemabstractLinear-response time-dependent Density Functional Theory (LR-TDDFT) is a widely used method for accurately predicting the excited-state properties of physical systems. Previous works have attempted to accelerate LR-TDDFT using heterogeneous systems such as GPUs, FPGAs, and the Sunway architecture. However, a major drawback of these approaches is the constant data movement between host memory and the memory of the heterogeneous systems, which results in substantial data movement overhead. Moreover, these works focus primarily on optimizing the compute-intensive portions of LR-TDDFT, despite the fact that the calculation steps are fundamentally memory-bound.To address these challenges, we propose NDFT, a Near-Data Density Functional Theory framework. Specifically, we design a novel task partitioning and scheduling mechanism to offload each part of LR-TDDFT to the most suitable computing units within a CPU-NDP system. Additionally, we implement a hardware/software co-optimization of a critical kernel in LR-TDDFT to further enhance performance on the CPU-NDP system. Our results show that NDFT achieves performance improvements of 5.2x and 2.5x over CPU and GPU baselines, respectively, on a large physical system. Qingcai Jiang, Buxin Tu, Junshi Chen 0003, Hong An |
DAC | 5 |
| 2025 | NDPage: Efficient Address Translation for Near-Data Processing Architectures via Tailored Page TableabstractNear-Data Processing (NDP) has been a promising architectural paradigm to address the memory wall problem for data-intensive applications. Practical implementation of NDP architectures calls for system support for better programmability, where having virtual memory (VM) is critical. Modern computing systems incorporate a 4-level page table design to support address translation in VM. However, simply adopting an existing 4-level page table in NDP systems causes significant address translation overhead because (1) NDP applications generate a lot of address translations, and (2) the limited L1 cache in NDP systems cannot cover the accesses to page table entries (PTEs). We extensively analyze the 4-level page table design in the NDP scenario and observe that (1) the memory access to page table entries is highly irregular, thus cannot benefit from the L1 cache, and (2) the last two levels of page tables are nearly fully occupied. Based on our observations, we propose NDPage, an efficient page table design tailored for NDP systems. The key mechanisms of NDPage are (1) an L1 cache bypass mechanism for PTEs that not only accelerates the memory accesses of PTEs but also prevents the pollution of PTEs in the cache system, and (2) a flattened page table design that merges the last two levels of page tables, allowing the page table to enjoy the flexibility of a 4KB page while reducing the number of PTE accesses. We evaluate NDPage using a variety of data-intensive work-loads. Our evaluation shows that in a single-core NDP system, NDPage improves the end-to-end performance over the state-of-the-art address translation mechanism of 14.3 %; in 4-core and 8-core NDP systems, NDPage enhances the performance of 9.8% and 30.5 %, respectively. Qingcai Jiang, Buxin Tu, Hong An |
DATE | 3 |
| 2025 | Uniform Dense Blocking for Efficient Sparse LU Factorization in First-Principles Materials Simulation
Junshi Chen 0003, Longsheng Song, Haijie Hou, Dongdong Tan, Yueqiang He, Wentiao Wu, Sihan Lu, Hong An |
Euro-Par (3) | 9 |
| 2025 | HManyCore-Sim: Heterogeneous Simulator for Deeply Fused Many-Core Processor
Liyi Wang, Hong An, Xiahui Hu, Yiyun Yin, Qinglin Liu |
ICA3PP (6) | 2 |
| 2025 | Carver: Learning to Reconstruct Right Ventricle from Sparse Multi-View 2D EchocardiogramsabstractAccurate 3D reconstruction of the right ventricle from multi-view echocardiograms is crucial for the quantitative diagnosis of cardiac diseases. However, existing methods often fail to deliver satisfactory results due to the structural complexity of the right ventricle and the sparsity of non-parallel ultrasound views. In this paper, we propose an efficient reconstruction method named Carver, which redefines 3D reconstruction of the right ventricle as a voxel-wise dense prediction task for the first time. The core idea lies in using a deep neural network to learn the end-to-end deformation from a coarse geometric convex hull to a fine structure of the right ventricle, similar to carving. To improve the reconstruction accuracy and robustness, we design a dual-aware network incorporating prior contour information to enhance learning representation. We conduct extensive experiments on an echocardiography dataset containing 1,278 instances to validate the effectiveness of the proposed method. Experimental results demonstrate that Carver outperforms existing state-of-the-art methods, achieving a Volume Similarity (VS) of 98.75%, a Dice Similarity Coefficient (DSC) of 97.80%, a Hausdorff Distance (HD) of 4.96, and a Root Mean Square Error (RMSE) of 0.013 for the ejection fraction, while maintaining considerable robustness even with sparser inputs. The code is available at https://github.com/ustclyd/Carver. Jun Shi 0007, Zhaohui Wang 0002, Tiantong Wang, Minfan Zhao, Junshi Chen 0003, Hong An |
ICASSP | 8 |
| 2025 | PromptSeg: Learning to Segment Medical Image via Visual PromptsabstractDeep learning has made remarkable medical image segmentation advancements, yet its generalization capability across tasks remains challenging. The variety of task objectives, disease-dependent labeling variations, and multi-center data contribute to the poor generalization capacity of task-specific models on unseen tasks, necessitating domain adaptation or fine-tuning. This typically involves data annotation and network retraining, limiting the application of deep learning in clinical practice. In this study, we propose PromptSeg, an innovative Transformer-based unified segmentation framework for general medical image segmentation tasks. PromptSeg aims to utilize the provided visual prompts to recognize task patterns and learn contextual representations, thereby breaking the restrictions of the task-specific paradigm. When faced with unseen datasets or segmentation targets during inference, our method only requires a few annotated prompt pairs to understand the task and segment the query images without retraining, alleviating the need for large-scale annotation data. The experimental results demonstrate that our method outperforms existing state-of-the-art methods and exhibits high generalization capability on multiple unseen tasks. The source code is available at https://github.com/MinfanZhao/PromptSeg. Minfan Zhao, Jun Shi 0007, Zhaohui Wang 0002, Junshi Chen 0003, Hong An |
ICASSP | 6 |
| 2025 | CIExplorer: Microarchitecture-Aware Exploration for Tightly Integrated Custom InstructionabstractExtending existing architectures with customized instruction extensions is emerging to achieve high performance and energy efficiency for specific applications.Automated discovery of custom instructions (CIs) is well-studied nowadays, which requires exploring combinations of different types and quantities of operations, resulting in a vast search space.However, previous works typically use microarchitectureagnostic cost models, leading to suboptimal CIs that may degrade performance.They leverage graph isomorphism to reduce area overhead, but few of them consider its potential to benefit performance-oriented exploration.To this end, we present CIExplorer, a framework for adaptive CI exploration. Qingcai Jiang, Jun Shi 0007, Junshi Chen 0003, Hong An, Xulong Tang, Honghui Yuan |
ICS | 7 |
| 2025 | AMALI: An Analytical Model for Accurately Modeling LLM Inference on Modern GPUsabstractLarge language model (LLM) inference applications are surging in recent years, which largely relies on modern GPUs.On the other hand, GPU analytical model is a commonly used tool for architects to precisely identify bottlenecks quickly with deep insights.However, existing GPU analytical models fall short of accurately modeling LLM inference applications on modern GPUs, because of unsuitable tensor core modeling, ignoring constant cache as well as instruction cache modeling and abstracting away important details for LLM inference applications.To address this problem, we propose a novel analytical model dubbed AMALI to accurately model LLM inference on modern GPUs with three innovations.First, we develop an instruction modifier and throughput based tensor core model by accurately capturing the math pipe throttle stalls to enhance the architecture modeling for modern GPUs.Second, we propose analytical models for constant cache and instruction cache by developing micro-benchmarks to measure CUDA kernel launching latencies.This significantly improves AMALI's accuracy compared to real GPU hardware.Finally, we design a multi-warp model by leveraging warp instruction number distribution to reflect LLM inference application characteristics.We validate AMALI on an A100 GPU by using typical LLM inference applications.The results show that AMALI reduces the MAPE (mean absolute percentage error) from 127.56% to 23.59% Shiheng Cao, Junmin Wu, Junshi Chen 0003, Hong An, Zhibin Yu 0001 |
ISCA | 4 |
| 2025 | Million-Atom Ab Initio Electron Dynamics: Discontinuous Galerkin Real-Time Time-Dependent Density Functional TheoryabstractOver the past decades, first-principles real-time time dependent density functional theory(rt-TDDFT) simulations have been limited to systems with only thousands of atoms. We propose a novel method based on the discontinuous Galerkin adaptive local basis, significantly reducing global communication in rt-TDDFT. We further introduce a tensor compression technique that leverages basis locality to avoid repeated evaluation of multi-center integrals in hybrid functionals, greatly reducing computational cost. To overcome the projection bottleneck in our basis sets, we design a fused Gemm-Reduce operation that achieves several times higher floating-point efficiency than standard BLAS combination. Our implementation reaches 34.8% of theoretical peak performance on 524,288 CGs of the New Sunway supercomputer and simulates electronic dynamics of systems with over one million atoms for both local-semi-local and hybrid functionals. This work improves computational scale by two orders of magnitude, opening new possibilities for exploring ultrafast dynamics in large-scale materials and nanophotonic devices. Junwei Feng, Junshi Chen 0003, Xinming Qin, Lingyun Wan, Wentiao Wu, Bingkun Hou, Yexuan Lin, Zechuan Zhang, Weile Jia, Hong An, Jinlong Yang 0003, Wei Hu 0006 |
SC | 15 |
| 2025 | Matrix Is All You Need: Rearchitecting Quantum Chemistry to Scale on AI AcceleratorsabstractScientific computing remains fundamentally misaligned with the execution paradigm of modern AI accelerators, which rely on structured, low-precision matrix operations for performance and scalability. Quantum chemistry exemplifies this gap through three core scalability limits: irregular computational patterns, fragmented hardware utilization, and limited scientific reach. Haozhi Han, Kun Li 0016, Fusong Ju, Qi Li 0039, Hong An, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004 |
SC | 5 |
| 2025 | SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity TransformationabstractSparse Tensor Cores offer exceptional performance gains for AI workloads by exploiting structured 2:4 sparsity. However, their potential remains untapped for core scientific workloads such as stencil computations, which exhibit irregular sparsity patterns. Qi Li 0039, Kun Li 0016, Haozhi Han, Yunquan Zhang, Junshi Chen 0003, Hong An, Ting Cao 0003, Mao Yang 0004 |
SC | 8 |
| 2025 | swPredictor: A data-driven performance model for distributed data parallelism training on large-scale HPC clusters
Xianyu Zhu, Ruohan Wu, Junshi Chen 0003, Hong An |
Perform. Evaluation | 4 |
| 2025 | PWDFT-SW: Extending the Limit of Plane-Wave DFT Calculations to 16K Atoms on the New Sunway SupercomputerabstractFirst-principles density functional theory (DFT) with plane wave (PW) basis set is the most widely used method in quantum mechanical material simulations due to its advantages in accuracy and universality. However, a perceived drawback of PW-based DFT calculations is their substantial computational cost and memory usage, which currently limits their ability to simulate large-scale complex systems containing thousands of atoms. This situation is exacerbated in the new Sunway supercomputer, where each process is limited to a mere 16 GB of memory. Herein, we present a novel parallel implementation of plane wave density functional theory on the new Sunway supercomputer (PWDFT-SW). PWDFT-SW fully extracts the benefits of Sunway supercomputer by extensively refactoring and calibrating our algorithms to align with the system characteristics of the Sunway system. Through extensive numerical experiments, we demonstrate that our methods can substantially decrease both computational costs and memory usage. Our optimizations translate to a speedup of 64.8x for a physical system containing 4,096 silicon atoms, enabling us to push the limit of PW-based DFT calculations to large-scale systems containing 16,384 carbon atoms. Qingcai Jiang, Zhenwei Cao, Junshi Chen 0003, Xinming Qin, Wei Hu 0006, Hong An, Jinlong Yang 0003 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2024 | A3PIM: An Automated, Analytic and Accurate Processing-in-Memory OffloaderabstractThe performance gap between memory and processor has grown rapidly. Consequently, the energy and wall-clock time costs associated with moving data between the CPU and main memory predominate the overall computational cost. The Processing-in-Memory (PIM) paradigm emerges as a promising architecture that mitigates the need for extensive data movements by strategically positioning computing units proximate to the memory. Despite the abundant efforts devoted to building a robust and highly-available PIM system, identifying PIM-friendly segments of applications poses significant challenges due to the lack of a comprehensive tool to evaluate the intrinsic memory access pattern of the segment. To tackle this challenge, we propose A3PIM11The code and benchmarks of this work are opened-sourced in: https://github.com/ACSA-PIM/A3PIM: an Automated, Analytic and Accurate Processing-in-Memory offloader. We sys-tematically consider the cross-segment data movement and the intrinsic memory access pattern of each code segment via static code analyzer. We evaluate A 3PIM across a wide range of real-world workloads including GAP and PrIM benchmarks and achieve an average speedup of 2.63x and 4.45x (up to 7.14x and lO.64x) when compared to CPU-only and PIM-only executions, respectively. Qingcai Jiang, Shaojie Tan, Junshi Chen 0003, Hong An |
DATE | 4 |
| 2024 | swYAKL: A Data Parallel Runtime on Manycore ArchitectureabstractThe traditional architecture of supercomputers comprises a control unit and heterogeneous accelerators with distinct memory spaces, necessitating developers to code separately for each accelerator. The notion of performance portability has been introduced to facilitate the swift adaptation of a singular codebase across multiple heterogeneous accelerators, enabling the same code to efficiently harness the performance of various platforms with minimal alterations. Frameworks such as Kokkos, RAJA, and YAKL accomplish this through a data parallel model, primarily targeting SIMT devices but also applicable to SIMD devices. With the advent of manycore architectures that provide high levels of parallelism and performance, it becomes imperative to extend performance portability to these architectures, which currently require users to manually partition and map tasks to fully exploit their capabilities. This paper introduces a data parallel programming runtime for the Sunway manycore architecture and adapts YAKL to this platform, termed swYAKL. This runtime capitalizes on the computational characteristics of Sunway, including task division and mapping, and supports Sunway’s CPE LDM and native vectorization. It employs a specialized task division method for multi-dimensional stencil kernels to leverage the DMA capabilities of Sunway’s CPE. Performance evaluations using several common computational kernels and a proxy application named miniWeather indicate that swYAKL achieves speedups exceeding 100x compared to execution on Sunway’s MPE for various workloads, and under specific conditions, it surpasses the performance of the NVIDIA A100 GPU. Yanwei Ye, Junshi Chen 0003, Hong Qian, Kunxian Lin, Yuanhang Li, Hong An |
HPCC | 6 |
| 2024 | Multi-level Load Balancing Strategies for Massively Parallel Smoothed Particle Hydrodynamics SimulationabstractIn the field of computational fluid dynamics, Smoothed Particle Hydrodynamics (SPH) serves as a powerful tool for investigating complex fluid interactions and instabilities. For the practical SPH simulation of large-scale fluid phenomena such as tsunamis, volcanic eruptions, and planetary collisions, it typically requires billions of particles, as the numerical resolution increases proportionally with the number of particles. To efficiently conduct large-scale SPH simulations on modern supercomputers with massive many-core processors, we propose a novel SPH implementation leveraging multi-level parallelism and a corresponding three-level load balancing strategy. Our load balancing approach comprises: (1) a process-level domain decomposition algorithm based on an improved 1D partitioning exact algorithm; (2) an adaptive recursive cell subdivision method; (3) a fine-grained dynamic thread-level task scheduling strategy. Our experiment uses 1 billion particles to simulate converging Richtmyer–Meshkov instability and verifies the effect of load balancing on new Sunway supercomputer. As the shockwave converges on the central interface area, our load balancing strategy breaks the bottleneck constraints on the slowest node, increases the balance of computational loads between nodes from 30.01% to 91.48%, and achieves a 2.8 × improvement in computational performance. Finally, our implementation enables each CPU to handle 10 million particles and scale from 1 CPU to 100,000 CPUs (in total 39 million cores with 1 trillion particles) with a performance of 80.4% parallel efficiency. Ziyu Zhang 0003, Yang Zhao 0040, Junshi Chen 0003, Hong An, Zhanming Wang, Longkui Chen |
ICPP | 5 |
| 2024 | DB-SpGEMM: A Massively Distributed Block-Sparse Matrix-Matrix Multiplication for Linear-Scaling DFT CalculationsabstractLinear-scaling <?TeX $\mathcal {O}(N)$?> Math 1 density functional theory (DFT) represents a significant advancement in the field of computational materials science, especially for simulations of large systems where traditional cubic-scaling methods become computationally prohibitive. The core operation in <?TeX $\mathcal {O}(N)$?> Math 2 methods is sparse general matrix-matrix multiplication (SpGEMM), which is the major performance bottleneck. To enhance the computational efficiency of SpGEMM, it is crucial to consider the inherent sparse pattern of these matrices. Targeting block-sparse matrices with moderate block sizes and regular block shapes, we have developed a distributed block-sparse matrix-matrix multiplication (DB-SpGEMM) algorithm for large-scale DFT calculations. Through deep optimizations in distributed matrix storage, computational task decomposition, asynchronous task scheduling, and load balancing, we have implemented a linear-scaling method based on this algorithm within the discontinuous Galerkin density functional theory (DGDFT). On the new Sunway supercomputer, our approach achieves a 8 ∼ 10x speedup compared to the original version on monolayer phosphorene systems, and demonstrates superior scalability. Junshi Chen 0003, Yang Zhao 0040, Longsheng Song, Xinming Qin, Hong An |
ICPP | 6 |
| 2024 | Predictive Accuracy-Based Active Learning for Medical Image Segmentation
Jun Shi 0007, Shulan Ruan, Minfan Zhao, Hong An, Xudong Xue |
IJCAI | 5 |
| 2024 | Pushing the Limit of Quantum Mechanical Simulation to the Raman Spectra of a Biological System with 100 Million AtomsabstractRaman spectroscopy offers invaluable insights into the chemical composition and structural characteristics of various materials, making it a powerful tool for structural analysis. However, accurate quantum mechanical simulations of Raman spectra for large systems, such as biological materials, have been limited due to immense computational costs and technical challenges. In this study, we developed efficient algorithms and optimized implementations on heterogeneous computing architectures to enable fast and highly scalable ab initio simulations of Raman spectra for large-scale biological systems with up to 100 million atoms. Our simulations have achieved nearly linear strong and weak scaling on two cutting-edge high-performance computing systems, with peak FP64 performances reaching 400 PFLOPS on 96,000 nodes of new Sunway supercomputer and 85 PFLOPS on 6,000 node of ORISE supercomputer. These advances provide promising prospects for extending quantum mechanical simulations to biological systems. Honghui Shang, Ying Liu 0055, Zhikun Wu, Zhenchuan Chen, Jinfeng Liu 0004, Meiyue Shao, Yingzhou Li, Bowen Kan, Huimin Cui, Xiaobing Feng 0002, Yunquan Zhang, Donald G. Truhlar, Hong An, Xiao He 0004, Jinlong Yang 0003 |
SC | 13 |
| 2024 | Enabling 13K-Atom Excited-State GW Calculations via Low-Rank Approximations and HPC on the New Sunway SupercomputerabstractGW approximation is a powerful approach to accurately describe the excited-state of semiconductors. However, GW incurs high computational cost $\mathcal{O}\left(N^{4}\right)$ and large memory usage $\mathcal{O}\left(N^{3}\right)$, limiting its applications to thousands of (2,742) atoms even on leadership supercomputers. Herein we present a massively parallel implementation of accurate and efficient cubic-scaling plane-wave GW calculations by using low-rank approximations and high-performance computing on leadership supercomputers. By using a series of low rank approximations, we can reduce the expensive GW calculations to the cubic-scaling computational cost $\mathcal{O}\left(N^{3}\right)$ and quadratic memory usage $\mathcal{O}\left(N^{2}\right)$. With the help of parallel and communication optimization, the plane-wave GW calculations gain an overall speedup of over 70x and efficiently scale up to 13,824 atoms within a few minutes using 449,280 cores on new Sunway supercomputer. This accomplishment paves the way for excited-state quantum mechanical material simulations at mesoscopic scale (10K atoms) and for the design of next-generation semiconductor devices. Wentiao Wu, Zhengbang Zhou, Qingcai Jiang, Junwei Feng, Xinming Qin, Huanhuan Ma, Zhenwei Cao, Junshi Chen 0003, Xinyong Meng, Bingkun Hou, Yuanfan Xiong, Linhao Wang, Yixuan Sun, Hong An, Jinlong Yang 0003, Wei Hu 0006 |
SC | 15 |
| 2024 | Uncovering the performance bottleneck of modern HPC processor with static code analyzer: a case study on Kunpeng 920
Shaojie Tan, Qingcai Jiang, Zhenwei Cao, Junshi Chen 0003, Hong An |
CCF Trans. High Perform. Comput. | 6 |
| 2024 | Extending the limit of LR-TDDFT on two different approaches: Numerical algorithms and new Sunway heterogeneous supercomputerabstractFirst-principles time-dependent density functional theory (TDDFT) is a powerful tool to accurately describe the excited-state properties of molecules and solids in condensed matter physics , computational chemistry, and materials science. However, a perceived drawback in TDDFT calculations is its ultrahigh computational cost O ( N 5 ∼ N 6 ) and large memory usage O ( N 4 ) especially for plane-wave basis set, confining its applications to large systems containing thousands of atoms. Here, we present a massively parallel implementation of linear-response TDDFT (LR-TDDFT) and accelerate LR-TDDFT in two different aspects: (1) numerical algorithms on the X86 supercomputer and (2) optimizations on the heterogeneous architecture of the new Sunway supercomputer. Furthermore, we carefully design the parallel data and task distribution schemes to accommodate the physical nature of different computation steps. By utilizing these two different methods, our implementation can gain an overall speedup of 10x and 80x and efficiently scales to large systems up to 4096 and 2744 atoms within dozens of seconds. Qingcai Jiang, Zhenwei Cao, Xinhui Cui, Lingyun Wan, Xinming Qin, Huanqi Cao, Hong An, Junshi Chen 0003, Jie Liu 0069, Wei Hu 0006, Jinlong Yang 0003 |
Parallel Comput. | 7 |
| 2024 | Gene expression bias between the subgenomes of allopolyploid hybrids is an emergent property of the kinetics of expressionabstractHybridization coupled to polyploidy, or allopolyploidy, has dramatically shaped the evolution of flowering plants, teleost fishes, and other lineages. Studies of recently formed allopolyploid plants have shown that the two subgenomes that merged to form that new allopolyploid do not generally express their genes equally. Instead, one of the two subgenomes expresses its paralogs more highly on average. Meanwhile, older allopolyploidy events tend to show biases in duplicate losses, with one of the two subgenomes retaining more genes than the other. Since reduced expression is a pathway to duplicate loss, understanding the origins of expression biases may help explain the origins of biased losses. Because we expect gene expression levels to experience stabilizing selection, our conceptual frameworks for how allopolyploid organisms form tend to assume that the new allopolyploid will show balanced expression between its subgenomes. It is then necessary to invoke phenomena such as differences in the suppression of repetitive elements to explain the observed expression imbalances. Here we show that, even for phenotypically identical diploid progenitors, the inherent kinetics of gene expression give rise to biases between the expression levels of the progenitor genes in the hybrid. Some of these biases are expected to be gene-specific and not give rise to global differences in progenitor gene expression. However, particularly in the case of allopolyploids formed from progenitors with different genome sizes, global expression biases favoring one subgenome are expected immediately on formation. Hence, expression biases are arguably the expectation upon allopolyploid formation rather than a phenomenon needing explanation. In the future, a deeper understanding of the kinetics of allopolyploidy may allow us to better understand both biases in duplicate losses and hybrid vigor. Hong An, J. Chris Pires, Gavin C. Conant |
PLoS Comput. Biol. | 1 |
| 2024 | SWattention: designing fast and memory-efficient attention for a new Sunway SupercomputerabstractAbstract In the past few years, Transformer-based large language models (LLM) have become the dominant technology in a series of applications. To scale up the sequence length of the Transformer, FlashAttention is proposed to compute exact attention with reduced memory requirements and faster execution. However, implementing the FlashAttention algorithm on the new generation Sunway Supercomputer faces many constraints such as the unique heterogeneous architecture and the limited memory bandwidth. This work proposes SWattention, a highly efficient method for computing the exact attention on the SW26010pro processor. To fully utilize the 6 core groups (CG) and 64 cores per CG on the processor, we design a two-level parallel task partition strategy. Asynchronous memory access is employed to ensure that memory access overlaps with computation. Additionally, a tiling strategy is introduced to determine optimal SRAM block sizes. Compared with the standard attention, SWattention achieves around 2.0x speedup for FP32 training and 2.5x speedup for mixed-precision training. The sequence lengths range from 1k to 8k and scale up to 16k without being out of memory. As for the end-to-end performance, SWattention achieves up to 1.26x speedup for training GPT-style models, which demonstrates that SWattention enables longer sequence length for LLM training. Ruohan Wu, Xianyu Zhu, Junshi Chen 0003, Tianyu Zheng, Xin Liu 0081, Hong An |
J. Supercomput. | 7 |
| 2023 | SWSPH: A Massively Parallel SPH Implementation for Hundred-Billion-Particle Simulation on New Sunway Supercomputer
Ziyu Zhang 0003, Junshi Chen 0003, Zhanming Wang, Jineng Yao, Shenghong Huang, Hong An |
Euro-Par | 7 |
| 2023 | H-DenseFormer: An Efficient Hybrid Densely Connected Transformer for Multimodal Tumor Segmentation
Jun Shi 0007, Hongyu Kan, Shulan Ruan, Minfan Zhao, Zhaohui Wang 0002, Hong An, Xudong Xue |
MICCAI (4) | 8 |
| 2023 | Establishing a Modeling System in 3-km Horizontal Resolution for Global Atmospheric Circulation triggered by Submarine Volcanic Eruptions with 400 Billion Smoothed Particle HydrodynamicsabstractPeople are increasingly concerned about how tectonic processes affect climate and vice versa. We establish a cross-sphere modeling system for volcanic eruptions and atmosphere circulation on a new Sunway supercomputer with a spatial resolution from 10m locally to 3km globally, using an improved multimedium and multiphase smoothed particle hydrodynamics (SPH) combined with a fully coupled meteorology-chemistry global atmospheric modeling scheme. We achieve 400 billion particles and 80% parallel efficiency using 39,000,000 processor cores. The simulation captures the whole dynamic process of the Tonga eruption from shock waves, earthquakes, tsunamis, mushroom clouds to the following 6--7 days of transport and diffusion of ash and water vapor, and preliminarily obtains the influence effect of full coupling of volcano, earthquake, ocean and atmosphere. This work is of great significance for deeply understanding the interaction between tectonic processes and climate change, and establishing an early warning simulation system for similar global hazard events. Shenghong Huang, Junshi Chen 0003, Ziyu Zhang 0003, Hong An, Yan Hu 0004, Zhanming Wang, Longkui Chen, Jineng Yao, Yang Zhao 0040, Dongning Jia, Changming Song, Xisheng Luo, Xiaobin He, Dexun Chen |
SC | 6 |
| 2023 | High performance computing for first-principles Kohn-Sham density functional theory towards exascale supercomputers
Xinming Qin, Junshi Chen 0003, Zhaolong Luo, Lingyun Wan, Jielan Li, Shizhe Jiao, Qingcai Jiang, Wei Hu 0006, Hong An, Jinlong Yang 0003 |
CCF Trans. High Perform. Comput. | 10 |
| 2023 | swMPAS-A: Scaling MPAS-A to 39 Million Heterogeneous Cores on the New Generation Sunway SupercomputerabstractWith the computing power of High-Performance Computing (HPC) systems having stepped into the exascale era, more complex problems can be solved with scientific applications on a large scale. However, due to the significant performance gap between computing nodes and storage subsystems, suboptimal design for the Input/Output (I/O) module will significantly impede the efficiency of scientific applications, especially for the ubiquitous atmosphere applications. Two-phase I/O implemented in N-to-1 mode creates a serious bottleneck that hinders the scalability for the Model for Prediction Across Scales-Atmosphere (MPAS-A) on the new generation Sunway supercomputer. To address the I/O problem, we apply a custom data reorganization method to enable N-to-M I/O mode to exploit the parallel file system's performance and limit the data transfer among MPI ranks to a restricted scope to alleviate communication overhead. Moreover, we have conducted several methods to accelerate the computations, including the redesign for tracer transport, a hybrid buffering scheme, and a three-level parallelization scheme, which allows MPAS-A to use all heterogeneous computing resources efficiently. Experimental results show admirable scalability and efficiency of our I/O method, which achieves speedups of 41× and 58.9× for input and output compared with the raw I/O method on 30,000 MPI ranks. By scaling MPAS-A to 39 million heterogeneous cores, we demonstrate the necessity of a well-constructed I/O module for a real-world atmosphere application. Speed tests show that our optimization methods obtain good results for computations, and MPAS-A achieves a speed of 0.82 Simulated Day per Hour (SDPH) and 0.76 parallel efficiency of strong scaling with 600,000 MPI ranks. Junshi Chen 0003, Jiawang Feng, Hong An |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | Accelerating Parallel First-Principles Excited-State Calculation by Low-Rank Approximation with K-Means ClusteringabstractFirst-principles time-dependent density functional theory (TDDFT) is a powerful tool to accurately describe the excited-state properties of molecules and solids in condensed matter physics, computational chemistry and materials science. However, a perceived drawback in TDDFT calculations is its ultrahigh computational cost and large memory usage especially for plane-wave basis set, confining its applications to large systems containing thousands of atoms. Here, we present a massively parallel implementation of linear-response TDDFT (LR-TDDFT) and reduce the complexity to by combining K-Means clustering based low-rank approximation with iterative eigensolve algorithm. Furthermore, we carefully design the parallel data and task distribution schemes to accommodate with the physical nature in different steps of the computation, also, several optimization methods are employed to effectively handle the matrix operations and data communications of constructing and diagonalizing the LR-TDDFT Hamiltonian. In particular, our method can significantly reduce the cost of computation and memory by nearly 2 orders of magnitude compared to conventional LR-TDDFT calculations. Numerical results demonstrate that our implementation can gain an overall speedup of 10x and efficiently scale up to 12,288 CPU cores for large systems up to 4,096 atoms within dozens of seconds. Qingcai Jiang, Jielan Li, Junshi Chen 0003, Xinming Qin, Lingyun Wan, Jinlong Yang 0003, Jie Liu 0069, Wei Hu 0006, Hong An |
ICPP | 9 |
| 2022 | 2.5 Million-Atom Ab Initio Electronic-Structure Simulation of Complex Metallic Heterostructures with DGDFTabstractOver the past three decades, ab initio electronic structure calculations of large, complex and metallic systems are limited to tens of thousands of atoms in computational accuracy and efficiency on leadership supercomputers. We present a massively parallel discontinuous Galerkin density functional theory (DGDFT) implementation, which adopts adaptive local basis functions to discretize the Kohn-Sham equation, resulting in a block-sparse Hamiltonian matrix. A highly efficient pole expansion and selected inversion (PEXSI) sparse direct solver is implemented in DGDFT to achieve O(N1.5) scaling for quasi two-dimensional systems. DGDFT allows us to compute the electronic structures of complex metallic heterostructures with 2.5 million atoms (17.2 million electrons) using 35.9 million cores on the new Sunway supercomputer. The peak performance of PEXSI can achieve 64 PFLOPS (~5% of theoretical peak), which is un-precedented for sparse direct solvers. This accomplishment paves the way for quantum mechanical simulations into mesoscopic scale for designing next-generation electronic devices. Wei Hu 0006, Hong An, Zhuoqiang Guo, Qingcai Jiang, Xinming Qin, Junshi Chen 0003, Weile Jia, Chao Yang 0001, Zhaolong Luo, Jielan Li, Wentiao Wu, Guangming Tan, Dongning Jia, Qinglin Lu, Yeqi Huang, Liyi Wang, Jinlong Yang 0003 |
SC | 2 |
| 2022 | AI for Quantum Mechanics: High Performance Quantum Many-Body Simulations via Deep LearningabstractSolving quantum many-body problems is one of the most fascinating research fields in condensed matter physics. An efficient numerical method is crucial to understand the mechanism of novel physics, such as the high Tc superconductivity, as one has to find the optimal solution in the exponentially large Hilbert space. The development of Artificial Intelligence (AI) provides a unique opportunity to solve the quantum many-body problems, but there is still a large gap from the goal. In this work, we present a novel computational framework, and adapt it to the Sunway supercomputer. With highly efficient scalability up to 40 million heterogeneous cores, we can drastically increase the number of variational parameters, which greatly improves the accuracy of the solutions. The investigations of the spin-1/2 J1-J2 model and the t-J model achieve unprecedented accuracy and time-to-solution far beyond the previous state of the art. Xuncheng Zhao, Mingfan Li, Junshi Chen 0003, Meijia Zhao, Hong An, Lixin He |
SC | 9 |
| 2022 | Bridging the Gap between Deep Learning and Frustrated Quantum Spin System for Extreme-Scale Simulations on New Generation of Sunway SupercomputerabstractEfficient numerical methods are promising tools for delivering unique insights into the fascinating properties of physics, such as the highly frustrated quantum many-body systems. However, the computational complexity of obtaining the wave functions for accurately describing the quantum states increases exponentially with respect to particle number. Here we present a novel convolutional neural network (CNN) for simulating the two-dimensional highly frustrated spin-$1/2$$J_1-J_2$Heisenberg model, meanwhile the simulation is performed at an extreme scale system with low cost and high scalability. By ingenious employment of transfer learning and CNN’s translational invariance, we successfully investigate the quantum system with the lattice size up to$24\times 24$, within 30 million cores of the new generation of sunway supercomputer. The final achievement demonstrates the effectiveness of CNN-based representation of quantum-state and brings the state-of-the-art record up to a brand-new level from both aspects of remarkable accuracy and unprecedented scales. Mingfan Li, Junshi Chen 0003, Qingcai Jiang, Xuncheng Zhao, Rongfen Lin, Hong An, Lixin He |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2021 | DARNet: Dual-Attention Residual Network for Automatic Diagnosis of COVID-19 via CT ImagesabstractThe ongoing global pandemic of Coronavirus Disease 2019 (COVID-19) poses a serious threat to public health and the economy. Rapid and accurate diagnosis of COVID-19 is essential to prevent the further spread of the disease and reduce its mortality. Chest Computed tomography (CT) is an effective tool for the early diagnosis of lung diseases including pneumonia. However, detecting COVID-19 from CT is demanding and prone to human errors as some early-stage patients may have negative findings on images. Recently, many deep learning methods have achieved impressive performance in this regard. Despite their effectiveness, most of these methods underestimate the rich spatial information preserved in the 3D structure or suffer from the propagation of errors. To address this problem, we propose a Dual-Attention Residual Network (DARNet) to automatically identify COVID-19 from other common pneumonia (CP) and healthy people using 3D chest CT images. Specifically, we design a dual-attention module consisting of channel-wise attention and depth-wise attention mechanisms. The former is utilized to enhance channel independence, while the latter is developed to recalibrate the depth-level features. Then, we integrate them in a unified manner to extract and refine the features at different levels to further improve the diagnostic performance. We evaluate DARNet on a large public CT dataset and obtain superior performance. Besides, the ablation study and visualization analysis prove the effectiveness and interpretability of the proposed method. Jun Shi 0007, Huite Yi, Shulan Ruan, Zhaohui Wang 0002, Hong An |
BIBM | 6 |
| 2021 | Symplectic structure-preserving particle-in-cell whole-volume simulation of tokamak plasmas to 111.3 trillion particles and 25.7 billion gridsabstractWe employ our recently developed explicit 2nd-order charge-conservative symplectic electromagnetic particle-in-cell (PIC) scheme in the cylindrical mesh to simulate the whole-volume magnetic confinement toroidal plasmas on the new Sunway supercomputer. From a large-scale simulation of magneticized toroidal plasma with 111.3 trillion particles and 25.7 billion grids, we have obtained a sustained performance exceeding 201.1 PFLOP/s (double precision) with the fastest iteration step achieving 298.2 PFLOP/s (double precision). For the first time, unprecedented high resolution evolution of 6D electromagnetic fully kinetic plasmas based on 2D equilibrium profiles from Experimental Advanced Superconducting Tokamak (EAST) and designed operation state of China Fusion Engineering Test Reactor (CFETR) are presented, and edge micro-instabilities can be investigated directly. This shows the possibility to study crucial problems and phenomena in the magnetic confinement toroidal plasma directly using the symplectic electromagnetic fully kinetic PIC method on world's leading supercomputers. Jianyuan Xiao, Junshi Chen 0003, Jiangshan Zheng, Hong An, Shenghong Huang, Chao Yang 0001, Ziyu Zhang 0003, Yeqi Huang, Wenting Han, Xin Liu 0081, Dexun Chen, Ge Zhuang, Qiang Chen 0005 |
SC | 4 |
| 2021 | swFLOW: A large-scale distributed framework for deep learning on Sunway TaihuLight supercomputer
Mingfan Li, Junshi Chen 0003, José Monsalve Diaz, Rongfen Lin, Guang R. Gao, Hong An |
Inf. Sci. | 9 |
| 2021 | Towards Efficient Short-Range Pair Interaction on Sunway Many-Core Architecture
Junshi Chen 0003, Hong An, Wenting Han, Zeng Lin, Xin Liu 0081 |
J. Comput. Sci. Technol. | 2 |
| 2020 | RDMA-Based Apache Storm for High-Performance Stream Data Processing
Ziyu Zhang 0003, Zitan Liu, Qingcai Jiang, Junshi Chen 0003, Hong An |
NPC | 6 |
| 2020 | Distributed deep learning system for cancerous region detection on Sunway TaihuLight
Guofeng Lv, Mingfan Li, Hong An, Junshi Chen 0003, Wenting Han, Rongfen Lin |
CCF Trans. High Perform. Comput. | 3 |
| 2019 | DDP-B: A Distributed Dynamic Parallel Framework for Meta-genomics Binary Similarity
Mengxian Chi, Hong An |
NPC | 4 |
| 2019 | Degree-of-Node Task Scheduling of Fine-Grained Parallel Programs on Heterogeneous Systems
Mingfan Li, Chengfan Jia, Hong An |
J. Comput. Sci. Technol. | 5 |
| 2019 | CARS: A contention-aware scheduler for efficient resource management of HPC storage systems
Weihao Liang, Yong Chen 0001, Jialin Liu 0002, Hong An |
Parallel Comput. | 4 |
| 2018 | PEPS++: Towards Extreme-Scale Simulations of Strongly Correlated Quantum Many-Particle Models on Sunway TaihuLightabstractThe study of strongly frustrated magnetic systems has drawn great attentions from both theoretical and experimental physics. Efficient simulations of these models are essential for understanding their exotic properties. Here we present PEPS++, a novel computational paradigm for simulating frustrated magnetic systems and other strongly correlated quantum many-body systems. PEPS++ can accurately solve these models at the extreme scale with low cost and high scalability on modern heterogeneous supercomputers. We implement PEPS++ on Sunway TaihuLight based on a carefully designed tensor computation library for manipulating high-rank tensors and optimize it by invoking various high-performance matrix and tensor operations. By solving a 2D strongly frustrated$J_1$-$J_2$model with over ten million cores, PEPS++ demonstrates the capability of simulating strongly correlated quantum many-body problems at unprecedented scales with accuracy and time-to-solution far beyond the previous state of the art. Lixin He, Hong An, Chao Yang 0002, Junshi Chen 0003, Weihao Liang, Shao-Jun Dong, Qiao Sun 0005, Wenting Han, Yongjian Han, Wenjun Yao |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Pipelining Computation and Optimization Strategies for Scaling GROMACS on the Sunway Many-Core Processor
Hong An, Junshi Chen 0003, Weihao Liang, Qingqing Xu, Yong Chen 0001 |
ICA3PP | 2 |
| 2016 | Parallelizing Back Propagation Neural Network on Speculative MulticoresabstractApplications typically exhibit extremely different performance characteristics depending on the accelerator. Back propagation neural network (BPNN) has been parallelized into different platforms. However, it has not yet been explored on speculative multicore architecture thoroughly. This paper presents a study of parallelizing BPNN on a speculative multicore architecture, including its speculative execution model, hardware design and programming model. The implementation was analyzed with seven well-known benchmark data sets. Furthermore, it trades off several important design factors in coming speculative multicore architecture. The experimental results show that: (1) the BPNN performs well on speculative multicore platform. It can achieve similar speedup (17.7x to 57.4x) compared with graphics processors (GPU) while provides a more friendly programmability. (2) 64 cores' computing resources can be used efficiently and 4k is the proper speculative buffer capacity in the model. Yaobin Wang, Hong An, Zhiqin Liu, Dongmei Zhao |
ICPADS | 2 |
| 2015 | Optimization of Binomial Option Pricing on Intel MIC Heterogeneous System
Weihao Liang, Hong An, Yichao Cheng |
ICA3PP (3) | 2 |
| 2015 | Parallelizing Block Cryptography Algorithms on Speculative Multicores
Yaobin Wang, Hong An, Zhiqin Liu, Qingfeng Wang 0004 |
ICA3PP (1) | 2 |
| 2015 | Local State Reusing for Efficient Model Checking of Multithreaded Programs
Junrui Zhou, Hong An, Junshi Chen 0003 |
ICA3PP (4) | 2 |
| 2015 | Optimization and Analysis of Parallel Back Propagation Neural Network on GPU Using CUDA
Yaobin Wang, Pingping Tang, Hong An, Zhiqin Liu |
ICONIP (3) | 3 |
| 2014 | Understanding the SIMD Efficiency of Graph Traversal on GPU
Yichao Cheng, Hong An, Zhitao Chen, Xia Jiang |
ICA3PP (1) | 2 |
| 2014 | Efficient execution of speculative threads and transactions with hardware transactional memory
Gongming Li, Hong An, Qi Li 0034, Bobin Deng, Wenbo Dai |
Future Gener. Comput. Syst. | 2 |
| 2012 | SeTM: Efficient Execution of Speculative Threads with Hardware Transactional MemoryabstractThread-Level Speculation (TLS) was researched to automatically parallelize portions of serial programs for execution, and transactional memory (TM) was studied as a promising alternative of lock for parallel programming due to its simplicity. Both TLS and TM require similar underlying support. In the paper, we present SeTM (Sequential Transactional Memory), a hardware enhanced TM system which supports TLS at minor extra cost. Signature is an effective way to buffer speculative states in TM and TLS. But it cripples TM and TLS performance due to its false-positive in terms of conflict detection, especially for conflict-intensive TLS. SeTM adopts R/W bits and signature concurrently to ameliorate this bad influence. Additionally, SeTM introduces fast rollback mechanism, which provides fast abort recovery for eager log-based HTM and TLS. The most important contribution of SeTM is conflict-tolerant mechanism, which tolerates some ambiguous data conflicts in TLS. Six representative benchmarks have been adopted to evaluate our model. Our experimental results show that our scheme improves the execution performance of most tested codes at a modest hardware cost. For a set of important scientific loops, we report the highest speedup of 6.5 with 15 cores. Besides, experimental results also show good scalability of SeTM system. Gongming Li, Hong An, Qi Li 0034, Bobin Deng, Wenbo Dai |
ICPADS | 2 |
| 2012 | Distributed replay protocol for distributed uniprocessorsabstractData speculation technique has been heavily exploited in various scenarios of architecture design. It bridges the time or space gap between data producer and data consumer, which gives opportunities to processors to gain significant speedups. However, large instruction windows, deep pipeline and increasing latency of on-chip communication make data misspeculation very expensive in modern processors. Mengjie Mao, Hong An, Bobin Deng, Xuechao Wei, Wenting Han |
ICS | 2 |
| 2012 | CRQ-based fair scheduling on composable multicore architecturesabstractAs different workloads require different processor resources for better execution efficiency, recent work has proposed composable chip multiprocessors (CCMPs), which provide the capability to configure different number and types of processing cores at system runtime. However, such composable architecture poses a new significant challenge to system scheduler, that is, how to ensure priority-based performance for each task (i.e. fairness), while exploiting the benefits of composability by dynamically changing the hardware configurations to match the parallelism requirements in running tasks (i.e. resource allocation). Current multicore schedulers fail to address this problem, as they traditionally assume fixed number and types of cores. Hong An, Haibo Zhang 0005, Xiufeng Sui |
ICS | 2 |
| 2012 | FlexBFS: a parallelism-aware implementation of breadth-first search on GPUabstractIn this paper, we present FlexBFS, a parallelism-aware implementation for breadth-first search on GPU. Our implementation can adjust the computation resources according to the feedback of available parallelism dynamically. We also optimized our program in three ways: (1)a simplified two-level queue management,(2)a combined kernel strategy and (3)a high-degree vertices specialization approach. Our experimental results show that it can achieve 3~20 times speedup against the fastest serial version, and can outperform the TBB based multi-threading CPU version and the previous most effective GPU version on all types of input graphs. Gu Liu, Hong An, Wenting Han, Xuechao Wei, Xulong Tang |
PPoPP | 2 |
| 2011 | A Priority-Aware NoC to Reduce Squashes in Thread Level Speculation for Chip MultiprocessorsabstractThread Level Speculation (TLS) is a technique aims at boosting the performance of sequential programs running on Chip Multiprocessors (CMPs) by automatically parallelizing them. It exempts programmers from the heavy task of parallel programming. But its performance may suffer from frequent squashing caused by inter-thread data dependency violation. In this paper, we propose a Network-on-Chip (NoC) in CMP that employs a priority-aware packet arbitration policy. Packet scheduling guided by such policy reduces the occurrence of TLS squashes. Simulation results with 5 applications show that our policy reduces squashes by 22% in best case and 15% on average. Moreover, our priority aware approach could be generalized to similar scenarios in which different threads running on CMP manifest different priorities. Wenbo Dai, Hong An, Qi Li 0034, Gongming Li, Bobin Deng, Shilei Wu |
ISPA | 2 |
| 2011 | A Non-blocking Programming Framework for Pipeline Application on Multi-core PlatformabstractMany applications meet certain programming patterns like pipeline, fork-join, do-all etc. While tools such as OS threads and OpenMP allow programmers only to express task or data parallelism, special support for programming patterns is distinctly lacking. Intel threading building blocks (TBB) is developed to address this problem, but its scheduler is general and not optimized for any of its parallel algorithms which include pipeline specially. In this paper, we provide a non-blocking framework for pipeline application on multi-core platform. We target linear pipeline in which each filter has one entrance and one exit. We design a novel work-stealing scheduler optimized specially for pipeline application: first, priority based stealing, priority is calculated for each filter in pipeline so that a worker can find the optimal "victim" easily when it needs to steal, second, multiple tasks can be stolen at a time so that much stealing time is reduced. A non-block queue is used to store intermediate result to reduce lock overhead and increase scalability. We apply our framework to four case studies, including text filter, two fish, ferret, ded up. And our framework reduces execution time of TBB by 72% in best case and 20% on average on an 8 core machine. Hong An, Gu Liu, Wenting Han, Mu Xu, Qi Li 0034 |
ISPA | 2 |
| 2011 | CHMasters: A Scalable and Speed-Efficient Metadata Service in Distributed File SystemabstractDistributed file system (DFS) is playing important roles of supporting large distributed data-intensive applications to meet storage needs. Typically, the design of DFS, such as GFS in Google, DMS in Cisco and TFS in Alibaba, is driven by observations of specific application workloads, internal demands and technological environment. In such systems, the metadata service is a critical factor that can affect the file system performance and availability to a great degree. Five requirements have been summarized for the metadata service: location transparent file service, smart director, efficient speed, strong scalability and friendly collaborator. In this paper, we present metadata service module called CH Masters in our DFS. Consistent hashing protocol is used to relieve potential hot spots on name servers. Files' metadata and master nodes are mapped into the same hash space by consistent hash function. And then files' metadata are scattered to master nodes by clockwise "closest" principle. Chunk server acts as a client when report its chunks info. Only a small proportion of files' metadata will be rehashed when master nodes state change. A new scalable file mapping strategy is also proposed to map file sizes from few MB to several GB efficiently. After intensive experiments, it shows CH Masters is satisfying the above five requirements. Junrui Zhou, Hong An |
PDCAT | 4 |
| 2010 | Dynamic Resource Tuning for Flexible Core Chip Multiprocessors
Yongqing Ren, Hong An, Yaobin Wang |
ICA3PP (2) | 2 |
| 2010 | Pattern-Unit Based Regular Expression Matching with Reconfigurable Function Unit
Hong An, Peng Li 0031, Tao Wang 0004, Zhihong Yu |
ICCSA (4) | 2 |
| 2010 | FACRA: Flexible-Core Architecture Chip Resource AbstractorabstractA family of flexible-core chip multiprocessors (FCMPs) has been recently proposed to allow simple, identical physical cores to be aggregated dynamically to form larger and more powerful logical processors. However, such flexible-core architecture faces a new significant scheduling problem in the operating system, which traditionally assumes only fixed-number and fixed-granularity processors. This paper proposes a framework, called FACRA, that employs low-level runtime software to simplify OS resource allocation and process scheduling on FCMPs. Through exporting a simple, uniform processor abstraction on flexible-core chip resource, FACRA provides a set of functions with uniform interface for system-level scheduling on FCMPs. To verify the design, FACRA is built on TFlex (a typical FCMP) in our experiments, and two well known process schedulers, round-robin and dynamic-priority scheduler of Linux 2.6.11, are modified to schedule on TFlex. The evaluation results demonstrate that FACRA can efficiently simplify OS resource allocation and process scheduling on FCMPs with negligible performance loss. Hong An, Yongqing Ren, Mengjie Mao, Mu Xu, Qi Li 0034 |
PDCAT | 2 |
| 2009 | The Mapping Framework and Optimizing Strategy for Block Cryptography Algorithms on Cell Broadband EngineabstractThe Cell Broadband Engine is a typical heterogeneous chip multiprocessor which provides potential high performance for computing-intensive applications. Our researches focus on how to use Cell to speed up block cryptography applications. In this paper, we propose a mapping framework for block cryptography working in ECB mode and corresponding optimizing strategy. We take four algorithms(RC5, 3DES, AES, and Twofish) as benchmark and implement these four algorithms using Cell programming language. In order to enhance the performance, we present an optimizing strategy and evaluate the effects of the optimizing methods including compiler optimization, dual buffering, vectorization, and loop unrolling. The experiments indicate that all these four algorithms can obtain 5-20 times speedup compared with traditional processors, which shows that our mapping framework and optimizing strategy are effective for the block cryptography algorithms. Mu Xu, Hong An, Gu Liu, Yaobin Wang, Ping Yao, Xiurui Hao, Wenting Han |
PDCAT | 2 |
| 2007 | Balancing Thread Partition for Efficiently Exploiting Speculative Thread-Level Parallelism
Yaobin Wang, Hong An, Yongqing Ren |
APPT | 2 |