VLDB 2026 Research / reviewers in the wild / expert
Chun Huang 0006
dblp:64/2438-6
· DBLP profile ↗
46ranked-venue papers
0as first author
36since 2021 · last 2026
0000-0002-0317-8192ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 28 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Theory of computation · 2 · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Compensated Estrin Scheme for Accurate Polynomial Evaluation
Guangping Yu, Stef Graillat, Hao Jiang 0001, Chun Huang 0006, Tao Tang 0001 |
CASC | 4 |
| 2026 | Adaptive Data Augmentation with Bayesian Optimization for Basic Block Throughput Prediction
Xiabing Hu, Xuezheng Xu, Deheng Yang, Chun Huang 0006 |
Euro-Par (1) | 5 |
| 2026 | Uni-STC: Unified Sparse Tensor CoreabstractModern processors are increasingly adopting tensor cores as key computational units. Compared to existing designs for dense and structured sparsity, recent dual-side sparse tensor cores have evolved to support general sparsity. However, existing methods still face limitations on generality (incomplete sparse kernel support prevents broad applicability) and performance (outer-product/row-row schemes yield unsatisfactory hardware utilisation, data reuse, and energy efficiency). In this paper, we propose Uni-STC, a unified sparse tensor core that delivers high-performance dataflows for four key sparse kernels: sparse matrix-vector multiplication (SpMV), sparse matrixsparse vector multiplication (SpMSpV), sparse matrix-multiple vector multiplication (SpMM), and sparse general matrix-matrix multiplication (SpGEMM). To efficiently support these diverse sparse workloads, we first introduce BBC, a unified sparse format co-designed with Uni-STC's dataflow. We then design UniSTC's architecture supporting (1) fine-grained task partitioning to improve resource utilisation, (2) parallel sparse-tile processing to enhance data reuse, and (3) a dynamic network to reduce intermediate data movement and energy consumption. Evaluated across 2893 SuiteSparse and 302 DLMC matrices, Uni-STC demonstrates significant improvements, outperforming the state-of-the-art RM-STC with a$2.21 \times$geomean speedup and$2.96 \times$higher energy efficiency. Haocheng Lian, Meichen Dong, Yijie Nie, Junzhong Shen, Chun Huang 0006, Bingcai Sui, Weifeng Liu 0002 |
HPCA | 9 |
| 2026 | BAAS: A Bidirectional Aggregation and Affinity-Aware Scheduling Framework for Parallelizing Sparse Matrix Computations
Qingyang Zhang 0009, Chuanfu Xu, Chun Huang 0006, Zhimeng Han, Jie Liu 0002 |
IPDPS | 4 |
| 2026 | Optimizing small matrix multiplications via batch grouping on multi-core DSPs
Xiaotian Chen, Jianbin Fang, Peng Zhang 0061, Chun Huang 0006 |
CCF Trans. High Perform. Comput. | 5 |
| 2026 | High-performance matrix multiplication micro-kernel generation for GPDSPs via MLIR progressive lowering
Kainan Yu, Peng Zhang 0061, Chun Huang 0006, Jianbin Fang |
CCF Trans. High Perform. Comput. | 4 |
| 2025 | Function-level Optimization Automatic Tuner for Numerical ProgramsabstractNumerical programs are widely used in high-performance computing, graphics, finance, deep learning, and other fields. The use of high-precision floating-point numbers can ensure the accuracy and robustness of the program and reduce the accumulation of errors. However, this approach also increases the program’s execution time, memory usage, and energy consumption. Therefore, the reasonable use of mixed precision is beneficial to balance the accuracy and performance of program results. To look for efficient mixed-precision configurations, we propose an LLVM toolchain, FuncTuner, that obtains the program transformation range by identifying the #pragma directive in the front end of Clang, uses a recursive search algorithm to access the configuration information of functions in the range, and uses LLVM Pass to modify the LLVM IR of the program according to the given configuration, pioneering the use of function-level mixed-precision optimization. Our approach achieves significant results in the HPL-AI program, with a maximum floating-point performance of 24.78 GFLOPs at fp16 precision and a performance improvement of 304.73 % at the max scale. Xinni Liu, Guangping Yu, Hengbiao Yu, Xin Yi 0002, Chun Huang 0006 |
APSEC | 6 |
| 2025 | Selection of Supervised Learning-Based Sparse Matrix Reordering Algorithms
Tao Tang 0001, Youfu Jiang, Yingbo Cui 0001, Jianbin Fang, Peng Zhang 0061, Lin Peng 0001, Chun Huang 0006 |
HiPC | 7 |
| 2025 | Optimizing Direct Convolutions on High-Performance Multi-Core DSPsabstractConvolution operations form the computational backbone of deep learning inference but often become performance bottlenecks on conventional architectures. While multi-core Digital Signal Processors (DSPs) offer energy-efficient alternatives through long vector units and software-managed memory hierarchies, existing convolution optimizations designed for CPUs/GPUs underperform due to architectural mismatches in memory systems and execution pipelines. We present mtConv, an optimizing convolution method for multi-core DSPs. mtConv achieves high performance by exploiting data reuse, managing on-chip memory, and designing efficient micro-kernels. It maximizes the overlap between computation and communication to hide data transfer latency, leveraging DSPs’ long vector units and hierarchical scratchpad memories. We evaluate mtConv against state-of-the-art convolution optimizations on DSPs. Experimental results show that mtConv delivers the best overall performance across various convolution layers, achieving up to 93.25% of the hardware’s peak performance on a single DSP core and 92.31% when using all 8 cores of a single DSP cluster. Xiaotian Chen, Jianbin Fang, Peng Zhang 0061, Yonggang Che, Chun Huang 0006, Jie Ren 0007 |
ICPP | 6 |
| 2025 | Fine-Grained Global Search for Inputs Triggering Floating-Point Exceptions in Gpu ProgramsabstractFloating-point exceptions are hard to avoid and can cause disastrous consequences. However, testing methods for floating-point exceptions in GPU programs are currently quite limited due to their closed-source nature. Existing tools, even the state-of-the-art Xscope, still exhibit low search efficiency and poor input coverage. In this paper, we combine interval-wise random sampling and Markov Chain Monte Carlo (MCMC) sampling in a synergistic way to efficiently detect exception-inducing inputs in GPU programs. To improve the search efficiency, based on the bit patterns of exceptional floating-point values, we propose a floating-point format-aware input space partitioning method for random sampling and define a unified fitness function for MCMC sampling. We implement our approach in a tool DFEG and demonstrate it on 76 functions from the CUDA Math Library, HPC programs, and FPBench. DFEG outperforms Xscope in terms of both effectiveness and efficiency. DFEG finds$949 \times$more exceptions than Xscope and detects new exceptions in 9 functions where Xscope fails. Moreover, compared to Xscope, DFEG achieves an average$34 \times$speedup. Xin Yi 0002, Hengbiao Yu, Liqian Chen, Xiaoguang Mao, Ji Wang 0001, Chun Huang 0006, Deheng Yang |
IPDPS | 6 |
| 2025 | An empirical study of error-free transformations for enhancing mathematical function precision
Dongting Chen, Jie Shen 0003, Chun Huang 0006, Xin Yi 0002 |
CCF Trans. High Perform. Comput. | 3 |
| 2025 | An empirical performance evaluation of SYCL on ARM multi-core processors
Hanzheng Liang, Chencheng Deng, Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Chun Huang 0006 |
CCF Trans. High Perform. Comput. | 6 |
| 2025 | Gator: Accelerating Graph Attention Networks by Jointly Optimizing Attention and Graph ProcessingabstractGraph attention networks (GATs) have advanced performance in various application domains by introducing the attention mechanism into the graph neural networks (GNNs). The inefficiency of running GATs on CPUs or GPUs necessitates specialized hardware designs. Unfortunately, previous specialized architecture designs have focused on either the GNN architecture or the attention mechanism, resulting in limited performance and leaving ample room for improvement. This article presents Gator , a joint optimization approach with software–hardware co-designs for GAT inference. On the software level, Gator leverages degree-weighted graph partitioning and parameter-adaptive feature selection techniques to preprocess the input graph data, mining subgraph-level parallelism and mitigating the computation bottleneck of the dedicated dataflow. On the hardware level, Gator designs a unified processing engine to support various kernels by extracting a common computation pattern and a dimension-aware microarchitecture for efficient partial sum reduction. Extensive experiments show that our approach can achieve 11.5× more efficiency compared to NVIDIA RTX 4090 and provide a speedup of 3× to 9.4×, along with a 2.6× to 4.7× reduction in memory traffic, when compared to six state-of-the-art methods, with minimal accuracy loss. Xiaobo Lu, Jianbin Fang, Lin Peng 0001, Chun Huang 0006, Zixiao Yu |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUsabstractConvolution kernels are widely seen in high-performance computing (HPC) and deep learning (DL) workloads and are often responsible for performance bottlenecks. Prior works have demonstrated that the direct convolution approach can outperform the conventional convolution implementation. Although well-studied, the existing approaches for direct convolution are either incompatible with the mainstream DL data layouts or lead to suboptimal performance. We designnDirect2, a novel direct convolution approach that targets multi-core CPUs commonly found in smartphones and HPC systems.nDirect2is compatible with the data layout formats used by mainstream DL frameworks and offers new optimizations for the computational kernel, data packing, advanced operator fusion, and parallelization. We evaluatenDirect2by applying it to representative convolution kernels and demonstrating how well it performs on four distinct ARM-based CPUs and an X86-based CPU. Experimental results show thatnDirect2outperforms four state-of-the-art convolution approaches across most evaluation cases and hardware architectures. Weiling Yang, Jianbin Fang, Dezun Dong, Zhengbin Pang, Runxi He, Peng Zhang 0061, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Jie Ren 0007 |
IEEE Trans. Computers | 9 |
| 2024 | VLASPH: Smoothed Particle Hydrodynamics on VLA SIMD Architectures
Xiaokang Fan, Zhen Ge, Tao Tang 0001, Chun Huang 0006, Lin Peng 0001, Canqun Yang |
Euro-Par (3) | 5 |
| 2024 | Parallel Optimization for Accelerating the Generation of Correctly Rounded Elementary FunctionsabstractCorrectly rounded elementary mathematical functions are crucial for numerical computations and scientific applications. Generating these functions accurately is a challenging task. The latest methods automate this process by transforming the problem of generating correctly rounded elementary mathematical functions into a linear programming problem. However, this generation process is serial, and the inefficiency of serialization hinders the creation of new elementary mathematical functions and limits the broader application of the technique. Xianglin Wang, Xin Yi 0002, Hengbiao Yu, Chun Huang 0006, Lin Peng 0001 |
ICPP | 4 |
| 2024 | Optimizing Stencil Computation on Multi-core DSPsabstractStencil is a common computation pattern in high-performance computing (HPC) applications. While extensive work has been proposed to optimize stencil kernels on CPUs and GPUs, there is no consensus on how to best optimize stencils on multi-core Digital Signal Processors (DSPs) used in emerging HPC systems. This paper shares our experience in optimizing stencil kernels on multi-core DSPs. Our approach combines coarse and fine-grained parallel optimization techniques to enhance the performance of stencil computations. Our optimizations include a vectorization-enabled micro-kernel to utilize instruction parallelism, a memory-aware data reuse strategy to maximize data locality across multiple memory levels and a triple-buffering mechanism to overlap computation and memory communications. Experimental results show that our approach can effectively utilize the memory bandwidth and the computation capability of the underlying hardware. Our integrated optimizations can yield a 3.72x speedup over the 16-core CPU counterpart. Fugeng Zhu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Yonggang Che, Kainan Yu, Jing Xie 0023, Chun Huang 0006, Jie Ren 0007 |
ICPP | 9 |
| 2024 | Optimizing General Matrix Multiplications on Modern Multi-core DSPsabstractGeneral Matrix Multiplication (GEMM) is a key subprogram in high-performance computing (HPC) and deep learning workloads. With the rising significance of power and energy consumption in HPC systems, accelerators based on Digital Signal Processors (DSPs) have been integrated into general-purpose HPC systems. Due to the architecture disparities, the GEMM optimization techniques used on conventional multi-core CPUs and GPGPUs are not always applicable to DSPs. This paper shares our experience in optimizing GEMM on multi-core GPDSPs, using a CPU-DSP processor as a case study. Our approach employs a range of techniques to optimize performance for DSP architectures. These include data partitioning, three-level pipelining, dedicated micro-kernel design, and improved vector reduction. These optimizations maximize the overlap between computation and communication while fully exploiting the capabilities of floating-point arithmetic units to achieve high performance. Our experimental results demonstrate that the performance attained by our optimization is up to 96% of the theoretical peak performance of the hardware. Kainan Yu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Dezun Dong, Ruibo Wang, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Zheng Wang 0001 |
IPDPS | 8 |
| 2024 | Efficient compiler optimization by modeling passes dependenceabstractAbstract Selecting the optimal combination of compiler passes is a significant challenge to enhance performance and reduce the code size of compiled binaries. While a well-selected sequence of compiler passes can yield considerable benefits, the large number of potential combinations and the scarcity of effective ones make this task prohibitively complex. To tackle this problem, we propose a novel approach to group compiler passes into a small set of sub-sequences. This approach translates the task of identifying the right compiler passes combination into determining the appropriate combination of these sub-sequences. We apply our approach to CBench and PolyBench, demonstrating remarkable performance improvements. Our approach enhances runtime performance by 22% compared to the default LLVM ‘O3’ option, and achieves a code size reduction of 24% compared to the ‘Oz’ option. Our approach also outperforms state-of-the-art across various optimization tasks and hardware platforms. Jianbin Fang, Ting Wang 0009, Jing Xie 0023, Chun Huang 0006 |
CCF Trans. High Perform. Comput. | 5 |
| 2024 | thSORT: an efficient parallel sorting algorithm on multi-core DSPs
Mouzhi Yang, Peng Zhang 0061, Jianbin Fang, Weifeng Liu 0002, Chun Huang 0006 |
CCF Trans. High Perform. Comput. | 5 |
| 2024 | Mentor: A Memory-Efficient Sparse-dense Matrix Multiplication Accelerator Based on Column-Wise ProductabstractSparse-dense matrix multiplication (SpMM) is the performance bottleneck of many high-performance and deep-learning applications, making it attractive to design specialized SpMM hardware accelerators. Unfortunately, existing hardware solutions do not take full advantage of data reuse opportunities of the input and output matrices or suffer from irregular memory access patterns. Their strategies increase the off-chip memory traffic and bandwidth pressure, leaving much room for improvement. We present Mentor , a new approach to designing SpMM accelerators. Our key insight is that column-wise dataflow, while rarely exploited in prior works, can address these issues in SpMM computations. Mentor is a software-hardware co-design approach for leveraging column-wise dataflow to improve data reuse and regular memory accesses of SpMM. On the software level, Mentor incorporates a novel streaming construction scheme to preprocess the input matrix for enabling a streaming access pattern. On the hardware level, it employs a fully pipelined design to unlock the potential of column-wise dataflow further. The design of Mentor is underpinned by a carefully designed analytical model to find the tradeoff between performance and hardware resources. We have implemented an FPGA prototype of Mentor . Experimental results show that Mentor achieves speedup by geomean 2.05× (up to 3.98×), reduces the memory traffic by geomean 2.92× (up to 4.93×), and improves bandwidth utilization by geomean 1.38× (up to 2.89×), compared with the state-of-the-art hardware solutions. Xiaobo Lu, Jianbin Fang, Lin Peng 0001, Chun Huang 0006, Zidong Du, Yongwei Zhao 0001, Zheng Wang 0079 |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | SNCL: a supernode OpenCL implementation for hybrid computing arrays
Tao Tang 0001, Kai Lu 0001, Lin Peng 0001, Yingbo Cui 0001, Jianbin Fang, Chun Huang 0006, Ruibo Wang, Canqun Yang, Yifei Guo |
J. Supercomput. | 6 |
| 2024 | Optimizing Multi-Grid Preconditioned Conjugate Gradient Method on Multi-CoresabstractMultigrid preconditioned conjugate gradient (MGPCG) is commonly used in high-performance computing (HPC) workloads. However, MGPCG is notoriously challenging to optimize since most of its computation kernels are memory-bounded with low arithmetic intensity and non-trivial communication patterns among parallel processes. This article presents new techniques to improve the data locality and reduce the communication overhead of MGPCG by first merging the kernels of multigrid (MG). We then develop an asynchronous neighboring communication algorithm to reduce the data communications across parallel processes. We demonstrated the benefits of our approach by applying it to the high-performance conjugate gradient (HPCG) benchmark and integrating it with a real-life algebraic multigrid package. We test the resulting software implementations on three ARMv8 and one Intel Xeon system. Experimental results show that our approach leads to a 1.62x-2.54x speedup over the engineer- and vendor-tuned HPCG implementations across various workloads and platforms. Xiaojian Yang, Shengguo Li, Dezun Dong, Chun Huang 0006, Zheng Wang 0079 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | Accelerating Type Confusion Detection by Identifying Harmless Type CastingsabstractC++ allows reinterpretation of memory objects via type casting, which facilitates easier manipulation of class fields and virtual methods inside the class hierarchy. However, misinterpretation of memory objects, which is called type confusion, can result in illegal access of class fields or methods. Type confusion accounts for many security vulnerabilities for programs written in C++. Previous type confusion detection techniques report a type confusion bug when an object of a parent class is casted to a child class. However, a downcast is safe as long as no illegal fields or methods are accessed. This paper presents Harmless Type Casting Detection (htade), which identifies safe downcast instructions and removes redundant runtime verifications before them by analyzing the type and access information of casted objects. We evaluated htade against 11 SPEC CPU 2006/2017 C++ programs. Compared with LLVM-CFI, htade can reduce the runtime performance overhead by 58.98% on average. Xiaokang Fan, Chun Huang 0006, Canqun Yang, Fa Li |
CF | 3 |
| 2023 | Optimizing Multi-grid Computation and Parallelization on Multi-coresabstractMultigrid algorithms are widely used to solve large-scale sparse linear systems, which is essential for many high-performance workloads. The symmetric Gauss-Seidel (SYMGS) method is often responsible for the performance bottleneck of MG. This paper presents new methods to parallelize and enhance the computation and parallelization efficiency of the SYMGS and MG algorithms on multi-core CPUs. Our solution employs a matrix splitting strategy and a revised computation formula to decrease the computation operations and memory accesses in SYMGS. With this new SYMGS strategy, we can then merge the two most time-consuming components of MG. On top of these, we propose a new asynchronous parallelization scheme to reduce the synchronization overhead when parallelizing SYMGS. We demonstrate the benefit of our techniques by integrating them with the HPCG benchmark and two real-life applications. Evaluation conducted on four architectures, including three ARMv8 and one x86, shows that our techniques greatly surpass the performance of engineer- and vendor-tuned implementations across various workloads and platforms. Xiaojian Yang, Shengguo Li, Dezun Dong, Chun Huang 0006, Zheng Wang 0001 |
ICS | 5 |
| 2023 | Optimizing Direct Convolutions on ARM Multi-CoresabstractConvolution kernels are widely seen in deep learning workloads and are often responsible for performance bottlenecks. Recent research has demonstrated that a direct convolution approach can outperform the traditional convolution implementation based on tensor-to-matrix conversions. However, existing approaches for direct convolution still have room for performance improvement. We present nDirect, a new direct convolution approach that targets ARM-based multi-core CPUs commonly found in smartphones and HPC systems. nDirect is designed to be compatible with the data layout formats used by mainstream deep learning frameworks but offers new optimizations for the computational kernel, data packing, and parallelization. We evaluate nDirect by applying it to representative convolution kernels and demonstrating its performance on four distinct ARM multi-core CPU platforms. We compare nDirect against state-of-the-art convolution optimization techniques. Experimental results show that nDirect gives the best overall performance across evaluation scenarios and platforms. Weiling Yang, Jianbin Fang, Dezun Dong, Chun Huang 0006, Peng Zhang 0061, Tao Tang 0001, Zheng Wang 0001 |
SC | 5 |
| 2023 | Efficient Generation of Floating-Point Inputs for Compiler-Induced VariabilityabstractIn scientific computation, developers usually exploit the compiler to improve the performance of floating-point programs. However, many compiler optimizations might affect the floating-point behavior, which can cause numerical variations. This paper proposes an efficient generation method of floating-point inputs for compiler-induced variability. Specifically, we formulate the problem of generating high variability-inducing inputs as a mathematical optimization problem and solve it through input space partition and Markov Chain Monte Carlo (MCMC) sampling. To improve the sampling efficiency, besides the result variation, we utilize the difference between the execution traces of floating-point instructions to guide the search. We have implemented our approach in the tool CIV. Compared to the state-of-the-art method, CIV achieves an average 11x speedup for generating an equivalent or better input to trigger large result variations. Moreover, CIV finds better inputs for 100% programs and has better stability for detecting large result variations. The experimental results demonstrate the effectiveness and efficiency of our approach. Hengbiao Yu, Xin Yi 0002, Banghu Yin, Fa Li, Zhenbang Chen 0001, Chun Huang 0006 |
SANER | 6 |
| 2023 | wrBench: Comparing Cache Architectures and Coherency Protocols on ARMv8 Many-Core Systems
Wanrong Gao, Jianbin Fang, Chun Huang 0006, Chuanfu Xu, Zheng Wang 0001 |
J. Comput. Sci. Technol. | 3 |
| 2023 | A Quantitative Evaluation of Vector Transcendental Functions on ARMv8-Based Processors
Jie Shen 0003, Biao Long, Chun Huang 0006 |
J. Comput. Sci. Technol. | 3 |
| 2023 | Programming bare-metal accelerators with heterogeneous threading models: a case study of Matrix-3000abstractAs the hardware industry moves toward using specialized heterogeneous many-core processors to avoid the effects of the power wall, software developers are finding it hard to deal with the complexity of these systems. In this paper, we share our experience of developing a programming model and its supporting compiler and libraries for Matrix-3000, which is designed for next-generation exascale supercomputers but has a complex memory hierarchy and processor organization. To assist its software development, we have developed a software stack from scratch that includes a low-level programming interface and a high-level OpenCL compiler. Our low-level programming model offers native programming support for using the bare-metal accelerators of Matrix-3000, while the high-level model allows programmers to use the OpenCL programming standard. We detail our design choices and highlight the lessons learned from developing system software to enable the programming of bare-metal accelerators. Our programming models have been deployed in the production environment of an exascale prototype system. Jianbin Fang, Peng Zhang 0061, Chun Huang 0006, Tao Tang 0001, Kai Lu 0001, Ruibo Wang, Zheng Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2022 | MT-3000: a heterogeneous multi-zone processor for HPC
Kai Lu 0001, Yang Guo 0003, Chun Huang 0006, Sheng Liu 0001, Ruibo Wang, Jianbin Fang, Tao Tang 0001, Zhaoyun Chen, Biwei Liu, Zhong Liu 0003, Yuanwu Lei, Haiyan Sun |
CCF Trans. High Perform. Comput. | 4 |
| 2022 | Efficient Data Redistribution Algorithms From Irregular to Block Cyclic Data DistributionabstractIn this paper, we propose some efficient data redistribution algorithms for redistributing matrices from 1D or 2D irregular format to block cyclic data distribution (BCDD) format, which can be much faster than the BLACS routinePXGEMR2D. These algorithms can be used to combine direct methods with iterative methods. The proposed algorithms divide the communication into two phases: one for processes in the same column and the other for processes in the same row, and the whole data redistribution task is divided into several independent sub-communications. The communication time can be reduced a lot compared with BLACS. Performance results show that our algorithms can be$2\times$–$5\times$faster than the BLACS routinePXGEMR2Dwhen using 4096 processes and the experiments are performed on Tianhe-2A supercomputer. Shengguo Li, Hao Jiang 0001, Dezun Dong, Chun Huang 0006, Jie Liu 0002, Xia Liao, Xuguang Chen |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | Optimizing Barrier Synchronization on ARMv8 Many-Core ArchitecturesabstractSynchronization operations are commonly seen in OpenMP programs where a parallel construct often works with an explicit or implicit barrier operation. While OpenMP synchronization has been extensively studied on the traditional x86 CPU architectures, there is little work on understanding OpenMP barrier synchronization operations on ARMv8 high-performance many-cores. This paper presents the first comprehensive performance study on OpenMP barrier implementations on emerging ARMvS-based many-cores. We evaluate seven representative barrier algorithms on three distinct ARMv8 architectures: Phytium 2000+, ThunderX2, and Kunpeng920. We empirically show that the existing synchronization implementations exhibit poor scalability on ARMv8 architectures compared to the x86 counterpart. We then propose various optimization strategies for improving these widely used synchronization algorithms on each platform. We showcase that our optimizations yield 12.6x performance improvement over the GCC implementation and 4.7x improvement over the LLVM implementation, translating to 1.6x improvement over the state-of-the-art best-performing algorithm. We share our experience and practical insights on optimizing OpenMP synchronization operations on emerging ARMv8 multi-core CPU architectures. Wanrong Gao, Jianbin Fang, Chun Huang 0006, Chuanfu Xu, Zheng Wang 0001 |
CLUSTER | 3 |
| 2021 | Large-Scale Parallel Alignment Algorithm for SMRT Reads
Yingbo Cui 0001, Peng Zhang 0061, Tao Tang 0001, Lin Peng 0001, Chun Huang 0006, Canqun Yang, Xiangke Liao |
ICA3PP (2) | 8 |
| 2021 | VISPR-online: a web-based interactive tool to visualize CRISPR screening experimentsabstractBACKGROUND: VISPR is an interactive visualization and analysis framework for CRISPR screening experiments. However, it only supports the output of MAGeCK, and requires installation and manual configuration. Furthermore, VISPR is designed to run on a single computer, and data sharing between collaborators is challenging. RESULTS: To make the tool easily accessible to the community, we present VISPR-online, a web-based general application allowing users to visualize, explore, and share CRISPR screening data online with a few simple steps. VISPR-online provides an exploration of screening results and visualization of read count changes. Apart from MAGeCK, VISPR-online supports two more popular CRISPR screening analysis tools: BAGEL and JACKS. It provides an interactive environment for exploring gene essentiality, viewing guide RNA (gRNA) locations, and allowing users to resume and share screening results. CONCLUSIONS: VISPR-online allows users to visualize, explore and share CRISPR screening data online. It is freely available at http://vispr-online.weililab.org , while the source code is available at https://github.com/lemoncyb/VISPR-online . Yingbo Cui 0001, Johannes Köster, Xiangke Liao, Shaoliang Peng, Tao Tang 0001, Chun Huang 0006, Canqun Yang |
BMC Bioinform. | 7 |
| 2021 | Performance Evaluation of Memory-Centric ARMv8 Many-Core Architectures: A Case Study with Phytium 2000+
Jianbin Fang, Xiangke Liao, Chun Huang 0006, Dezun Dong |
J. Comput. Sci. Technol. | 3 |
| 2020 | Symbolic verification of message passing interface programsabstractMessage passing is the standard paradigm of programming in high-performance computing. However, verifying Message Passing Interface (MPI) programs is challenging, due to the complex program features (such as non-determinism and non-blocking operations). In this work, we present MPI symbolic verifier (MPI-SV), the first symbolic execution based tool for automatically verifying MPI programs with non-blocking operations. MPI-SV combines symbolic execution and model checking in a synergistic way to tackle the challenges in MPI program verification. The synergy improves the scalability and enlarges the scope of verifiable properties. We have implemented MPI-SV1 and evaluated it with 111 real-world MPI verification tasks. The pure symbolic execution-based technique successfully verifies 61 out of the 111 tasks (55%) within one hour, while in comparison, MPI-SV verifies 100 tasks (90%). On average, compared with pure symbolic execution, MPI-SV achieves 19x speedups on verifying the satisfaction of the critical property and 5x speedups on finding violations. Hengbiao Yu, Zhenbang Chen 0001, Xianjin Fu, Ji Wang 0001, Zhendong Su 0001, Jun Sun 0001, Chun Huang 0006, Wei Dong 0006 |
ICSE | 7 |
| 2020 | Symbolic Verification of MPI Programs with Non-deterministic Synchronizations
Hengbiao Yu, Zhenbang Chen 0001, Chun Huang 0006, Ji Wang 0001 |
SETTA | 3 |
| 2020 | Parallel programming models for heterogeneous many-cores: a comprehensive survey
Jianbin Fang, Chun Huang 0006, Tao Tang 0001, Zheng Wang 0001 |
CCF Trans. High Perform. Comput. | 2 |
| 2020 | Optimizing Streaming Parallelism on Heterogeneous Many-Core ArchitecturesabstractAs many-core accelerators keep integrating more processing units, it becomes increasingly more difficult for a parallel application to make effective use of all available resources. An effective way of improving hardware utilization is to exploit spatial and temporal sharing of the heterogeneous processing units by multiplexing computation and communication tasks - a strategy known as heterogeneous streaming. Achieving effective heterogeneous streaming requires carefully partitioning hardware among tasks, and matching the granularity of task parallelism to the resource partition. However, finding the right resource partitioning and task granularity is extremely challenging, because there is a large number of possible solutions and the optimal solution varies across programs and datasets. This article presents an automatic approach to quickly derive a good solution for hardware resource partition and task granularity for task-based parallel applications on heterogeneous many-core architectures. Our approach employs a performance model to estimate the resulting performance of the target application under a given resource partition and task granularity configuration. The model is used as a utility to quickly search for a good configuration at runtime. Instead of hand-crafting an analytical model that requires expert insights into low-level hardware details, we employ machine learning techniques to automatically learn it. We achieve this by first learning a predictive model offline using training programs. The learned model can then be used to predict the performance of any unseen program at runtime. We apply our approach to 39 representative parallel applications and evaluate it on two representative heterogeneous many-core platforms: a CPU-XeonPhi platform and a CPU-GPU platform. Compared to the single-stream version, our approach achieves, on average, a 1.6x and 1.1x speedup on the XeonPhi and the GPU platform, respectively. These results translate to over 93 percent of the performance delivered by a theoretically perfect predictor. Peng Zhang 0061, Jianbin Fang, Canqun Yang, Chun Huang 0006, Tao Tang 0001, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Non-Ergodic Convergence Analysis of Heavy-Ball AlgorithmsabstractIn this paper, we revisit the convergence of the Heavy-ball method, and present improved convergence complexity results in the convex setting. We provide the first non-ergodic O(1/k) rate result of the Heavy-ball algorithm with constant step size for coercive objective functions. For objective functions satisfying a relaxed strongly convex condition, the linear convergence is established under weaker assumptions on the step size and inertial parameter than made in the existing literature. We extend our results to multi-block version of the algorithm with both the cyclic and stochastic update rules. In addition, our results can also be extended to decentralized optimization, where the ergodic analysis is not applicable. Tao Sun 0005, Penghang Yin, Dongsheng Li 0001, Chun Huang 0006, Lei Guan 0001, Hao Jiang 0001 |
AAAI | 4 |
| 2019 | A Skewness-Aware Matrix Factorization Approach for Mesh-Structured Cloud ServicesabstractOnline cloud services need to fulfill clients' requests scalably and fast. State-of-the-art cloud services are increasingly deployed as a distributed service mesh. Service to service communication is frequent in the mesh. Unfortunately, problematic events may occur between any pair of nodes in the mesh, therefore, it is vital to maximize the network visibility. A state-of-the-art approach is to model pairwise RTTs based on a latent factor model represented as a low-rank matrix factorization. A latent factor corresponds to a rank-1 component in the factorization model, and is shared by all node pairs. However, different node pairs usually experience a skewed set of hidden factors, which should be fully considered in the model. In this paper, we propose a skewness-aware matrix factorization method named SMF. We decompose the matrix factorization into basic units of rank-one latent factors, and progressively combine rank-one factors for different node pairs. We present a unifying framework to automatically and adaptively select the rank-one factors for each node pair, which not only preserves the low rankness of the matrix model, but also adapts to skewed network latency distributions. Over real-world RTT data sets, SMF significantly improves the relative error by a factor of 0.2 x to 10 x, converges fast and stably, and compactly captures fine-grained local and global network latency structures. Yongquan Fu, Dongsheng Li 0001, Pere Barlet-Ros, Chun Huang 0006, Zhen Huang 0006, Huayou Su |
IEEE/ACM Trans. Netw. | 4 |
| 2018 | MOCL: an efficient openCL implementation for the matrix-2000 architectureabstractThis paper presents the design and implementation of an Open Computing Language (OpenCL) framework for the Matrix-2000 many-core architecture. This architecture is designed to replace the Intel XeonPhi accelerators of the TianHe-2 supercomputer. We share our experience and insights on how to design an effective OpenCL system for this new hardware accelerator. We propose a set of new analysis and optimizations to unlock the potential of the hardware. We extensively evaluate our approach using a wide range of OpenCL benchmarks on a single and multiple computing nodes. We present our design choices and provide guidance how to optimize code on the new Matrix-2000 architecture. Peng Zhang 0061, Tao Tang 0001, Jianbin Fang, Chun Huang 0006, Canqun Yang, Zheng Wang 0001 |
CF | 4 |
| 2015 | Implementation of an Accurate and Efficient Compensated DGEMM for 64-bit ARMv8 Multi-Core ProcessorsabstractThis paper presents an implementation of an accurate and efficient compensated Double-precision General Matrix Multiplication (DGEMM) based on OpenBLAS for 64-bit ARMv8 multi-core processors. Due to cancellation phenomena in floating point arithmetic, the results of DGEMM may not be as accurate as expected. In order to increase the accuracy of DGEMM, we compensate the error introduced by its dot product kernel (GEBP) by applying an error-free transformation to rewrite the kernel in assembly language. We optimize the computations in the inner kernel through exploiting loop unrolling, instruction scheduling and software-implemented register rotation to exploit instruction level parallelism (ILP). We also conduct a priori error analysis of the derived CompDGEMM. Our compensated DGEMM is as accurate as the existing quadruple precision GEMM using MBLAS, but is up to 6.4x faster. Our parallel implementation achieves good performance and scalability under varying thread counts across a range of matrix sizes evaluated. Hao Jiang 0001, Feng Wang 0050, Kuan Li, Canqun Yang, Kejia Zhao, Chun Huang 0006 |
ICPADS | 6 |
| 2015 | Poster: Symbolic Execution of MPI ProgramsabstractMPI is widely used in high performance computing. In this extended abstract, we report our current status of analyzing MPI programs. Our method can provide coverage of both input and non-determinism for MPI programs with mixed blocking and non-blocking operations. In addition, to improve the scalability further, a deadlock-oriented guiding method for symbolic execution is proposed. We have implemented our methods, and the preliminary experimental results are promising. Xianjin Fu, Zhenbang Chen 0001, Hengbiao Yu, Chun Huang 0006, Wei Dong 0006, Ji Wang 0001 |
ICSE (2) | 4 |
| 2014 | Synchronization Error Detection of MPI Programs by Symbolic ExecutionabstractAsynchrony based overlapping of computation and communication is commonly used in MPI applications. However, this overlapping introduces synchronization errors frequently in asynchronous MPI programming. In this paper, we propose a symbolic execution based method for detecting input-related synchronization errors. The path space of an MPI program is systematically explored, and the related operations of the synchronization errors in the program are checked specifically. In addition, two optimizations are proposed to improve the efficiency. We have implemented our method as a prototype tool based on the symbolic executor Cloud9. The results of the extensive experiments indicate the effectiveness of our method. Xianjin Fu, Zhenbang Chen 0001, Chun Huang 0006, Wei Dong 0006, Ji Wang 0001 |
APSEC (1) | 3 |