VLDB 2026 Research / reviewers in the wild / expert
Tao Tang 0001
dblp:35/1524-1
· DBLP profile ↗
46ranked-venue papers
3as first author
23since 2021 · last 2026
0009-0009-2883-6997ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 8 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Compensated Estrin Scheme for Accurate Polynomial Evaluation
Guangping Yu, Stef Graillat, Hao Jiang 0001, Chun Huang 0006, Tao Tang 0001 |
CASC | 5 |
| 2026 | Optimizing Long-Read Sequence Alignment on a CPU-DSPs Heterogeneous Processor
Xinjie An, Yifei Guo, Tao Tang 0001, Canqun Yang, Xiangke Liao, Yingbo Cui 0001 |
IEEE Trans. Computers | 4 |
| 2025 | SVScope: Structural Variation Detection for Short Reads via Multi-Source Fusion and Visual FilteringabstractStructural variations (SVs) are one of the major sources of genomic diversity and are closely associated with human disease. Existing short-read-based SV detection tools often rely on limited alignment features, which restricts their ability to fully capture variation signals. While many multi-source fusion methods have improved recall rates, they also introduce a large number of false positives. To address these issues, we present SVScope, an integrated SV detection tool that fuses multi-source signals, structure-sensitive image encoding, and deep visual filtering. It first consolidates candidate variations derived from complementary alignment signals across tools into a standardized candidate set via a harmonized detection and merging pipeline. To further improve specificity, SVScope utilizes a dynamic window to extract alignment features from candidate regions, then encodes them into a seven-channel image with CIGAR operations and read pair orientations to capture local structural details. It also combines a convolutional neural network with an embedded attention mechanism to enhance effective signals and suppress redundant noise, thereby achieving precise filtering of false positives. Benchmarking on real datasets demonstrates that SVScope consistently improves recall compared to individual detection tools while substantially enhancing overall precision through deep learning-based filtering. These results highlight SVScope's capacity to balance sensitivity and specificity, offering a precise and robust solution for SV analysis in large-scale short-read sequencing studies. The code and documentation of SVScope are publicly available at https://github.com/nudt-bioinfo/SVScope Weiming Xiang 0003, Tao Tang 0001, Yingbo Cui 0001 |
BIBM | 5 |
| 2025 | GESA: A Transformer-CNN Hybrid Framework for Sequence-to-Graph Alignment in Highly Divergent Genomic RegionsabstractModern genomics faces challenges from “reference bias” in linear genomes, prompting the adoption of pangenomic graphs to integrate multi-allelic variations. Sequence-to-graph alignment is a fundermental procedure in many pangenomic analyses. However, the alignment in complex topologies like cyclic graphs and highly polymorphic regions remains difficult due to path branch explosion and computational complexity. In this paper, we propose GESA, a sequence-to-graph alignment framework for sequences in highly divergent genomic regions. GESA adopts a hybrid strategy integrating haplotype-guided path linearization to organize topological information, thereby reducing information loss and potential path branch explosion. It employs a Transformer-CNN contrastive learning strategy to further capture global and local genomic features, enabling the identification of genetic characteristics in complex regions across the entire genome. Finally, a hierarchical vector-space retrieval technique is used to simplify the complex graph alignment computation into linear alignments on multiple sequences through vector similarity retrieval algorithms. GESA achieves an alignment ratio of 0.79 in cyclic graphs within the complex MHC region, outperforming Minigraph and GraphAligner by$4.3 \times$and$3.3 \times$, respectively. GESA lays a foundation for the future development of deep learning model applications in the field of pangenome graph alignment. The GESA code is available at https://github.com/nudt-bioinfo/GESA. Chenchen Peng, Canqun Yang, Yifei Guo, Tao Tang 0001, Yingbo Cui 0001 |
BIBM | 5 |
| 2025 | FusionSVFilter: A Deep-Learning Based Fast Structural Variation Filtering Tool for Long ReadsabstractStructural variations (SVs) play a critical role in species diversity, biological evolution, and human diseases. Although third-generation sequencing technology has enhanced the ability to detect long and complex SVs through long reads that can span complex genomic regions, it still suffers from high false-positive (FP) SV calls due to three factors: the complexity of SVs, limitations of detection algorithms, and relatively high single-base error rates in long reads. To address this challenge, we propose FusionSVFilter, a deep learning-based algorithm for SV filtering. It first transforms genomic sequence features into multi-level grayscale images via sequence-to-image encoding, which captures the structural complexity of SVs. Moreover, the encoder is accelerated with parallel computing to improve speed. Afterward, these grayscale images are enhanced to retain key features and converted into RGB format to meet the input requirements of the model. In addition, FusionSVFilter further leverages transfer learning, initializing with a pre-trained ResNet50 model that is fine-tuned on our curated dataset to recognize SV-specific patterns. Experimental results show that FusionSVFilter sub-stantially reduces FP calls while maintaining true-positive (TP) detection at near-constant levels. The code and documentation are available at: https://github.com/nudt-bioinfo/FusionSVFilter. Xinghai Zeng, Tao Tang 0001, Shijie Li 0002, Yingbo Cui 0001 |
BIBM | 2 |
| 2025 | Selection of Supervised Learning-Based Sparse Matrix Reordering Algorithms
Tao Tang 0001, Youfu Jiang, Yingbo Cui 0001, Jianbin Fang, Peng Zhang 0061, Lin Peng 0001, Chun Huang 0006 |
HiPC | 1 |
| 2025 | MMF-SV: A Multi-Modal Feature Fusion-Based Structural Variant CallerabstractStructural variant (SV) calling plays a critical role in understanding genome diversity and disease mechanisms. Although deep learning techniques have been increasingly applied to SV identification, existing general-purpose models still face significant challenges, including incomplete extraction of alignment signals, limited accuracy and efficiency, and poor performance in highly polymorphic or structurally complex genomic regions. These limitations lead to suboptimal detection accuracy in current SV callers. In this work, we present MMF-SV, a multi-modal feature fusion-based model (MMF) for SV calling. MMF-SV integrates matching patterns and statistical information from CIGAR signals with textual features extracted from alignment information, enabling comprehensive representation of diverse SV signals. We trained MMF-SV using CLIP, and the trained model achieved over 96% F1 score for classifying various types of variations. We validated the stability and robustness of the MMF-SV model through 5-fold cross-validation. Compared to existing long-read SV callers, MMF-SV achieves higher accuracy and can be effectively integrated with them to significantly reduce the number of false positives in the calling results. Canqun Yang, Haoang Chi, Tao Tang 0001, Weiming Xiang 0003, Yingbo Cui 0001 |
ACM Multimedia | 4 |
| 2025 | Fast noisy long read alignment with multi-level parallelismabstractBACKGROUND: The advent of Single Molecule Real-Time (SMRT) sequencing has overcome many limitations of second-generation sequencing, such as limited read lengths, PCR amplification biases. However, longer reads increase data volume exponentially and high error rates make many existing alignment tools inapplicable. Additionally, a single CPU's performance bottleneck restricts the effectiveness of alignment algorithms for SMRT sequencing. RESULTS: To address these challenges, we introduce ParaHAT, a parallel alignment algorithm for noisy long reads. ParaHAT utilizes vector-level, thread-level, process-level, and heterogeneous parallelism. We redesign the dynamic programming matrices layouts to eliminate data dependency in the base-level alignment, enabling effective vectorization. We further enhance computational speed through heterogeneous parallel technology and implement the algorithm for multi-node computing using MPI, overcoming the computational limits of a single node. CONCLUSIONS: Performance evaluations show that ParaHAT got a 10.03x speedup in base-level alignment, with a parallel acceleration ratio and weak scalability metric of 94.61 and 98.98% on 128 nodes, respectively. Canqun Yang, Chenchen Peng, Yifei Guo, Tao Tang 0001, Yingbo Cui 0001 |
BMC Bioinform. | 6 |
| 2025 | An empirical performance evaluation of SYCL on ARM multi-core processors
Hanzheng Liang, Chencheng Deng, Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Chun Huang 0006 |
CCF Trans. High Perform. Comput. | 5 |
| 2025 | nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUsabstractConvolution kernels are widely seen in high-performance computing (HPC) and deep learning (DL) workloads and are often responsible for performance bottlenecks. Prior works have demonstrated that the direct convolution approach can outperform the conventional convolution implementation. Although well-studied, the existing approaches for direct convolution are either incompatible with the mainstream DL data layouts or lead to suboptimal performance. We designnDirect2, a novel direct convolution approach that targets multi-core CPUs commonly found in smartphones and HPC systems.nDirect2is compatible with the data layout formats used by mainstream DL frameworks and offers new optimizations for the computational kernel, data packing, advanced operator fusion, and parallelization. We evaluatenDirect2by applying it to representative convolution kernels and demonstrating how well it performs on four distinct ARM-based CPUs and an X86-based CPU. Experimental results show thatnDirect2outperforms four state-of-the-art convolution approaches across most evaluation cases and hardware architectures. Weiling Yang, Jianbin Fang, Dezun Dong, Zhengbin Pang, Runxi He, Peng Zhang 0061, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Jie Ren 0007 |
IEEE Trans. Computers | 8 |
| 2025 | PVGwfa: a multi-level parallel sequence-to-graph alignment algorithm
Chenchen Peng, Shengbo Tang, Yifei Guo, Canqun Yang, Tao Tang 0001, Yingbo Cui 0001 |
J. Supercomput. | 6 |
| 2024 | WFA-vect: a SIMD wavefront algorithm for gap-affine pairwise alignmentabstractSequence alignment is the core of many bioinformatics tasks such as read mapping, genome assembly, variant detection and so on. With the advent of the third generation sequencing, classical dynamic programming-based alignment algorithms face challenges in efficiently handling these long reads. To address this issue, we present WFA-vect, a SIMD-based fast sequence alignment algorithm based on WFA. In WFA-vect, we introduce load synchronous and mask-based branch strategies to make the algorithm more suitable for vectorization. The load synchronous equalizes the load across different vector units to facilitate vectorization. The mask-based branch uses branch masking to bypass branch, avoiding pipeline hazards. To avoid binding the SIMD algorithm to specific hardware, we design a universal vectorization framework, which allows researchers to quickly port WFA-vect to other platforms without needing to understand the details of the algorithm. WFA-vect attains a peak speedup of 3.87× and 3.98× for data with error rates of 1% and 20%, respectively, compared to the scalar algorithm, while maintaining the alignment result consistent. The code and documentation of WFA-vect are publicly available at https://github.com/nudt-bioinfo/WFA-vect. Yifei Guo, Tao Tang 0001, Qingzhe Wang, Canqun Yang, Chenchen Peng, Yingbo Cui 0001 |
BIBM | 2 |
| 2024 | VLASPH: Smoothed Particle Hydrodynamics on VLA SIMD Architectures
Xiaokang Fan, Zhen Ge, Tao Tang 0001, Chun Huang 0006, Lin Peng 0001, Canqun Yang |
Euro-Par (3) | 4 |
| 2024 | A Motion Trace Decomposition-based overset grid method for parallel CFD simulations with moving boundariesabstractThe overset grid method is widely employed to solve moving boundary problems in numerical simulations. However, the heavy and inevitable communication resulting from boundary movements severely impedes the improvement of parallel efficiency. This paper proposes a Motion Trace Decomposition (MTD) method to alleviate this issue. The MTD method minimizes communication overhead between processors by decomposing sub-grids and distributing them according to the object motion trajectory, negating the need to reproduce communication areas when boundaries move. Various tests were conducted to evaluate the MTD method, incorporating diverse motion types, such as displacement and rotation. Results from experimental simulations with 1.9 × 106 grid cells indicate that the proposed method enhances the parallel efficiency of the assembly process by up to 20.35% using 72 processors. These findings showcase the significant potential of the MTD method in alleviating communication challenges associated with simulating moving boundary problems using overset grids. Chao Li 0070, Xi Yang 0020, Tao Tang 0001, Canqun Yang |
ICPP | 6 |
| 2024 | Optimizing Stencil Computation on Multi-core DSPsabstractStencil is a common computation pattern in high-performance computing (HPC) applications. While extensive work has been proposed to optimize stencil kernels on CPUs and GPUs, there is no consensus on how to best optimize stencils on multi-core Digital Signal Processors (DSPs) used in emerging HPC systems. This paper shares our experience in optimizing stencil kernels on multi-core DSPs. Our approach combines coarse and fine-grained parallel optimization techniques to enhance the performance of stencil computations. Our optimizations include a vectorization-enabled micro-kernel to utilize instruction parallelism, a memory-aware data reuse strategy to maximize data locality across multiple memory levels and a triple-buffering mechanism to overlap computation and memory communications. Experimental results show that our approach can effectively utilize the memory bandwidth and the computation capability of the underlying hardware. Our integrated optimizations can yield a 3.72x speedup over the 16-core CPU counterpart. Fugeng Zhu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Yonggang Che, Kainan Yu, Jing Xie 0023, Chun Huang 0006, Jie Ren 0007 |
ICPP | 5 |
| 2024 | Optimizing General Matrix Multiplications on Modern Multi-core DSPsabstractGeneral Matrix Multiplication (GEMM) is a key subprogram in high-performance computing (HPC) and deep learning workloads. With the rising significance of power and energy consumption in HPC systems, accelerators based on Digital Signal Processors (DSPs) have been integrated into general-purpose HPC systems. Due to the architecture disparities, the GEMM optimization techniques used on conventional multi-core CPUs and GPGPUs are not always applicable to DSPs. This paper shares our experience in optimizing GEMM on multi-core GPDSPs, using a CPU-DSP processor as a case study. Our approach employs a range of techniques to optimize performance for DSP architectures. These include data partitioning, three-level pipelining, dedicated micro-kernel design, and improved vector reduction. These optimizations maximize the overlap between computation and communication while fully exploiting the capabilities of floating-point arithmetic units to achieve high performance. Our experimental results demonstrate that the performance attained by our optimization is up to 96% of the theoretical peak performance of the hardware. Kainan Yu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Dezun Dong, Ruibo Wang, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Zheng Wang 0001 |
IPDPS | 7 |
| 2024 | CSV-Filter: a deep learning-based comprehensive structural variant filtering method for both short and long readsabstractMOTIVATION: Structural variants (SVs) play an important role in genetic research and precision medicine. As existing SV detection methods usually contain a substantial number of false positive calls, approaches to filter the detection results are needed. RESULTS: We developed a novel deep learning-based SV filtering tool, CSV-Filter, for both short and long reads. CSV-Filter uses a novel multi-level grayscale image encoding method based on CIGAR strings of the alignment results and employs image augmentation techniques to improve SV feature extraction. CSV-Filter also utilizes self-supervised learning networks for transfer as classification models, and employs mixed-precision operations to accelerate training. The experiments showed that the integration of CSV-Filter with popular SV detection tools could considerably reduce false positive SVs for short and long reads, while maintaining true positive SVs almost unchanged. Compared with DeepSVFilter, a SV filtering tool for short reads, CSV-Filter could recognize more false positive calls and support long reads as an additional feature. AVAILABILITY AND IMPLEMENTATION: https://github.com/xzyschumacher/CSV-Filter. Weiming Xiang 0003, Qingzhe Wang, Xingze Li, Junyu Gao 0005, Tao Tang 0001, Canqun Yang, Yingbo Cui 0001 |
Bioinform. | 7 |
| 2024 | SNCL: a supernode OpenCL implementation for hybrid computing arrays
Tao Tang 0001, Kai Lu 0001, Lin Peng 0001, Yingbo Cui 0001, Jianbin Fang, Chun Huang 0006, Ruibo Wang, Canqun Yang, Yifei Guo |
J. Supercomput. | 1 |
| 2023 | Optimizing Direct Convolutions on ARM Multi-CoresabstractConvolution kernels are widely seen in deep learning workloads and are often responsible for performance bottlenecks. Recent research has demonstrated that a direct convolution approach can outperform the traditional convolution implementation based on tensor-to-matrix conversions. However, existing approaches for direct convolution still have room for performance improvement. We present nDirect, a new direct convolution approach that targets ARM-based multi-core CPUs commonly found in smartphones and HPC systems. nDirect is designed to be compatible with the data layout formats used by mainstream deep learning frameworks but offers new optimizations for the computational kernel, data packing, and parallelization. We evaluate nDirect by applying it to representative convolution kernels and demonstrating its performance on four distinct ARM multi-core CPU platforms. We compare nDirect against state-of-the-art convolution optimization techniques. Experimental results show that nDirect gives the best overall performance across evaluation scenarios and platforms. Weiling Yang, Jianbin Fang, Dezun Dong, Chun Huang 0006, Peng Zhang 0061, Tao Tang 0001, Zheng Wang 0001 |
SC | 7 |
| 2023 | Programming bare-metal accelerators with heterogeneous threading models: a case study of Matrix-3000abstractAs the hardware industry moves toward using specialized heterogeneous many-core processors to avoid the effects of the power wall, software developers are finding it hard to deal with the complexity of these systems. In this paper, we share our experience of developing a programming model and its supporting compiler and libraries for Matrix-3000, which is designed for next-generation exascale supercomputers but has a complex memory hierarchy and processor organization. To assist its software development, we have developed a software stack from scratch that includes a low-level programming interface and a high-level OpenCL compiler. Our low-level programming model offers native programming support for using the bare-metal accelerators of Matrix-3000, while the high-level model allows programmers to use the OpenCL programming standard. We detail our design choices and highlight the lessons learned from developing system software to enable the programming of bare-metal accelerators. Our programming models have been deployed in the production environment of an exascale prototype system. Jianbin Fang, Peng Zhang 0061, Chun Huang 0006, Tao Tang 0001, Kai Lu 0001, Ruibo Wang, Zheng Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2022 | MT-3000: a heterogeneous multi-zone processor for HPC
Kai Lu 0001, Yang Guo 0003, Chun Huang 0006, Sheng Liu 0001, Ruibo Wang, Jianbin Fang, Tao Tang 0001, Zhaoyun Chen, Biwei Liu, Zhong Liu 0003, Yuanwu Lei, Haiyan Sun |
CCF Trans. High Perform. Comput. | 8 |
| 2021 | Large-Scale Parallel Alignment Algorithm for SMRT Reads
Yingbo Cui 0001, Peng Zhang 0061, Tao Tang 0001, Lin Peng 0001, Chun Huang 0006, Canqun Yang, Xiangke Liao |
ICA3PP (2) | 6 |
| 2021 | VISPR-online: a web-based interactive tool to visualize CRISPR screening experimentsabstractBACKGROUND: VISPR is an interactive visualization and analysis framework for CRISPR screening experiments. However, it only supports the output of MAGeCK, and requires installation and manual configuration. Furthermore, VISPR is designed to run on a single computer, and data sharing between collaborators is challenging. RESULTS: To make the tool easily accessible to the community, we present VISPR-online, a web-based general application allowing users to visualize, explore, and share CRISPR screening data online with a few simple steps. VISPR-online provides an exploration of screening results and visualization of read count changes. Apart from MAGeCK, VISPR-online supports two more popular CRISPR screening analysis tools: BAGEL and JACKS. It provides an interactive environment for exploring gene essentiality, viewing guide RNA (gRNA) locations, and allowing users to resume and share screening results. CONCLUSIONS: VISPR-online allows users to visualize, explore and share CRISPR screening data online. It is freely available at http://vispr-online.weililab.org , while the source code is available at https://github.com/lemoncyb/VISPR-online . Yingbo Cui 0001, Johannes Köster, Xiangke Liao, Shaoliang Peng, Tao Tang 0001, Chun Huang 0006, Canqun Yang |
BMC Bioinform. | 6 |
| 2020 | Parallel programming models for heterogeneous many-cores: a comprehensive survey
Jianbin Fang, Chun Huang 0006, Tao Tang 0001, Zheng Wang 0001 |
CCF Trans. High Perform. Comput. | 3 |
| 2020 | clMF: A fine-grained and portable alternating least squares algorithm for parallel matrix factorization
Jing Chen 0038, Jianbin Fang, Weifeng Liu 0002, Tao Tang 0001, Canqun Yang |
Future Gener. Comput. Syst. | 4 |
| 2020 | Optimizing Streaming Parallelism on Heterogeneous Many-Core ArchitecturesabstractAs many-core accelerators keep integrating more processing units, it becomes increasingly more difficult for a parallel application to make effective use of all available resources. An effective way of improving hardware utilization is to exploit spatial and temporal sharing of the heterogeneous processing units by multiplexing computation and communication tasks - a strategy known as heterogeneous streaming. Achieving effective heterogeneous streaming requires carefully partitioning hardware among tasks, and matching the granularity of task parallelism to the resource partition. However, finding the right resource partitioning and task granularity is extremely challenging, because there is a large number of possible solutions and the optimal solution varies across programs and datasets. This article presents an automatic approach to quickly derive a good solution for hardware resource partition and task granularity for task-based parallel applications on heterogeneous many-core architectures. Our approach employs a performance model to estimate the resulting performance of the target application under a given resource partition and task granularity configuration. The model is used as a utility to quickly search for a good configuration at runtime. Instead of hand-crafting an analytical model that requires expert insights into low-level hardware details, we employ machine learning techniques to automatically learn it. We achieve this by first learning a predictive model offline using training programs. The learned model can then be used to predict the performance of any unseen program at runtime. We apply our approach to 39 representative parallel applications and evaluate it on two representative heterogeneous many-core platforms: a CPU-XeonPhi platform and a CPU-GPU platform. Compared to the single-stream version, our approach achieves, on average, a 1.6x and 1.1x speedup on the XeonPhi and the GPU platform, respectively. These results translate to over 93 percent of the performance delivered by a theoretically perfect predictor. Peng Zhang 0061, Jianbin Fang, Canqun Yang, Chun Huang 0006, Tao Tang 0001, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | MOCL: an efficient openCL implementation for the matrix-2000 architectureabstractThis paper presents the design and implementation of an Open Computing Language (OpenCL) framework for the Matrix-2000 many-core architecture. This architecture is designed to replace the Intel XeonPhi accelerators of the TianHe-2 supercomputer. We share our experience and insights on how to design an effective OpenCL system for this new hardware accelerator. We propose a set of new analysis and optimizations to unlock the potential of the hardware. We extensively evaluate our approach using a wide range of OpenCL benchmarks on a single and multiple computing nodes. We present our design choices and provide guidance how to optimize code on the new Matrix-2000 architecture. Peng Zhang 0061, Tao Tang 0001, Jianbin Fang, Chun Huang 0006, Canqun Yang, Zheng Wang 0001 |
CF | 2 |
| 2018 | Auto-tuning Streamed Applications on Intel Xeon PhiabstractMany-core accelerators, as represented by the XeonPhi coprocessors and GPGPUs, allow software to exploit spatial and temporal sharing of computing resources to improve the overall system performance. To unlock this performance potential requires software to effectively partition the hardware resource to maximize the overlap between host-device communication and accelerator computation, and to match the granularity of task parallelism to the resource partition. However, determining the right resource partition and task parallelism on a per program, per dataset basis is challenging. This is because the number of possible solutions is huge, and the benefit of choosing the right solution may be large, but mistakes can seriously hurt the performance. In this paper, we present an automatic approach to determine the hardware resource partition and the task granularity for any given streamed application, targeting the Intel XeonPhi architecture. Instead of hand-crafting the heuristic for which the process will have to repeat for each hardware generation, we employ machine learning techniques to automatically learn it. We achieve this by first learning a predictive model offline using training programs; we then use the learned model to predict the resource partition and task granularity for any unseen programs at runtime. We apply our approach to 23 representative parallel applications and evaluate it on a CPU-XeonPhi mixed heterogenous many-core platform. Our approach achieves, on average, a 1.6x (upto 5.6x) speedup, which translates to 94.5% of the performance delivered by a theoretically perfect predictor. Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Canqun Yang, Zheng Wang 0001 |
IPDPS | 3 |
| 2018 | Orchestrating parallel detection of strongly connected components on GPUs
Xuhao Chen 0001, Cheng Chen 0005, Jie Shen 0003, Jianbin Fang, Tao Tang 0001, Canqun Yang, Zhiying Wang 0003 |
Parallel Comput. | 5 |
| 2017 | Efficient and high-quality sparse graph coloring on GPUsabstractSummary Graph coloring has been broadly used to discover concurrency in parallel computing. To speed up graph coloring for large‐scale datasets, parallel algorithms have been proposed to leverage modern GPUs. Existing GPU implementations either have limited performance or yield unsatisfactory coloring quality (too many colors assigned). We present a work‐efficient parallel graph coloring implementation on GPUs with good coloring quality. Our approach uses the speculative greedy scheme, which inherently yields better quality than the method of finding maximal independent set . To achieve high performance on GPUs, we refine the algorithm to leverage efficient operators and alleviate conflicts. We also incorporate common optimization techniques to further improve performance. Our method is evaluated with both synthetic and real‐world sparse graphs on the NVIDIA GPU. Experimental results show that our proposed implementation achieves averaged 4.1 × (up to 8.9 × ) speedup over the serial implementation. It also outperforms the existing GPU implementation from the NVIDIA CUSPARSE library (2.2 × average speedup), while yielding much better coloring quality than CUSPARSE. Xuhao Chen 0001, Pingfan Li, Jianbin Fang, Tao Tang 0001, Zhiying Wang 0003, Canqun Yang |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | An Energy-Efficient Implementation of LU Factorization on Heterogeneous SystemsabstractEnergy consumption is increasingly becoming a critical issue in HPC. There is a broad consensus that future exascale-computing will be strongly constrained by energy consumption. Heterogeneous systems usually feature higher energy efficiency than homogeneous ones since the former employ coprocessors that provide higher GFlops/Watt than CPUs. Thus, it is of great importance to better utilize the coprocessors from an energy-efficiency standpoint. Dense LU factorization (LU) is a critical kernel that is widely used to solve dense linear algebra problems. However, existingheterogeneous implementations are typically designed to be CPU-centered, which rely highly on CPUs and thus suffer from large data transfer overheads via PCIe, hurting the energy efficiency of the entire computer system. We present a coprocessor-resident implementation of LU for a heterogeneous platform to improve energy efficiency without impeding performance by relieving the CPUs from performing unnecessary computations and reducing excessive data transfers via PCIe. In addition, several optimizations are judiciously employed to overlap the computation and communication between the CPUs and coprocessors. Validation on the Tianhe-2 supercomputer shows that our LU implementation gains higher performance, achieves higher energy efficiency, and features a better scalability than Intel MKL. Canqun Yang, Cheng Chen 0005, Tao Tang 0001, Xuhao Chen 0001, Jianbin Fang, Jingling Xue |
ICPADS | 3 |
| 2016 | Streaming Applications on Heterogeneous Platforms
Zhaokui Li, Jianbin Fang, Tao Tang 0001, Xuhao Chen 0001, Canqun Yang |
NPC | 3 |
| 2014 | OpenMC: Towards Simplifying Programming for TianHe Supercomputers
Xiangke Liao, Canqun Yang, Tao Tang 0001, Huizhan Yi, Feng Wang 0050, Jingling Xue |
J. Comput. Sci. Technol. | 3 |
| 2013 | Exploiting hierarchy parallelism for molecular dynamics on a petascale heterogeneous system
Canqun Yang, Tao Tang 0001, Liquan Xiao |
J. Parallel Distributed Comput. | 3 |
| 2012 | MPtostream: an OpenMP compiler for CPU-GPU heterogeneous parallel systems
Xuejun Yang, Tao Tang 0001, Guibin Wang, Jia Jia 0004, Xinhai Xu |
Sci. China Inf. Sci. | 2 |
| 2011 | Cache Miss Analysis for GPU Programs Based on Stack Distance ProfileabstractUsing the graphics processing unit (GPU) to accelerate the general purpose computation has attracted much attention from both the academia and industry due to GPU's powerful computing capacity. Thus optimization of GPU programs has become a popular research direction. In order to support the general purpose computing more efficiently, GPU has integrated the general data cache to replace the existing software-managed on-chip memory. Consequently, improving the usage of the data cache becomes of vital importance to improve the performance of the GPU programs. The foundation of cache locality optimizations is efficient analysis and prediction of the cache behavior. Unfortunately, existing cache miss analysis models are based on sequential programs and thus cannot be used to analyze the GPU programs directly. In this paper, based on the deep analysis of GPU's execution model, we propose, for the first time, a cache miss analysis model for the GPU programs. We divide the problem into two subproblems: stack distance profile analysis of single thread block and cache contention analysis of multiple thread blocks. The experimental results from nine typical application kernels in the scientific computing field illustrate that our method is efficient and can be used to guide the cache locality optimizations for the GPU programs. Tao Tang 0001, Xuejun Yang, Yisong Lin |
ICDCS | 1 |
| 2011 | Power Optimization for GPU Programs Based on Software PrefetchingabstractGPUs render higher computing unit density than contemporary CPUs and thus exhibit much higher power consumption despite its higher power efficiency. The power consumption has become an important issue that impacts CPU's applications, thereby necessitating the low power optimization technology for GPUs. Software prefetching is an efficient way to alleviate the memory wall problem which overlaps the computing and memory access latencies. However, software prefetching will cause some power overhead because it increases the number and density of the instructions. Thus, we should consider the balance between the performance income and the power overhead when applying the optimization. To address this problem, in this paper we first analyze the multi-thread execution model of GPU and validate the potential space of software prefetching optimization. Then we give the software prefetching method for GPU programs to improve the performance. Aiming at two different objects: energy optimization under performance constraint and performance optimization under power constraint, we discuss the optimization methods based on software prefetching and dynamic voltage scaling technologies. The experimental results show that our method can efficiently optimize the energy consumption (performance) under the performance (power) constraint. Yisong Lin, Tao Tang 0001, Guibin Wang |
TrustCom | 2 |
| 2010 | Improving scratchpad allocation with demand-driven data tilingabstractExisting scratchpad memory (SPM) allocation algorithms for arrays, whether they rely on well-crafted heuristics or resort to integer linear programming (ILP) techniques, typically assume that every array is small enough to fit directly into the SPM. As a result, some arrays have to be spilled entirely to the off-chip memory in order to make room for other arrays to stay in the SPM, resulting in sometimes poor SPM utilization. Xuejun Yang, Li Wang 0027, Jingling Xue, Tao Tang 0001, Xiaoguang Ren, Sen Ye |
CASES | 4 |
| 2009 | Program Optimization of Array-Intensive SPEC2k Benchmarks on Multithreaded GPU Using CUDA and Brook+abstractGraphic Processing Unit (GPU), with many light-weight data-parallel cores, can provide substantial parallel computing power to accelerate several general purpose applications. Both the AMD and NVIDIA corps provide their specific high performance GPUs and software platforms. As the floating-point computing capacity increases continually, the problem of ``memory-wall'' becomes more serious, especially for array-intensive applications. In this paper, we optimize and implement two SPEC2k benchmarks mgrid and swim on multithreaded GPU using CUDA and Brook+. In order to reduce the pressure on off-chip memory, we make use of data locality in multi-level memory hierarchies and hide long memory access latency via double-buffers. To balance inter-thread parallelism and intra-thread locality, we further tune thread granularity for each kernel and empirically study the best equilibrium point for this problem. Flow control instruction can significantly impact the effective instruction throughput. Oriented to this problem, we introduce a diverge elimination technology to convert condition expression into computing operation. Through all the optimizations, we gain the speedup of 10×-34× to the CPU implementation on the GPUs of AMD and NVIDIA respectively. Finally, we summarize and compares the GPUs from AMD and NVIDIA in hardware and software. Guibin Wang, Tao Tang 0001, Xudong Fang, Xiaoguang Ren |
ICPADS | 2 |
| 2009 | Program Optimization of Stencil Based Application on the GPU-Accelerated SystemabstractGraphic Processing Unit (GPU), with many light-weight data-parallel cores, can provide substantial parallel computational power to accelerate general purpose applications. But the powerful computing capacity could not be fully utilized for memory-intensive applications, which are limited by off-chip memory bandwidth and latency. Stencil computation has abundant parallelism and low computational intensity which make it a useful architectural evaluation benchmark. In this paper, we propose some memory optimizations for a stencil based application mgrid from SPEC 2K benchmarks. Through exploiting data locality in 3-level memory hierarchies and tuning the thread granularity, we reduce the pressure on the off-chip memory bandwidth. To hide the long off-chip memory access latency, we further prefetch data during computation through double-buffer. In order to fully exploit the CPU-GPU heterogeneous system, we redistribute the computation between these two computing resource. Through all these optimizations, we gain 24.2x speedup compared to the simple mapping version, and get as high as 34.3x speedup when compared with a CPU implementation. Guibin Wang, Xuejun Yang, Ying Zhang 0032, Tao Tang 0001, Xudong Fang |
ISPA | 4 |
| 2009 | SRF Coloring: Stream Register File Allocation via Graph Coloring
Xuejun Yang, Yu Deng 0001, Li Wang 0027, Xiaobo Yan, Jing Du 0002, Ying Zhang 0032, Guibin Wang, Tao Tang 0001 |
J. Comput. Sci. Technol. | 8 |
| 2008 | Optimizing scientific application loops on stream processorsabstractThis paper describes a graph coloring compiler framework to allocate on-chip SRF(Stream Register File) storage for optimizing scientific applications on stream processors. Our framework consists of first applying enabling optimizations such as loop unrolling to expose stream reuse and opportunities for maximizing parallelism, i.e., overlapping kernel execution and memory transfers.Then the three SRF management tasks are solved in a unified manner via graph coloring: (1) placing streams in the SRF, (2) exploiting stream use, and (3) maximizing parallelism. We evaluate the performance of our compiler framework by actually running nine representative scientific computing kernels on our FT64 stream processor. Our preliminary results show that compiler management achieves an average speedup of 2.3x compared to First-Fit allocation. In comparison with the performance results obtained from running these benchmarks on Itanium 2, an average speedup of 2.1x is observed. Li Wang 0027, Xuejun Yang, Jingling Xue, Yu Deng 0001, Xiaobo Yan, Tao Tang 0001, Quan Hoang Nguyen 0001 |
LCTES | 6 |
| 2007 | Implementation and Evaluation of Jacobi Iteration on the Imagine Stream Processor
Jing Du 0002, Xuejun Yang, Tao Tang 0001, Guibin Wang |
HiPC | 4 |
| 2007 | Evaluation of Transcendental Functions on Imagine ArchitectureabstractThe fast and accurate evaluation of transcendental functions (e.g. exp, log, sin, and atan) is quite important in many domains. We implement a software inline function library that can be called from KernelC programming language to compute 8 typical functions on Imagine architecture. By exploiting some of the key features of Imagine architecture, we have been able to provide single precision transcendental functions that are very accurate yet can typically be evaluated to get 16 function values in between 18 and 43 clock cycles. In this paper, we also discuss the algorithms and implementation details of these functions. Xiaobo Yan, Tao Tang 0001, Yu Deng 0001, Jing Du 0002, Xuejun Yang |
ICPP | 2 |
| 2007 | Architecture-Based Optimization for Mapping Scientific Applications to Imagine
Jing Du 0002, Xuejun Yang, Guibin Wang, Tao Tang 0001 |
ISPA | 4 |
| 2007 | Implementation and Optimization of Sparse Matrix-Vector Multiplication on Imagine Stream Processor
Li Wang 0027, Xuejun Yang, Guibin Wang, Xiaobo Yan, Yu Deng 0001, Jing Du 0002, Ying Zhang 0032, Tao Tang 0001 |
ISPA | 8 |