VLDB 2026 Research / reviewers in the wild / expert
Canqun Yang
dblp:92/2942
· DBLP profile ↗
74ranked-venue papers
5as first author
34since 2021 · last 2026
0009-0008-4757-2475ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 7 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DFS-PINN: A Dynamic Feature Separation Physics-Informed Neural Network
Zhuo Zhang 0020, Wei Wang 0130, Hongzhou Wu, Xi Yang 0020, Canqun Yang |
Comput. Aided Des. | 7 |
| 2026 | Legend-KINN: A legendre polynomial-based Kolmogorov-Arnold-informed neural network for efficient PDE solving
Zhuo Zhang 0020, Wei Wang 0130, Yanxu Zhong, Canqun Yang, Xi Yang 0020 |
Expert Syst. Appl. | 6 |
| 2026 | Dual-Gradient Co-optimization for Robust Physics-Informed Neural Networks
Seongwon Kang, Xi Yang 0020, Dongseok Kim, Yifu Gao, Canqun Yang, Chongam Kim |
Knowl. Based Syst. | 7 |
| 2026 | Optimizing Long-Read Sequence Alignment on a CPU-DSPs Heterogeneous Processor
Xinjie An, Yifei Guo, Tao Tang 0001, Canqun Yang, Xiangke Liao, Yingbo Cui 0001 |
IEEE Trans. Computers | 5 |
| 2026 | Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and SegmentationabstractReferring multi-object tracking (RMOT) is an emerging cross-modal task that aims to locate an arbitrary number of target objects and maintain their identities referred by a language expression in a video. This intricate task involves the reasoning of linguistic and visual modalities, along with the temporal association of target objects. However, the seminal work relies on loose feature fusion and neglects long-term information. In this study, we introduce a compact Transformer-based method, termed TenRMOT. We conduct feature fusion at both encoding and decoding stages to fully exploit the advantages of Transformer architecture. Specifically, we incrementally perform cross-modal fusion layer-by-layer during the encoding phase. In the decoding phase, we utilize language-guided queries to probe memory features for accurate prediction of the desired objects. Moreover, we introduce a query update module that explicitly leverages temporal prior information of the tracked objects to enhance the consistency of their trajectories. In addition, we introduce a novel task called Referring Multi-Object Tracking and Segmentation (RMOTS) and construct a new dataset named Ref-KITTI Segmentation. Our dataset consists of 18 videos with 818 expressions, and each expression averages 10.7 masks, which poses a greater challenge compared to the typical single mask in most existing referring video segmentation datasets. TenRMOT demonstrates superior performance on both the referring multi-object tracking and the segmentation tasks. Changcheng Xiao, Qiong Cao, Xiang Zhang 0008, Tao Wang 0006, Canqun Yang, Long Lan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | GESA: A Transformer-CNN Hybrid Framework for Sequence-to-Graph Alignment in Highly Divergent Genomic RegionsabstractModern genomics faces challenges from “reference bias” in linear genomes, prompting the adoption of pangenomic graphs to integrate multi-allelic variations. Sequence-to-graph alignment is a fundermental procedure in many pangenomic analyses. However, the alignment in complex topologies like cyclic graphs and highly polymorphic regions remains difficult due to path branch explosion and computational complexity. In this paper, we propose GESA, a sequence-to-graph alignment framework for sequences in highly divergent genomic regions. GESA adopts a hybrid strategy integrating haplotype-guided path linearization to organize topological information, thereby reducing information loss and potential path branch explosion. It employs a Transformer-CNN contrastive learning strategy to further capture global and local genomic features, enabling the identification of genetic characteristics in complex regions across the entire genome. Finally, a hierarchical vector-space retrieval technique is used to simplify the complex graph alignment computation into linear alignments on multiple sequences through vector similarity retrieval algorithms. GESA achieves an alignment ratio of 0.79 in cyclic graphs within the complex MHC region, outperforming Minigraph and GraphAligner by$4.3 \times$and$3.3 \times$, respectively. GESA lays a foundation for the future development of deep learning model applications in the field of pangenome graph alignment. The GESA code is available at https://github.com/nudt-bioinfo/GESA. Chenchen Peng, Canqun Yang, Yifei Guo, Tao Tang 0001, Yingbo Cui 0001 |
BIBM | 2 |
| 2025 | UP-Bern: A Unified Progressive Transition Framework for Biomedical Entity Recognition and NormalizationabstractAccurate recognition of biomedical entities (e.g., diseases, drugs, proteins, and genes) from literature, along with their normalization using standardized biomedical vocabularies, is crucial for facilitating various downstream tasks and boosting significant advancements in further biomedical research. However, most current studies still struggle with boundary inconsistencies stemming from separate decoding stages. Additionally, such models often neglect valuable semantic information encapsulated in the vocabulary's candidate concepts (e.g., surface forms of text), which is vital for effective entity normalization. In this paper, we propose a unified progressive transition framework named UP-Bern, which progressively constructs biomedical entity recognition and normalization outputs through an action sequence prediction process. This framework promotes joint modeling through shared input representations and concurrently optimizes outputs within a unified search space via state transitions. Moreover, we integrate an attention mechanism to fully harness the surface form information of candidate concepts, thereby enhancing the accuracy of entity normalization. We evaluated the proposed approach on four public biomedical datasets against 11 established methods, demonstrating consistent and notable improvements in performance. Canqun Yang, Siqi Wang 0001 |
ICDM | 2 |
| 2025 | MMF-SV: A Multi-Modal Feature Fusion-Based Structural Variant CallerabstractStructural variant (SV) calling plays a critical role in understanding genome diversity and disease mechanisms. Although deep learning techniques have been increasingly applied to SV identification, existing general-purpose models still face significant challenges, including incomplete extraction of alignment signals, limited accuracy and efficiency, and poor performance in highly polymorphic or structurally complex genomic regions. These limitations lead to suboptimal detection accuracy in current SV callers. In this work, we present MMF-SV, a multi-modal feature fusion-based model (MMF) for SV calling. MMF-SV integrates matching patterns and statistical information from CIGAR signals with textual features extracted from alignment information, enabling comprehensive representation of diverse SV signals. We trained MMF-SV using CLIP, and the trained model achieved over 96% F1 score for classifying various types of variations. We validated the stability and robustness of the MMF-SV model through 5-fold cross-validation. Compared to existing long-read SV callers, MMF-SV achieves higher accuracy and can be effectively integrated with them to significantly reduce the number of false positives in the calling results. Canqun Yang, Haoang Chi, Tao Tang 0001, Weiming Xiang 0003, Yingbo Cui 0001 |
ACM Multimedia | 2 |
| 2025 | Fast noisy long read alignment with multi-level parallelismabstractBACKGROUND: The advent of Single Molecule Real-Time (SMRT) sequencing has overcome many limitations of second-generation sequencing, such as limited read lengths, PCR amplification biases. However, longer reads increase data volume exponentially and high error rates make many existing alignment tools inapplicable. Additionally, a single CPU's performance bottleneck restricts the effectiveness of alignment algorithms for SMRT sequencing. RESULTS: To address these challenges, we introduce ParaHAT, a parallel alignment algorithm for noisy long reads. ParaHAT utilizes vector-level, thread-level, process-level, and heterogeneous parallelism. We redesign the dynamic programming matrices layouts to eliminate data dependency in the base-level alignment, enabling effective vectorization. We further enhance computational speed through heterogeneous parallel technology and implement the algorithm for multi-node computing using MPI, overcoming the computational limits of a single node. CONCLUSIONS: Performance evaluations show that ParaHAT got a 10.03x speedup in base-level alignment, with a parallel acceleration ratio and weak scalability metric of 94.61 and 98.98% on 128 nodes, respectively. Canqun Yang, Chenchen Peng, Yifei Guo, Tao Tang 0001, Yingbo Cui 0001 |
BMC Bioinform. | 2 |
| 2025 | GCPNet: An interpretable Generic Crystal Pattern graph neural Network for predicting material properties
Hengda Gao, Genglin Li, Chao Li 0070, Canqun Yang |
Neural Networks | 5 |
| 2025 | PVGwfa: a multi-level parallel sequence-to-graph alignment algorithm
Chenchen Peng, Shengbo Tang, Yifei Guo, Canqun Yang, Tao Tang 0001, Yingbo Cui 0001 |
J. Supercomput. | 5 |
| 2025 | Real-Time Anomaly Detection for Large-Scale Network DevicesabstractWith the booming of large-scale network devices, anomaly detection on multivariate time series (MTS), such as a combination of CPU utilization, average response time, and network packet loss, is important for system reliability. Although a collection of learning-based approaches have been designed for this purpose, our study shows that these approaches suffer from long initialization time for sufficient training data. Our previously proposed JumpStarter model stands as a MTS anomaly detection method characterized by its brief initialization time and commendable detection performance. However, it suffers from high computational cost and inappropriateness for periodic MTS. In this paper, we propose VersaGuardian, which introduces the Dynamic Mode Decomposition technique to MTS anomaly detection for diverse types of MTS in a rapidly initialized, computationally efficient manner. With real-world MTS datasets collected from three companies, our results show that VersaGuardian achieves an average F1 score of 94.42%, significantly outperforming the popular anomaly detection algorithms, with a much shorter initialization time of 20 minutes and detection time of 15.28 milliseconds. Shenglin Zhang, Junhua Kuang, Canqun Yang |
IEEE Trans. Netw. | 5 |
| 2024 | WFA-vect: a SIMD wavefront algorithm for gap-affine pairwise alignmentabstractSequence alignment is the core of many bioinformatics tasks such as read mapping, genome assembly, variant detection and so on. With the advent of the third generation sequencing, classical dynamic programming-based alignment algorithms face challenges in efficiently handling these long reads. To address this issue, we present WFA-vect, a SIMD-based fast sequence alignment algorithm based on WFA. In WFA-vect, we introduce load synchronous and mask-based branch strategies to make the algorithm more suitable for vectorization. The load synchronous equalizes the load across different vector units to facilitate vectorization. The mask-based branch uses branch masking to bypass branch, avoiding pipeline hazards. To avoid binding the SIMD algorithm to specific hardware, we design a universal vectorization framework, which allows researchers to quickly port WFA-vect to other platforms without needing to understand the details of the algorithm. WFA-vect attains a peak speedup of 3.87× and 3.98× for data with error rates of 1% and 20%, respectively, compared to the scalar algorithm, while maintaining the alignment result consistent. The code and documentation of WFA-vect are publicly available at https://github.com/nudt-bioinfo/WFA-vect. Yifei Guo, Tao Tang 0001, Qingzhe Wang, Canqun Yang, Chenchen Peng, Yingbo Cui 0001 |
BIBM | 4 |
| 2024 | VLASPH: Smoothed Particle Hydrodynamics on VLA SIMD Architectures
Xiaokang Fan, Zhen Ge, Tao Tang 0001, Chun Huang 0006, Lin Peng 0001, Canqun Yang |
Euro-Par (3) | 7 |
| 2024 | A Vectorized Sequence-to-Graph Alignment Algorithm
Chenchen Peng, Shengbo Tang, Yifei Guo, Canqun Yang, Yingbo Cui 0001 |
ICA3PP (1) | 5 |
| 2024 | A Motion Trace Decomposition-based overset grid method for parallel CFD simulations with moving boundariesabstractThe overset grid method is widely employed to solve moving boundary problems in numerical simulations. However, the heavy and inevitable communication resulting from boundary movements severely impedes the improvement of parallel efficiency. This paper proposes a Motion Trace Decomposition (MTD) method to alleviate this issue. The MTD method minimizes communication overhead between processors by decomposing sub-grids and distributing them according to the object motion trajectory, negating the need to reproduce communication areas when boundaries move. Various tests were conducted to evaluate the MTD method, incorporating diverse motion types, such as displacement and rotation. Results from experimental simulations with 1.9 × 106 grid cells indicate that the proposed method enhances the parallel efficiency of the assembly process by up to 20.35% using 72 processors. These findings showcase the significant potential of the MTD method in alleviating communication challenges associated with simulating moving boundary problems using overset grids. Chao Li 0070, Xi Yang 0020, Tao Tang 0001, Canqun Yang |
ICPP | 7 |
| 2024 | Giving Every Modality a Voice in Microservice Failure Diagnosis via Multimodal Adaptive OptimizationabstractMicroservice systems are inherently complex and prone to failures, which can significantly impact user experience. Existing diagnostic approaches based on single-modal data such as logs, metrics, or traces cannot comprehensively capture failure patterns. For those multimodal data-based failure diagnosis methods, the dominant modality can overshadow others, hindering low-yield modalities from fully leveraging their characteristics. This paper proposes Medicine, a modal-independent microservice failure diagnosis framework based on multimodal adaptive optimization. It encodes different modalities separately to retain their unique features and employs adaptive optimization to adjust the learning pace between modalities, thereby enhancing overall diagnostic performance. Experimental results demonstrate that Medicine outperforms existing single-modal and multimodal diagnostic approaches on three public datasets, with F1-score improving by 15.72% to 70.84%. Even in cases where individual modal data is missing or of lower quality, Medicine maintains high diagnostic accuracy. Shenglin Zhang, Zedong Jia, Jinrui Sun, Minghua Ma, Zhengdan Li, Yongqian Sun, Canqun Yang, Dan Pei |
ASE | 8 |
| 2024 | CSV-Filter: a deep learning-based comprehensive structural variant filtering method for both short and long readsabstractMOTIVATION: Structural variants (SVs) play an important role in genetic research and precision medicine. As existing SV detection methods usually contain a substantial number of false positive calls, approaches to filter the detection results are needed. RESULTS: We developed a novel deep learning-based SV filtering tool, CSV-Filter, for both short and long reads. CSV-Filter uses a novel multi-level grayscale image encoding method based on CIGAR strings of the alignment results and employs image augmentation techniques to improve SV feature extraction. CSV-Filter also utilizes self-supervised learning networks for transfer as classification models, and employs mixed-precision operations to accelerate training. The experiments showed that the integration of CSV-Filter with popular SV detection tools could considerably reduce false positive SVs for short and long reads, while maintaining true positive SVs almost unchanged. Compared with DeepSVFilter, a SV filtering tool for short reads, CSV-Filter could recognize more false positive calls and support long reads as an additional feature. AVAILABILITY AND IMPLEMENTATION: https://github.com/xzyschumacher/CSV-Filter. Weiming Xiang 0003, Qingzhe Wang, Xingze Li, Junyu Gao 0005, Tao Tang 0001, Canqun Yang, Yingbo Cui 0001 |
Bioinform. | 8 |
| 2024 | HFN: Heterogeneous feature network for multivariate time series anomaly detection
Chengkun Wu, Canqun Yang, Qiucheng Miao, Xiandong Ma |
Inf. Sci. | 3 |
| 2024 | Relation representation based on private and shared features for adaptive few-shot link prediction
Canqun Yang |
J. Intell. Inf. Syst. | 2 |
| 2024 | SNCL: a supernode OpenCL implementation for hybrid computing arrays
Tao Tang 0001, Kai Lu 0001, Lin Peng 0001, Yingbo Cui 0001, Jianbin Fang, Chun Huang 0006, Ruibo Wang, Canqun Yang, Yifei Guo |
J. Supercomput. | 8 |
| 2024 | Diagnosing Performance Issues for Large-Scale Microservice Systems With Heterogeneous GraphabstractThe availability of microservice systems is critical to business operations and corporate reputation. However, the dynamics and complexity of microservice systems introduce significant challenges to the performance issue diagnosis of large-scale microservice systems. After investigating hundreds of real-world performance issue cases in Tencent, we find that previous troubleshooting approaches fail to accurately localize root causes because they overlook the inconsistency between causality and calling relationships. Therefore, we propose a novel approach, MicroDig, to diagnose performance issues for large-scale microservice systems. Specifically, MicroDig constructs a heterogeneous propagation graph to capture the causal relationships between calls and microservices. It then conducts a heterogeneity-oriented random walk (HORW) to pinpoint the culprit microservice. Extensive evaluation experiments have been conducted to evaluate MicroDig's performance on 60 real-world performance issues collected from Tencent, 80 manually injected ones collected from a widely used open-source microservice system and 128 performance issues collected from an e-commerce system used by a top-tier global commercial bank. MicroDig achieves 94.1%, 85.5% and 93.8% top-3 accuracy on the three datasets, respectively, significantly outperforming six popular baseline methods. Additionally, we have shared our success stories and learned lessons from the deployment of MicroDig in Tencent. Xianglin Lu, Shenglin Zhang, Jiaqi Luan, Yingke Li, Mingjie Li 0005, Zeyan Li 0001, Qingyang Yu, Hucheng Xie, Chenyuan Hu, Canqun Yang, Dan Pei |
IEEE Trans. Serv. Comput. | 12 |
| 2023 | DrugProtKGE: Weakly Supervised Knowledge Graph Embedding for Highly-Effective Drug-Protein Interaction RepresentationabstractWith the exponential growth of biomedical knowledge in unstructured text repositories such as PubMed, it is imminent to establish a knowledge graph-style, efficient searchable and targeted database that can support the need of information retrieval from researchers and clinicians. To mine knowledge from graph databases, most previous methods view a triple in a graph (see Fig. 1) as the basic processing unit and embed the triplet element (i.e. drugs/chemicals, proteins/genes and their interaction) as separated embedding matrices, which cannot capture the semantic correlation among triple elements. To remedy the loss of semantic correlation caused by disjoint embeddings, we propose a novel approach to learn triple embeddings by combining entities and interactions into a unified representation. Furthermore, traditional methods usually learn triple embeddings from scratch, which cannot take advantage of the rich domain knowledge embedded in pre-trained models, and is also another significant reason for the fact that they cannot distinguish the differences implied by the same entity in the multi-interaction triples. In this paper, we propose a novel fine-tuning based approach to learn better triple embeddings by creating weakly supervised signals from pre-trained knowledge graph embeddings. The method automatically samples triples from knowledge graphs and estimates their pairwise similarity from pre-trained embedding models. The triples are then fed pairwise into a Siamese-like neural architecture, where the triple representation is fine-tuned in the manner bootstrapped by triple similarity scores. Finally, we demonstrate that triple embeddings learned with our method can be readily applied to several downstream applications (e.g. triple classification and triple clustering). We evaluated the proposed method on two open-source drug-protein knowledge graphs constructed from PubMed abstracts, as provided by BioCreative. Our method achieves consistent improvement in both triple classification and triple clustering tasks when compared to other state-of-the-art triple embedding methods, with an average 35% improvement of F1 score for the multi-interaction triples. Siqi Wang 0001, Xi Yang 0020, Xinyuan Qiu, Chengkun Wu, Yingbo Cui 0001, Canqun Yang |
BIBM | 7 |
| 2023 | Accelerating Type Confusion Detection by Identifying Harmless Type CastingsabstractC++ allows reinterpretation of memory objects via type casting, which facilitates easier manipulation of class fields and virtual methods inside the class hierarchy. However, misinterpretation of memory objects, which is called type confusion, can result in illegal access of class fields or methods. Type confusion accounts for many security vulnerabilities for programs written in C++. Previous type confusion detection techniques report a type confusion bug when an object of a parent class is casted to a child class. However, a downcast is safe as long as no illegal fields or methods are accessed. This paper presents Harmless Type Casting Detection (htade), which identifies safe downcast instructions and removes redundant runtime verifications before them by analyzing the type and access information of casted objects. We evaluated htade against 11 SPEC CPU 2006/2017 C++ programs. Compared with LLVM-CFI, htade can reduce the runtime performance overhead by 58.98% on average. Xiaokang Fan, Chun Huang 0006, Canqun Yang, Fa Li |
CF | 4 |
| 2023 | An Improved Parallel Overset Grid Method for Fluid Simulation with Moving BoundaryabstractThe Overset Grid method is a promising computational approach for tackling the challenging moving boundary problems in Computational Fluid Dynamics (CFD) simulations. The computational efficiency and accuracy of the method are critically dependent on the effectiveness of the Overset Grid Assembly (OGA) process. However, the OGA process is plagued by unavoidable issues of load imbalance and communication overheads, which adversely impact the parallel efficiency of the method, particularly when dealing with sub-grids in motion. This paper proposes an improved parallel assembly approach as an effective alternative to address these challenges. Specifically, we introduce a Balanced Merging After Decomposition (BMAD) approach, which ensures that each processor possesses a uniform number of cells from each sub-grid after partitioning and a consistent donor search time. In addition, we deploy a fine-grained list to reduce the data transfer domain, thereby minimizing communication redundancy and cost. We validate the efficiency of our approach in the case of a moving Autonomous Underwater Vehicle (AUV). Experimental results in 3 × 106 grid cells indicate that the proposed approach reduces the parallel computational cost of the OGA process by an average of 21.9% and the speedup has increased by 23.9% with 128 processors. Additionally, it demonstrated equally effective and stable performance in tests using 6 × 106 grid cells, especially achieving the highest speedup of 55.0 with 256 processors. Chao Li 0070, Yi Liu 0083, Canqun Yang |
ICPP | 8 |
| 2023 | A large scale parallel fluid-structure interaction computing platform for simulating structural responses to a detonation shockabstractAbstract Due to the intrinsic nature of multi‐physics, it is prohibitively complex to design and implement a simulation software platform for study of structural responses to a detonation shock. In this article, a partitioned fluid‐structure interaction computing platform is designed for parallel simulating structural responses to a detonation shock. The detonation and wave propagation are modeled in an open‐source multi‐component solver based on OpenFOAM and blastFoam, and the structural responses are simulated through the finite element library deal.II. To capture the interaction dynamics between the fluid and the structure, both solvers are adapted to preCICE. For improving the parallel performance of the computing platform, the inter‐solver data is exchanged by peer‐to‐peer communications and the intermediate server in conventional multi‐physics software is eliminated. Furthermore, the coupled solver with detonation support has been deployed on a computing cluster after considering the distributed data storage and load‐balancing between solvers. The 3D numerical result of structural responses to a detonation shock is presented and analyzed. On 256 processor cores, the speedup ratio of the simulations for a detonation shock reach 178.0 with 5.1 million of mesh cells and the parallel efficiency achieve 69.5%. The results demonstrate good potential of massively parallel simulations. Overall, a general‐purpose fluid‐structure interaction software platform with detonation support is proposed by integrating open source codes. And this work has important practical significance for engineering application in fields of construction blasting, mining, and so forth. Chao Li 0070, Yi Liu 0083, Sijiang Fan, Canqun Yang |
Softw. Pract. Exp. | 7 |
| 2022 | Stgat-Mad : Spatial-Temporal Graph Attention Network For Multivariate Time Series Anomaly DetectionabstractAnomaly detection in multivariate time series data is challenging due to complex temporal and feature correlations. This paper proposes a novel unsupervised multi-scale stacked spatial-temporal graph attention network for multivariate time series anomaly detection (STGAT-MAD). The core of our framework is to coherently capture the feature and temporal correlations among multivariate time-series data by stackable STGAT networks. Meanwhile, a multi-scale input network is exploited to capture the temporal correlations in different time-scales. Besides, a new dataset derived from a real-world wind farm is built and released for multivariate time series anomaly detection. Experiments on the proprietary dataset and three public datasets show that our method significantly outperforms existing baseline approaches, and provides interpretability for anomaly location. Siqi Wang 0001, Xiandong Ma, Chengkun Wu, Canqun Yang, Detian Zeng, Shi-Lin Wang |
ICASSP | 5 |
| 2022 | ParallelDualSPHysics: supporting efficient parallel fluid simulations through MPI-enabled SPH methodabstractSmoothed Particle Hydrodynamics (SPH) is a classical mesh-free particle method which has been successfully applied in the field of Computational Fluid Dynamics (CFD). Its advantages over traditional mesh-based methods have made it very popular in simulating problems involving large deformation and free-surface flow. The high computational cost of the SPH method has obstructed its vast application. A lot of research effort has been devoted to accelerating the SPH method using GPU and multi threading. However, developing efficient parallel SPH algorithms on modern high-performance computers (HPCs) remains significantly challenging, especially for simulating real-world engineering problems involving hundreds of millions of particles. In this paper, we proposed an MPI-enabled parallel SPH algorithm and developed the ParallelDualSPHysics1, an open-source software supporting efficient parallel fluid simulations. Based on an efficient domain decomposition scheme, the essential data structure and algorithms of DualSPHysics were refactored to build the parallel version. For collaborating with evenly distributed particles on a distributed-memory HPC system, the parallel particle interaction and particle update modules were introduced, which enabled the SPH solver to synchronize computations among multiple processors using MPI. In addition, the redesigned pre-processing and post-processing capabilities of the ParallelDualSPHysics supported the applications of this software in a wide range of areas. Real-life test cases with up to 120 million particles were simulated and analyzed on a modern HPC system. The results showed that the parallel efficiency of ParallelDualSPHysics exceeds 90 with up to 1024 CPU cores. It indicated that ParallelDualSPHysics has the potential for large-scale engineering applications. Xiaokang Fan, Chao Li 0070, Kelvin K. L. Wong, Yi Liu 0083, Canqun Yang |
ICPP | 9 |
| 2022 | Private and Shared Feature Extractors Based on Hierarchical Neighbor Encoder for Adaptive Few-Shot Knowledge Graph CompletionabstractWhile Knowledge Graphs (KGs) have been applied in many AI tasks, KGs are known for being incomplete with many missing facts. Previous works rely on a large number of training data for KG completion. However, there are often few entity pairs available for most relations in KGs. In this paper, we propose a Few-shot Knowledge Graph Completion (FKGC) model, named Private and Shared feature extractors based on Hierarchical neighbor encoder for Adaptive few-shot knowledge graph completion (PSHA). In the PSHA model, we first exploit the hierarchical attention mechanism to extract the inherent and valuable hidden information of the neighborhood surrounding the entity. Following that, we adopt a private feature extractor to extract the private features of relation information of the entity pairs, and then a shared feature extractor is used to extract the shared features of the entity pairs of the support set. In addition, an adaptive aggregator aggregates entity pairs of the support set about the query. We conduct experiments on the 2-shot and 5-shot of the NELL-One and CoDEx-S-One dataset. The experimental results show that the PSHA outperforms the existing FKGC models in both scenarios. Canqun Yang |
ICTAI | 1 |
| 2021 | Large-Scale Parallel Alignment Algorithm for SMRT Reads
Yingbo Cui 0001, Peng Zhang 0061, Tao Tang 0001, Lin Peng 0001, Chun Huang 0006, Canqun Yang, Xiangke Liao |
ICA3PP (2) | 9 |
| 2021 | CNN+LSTM Accelerated Turbulent Flow Simulation with Link-Wise Artificial Compressibility MethodabstractThe simulation of turbulent flow, the most common form of fluid, is indispensable in computational fluid dynamics (CFD). The synthetic eddy method (SEM) generates the turbulent inflow and is adopted as the inlet boundary condition of simulation. However, SEM is time-consuming and can significantly slow down the simulation process which occupies 58% of the whole computational time. This is highly inefficient especially since SEM is only used as the inlet. In this paper we propose an efficient alternative. In particular, we leverage CNN+LSTM to replace SEM to obtain the turbulence statistics and combine it with link-wise artificial compressibility method (LW-ACM), which is a fast numerical method of CFD. We validate the predicted results by CNN+LSTM and prove that our model can provide the correct turbulence statistics even after a long time. Experiment results show that our CNN+LSTM module achieves over 15 × speedup compared with SEM, which greatly reduces the time consumption of turbulent inflow generation (from 58% to 7%). As a result, the whole time of turbulent flow simulation is more than halved. Compared with a newly released GPU-accelerated standard lattice Boltzmann method solver, our combination of CNN+LSTM and LW-ACM is about 8.6 × faster. Among all studies reported to date, our work is the fastest implementation for simulating turbulent channel flow, an important step for the field of fast CFD analysis. Sijiang Fan, Jiawei Fei, Canqun Yang, Alistair Revell |
ICPP | 4 |
| 2021 | VISPR-online: a web-based interactive tool to visualize CRISPR screening experimentsabstractBACKGROUND: VISPR is an interactive visualization and analysis framework for CRISPR screening experiments. However, it only supports the output of MAGeCK, and requires installation and manual configuration. Furthermore, VISPR is designed to run on a single computer, and data sharing between collaborators is challenging. RESULTS: To make the tool easily accessible to the community, we present VISPR-online, a web-based general application allowing users to visualize, explore, and share CRISPR screening data online with a few simple steps. VISPR-online provides an exploration of screening results and visualization of read count changes. Apart from MAGeCK, VISPR-online supports two more popular CRISPR screening analysis tools: BAGEL and JACKS. It provides an interactive environment for exploring gene essentiality, viewing guide RNA (gRNA) locations, and allowing users to resume and share screening results. CONCLUSIONS: VISPR-online allows users to visualize, explore and share CRISPR screening data online. It is freely available at http://vispr-online.weililab.org , while the source code is available at https://github.com/lemoncyb/VISPR-online . Yingbo Cui 0001, Johannes Köster, Xiangke Liao, Shaoliang Peng, Tao Tang 0001, Chun Huang 0006, Canqun Yang |
BMC Bioinform. | 8 |
| 2021 | Mining microbe-disease interactions from literature via a transfer learning modelabstractBACKGROUND: Interactions of microbes and diseases are of great importance for biomedical research. However, large-scale of microbe-disease interactions are hidden in the biomedical literature. The structured databases for microbe-disease interactions are in limited amounts. In this paper, we aim to construct a large-scale database for microbe-disease interactions automatically. We attained this goal via applying text mining methods based on a deep learning model with a moderate curation cost. We also built a user-friendly web interface that allows researchers to navigate and query required information. RESULTS: Firstly, we manually constructed a golden-standard corpus and a sliver-standard corpus (SSC) for microbe-disease interactions for curation. Moreover, we proposed a text mining framework for microbe-disease interaction extraction based on a pretrained model BERE. We applied named entity recognition tools to detect microbe and disease mentions from the free biomedical texts. After that, we fine-tuned the pretrained model BERE to recognize relations between targeted entities, which was originally built for drug-target interactions or drug-drug interactions. The introduction of SSC for model fine-tuning greatly improved detection performance for microbe-disease interactions, with an average reduction in error of approximately 10%. The MDIDB website offers data browsing, custom searching for specific diseases or microbes, and batch downloading. CONCLUSIONS: Evaluation results demonstrate that our method outperform the baseline model (rule-based PKDE4J) with an average [Formula: see text]-score of 73.81%. For further validation, we randomly sampled nearly 1000 predicted interactions by our model, and manually checked the correctness of each interaction, which gives a 73% accuracy. The MDIDB webiste is freely avaliable throuth http://dbmdi.com/index/. Chengkun Wu, Xinyi Xiao, Canqun Yang, Jiacai Yi |
BMC Bioinform. | 3 |
| 2021 | BALS: Blocked Alternating Least Squares for Parallel Sparse Matrix Factorization on GPUsabstractMatrix factorization on sparse matrices has been proven to be an effective approach for data mining and machine learning. However, the prior parallel implementations for matrix factorization fail to capture the internal social property embedded in real-world use cases. This article presents an efficient implementation of the alternative least squares (ALS) algorithm calledBALSbuilt on top of a new sparse matrix format for parallel matrix factorization. The BALS storage format organizes the sparse matrix into 2D tiles to avoid repeated data loads and improve data reuses. We further propose a data reordering technique to sort sparse matrices according to nonzeros. The experimental results show that BALS can yield a superior performance than state-of-the-art implementations, i.e., our BALS generally runs faster than Gates’ implementation over different latent feature sizes, with a speedup of up to 2.08× on K20C, 3.72× on TITAN X and 3.13× on TITAN RTX. When compared with alternative matrix factorization algorithms, our BALS consistently outperforms CDMF, cuMF_CCD, and cuMF_SGD over various latent feature sizes and datasets. The reordering technique can provide an extra improvement of up to 23.68 percent on K20C, 19.87 percent on TITAN X and 20.38 percent on TITAN RTX. Jing Chen 0038, Jianbin Fang, Weifeng Liu 0002, Canqun Yang |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | CGINet: graph convolutional network-based model for identifying chemical-gene interaction in an integrated multi-relational graphabstractBACKGROUND: Elucidation of interactive relation between chemicals and genes is of key relevance not only for discovering new drug leads in drug development but also for repositioning existing drugs to novel therapeutic targets. Recently, biological network-based approaches have been proven to be effective in predicting chemical-gene interactions. RESULTS: We present CGINet, a graph convolutional network-based method for identifying chemical-gene interactions in an integrated multi-relational graph containing three types of nodes: chemicals, genes, and pathways. We investigate two different perspectives on learning node embeddings. One is to view the graph as a whole, and the other is to adopt a subgraph view that initial node embeddings are learned from the binary association subgraphs and then transferred to the multi-interaction subgraph for more focused learning of higher-level target node representations. Besides, we reconstruct the topological structures of target nodes with the latent links captured by the designed substructures. CGINet adopts an end-to-end way that the encoder and the decoder are trained jointly with known chemical-gene interactions. We aim to predict unknown but potential associations between chemicals and genes as well as their interaction types. CONCLUSIONS: We study three model implementations CGINet-1/2/3 with various components and compare them with baseline approaches. As the experimental results suggest, our models exhibit competitive performances on identifying chemical-gene interactions. Besides, the subgraph perspective and the latent link both play positive roles in learning much more informative node embeddings and can lead to improved prediction. Wei Wang 0130, Xi Yang 0020, Chengkun Wu, Canqun Yang |
BMC Bioinform. | 4 |
| 2020 | clMF: A fine-grained and portable alternating least squares algorithm for parallel matrix factorization
Jing Chen 0038, Jianbin Fang, Weifeng Liu 0002, Tao Tang 0001, Canqun Yang |
Future Gener. Comput. Syst. | 5 |
| 2020 | Correlation maximization machine for multi-modalities multiclass classification
Canqun Yang, Naiyang Guan |
Pattern Anal. Appl. | 1 |
| 2020 | High-Scalable Collaborated Parallel Framework for Large-Scale Molecular Dynamic Simulation on Tianhe-2 SupercomputerabstractMolecular dynamics (MD) is a computer simulation method of studying physical movements of atoms and molecules that provide detailed microscopic sampling on molecular scale. With the continuous efforts and improvements, MD simulation gained popularity in materials science, biochemistry and biophysics with various application areas and expanding data scale. Assisted Model Building with Energy Refinement (AMBER) is one of the most widely used software packages for conducting MD simulations. However, the speed of AMBER MD simulations for system with millions of atoms in microsecond scale still need to be improved. In this paper, we propose a parallel acceleration strategy for AMBER on the Tianhe-2 supercomputer. The parallel optimization of AMBER is carried out on three different levels: fine grained OpenMP parallel on a single CPU, single node CPU/MIC parallel optimization and multi-node multi-MIC collaborated parallel acceleration. By the three levels of parallel acceleration strategy above, we achieved the highest speedup of 25-33 times compared with the original program. Shaoliang Peng, Xiaoyu Zhang 0008, Wenhe Su, Yutong Lu, Xiangke Liao, Kai Lu 0001, Canqun Yang, Jie Liu 0002, Weiliang Zhu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2020 | Optimizing Streaming Parallelism on Heterogeneous Many-Core ArchitecturesabstractAs many-core accelerators keep integrating more processing units, it becomes increasingly more difficult for a parallel application to make effective use of all available resources. An effective way of improving hardware utilization is to exploit spatial and temporal sharing of the heterogeneous processing units by multiplexing computation and communication tasks - a strategy known as heterogeneous streaming. Achieving effective heterogeneous streaming requires carefully partitioning hardware among tasks, and matching the granularity of task parallelism to the resource partition. However, finding the right resource partitioning and task granularity is extremely challenging, because there is a large number of possible solutions and the optimal solution varies across programs and datasets. This article presents an automatic approach to quickly derive a good solution for hardware resource partition and task granularity for task-based parallel applications on heterogeneous many-core architectures. Our approach employs a performance model to estimate the resulting performance of the target application under a given resource partition and task granularity configuration. The model is used as a utility to quickly search for a good configuration at runtime. Instead of hand-crafting an analytical model that requires expert insights into low-level hardware details, we employ machine learning techniques to automatically learn it. We achieve this by first learning a predictive model offline using training programs. The learned model can then be used to predict the performance of any unseen program at runtime. We apply our approach to 39 representative parallel applications and evaluate it on two representative heterogeneous many-core platforms: a CPU-XeonPhi platform and a CPU-GPU platform. Compared to the single-stream version, our approach achieves, on average, a 1.6x and 1.1x speedup on the XeonPhi and the GPU platform, respectively. These results translate to over 93 percent of the performance delivered by a theoretically perfect predictor. Peng Zhang 0061, Jianbin Fang, Canqun Yang, Chun Huang 0006, Tao Tang 0001, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | The Communication-Overlapped Hybrid Decomposition Parallel Algorithm for Multi-Scale Fluid SimulationsabstractThe MCDPar (Parallel algorithm for multi-scale simulations based on Mesh and BCF Decomposition) algorithm significantly reduced the execution time and improved the parallel scalability for the multi-scale fluid simulations. However, the performance bottleneck still exists for extremely large-scale parallel simulations. In this paper, we designed a communication-overlapped hybrid decomposition parallel algorithm to improve the performance of the original MCDPar on large-scale clusters. Through non-blocking communication and code scheduling, the communication overhead between the master and slave groups have been overlapped with the computation of more microscopic configuration fields for the master process. Thus the parallel efficiency and scalability of the multi-scale solver could be improved on large-scale parallel simulations. In the test case with the number of configuration fields NBCF = 1000 and mesh cells Ncell = 64000, the communication percentage between the corresponding master and slave processes is reduced by 39.71%. In the test case with NBCF = 3000 and Ncell = 64000, the time cost of the fastest execution is reduced by 31.13% using the communication-overlapped algorithm, which offers a better parallel scaling on 256 cores compared to original 128 cores. Yi Liu 0083, Chao Li 0070, Canqun Yang, Xinbiao Gan, Peng Zhang 0061, Sijiang Fan |
ICPP | 4 |
| 2019 | GARDENIA: A Graph Processing Benchmark Suite for Next-Generation AcceleratorsabstractThis article presents the Graph Algorithm Repository for Designing Next-generation Accelerators (GARDENIA), a benchmark suite for studying irregular graph algorithms on massively parallel accelerators. Applications with limited control and data irregularity are the main focus of existing generic benchmarks for accelerators, while available graph processing benchmarks do not apply state-of-the-art algorithms and/or optimization techniques. GARDENIA includes emerging graph processing workloads from graph analytics, sparse linear algebra, and machine-learning domains, which mimic massively multithreaded commercial programs running on modern large-scale datacenters. Our characterization shows that GARDENIA exhibits irregular microarchitectural behavior, which is quite different from structured workloads and straightforward-implemented graph benchmarks. Zhen Xu 0004, Xuhao Chen 0001, Jie Shen 0003, Yang Zhang 0026, Cheng Chen 0005, Canqun Yang |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2019 | SCP: Shared Cache Partitioning for High-Performance GEMMabstractGEneral Matrix Multiply (GEMM) is the most fundamental computational kernel routine in the BLAS library. To achieve high performance, in-memory data must be prefetched into fast on-chip caches before they are used. Two techniques, software prefetching and data packing, have been used to effectively exploit the capability of on-chip least recent used (LRU) caches, which are popular in traditional high-performance processors used in high-end servers and supercomputers. However, the market has recently witnessed a new diversity in processor design, resulting in high-performance processors equipped with shared caches with non-LRU replacement policies. This poses a challenge to the development of high-performance GEMM in a multithreaded context. As several threads try to load data into a shared cache simultaneously, interthread cache conflicts will increase significantly. We present a Shared Cache Partitioning (SCP) method to eliminate interthread cache conflicts in the GEMM routines, by partitioning a shared cache into physically disjoint sets and assigning different sets to different threads. We have implemented SCP in the OpenBLAS library and evaluated it on Phytium 2000+, a 64-core AArch64 processor with private LRU L1 caches and shared pseudo-random L2 caches (per four-core cluster). Our evaluation shows that SCP has effectively reduced the conflict misses in both L1 and L2 caches in a highly optimized GEMM implementation, resulting in an improvement of its performance by 2.75% to 6.91%. Xing Su 0004, Xiangke Liao, Hao Jiang 0001, Canqun Yang, Jingling Xue |
ACM Trans. Archit. Code Optim. | 4 |
| 2019 | Toward fault-tolerant hybrid programming over large-scale heterogeneous clusters via checkpointing/restart optimization
Cheng Chen 0005, Yunfei Du 0001, Ke Zuo, Jianbin Fang, Canqun Yang |
J. Supercomput. | 5 |
| 2019 | Application-aware NoC management in GPUs multitasking
Zhen Xu 0004, Xia Zhao 0004, Zhiying Wang 0003, Canqun Yang |
J. Supercomput. | 4 |
| 2018 | MOCL: an efficient openCL implementation for the matrix-2000 architectureabstractThis paper presents the design and implementation of an Open Computing Language (OpenCL) framework for the Matrix-2000 many-core architecture. This architecture is designed to replace the Intel XeonPhi accelerators of the TianHe-2 supercomputer. We share our experience and insights on how to design an effective OpenCL system for this new hardware accelerator. We propose a set of new analysis and optimizations to unlock the potential of the hardware. We extensively evaluate our approach using a wide range of OpenCL benchmarks on a single and multiple computing nodes. We present our design choices and provide guidance how to optimize code on the new Matrix-2000 architecture. Peng Zhang 0061, Tao Tang 0001, Jianbin Fang, Chun Huang 0006, Canqun Yang, Zheng Wang 0001 |
CF | 5 |
| 2018 | UHCL-Darknet: An OpenCL-based Deep Neural Network Framework for Heterogeneous Multi-/Many-core ClustersabstractAs the majority of popular deep neural network (DNN) frameworks focus on a closed format CUDA implementations based on one or more NVIDIA GPUs, they cannot efficiently leverage other devices in cluster mode to accelerate the training and inference of DNNs except NVIDIA GPUs. To accelerate DNNs using heterogeneous multi-/many-core clusters, we propose an OpenCL-based DNN framework called UHCL-Darknet. First, we design a unified OpenCL platform model for the heterogeneous cluster called UHCL, and an adaptive runtime system with the affinity-based dynamic scheduler for UHCL, enabling transparent utilization of a wide variety of vendor-specific OpenCL devices in the heterogeneous cluster. Then, we extend Darknet to UHCL by introducing the parallel optimization of DNNs, such as paralleling Winogrand-based convolutions and auto-tuning parameterized OpenCL kernels. The training and inference of art-of-data DNN models (e.g., YOLOv2, ResNet-50, and DenseNet-201) are executed on an experimental heterogeneous cluster. Results show that UHCL-Darknet is a scalable and portable DNN framework for heterogeneous clusters, and achieves 1.9x and 2.2x speedups on average respectively for the image throughput of data-parallel training and inference on the experimental heterogeneous cluster. Longlong Liao, Kenli Li 0001, Keqin Li 0001, Canqun Yang, Qi Tian 0001 |
ICPP | 4 |
| 2018 | Auto-tuning Streamed Applications on Intel Xeon PhiabstractMany-core accelerators, as represented by the XeonPhi coprocessors and GPGPUs, allow software to exploit spatial and temporal sharing of computing resources to improve the overall system performance. To unlock this performance potential requires software to effectively partition the hardware resource to maximize the overlap between host-device communication and accelerator computation, and to match the granularity of task parallelism to the resource partition. However, determining the right resource partition and task parallelism on a per program, per dataset basis is challenging. This is because the number of possible solutions is huge, and the benefit of choosing the right solution may be large, but mistakes can seriously hurt the performance. In this paper, we present an automatic approach to determine the hardware resource partition and the task granularity for any given streamed application, targeting the Intel XeonPhi architecture. Instead of hand-crafting the heuristic for which the process will have to repeat for each hardware generation, we employ machine learning techniques to automatically learn it. We achieve this by first learning a predictive model offline using training programs; we then use the learned model to predict the resource partition and task granularity for any unseen programs at runtime. We apply our approach to 23 representative parallel applications and evaluate it on a CPU-XeonPhi mixed heterogenous many-core platform. Our approach achieves, on average, a 1.6x (upto 5.6x) speedup, which translates to 94.5% of the performance delivered by a theoretically perfect predictor. Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Canqun Yang, Zheng Wang 0001 |
IPDPS | 4 |
| 2018 | Collaborative Subspace Graph Hashing for Cross-modal RetrievalabstractCurrent hashing methods for cross-modal retrieval generally attempt to learn the separate modality-specific transformation matrices to embed multi-modality data into a latent common subspace, and usually ignore the fact that respecting the diversity of multi-modality features in the latent subspace could be beneficial for retrieval improvements. To this, we propose a collaborative subspace graph hashing method (CSGH) to perform a two-stage collaborative learning framework for cross-modal retrieval. Particularly, CSGH first embeds multi-modality data into separate latent subspaces through individual modality-specific transformation matrices, and then connects these latent subspaces to a common Hamming space through a shared transformation matrix. In this framework, CSGH considers the modality-specific neighborhood structure and the cross-modal correlation within multi-modality data through the Laplacian regularization and the graph based correlation constraint, respectively. To solve CSGH, we develop an alternative procedure to optimize it, and fortunately, each sub-problem of CSGH has the elegant analytical solution. Experiments of cross-modal retrieval on Wiki, NUS-WIDE, Flickr25K and Flickr1M datasets show the effectiveness of CSGH compared with the state-of-the-art cross-modal hashing methods. Xiang Zhang 0008, Guohua Dong, Yimo Du, Chengkun Wu, Zhigang Luo, Canqun Yang |
ICMR | 6 |
| 2018 | A hybrid deep learning CNN-ELM for age and gender classification
Mingxing Duan, Kenli Li 0001, Canqun Yang, Keqin Li 0001 |
Neurocomputing | 3 |
| 2018 | Moving from exascale to zettascale computing: challenges and techniquesabstractHigh-performance computing (HPC) is essential for both traditional and emerging scientific fields, enabling scientific activities to make progress. With the development of high-performance computing, it is foreseeable that exascale computing will be put into practice around 2020. As Moore’s law approaches its limit, high-performance computing will face severe challenges when moving from exascale to zettascale, making the next 10 years after 2020 a vital period to develop key HPC techniques. In this study, we discuss the challenges of enabling zettascale computing with respect to both hardware and software. We then present a perspective of future HPC technology evolution and revolution, leading to our main recommendations in support of zettascale computing in the coming future. Xiangke Liao, Kai Lu 0001, Canqun Yang, Jin-wen Li, Yuan Yuan 0034, Libo Huang 0002, Pingjing Lu, Jianbin Fang, Jie Shen 0003 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2018 | Orchestrating parallel detection of strongly connected components on GPUs
Xuhao Chen 0001, Cheng Chen 0005, Jie Shen 0003, Jianbin Fang, Tao Tang 0001, Canqun Yang, Zhiying Wang 0003 |
Parallel Comput. | 6 |
| 2017 | Automatic density clustering with multiple kernels for high-dimension bioinformatics dataabstractClustering is an effective method for data analysis and can be exploited to unknown features of data samples, its applications range from data mining to bioinformatics analysis. Several clustering approaches have been proposed in order to obtain a better trade-off between accuracy and efficiency of the clustering process. It is well-known that no existing clustering algorithm completely satisfies both accuracy and efficiency requirements, thus we propose a clustering algorithm called ADCMK (for Automatic Density Clustering with Multiple Kernels) exhibiting higher quality than the density ones proposed so far, while allowing users to cluster efficiently without determining parameters manually. The algorithm consists of learning optimal combined kernel, reducing dimensionality with the optimal kernel, automatically detecting cluster centroids with outliers test, assigning clusters and visualizing results. The proposed clustering algorithm is extensively tested on several well-known high-dimensional dataset in biomedical and bioinformatics field. The results show that the new algorithm tends to automatically produce clusters of better quality than other density clustering algorithms. Longlong Liao, Kenli Li 0001, Keqin Li 0001, Qi Tian 0001, Canqun Yang |
BIBM | 5 |
| 2017 | Dependency-based long short term memory network for drug-drug interaction extractionabstractBACKGROUND: Drug-drug interaction extraction (DDI) needs assistance from automated methods to address the explosively increasing biomedical texts. In recent years, deep neural network based models have been developed to address such needs and they have made significant progress in relation identification. METHODS: We propose a dependency-based deep neural network model for DDI extraction. By introducing the dependency-based technique to a bi-directional long short term memory network (Bi-LSTM), we build three channels, namely, Linear channel, DFS channel and BFS channel. All of these channels are constructed with three network layers, including embedding layer, LSTM layer and max pooling layer from bottom up. In the embedding layer, we extract two types of features, one is distance-based feature and another is dependency-based feature. In the LSTM layer, a Bi-LSTM is instituted in each channel to better capture relation information. Then max pooling is used to get optimal features from the entire encoding sequential data. At last, we concatenate the outputs of all channels and then link it to the softmax layer for relation identification. RESULTS: To the best of our knowledge, our model achieves new state-of-the-art performance with the F-score of 72.0% on the DDIExtraction 2013 corpus. Moreover, our approach obtains much higher Recall value compared to the existing methods. CONCLUSIONS: The dependency-based Bi-LSTM model can learn effective relation information with less feature engineering in the task of DDI extraction. Besides, the experimental results show that our model excels at balancing the Precision and Recall values. Wei Wang 0130, Xi Yang 0020, Canqun Yang, Xiang Zhang 0008, Chengkun Wu |
BMC Bioinform. | 3 |
| 2017 | Efficient and high-quality sparse graph coloring on GPUsabstractSummary Graph coloring has been broadly used to discover concurrency in parallel computing. To speed up graph coloring for large‐scale datasets, parallel algorithms have been proposed to leverage modern GPUs. Existing GPU implementations either have limited performance or yield unsatisfactory coloring quality (too many colors assigned). We present a work‐efficient parallel graph coloring implementation on GPUs with good coloring quality. Our approach uses the speculative greedy scheme, which inherently yields better quality than the method of finding maximal independent set . To achieve high performance on GPUs, we refine the algorithm to leverage efficient operators and alleviate conflicts. We also incorporate common optimization techniques to further improve performance. Our method is evaluated with both synthetic and real‐world sparse graphs on the NVIDIA GPU. Experimental results show that our proposed implementation achieves averaged 4.1 × (up to 8.9 × ) speedup over the serial implementation. It also outperforms the existing GPU implementation from the NVIDIA CUSPARSE library (2.2 × average speedup), while yielding much better coloring quality than CUSPARSE. Xuhao Chen 0001, Pingfan Li, Jianbin Fang, Tao Tang 0001, Zhiying Wang 0003, Canqun Yang |
Concurr. Comput. Pract. Exp. | 6 |
| 2016 | mAMBER: A CPU/MIC collaborated parallel framework for AMBER on Tianhe-2 supercomputerabstractMolecular dynamics (MD) is a computer simulation method of studying physical movements of atoms and molecules that provide detailed microscopic sampling on molecular scale. With the continuous efforts and improvements, MD simulation gained popularity in materials science, biochemistry and biophysics with various application areas and expanding data scale. Assisted Model Building with Energy Refinement (AMBER) is one of the most widely used software packages for conducting MD simulations. However, the speed of AMBER MD simulations for system with millions of atoms in microsecond scale still need to be improved. In this paper, we propose a parallel acceleration strategy for AMBER on Tianhe-2 supercomputer. The parallel optimization of AMBER is carried out on three different levels: fine grained OpenMP parallel on a single MIC, single-node CPU/MIC collaborated parallel optimization and multi-node multi-MIC collaborated parallel acceleration. By the three levels of parallel acceleration strategy above, we achieved the highest speedup of 25-33 times compared with the original program. Source Code: https://github.com/tianhe2/mAMBER. Shaoliang Peng, Xiaoyu Zhang 0008, Yutong Lu, Xiangke Liao, Kai Lu 0001, Canqun Yang, Jie Liu 0002, Weiliang Zhu |
BIBM | 6 |
| 2016 | An Energy-Efficient Implementation of LU Factorization on Heterogeneous SystemsabstractEnergy consumption is increasingly becoming a critical issue in HPC. There is a broad consensus that future exascale-computing will be strongly constrained by energy consumption. Heterogeneous systems usually feature higher energy efficiency than homogeneous ones since the former employ coprocessors that provide higher GFlops/Watt than CPUs. Thus, it is of great importance to better utilize the coprocessors from an energy-efficiency standpoint. Dense LU factorization (LU) is a critical kernel that is widely used to solve dense linear algebra problems. However, existingheterogeneous implementations are typically designed to be CPU-centered, which rely highly on CPUs and thus suffer from large data transfer overheads via PCIe, hurting the energy efficiency of the entire computer system. We present a coprocessor-resident implementation of LU for a heterogeneous platform to improve energy efficiency without impeding performance by relieving the CPUs from performing unnecessary computations and reducing excessive data transfers via PCIe. In addition, several optimizations are judiciously employed to overlap the computation and communication between the CPUs and coprocessors. Validation on the Tianhe-2 supercomputer shows that our LU implementation gains higher performance, achieves higher energy efficiency, and features a better scalability than Intel MKL. Canqun Yang, Cheng Chen 0005, Tao Tang 0001, Xuhao Chen 0001, Jianbin Fang, Jingling Xue |
ICPADS | 1 |
| 2016 | Streaming Applications on Heterogeneous Platforms
Zhaokui Li, Jianbin Fang, Tao Tang 0001, Xuhao Chen 0001, Canqun Yang |
NPC | 5 |
| 2016 | Accelerator-Centered Programming on Heterogeneous SystemsabstractParallel many cores contribute to heterogeneous architectures and achieve high computation throughput. Working as coprocessors and connected to general-purpose CPUs via PCIe, those special-purpose cores usually work as float computing accelerators (ACC). The popular programming models typically offload the computing intensive parts to accelerator then aggregate results, which would result in a great amount of data transfer via PCIe. In this paper, we introduce an ACC-centered model to leverage the limited bandwidth of PCIe, increase performance, reduce idle time of ACC. In order to realize dada-near-computing, our ACC-centered model arms to program centered on ACC and the control intensive parts are offloaded to CPU. Both CPU and ACC are devoted to higher performance with their architect feature. Validation on the Tianhe-2 supercomputer shows that the implementation of ACC-centered LU competes with the highly optimized Intel MKL hybrid implementation and achieves about 5× speedup versus the CPU version. Cheng Chen 0005, Yunfei Du 0001, Canqun Yang |
PDCAT | 3 |
| 2016 | Accurate Evaluation of Bivariate PolynomialsabstractPolynomials are widely used in scientific computing and engineering. In this paper, we present an accurate and fast compensated algorithm to evaluate bivariate polynomials with floating-point coefficients. This algorithm is applying error free transformations to the bivariate Horner scheme and sum the final decomposition accurately. We also prove the proposed algorithm's accuracy with forward error analysis that the accuracy of the computed result is similar to the result computed by the bivariate Horner scheme in twice the working precision. Numerical experiments illustrate the behavior and it has higher efficiency than the bivariate Horner scheme implemented in double-double library. Peibing Du, Hao Jiang 0001, Housen Li, Lizhi Cheng, Canqun Yang |
PDCAT | 5 |
| 2016 | Bilateral Sampling Randomized Singular Value DecompositionabstractDesigning fast singular value decomposition (SVD) is significantly interesting in applications. The random direct SVD (RSVD) has provided a fast scheme to compute the well-approximate SVD by unilateral randomized sampling. In this paper, we present an efficient random algorithm in a bilateral sampling way. We also prove that the proposed algorithms can be bounded well and have less computational complexity compared to RSVD when the objective matrix is approximately square. Numerical experiments on graph Laplacian and Hilbert matrix demonstrate the efficiency and stability of the proposed methods. Hao Jiang 0001, Peibing Du, Tao Sun 0005, Housen Li, Lizhi Cheng, Canqun Yang |
PDCAT | 6 |
| 2015 | mAMBER: Accelerating Explicit Solvent Molecular Dynamic with Intel Xeon Phi Many-Integrated Core CoprocessorsabstractMolecular dynamics (MD) is a computer simulation of physical movements of atoms and molecules, which is a very important research technique for the study of biological and chemical systems at micro-scale. Assisted Model Building with Energy Refinement (AMBER) is one of the most commonly used software for MD. However, the microsecond MD simulation of large-scale atom system requires a lot of computation power. In this paper, we propose mAMBER: an Intel Xeon Phi Many-Integrated Core (MIC) Coprocessors accelerated implementation of explicit solvent all-atom classical molecular dynamics (MD) within the AMBER program package. We mAMBER also includes new parallel algorithm using CPUs and MIC coprocessors on Tianhe-2 supercomputer. With several optimizing techniques including CPU/MIC collaborated parallelization, factorization and asynchronous data transfer framework, we can accelerate the sander program of AMBER (version 12) in 'offload' mode, and achieves a 4.17-fold overall speedup compared with the CPU-only sander program. Shaoliang Peng, Canqun Yang, Chengkun Wu, Haiqiang Wang, Weiliang Zhu, Jinan Wang |
CCGRID | 3 |
| 2015 | The Challenge of Scaling Genome Big Data Analysis Software on TH-2 SupercomputerabstractWhole genome re-sequencing plays a crucial role in biomedical studies. The emergence of genomic big data calls for an enormous amount of computing power. However, current computational methods are inefficient in utilizing available computational resources. In this paper, we address this challenge by optimizing the utilization of the fastest supercomputer in the world - TH-2 supercomputer. TH-2 is featured by its neo-heterogeneous architecture, in which each compute node is equipped with 2 Intel Xeon CPUs and 3 Intel Xeon Phi coprocessors. The heterogeneity and the massive amount of data to be processed pose great challenges for the deployment of the genome analysis software pipeline on TH-2. Runtime profiling shows that SOAP3-dp and SOAPsnp are the most time-consuming components (up to 70% of total runtime) in a typical genome-analyzing pipeline. To optimize the whole pipeline, we first devise a number of parallel and optimization strategies for SOAP3-dp and SOAPsnp, respectively targeting each node to fully utilize all sorts of hardware resources provided both by CPU and MIC. We also employ a few scaling methods to reduce communication between different nodes. We then scaled up our method on TH-2. With 8192 nodes, the whole analyzing procedure took 8.37 hours to finish the analysis of a 300 TB dataset of whole genome sequences from 2,000 human beings, which can take as long as 8 months on a commodity server. The speedup is about 700x. Shaoliang Peng, Xiangke Liao, Canqun Yang, Yutong Lu, Jie Liu 0002, Yingbo Cui 0001, Chengkun Wu, Bingqiang Wang |
CCGRID | 3 |
| 2015 | FT-Offload: A Scalable Fault-Tolerance Programing Model on MIC Cluster
Cheng Chen 0005, Yunfei Du 0001, Zhen Xu 0004, Canqun Yang |
ICA3PP (4) | 4 |
| 2015 | Implementation of an Accurate and Efficient Compensated DGEMM for 64-bit ARMv8 Multi-Core ProcessorsabstractThis paper presents an implementation of an accurate and efficient compensated Double-precision General Matrix Multiplication (DGEMM) based on OpenBLAS for 64-bit ARMv8 multi-core processors. Due to cancellation phenomena in floating point arithmetic, the results of DGEMM may not be as accurate as expected. In order to increase the accuracy of DGEMM, we compensate the error introduced by its dot product kernel (GEBP) by applying an error-free transformation to rewrite the kernel in assembly language. We optimize the computations in the inner kernel through exploiting loop unrolling, instruction scheduling and software-implemented register rotation to exploit instruction level parallelism (ILP). We also conduct a priori error analysis of the derived CompDGEMM. Our compensated DGEMM is as accurate as the existing quadruple precision GEMM using MBLAS, but is up to 6.4x faster. Our parallel implementation achieves good performance and scalability under varying thread counts across a range of matrix sizes evaluated. Hao Jiang 0001, Feng Wang 0050, Kuan Li, Canqun Yang, Kejia Zhao, Chun Huang 0006 |
ICPADS | 4 |
| 2015 | Design and Implementation of a Highly Efficient DGEMM for 64-Bit ARMv8 Multi-core ProcessorsabstractThis paper presents the design and implementation of a highly efficient Double-precision General Matrix Multiplication (DGEMM) based on Open BLAS for 64-bit ARMv8 eight-core processors. We adopt a theory-guided approach by first developing a performance model for this architecture and then using it to guide our exploration. The key enabler for a highly efficient DGEMM is a highly-optimized inner kernel GEBP developed in assembly language. We have obtained GEBP by (1) maximizing its compute-to-memory access ratios across all levels of the memory hierarchy in the ARMv8 architecture with its performance-critical block sizes being determined analytically, and (2) optimizing its computations through exploiting loop unrolling, instruction scheduling and software-implemented register rotation and taking advantage of A64 instructions to support efficient FMA operations, data transfers and prefetching. We have compared our DGEMM implemented in Open BLAS with another implemented in ATLAS (also in terms of a highly-optimized GEBP in assembly). Our implementation outperforms the one in ALTAS by improving the peak performance (efficiency) of DGEMM from 3.88 Gflops (80.9%) to 4.19 Gflops (87.2%) on one core and from 30.4 Gflops (79.2%) to 32.7 Gflops (85.3%) on eight cores. These results translate into substantial performance (efficiency) improvements by 7.79% on one core and 7.70% on eight cores. In addition, the efficiency of our implementation on one core is very close to the theoretical upper bound 91.5% obtained from micro-benchmarking. Our parallel implementation achieves good performance and scalability under varying thread counts across a range of matrix sizes evaluated. Feng Wang 0050, Hao Jiang 0001, Ke Zuo, Xing Su 0004, Jingling Xue, Canqun Yang |
ICPP | 6 |
| 2014 | HPCG: Preliminary Evaluation and Optimization on Tianhe-2 CPU-only NodesabstractHPCG has become a new metric for the design and ranking of HPC. By incorporating a local symmetric Gauss-Seidel preconditioned, HPCG implements the Conjugate Gradient method to solve a sparse linear system. HPCG performs poorly with irregular memory access and may consume a great deal of MPI resources when it is executed on supercomputers. This paper focuses on optimizing SpMV and the Gauss-Seidel preconditioned, the two most important kernels in HPCG. By evaluating the performance impacts of several representative sparse matrix formats, ELLPACK is selected due to its suitability for SIMD, resulting in a speedup of 2.3x for the SpMV kernel. Multi-coloring is performed for Gauss-Seidel, resulting in a speedup of 7.3x over the reference implementation. The CG convergence rate may also be improved after multi-coloring. Our experimental results show that our optimization process works well on supercomputers, achieving 6.5 Gflops on a CPU-only node. This has boosted the total HPCG Gflops by about 7x, giving rise to 80,151 Gflops on 8192 CPU-only Tianhe-2 nodes. Cheng Chen 0005, Yunfei Du 0001, Hao Jiang 0001, Ke Zuo, Canqun Yang |
SBAC-PAD | 5 |
| 2014 | MilkyWay-2 supercomputer: system and application
Xiangke Liao, Liquan Xiao, Canqun Yang, Yutong Lu |
Frontiers Comput. Sci. | 3 |
| 2014 | OpenMC: Towards Simplifying Programming for TianHe Supercomputers
Xiangke Liao, Canqun Yang, Tao Tang 0001, Huizhan Yi, Feng Wang 0050, Jingling Xue |
J. Comput. Sci. Technol. | 2 |
| 2013 | Exploiting hierarchy parallelism for molecular dynamics on a petascale heterogeneous system
Canqun Yang, Tao Tang 0001, Liquan Xiao |
J. Parallel Distributed Comput. | 2 |
| 2012 | Parallelizing SOR for GPGPUs using alternate loop tiling
Peng Di, Hui Wu 0001, Jingling Xue, Feng Wang 0050, Canqun Yang |
Parallel Comput. | 5 |
| 2011 | Optimizing Linpack Benchmark on GPU-Accelerated Petascale Supercomputer
Feng Wang 0050, Canqun Yang, Yunfei Du 0001, Juan Chen 0001, Huizhan Yi, Weixia Xu 0001 |
J. Comput. Sci. Technol. | 2 |
| 2010 | Adaptive Optimization for Petascale Heterogeneous CPU/GPU ComputingabstractIn this paper, we describe our experiment developing an implementation of the Linpack benchmark for TianHe-1, a petascale CPU/GPU supercomputer system, the largest GPU-accelerated system ever attempted before. An adaptive optimization framework is presented to balance the workload distribution across the GPUs and CPUs with the negligible runtime overhead, resulting in the better performance than the static or the training partitioning methods. The CPU-GPU communication overhead is effectively hidden by a software pipelining technique, which is particularly useful for large memory-bound applications. Combined with other traditional optimizations, the Linpack we optimized using the adaptive optimization framework achieved 196.7 GFLOPS on a single compute element of TianHe-1. This result is 70.1% of the peak compute capability and 3.3 times faster than the result using the vendor's library. On the full configuration of TianHe-1 our optimizations resulted in a Linpack performance of 0.563PFLOPS, which made TianHe-1 the 5th fastest supercomputer on the Top500 list released in November 2009. Canqun Yang, Feng Wang 0050, Yunfei Du 0001, Juan Chen 0001, Jie Liu 0002, Huizhan Yi, Kai Lu 0001 |
CLUSTER | 1 |
| 2010 | TH-1: China's first petaflop supercomputer
Xuejun Yang, Xiangke Liao, Weixia Xu 0001, Junqiang Song, Qingfeng Hu, Jinshu Su, Liquan Xiao, Kai Lu 0001, Qiang Dou, Juping Jiang, Canqun Yang |
Frontiers Comput. Sci. China | 11 |
| 2009 | Solving 2D Nonlinear Unsteady Convection-Diffusion Equations on Heterogenous Platforms with Multiple GPUsabstractSolving complex convection-diffusion equations is very important to many practical mathematical and physical problems. After the finite difference discretization, most of the time for equations solution is spent on sparse linear equation solvers. In this paper, our goal is to solve 2D Nonlinear Unsteady Convection-Diffusion Equations by accelerating an iterative algorithm named Jacobi-preconditioned QMRCGSTAB on a heterogenous platform, which is composed of a multi-core processor and multiple GPUs. Firstly, a basic implementation and evaluation for adapting the problem to this kind of platform is given. Then, we propose two optimization methods to improve the performance: kernel merging method and matrix boundary data processing. Our experimental evaluation on an AMD Opteron(tm) quad-core processor 2380 linked to an NVIDIA Tesla S1070 platform with four GPUs delivers the peak performance of 33 GFLOPS (double precision), which is a speedup of close to a factor 32 compared to the same problem running on 4 cores of the same CPU. Canqun Yang, Zhen Ge, Juan Chen 0001, Feng Wang 0050, Yunfei Du 0001 |
ICPADS | 1 |