EDBT 2026 Demo / reviewers in the wild / expert
Yanjie Wei
dblp:18/4720
· DBLP profile ↗
55ranked-venue papers
0as first author
41since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 34 since 2021Systems, architecture and hardware · 13 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Spatiotemporal Graph Learning of EEG for Multi-type Epileptiform Event Detection
Yanjie Wei, Linxia Xiao, Shizhong Lian |
ISBRA (1) | 3 |
| 2026 | AlloMut: Dual-View Self-Supervised Representation Learning for Predicting the Pathogenicity of Allosteric-Site Mutations
Ziqi Xu 0003, Haowen Zhao, Shaozhen Cai, Siya Huang, Nuoxian Xu, Yanjie Wei |
ISBRA (2) | 8 |
| 2026 | HiFuseVEP: Hierarchical Fusion of Multi-source Protein Language Model Embeddings for Missense Variant Effect Prediction
Ganzhang Zheng, Shaozhen Cai, Haowen Zhao, Siya Huang, Nuoxian Xu, Yanjie Wei |
ISBRA (2) | 8 |
| 2026 | FusionODE: Biochemical Function Prediction from Microscopy Images via Multimodal Continuous-Time Learning
Shengqi Zhou, Jijian Long, Qiucheng Miao, Jovial Niyogisubizo, Yanjie Wei |
ISBRA (2) | 8 |
| 2026 | BioRxnReasoner: Multi-agent Reasoning for Biochemical Reaction Diagram Question Answering
Qiucheng Miao, Shengqi Zhou, Yanjie Wei |
ISBRA (2) | 6 |
| 2026 | RabbitVar: Ultra-fast and accurate somatic small-variant calling on multi-core architectures
Hao Zhang 0142, Lin Gan 0001, Zekun Yin, Lifeng Yan, Honglei Song, Qixin Chang, Yanjie Wei, Beifang Niu, Bertil Schmidt |
Future Gener. Comput. Syst. | 8 |
| 2026 | From indoor to outdoor: Unsupervised domain adaptive gait recognition
Likai Wang 0002, Wei Feng 0005, Rui-Ze Han, Xiangqun Zhang 0003, Yanjie Wei, Song Wang 0002 |
Pattern Recognit. | 5 |
| 2025 | Task-Adaptive Refined Reinforcement Learning with Granular Reward Shaping for Biomedical Information ExtractionabstractThe surge of biomedical literature and omics data calls for automated knowledge extraction, and LLMs show strong potential for this task. However, existing LLMs still struggle with structured tasks, as they are primarily optimized for generating free-form text rather than adhering to schema-constrained outputs. LLM alignment methods often fail to generalize across biomedical tasks due to semantic ambiguity and unstable policy optimization. To address these challenges, we introduce GRASP (Group-Relative Adaptive Structured Prompting), a unified framework for biomedical information extraction. GRASP combines a hierarchical task-aware prompt design that explicitly encodes task semantics with a novel Group-Relative Policy Optimization strategy, enabling fine-grained, semantically sensitive reward modeling. This approach resolves task inter-ference and enhances the fidelity of structured outputs across diverse extraction tasks. Extensive experiments across biomedical and general benchmarks demonstrate that GRASP achieves state-of-the-art performance in structural accuracy and semantic consistency, with up to 11 % and 7 % relative improvements in Micro F1 for NER and relation extraction, respectively, over competitive baselines. Our code is publicly available at GitHub. Qiucheng Miao, Jintao Meng 0001, Yanjie Wei |
BIBM | 4 |
| 2025 | FlexiCell: Deep Learning with Learnable Adaptive Filtering and Dual Attention for Cell SegmentationabstractAccurate cell segmentation remains challenging due to morphological variations, diverse imaging modalities, and unclear cellular boundaries. Existing deep learning (DL) methods struggle to extract features adaptively across heterogeneous cellular environments, thereby limiting generalization capacity. To address these challenges, we propose FlexiCell, a novel adaptive segmentation framework that integrates a learnable adaptive filter with dual attention mechanisms. The core innovation lies in the FlexiFilter approach, which combines standard convolution with adaptive residual learning through learnable mixing parameters. These parameters dynamically balance input preservation and feature enhancement. FlexiCell employs multi-scale FlexiFilter blocks with varying kernel sizes, channel and spatial attention networks, and a dedicated boundary extractor for precise edge detection. Extensive experiments demonstrate superior performance compared to benchmark models, achieving 3.8% improvement in detection accuracy and 5.5 % in segmentation quality on our newly developed induced pluripotent stem (iPS) cell datasets. Further evaluation on standardized Cell Tracking Challenge (CTC) benchmarks confirms state-of-the-art performance on mesenchymal stem cells and glioblastoma datasets, outperforming established CTC methods. The framework demonstrates robust generalization across fluorescence, phase contrast, and differential interference contrast microscopy, without requiring dataset-specific optimization. Codes are available at https://github.com/jovialniyo93/FlexiCell. Jovial Niyogisubizo, Keliang Zhao, Shengqi Zhou, Rui-Ze Han, Jintao Meng 0001, Wenhui Xi, Yanjie Wei |
BIBM | 7 |
| 2025 | MVmamba Deciphers Missense Variant Pathogenicity via Enhanced Bi-Mamba and Structure-Informed Protein Language ModelabstractMissense mutations are common genetic variations that can alter protein functions, and distinguishing pathogenic from benign variants remains challenging despite computational advances. Here, we propose MVmamba, a novel computational approach for predicting disease-associated missense variants, based on large-scale protein language model embeddings and a gate-wave based Bi-Mamba network. MVmamba integrates global and local embeddings from pretrained structural protein language models, overcoming limitations of sequence-only models and manual structural feature extraction. Its GateWave Bi-Mamba module enhances global feature extraction via wavelet transform and dynamically fuses multi-modal features through gating mechanisms, while integrating allele frequency data enriches predictive clues from population genetics. Ablation studies validate the effectiveness of key components, confirming its robustness. Evaluated on an independent test set with 18,731 clinical variants, MVmamba achieves an AUC of 0.901, AUPR of 0.848, MCC of 0.656, and F1-score of 0.772, outperforming 21 state-of-the-art methods. The data and codes for MVmamba are available at https://github.com/mjcoo/MVmamba for academic use. Ganzhang Zheng, Ziqi Xu 0003, Zhenni Huang, Yanjie Wei |
BIBM | 7 |
| 2025 | NM-SpMM: Accelerating Matrix Multiplication Using N: M Sparsity with GPGPUabstractDeep learning demonstrates effectiveness across a wide range of tasks. However, the dense and over-parameterized nature of these models results in significant resource consumption during deployment. In response to this issue, weight pruning, particularly through$N: M$sparsity matrix multiplication, offers an efficient solution by transforming dense operations into semisparse ones.$N: M$sparsity provides an option for balancing performance and model accuracy, but introduces more complex programming and optimization challenges. To address these issues, we design a systematic top-down performance analysis model for$N: M$sparsity. Meanwhile, NM-SpMM is proposed as an efficient general$N: M$sparsity implementation. Based on our performance analysis, NM-SpMM employs a hierarchical blocking mechanism as a general optimization to enhance data locality, while memory access optimization and pipeline design are introduced as sparsity-aware optimization, allowing it to achieve close-to-theoretical peak performance across different sparsity levels. Experimental results show that NM-SpMM is 2.1x faster than nmSPARSE (the state-of-the-art for general$N: M$sparsity) and 1.4× to 6.3× faster than cuBLAS's dense GEMM operations, closely approaching the theoretical maximum speedup resulting from the reduction in computation due to sparsity. NM-SpMM is open source and publicly available at https://github.com/M-H482/NM-SpMM. Du Wu, Zhelang Deng, Jintao Meng 0001, Wenxi Zhu, Bingqiang Wang, Amelie Chi Zhou, Peng Chen 0035, Minwen Deng, Yanjie Wei, Shengzhong Feng, Yi Pan 0001 |
IPDPS | 12 |
| 2025 | An Efficient Parallel List Ranking Algorithm for Graph Concatenation on BSP Graph System
Maocheng Cao, Zhelang Deng, Qiucheng Miao, Jintao Meng 0001, Yanjie Wei, Jiefeng Cheng |
ISBRA (2) | 5 |
| 2025 | AlloPED: Leveraging Protein Language Models and Structure Features for Allosteric Site Prediction
Xiaochuan Chen, Jianqiang Zheng, Zhenni Huang, Ziqi Xu 0003, Junye Huang, Yueyi Tan, Yanjie Wei |
ISBRA (2) | 7 |
| 2025 | CircRNA Profiles Analysis of Neuroblastoma for Identification of Drug TargetsabstractNeuroblastoma is a prevalent pediatric tumor with a low 5-year survival rate among high-risk patients, and the prognosis remains poor despite available therapeutic interventions. Therefore, identifying novel and effective therapeutic targets is critical for improving outcomes in these patients. In this study, we performed an integrative analysis of two neuroblastoma circRNA sequencing datasets to identify potential drug targets. By comparing circRNA expression levels between neuroblastoma tissues and adjacent normal tissues, we identified differentially expressed circRNAs and subsequently predicted 30 hub circRNAs through Weighted Gene Co-expression Network Analysis. To elucidate the functional roles of these circRNAs, we investigated their interactions with RNA-binding proteins. The results suggest that hsa_circ_0051680 and hsa_circ_0006107 may influence neuroblastoma progression through interactions with FUS and IGF2BP1, respectively. Furthermore, we analyzed the translational potential of the hub circRNAs, revealing that six circRNAs encode proteins with complex secondary structures. Molecular docking analysis identified five high-affinity complexes between circRNA-encoded proteins (hsa_circ_0000786, hsa_circ_0005087, hsa_circ_0006867) and their corresponding ligands. These circRNA-derived proteins present promising novel drug targets for both the diagnosis and treatment of neuroblastoma. Zhen Ju, Godfrey Chi-Fung Chan, Jintao Meng 0001, Wenhui Xi, Yanjie Wei |
IEEE Trans. Comput. Biol. Bioinform. | 9 |
| 2025 | A New Benchmark and Algorithm for Clothes-Changing Video Person Re-IdentificationabstractPerson re-identification (Re-ID) is a classical computer vision task and has significant applications for public security and information forensics. Recently, long-term Re-ID with clothes-changing has attracted increasing attention. However, existing methods mainly focus on image-based setting, where richer temporal information is overlooked. In this paper, we focus on the relatively new yet practical problem of Clothes-Changing Video-based Re-ID (CCVReID), which is less studied. First, given the dataset shortage, we build two new benchmark datasets for CCVReID problem, including a large-scale synthetic video dataset and a real-world one, both containing human sequences with various clothing changes. Moreover, we systematically study this problem by simultaneously considering the classical appearance feature and temporal feature contained in the video. We develop a dual-branch fusion framework that makes use of the information from both clothes-aware appearance feature and clothes-free gait feature. For better information fusion, a confidence-guided re-ranking strategy is proposed to adaptively balance the weight of these two categories of features. We have released the benchmark and code proposed in this work to the public athttps://github.com/kkw98/CCVReID. Likai Wang 0002, Xiangqun Zhang 0003, Rui-Ze Han, Yanjie Wei, Song Wang 0002, Wei Feng 0005 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | An In-Depth Assessment of Sequence Clustering Software in Bioinformatics
Zhen Ju, Xuelei Li, Jintao Meng 0001, Wenhui Xi, Yanjie Wei |
ISBRA (1) | 6 |
| 2024 | PmmNDD: Predicting the Pathogenicity of Missense Mutations in Neurodegenerative Diseases via Ensemble Learning
Xijian Li, Runxuan Tang, Guangcheng Xiao, Xiaochuan Chen, Ruilin He, Zhaolei Zhang, Jiana Luo, Yanjie Wei, Yijun Mao |
ISBRA (3) | 9 |
| 2024 | autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm ArchitecturesabstractThis paper presents an open-source library that pushes the limits of performance portability for irregular General Matrix Multiplication (GEMM) on the widely-used Arm architectures. Our library, autoGEMM, is designed to support a wide range of Arm processors: from edge devices to HPCgrade CPUs. autoGEMM generates optimized kernels for various hardware configurations by auto-combining fragments of autogenerated micro-kernels that employ hand-written optimizations to maximize computational efficiency. We optimize the kernel pipeline by tuning the register reuse and the data load/store overlapping. In addition, we use a dynamic tiling scheme to generate balanced tile shapes. Finally, we position autoGEMM on top of the TVM framework where our dynamic tiling scheme prunes the search space for TVM to identify the optimal combination of parameters for code optimization. Evaluations on five different classes of Arm chips demonstrate the advantages of autoGEMM. For small matrices, autoGEMM achieves 98% of peak and up to 2.0x speedup over state-of-the-art libraries such as LIBXSMM and LibShalom. For irregular matrices (i.e. tall skinny and long rectangles), autoGEMM is 1.3-2.0x faster than widely-used libraries such as OpenBLAS and Eigen. autoGEMM is publicly available at: https://github.com/wudu98/autoGEMM. Du Wu, Jintao Meng 0001, Wenxi Zhu, Minwen Deng, Xiao Wang 0004, Tao Luo 0014, Mohamed Wahib, Yanjie Wei |
SC | 8 |
| 2024 | A new paradigm for applying deep learning to protein-ligand interaction predictionabstractProtein-ligand interaction prediction presents a significant challenge in drug design. Numerous machine learning and deep learning (DL) models have been developed to accurately identify docking poses of ligands and active compounds against specific targets. However, current models often suffer from inadequate accuracy or lack practical physical significance in their scoring systems. In this research paper, we introduce IGModel, a novel approach that utilizes the geometric information of protein-ligand complexes as input for predicting the root mean square deviation of docking poses and the binding strength (pKd, the negative value of the logarithm of binding affinity) within the same prediction framework. This ensures that the output scores carry intuitive meaning. We extensively evaluate the performance of IGModel on various docking power test sets, including the CASF-2016 benchmark, PDBbind-CrossDocked-Core and DISCO set, consistently achieving state-of-the-art accuracies. Furthermore, we assess IGModel's generalizability and robustness by evaluating it on unbiased test sets and sets containing target structures generated by AlphaFold2. The exceptional performance of IGModel on these sets demonstrates its efficacy. Additionally, we visualize the latent space of protein-ligand interactions encoded by IGModel and conduct interpretability analysis, providing valuable insights. This study presents a novel framework for DL-based prediction of protein-ligand interactions, contributing to the advancement of this field. The IGModel is available at GitHub repository https://github.com/zchwang/IGModel. Zechen Wang, Sheng Wang 0001, Yanjie Wei, Yuguang Mu, Liangzhen Zheng |
Briefings Bioinform. | 5 |
| 2024 | OpenDock: a pytorch-based open-source framework for protein-ligand docking and modellingabstractMOTIVATION: Molecular docking is an invaluable computational tool with broad applications in computer-aided drug design and enzyme engineering. However, current molecular docking tools are typically implemented in languages such as C++ for calculation speed, which lack flexibility and user-friendliness for further development. Moreover, validating the effectiveness of external scoring functions for molecular docking and screening within these frameworks is challenging, and implementing more efficient sampling strategies is not straightforward. RESULTS: To address these limitations, we have developed an open-source molecular docking framework, OpenDock, based on Python and PyTorch. This framework supports the integration of multiple scoring functions; some can be utilized during molecular docking and pose optimization, while others can be used for post-processing scoring. In terms of sampling, the current version of this framework supports simulated annealing and Monte Carlo optimization. Additionally, it can be extended to include methods such as genetic algorithms and particle swarm optimization for sampling docking poses and protein side chain orientations. Distance constraints are also implemented to enable covalent docking, restricted docking or distance map constraints guided pose sampling. Overall, this framework serves as a valuable tool in drug design and enzyme engineering, offering significant flexibility for most protein-ligand modelling tasks. AVAILABILITY AND IMPLEMENTATION: OpenDock is publicly available at: https://github.com/guyuehuo/opendock. Qiuyue Hu, Zechen Wang, Jintao Meng 0001, Yuguang Mu, Sheng Wang 0001, Liangzhen Zheng, Yanjie Wei |
Bioinform. | 9 |
| 2024 | SeedHit: A GPU Friendly Pre-Align Filtering AlgorithmabstractThe amount of genetic data generated by Next Generation Sequencing (NGS) technologies grows faster than Moore's law. This necessitates the development of efficient NGS data processing and analysis algorithms. A filter before the computationally-costly analysis step can significantly reduce the run time of the NGS data analysis. As GPUs are orders of magnitude more powerful than CPUs, this paper proposes a GPU-friendly pre-align filtering algorithm named SeedHit for the fast processing of NGS data. Inspired by BLAST, SeedHit counts seed hits between two sequences to determine their similarity. In SeedHit, a nucleic acid in a gene sequence is presented in binary format. By packaging data and generating a lookup table that fits into the L1 cache, SeedHit is GPU-friendly and high-throughput. Using three 16 s rRNA datasets from Greengenes as input SeedHit can reject 84%-89% dissimilar sequence pairs on average when the similarity is 0.9-0.99. The throughput of SeedHit achieved 1 T/s (Tera base per second) on 3080 Ti. Compared with the other two GPU-based filtering algorithms, GateKeeper and SneakySnake, SeedHit has the highest rejection rate and throughput. By incorporating SeedHit into our in-house clustering algorithm nGIA, the modified nGIA achieved a 1.6-2.1 times speedup compared to the original version. Zhen Ju, Xuelei Li, Jintao Meng 0001, Yanjie Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | A Novel Deep Learning Approach Featuring Graph-Based Algorithm for Cell Segmentation and TrackingabstractThe precise segmentation and tracking of cells in microscopy image sequences play a pivotal role in biomedical research, facilitating the study of tissue, organ, and organism development. However, manual segmentation and tracking of cells is time-consuming and often require professional experiences. Besides, segmenting cells in the images with a low signal-to-noise ratio remains difficult. While deep learning (DL) has become a common method for cell segmentation, few DL- based methods address concurrent cell segmentation and tracking. In this paper, we propose a novel DL approach featuring graph-based tracking for cell segmentation and tracking in microscopy images. We combine Deeplabv3+ for semantic segmentation and ResNet50 for enhanced feature extraction, enabling comprehensive cell detection and instance segmentation. Post-processing, involving non-maxima suppression and outlier detection, refines predictions and produces final segmentation. The tracking method is based on the relative position of graph nodes to track segmented cells, encompassing cell division and apoptosis. We conduct our experiments on the induced pluripotent stem (iPS) cell datasets, and the results show that the segmentation and tracking performance of our method yields superior performance compared to the benchmark models. More specifically, our approach achieved DET values of 0.955 and 0.913, TRA values of 0.951 and 0.906, and SEG values of 0.690 and 0.665 on two iPS dataset videos, respectively. Additionally, the performance assessment encompasses four real microscopy datasets from the Cell Tracking Challenge (CTC). Our method greatly reduces the cost of manual labeling and labor-intensive costs. Keliang Zhao, Jovial Niyogisubizo, Linxia Xiao, Yi Pan 0001, Didi Rosiyadi, Yanjie Wei |
BIBM | 7 |
| 2023 | PCPI: Prediction of circRNA and Protein Interaction Using Machine Learning Method
Md. Tofazzal Hossain, Md. Selim Reza, Xuelei Li, Yin Peng, Shengzhong Feng, Yanjie Wei |
ISBRA | 6 |
| 2023 | Identification and Functional Annotation of circRNAs in Neuroblastoma Based on Bioinformatics
Md. Tofazzal Hossain, Zhen Ju, Wenhui Xi, Yanjie Wei |
ISBRA | 5 |
| 2023 | A fully differentiable ligand pose optimization framework guided by deep learning and a traditional scoring functionabstractThe recently reported machine learning- or deep learning-based scoring functions (SFs) have shown exciting performance in predicting protein-ligand binding affinities with fruitful application prospects. However, the differentiation between highly similar ligand conformations, including the native binding pose (the global energy minimum state), remains challenging that could greatly enhance the docking. In this work, we propose a fully differentiable, end-to-end framework for ligand pose optimization based on a hybrid SF called DeepRMSD+Vina combined with a multi-layer perceptron (DeepRMSD) and the traditional AutoDock Vina SF. The DeepRMSD+Vina, which combines (1) the root mean square deviation (RMSD) of the docking pose with respect to the native pose and (2) the AutoDock Vina score, is fully differentiable; thus is capable of optimizing the ligand binding pose to the energy-lowest conformation. Evaluated by the CASF-2016 docking power dataset, the DeepRMSD+Vina reaches a success rate of 94.4%, which outperforms most reported SFs to date. We evaluated the ligand conformation optimization framework in practical molecular docking scenarios (redocking and cross-docking tasks), revealing the high potentialities of this framework in drug design and discovery. Structural analysis shows that this framework has the ability to identify key physical interactions in protein-ligand binding, such as hydrogen-bonding. Our work provides a paradigm for optimizing ligand conformations based on deep learning algorithms. The DeepRMSD+Vina model and the optimization framework are available at GitHub repository https://github.com/zchwang/DeepRMSD-Vina_Optimization. Zechen Wang, Liangzhen Zheng, Sheng Wang 0001, Mingzhi Lin, Adams Wai-Kin Kong, Yuguang Mu, Yanjie Wei |
Briefings Bioinform. | 8 |
| 2023 | JCcirc: circRNA full-length sequence assembly through integrated junction contigsabstractRecent studies have shed light on the potential of circular RNA (circRNA) as a biomarker for disease diagnosis and as a nucleic acid vaccine. The exploration of these functionalities requires correct circRNA full-length sequences; however, existing assembly tools can only correctly assemble some circRNAs, and their performance can be further improved. Here, we introduce a novel feature known as the junction contig (JC), which is an extension of the back-splice junction (BSJ). Leveraging the strengths of both BSJ and JC, we present a novel method called JCcirc (https://github.com/cbbzhang/JCcirc). It enables efficient reconstruction of all types of circRNA full-length sequences and their alternative isoforms using splice graphs and fragment coverage. Our findings demonstrate the superiority of JCcirc over existing methods on human simulation datasets, and its average F1 score surpasses CircAST by 0.40 and outperforms both CIRI-full and circRNAfull by 0.13. For circRNAs below 400 bp, 400-800 bp, 800 bp-1200 bp and above 1200 bp, the correct assembly rates are 0.13, 0.09, 0.04 and 0.03 higher, respectively, than those achieved by existing methods. Moreover, JCcirc also outperforms existing assembly tools on other five model species datasets and real sequencing datasets. These results show that JCcirc is a robust tool for accurately assembling circRNA full-length sequences, laying the foundation for the functional analysis of circRNAs. Zhen Ju, Yin Peng, Yi Pan 0001, Wenhui Xi, Yanjie Wei |
Briefings Bioinform. | 7 |
| 2023 | Guest Editors' Introduction to the Special Section on Bioinformatics Research and Applications
Zhipeng Cai 0001, Min Li 0007, Pavel Skums, Yanjie Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | RabbitFX: Efficient Framework for FASTA/Q File Parsing on Modern Multi-Core PlatformsabstractThe continuous growth of generated sequencing data leads to the development of a variety of associated bioinformatics tools. However, many of them are not able to fully exploit the resources of modern multi-core systems since they are bottlenecked by parsing files leading to slow execution times. This motivates the design of an efficient method for parsing sequencing data that can exploit the power of modern hardware, especially for modern CPUs with fast storage devices. We have developed RabbitFX, a fast, efficient, and easy-to-use framework for processing biological sequencing data on modern multi-core platforms. It can efficiently read FASTA and FASTQ files by combining a lightweight parsing method by means of an optimized formatting implementation. Furthermore, we provide user-friendly and modularized C++ APIs that can be easily integrated into applications in order to increase their file parsing speed. As proof-of-concept, we have integrated RabbitFX into three I/O-intensive applications: fastp, Ktrim, and Mash. Our evaluation shows that the inclusion of RabbitFX leads to speedups of at least 11.6 (6.6), 2.4 (2.4), and 3.7 (3.2) compared to the original versions on plain (gzip-compressed) files, respectively. These case studies demonstrate that RabbitFX can be easily integrated into a variety of NGS analysis tools to significantly reduce associated runtimes. It is open source software available at https://github.com/RabbitBio/RabbitFX. Hao Zhang 0142, Honglei Song, Xiaoming Xu 0004, Qixin Chang, Yanjie Wei, Zekun Yin, Bertil Schmidt |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2022 | RabbitQCPlus: More Efficient Quality Control for Sequencing DataabstractAssessing the quality of sequencing data plays a crucial role in downstream data analysis. However, existing tools often achieve sub-optimal efficiency, especially when dealing with compressed files or performing complicated quality control operations such as over-representation analysis. We present RabbitQCPlus, an ultra-efficient quality control tool for modern multi-core systems. RabbitQCPlus uses vectorization, memory copy reduction, parallel (de)compression, and optimized data structures to achieve substantial performance gains. It is 1.1 to 5.4 times faster when performing basic quality control operations compared to state-of-the-art applications yet requires fewer compute resources. Moreover, RabbitQCPlus is at least 4 times faster than other applications when processing gzip-compressed FASTQ files. Furthermore, it takes less than 4 minutes to process 280GB of plain FASTQ sequencing data, while other applications take at least 22 minutes on a 48-core server when enabling the per-read over-representation analysis. C++ sources are available at https://github.com/RabbitBio/RabbitQCPlus. Lifeng Yan, Zekun Yin, Hao Zhang 0142, Zhan Zhao, André Müller, Robin Kobus, Yanjie Wei, Beifang Niu, Bertil Schmidt |
BIBM | 8 |
| 2022 | Simulating Spiking Neural Networks Based on SW26010pro
Xuelei Li, Jintao Meng 0001, Yi Pan 0001, Yanjie Wei |
ISBRA | 5 |
| 2022 | Generating and screening de novo compounds against given targets using ultrafast deep learning models as core componentsabstractDeep learning is an artificial intelligence technique in which models express geometric transformations over multiple levels. This method has shown great promise in various fields, including drug development. The availability of public structure databases prompted the researchers to use generative artificial intelligence models to narrow down their search of the chemical space, a novel approach to chemogenomics and de novo drug development. In this study, we developed a strategy that combined an accelerated LSTM_Chem (long short-term memory for de novo compounds generation), dense fully convolutional neural network (DFCNN), and docking to generate a large number of de novo small molecular chemical compounds for given targets. To demonstrate its efficacy and applicability, six important targets that account for various human disorders were used as test examples. Moreover, using the M protease as a proof-of-concept example, we find that iteratively training with previously selected candidates can significantly increase the chance of obtaining novel compounds with higher and higher predicted binding affinities. In addition, we also check the potential benefit of obtaining reliable final de novo compounds with the help of MD simulation and metadynamics simulation. The generation of de novo compounds and the discovery of binders against various targets proposed here would be a practical and effective approach. Assessing the efficacy of these top de novo compounds with biochemical studies is promising to promote related drug development. Konda Mani Saravanan, Yanjie Wei, Yi Pan 0001, John Z. H. Zhang |
Briefings Bioinform. | 4 |
| 2022 | Improving protein-ligand docking and screening accuracies by incorporating a scoring function correction termabstractScoring functions are important components in molecular docking for structure-based drug discovery. Traditional scoring functions, generally empirical- or force field-based, are robust and have proven to be useful for identifying hits and lead optimizations. Although multiple highly accurate deep learning- or machine learning-based scoring functions have been developed, their direct applications for docking and screening are limited. We describe a novel strategy to develop a reliable protein-ligand scoring function by augmenting the traditional scoring function Vina score using a correction term (OnionNet-SFCT). The correction term is developed based on an AdaBoost random forest model, utilizing multiple layers of contacts formed between protein residues and ligand atoms. In addition to the Vina score, the model considerably enhances the AutoDock Vina prediction abilities for docking and screening tasks based on different benchmarks (such as cross-docking dataset, CASF-2016, DUD-E and DUD-AD). Furthermore, our model could be combined with multiple docking applications to increase pose selection accuracies and screening abilities, indicating its wide usage for structure-based drug discoveries. Furthermore, in a reverse practice, the combined scoring strategy successfully identified multiple known receptors of a plant hormone. To summarize, the results show that the combination of data-driven model (OnionNet-SFCT) and empirical scoring function (Vina score) is a good scoring strategy that could be useful for structure-based drug discoveries and potentially target fishing in future. Liangzhen Zheng, Jintao Meng 0001, Haidong Lan, Zechen Wang, Mingzhi Lin, Yanjie Wei, Yuguang Mu |
Briefings Bioinform. | 9 |
| 2022 | RabbitV: fast detection of viruses and microorganisms in sequencing data on multi-core architecturesabstractMOTIVATION: Detection and identification of viruses and microorganisms in sequencing data plays an important role in pathogen diagnosis and research. However, existing tools for this problem often suffer from high runtimes and memory consumption. RESULTS: We present RabbitV, a tool for rapid detection of viruses and microorganisms in Illumina sequencing datasets based on fast identification of unique k-mers. It can exploit the power of modern multi-core CPUs by using multi-threading, vectorization and fast data parsing. Experiments show that RabbitV outperforms fastv by a factor of at least 42.5 and 14.4 in unique k-mer generation (RabbitUniq) and pathogen identification (RabbitV), respectively. Furthermore, RabbitV is able to detect COVID-19 from 40 samples of sequencing data (255 GB in FASTQ format) in only 320 s. AVAILABILITY AND IMPLEMENTATION: RabbitUniq and RabbitV are available at https://github.com/RabbitBio/RabbitUniq and https://github.com/RabbitBio/RabbitV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hao Zhang 0142, Qixin Chang, Zekun Yin, Xiaoming Xu 0004, Yanjie Wei, Bertil Schmidt |
Bioinform. | 5 |
| 2022 | nGIA: A novel Greedy Incremental Alignment based algorithm for gene sequence clustering
Zhen Ju, Jintao Meng 0001, Jianping Fan 0002, Yi Pan 0001, Xuelei Li, Yanjie Wei |
Future Gener. Comput. Syst. | 9 |
| 2022 | Automatic Generation of High-Performance Convolution Kernels on ARM CPUs for Deep LearningabstractWe presentFastConv, a template-based code auto-generation open-source library that can automatically generate high-performance deep learning convolution kernels of arbitrary matrices/tensors shapes. FastConv is based on the Winograd algorithm, which is reportedly the highest performing algorithm for the time-consuming layers of convolutional neural networks. ARM CPUs cover a wide range of designs and specifications, from embedded devices to HPC-grade CPUs. The leads to the dilemma of how to consistently optimize Winograd-based convolution solvers for convolution layers of different shapes. FastConv addresses this problem by using templates to auto-generate multiple shapes of tuned kernels variants suitable for skinny tall matrices. As a performance portable library, FastConv transparently searches for the best combination of kernel shapes, cache tiles, scheduling of loop orders, packing strategies, access patterns, and online/offline computations. Auto-tuning is used to search the parameter configuration space for the best performance for a given target architecture and problem size. Results show 1.02x to 1.40x, 1.14x to 2.17x, and 1.22x and 2.48x speedup is achieved over NNPACK, ARM NN, and FeatherCNN on Kunpeng 920. Furthermore, performance portability experiments with various convolution shapes show that FastConv achieves 1.2x to 1.7x speedup and 2x to 22x speedup over NNPACK and ARM NN inference engine using Winograd on Kunpeng 920. CPU performance portability evaluation on VGG–16 show an average speedup over NNPACK of 1.42x, 1.21x, 1.26x, 1.37x, 2.26x, and 11.02x on Kunpeng 920, Snapdragon 835, 855, 888, Apple M1, and AWS Graviton2, respectively. Jintao Meng 0001, Chen Zhuang, Peng Chen 0035, Mohamed Wahib, Bertil Schmidt, Xiao Wang 0004, Haidong Lan, Dou Wu, Minwen Deng, Yanjie Wei, Shengzhong Feng |
IEEE Trans. Parallel Distributed Syst. | 10 |
| 2021 | A novel virtual drug screening pipeline with deep-leaning as core component identifies inhibitor of pancreatic alpha-amylaseabstractVirtual drug screening that provides possible drug candidates facilitates early-stage drug discovery. It works by large scale predicting native-like protein-ligand complexes (PLC) from an abundance of docking decoys. Many affinity predicting models currently in use fail to provide reliable prediction because of a lack of non-binding data during model training, lost critical physical-chemical features, and difficulties in learning abstract information with limited neural layers. In this paper, we developed a deep learning model, DeepBindBC for classifying putative ligands as binding or non-binding. Our model incorporates information of non-binding interactions, making it more suitable for real applications. ResNet model architecture and more detailed atom type representation guarantee implicit features can be learned more accurately. DeepBindBC identified a novel human pancreatic $\alpha$-amylase binder validated by a fluorescence spectral experiment (Ka $=1.0\times 10^{5}\mathrm{M}$). Furthermore, we proposed a virtual screening pipeline by incorporating multiple complementary methods, such as DFCNN, Autodock vina docking, DeepBindBC, and pocket molecular dynamics simulation. Three potential inhibitors of pancreatic $\alpha$-amylase were identified by the proposed pipeline, and interestingly most of them contain glycan groups. Additionally, an online webserver based on the model is available at http://cbblab.siat.ac.cn/DeepBindBC/index.php for the convenience of the users. Konda Mani Saravanan, Linbu Liao, Hao Wu 0003, Haishan Zhang, Yi Pan 0001, Xuli Wu, Yanjie Wei |
BIBM | 10 |
| 2021 | The Classification System and Biomarkers for Autism Spectrum Disorder: A Machine Learning Approach
Zhongyang Dai, Haishan Zhang, Feifei Lin, Shengzhong Feng, Yanjie Wei, Jiaxiu Zhou |
ISBRA | 5 |
| 2021 | An Efficient Greedy Incremental Sequence Clustering Algorithm
Zhen Ju, Jingtao Meng, Xuelei Li, Jianping Fan 0002, Yi Pan 0001, Yanjie Wei |
ISBRA | 9 |
| 2021 | RabbitMash: accelerating hash-based genome analysis on modern multi-core architecturesabstractMOTIVATION: Mash is a popular hash-based genome analysis toolkit with applications to important downstream analyses tasks such as clustering and assembly. However, Mash is currently not able to fully exploit the capabilities of modern multi-core architectures, which in turn leads to high runtimes for large-scale genomic datasets. RESULTS: We present RabbitMash, an efficient highly optimized implementation of Mash which can take full advantage of modern hardware including multi-threading, vectorization and fast I/O. We show that our approach achieves speedups of at least 1.3, 9.8, 8.5 and 4.4 compared to Mash for the operations sketch, dist, triangle and screen, respectively. Furthermore, RabbitMash is able to compute the all-versus-all distances of 100 321 genomes in <5 min on a 40-core workstation while Mash requires over 40 min. AVAILABILITY AND IMPLEMENTATION: RabbitMash is available at https://github.com/ZekunYin/RabbitMash. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zekun Yin, Xiaoming Xu 0004, Jinxiao Zhang, Yanjie Wei, Bertil Schmidt |
Bioinform. | 4 |
| 2021 | RabbitQC: high-speed scalable quality control for sequencing dataabstractMOTIVATION: Modern sequencing technologies continue to revolutionize many areas of biology and medicine. Since the generated datasets are error-prone, downstream applications usually require quality control methods to pre-process FASTQ files. However, existing tools for this task are currently not able to fully exploit the capabilities of computing platforms leading to slow runtimes. RESULTS: We present RabbitQC, an extremely fast integrated quality control tool for FASTQ files, which can take full advantage of modern hardware. It includes a variety of operations and supports different sequencing technologies (Illumina, Oxford Nanopore and PacBio). RabbitQC achieves speedups between one and two orders-of-magnitude compared to other state-of-the-art tools. AVAILABILITY AND IMPLEMENTATION: C++ sources and binaries are available at https://github.com/ZekunYin/RabbitQC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zekun Yin, Hao Zhang 0142, Meiyang Liu, Honglei Song, Haidong Lan, Yanjie Wei, Beifang Niu, Bertil Schmidt |
Bioinform. | 7 |
| 2021 | Evaluation of residue-residue contact prediction methods: From retrospective to prospectiveabstractSequence-based residue contact prediction plays a crucial role in protein structure reconstruction. In recent years, the combination of evolutionary coupling analysis (ECA) and deep learning (DL) techniques has made tremendous progress for residue contact prediction, thus a comprehensive assessment of current methods based on a large-scale benchmark data set is very needed. In this study, we evaluate 18 contact predictors on 610 non-redundant proteins and 32 CASP13 targets according to a wide range of perspectives. The results show that different methods have different application scenarios: (1) DL methods based on multi-categories of inputs and large training sets are the best choices for low-contact-density proteins such as the intrinsically disordered ones and proteins with shallow multi-sequence alignments (MSAs). (2) With at least 5L (L is sequence length) effective sequences in the MSA, all the methods show the best performance, and methods that rely only on MSA as input can reach comparable achievements as methods that adopt multi-source inputs. (3) For top L/5 and L/2 predictions, DL methods can predict more hydrophobic interactions while ECA methods predict more salt bridges and disulfide bonds. (4) ECA methods can detect more secondary structure interactions, while DL methods can accurately excavate more contact patterns and prune isolated false positives. In general, multi-input DL methods with large training sets dominate current approaches with the best overall performance. Despite the great success of current DL methods must be stated the fact that there is still much room left for further improvement: (1) With shallow MSAs, the performance will be greatly affected. (2) Current methods show lower precisions for inter-domain compared with intra-domain contact predictions, as well as very high imbalances in precisions between intra-domains. (3) Strong prediction similarities between DL methods indicating more feature types and diversified models need to be developed. (4) The runtime of most methods can be further optimized. Zhendong Bei, Wenhui Xi, Min Hao 0002, Zhen Ju, Konda Mani Saravanan, Yanjie Wei |
PLoS Comput. Biol. | 9 |
| 2020 | Protein Interresidue Contact Prediction Based on Deep Learning and Massive Features from Multi-sequence Alignment
Hing-Fung Ting, Yanjie Wei |
PDCAT | 4 |
| 2020 | A novel virtual screening procedure identifies Pralatrexate as inhibitor of SARS-CoV-2 RdRp and it reduces viral replication in vitroabstractThe spread of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) virus poses serious threats to the global public health and leads to worldwide crisis. No effective drug or vaccine is readily available. The viral RNA-dependent RNA polymerase (RdRp) is a promising therapeutic target. A hybrid drug screening procedure was proposed and applied to identify potential drug candidates targeting RdRp from 1906 approved drugs. Among the four selected market available drug candidates, Pralatrexate and Azithromycin were confirmed to effectively inhibit SARS-CoV-2 replication in vitro with EC50 values of 0.008μM and 9.453 μM, respectively. For the first time, our study discovered that Pralatrexate is able to potently inhibit SARS-CoV-2 replication with a stronger inhibitory activity than Remdesivir within the same experimental conditions. The paper demonstrates the feasibility of fast and accurate anti-viral drug screening for inhibitors of SARS-CoV-2 and provides potential therapeutic agents against COVID-19. Junxin Li, Konda Mani Saravanan, Jinli Wei, Justin Tze-Yang Ng, Md. Tofazzal Hossain, Maoxuan Liu, Xiaohu Ren, Yi Pan 0001, Yin Peng, Xiaochun Wan, Yingxia Liu, Yanjie Wei |
PLoS Comput. Biol. | 17 |
| 2019 | A novel machine learning based approach for iPS progenitor cell identificationabstractIdentification of induced pluripotent stem (iPS) progenitor cells, the iPS forming cells in early stage of reprogramming, could provide valuable information for studying the origin and underlying mechanism of iPS cells. However, it is very difficult to identify experimentally since there are no biomarkers known for early progenitor cells, and only about 6 days after reprogramming initiation, iPS cells can be experimentally determined via fluorescent probes. What is more, the ratio of progenitor cells during early reprograming period is below 5%, which is too low to capture experimentally in the early stage. In this paper, we propose a novel computational approach for the identification of iPS progenitor cells based on machine learning and microscopic image analysis. Firstly, we record the reprogramming process using a live cell imaging system after 48 hours of infection with retroviruses expressing Oct4, Sox2 and Klf4, later iPS progenitor cells and normal murine embryonic fibroblasts (MEFs) within 3 to 5 days after infection are labeled by retrospectively tracing the time-lapse microscopic image. We then calculate 11 types of cell morphological and motion features such as area, speed, etc., and select best time windows for modeling and perform feature selection. Finally, a prediction model using XGBoost is built based on the selected six types of features and best time windows. Our model allows several missing values/frames in the sample datasets, thus it is applicable to a wide range of scenarios. Cross-validation, holdout validation and independent test experiments show that the minimum precision is above 52%, that is, the ratio of predicted progenitor cells within 3 to 5 days after viral infection is above 52%. The results also confirm that the morphology and motion pattern of iPS progenitor cells is different from that of normal MEFs, which helps with the machine learning methods for iPS progenitor cell identification. Haishan Zhang, Ximing Shao, Yin Peng, Yanning Teng, Konda Mani Saravanan, Hongchang Li, Yanjie Wei |
PLoS Comput. Biol. | 8 |
| 2018 | SPECTR: Scalable Parallel Short Read Error Correction on Multi-core and Many-core ArchitecturesabstractModern high throughput sequencing platforms can produce large amounts of short read DNA data at low cost. Error correction is an important but time-consuming initial step when processing this data in order to improve the quality of downstream analyses. In this paper, we present a Scalable Parallel Error CorrecToR designed to improve the throughput of DNA error correction for Illumina reads on various parallel platforms. Our design is based on a k-spectrum approach where a Bloom filter is frequently probed as a key operation and is optimized towards AVX-512-based multi-core CPUs, Xeon Phi many-cores (both KNC and KNL), and heterogeneous compute clusters. A number of architecture-specific optimizations are employed to achieve high performance such as memory alignment, vectorized Bloom filter probing, and a stack-based iteration to eliminate recursion. Our experiments show that our optimizations result in speedups of up to 2.8, 5.2, and 9.3 on a CPU (Xeon W-2123), a KNC-based Xeon Phi (31S1P), and a KNL-based Xeon Phi (7210), respectively, compared to a multi-threaded CPU reference implementation for the error correction stage. Furthermore, when executed on the same hardware, SPECTR achieves a speedup of up to 1.7, 2.1, 2.4, and 6.4, compared to the state-of-the-art tools Lighter, BLESS2, RECKONER, and Musket, respectively. In addition, our MPI implementation exhibits an efficiency of around 86% when executed on 32 nodes of the Tianhe-2 supercomputer. SPECTR is available at https://github.com/Xu-Kai/SPECTR. Robin Kobus, Yuandong Chan, Ping Gao 0005, Xiangxu Meng, Yanjie Wei, Bertil Schmidt |
ICPP | 6 |
| 2017 | Scalable Assembly for Massive Genomic GraphsabstractScientists increasingly want to assemble large genomes, metagenomes, and large numbers of individual genomes. In order to meet the demand for processing these huge datasets, parallel genome assembly is a vital step. Among all the parallel genome assemblers, de Bruijn graph based ones are most popular. However, the size of de Bruijn graph is determined by the number of distinct kmers used in the algorithm, thus redundant kmers in the genome datasets donot contribute to the graph size. The scalability of genome assemblers is influenced directly by the distinct kmers in the dataset or de Bruijn graph size, rather than the input dataset size. In order to assembly large genomes, we have artificially created 16 datasets of 4 Terabytes in total from the human reference genome. The human reference genome is firstly mutated with a 5% mutation rate, and then subjected to a genome sequencing data simulator ART. The simulated datasets have linearly increasing number of distinct kmers as the size/number of the combined datasets increases. We then evaluate all five time-consuming steps of the SWAP-Assembler 2.0 (SWAP2) using these 16 simulated datasets. Compared with our previous experiment on 1000 human dataset with fixed de Bruijn graph size, the weak-scaling test shows that SWAP2 can scale well from 1024 cores using one dataset to 16,384 cores. The percentage of time usage for all five steps of SWAP2 is fixed, and total time usage is also constant. The result showed that the time usage of graph simplification occupied almost 75% of the total time usage, which will be subject to further optimization for future work. Jintao Meng 0001, Jianqiu Ge, Yanjie Wei, Pavan Balaji, Bingqiang Wang |
CCGrid | 4 |
| 2017 | Bloomfish: A Highly Scalable Distributed K-mer Counting FrameworkabstractK-mer counting is a fundamental operation in DNA research and genome analytics; its application includes estimating genome assembly, understanding similarities in genomic samples, and merging a newly processed genome with a reference genome. As the genome dataset becomes larger and larger, designing a highly optimized distributed-memory implementation becomes more and more important. Current distributed-memory solutions have two limitations: they have a high memory footprint, and they do not provide advanced optimizations for loading enormous genome datasets into memory. Based on these observations, we present Bloomfish, a distributed, memory-efficient, scalable solution to the limits of current work. To keep a low memory footprint, Bloomfish leverages the compact hash array design of the single-node Jellyfish system and the optimized workflow of the high-performance MapReduce framework Mimir. We have also codesigned Mimir's I/O to efficiently load enormous datasets. We ran Bloomfish on the Tianhe-2 supercomputer with large sequence datasets (up to 24 TB). Our results show that Bloomfish achieves unprecedented scalability in genome analytics. Yanfei Guo, Yanjie Wei, Bingqiang Wang, Yutong Lu, Pietro Cicotti, Pavan Balaji, Michela Taufer |
ICPADS | 3 |
| 2016 | SWAP-Assembler 2: Optimization of De Novo Genome Assembler at Extreme ScaleabstractIn this paper, we analyze and optimize the most time-consuming steps of the SWAP-Assembler, a parallel genome assembler, so that it can scale to a large number of cores for huge genomes with sequencing data ranging from terabyes to petabytes. Performance analysis results show that the most time-consuming steps are input parallelization, k-mer graph construction, and graph simplification (edge merging). For the input parallelization, the input data is divided into virtual fragments with nearly equal size, and the start position and end position of each fragment are automatically separated at the beginning of the reads. In k-mer graph construction, in order to improve the communication efficiency, the message size is kept constant between any two processes by proportionally increasing the number of nucleotides to the number of processes in the input parallelization step for each round. The memory usage is also decreased because only a small part of the input data is processed in each round. With graph simplification, the communication protocol reduces the number of communication loops from four to two loops and decreases the idle communication time. The optimized assembler is denoted SWAP-Assembler 2 (SWAP2). In our experiments using a 1000 Genomes project dataset of 4 terabytes (the largest dataset ever used for assembling) on the supercomputer Mira, the results show that SWAP2 scales to 131,072 cores with an efficiency of 40%. We also compared our work with both the HipMer assembler and the SWAP-Assembler. On the Yanhuang dataset of 300 gigabytes, SWAP2 shows a 3X speedup and 4X better scalability compared with the HipMer assembler and is 45 times faster than the SWAP-Assembler. The SWAP2 software is available at https://sourceforge.net/projects/swapassembler. Jintao Meng 0001, Pavan Balaji, Yanjie Wei, Bingqiang Wang, Shengzhong Feng |
ICPP | 4 |
| 2015 | SWAP-Assembler 2: Scalable Genome Assembler towards Millions of Cores - Practice and ExperienceabstractThere is widening gap between the throughput of massive parallel sequencing machines and the ability to analyze these huge sequencing data, which can be Tara bytes or even Peta bytes. Previously our assembly tool, SWAP-Assembler, can scale to 2048 cores on TianHe 1A for human Yanhuang genome. This work is to further scale SWAP-Assembler to millions of cores on Mira. SWAP-Assembler can be divided into 5 steps, and the most time consuming steps are input parallelization, kmer graph construction, graph simplification (edge merging). We optimize these three steps to keep the percentage of time usage in each step constant when the number of cores increases. For the input parallelization step, the input data is divided into virtual fragments with almost equal size, the begin position and end position for each fragment is automatically separated at the beginning symbol of reads. This data blocking strategy plays a central role in adjusting the data size to keep the communication and memory efficiency for the subsequent steps. In kmer graph construction, to prevent the communication efficiency degradation, the message size is kept constant (about 8k bytes) between any two processes by proportionally increasing the number of nucleotides to the number of processes in the input parallelization step in each round. The memory usage can be also benefited, as only a small part of the input data is processed in each round. Within graph simplification, the major improvement is to combine messages sending & receiving between its two neighbors into one loop in the communication protocol. After integrated with the above optimizations, the new assembly tool is denoted as SWAP-Assembler 2 or SWAP2 for short. In our experiment for 1k human genome dataset, the modified SWAP-Assembler 2 can scale to 16k cores with parallel efficiency of 70%. Jintao Meng 0001, Yanjie Wei, Pavan Balaji |
CCGRID | 2 |
| 2015 | MPI+Threads: runtime contention and remediesabstractHybrid MPI+Threads programming has emerged as an alternative model to the “MPI everywhere” model to better handle the increasing core density in cluster nodes. While the MPI standard allows multithreaded concurrent communication, such flexibility comes with the cost of maintaining thread safety within the MPI implementation, typically implemented using critical sections. In contrast to previous works that studied the importance of critical-section granularity in MPI implementations, in this paper we investigate the implication of critical-section arbitration on communication performance. We first analyze the MPI runtime when multithreaded concurrent communication takes place on hierarchical memory systems. Our results indicate that the mutex-based approach that most MPI implementations use today can incur performance penalties due to unfair arbitration. We then present methods to mitigate these penalties with a first-come, first-served arbitration and a priority locking scheme that favors threads doing useful work. Through evaluations using several benchmarks and applications, we demonstrate up to 5-fold improvement in performance. Abdelhalim Amer, Huiwei Lu, Yanjie Wei, Pavan Balaji, Satoshi Matsuoka |
PPoPP | 3 |
| 2014 | SWAP-Assembler: scalable and efficient genome assembly towards thousands of coresabstractBACKGROUND: There is a widening gap between the throughput of massive parallel sequencing machines and the ability to analyze these sequencing data. Traditional assembly methods requiring long execution time and large amount of memory on a single workstation limit their use on these massive data. RESULTS: This paper presents a highly scalable assembler named as SWAP-Assembler for processing massive sequencing data using thousands of cores, where SWAP is an acronym for Small World Asynchronous Parallel model. In the paper, a mathematical description of multi-step bi-directed graph (MSG) is provided to resolve the computational interdependence on merging edges, and a highly scalable computational framework for SWAP is developed to automatically preform the parallel computation of all operations. Graph cleaning and contig extension are also included for generating contigs with high quality. Experimental results show that SWAP-Assembler scales up to 2048 cores on Yanhuang dataset using only 26 minutes, which is better than several other parallel assemblers, such as ABySS, Ray, and PASHA. Results also show that SWAP-Assembler can generate high quality contigs with good N50 size and low error rate, especially it generated the longest N50 contig sizes for Fish and Yanhuang datasets. CONCLUSIONS: In this paper, we presented a highly scalable and efficient genome assembly software, SWAP-Assembler. Compared with several other assemblers, it showed very good performance in terms of scalability and contig quality. This software is available at: https://sourceforge.net/projects/swapassembler. Jintao Meng 0001, Bingqiang Wang, Yanjie Wei, Shengzhong Feng, Pavan Balaji |
BMC Bioinform. | 3 |
| 2013 | An Energy Efficient Clustering Scheme for Data Aggregation in Wireless Sensor Networks
Jintao Meng 0001, Jian-Rui Yuan, Shengzhong Feng, Yanjie Wei |
J. Comput. Sci. Technol. | 4 |
| 2012 | DGraph: Algorithms for Shortgun Reads Assembly Using De Bruijn Graph
Jintao Meng 0001, Jianrui Yuan, Jiefeng Cheng, Yanjie Wei, Shengzhong Feng |
NPC | 4 |
| 2012 | Small World Asynchronous Parallel Model for Genome Assembly
Jintao Meng 0001, Jianrui Yuan, Jiefeng Cheng, Yanjie Wei, Shengzhong Feng |
NPC | 4 |
| 2012 | A novel hierarchical clustering algorithm for gene sequencesabstractBACKGROUND: Clustering DNA sequences into functional groups is an important problem in bioinformatics. We propose a new alignment-free algorithm, mBKM, based on a new distance measure, DMk, for clustering gene sequences. This method transforms DNA sequences into the feature vectors which contain the occurrence, location and order relation of k-tuples in DNA sequence. Afterwards, a hierarchical procedure is applied to clustering DNA sequences based on the feature vectors. RESULTS: The proposed distance measure and clustering method are evaluated by clustering functionally related genes and by phylogenetic analysis. This method is also compared with BlastClust, CD-HIT-EST and some others. The experimental results show our method is effective in classifying DNA sequences with similar biological characteristics and in discovering the underlying relationship among the sequences. CONCLUSIONS: We introduced a novel clustering algorithm which is based on a new sequence similarity measure. It is effective in classifying DNA sequences with similar biological characteristics and in discovering the relationship among the sequences. Qingshan Jiang, Yanjie Wei, Shengrui Wang |
BMC Bioinform. | 3 |