EDBT 2026 Demo / reviewers in the wild / expert
Jintao Meng 0001
dblp:56/3390-1
· DBLP profile ↗
28ranked-venue papers
8as first author
19since 2021 · last 2025
0000-0002-6208-4102ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Task-Adaptive Refined Reinforcement Learning with Granular Reward Shaping for Biomedical Information ExtractionabstractThe surge of biomedical literature and omics data calls for automated knowledge extraction, and LLMs show strong potential for this task. However, existing LLMs still struggle with structured tasks, as they are primarily optimized for generating free-form text rather than adhering to schema-constrained outputs. LLM alignment methods often fail to generalize across biomedical tasks due to semantic ambiguity and unstable policy optimization. To address these challenges, we introduce GRASP (Group-Relative Adaptive Structured Prompting), a unified framework for biomedical information extraction. GRASP combines a hierarchical task-aware prompt design that explicitly encodes task semantics with a novel Group-Relative Policy Optimization strategy, enabling fine-grained, semantically sensitive reward modeling. This approach resolves task inter-ference and enhances the fidelity of structured outputs across diverse extraction tasks. Extensive experiments across biomedical and general benchmarks demonstrate that GRASP achieves state-of-the-art performance in structural accuracy and semantic consistency, with up to 11 % and 7 % relative improvements in Micro F1 for NER and relation extraction, respectively, over competitive baselines. Our code is publicly available at GitHub. Qiucheng Miao, Jintao Meng 0001, Yanjie Wei |
BIBM | 3 |
| 2025 | FlexiCell: Deep Learning with Learnable Adaptive Filtering and Dual Attention for Cell SegmentationabstractAccurate cell segmentation remains challenging due to morphological variations, diverse imaging modalities, and unclear cellular boundaries. Existing deep learning (DL) methods struggle to extract features adaptively across heterogeneous cellular environments, thereby limiting generalization capacity. To address these challenges, we propose FlexiCell, a novel adaptive segmentation framework that integrates a learnable adaptive filter with dual attention mechanisms. The core innovation lies in the FlexiFilter approach, which combines standard convolution with adaptive residual learning through learnable mixing parameters. These parameters dynamically balance input preservation and feature enhancement. FlexiCell employs multi-scale FlexiFilter blocks with varying kernel sizes, channel and spatial attention networks, and a dedicated boundary extractor for precise edge detection. Extensive experiments demonstrate superior performance compared to benchmark models, achieving 3.8% improvement in detection accuracy and 5.5 % in segmentation quality on our newly developed induced pluripotent stem (iPS) cell datasets. Further evaluation on standardized Cell Tracking Challenge (CTC) benchmarks confirms state-of-the-art performance on mesenchymal stem cells and glioblastoma datasets, outperforming established CTC methods. The framework demonstrates robust generalization across fluorescence, phase contrast, and differential interference contrast microscopy, without requiring dataset-specific optimization. Codes are available at https://github.com/jovialniyo93/FlexiCell. Jovial Niyogisubizo, Keliang Zhao, Shengqi Zhou, Rui-Ze Han, Jintao Meng 0001, Wenhui Xi, Yanjie Wei |
BIBM | 5 |
| 2025 | NM-SpMM: Accelerating Matrix Multiplication Using N: M Sparsity with GPGPUabstractDeep learning demonstrates effectiveness across a wide range of tasks. However, the dense and over-parameterized nature of these models results in significant resource consumption during deployment. In response to this issue, weight pruning, particularly through$N: M$sparsity matrix multiplication, offers an efficient solution by transforming dense operations into semisparse ones.$N: M$sparsity provides an option for balancing performance and model accuracy, but introduces more complex programming and optimization challenges. To address these issues, we design a systematic top-down performance analysis model for$N: M$sparsity. Meanwhile, NM-SpMM is proposed as an efficient general$N: M$sparsity implementation. Based on our performance analysis, NM-SpMM employs a hierarchical blocking mechanism as a general optimization to enhance data locality, while memory access optimization and pipeline design are introduced as sparsity-aware optimization, allowing it to achieve close-to-theoretical peak performance across different sparsity levels. Experimental results show that NM-SpMM is 2.1x faster than nmSPARSE (the state-of-the-art for general$N: M$sparsity) and 1.4× to 6.3× faster than cuBLAS's dense GEMM operations, closely approaching the theoretical maximum speedup resulting from the reduction in computation due to sparsity. NM-SpMM is open source and publicly available at https://github.com/M-H482/NM-SpMM. Du Wu, Zhelang Deng, Jintao Meng 0001, Wenxi Zhu, Bingqiang Wang, Amelie Chi Zhou, Peng Chen 0035, Minwen Deng, Yanjie Wei, Shengzhong Feng, Yi Pan 0001 |
IPDPS | 6 |
| 2025 | An Efficient Parallel List Ranking Algorithm for Graph Concatenation on BSP Graph System
Maocheng Cao, Zhelang Deng, Qiucheng Miao, Jintao Meng 0001, Yanjie Wei, Jiefeng Cheng |
ISBRA (2) | 4 |
| 2025 | A Sample-Free Compilation Framework for Efficient Dynamic Tensor ComputationabstractDynamic-shape tensor computation poses challenges for shape-specific compilation due to variable input dimensions. Existing compilers rely on shape samples, incurring high tuning costs and performance degradation on unseen inputs. We present Helix, a dynamic tensor compilation framework with sample-free compilation and architecture-guided optimization to achieve both compilation efficiency and shape-general performance. To avoid shape sampling, Helix constructs shape-agnostic compilation by decomposing computations across architectural layers. A bidirectional strategy combines top-down abstraction to align tensor computations with architectural hierarchies, and bottom-up kernel construction to build efficient execution strategies from reusable, architecture-aligned micro-kernels. A hybrid analyzer ensures accuracy through profiling at lower architectural levels, and achieves scalability through architecture-informed modeling at higher levels and runtime. This hierarchical design eliminates shape-specific tuning and enables shape-adaptive execution. Evaluations conducted on x86 CPUs, ARM CPUs, and NVIDIA GPUs demonstrate that Helix reduces compilation time by 174 × over the existing compilers and delivers 2.26 × and 3.29 × execution speedups over vendor libraries and dynamic-shape compilers, respectively. Yangjie Zhou 0001, Weihao Cui, Zihan Liu 0002, Peng Chen 0035, Mohamed Wahib, Cong Guo 0003, Siyuan Feng 0007, Jintao Meng 0001, Haidong Lan, Jingwen Leng, Yun Lin 0001, Jin Song Dong 0001, Wenxi Zhu, Minwen Deng |
SC | 10 |
| 2025 | CircRNA Profiles Analysis of Neuroblastoma for Identification of Drug TargetsabstractNeuroblastoma is a prevalent pediatric tumor with a low 5-year survival rate among high-risk patients, and the prognosis remains poor despite available therapeutic interventions. Therefore, identifying novel and effective therapeutic targets is critical for improving outcomes in these patients. In this study, we performed an integrative analysis of two neuroblastoma circRNA sequencing datasets to identify potential drug targets. By comparing circRNA expression levels between neuroblastoma tissues and adjacent normal tissues, we identified differentially expressed circRNAs and subsequently predicted 30 hub circRNAs through Weighted Gene Co-expression Network Analysis. To elucidate the functional roles of these circRNAs, we investigated their interactions with RNA-binding proteins. The results suggest that hsa_circ_0051680 and hsa_circ_0006107 may influence neuroblastoma progression through interactions with FUS and IGF2BP1, respectively. Furthermore, we analyzed the translational potential of the hub circRNAs, revealing that six circRNAs encode proteins with complex secondary structures. Molecular docking analysis identified five high-affinity complexes between circRNA-encoded proteins (hsa_circ_0000786, hsa_circ_0005087, hsa_circ_0006867) and their corresponding ligands. These circRNA-derived proteins present promising novel drug targets for both the diagnosis and treatment of neuroblastoma. Zhen Ju, Godfrey Chi-Fung Chan, Jintao Meng 0001, Wenhui Xi, Yanjie Wei |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2024 | An In-Depth Assessment of Sequence Clustering Software in Bioinformatics
Zhen Ju, Xuelei Li, Jintao Meng 0001, Wenhui Xi, Yanjie Wei |
ISBRA (1) | 4 |
| 2024 | autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm ArchitecturesabstractThis paper presents an open-source library that pushes the limits of performance portability for irregular General Matrix Multiplication (GEMM) on the widely-used Arm architectures. Our library, autoGEMM, is designed to support a wide range of Arm processors: from edge devices to HPCgrade CPUs. autoGEMM generates optimized kernels for various hardware configurations by auto-combining fragments of autogenerated micro-kernels that employ hand-written optimizations to maximize computational efficiency. We optimize the kernel pipeline by tuning the register reuse and the data load/store overlapping. In addition, we use a dynamic tiling scheme to generate balanced tile shapes. Finally, we position autoGEMM on top of the TVM framework where our dynamic tiling scheme prunes the search space for TVM to identify the optimal combination of parameters for code optimization. Evaluations on five different classes of Arm chips demonstrate the advantages of autoGEMM. For small matrices, autoGEMM achieves 98% of peak and up to 2.0x speedup over state-of-the-art libraries such as LIBXSMM and LibShalom. For irregular matrices (i.e. tall skinny and long rectangles), autoGEMM is 1.3-2.0x faster than widely-used libraries such as OpenBLAS and Eigen. autoGEMM is publicly available at: https://github.com/wudu98/autoGEMM. Du Wu, Jintao Meng 0001, Wenxi Zhu, Minwen Deng, Xiao Wang 0004, Tao Luo 0014, Mohamed Wahib, Yanjie Wei |
SC | 2 |
| 2024 | REXIO: Indexing for Low Write Amplification by Reducing Extra I/Os in Key-Value Store Under Mixed Read/Write Workloads
Qiang Qu 0001, Nan Han, Zhelang Deng, Yizhuo Ma, Jintao Meng 0001 |
WISE (1) | 7 |
| 2024 | OpenDock: a pytorch-based open-source framework for protein-ligand docking and modellingabstractMOTIVATION: Molecular docking is an invaluable computational tool with broad applications in computer-aided drug design and enzyme engineering. However, current molecular docking tools are typically implemented in languages such as C++ for calculation speed, which lack flexibility and user-friendliness for further development. Moreover, validating the effectiveness of external scoring functions for molecular docking and screening within these frameworks is challenging, and implementing more efficient sampling strategies is not straightforward. RESULTS: To address these limitations, we have developed an open-source molecular docking framework, OpenDock, based on Python and PyTorch. This framework supports the integration of multiple scoring functions; some can be utilized during molecular docking and pose optimization, while others can be used for post-processing scoring. In terms of sampling, the current version of this framework supports simulated annealing and Monte Carlo optimization. Additionally, it can be extended to include methods such as genetic algorithms and particle swarm optimization for sampling docking poses and protein side chain orientations. Distance constraints are also implemented to enable covalent docking, restricted docking or distance map constraints guided pose sampling. Overall, this framework serves as a valuable tool in drug design and enzyme engineering, offering significant flexibility for most protein-ligand modelling tasks. AVAILABILITY AND IMPLEMENTATION: OpenDock is publicly available at: https://github.com/guyuehuo/opendock. Qiuyue Hu, Zechen Wang, Jintao Meng 0001, Yuguang Mu, Sheng Wang 0001, Liangzhen Zheng, Yanjie Wei |
Bioinform. | 3 |
| 2024 | SeedHit: A GPU Friendly Pre-Align Filtering AlgorithmabstractThe amount of genetic data generated by Next Generation Sequencing (NGS) technologies grows faster than Moore's law. This necessitates the development of efficient NGS data processing and analysis algorithms. A filter before the computationally-costly analysis step can significantly reduce the run time of the NGS data analysis. As GPUs are orders of magnitude more powerful than CPUs, this paper proposes a GPU-friendly pre-align filtering algorithm named SeedHit for the fast processing of NGS data. Inspired by BLAST, SeedHit counts seed hits between two sequences to determine their similarity. In SeedHit, a nucleic acid in a gene sequence is presented in binary format. By packaging data and generating a lookup table that fits into the L1 cache, SeedHit is GPU-friendly and high-throughput. Using three 16 s rRNA datasets from Greengenes as input SeedHit can reject 84%-89% dissimilar sequence pairs on average when the similarity is 0.9-0.99. The throughput of SeedHit achieved 1 T/s (Tera base per second) on 3080 Ti. Compared with the other two GPU-based filtering algorithms, GateKeeper and SneakySnake, SeedHit has the highest rejection rate and throughput. By incorporating SeedHit into our in-house clustering algorithm nGIA, the modified nGIA achieved a 1.6-2.1 times speedup compared to the original version. Zhen Ju, Xuelei Li, Jintao Meng 0001, Yanjie Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | PERKS: a Locality-Optimized Execution Model for Iterative Memory-bound GPU ApplicationsabstractIterative memory-bound solvers commonly occur in HPC codes. Typical GPU implementations have a loop on the host side that invokes the GPU kernel as much as time/algorithm steps there are. The termination of each kernel implicitly acts the barrier required after advancing the solution every time step. We propose an execution model for running memory-bound iterative GPU kernels: PERsistent KernelS (PERKS). In this model, the time loop is moved inside persistent kernel, and device-wide barriers are used for synchronization. We then reduce the traffic to device memory by caching subset of the output in each time step in the unused registers and shared memory. PERKS can be generalized to any iterative solver: they largely independent of the solver's implementation. We explain the design principle of PERKS and demonstrate effectiveness of PERKS for a wide range of iterative 2D/3D stencil benchmarks (geomean speedup of 2.12x for 2D stencils and 1.24x for 3D stencils over state-of-art libraries), and a Krylov subspace conjugate gradient solver (geomean speedup of 4.86x in smaller SpMV datasets from SuiteSparse and 1.43x in larger SpMV datasets over a state-of-art library). All PERKS-based implementations available at: https://github.com/neozhang307/PERKS. Lingqi Zhang 0001, Mohamed Wahib, Peng Chen 0035, Jintao Meng 0001, Xiao Wang 0004, Toshio Endo, Satoshi Matsuoka |
ICS | 4 |
| 2023 | Revisiting Temporal Blocking Stencil OptimizationsabstractIterative stencils are used widely across the spectrum of High Performance Computing (HPC) applications. Many efforts have been put into optimizing stencil GPU kernels, given the prevalence of GPU-accelerated supercomputers. To improve the data locality, temporal blocking is an optimization that combines a batch of time steps to process them together. Under the observation that GPUs are evolving to resemble CPUs in some aspects, we revisit temporal blocking optimizations for GPUs. We explore how temporal blocking schemes can be adapted to the new features in the recent Nvidia GPUs, including large scratchpad memory, hardware prefetching, and device-wide synchronization. We propose a novel temporal blocking method, EBISU, which champions low device occupancy to drive aggressive deep temporal blocking on large tiles that are executed tile-by-tile. We compare EBISU with state-of-the-art temporal blocking libraries: STENCILGEN and AN5D. We also compare with state-of-the-art stencil auto-tuning tools that are equipped with temporal blocking optimizations: ARTEMIS and DRSTENCIL. Over a wide range of stencil benchmarks, EBISU achieves speedups up to 2.53x and a geometric mean speedup of 1.49x over the best state-of-the-art performance in each stencil benchmark. Lingqi Zhang 0001, Mohamed Wahib, Peng Chen 0035, Jintao Meng 0001, Xiao Wang 0004, Toshio Endo, Satoshi Matsuoka |
ICS | 4 |
| 2023 | Simeuro: A Hybrid CPU-GPU Parallel Simulator for Neuromorphic Computing ChipsabstractWith the success of deep learning, there have been numerous efforts to build hardware for it. One approach that is gaining momentum is neuromorphic computing with spiking neural networks (SNNs), which are multiplication-free and open the possibility of using analog computing via novel technologies. However, to design effective and efficient hardware for such architectures, a fast and accurate software simulator is key. This article presents Simeuro, a fast and scalable system-level simulator for SNN models used in neuromorphic accelerators. The simulator uses spike-level details and configurable architectural constraints that are independent of the underlying hardware implementation. Simeuro supports a wide range of features including analog computing, novel memory (currently, RRAM is supported), and a full network-on-chip. The simulator can provide detailed simulation results such as routing statistics, energy consumption, delay, and accuracy of arbitrarily defined SNN architectures. Our simulator leverages a CPU-GPU hybrid environment to expedite the simulation by scaling out to multi-nodes equipped with multi-GPUs. We are able to conduct core simulations for a system-scale SNN chip of 20,000 neuromorphic cores on up to 512 A100 GPUs in a few minutes. Huaipeng Zhang, Nhut-Minh Ho, Dogukan Yigit Polat, Peng Chen 0035, Mohamed Wahib, Truong Thao Nguyen, Jintao Meng 0001, Rick Siow Mong Goh, Satoshi Matsuoka, Tao Luo 0014, Weng-Fai Wong |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2022 | Efficient Phase-Functioned Real-time Character Control in Mobile Games: A TVM Enabled ApproachabstractIn this paper, we propose a highly efficient computing method for game character control with phase-functioned neural networks (PFNN). The primary challenge to accelerate PFNN on mobile platforms is that PFNN dynamically produces weight matrices with an argument, phase, which is individual to each game character. Therefore existing libraries that generally assume frozen weight matrices are inefficient to accelerate PFNN. The situation becomes even worse when multiple characters are present. To address the challenges, we reformulate the equations and leverage the deep learning compiler stack TVM to build a cross-platform, high-performance implementation. Evaluations reveal that our solutions deliver close-to-peak performance on various platforms, from high-performance servers to energy-efficient mobile platforms. This work is publicly available at https://github.com/turbo0628/pfnn_tvm. Haidong Lan, Wenxi Zhu, Du Wu, Xinghui Fu, Liu Wei, Jintao Meng 0001, Minwen Deng |
ICPP | 9 |
| 2022 | Simulating Spiking Neural Networks Based on SW26010pro
Xuelei Li, Jintao Meng 0001, Yi Pan 0001, Yanjie Wei |
ISBRA | 3 |
| 2022 | Improving protein-ligand docking and screening accuracies by incorporating a scoring function correction termabstractScoring functions are important components in molecular docking for structure-based drug discovery. Traditional scoring functions, generally empirical- or force field-based, are robust and have proven to be useful for identifying hits and lead optimizations. Although multiple highly accurate deep learning- or machine learning-based scoring functions have been developed, their direct applications for docking and screening are limited. We describe a novel strategy to develop a reliable protein-ligand scoring function by augmenting the traditional scoring function Vina score using a correction term (OnionNet-SFCT). The correction term is developed based on an AdaBoost random forest model, utilizing multiple layers of contacts formed between protein residues and ligand atoms. In addition to the Vina score, the model considerably enhances the AutoDock Vina prediction abilities for docking and screening tasks based on different benchmarks (such as cross-docking dataset, CASF-2016, DUD-E and DUD-AD). Furthermore, our model could be combined with multiple docking applications to increase pose selection accuracies and screening abilities, indicating its wide usage for structure-based drug discoveries. Furthermore, in a reverse practice, the combined scoring strategy successfully identified multiple known receptors of a plant hormone. To summarize, the results show that the combination of data-driven model (OnionNet-SFCT) and empirical scoring function (Vina score) is a good scoring strategy that could be useful for structure-based drug discoveries and potentially target fishing in future. Liangzhen Zheng, Jintao Meng 0001, Haidong Lan, Zechen Wang, Mingzhi Lin, Yanjie Wei, Yuguang Mu |
Briefings Bioinform. | 2 |
| 2022 | nGIA: A novel Greedy Incremental Alignment based algorithm for gene sequence clustering
Zhen Ju, Jintao Meng 0001, Jianping Fan 0002, Yi Pan 0001, Xuelei Li, Yanjie Wei |
Future Gener. Comput. Syst. | 3 |
| 2022 | Automatic Generation of High-Performance Convolution Kernels on ARM CPUs for Deep LearningabstractWe presentFastConv, a template-based code auto-generation open-source library that can automatically generate high-performance deep learning convolution kernels of arbitrary matrices/tensors shapes. FastConv is based on the Winograd algorithm, which is reportedly the highest performing algorithm for the time-consuming layers of convolutional neural networks. ARM CPUs cover a wide range of designs and specifications, from embedded devices to HPC-grade CPUs. The leads to the dilemma of how to consistently optimize Winograd-based convolution solvers for convolution layers of different shapes. FastConv addresses this problem by using templates to auto-generate multiple shapes of tuned kernels variants suitable for skinny tall matrices. As a performance portable library, FastConv transparently searches for the best combination of kernel shapes, cache tiles, scheduling of loop orders, packing strategies, access patterns, and online/offline computations. Auto-tuning is used to search the parameter configuration space for the best performance for a given target architecture and problem size. Results show 1.02x to 1.40x, 1.14x to 2.17x, and 1.22x and 2.48x speedup is achieved over NNPACK, ARM NN, and FeatherCNN on Kunpeng 920. Furthermore, performance portability experiments with various convolution shapes show that FastConv achieves 1.2x to 1.7x speedup and 2x to 22x speedup over NNPACK and ARM NN inference engine using Winograd on Kunpeng 920. CPU performance portability evaluation on VGG–16 show an average speedup over NNPACK of 1.42x, 1.21x, 1.26x, 1.37x, 2.26x, and 11.02x on Kunpeng 920, Snapdragon 835, 855, 888, Apple M1, and AWS Graviton2, respectively. Jintao Meng 0001, Chen Zhuang, Peng Chen 0035, Mohamed Wahib, Bertil Schmidt, Xiao Wang 0004, Haidong Lan, Dou Wu, Minwen Deng, Yanjie Wei, Shengzhong Feng |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | FeatherCNN: Fast Inference Computation with TensorGEMM on ARM ArchitecturesabstractDeep Learning is ubiquitous in a wide field of applications ranging from research to industry. In comparison to timeconsuming iterative training of convolutional neural networks (CNNs), inference is a relatively lightweight operation making it amenable to execution on mobile devices. Nevertheless, lower latency and higher computation efficiency are crucial to allow for complex models and prolonged battery life. Addressing the aforementioned challenges, we propose FeatherCNN- a fast inference library for ARM CPUs - targeting the performance ceiling of mobile devices. FeatherCNN employs three key techniques: 1) A highly efficient TensorGEMM (generalized matrix multiplication) routine is applied to accelerate Winograd convolution on ARM CPUs, 2) General layer optimization based on custom high performance kernels improves both the computational efficiency and locality of memory access patterns for non-Winograd layers. 3) The framework design emphasizes joint layer-wise optimization using layer fusion to remove redundant calculations and memory movements. Performance evaluation reveals that FeatherCNN significantly outperforms state-ofthe-art libraries. A forward propagation pass of VGG-16 on a 64-core ARM server is 48, 14, and 12 times faster than Caffe using OpenBLAS, Caffe2 using Eigen, and NNPACK, respectively. In addition, FeatherCNN is 3.19 times faster than the recently released TensorFlow Lite library on an iPhone 7 plus. In terms of GEMM performance, FeatherCNN achieves 14.8 and 39.0 percent higher performance than Apple's Accelerate framework on an iPhone 7 plus and Eigen on a Samsung Galaxy S8, respectively. The source code of FeatherCNN library is publicly available at https://github.com/tencent/feathercnn. Haidong Lan, Jintao Meng 0001, Christian Hundt 0002, Bertil Schmidt, Minwen Deng, Yu Qiao 0001, Shengzhong Feng |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Scalable Assembly for Massive Genomic GraphsabstractScientists increasingly want to assemble large genomes, metagenomes, and large numbers of individual genomes. In order to meet the demand for processing these huge datasets, parallel genome assembly is a vital step. Among all the parallel genome assemblers, de Bruijn graph based ones are most popular. However, the size of de Bruijn graph is determined by the number of distinct kmers used in the algorithm, thus redundant kmers in the genome datasets donot contribute to the graph size. The scalability of genome assemblers is influenced directly by the distinct kmers in the dataset or de Bruijn graph size, rather than the input dataset size. In order to assembly large genomes, we have artificially created 16 datasets of 4 Terabytes in total from the human reference genome. The human reference genome is firstly mutated with a 5% mutation rate, and then subjected to a genome sequencing data simulator ART. The simulated datasets have linearly increasing number of distinct kmers as the size/number of the combined datasets increases. We then evaluate all five time-consuming steps of the SWAP-Assembler 2.0 (SWAP2) using these 16 simulated datasets. Compared with our previous experiment on 1000 human dataset with fixed de Bruijn graph size, the weak-scaling test shows that SWAP2 can scale well from 1024 cores using one dataset to 16,384 cores. The percentage of time usage for all five steps of SWAP2 is fixed, and total time usage is also constant. The result showed that the time usage of graph simplification occupied almost 75% of the total time usage, which will be subject to further optimization for future work. Jintao Meng 0001, Jianqiu Ge, Yanjie Wei, Pavan Balaji, Bingqiang Wang |
CCGrid | 1 |
| 2016 | SWAP-Assembler 2: Optimization of De Novo Genome Assembler at Extreme ScaleabstractIn this paper, we analyze and optimize the most time-consuming steps of the SWAP-Assembler, a parallel genome assembler, so that it can scale to a large number of cores for huge genomes with sequencing data ranging from terabyes to petabytes. Performance analysis results show that the most time-consuming steps are input parallelization, k-mer graph construction, and graph simplification (edge merging). For the input parallelization, the input data is divided into virtual fragments with nearly equal size, and the start position and end position of each fragment are automatically separated at the beginning of the reads. In k-mer graph construction, in order to improve the communication efficiency, the message size is kept constant between any two processes by proportionally increasing the number of nucleotides to the number of processes in the input parallelization step for each round. The memory usage is also decreased because only a small part of the input data is processed in each round. With graph simplification, the communication protocol reduces the number of communication loops from four to two loops and decreases the idle communication time. The optimized assembler is denoted SWAP-Assembler 2 (SWAP2). In our experiments using a 1000 Genomes project dataset of 4 terabytes (the largest dataset ever used for assembling) on the supercomputer Mira, the results show that SWAP2 scales to 131,072 cores with an efficiency of 40%. We also compared our work with both the HipMer assembler and the SWAP-Assembler. On the Yanhuang dataset of 300 gigabytes, SWAP2 shows a 3X speedup and 4X better scalability compared with the HipMer assembler and is 45 times faster than the SWAP-Assembler. The SWAP2 software is available at https://sourceforge.net/projects/swapassembler. Jintao Meng 0001, Pavan Balaji, Yanjie Wei, Bingqiang Wang, Shengzhong Feng |
ICPP | 1 |
| 2015 | SWAP-Assembler 2: Scalable Genome Assembler towards Millions of Cores - Practice and ExperienceabstractThere is widening gap between the throughput of massive parallel sequencing machines and the ability to analyze these huge sequencing data, which can be Tara bytes or even Peta bytes. Previously our assembly tool, SWAP-Assembler, can scale to 2048 cores on TianHe 1A for human Yanhuang genome. This work is to further scale SWAP-Assembler to millions of cores on Mira. SWAP-Assembler can be divided into 5 steps, and the most time consuming steps are input parallelization, kmer graph construction, graph simplification (edge merging). We optimize these three steps to keep the percentage of time usage in each step constant when the number of cores increases. For the input parallelization step, the input data is divided into virtual fragments with almost equal size, the begin position and end position for each fragment is automatically separated at the beginning symbol of reads. This data blocking strategy plays a central role in adjusting the data size to keep the communication and memory efficiency for the subsequent steps. In kmer graph construction, to prevent the communication efficiency degradation, the message size is kept constant (about 8k bytes) between any two processes by proportionally increasing the number of nucleotides to the number of processes in the input parallelization step in each round. The memory usage can be also benefited, as only a small part of the input data is processed in each round. Within graph simplification, the major improvement is to combine messages sending & receiving between its two neighbors into one loop in the communication protocol. After integrated with the above optimizations, the new assembly tool is denoted as SWAP-Assembler 2 or SWAP2 for short. In our experiment for 1k human genome dataset, the modified SWAP-Assembler 2 can scale to 16k cores with parallel efficiency of 70%. Jintao Meng 0001, Yanjie Wei, Pavan Balaji |
CCGRID | 1 |
| 2014 | SWAP-Assembler: scalable and efficient genome assembly towards thousands of coresabstractBACKGROUND: There is a widening gap between the throughput of massive parallel sequencing machines and the ability to analyze these sequencing data. Traditional assembly methods requiring long execution time and large amount of memory on a single workstation limit their use on these massive data. RESULTS: This paper presents a highly scalable assembler named as SWAP-Assembler for processing massive sequencing data using thousands of cores, where SWAP is an acronym for Small World Asynchronous Parallel model. In the paper, a mathematical description of multi-step bi-directed graph (MSG) is provided to resolve the computational interdependence on merging edges, and a highly scalable computational framework for SWAP is developed to automatically preform the parallel computation of all operations. Graph cleaning and contig extension are also included for generating contigs with high quality. Experimental results show that SWAP-Assembler scales up to 2048 cores on Yanhuang dataset using only 26 minutes, which is better than several other parallel assemblers, such as ABySS, Ray, and PASHA. Results also show that SWAP-Assembler can generate high quality contigs with good N50 size and low error rate, especially it generated the longest N50 contig sizes for Fish and Yanhuang datasets. CONCLUSIONS: In this paper, we presented a highly scalable and efficient genome assembly software, SWAP-Assembler. Compared with several other assemblers, it showed very good performance in terms of scalability and contig quality. This software is available at: https://sourceforge.net/projects/swapassembler. Jintao Meng 0001, Bingqiang Wang, Yanjie Wei, Shengzhong Feng, Pavan Balaji |
BMC Bioinform. | 1 |
| 2013 | Improved Parallel Processing of Massive De Bruijn Graph for Genome Assembly
Jiefeng Cheng, Jintao Meng 0001, Bingqiang Wang, Shengzhong Feng |
APWeb | 3 |
| 2013 | An Energy Efficient Clustering Scheme for Data Aggregation in Wireless Sensor Networks
Jintao Meng 0001, Jian-Rui Yuan, Shengzhong Feng, Yanjie Wei |
J. Comput. Sci. Technol. | 1 |
| 2012 | DGraph: Algorithms for Shortgun Reads Assembly Using De Bruijn Graph
Jintao Meng 0001, Jianrui Yuan, Jiefeng Cheng, Yanjie Wei, Shengzhong Feng |
NPC | 1 |
| 2012 | Small World Asynchronous Parallel Model for Genome Assembly
Jintao Meng 0001, Jianrui Yuan, Jiefeng Cheng, Yanjie Wei, Shengzhong Feng |
NPC | 1 |