Pinglu Zhang

dblp:344/2128 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0002-1788-3084ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TAXICF: Efficient Index Construction and Accurate Metagenomic Classification with Interleaved Cuckoo Filters
Qinzhong Tian, Pinglu Zhang, Quan Zou 0001
ICIC (30)2
2026 High-Quality Large-Scale Viral Genome Multiple Sequence Alignment with Updated HAlign4
Pinglu Zhang, Qinzhong Tian, Yixiao Zhai, Quan Zou 0001
ICIC (30)1
2025 ReAlign-Star: An Optimized Realignment Method for Multiple Sequence Alignment, Targeting Star Algorithm Tools
Yixiao Zhai, Pinglu Zhang, Yi Liu 0112, Quan Zou 0001
ICIC (26)2
2025 ReAlign-P: a vertical iterative realignment method for protein multiple sequence alignment
abstract
MOTIVATION: Reliable protein multiple sequence alignment (MSA) is essential for downstream biomedical research and directly impacts the accuracy of analytical results. However, protein sequences often exhibit low similarity and complex alignment patterns, and existing general alignment tools frequently fall short in terms of accuracy. Many current realignment methods are outdated, suffering from issues such as code obsolescence and inadequate precision. As a result, there is a pressing need for realignment methods that can better address these challenges. RESULTS: This study introduces ReAlign-P, a realignment tool designed specifically for protein MSA. ReAlign-P first divides the initial alignment into three regions and applies a novel vertical iterative realignment strategy to optimize the more conserved middle region. This method is by default compatible with MUSCLE5 for realignment, leading to a significant improvement in accuracy. We evaluated initial alignments generated using 10 different MSA parameter configurations across four protein benchmark datasets. The results demonstrate that ReAlign-P consistently outperforms or matches the quality of the initial alignments in all cases. In contrast, RASCAL-the only other currently functional protein realignment tool-sometimes even reduces alignment quality. ReAlign-P not only delivers more substantial improvements but also exhibits greater stability, effectively addressing the gap in available protein realignment tools. AVAILABILITY AND IMPLEMENTATION: The source code and test data for ReAlign-P are available on GitHub (https://github.com/malabz/ReAlign-P).
Yixiao Zhai, Pinglu Zhang, Quan Zou 0001, Ximei Luo
Bioinform.2
2024 HAlign 4: a new strategy for rapidly aligning millions of sequences
abstract
MOTIVATION: HAlign is a high-performance multiple sequence alignment software based on the star alignment strategy, which is the preferred choice for rapidly aligning large numbers of sequences. HAlign3, implemented in Java, is the latest version capable of aligning an ultra-large number of similar DNA/RNA sequences. However, HAlign3 still struggles with long sequences and extremely large numbers of sequences. RESULTS: To address this issue, we have implemented HAlign4 in C++. In this version, we replaced the original suffix tree with Burrows-Wheeler Transform and introduced the wavefront alignment algorithm to further optimize both time and memory efficiency. Experiments show that HAlign4 significantly outperforms HAlign3 in runtime and memory usage in both single-threaded and multi-threaded configurations, while maintains high alignment accuracy comparable to MAFFT. HAlign4 can complete the alignment of 10 million coronavirus disease 2019 (COVID-19) sequences in about 12 min and 300 GB of memory using 96 threads, demonstrating its efficiency and practicality for large-scale alignment on standard workstations. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/malabz/HAlign-4, dataset is available at https://zenodo.org/records/13934503.
Tong Zhou 0016, Pinglu Zhang, Quan Zou 0001, Wu Han
Bioinform.2
2024 FMAlign2: a novel fast multiple nucleotide sequence alignment method for ultralong datasets
abstract
MOTIVATION: In bioinformatics, multiple sequence alignment (MSA) is a crucial task. However, conventional methods often struggle with aligning ultralong sequences. To address this issue, researchers have designed MSA methods rooted in a vertical division strategy, which segments sequence data for parallel alignment. A prime example of this approach is FMAlign, which utilizes the FM-index to extract common seeds and segment the sequences accordingly. RESULTS: FMAlign2 leverages the suffix array to identify maximal exact matches, redefining the approach of FMAlign from searching for global chains to partial chains. By using a vertical division strategy, large-scale problem is deconstructed into manageable tasks, enabling parallel execution of subMSA. Furthermore, sequence-profile alignment and refinement are incorporated to concatenate subsets, yielding the final result seamlessly. Compared to FMAlign, FMAlign2 markedly augments the segmentation of sequences and significantly reduces the time while maintaining accuracy, especially on ultralong datasets. Importantly, FMAlign2 enhances existing MSA methods by conferring the capability to handle sequences reaching billions in length within an acceptable time frame. AVAILABILITY AND IMPLEMENTATION: Source code and datasets are available at https://github.com/malabz/FMAlign2 and https://zenodo.org/records/10435770.
Pinglu Zhang, Huan Liu 0024, Yanming Wei, Yixiao Zhai, Qinzhong Tian, Quan Zou 0001
Bioinform.1
2024 TPMA: A two pointers meta-alignment tool to ensemble different multiple nucleic acid sequence alignments
abstract
Accurate multiple sequence alignment (MSA) is imperative for the comprehensive analysis of biological sequences. However, a notable challenge arises as no single MSA tool consistently outperforms its counterparts across diverse datasets. Users often have to try multiple MSA tools to achieve optimal alignment results, which can be time-consuming and memory-intensive. While the overall accuracy of certain MSA results may be lower, there could be local regions with the highest alignment scores, prompting researchers to seek a tool capable of merging these locally optimal results from multiple initial alignments into a globally optimal alignment. In this study, we introduce Two Pointers Meta-Alignment (TPMA), a novel tool designed for the integration of nucleic acid sequence alignments. TPMA employs two pointers to partition the initial alignments into blocks containing identical sequence fragments. It selects blocks with the high sum of pairs (SP) scores to concatenate them into an alignment with an overall SP score superior to that of the initial alignments. Through tests on simulated and real datasets, the experimental results consistently demonstrate that TPMA outperforms M-Coffee in terms of aSP, Q, and total column (TC) scores across most datasets. Even in cases where TPMA's scores are comparable to M-Coffee, TPMA exhibits significantly lower running time and memory consumption. Furthermore, we comprehensively assessed all the MSA tools used in the experiments, considering accuracy, time, and memory consumption. We propose accurate and fast combination strategies for small and large datasets, which streamline the user tool selection process and facilitate large-scale dataset integration. The dataset and source code of TPMA are available on GitHub (https://github.com/malabz/TPMA).
Yixiao Zhai, Jiannan Chao, Yizheng Wang, Pinglu Zhang, Furong Tang, Quan Zou 0001
PLoS Comput. Biol.4
2023 PS-Mixer: A Polar-Vector and Strength-Vector Mixer Model for Multimodal Sentiment Analysis
Pinglu Zhang, Jiading Ling, Zhenguo Yang, Lap-Kei Lee, Wenyin Liu
Inf. Process. Manag.2