EDBT 2026 Demo / reviewers in the wild / expert
Yixiao Zhai
dblp:355/0223
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-Quality Large-Scale Viral Genome Multiple Sequence Alignment with Updated HAlign4
Pinglu Zhang, Qinzhong Tian, Yixiao Zhai, Quan Zou 0001 |
ICIC (30) | 3 |
| 2025 | ReAlign-Star: An Optimized Realignment Method for Multiple Sequence Alignment, Targeting Star Algorithm Tools
Yixiao Zhai, Pinglu Zhang, Yi Liu 0112, Quan Zou 0001 |
ICIC (26) | 1 |
| 2025 | FORAlign: accelerating gap-affine DNA pairwise sequence alignment using FOR-blocks based on Four Russians approach with linear space complexityabstractPairwise sequence alignment (PSA) serves as the cornerstone in computational bioinformatics, facilitating multiple sequence alignment and phylogenetic analysis. This paper introduces the FORAlign algorithm, leveraging the Four Russians algorithm with identical upper-bound time and space complexity as the Hirschberg divide-and-conquer PSA algorithm, aimed at accelerating Hirschberg PSA algorithm in parallel. Particularly notable is its capability to achieve up to 16.79 times speedup when aligning sequences with low sequence similarity, compared to the conventional Needleman-Wunsch PSA method using non-heuristic methods. Empirical evaluations underscore FORAlign's superiority over existing wavefront alignment (WFA) series software, especially in scenarios characterized by low sequence similarity during PSA tasks. Our method is capable of directly aligning monkeypox sequences with other sequences using non-heuristic methods. The algorithm was implemented within the FORAlign library, providing functionality for PSA and foundational support for multiple sequence alignment and phylogenetic trees. The FORAlign library is freely available at https://github.com/malabz/FORAlign. Yanming Wei, Tong Zhou 0016, Yixiao Zhai, Liang Yu 0002, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2025 | ReAlign-P: a vertical iterative realignment method for protein multiple sequence alignmentabstractMOTIVATION: Reliable protein multiple sequence alignment (MSA) is essential for downstream biomedical research and directly impacts the accuracy of analytical results. However, protein sequences often exhibit low similarity and complex alignment patterns, and existing general alignment tools frequently fall short in terms of accuracy. Many current realignment methods are outdated, suffering from issues such as code obsolescence and inadequate precision. As a result, there is a pressing need for realignment methods that can better address these challenges. RESULTS: This study introduces ReAlign-P, a realignment tool designed specifically for protein MSA. ReAlign-P first divides the initial alignment into three regions and applies a novel vertical iterative realignment strategy to optimize the more conserved middle region. This method is by default compatible with MUSCLE5 for realignment, leading to a significant improvement in accuracy. We evaluated initial alignments generated using 10 different MSA parameter configurations across four protein benchmark datasets. The results demonstrate that ReAlign-P consistently outperforms or matches the quality of the initial alignments in all cases. In contrast, RASCAL-the only other currently functional protein realignment tool-sometimes even reduces alignment quality. ReAlign-P not only delivers more substantial improvements but also exhibits greater stability, effectively addressing the gap in available protein realignment tools. AVAILABILITY AND IMPLEMENTATION: The source code and test data for ReAlign-P are available on GitHub (https://github.com/malabz/ReAlign-P). Yixiao Zhai, Pinglu Zhang, Quan Zou 0001, Ximei Luo |
Bioinform. | 1 |
| 2024 | FMAlign2: a novel fast multiple nucleotide sequence alignment method for ultralong datasetsabstractMOTIVATION: In bioinformatics, multiple sequence alignment (MSA) is a crucial task. However, conventional methods often struggle with aligning ultralong sequences. To address this issue, researchers have designed MSA methods rooted in a vertical division strategy, which segments sequence data for parallel alignment. A prime example of this approach is FMAlign, which utilizes the FM-index to extract common seeds and segment the sequences accordingly. RESULTS: FMAlign2 leverages the suffix array to identify maximal exact matches, redefining the approach of FMAlign from searching for global chains to partial chains. By using a vertical division strategy, large-scale problem is deconstructed into manageable tasks, enabling parallel execution of subMSA. Furthermore, sequence-profile alignment and refinement are incorporated to concatenate subsets, yielding the final result seamlessly. Compared to FMAlign, FMAlign2 markedly augments the segmentation of sequences and significantly reduces the time while maintaining accuracy, especially on ultralong datasets. Importantly, FMAlign2 enhances existing MSA methods by conferring the capability to handle sequences reaching billions in length within an acceptable time frame. AVAILABILITY AND IMPLEMENTATION: Source code and datasets are available at https://github.com/malabz/FMAlign2 and https://zenodo.org/records/10435770. Pinglu Zhang, Huan Liu 0024, Yanming Wei, Yixiao Zhai, Qinzhong Tian, Quan Zou 0001 |
Bioinform. | 4 |
| 2024 | SBSM-Pro: support bio-sequence machine for proteins
Yizheng Wang, Yixiao Zhai, Yijie Ding, Quan Zou 0001 |
Sci. China Inf. Sci. | 2 |
| 2024 | TPMA: A two pointers meta-alignment tool to ensemble different multiple nucleic acid sequence alignmentsabstractAccurate multiple sequence alignment (MSA) is imperative for the comprehensive analysis of biological sequences. However, a notable challenge arises as no single MSA tool consistently outperforms its counterparts across diverse datasets. Users often have to try multiple MSA tools to achieve optimal alignment results, which can be time-consuming and memory-intensive. While the overall accuracy of certain MSA results may be lower, there could be local regions with the highest alignment scores, prompting researchers to seek a tool capable of merging these locally optimal results from multiple initial alignments into a globally optimal alignment. In this study, we introduce Two Pointers Meta-Alignment (TPMA), a novel tool designed for the integration of nucleic acid sequence alignments. TPMA employs two pointers to partition the initial alignments into blocks containing identical sequence fragments. It selects blocks with the high sum of pairs (SP) scores to concatenate them into an alignment with an overall SP score superior to that of the initial alignments. Through tests on simulated and real datasets, the experimental results consistently demonstrate that TPMA outperforms M-Coffee in terms of aSP, Q, and total column (TC) scores across most datasets. Even in cases where TPMA's scores are comparable to M-Coffee, TPMA exhibits significantly lower running time and memory consumption. Furthermore, we comprehensively assessed all the MSA tools used in the experiments, considering accuracy, time, and memory consumption. We propose accurate and fast combination strategies for small and large datasets, which streamline the user tool selection process and facilitate large-scale dataset integration. The dataset and source code of TPMA are available on GitHub (https://github.com/malabz/TPMA). Yixiao Zhai, Jiannan Chao, Yizheng Wang, Pinglu Zhang, Furong Tang, Quan Zou 0001 |
PLoS Comput. Biol. | 1 |
| 2024 | Efficient and Privacy-Preserving Outsourcing of Gradient Boosting Decision Tree InferenceabstractRecently, outsourcing machine learning inference services to the cloud has become increasingly popular. The inference process, however, remains an open question onhow to effectively protect the model owner's proprietary model, the user's sensitive data, and prediction results. In this work, we propose an efficient and comprehensive privacy-preserving framework for outsourcing Gradient Boosting Decision Tree (GBDT) inference utilizing pseudorandom function and additively homomorphic encryption. Specifically, we first design a transformation method for GBDT to protect the node and structure privacy of the owner's model. On top of the protected model, we further propose customized comparison and random trees permutation protocols, which substantially boost the computation and reduce the communication cost of the outsourcing inference, while preventing the user from inferring privacy associated with GBDT. Besides, we provide rigorous security analysis, and extensive experiments on 7 real-world datasets and various models demonstrating that our scheme achieves up to 36 times less runtime and 69 times less communication compared to the state-of-the-arts. Shuai Yuan 0009, Hongwei Li 0001, Xinyuan Qian 0002, Meng Hao 0001, Yixiao Zhai, Guowen Xu |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | Toward Efficient and End-to-End Privacy-Preserving Distributed Gradient Boosting Decision TreesabstractGradient Boosting Decision Trees (GBDTs) are popular machine learning models due to its simplicity, effectiveness, and interpretability. Recently, to alleviate serious privacy leakages in conventional centralized methods, researchers have proposed several privacy-preserving distributed GBDT solutions. However, those approaches still suffer from either insufficient privacy protection or significant runtime and communication overhead. In this paper, we propose an efficient and end-to-end privacy-preserving distributed GBDT framework, called PPD-GBDT, which uses differential privacy, polynomial approximation, and fully homomorphic encryption to achieve comprehensive privacy protection. Specifically, during the boosting phase, we design a novel model preparation method to improve the efficiency of prediction with acceptably slight accuracy/RMSE loss while preventing data owners' corruption. On the other hand, for the prediction phase, we propose a customized secure prediction method, which effectively prevents the malicious server from stealing private information. Besides, we conduct extensive experiments on six datasets and compare with three prior schemes. Evaluation results show that our privacy-preserving scheme achieves lower runtime and up to 40× less communication overhead compared to the state-of-the-arts. Shuai Yuan 0009, Hongwei Li 0001, Xinyuan Qian 0002, Meng Hao 0001, Yixiao Zhai |
ICC | 5 |