Beifang Niu

dblp:90/7969 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-7448-7793ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 3 first-author · 9 since 2021Systems, architecture and hardware · 3 · 2 since 2021
YearPublicationVenuePosition
2026 RabbitVar: Ultra-fast and accurate somatic small-variant calling on multi-core architectures
Hao Zhang 0142, Lin Gan 0001, Zekun Yin, Lifeng Yan, Honglei Song, Qixin Chang, Yanjie Wei, Beifang Niu, Bertil Schmidt
Future Gener. Comput. Syst.9
2025 Review of deep learning-based pathological image classification: From task-specific models to foundation models
Haijing Luan, Kaixing Yang, Taiyuan Hu, Jifang Hu, Jiayin He, Rui Yan 0009, Xiaobing Guo, Niansong Qian, Beifang Niu
Future Gener. Comput. Syst.11
2024 Transformer-Based Multi-Scale Fusion for Robust Predicting Microsatellite Instability from Pathological Images
abstract
Microsatellite instability (MSI) is a crucial biomarker for guiding the efficacy of immunotherapy and adjuvant chemotherapy, making its detection essential for effective cancer treatment and prognosis. Traditional MSI prediction methods encounter challenges including high costs and limited accuracy under low tumor purity conditions. Recent advancements have explored deep learning for MSI prediction from pathological images, yet these approaches often overlook the multi-scale nature of pathological images and specific pathological features critical for MSI diagnosis. In this study, we proposed MSIscope, a novel Transformer-based method for detecting MSI from pathological images by fusing multi-scale pathological image information. Our approach consists of three key components: 1) ROI selection: we design a region of interest (ROI) selector based on convolutional neural networks and attention mechanisms, selecting tumor regions and important non-tumor regions as our focus; 2) Multi-scale vision expansion and feature extraction: we develop an algorithm that captures a broader view centered on a specified area to obtain a multi-scale field of view. The CTransPath feature extractor is then used to extract features from the image; 3) Multi-scale fusion Transformer: we propose a multi-scale feature aggregator (MS-Transformer) to aggregate contextual features across regions and scales. Our method was experimentally validated on public datasets, achieving an AU-ROC of 0.911 on the TCGA pan-cancer dataset and 0.887 on the TCGA-CRC dataset, surpassing existing methods. Additionally, it maintains high AUROC on datasets with lower tumor purity, outperforming current approaches. These results highlight the potential of MSIscope as an robust method for MSI prediction.
Taiyuan Hu, Haijing Luan, Rui Yan 0009, Jifang Hu, Kaixing Yang, Xinyin Han, Weier Liu, Jiayin He, Xiaohong Duan, Fa Zhang 0001, Beifang Niu
BIBM12
2024 A comprehensive performance evaluation, comparison, and integration of computational methods for detecting and estimating cross-contamination of human samples in cancer next-generation sequencing analysis
Huijuan Chen, Lili Cai, Yali Hu, Xue Leng, Dongjie Fan, Beifang Niu, Qiming Zhou
J. Biomed. Informatics10
2023 Multi-class Cancer Classification of Whole Slide Images Through Transformer and Multiple Instance Learning
Haijing Luan, Taiyuan Hu, Jifang Hu, Detao Ji, Jiayin He, Xiaohong Duan, Chunyan Yang, Yajun Gao, Beifang Niu
ISBRA11
2022 RabbitQCPlus: More Efficient Quality Control for Sequencing Data
abstract
Assessing the quality of sequencing data plays a crucial role in downstream data analysis. However, existing tools often achieve sub-optimal efficiency, especially when dealing with compressed files or performing complicated quality control operations such as over-representation analysis. We present RabbitQCPlus, an ultra-efficient quality control tool for modern multi-core systems. RabbitQCPlus uses vectorization, memory copy reduction, parallel (de)compression, and optimized data structures to achieve substantial performance gains. It is 1.1 to 5.4 times faster when performing basic quality control operations compared to state-of-the-art applications yet requires fewer compute resources. Moreover, RabbitQCPlus is at least 4 times faster than other applications when processing gzip-compressed FASTQ files. Furthermore, it takes less than 4 minutes to process 280GB of plain FASTQ sequencing data, while other applications take at least 22 minutes on a 48-core server when enabling the per-read over-representation analysis. C++ sources are available at https://github.com/RabbitBio/RabbitQCPlus.
Lifeng Yan, Zekun Yin, Hao Zhang 0142, Zhan Zhao, André Müller, Robin Kobus, Yanjie Wei, Beifang Niu, Bertil Schmidt
BIBM9
2022 OncoPubMiner: a platform for mining oncology publications
abstract
Updated and expert-quality knowledge bases are fundamental to biomedical research. A knowledge base established with human participation and subject to multiple inspections is needed to support clinical decision making, especially in the growing field of precision oncology. The number of original publications in this field has risen dramatically with the advances in technology and the evolution of in-depth research. Consequently, the issue of how to gather and mine these articles accurately and efficiently now requires close consideration. In this study, we present OncoPubMiner (https://oncopubminer.chosenmedinfo.com), a free and powerful system that combines text mining, data structure customisation, publication search with online reading and project-centred and team-based data collection to form a one-stop 'keyword in-knowledge out' oncology publication mining platform. The platform was constructed by integrating all open-access abstracts from PubMed and full-text articles from PubMed Central, and it is updated daily. OncoPubMiner makes obtaining precision oncology knowledge from scientific articles straightforward and will assist researchers in efficiently developing structured knowledge base systems and bring us closer to achieving precision oncology goals.
Jifang Hu, Xiaohong Duan, Niuben Song, Jincheng Zhai, Junyan Su, Zhongjia Guo, Hexiang Li, Qiming Zhou, Beifang Niu
Briefings Bioinform.15
2021 MSIsensor-ct: microsatellite instability detection using cfDNA sequencing data
abstract
MOTIVATION: Microsatellite instability (MSI) is a promising biomarker for cancer prognosis and chemosensitivity. Techniques are rapidly evolving for the detection of MSI from tumor-normal paired or tumor-only sequencing data. However, tumor tissues are often insufficient, unavailable, or otherwise difficult to procure. Increasing clinical evidence indicates the enormous potential of plasma circulating cell-free DNA (cfNDA) technology as a noninvasive MSI detection approach. RESULTS: We developed MSIsensor-ct, a bioinformatics tool based on a machine learning protocol, dedicated to detecting MSI status using cfDNA sequencing data with a potential stable MSIscore threshold of 20%. Evaluation of MSIsensor-ct on independent testing datasets with various levels of circulating tumor DNA (ctDNA) and sequencing depth showed 100% accuracy within the limit of detection (LOD) of 0.05% ctDNA content. MSIsensor-ct requires only BAM files as input, rendering it user-friendly and readily integrated into next generation sequencing (NGS) analysis pipelines. AVAILABILITY: MSIsensor-ct is freely available at https://github.com/niu-lab/MSIsensor-ct. SUPPLEMENTARY INFORMATION: Supplementary data are available at Briefings in Bioinformatics online.
Xinyin Han, Shuying Zhang, Daniel Cui Zhou, Danyang Yuan, Jiayin He, Xiaohong Duan, Michael C. Wendl, Beifang Niu
Briefings Bioinform.12
2021 Comprehensive fundamental somatic variant calling and quality management strategies for human cancer genomes
abstract
Next-generation sequencing (NGS) technology has revolutionised human cancer research, particularly via detection of genomic variants with its ultra-high-throughput sequencing and increasing affordability. However, the inundation of rich cancer genomics data has resulted in significant challenges in its exploration and translation into biological insights. One of the difficulties in cancer genome sequencing is software selection. Currently, multiple tools are widely used to process NGS data in four stages: raw sequence data pre-processing and quality control (QC), sequence alignment, variant calling and annotation and visualisation. However, the differences between these NGS tools, including their installation, merits, drawbacks and application, have not been fully appreciated. Therefore, a systematic review of the functionality and performance of NGS tools is required to provide cancer researchers with guidance on software and strategy selection. Another challenge is the multidimensional QC of sequencing data because QC can not only report varied sequence data characteristics but also reveal deviations in diverse features and is essential for a meaningful and successful study. However, monitoring of QC metrics in specific steps including alignment and variant calling is neglected in certain pipelines such as the 'Best Practices Workflows' in GATK. In this review, we investigated the most widely used software for the fundamental analysis and QC of cancer genome sequencing data and provided instructions for selecting the most appropriate software and pipelines to ensure precise and efficient conclusions. We further discussed the prospects and new research directions for cancer genomics.
Shanyu Chen, Xinyin Han, Zhipeng He 0003, Danyang Yuan, Shuying Zhang, Xiaohong Duan, Beifang Niu
Briefings Bioinform.9
2021 Comprehensive review and evaluation of computational methods for identifying FLT3-internal tandem duplication in acute myeloid leukaemia
abstract
Internal tandem duplication (ITD) of FMS-like tyrosine kinase 3 (FLT3-ITD) constitutes an independent indicator of poor prognosis in acute myeloid leukaemia (AML). AML with FLT3-ITD usually presents with poor treatment outcomes, high recurrence rate and short overall survival. Currently, polymerase chain reaction and capillary electrophoresis are widely adopted for the clinical detection of FLT3-ITD, whereas the length and mutation frequency of ITD are evaluated using fragment analysis. With the development of sequencing technology and the high incidence of FLT3-ITD mutations, a multitude of bioinformatics tools and pipelines have been developed to detect FLT3-ITD using next-generation sequencing data. However, systematic comparison and evaluation of the methods or software have not been performed. In this study, we provided a comprehensive review of the principles, functionality and limitations of the existing methods for detecting FLT3-ITD. We further compared the qualitative and quantitative detection capabilities of six representative tools using simulated and biological data. Our results will provide practical guidance for researchers and clinicians to select the appropriate FLT3-ITD detection tools and highlight the direction of future developments in this field. Availability: A Docker image with several programs pre-installed is available at https://github.com/niu-lab/docker-flt3-itd to facilitate the application of FLT3-ITD detection tools.
Danyang Yuan, Xinyin Han, Chunyan Yang, Shuying Zhang, Haijing Luan, Jiayin He, Xiaohong Duan, Qiming Zhou, Sujun Gao, Beifang Niu
Briefings Bioinform.14
2021 RabbitQC: high-speed scalable quality control for sequencing data
abstract
MOTIVATION: Modern sequencing technologies continue to revolutionize many areas of biology and medicine. Since the generated datasets are error-prone, downstream applications usually require quality control methods to pre-process FASTQ files. However, existing tools for this task are currently not able to fully exploit the capabilities of computing platforms leading to slow runtimes. RESULTS: We present RabbitQC, an extremely fast integrated quality control tool for FASTQ files, which can take full advantage of modern hardware. It includes a variety of operations and supports different sequencing technologies (Illumina, Oxford Nanopore and PacBio). RabbitQC achieves speedups between one and two orders-of-magnitude compared to other state-of-the-art tools. AVAILABILITY AND IMPLEMENTATION: C++ sources and binaries are available at https://github.com/ZekunYin/RabbitQC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zekun Yin, Hao Zhang 0142, Meiyang Liu, Honglei Song, Haidong Lan, Yanjie Wei, Beifang Niu, Bertil Schmidt
Bioinform.8
2020 Digital Currency Investment Strategy Framework Based on Ranking
Chuangchuang Dai, Xueying Yang, Meikang Qiu, Xiaobing Guo, Zhonghua Lu, Beifang Niu
ICA3PP (3)6
2020 HotSpot3D web server: an integrated resource for mutation analysis in protein 3D structures
abstract
MOTIVATION: HotSpot3D is a widely used software for identifying mutation hotspots on the 3D structures of proteins. To further assist users, we developed a new HotSpot3D web server to make this software more versatile, convenient and interactive. RESULTS: The HotSpot3D web server performs data pre-processing, clustering, visualization and log-viewing on one stop. Users can interactively explore each cluster and easily re-visualize the mutational clusters within browsers. We also provide a database that allows users to search and visualize proximal mutations from 33 cancers in the Cancer Genome Atlas. AVAILABILITY AND IMPLEMENTATION: http://niulab.scgrid.cn/HotSpot3D/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shanyu Chen, Xiaohong Duan, Beifang Niu
Bioinform.5
2019 DGCF: A Distributed Greedy Clustering Framework for Large-scale Genomic Sequences
abstract
Clustering is a very fundamental while time-consuming compute operation in biological sequence analysis. New sequencing technologies such as NGS and 3GS have dramatically increased both the dataset size and the length of a single read sequence. However, existing tools lack scalability for handling large-scale datasets as well as long sequences. A feasible solution to this problem is to use parallel and distributed systems. The efficient deployment of such systems, however, requires high parallelism in both software implementations as well as algorithmic optimizations. In this paper, we propose DGCF, a Distributed Greedy Clustering Framework which is capable to handle large-scale datasets and long sequences. Our framework adopts a greedy clustering strategy which overlaps communication with computation among many distributed computing nodes. We also design and implement a sparse suffix array (SSA)-based alignment algorithm that can support long sequences. Experiments show that our framework achieves near-linear speedups on a distributed memory cluster.
Zekun Yin, Xiaoming Xu 0004, Kaichao Fan, Weizhong Li 0002, Beifang Niu
BIBM7
2014 MSIsensor: microsatellite instability detection using paired tumor-normal sequence data
abstract
MOTIVATION: Microsatellite instability (MSI) is an important indicator of larger genome instability and has been linked to many genetic diseases, including Lynch syndrome. MSI status is also an independent prognostic factor for favorable survival in multiple cancer types, such as colorectal and endometrial. It also informs the choice of chemotherapeutic agents. However, the current PCR-electrophoresis-based detection procedure is laborious and time-consuming, often requiring visual inspection to categorize samples. We developed MSIsensor, a C++ program for automatically detecting somatic microsatellite changes. It computes length distributions of microsatellites per site in paired tumor and normal sequence data, subsequently using these to statistically compare observed distributions in both samples. Comprehensive testing indicates MSIsensor is an efficient and effective tool for deriving MSI status from standard tumor-normal paired sequence data. AVAILABILITY AND IMPLEMENTATION: https://github.com/ding-lab/msisensor
Beifang Niu, Kai Ye 0001, Qunyuan Zhang, Charles Lu 0002, Mingchao Xie, Michael D. McLellan, Michael C. Wendl
Bioinform.1
2013 MGAviewer: a desktop visualization tool for analysis of metagenomics alignment data
abstract
SUMMARY: Numerous metagenomics projects have produced tremendous amounts of sequencing data. Aligning these sequences to reference genomes is an essential analysis in metagenomics studies. Large-scale alignment data call for intuitive and efficient visualization tool. However, current tools such as various genome browsers are highly specialized to handle intraspecies mapping results. They are not suitable for alignment data in metagenomics, which are often interspecies alignments. We have developed a web browser-based desktop application for interactively visualizing alignment data of metagenomic sequences. This viewer is easy to use on all computer systems with modern web browsers and requires no software installation. AVAILABILITY: http://weizhongli-lab.org/mgaviewer
Zhengwei Zhu 0001, Beifang Niu, Sitao Wu, Shulei Sun, Weizhong Li 0002
Bioinform.2
2012 Ultrafast clustering algorithms for metagenomic sequence analysis
abstract
The rapid advances of high-throughput sequencing technologies dramatically prompted metagenomic studies of microbial communities that exist at various environments. Fundamental questions in metagenomics include the identities, composition and dynamics of microbial populations and their functions and interactions. However, the massive quantity and the comprehensive complexity of these sequence data pose tremendous challenges in data analysis. These challenges include but are not limited to ever-increasing computational demand, biased sequence sampling, sequence errors, sequence artifacts and novel sequences. Sequence clustering methods can directly answer many of the fundamental questions by grouping similar sequences into families. In addition, clustering analysis also addresses the challenges in metagenomics. Thus, a large redundant data set can be represented with a small non-redundant set, where each cluster can be represented by a single entry or a consensus. Artifacts can be rapidly detected through clustering. Errors can be identified, filtered or corrected by using consensus from sequences within clusters.
Weizhong Li 0002, Limin Fu, Beifang Niu, Sitao Wu, John C. Wooley
Briefings Bioinform.3
2012 CD-HIT: accelerated for clustering the next-generation sequencing data
abstract
SUMMARY: CD-HIT is a widely used program for clustering biological sequences to reduce sequence redundancy and improve the performance of other sequence analyses. In response to the rapid increase in the amount of sequencing data produced by the next-generation sequencing technologies, we have developed a new CD-HIT program accelerated with a novel parallelization strategy and some other techniques to allow efficient clustering of such datasets. Our tests demonstrated very good speedup derived from the parallelization for up to ∼24 cores and a quasi-linear speedup for up to ∼8 cores. The enhanced CD-HIT is capable of handling very large datasets in much shorter time than previous versions. AVAILABILITY: http://cd-hit.org. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Limin Fu, Beifang Niu, Zhengwei Zhu 0001, Sitao Wu, Weizhong Li 0002
Bioinform.2
2011 FR-HIT, a very fast program to recruit metagenomic reads to homologous reference genomes
abstract
SUMMARY: Fragment recruitment, a process of aligning sequencing reads to reference genomes, is a crucial step in metagenomic data analysis. The available sequence alignment programs are either slow or insufficient for recruiting metagenomic reads. We implemented an efficient algorithm, FR-HIT, for fragment recruitment. We applied FR-HIT and several other tools including BLASTN, MegaBLAST, BLAT, LAST, SSAHA2, SOAP2, BWA and BWA-SW to recruit four metagenomic datasets from different type of sequencers. On average, FR-HIT and BLASTN recruited significantly more reads than other programs, while FR-HIT is about two orders of magnitude faster than BLASTN. FR-HIT is slower than the fastest SOAP2, BWA and BWA-SW, but it recruited 1-5 times more reads. AVAILABILITY: http://weizhongli-lab.org/frhit.
Beifang Niu, Zhengwei Zhu 0001, Limin Fu, Sitao Wu, Weizhong Li 0002
Bioinform.1
2010 CD-HIT Suite: a web server for clustering and comparing biological sequences
abstract
UNLABELLED: CD-HIT is a widely used program for clustering and comparing large biological sequence datasets. In order to further assist the CD-HIT users, we significantly improved this program with more functions and better accuracy, scalability and flexibility. Most importantly, we developed a new web server, CD-HIT Suite, for clustering a user-uploaded sequence dataset or comparing it to another dataset at different identity levels. Users can now interactively explore the clusters within web browsers. We also provide downloadable clusters for several public databases (NCBI NR, Swissprot and PDB) at different identity levels. AVAILABILITY: Free access at http://cd-hit.org
Beifang Niu, Limin Fu, Weizhong Li 0002
Bioinform.2
2010 Artificial and natural duplicates in pyrosequencing reads of metagenomic data
abstract
BACKGROUND: Artificial duplicates from pyrosequencing reads may lead to incorrect interpretation of the abundance of species and genes in metagenomic studies. Duplicated reads were filtered out in many metagenomic projects. However, since the duplicated reads observed in a pyrosequencing run also include natural (non-artificial) duplicates, simply removing all duplicates may also cause underestimation of abundance associated with natural duplicates. RESULTS: We implemented a method for identification of exact and nearly identical duplicates from pyrosequencing reads. This method performs an all-against-all sequence comparison and clusters the duplicates into groups using an algorithm modified from our previous sequence clustering method cd-hit. This method can process a typical dataset in approximately 10 minutes; it also provides a consensus sequence for each group of duplicates. We applied this method to the underlying raw reads of 39 genomic projects and 10 metagenomic projects that utilized pyrosequencing technique. We compared the occurrences of the duplicates identified by our method and the natural duplicates made by independent simulations. We observed that the duplicates, including both artificial and natural duplicates, make up 4-44% of reads. The number of natural duplicates highly correlates with the samples' read density (number of reads divided by genome size). For high-complexity metagenomic samples lacking dominant species, natural duplicates only make up <1% of all duplicates. But for some other samples like transcriptomic samples, majority of the observed duplicates might be natural duplicates. CONCLUSIONS: Our method is available from http://cd-hit.org as a downloadable program and a web server. It is important not only to identify the duplicates from metagenomic datasets but also to distinguish whether they are artificial or natural duplicates. We provide a tool to estimate the number of natural duplicates according to user-defined sample types, so users can decide whether to retain or remove duplicates in their projects.
Beifang Niu, Limin Fu, Shulei Sun, Weizhong Li 0002
BMC Bioinform.1