Rendong Yang

dblp:47/8321 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Heterogeneous federated learning framework based on dynamic class awareness and gated collaborative optimization
Jinquan Zhang 0001, Rendong Yang, Yuncan Tang, Lina Ni
Knowl. Based Syst.2
2025 A comprehensive benchmark of tools for efficient genomic interval querying
abstract
Efficiently querying genomic intervals is fundamental to modern bioinformatics, enabling researchers to extract and analyze specific regions from large genomic datasets. While various tools have been developed for this purpose, there lacks a comprehensive comparison of their performance, memory usage, and practical utility. We present a systematic evaluation of genomic interval query tools using simulated datasets of varying sizes. Our benchmarking framework, segmeter, assesses both basic and complex interval queries, examining runtime performance, memory efficiency, and query precision across different tools. This comprehensive analysis provides insights into the strengths and limitations of different approaches to genomic interval querying, offering guidance for tool selection based on specific use cases and data requirements. The segmeter framework and all benchmark data are freely available, facilitating reproducibility and enabling researchers to conduct their own comparative analyses.
Richard A. Schäfer, Rendong Yang
Briefings Bioinform.2
2025 OctopuSV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis
abstract
MOTIVATION: Structural variants (SVs) influence gene regulation, disease progression, and diagnostics, yet integrating SV calls across platforms remains difficult due to inconsistent annotations, limited merging flexibility, and fragmented workflows. Ambiguous breakend (BND) annotations, which comprise many variant calls, are often discarded or misclassified, hindering variant characterization. Existing tools lack advanced merging operations essential for precise identification of disease-specific or somatic variants across samples or patient groups. Additionally, current SV analysis pipelines require extensive manual intervention and complex parameter tuning, compromising reproducibility and scalability. Addressing these gaps is crucial for improving the accuracy, interpretability, and clinical utility of SV analyses. RESULTS: We developed OctopuSV and TentacleSV to address these long-standing challenges in SV analysis. OctopuSV features a specialized BND correction module that converts ambiguous BND annotations into canonical SV types, recovering important variants that are often overlooked by existing tools. Additionally, it provides advanced set operations (difference, complement, custom-defined) that enable sophisticated variant filtering without programming expertise, critical for identifying tumor-specific SVs or variants unique to specific sample groups. TentacleSV completes our solution by automating the entire SV analysis process from raw sequencing data to high-confidence callsets, ensuring consistency and reproducibility across projects. Benchmarking across short-read and long-read platforms showed superior F1 score, complete SV type consistency compared to existing tools. Our framework enables experimental biologists and clinical researchers to perform sophisticated analyses ranging from cancer subtype-specific SV identification to multi-sample comparative studies without requiring specialized programming skills. AVAILABILITY AND IMPLEMENTATION: All codes are available at https://github.com/ylab-hi/OctopuSV; https://github.com/ylab-hi/TentacleSV.
Qingxiang Guo, Ting-You Wang, Abhirami Ramakrishnan, Rendong Yang
Bioinform.5
2024 Comparative analyses of gene networks mediating cancer metastatic potentials across lineage types
abstract
Studies have identified genes and molecular pathways regulating cancer metastasis. However, it remains largely unknown whether metastatic potentials of cancer cells from different lineage types are driven by the same or different gene networks. Here, we aim to address this question through integrative analyses of 493 human cancer cells' transcriptomic profiles and their metastatic potentials in vivo. Using an unsupervised approach and considering both gene coexpression and protein-protein interaction networks, we identify different gene networks associated with various biological pathways (i.e. inflammation, cell cycle, and RNA translation), the expression of which are correlated with metastatic potentials across subsets of lineage types. By developing a regularized random forest regression model, we show that the combination of the gene module features expressed in the native cancer cells can predict their metastatic potentials with an overall Pearson correlation coefficient of 0.90. By analyzing transcriptomic profile data from cancer patients, we show that these networks are conserved in vivo and contribute to cancer aggressiveness. The intrinsic expression levels of these networks are correlated with drug sensitivity. Altogether, our study provides novel comparative insights into cancer cells' intrinsic gene networks mediating metastatic potentials across different lineage types, and our results can potentially be useful for designing personalized treatments for metastatic cancers.
Emily Kunce Stroup, Ting-You Wang, Rendong Yang
Briefings Bioinform.4
2024 PxBLAT: an efficient python binding library for BLAT
abstract
BACKGROUND: With the surge in genomic data driven by advancements in sequencing technologies, the demand for efficient bioinformatics tools for sequence analysis has become paramount. BLAST-like alignment tool (BLAT), a sequence alignment tool, faces limitations in performance efficiency and integration with modern programming environments, particularly Python. This study introduces PxBLAT, a Python-based framework designed to enhance the capabilities of BLAT, focusing on usability, computational efficiency, and seamless integration within the Python ecosystem. RESULTS: PxBLAT demonstrates significant improvements over BLAT in execution speed and data handling, as evidenced by comprehensive benchmarks conducted across various sample groups ranging from 50 to 600 samples. These experiments highlight a notable speedup, reducing execution time compared to BLAT. The framework also introduces user-friendly features such as improved server management, data conversion utilities, and shell completion, enhancing the overall user experience. Additionally, the provision of extensive documentation and comprehensive testing supports community engagement and facilitates the adoption of PxBLAT. CONCLUSIONS: PxBLAT stands out as a robust alternative to BLAT, offering performance and user interaction enhancements. Its development underscores the potential for modern programming languages to improve bioinformatics tools, aligning with the needs of contemporary genomic research. By providing a more efficient, user-friendly tool, PxBLAT has the potential to impact genomic data analysis workflows, supporting faster and more accurate sequence analysis in a Python environment.
Rendong Yang
BMC Bioinform.2
2023 ScanNeo2: a comprehensive workflow for neoantigen detection and immunogenicity prediction from diverse genomic and transcriptomic alterations
abstract
MOTIVATION: Neoantigens, tumor-specific protein fragments, are invaluable in cancer immunotherapy due to their ability to serve as targets for the immune system. Computational prediction of these neoantigens from sequencing data often requires multiple algorithms and sophisticated workflows, which are currently restricted to specific types of variants, such as single-nucleotide variants or insertions/deletions. Nevertheless, other sources of neoantigens are often overlooked. RESULTS: We introduce ScanNeo2 an improved and fully automated bioinformatics pipeline designed for high-throughput neoantigen prediction from raw sequencing data. Unlike its predecessor, ScanNeo2 integrates multiple sources of somatic variants, including canonical- and exitron-splicing, gene fusion events, and various somatic variants. Our benchmark results demonstrate that ScanNeo2 accurately identifies neoantigens, providing a comprehensive and more efficient solution for neoantigen prediction. AVAILABILITY AND IMPLEMENTATION: ScanNeo2 is freely available at https://github.com/ylab-hi/ScanNeo2/ and is accompanied by instruction and application data.
Richard A. Schäfer, Qingxiang Guo, Rendong Yang
Bioinform.3
2022 ScanExitronLR: characterization and quantification of exitron splicing events in long-read RNA-seq data
abstract
SUMMARY: Exitron splicing is a type of alternative splicing where coding sequences are spliced out. Recently, exitron splicing has been shown to increase proteome plasticity and play a role in cancer. Long-read RNA-seq is well suited for quantification and discovery of alternative splicing events; however, there are currently no tools available for the detection and annotation of exitrons in long-read RNA-seq data. Here, we present ScanExitronLR, an application for the characterization and quantification of exitron splicing events in long-reads. From a BAM alignment file, reference genome and reference gene annotation, ScanExitronLR outputs exitron events at the individual transcript level. Outputs of ScanExitronLR can be used in downstream analyses of differential exitron splicing. In addition, ScanExitronLR optionally reports exitron annotations such as truncation or frameshift type, nonsense-mediated decay status and Pfam domain interruptions. We demonstrate that ScanExitronLR performs better on noisy long-reads than currently published exitron detection algorithms designed for short-read data. AVAILABILITY AND IMPLEMENTATION: ScanExitronLR is freely available at https://github.com/ylab-hi/ScanExitronLR and distributed as a pip package on the Python Package Index. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Joshua Fry, Rendong Yang
Bioinform.3
2019 ScanNeo: identifying indel-derived neoantigens using RNA-Seq data
abstract
SUMMARY: Insertion and deletion (indels) have been recognized as an important source generating tumor-specific mutant peptides (neoantigens). The focus of indel-derived neoantigen identification has been on leveraging DNA sequencing such as whole exome sequencing, with the effort of using RNA-seq less well explored. Here we present ScanNeo, a fast-streamlined computational pipeline for analyzing RNA-seq to predict neoepitopes derived from small to large-sized indels. We applied ScanNeo in a prostate cancer cell line and validated our predictions with matched mass spectrometry data. Finally, we demonstrated that indel neoantigens predicted from RNA-seq were associated with checkpoint inhibitor response in a cohort of melanoma patients. AVAILABILITY AND IMPLEMENTATION: ScanNeo is implemented in Python. It is freely accessible at the GitHub repository (https://github.com/ylab-hi/ScanNeo). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ting-You Wang, Li Wang 0141, Sk Kayum Alam, Luke H. Hoeppner, Rendong Yang
Bioinform.5
2011 LSPR: an integrated periodicity detection algorithm for unevenly sampled temporal microarray data
abstract
UNLABELLED: We propose a three-step periodicity detection algorithm named LSPR. Our method first preprocesses the raw time-series by removing the linear trend and filtering noise. In the second step, LSPR employs a Lomb-Scargle periodogram to estimate the periodicity in the time-series. Finally, harmonic regression is applied to model the cyclic components. Inferred periodic transcripts are selected by a false discovery rate procedure. We have applied LSPR to unevenly sampled synthetic data and two Arabidopsis diurnal expression datasets, and compared its performance with the existing well-established algorithms. Results show that LSPR is capable of identifying periodic transcripts more accurately than existing algorithms. AVAILABILITY: LSPR algorithm is implemented as MATLAB software and is available at http://bioinformatics.cau.edu.cn/LSPR.
Rendong Yang
Bioinform.1
2010 Analyzing circadian expression data by harmonic regression based on autoregressive spectral estimation
abstract
MOTIVATION: Circadian rhythms are prevalent in most organisms. Identification of circadian-regulated genes is a crucial step in discovering underlying pathways and processes that are clock-controlled. Such genes are largely detected by searching periodic patterns in microarray data. However, temporal gene expression profiles usually have a short time-series with low sampling frequency and high levels of noise. This makes circadian rhythmic analysis of temporal microarray data very challenging. RESULTS: We propose an algorithm named ARSER, which combines time domain and frequency domain analysis for extracting and characterizing rhythmic expression profiles from temporal microarray data. ARSER employs autoregressive spectral estimation to predict an expression profile's periodicity from the frequency spectrum and then models the rhythmic patterns by using a harmonic regression model to fit the time-series. ARSER describes the rhythmic patterns by four parameters: period, phase, amplitude and mean level, and measures the multiple testing significance by false discovery rate q-value. When tested on well defined periodic and non-periodic short time-series data, ARSER was superior to two existing and widely-used methods, COSOPT and Fisher's G-test, during identification of sinusoidal and non-sinusoidal periodic patterns in short, noisy and non-stationary time-series. Finally, analysis of Arabidopsis microarray data using ARSER led to identification of a novel set of previously undetected non-sinusoidal periodic transcripts, which may lead to new insights into molecular mechanisms of circadian rhythms. AVAILABILITY: ARSER is implemented by Python and R. All source codes are available from http://bioinformatics.cau.edu.cn/ARSER.
Rendong Yang
Bioinform.1