Juexin Wang

dblp:81/6726 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-2260-4310ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 scBSP: a fast and accurate tool for identifying spatially variable features from high-resolution spatial omics data
abstract
MOTIVATION: Emerging spatial omics technologies empower comprehensive exploration of biological systems from multi-omics perspectives in their native tissue location in 2D and 3D space. However, the limited sequencing depth, increasing spatial resolution, and growing spatial spots in spatial omics technologies present significant computational challenges in identifying biologically meaningful molecules with variable spatial distributions across various omics modalities. RESULTS: We introduce scBSP, an open-source, versatile, and user-friendly package for identifying spatially variable features in large-scale spatial omics data. scBSP demonstrates significantly enhanced computational efficiency, processing high-resolution spatial omics data within seconds, and exhibits robust cross-platform performance by consistently identifying spatially variable features with high reproducibility across various sequencing platforms. AVAILABILITY AND IMPLEMENTATION: scBSP is available for download from R CRAN at https://cran.r-project.org/web/packages/scBSP/index.html and PyPI at https://pypi.org/project/scbsp/.
Jinpu Li, Mauminah Raina, Ricardo Melo Ferreira, Michael Eadon, Qin Ma 0003, Juexin Wang, Dong Xu 0002
Bioinform.10
2025 Relation equivariant graph neural networks to explore the mosaic-like tissue architecture of kidney diseases on spatially resolved transcriptomics
abstract
MOTIVATION: Chronic kidney disease (CKD) and acute kidney injury (AKI) are prominent public health concerns affecting more than 15% of the global population. The ongoing development of spatially resolved transcriptomics (SRT) technologies presents a promising approach for discovering the spatial distribution patterns of gene expression within diseased tissues. However, existing computational tools are predominantly calibrated and designed on the ribbon-like structure of the brain cortex, presenting considerable computational obstacles in discerning highly heterogeneous mosaic-like tissue architectures in the kidney. Consequently, timely and cost-effective acquisition of annotation and interpretation in the kidney remains a challenge in exploring the cellular and morphological changes within renal tubules and their interstitial niches. RESULTS: We present an empowered graph deep learning framework, REGNN (Relation Equivariant Graph Neural Networks), designed for SRT data analyses on heterogeneous tissue structures. To increase expressive power in the SRT lattice using graph modeling, REGNN integrates equivariance to handle n-dimensional symmetries of the spatial area, while additionally leveraging Positional Encoding to strengthen relative spatial relations of the nodes uniformly distributed in the lattice. Given the limited availability of well-labeled spatial data, this framework implements both graph autoencoder and graph self-supervised learning strategies. On heterogeneous samples from different kidney conditions, REGNN outperforms existing computational tools in identifying tissue architectures within the 10× Visium platform. This framework offers a powerful graph deep learning tool for investigating tissues within highly heterogeneous expression patterns and paves the way to pinpoint underlying pathological mechanisms that contribute to the progression of complex diseases. AVAILABILITY AND IMPLEMENTATION: REGNN is publicly available at https://github.com/Mraina99/REGNN.
Mauminah Raina, Ricardo Melo Ferreira, Treyden Stansfield, Chandrima Modak, Ying-Hua Cheng, Hari Naga Sai Kiran Suryadevara, Dong Xu 0002, Michael Eadon, Qin Ma 0003, Juexin Wang
Bioinform.11
2022 Machine learning development environment for single-cell sequencing data analyses
abstract
Machine learning (ML) is transforming single-cell sequencing data analysis; however, the barriers of technology complexity and biology knowledge remain challenging for the involvement of the ML community in single-cell data analysis. Here we present an ML development environment for single-cell sequencing data analyses, together with a diverse set of realistic and accessible ML-Ready benchmark datasets. A cloud-based platform is built to dynamically scale workflows for collecting, processing, and managing various single-cell sequencing data to make them ML-ready. In addition, benchmarks for each problem formulation and a code-level and web-interface IDE for single-cell analysis method development are provided. These efforts provide an automated end-to-end single-cell analysis ML pipeline that simplifies and standardizes the process of single-cell data formatting, loading, model development, and model evaluation.
Yuexu Jiang, Cankun Wang, Clement Essien, Juexin Wang, Anjun Ma, Qin Ma 0003, Dong Xu 0002
BIBM5
2022 scGNN 2.0: a graph neural network tool for imputation and clustering of single-cell RNA-Seq data
abstract
MOTIVATION: Gene expression imputation has been an essential step of the single-cell RNA-Seq data analysis workflow. Among several deep-learning methods, the debut of scGNN gained substantial recognition in 2021 for its superior performance and the ability to produce a cell-cell graph. However, the implementation of scGNN was relatively time-consuming and its performance could still be optimized. RESULTS: The implementation of scGNN 2.0 is significantly faster than scGNN thanks to a simplified close-loop architecture. For all eight datasets, cell clustering performance was increased by 85.02% on average in terms of adjusted rand index, and the imputation Median L1 Error was reduced by 67.94% on average. With the built-in visualizations, users can quickly assess the imputation and cell clustering results, compare against benchmarks and interpret the cell-cell interaction. The expanded input and output formats also pave the way for custom workflows that integrate scGNN 2.0 with other scRNA-Seq toolkits on both Python and R platforms. AVAILABILITY AND IMPLEMENTATION: scGNN 2.0 is implemented in Python (as of version 3.8) with the source code available at https://github.com/OSU-BMBL/scGNN2.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haocheng Gu, Anjun Ma, Yang Li 0089, Juexin Wang, Dong Xu 0002, Qin Ma 0003
Bioinform.5
2021 The bioinformatics toolbox for circRNA discovery and analysis
abstract
Circular RNAs (circRNAs) are a unique class of RNA molecule identified more than 40 years ago which are produced by a covalent linkage via back-splicing of linear RNA. Recent advances in sequencing technologies and bioinformatics tools have led directly to an ever-expanding field of types and biological functions of circRNAs. In parallel with technological developments, practical applications of circRNAs have arisen including their utilization as biomarkers of human disease. Currently, circRNA-associated bioinformatics tools can support projects including circRNA annotation, circRNA identification and network analysis of competing endogenous RNA (ceRNA). In this review, we collected about 100 circRNA-associated bioinformatics tools and summarized their current attributes and capabilities. We also performed network analysis and text mining on circRNA tool publications in order to reveal trends in their ongoing development.
Liang Chen 0021, Changliang Wang, Huiyan Sun, Juexin Wang, Yanchun Liang 0001, Yan Wang 0028, Garry Wong
Briefings Bioinform.4
2018 G2S: a web-service for annotating genomic variants on 3D protein structures
abstract
Motivation: Accurately mapping and annotating genomic locations on 3D protein structures is a key step in structure-based analysis of genomic variants detected by recent large-scale sequencing efforts. There are several mapping resources currently available, but none of them provides a web API (Application Programming Interface) that supports programmatic access. Results: We present G2S, a real-time web API that provides automated mapping of genomic variants on 3D protein structures. G2S can align genomic locations of variants, protein locations, or protein sequences to protein structures and retrieve the mapped residues from structures. G2S API uses REST-inspired design and it can be used by various clients such as web browsers, command terminals, programming languages and other bioinformatics tools for bringing 3D structures into genomic variant analysis. Availability and implementation: The webserver and source codes are freely available at https://g2s.genomenexus.org. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Juexin Wang, Robert Sheridan, Selçuk Onur Sümer, Nikolaus Schultz, Dong Xu 0002, Jianjiong Gao
Bioinform.1
2017 SoyTSN: A web-based prediction tool for soybean tissue specific network within SoyKB
abstract
Soybean tissue-specific network helps identify and visualize the gene-gene relationships in various tissues [1]. We have built SoyTSN, a web-based tool for tissue-specific network prediction in soybean using 14 tissues RNA-Seq datasets including flower, root, nodule, leaf, stem, seed, etc. SoyTSN first combines multiple tissue specific RNA-Seq studies, and later Cross-Conditions Cluster Detection (C3D) algorithm [2] was applied to detect modules based on co-expression relationships across all the tissues. Following this soybean tissue-specific interactomes were inferred by combining tissue-specific expression and protein-protein interaction from the STRING database [3]. All these relationships between various soybean genes were collected and stored in Soybean Knowledge Base (SoyKB)[4] and can be queried using SoyTSN. For every query gene, SoyTSN computes and visualizes any of the 14 tissue specific networks both at the expression and interactome level. Users can compare the gene-gene relationship differences at different confidence levels across all the soybean tissues.
Juexin Wang, Zhen Lyu, Shakhawat Hossain, Gary Stacey, Dong Xu 0002, Trupti Joshi
BIBM1
2016 PGen: large-scale genomic variations analysis workflow and browser in SoyKB
abstract
BACKGROUND: With the advances in next-generation sequencing (NGS) technology and significant reductions in sequencing costs, it is now possible to sequence large collections of germplasm in crops for detecting genome-scale genetic variations and to apply the knowledge towards improvements in traits. To efficiently facilitate large-scale NGS resequencing data analysis of genomic variations, we have developed "PGen", an integrated and optimized workflow using the Extreme Science and Engineering Discovery Environment (XSEDE) high-performance computing (HPC) virtual system, iPlant cloud data storage resources and Pegasus workflow management system (Pegasus-WMS). The workflow allows users to identify single nucleotide polymorphisms (SNPs) and insertion-deletions (indels), perform SNP annotations and conduct copy number variation analyses on multiple resequencing datasets in a user-friendly and seamless way. RESULTS: We have developed both a Linux version in GitHub ( https://github.com/pegasus-isi/PGen-GenomicVariations-Workflow ) and a web-based implementation of the PGen workflow integrated within the Soybean Knowledge Base (SoyKB), ( http://soykb.org/Pegasus/index.php ). Using PGen, we identified 10,218,140 single-nucleotide polymorphisms (SNPs) and 1,398,982 indels from analysis of 106 soybean lines sequenced at 15X coverage. 297,245 non-synonymous SNPs and 3330 copy number variation (CNV) regions were identified from this analysis. SNPs identified using PGen from additional soybean resequencing projects adding to 500+ soybean germplasm lines in total have been integrated. These SNPs are being utilized for trait improvement using genotype to phenotype prediction approaches developed in-house. In order to browse and access NGS data easily, we have also developed an NGS resequencing data browser ( http://soykb.org/NGS_Resequence/NGS_index.php ) within SoyKB to provide easy access to SNP and downstream analysis results for soybean researchers. CONCLUSION: PGen workflow has been optimized for the most efficient analysis of soybean data using thorough testing and validation. This research serves as an example of best practices for development of genomics data analysis workflows by integrating remote HPC resources and efficient data management with ease of use for biological users. PGen workflow can also be easily customized for analysis of data in other species.
Saad M. Khan, Juexin Wang, Mats Rynge, Yuanxun Zhang, Shiyuan Chen, João V. Maldonado dos Santos, Babu Valliyodan, Prasad Calyam, Nirav C. Merchant, Henry T. Nguyen, Dong Xu 0002, Trupti Joshi
BMC Bioinform.3
2009 Immune Particle Swarm Optimization for Support Vector Regression on Forest Fire Prediction
Yan Wang 0028, Juexin Wang, Wei Du 0002, Chuncai Wang, Yanchun Liang 0001, Chunguang Zhou, Lan Huang 0002
ISNN (2)2