Yongjie Xu 0001

dblp:123/9257-1 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0002-6045-1626ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FGeneBERT: function-driven pre-trained gene language model for metagenomics
abstract
Metagenomic data, comprising mixed multi-species genomes, are prevalent in diverse environments like oceans and soils, significantly impacting human health and ecological functions. However, current research relies on K-mer, which limits the capture of structurally and functionally relevant gene contexts. Moreover, these approaches struggle with encoding biologically meaningful genes and fail to address the one-to-many and many-to-one relationships inherent in metagenomic data. To overcome these challenges, we introduce FGeneBERT, a novel metagenomic pre-trained model that employs a protein-based gene representation as a context-aware and structure-relevant tokenizer. FGeneBERT incorporates masked gene modeling to enhance the understanding of inter-gene contextual relationships and triplet enhanced metagenomic contrastive learning to elucidate gene sequence-function relationships. Pre-trained on over 100 million metagenomic sequences, FGeneBERT demonstrates superior performance on metagenomic datasets at four levels, spanning gene, functional, bacterial, and environmental levels and ranging from 1 to 213 k input sequences. Case studies of ATP synthase and gene operons highlight FGeneBERT's capability for functional recognition and its biological relevance in metagenomic research.
Chenrui Duan, Zelin Zang, Yongjie Xu 0001, Hang He, Siyuan Li 0002, Zhen Lei 0001, Ju-Sheng Zheng, Stan Z. Li
Briefings Bioinform.3
2025 Complex hierarchical structures analysis in single-cell data with Poincaré deep manifold transformation
abstract
Single-cell RNA sequencing (scRNA-seq) offers remarkable insights into cellular development and differentiation by capturing the gene expression profiles of individual cells. The role of dimensionality reduction and visualization in the interpretation of scRNA-seq data has gained widely acceptance. However, current methods face several challenges, including incomplete structure-preserving strategies and high distortion in embeddings, which fail to effectively model complex cell trajectories with multiple branches. To address these issues, we propose the Poincaré deep manifold transformation (PoincaréDMT) method, which maps high-dimensional scRNA-seq data to a hyperbolic Poincaré disk. This approach preserves global structure from a graph Laplacian matrix while achieving local structure correction through a structure module combined with data augmentation. Additionally, PoincaréDMT alleviates batch effects by integrating a batch graph that accounts for batch labels into the low-dimensional embeddings during network training. Furthermore, PoincaréDMT introduces the Shapley additive explanations method based on trained model to identify the important marker genes in specific clusters and cell differentiation process. Therefore, PoincaréDMT provides a unified framework for multiple key tasks essential for scRNA-seq analysis, including trajectory inference, pseudotime inference, batch correction, and marker gene selection. We validate PoincaréDMT through extensive evaluations on both simulated and real scRNA-seq datasets, demonstrating its superior performance in preserving global and local data structures compared to existing methods.
Yongjie Xu 0001, Zelin Zang, Bozhen Hu, Cheng Tan 0012, Jun Xia 0001, Stan Z. Li
Briefings Bioinform.1
2025 MuST: multiple-modality structure transformation for single-cell spatial transcriptomics
abstract
Spatial transcriptomics (ST) technologies have revolutionized the study of gene expression patterns in tissues by providing multimodal data, including transcriptomic (Tra.), spatial, and morphological modalities, thereby offering new opportunities to understand tissue biology beyond traditional Tra. However, we identify the modality bias phenomenon in ST data species, i.e. the inconsistent contribution of different modalities to the labels leads to a tendency for the analysis methods to retain the information of the dominant modality. How to mitigate the adverse effects of modality bias to satisfy various downstream tasks remains a fundamental challenge. This paper introduces Multiple-modality Structure Transformation, named MuST, a novel methodology to tackle the challenge. MuST integrates the multi-modality information contained in the ST data effectively into a uniform latent space to provide a foundation for all the downstream tasks. It learns intrinsic local structures by topology discovery strategy and topology fusion loss function to solve the inconsistencies among different modalities. Thus, these topology-based and deep learning techniques provide a solid foundation for a variety of analytical tasks while coordinating different modalities. The effectiveness of MuST is assessed by performance metrics and biological significance. The results show that it outperforms existing state-of-the-art methods with clear advantages in the precision of identifying and preserving structures of tissues and biomarkers. MuST offers a versatile toolkit for the intricate analysis of complex biological systems.
Zelin Zang, Yongjie Xu 0001, Chenrui Duan, Zhen Lei 0001, Stan Z. Li
Briefings Bioinform.3
2024 MuST: Maximizing the Latent Capacity of Spatial Transcriptomics Data with Multi-modality Structure Transformation
abstract
Spatial transcriptomics (ST) technologies have revolutionized the study of gene expression patterns in tissues by providing multimodality data in transcriptomic, spatial, and morphological, offering opportunities for understanding tissue biology beyond transcriptomics. However, we identify the modality bias phenomenon in ST data species, i.e., the inconsistent contribution of different modalities to the labels leads to a tendency for the analysis methods to retain the information of the dominant modality. How to mitigate the adverse effects of modality bias to satisfy various downstream tasks remains a fundamental challenge. This paper introduces Multiple-modality Structure Transformation, named MuST, a novel methodology to tackle the challenge. MuST integrates the multi-modality information contained in the ST data effectively into a uniform latent space to provide a foundation for all the downstream tasks. It learns intrinsic local structures by topology discovery strategy and topology fusion loss function to solve the inconsistencies among different modalities. Thus, these topology-based and deep learning techniques provide a solid foundation for a variety of analytical tasks while coordinating different modalities. The effectiveness of MuST is assessed by performance metrics and biological significance. The results show that it outperforms existing state-of-the-art methods with clear advantages in the precision of identifying and preserving structures of tissues and biomarkers. MuST offers a versatile toolkit for the intricate analysis of complex biological systems. The code is available at https://github.com/zangzelin/code_Must.
Zelin Zang, Yongjie Xu 0001, Chenrui Duan, Zhen Lei 0001, Stan Z. Li
BIBM3
2024 PhyloGen: Language Model-Enhanced Phylogenetic Inference via Graph Structure Generation
abstract
Phylogenetic trees elucidate evolutionary relationships among species, but phylogenetic inference remains challenging due to the complexity of combining continuous (branch lengths) and discrete parameters (tree topology). Traditional Markov Chain Monte Carlo methods face slow convergence and computational burdens. Existing Variational Inference methods, which require pre-generated topologies and typically treat tree structures and branch lengths independently, may overlook critical sequence features, limiting their accuracy and flexibility. We propose PhyloGen, a novel method leveraging a pre-trained genomic language model to generate and optimize phylogenetic trees without dependence on evolutionary models or aligned sequence constraints. PhyloGen views phylogenetic inference as a conditionally constrained tree structure generation problem, jointly optimizing tree topology and branch lengths through three core modules: (i) Feature Extraction, (ii) PhyloTree Construction, and (iii) PhyloTree Structure Modeling. Meanwhile, we introduce a Scoring Function to guide the model towards a more stable gradient descent. We demonstrate the effectiveness and robustness of PhyloGen on eight real-world benchmark datasets. Visualization results confirm PhyloGen provides deeper insights into phylogenetic relationships.
Chenrui Duan, Zelin Zang, Siyuan Li 0002, Yongjie Xu 0001, Stan Z. Li
NeurIPS4
2024 Learning Complete Protein Representation by Dynamically Coupling of Sequence and Structure
abstract
Learning effective representations is imperative for comprehending proteins and deciphering their biological functions. Recent strides in language models and graph neural networks have empowered protein models to harness primary or tertiary structure information for representation learning. Nevertheless, the absence of practical methodologies to appropriately model intricate inter-dependencies between protein sequences and structures has resulted in embeddings that exhibit low performance on tasks such as protein function prediction. In this study, we introduce CoupleNet, a novel framework designed to interlink protein sequences and structures to derive informative protein representations. CoupleNet integrates multiple levels and scales of features in proteins, encompassing residue identities and positions for sequences, as well as geometric representations for tertiary structures from both local and global perspectives. A two-type dynamic graph is constructed to capture adjacent and distant sequential features and structural geometries, achieving completeness at the amino acid and backbone levels. Additionally, convolutions are executed on nodes and edges simultaneously to generate comprehensive protein embeddings. Experimental results on benchmark datasets showcase that CoupleNet outperforms state-of-the-art methods, exhibiting particularly superior performance in low-sequence similarities scenarios, adeptly identifying infrequently encountered functions and effectively capturing remote homology relationships in proteins.
Bozhen Hu, Cheng Tan 0012, Jun Xia 0001, Yue Liu 0008, Lirong Wu, Jiangbin Zheng 0002, Yongjie Xu 0001, Yufei Huang 0002, Stan Z. Li
NeurIPS7
2024 ProtGO: Function-Guided Protein Modeling for Unified Representation Learning
abstract
Protein representation learning is indispensable for various downstream applications of artificial intelligence for bio-medicine research, such as drug design and function prediction. However, achieving effective representation learning for proteins poses challenges due to the diversity of data modalities involved, including sequence, structure, and function annotations. Despite the impressive capabilities of large language models in biomedical text modelling, there remains a pressing need for a framework that seamlessly integrates these diverse modalities, particularly focusing on the three critical aspects of protein information: sequence, structure, and function. Moreover, addressing the inherent data scale differences among these modalities is essential. To tackle these challenges, we introduce ProtGO, a unified model that harnesses a teacher network equipped with a customized graph neural network (GNN) and a Gene Ontology (GO) encoder to learn hybrid embeddings. Notably, our approach eliminates the need for additional functions as input for the student network, which shares the same GNN module. Importantly, we utilize a domain adaptation method to facilitate distribution approximation for guiding the training of the teacher-student framework. This approach leverages distributions learned from latent representations to avoid the alignment of individual samples. Benchmark experiments highlight that ProtGO significantly outperforms state-of-the-art baselines, clearly demonstrating the advantages of the proposed unified framework.
Bozhen Hu, Cheng Tan 0012, Yongjie Xu 0001, Zhangyang Gao, Jun Xia 0001, Lirong Wu, Stan Z. Li
NeurIPS3
2024 GNN Cleaner: Label Cleaner for Graph Structured Data
abstract
Graph Neural Network (GNN) has emerged as a predominant tool for graph data analysis. Despite their proliferation, the low-quality labels of many real-world graphs will undermine their performance dramatically. Existing studies on learning neural networks with noisy labels mainly focus on independent data and thus cannot fully exploit the structural information of graph data. Currently, there are few studies of robustness to noisy labels for graph-structured data even if this problem is commonly seen in real-world settings. To remedy this deficiency, we proposeGNN Cleanerwhich utilizes structural information of graph data to combat noisy labels. More specifically, a pseudo label is computed from the neighboring labels for each node in the training set via a modified version of label propagation. Additionally, a novel method is developed to learn to correct the labels adaptively and dynamically. Extensive experiments show that GNN Cleaner can train GNNs robustly and correct both the synthetic and real-world noisy labels even if the noise is severe. Moreover, GNN Cleaner is model-agnostic and can be combined with various GNNs to improve their robustness against label noise.
Jun Xia 0001, Yongjie Xu 0001, Cheng Tan 0012, Lirong Wu, Siyuan Li 0002, Stan Z. Li
IEEE Trans. Knowl. Data Eng.3
2024 DMT-EV: An Explainable Deep Network for Dimension Reduction
abstract
Dimension reduction (DR) is commonly utilized to capture the intrinsic structure and transform high-dimensional data into low-dimensional space while retaining meaningful properties of the original data. It is used in various applications, such as image recognition, single-cell sequencing analysis, and biomarker discovery. However, contemporary parametric-free and parametric DR techniques suffer from several significant shortcomings, such as the inability to preserve global and local features and the poor generalisation performance. On the other hand, regarding explainability, it is crucial to comprehend the embedding process, especially the contribution of each part to the embedding process, while understanding how each feature affects the embedding results that identify critical components and help diagnose the embedding process. To address these problems, we have developed a deep neural network method called DMT-EV, which provides not only excellent performance in structural maintainability but also explainability to the DR therein. DMT-EV starts with data augmentation and a manifold-based loss function to improve embedding performance. The explanation is based on saliency maps and aims to examine the trained DMT-EV parameters and contributions of components during the embedding process. The proposed techniques are integrated with a visual interface to help the user to adjust DMT-EV to achieve better DR performance and explainability. The interactive visual interface makes it easier to illustrate the data features, compare different DR techniques, and investigate DR. An in-depth experimental comparison shows that DMT-EV consistently outperforms the state-of-the-art methods in both performance measures and explainability.
Zelin Zang, Shenghui Cheng, Hanchen Xia, Yaoting Sun, Yongjie Xu 0001, Baigui Sun, Stan Z. Li
IEEE Trans. Vis. Comput. Graph.6
2023 Temporal Attention Unit: Towards Efficient Spatiotemporal Predictive Learning
abstract
Spatiotemporal predictive learning aims to generate future frames by learning from historical frames. In this paper, we investigate existing methods and present a general framework of spatiotemporal predictive learning, in which the spatial encoder and decoder capture intra-frame features and the middle temporal module catches inter-frame correlations. While the mainstream methods employ recurrent units to capture long-term temporal dependencies, they suffer from low computational efficiency due to their unparallelizable architectures. To parallelize the temporal module, we propose the Temporal Attention Unit (TAU), which decomposes temporal attention into intra-frame statical attention and inter-frame dynamical attention. Moreover, while the mean squared error loss focuses on intra-frame errors, we introduce a novel differential divergence regularization to take inter-frame variations into account. Extensive experiments demonstrate that the proposed method enables the derived model to achieve competitive performance on various spatiotemporal prediction benchmarks.
Cheng Tan 0012, Zhangyang Gao, Lirong Wu, Yongjie Xu 0001, Jun Xia 0001, Siyuan Li 0002, Stan Z. Li
CVPR4
2023 Wordreg: Mitigating the Gap between Training and Inference with Worst-Case Drop Regularization
abstract
Dropout has emerged as one of the most frequently used techniques for training deep neural networks (DNNs). Although effective, the sampled sub-model by random dropout during training is inconsistent with the full model (without dropout) during inference. To mitigate this undesirable gap, we propose WordReg, a simple yet effective regularization built on dropout that enforces the consistency between the outputs of different sub-models sampled by dropout. Specifically, WordReg first obtains the worst-case dropout by maximizing the divergence between the outputs with two sub-models with different random dropouts. And then, it encourages the agreements between the outputs of the two sub-models with worstcase divergence. Extensive experiments on diverse DNNs and tasks reveal that WordReg can achieve notable and consistent improvements over non-regularized models and yields some state-of-the-art results. Theoretically, we verify that WordReg can reduce the gap between training and inference.
Jun Xia 0001, Bozhen Hu, Cheng Tan 0012, Jiangbin Zheng 0002, Yongjie Xu 0001, Stan Z. Li
ICASSP6
2023 UDRN: Unified Dimensional Reduction Neural Network for feature selection and feature projection
Zelin Zang, Yongjie Xu 0001, Linyan Lu, Yulan Geng, Senqiao Yang, Stan Z. Li
Neural Networks2
2022 Conditional Local Convolution for Spatio-Temporal Meteorological Forecasting
abstract
Spatio-temporal forecasting is challenging attributing to the high nonlinearity in temporal dynamics as well as complex location-characterized patterns in spatial domains, especially in fields like weather forecasting. Graph convolutions are usually used for modeling the spatial dependency in meteorology to handle the irregular distribution of sensors' spatial location. In this work, a novel graph-based convolution for imitating the meteorological flows is proposed to capture the local spatial patterns. Based on the assumption of smoothness of location-characterized patterns, we propose conditional local convolution whose shared kernel on nodes' local space is approximated by feedforward networks, with local representations of coordinate obtained by horizon maps into cylindrical-tangent space as its input. The established united standard of local coordinate system preserves the orientation on geography. We further propose the distance and orientation scaling terms to reduce the impacts of irregular spatial distribution. The convolution is embedded in a Recurrent Neural Network architecture to model the temporal dynamics, leading to the Conditional Local Convolution Recurrent Network (CLCRN). Our model is evaluated on real-world weather benchmark datasets, achieving state-of-the-art performance with obvious improvements. We conduct further analysis on local pattern visualization, model's framework choice, advantages of horizon maps and etc. The source code is available at https://github.com/BIRD-TAO/CLCRN.
Zhangyang Gao, Yongjie Xu 0001, Lirong Wu, Stan Z. Li
AAAI3
2022 OT Cleaner: Label Correction as Optimal Transport
abstract
Datasets with noisy labels present challenges for training Deep Neural Networks (DNNs) with high generalization ability. An direct idea is to correct the noisy labels for robust learning. However, existing label correction methods can not handle with heavy noise or datasets with samples of many categories so well. We explain the reasons and introduce a global label distribution regularization to remedy these deficiencies. With this regularization, we convert the label correction to the Optimal Transport (OT) formulation and propose to utilize a fast version of the Sinkhorn-Knopp algorithm for finding an approximate solution efficiently at scale. Experiments on benchmark datasets with both synthetic and real-world label noise show that the superiority of our OT Cleaner in terms of both training efficiency and classification accuracy. The code is available at: https://github.com/junxia97/OT-Cleaner.
Jun Xia 0001, Cheng Tan 0012, Lirong Wu, Yongjie Xu 0001, Stan Z. Li
ICASSP4
2022 Deep manifold embedding of attributed graphs
Zelin Zang, Siyuan Li 0002, Di Wu 0057, Jianzhu Guo, Yongjie Xu 0001, Stan Z. Li
Neurocomputing5
2021 Ship Classification in SAR Images With Geometric Transfer Metric Learning
abstract
There are still many challenges to be resolved in the task of ship classification in synthetic aperture radar (SAR) images, such as limited number of labeled samples in SAR domain, large variance in the same subcategory, small variance among different subcategories, etc. Transfer metric learning (TML) has the potential to mitigate those issues in the domain of interest (target domain, TD) by leveraging knowledge/information from other related domains (source domain, SD). In this article, we proposed a novel TML method, termed as geometric transfer metric learning (GTML), which achieves discriminative information preservation (DIP), geometric structure preservation (GSP), and handles the domain shift (DS) simultaneously by integrating pairwise constraints (PC), joint distribution adaptation (JDA), and manifold regularization (MR) into a unified optimization function, aiming to make full use of their complementarity to improve SAR ship classification performance. In practice, we proposed two simple but effective optimization strategies, termed as GTML-A and GTML-R, to construct optimization function. We also proposed two solutions for two typical real-world application scenarios, that is, the task of: 1) zero-labeled sample (ZLS) and 2) scarce-labeled samples (SLS) in SAR domain. The experiments conducted on both tasks show that the proposed GTML outperforms most of state-of-the-art methods. Code is available athttps://github.com/sky-Yongjie-Xu/geometric-transfer-metric-learning.
Yongjie Xu 0001, Haitao Lang
IEEE Trans. Geosci. Remote. Sens.1
2019 Distribution Discrepancy Maximization Metric Learning for Ship Classification in Synthetic Aperture Radar Images
abstract
Supervised learning techniques are widely used in the task of ship classification in synthetic aperture radar (SAR) images in recent years. Learning distance metrics that describe the underlying distribution between data points based on the distance metric learning (DML) methods can further improve the performance of ship classification in SAR images. Traditional supervised DML methods usually learn distance metrics based on pairwise constraints, but ignore the importance of inter-class distribution discrepancy. In this study, we propose a novel DML method named distribution discrepancy maximization metric learning (DDMML) algorithm, which maximizes the maximum mean discrepancy (MMD) between different categories in the process of learning distance metrics. We adopt a high-resolution SAR ship database for experimental evaluation. The experimental results show that the proposed method outperforms the state-of-the-art DML methods.
Yongjie Xu 0001, Haitao Lang, Xiaopeng Chai
IGARSS1
2019 Discriminative Adaptation Regularization Framework-Based Transfer Learning for Ship Classification in SAR Images
abstract
Ship classification in synthetic-aperture radar (SAR) images is of great significance for dealing with various marine matters. Although traditional supervised learning methods have recently achieved dramatic successes, but they are limited by the insufficient labeled training data. This letter presents a novel unsupervised domain adaptation (DA) method, termed as discriminative adaptation regularization framework-based transfer learning (D-ARTL), to address the problem in case that there is no labeled training data available at all in the SAR image domain, i.e., target domain (TD). D-ARTL improves the original ARTL by adding a novel source discriminative information preservation (SDIP) regularization term. This improvement achieves an efficient transfer of interclass discriminative ability from source domain (SD) to TD, while achieving the alignment of cross-domain distributions. Extensive experiments have verified that D-ARTL outperforms state-of-the-art methods on the task of ship classification in SAR images by transferring the automatic identification system (AIS) information.
Yongjie Xu 0001, Haitao Lang, Lihui Niu, Chenguang Ge
IEEE Geosci. Remote. Sens. Lett.1
2018 Ship Classification in SAR Images Improved by AIS Knowledge Transfer
abstract
A major bottleneck in limiting the application of the existing methods of ship classification in synthetic aperture radar (SAR) images is the inadequate amount of labeled data available for training a classifier. However, generating ground truth involves expensive and time-consuming ground campaigns or is costly, since a high number of SAR image acquisition will be necessary. In contrast, an automatic identification system (AIS), which is an automatic tracking system used for monitoring maritime ships, can provide plenty of labeled ship samples that is relatively easier to be obtained. Inspired by these facts, this letter proposes to improve ship classification in SAR images by transferring AIS knowledge. We propose an improved multiclass adaptive support vector machine, combined with the naive geometric features (NGFs), to achieve transfer learning between the AIS domain and the SAR image domain. The experiments prove that the traditional method can be significantly improved by AIS information transfer, especially when only a few training samples in the SAR domain are available. In addition, it also shows that after feature selection, the performance of the proposed method can be close to that of the state of the art, even if by only using simpler NGFs and few training samples.
Haitao Lang, Siwen Wu, Yongjie Xu 0001
IEEE Geosci. Remote. Sens. Lett.3