Longyi Li

dblp:04/7650 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Annotating Spatial Multi-Omics Spot-Level Niche Types Using Bi-View Retrieval-Augmented Generation With SpotTypeLLM
abstract
Spatial multi-omics techniques generate extensive spot-level profiles without accompanying spot-type labels, forcing biologists into labor-intensive manual annotation. Although large language models (LLMs) promise automated annotation, they are poorly equipped to handle high-dimensional numeric inputs, struggle to convert complex spatial-omics structures into interpretable text, and lack the specialized biological knowledge needed to avoid hallucinations. Moreover, the intrinsic sparsity and heterogeneity of spatial omics data undermine robust feature extraction and accurate spot-level niche label assignment. To address these challenges, we propose SpotTypeLLM, a framework for annotating spatial multi-omics spot-level niche types using a Bi-view Retrieval-Augmented Generation (BiRAG) tailored for LLMs. Specifically, SpotTypeLLM encodes spatial multi-omics data through scLLM-based embeddings, graph convolutional networks (GCNs), and multi-view attention module, and then selects spot-specific genes to construct the prompts for LLMs. We validate the accuracy and practical utility of our framework by comparing it with ten state-of-the-art methods on labeled spatial multi-omics datasets. The results demonstrate the effectiveness of SpotTypeLLM in overcoming the limitations of current annotation methods and enhancing spatial multi-omics analysis.
Longyi Li, Liyan Dong, Decheng Li, Hanbo Liu, Yuheng Zhu, Xiyuan Mei, Hao Zhang 0064, Dong Xu 0002
IEEE Trans. Comput. Biol. Bioinform.1
2026 GALA: Integrating Weighted Graph Walks and Latent-Space Adversarial Training for Single-Cell Batch Alignment
abstract
Single-cell RNA sequencing (scRNA-seq) enables unprecedented exploration of cellular heterogeneity, yet technical variations across datasets introduce pervasive batch effects that severely compromise integrative analysis. We present Graph-based Adversarial Latent Alignment (GALA), a novel batch correction framework that synergistically integrates weighted graph random walks with latent space adversarial training to robustly align scRNA-seq data while preserving critical biological signals. GALA employs a Weighted Graph Mutual Nearest Neighbor (WGMNN) module that achieves up to 125% improvement in cross-batch cell pairing diversity and 48% increase in coverage compared to conventional approaches, substantially enhancing the detection of biologically meaningful correspondences between batches. These optimized cell pairs guide adversarial training within a carefully designed low-dimensional latent space, generating batch-agnostic representations that simultaneously eliminate technical artifacts while faithfully preserving biological variability. Evaluated across five benchmark datasets representing diverse real-world scenarios-including identical cell types, non-identical cell types, and complex multi-batch settings-GALA demonstrates robust performance in batch correction and biological signal preservation, achieving high F1 scores of 0.91, 0.82, 0.88, 0.82, and 0.76, respectively. GALA outperforms established methods including Seurat v4, Harmony, and Scanorama, with notable advantages in challenging datasets with heterogeneous cell populations. Aggregating performance across all datasets, GALA achieves the highest overall integration score and best average rank while maintaining high computational efficiency, demonstrating consistent superiority and strong robustness across diverse scRNA-seq integration scenarios.
Liyan Dong, Longyi Li, Hao Zhang 0064
IEEE Trans. Comput. Biol. Bioinform.3
2026 Snapshot Compressive Imaging via Degradation Cue and Spectral Latent Diffusion
abstract
The goal of snapshot spectral compressive imaging reconstruction is to recover the 3D hyperspectral image from a 2D measurement. However, current reconstruction methods still face significant challenges in fully leveraging degradation and image prior. Many methods estimate degradation solely from a single measurement rather than learning from the real imaging process, resulting in inaccurate prior modeling. Moreover, the high compression of the CASSI measurement leads to the loss of spectral-spatial context, and the existing priors fail to fully capture it - for instance, in complex scenarios (such as S5, S9 in Table I), the performance gap can be as high as 3 dB. To address these issues, this paper introduces a novel reconstruction method with Degradation Cue Learning and Spectral Latent Diffusion (DCL-SLD), which comprises two key components: the Degradation Cue Learning (DCL) module and the Spectral Latent Diffusion (SLD) module. In the spatial domain, the DCL module employs a pre-trained image encoder and a feature distribution transmission strategy to extract degraded information and integrate it into the feature, enabling reconstruction through learned visual context. In the spectral domain, the SLD module leverages a latent diffusion model based on spectral correlations to generate a low-rank vector representation, effectively preserving contextual relationships within the high-dimensional structure. By enhancing priors in both dimensions, the model significantly improves its ability to exploit contextual information for more accurate recovery. Extensive experimental results on both simulation and real datasets demonstrate the superior performance of DCL-SLD over state-of-the-art methods.
Mingjin Zhang, Longyi Li, Jie Guo 0009, Yunsong Li 0001
IEEE Trans. Image Process.2
2025 Tokenizing RNA Structures: A Discrete Generative Approach to Represent the RNA Folding Landscapes
abstract
RNA's diverse biological functions, including gene regulation, enzymatic catalysis, signal transduction, and even therapeutic activity, ultimately depend on its three-dimensional (3D) structure. Departing from the classical task of inferring structure directly from sequence, this study investigates how to efficiently encode and generate RNA 3D structures, thereby enabling systematic exploration of the molecules' vast folding landscape. We propose a generative framework, RNA-FrameEncoder (RNA-FE). First, a Vector-Quantised Variational AutoEncoder compresses continuous RNA backbones into sequences of discrete structural tokens, “verbalizing” the geometry while suppressing high-frequency noise. Next, a Transformer language model is trained on these token sequences to capture their joint distribution, and a decoder subsequently projects the tokens back into Cartesian coordinates. This design sidesteps the complexity of explicit equivariance modelling, lowers computational cost, and markedly enhances the representational capacity. Comprehensive benchmarks show that RNA-FE can unconditionally generate realistic, novel, and designable RNA structures without any sequence or structural priors. Overall, RNA-FE provides a scalable approach for efficiently charting RNA 3D structural space and opens a data-efficient avenue for a range of downstream applications. The code and pretrained models are publicly available at: https://github.com/users/Hanbo24-bit/RNA-FE.
Hanbo Liu, Hao Zhang 0064, Longyi Li, Xiyuan Mei, Enshuang Zhao, Yuheng Zhu, Yinfei Dai, Jingxun Cao, Dong Xu 0002
BIBM3
2025 RBTN: A Spatio Temporal Graph Neural Network for Integrating Single Cell and Spatial Transcriptomics During Magnaporthe Oryzae Infection in Rice
abstract
Rice blast, caused by the fungal pathogen Magnaporthe oryzae (M. oryzae), remains one of the most devastating threats to global rice production, capable of reducing yields by up to 30%. Current integration methods for single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) treat tissues as static snapshots, overlooking the temporal dynamics crucial for host-pathogen analysis. Here, we introduce Rice Blast Temporal Graph Neural Network (RBTN), the first computational framework specifically designed to integrate scRNA-seq and ST data across the three-dimensional spot-cell-time manifold during rice blast infection. RBTN constructs a unified graph whose nodes represent either individual cell transcriptomes or spatial spots sampled at three infection time points (0 h,$\text{1 2 h}, \text{2 4 h}$). It learns stage-specific mappings via bidirectional cosine loss and enforces temporal smoothness through a cosinebased consistency term. Alternating self-attention and temporal-attention layers capture local spatial organization and cross-time transitions. When applied to a public rice blast dataset, RBTN consistently delivers high-fidelity spatial reconstructions across different sets of highly variable marker genes (HVGs) and produces interpretable embeddings that jointly encode spatial, cellular, and temporal relationships. RBTN therefore establishes a robust, temporally-aware computational platform for elucidating the spatiotemporal dynamics of host-pathogen interactions during rice blast infection. All code in this paper is available at https://github.com/zhuyuheng111/RBTN.
Yuheng Zhu, Enshuang Zhao, Yinfei Dai, Hanbo Liu, Longyi Li, Qiyuan Mei, Hao Zhang 0064
BIBM6
2025 Multimodal Prior Learning with Double Constraint Alignment for Snapshot Spectral Compressive Imaging
abstract
The objective of snapshot spectral compressive imaging reconstruction is to recover the 3D hyperspectral image (HSI) from a 2D measurement. Existing methods either focus on network architecture design or simply introduce image-level prior to the model. However, these methods lack guiding information for accurate reconstruction. Recognizing that textual description contain rich semantic information that can significantly enhance details, this paper introduces a novel framework, CAMM, which integrates text information into the model to improve the performance. The framework comprises two key components: Fine-grained Alignment Module (FAM) and Multimodal Fusion Mamba (MFM). Specifically, FAM is used to reduce the knowledge gap between the RGB domain obtained by the pre-trained vision-language model and the HSI domain. Through the double constraints of distribution similarity and entropy, the adaptive alignment of different complexity features is realized, which makes the encoded features more accurate. MFM aims to identify the guiding effect of RGB features and text features on HSI in space and channel dimensions. Instead of fusing features directly, it integrates prior at image-level and text-level prior into Mamba's state-space equation, so that each scanning step can be accurately guided. This kind of positive feedback adjustment ensures the authenticity of the guiding information. To our knowledge, this is the first text-guided model for compressive spectral imaging. Extensive experimental results the public datasets demonstrate the superior performance of CAMM, validating the effectiveness of our proposed method.
Mingjin Zhang, Longyi Li, Fei Gao 0006, Qiming Zhang 0001, Jie Guo 0009
IJCAI2
2025 spaLLM: enhancing spatial domain analysis in multi-omics data through large language model integration
abstract
Spatial multi-omics technologies provide valuable data on gene expression from various omics in the same tissue section while preserving spatial information. However, deciphering spatial domains within spatial omics data remains challenging due to the sparse gene expression. We propose spaLLM, the first multi-omics spatial domain analysis method that integrates large language models to enhance data representation. Our method combines a pre-trained single-cell language model (scGPT) with graph neural networks and multi-view attention mechanisms to compensate for limited gene expression information in spatial omics while improving sensitivity and resolution within modalities. SpaLLM processes multiple spatial modalities, including RNA, chromatin, and protein data, potentially adapting to emerging technologies and accommodating additional modalities. Benchmarking against eight state-of-the-art methods across four different datasets and platforms demonstrates that our model consistently outperforms other advanced methods across multiple supervised evaluation metrics. The source code for spaLLM is freely available at https://github.com/liiilongyi/spaLLM.
Longyi Li, Liyan Dong, Hao Zhang 0064, Dong Xu 0002
Briefings Bioinform.1
2025 TVAE-RNA: ensemble-based RNA secondary structure prediction via transformer variational autoencoders
abstract
MOTIVATION: Accurate prediction of RNA secondary structure remains challenging due to the presence of pseudoknots, long-range dependencies, and limited labeled data. RESULTS: We propose TVAE, a novel framework that integrates a Transformer encoder with a Variational Autoencoder (VAE). The Transformer captures global dependencies in the sequence, while the VAE models structural variability by learning a probabilistic latent space. Unlike deterministic models, TVAE generates diverse and biologically plausible secondary structures, enabling more comprehensive structure discovery. To obtain discrete predictions, we introduce GHA-Pairing, a fast and biologically constrained base-pairing algorithm. TVAE demonstrates strong generalization across different RNA families and achieves state-of-the-art performance on benchmark datasets, reaching an F1 score of 0.89 and 83% accuracy, surpassing existing methods by 10%. These results highlight the advantage of probabilistic modeling for RNA structure prediction and its potential to enhance biological insights. AVAILABILITY AND IMPLEMENTATION: Code and pretrained models are available at https://github.com/mei-rna/TVAE-RNA. The released version of the dataset and models can also be accessed via DOI: 10.5281/zenodo.16946114.
Xiyuan Mei, Hanbo Liu, Yuheng Zhu, Enshuang Zhao, Longyi Li, Hao Zhang 0064
Bioinform.5
2025 TFS-Net: Temporal first simulation network for video saliency prediction
Longyi Li, Liyan Dong, Hao Zhang 0064, Zhengtai Zhang, Minghui Sun 0001
Expert Syst. Appl.1
2024 VmambaSCI: Dynamic Deep Unfolding Network with Mamba for Compressive Spectral Imaging
abstract
Snapshot spectral compressive imaging can capture spectral information across multiple wavelengths in one imaging. The coded aperture snapshot spectral imaging (CASSI) method, aims to recover 3D spectral cubes from 2D measurements. Most existing approaches employ a deep unfolding framework based on Transformer, which alternately address a data subproblem and a prior subproblem. However, these frameworks lack flexibility regarding the sensing matrix and inter-stage interactions. In addition, the quadratic computational complexity of global Transformer and the restricted receptive field of local Transformer impact reconstruction efficiency and accuracy. In this paper, we propose a dynamic deep unfolding network with mamba for compressive spectral imaging, called VmambaSCI. We integrate spatial-spectral information from the sensing matrix into the data module and utilizes spatial adaptive operations in the stage interaction of the prior module. Furthermore, recognizing that the imaging process causes aliasing of spatial and spectral information, we develop a dual-domain scanning mamba (DSMamba), featuring a novel spatial-channel scanning method for enhanced efficiency and accuracy. To our knowledge, VmambaSCI is the first Mamba-based model for compressive spectral imaging. Experimental results on the public databases, CAVE and KAIST, demonstrate the superiority of the proposed VmambaSCI over the state-of-the-art approaches.
Mingjin Zhang, Longyi Li, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001
ACM Multimedia2