Zhixiang Ren

dblp:21/8692 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-4104-3790ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 The limits of bio-molecular modeling with large language models: a cross-scale evaluation
abstract
MOTIVATION: The modeling of bio-molecular system across molecular scales remains a central challenge in scientific research. Large language models (LLMs) are increasingly applied to bio-molecular discovery, yet systematic evaluation across multi-scale biological problems and rigorous assessment of their tool-augmented capabilities remain limited. RESULTS: We reveal a systematic gap between LLM performance and mechanistic understanding through the proposed cross-scale bio-molecular benchmark: BioMol-LLM-Bench, a unified framework comprising 26 downstream tasks that covers 4 distinct difficulty levels, and computational tools are integrated for a more comprehensive evaluation. Evaluation on 13 representative models reveals 4 benchmark-specific observations: chain-of-thought-style training does not consistently improve performance on the evaluated biological tasks; the evaluated hybrid mamba-attention model shows strong performance on long bio-molecular sequence tasks; supervised fine-tuned models show task-specific specialization with reduced performance in some general settings; and current LLMs perform better on classification tasks than on challenging regression tasks under this benchmark setting. AVAILABILITY: Source code is available at https://github.com/AI-HPC-Research-Team/BioMol-LLM-Bench.
Yaxin Xu, Yue Zhou 0014, Zhengyu Ma, Fengwei An, Zhixiang Ren
Bioinform.6
2026 Concept-Driven Deep Learning for Enhanced Protein-Specific Molecular Generation
abstract
In recent years, deep learning techniques have made significant strides in molecular generation for specific targets, driving advancements in drug discovery. However, existing molecular generation methods present significant limitations: those operating at the atomic level often lack synthetic feasibility, drug-likeness, and interpretability, while fragment-based approaches frequently overlook comprehensive factors that influence protein–molecule interactions. To address these challenges, we propose a novel fragment-based molecular generation framework tailored for specific proteins. Our method begins by constructing a protein subpocket and molecular arm concept-based neural network, which systematically integrates interaction force information and geometric complementarity to sample molecular arms for specific protein subpockets. Subsequently, we introduce a diffusion model to generate molecular backbones that connect these arms, ensuring structural integrity and chemical diversity. Our approach improves synthetic feasibility and binding affinity, with a 4% increase in drug-likeness and a 6% improvement in synthetic feasibility. Furthermore, by integrating explicit interaction data through a concept-based model, our framework enhances interpretability, offering valuable insights into the molecular design process.
Taojie Kuang, Qianli Ma 0001, Athanasios V. Vasilakos, Yu Wang 0008, Qiang Shawn Cheng, Zhixiang Ren
ACM Trans. Knowl. Discov. Data6
2025 A multi-modal genomic knowledge distillation framework for drug response prediction
Shuang Ge, Shuqing Sun, Qiang Shawn Cheng, Zhixiang Ren
Appl. Intell.5
2025 Deep learning in single-cell and spatial transcriptomics data analysis: advances and challenges from a data science perspective
abstract
The development of single-cell and spatial transcriptomics has revolutionized our capacity to investigate cellular properties, functions, and interactions in both cellular and spatial contexts. Despite this progress, the analysis of single-cell and spatial omics data remains challenging. First, single-cell sequencing data are high-dimensional and sparse, and are often contaminated by noise and uncertainty, obscuring the underlying biological signal. Second, these data often encompass multiple modalities, including gene expression, epigenetic modifications, metabolite levels, and spatial locations. Integrating these diverse data modalities is crucial for enhancing prediction accuracy and biological interpretability. Third, while the scale of single-cell sequencing has expanded to millions of cells, high-quality annotated datasets are still limited. Fourth, the complex correlations of biological tissues make it difficult to accurately reconstruct cellular states and spatial contexts. Traditional feature engineering approaches struggle with the complexity of biological networks, while deep learning, with its ability to handle high-dimensional data and automatically identify meaningful patterns, has shown great promise in overcoming these challenges. Besides systematically reviewing the strengths and weaknesses of advanced deep learning methods, we have curated 21 datasets from nine benchmarks to evaluate the performance of 58 computational methods. Our analysis reveals that model performance can vary significantly across different benchmark datasets and evaluation metrics, providing a useful perspective for selecting the most appropriate approach based on a specific application scenario. We highlight three key areas for future development, offering valuable insights into how deep learning can be effectively applied to transcriptomic data analysis in biological, medical, and clinical settings.
Shuang Ge, Shuqing Sun, Qiang Shawn Cheng, Zhixiang Ren
Briefings Bioinform.5
2025 Generative prediction of real-world prevalent SARS-CoV-2 mutation with in silico virus evolution
abstract
Predicting the mutation prevalence trends of emerging viruses in the real world is an efficient means to update vaccines or drugs in advance. It is crucial to develop a computational method for the prediction of real-world prevalent SARS-CoV-2 mutations considering the impact of multiple selective pressures within and between hosts. Here, a deep-learning generative framework for real-world prevalent SARS-CoV-2 mutation prediction, named ViralForesight, is developed on top of protein language models and in silico virus evolution. Through the paradigm of host-to-herd in silico virus evolution, ViralForesight reproduced previous real-world prevalent SARS-CoV-2 mutations for multiple lineages with superior performance. More importantly, ViralForesight correctly predicted the future prevalent mutations that dominated the COVID-19 pandemic in the real world more than half a year in advance with in vitro experimental validation. Overall, ViralForesight demonstrates a proactive approach to the prevention of emerging viral infections, accelerating the process of discovering future prevalent mutations with the power of generative deep learning.
Xudong Liu 0001, Zhiwei Nie, Haorui Si, Xurui Shen, Yutian Liu 0004, Xiansong Huang, Tianyi Dong, Zhixiang Ren, Jie Chen 0001
Briefings Bioinform.9
2025 Predicting protein stability changes upon mutations with dual-view ensemble learning from single sequence
abstract
Predicting the protein stability changes upon mutations is one of the effective ways to improve the efficiency of protein engineering. Here, we propose a dual-view ensemble learning-based framework, DVE-stability, for mutation-induced protein stability change prediction from single sequence. DVE-stability integrates the global and local dependencies of mutations to capture the intramolecular interactions from two views through ensemble learning, in which a structural microenvironment simulation module is designed to indirectly introduce the information of structural microenvironment at the sequence level. DVE-stability achieved state-of-the-art prediction performance on seven single-point mutation benchmark datasets, and comprehensively surpassed other methods on five of them. Furthermore, DVE-stability outperformed other methods comprehensively through zero-shot inference on multiple-point mutation prediction task, demonstrating superior model generalizability to capture the epistasis of multiple-point mutations. More importantly, DVE-stability exhibited superior generalization performance in predicting rare beneficial mutations that are crucial for practical protein directed evolution scenarios. In addition, DVE-stability identified important intramolecular interactions via attention scores, demonstrating interpretable. Overall, DVE-stability provides a flexible and efficient tool for mutation-induced protein stability change prediction in an interpretable ensemble learning manner.
Zhiwei Nie, Yutian Liu 0004, Xiansong Huang, Peng Yang 0001, Zigang Li, Jie Fu 0001, Zhixiang Ren, Jie Chen 0001
Briefings Bioinform.11
2025 A self-feedback knowledge elicitation approach for chemical reaction predictions
Jun Tao 0002, Zhixiang Ren
Eng. Appl. Artif. Intell.3
2025 Deep learning methods for protein representation and function prediction: A comprehensive overview
Mingqing Wang, Zhiwei Nie, Yonghong He, Athanasios V. Vasilakos, Qiang Shawn Cheng, Zhixiang Ren
Eng. Appl. Artif. Intell.6
2025 Aligning sequence and structure representations leveraging protein domains for function prediction
Mingqing Wang, Zhiwei Nie, Yonghong He, Athanasios V. Vasilakos, Zhixiang Ren
Expert Syst. Appl.5
2025 Few-Shot Generalization to Novel Compounds in Single-Cell Drug Response via Graph-Infused Meta-Pretraining
abstract
Understanding drug responses at the single-cell level is crucial for identifying biomarkers and uncovering resistance mechanisms. However, existing models predominantly rely on genomic profiles, while overlooking drug structure-function relationships and showing limited generalization to novel drugs with distinct structures. To address this limitation, we propose a novel framework that integrates drug structural information with genomic data. Specifically, we develop a graph-aware Transformer to capture interatomic relations and generate joint representations linking atomic features to genomic profiles. To overcome the scarcity of single-cell drug response data, we propose a novel predictive framework that leverages prior knowledge from bulk RNA datasets through meta-pretraining and few-shot transfer learning. Furthermore, we introduce a position-based feature extraction network and a gene gradient attribution algorithm to identify key resistance genes and drug action pathways. Pre-trained on 223 drugs across 14 tissues and tested on seven single-cell datasets, our model achieves an approximate 5% improvement in accuracy for known drugs and about 20% increase in generalization to unseen drugs. This approach provides an effective method for studying drug resistance mechanisms at single-cell level, particularly for novel compounds.
Shuang Ge, Qiang Shawn Cheng, Shuqing Sun, Zhixiang Ren
IEEE Trans. Comput. Biol. Bioinform.6
2025 Discovery of Bioactive Constituents for Colitis From Traditional Chinese Medicine Prescription via Deep Neural Network
abstract
Colitis is a commonly encountered inflammatory disease in colon tissue, which can be triggered by various causes. Although a few ingredients in traditional Chinese medicine (TCM) have been identified as effective for the treatment of colitis, it remains a great challenge to discover the potential therapeutic bioactive constituents and their modes of action among thousands of ingredients in TCM prescriptions. To address this issue, we propose a pipeline that combines deep neural network (DNN) with network pharmacology to discover bioactive constituents. By integrating the herbal information network of 9,845 nodes and 161,950 edges, which includes detailed information on bioactive molecules and protein targets, with a prescription list expanded through a novel data augmentation strategy, the DNN can recommend diverse herbal combinations. Network pharmacology study revealed that the 10 most frequent constituents in recommended prescriptions were associated with multiple inflammatory signaling pathways. To verify the bioactive constituents in the recommended prescriptions, 5 selected constituents were administrated to BALB/c mice with colitis. Suppressive effects of disease progression and pro-inflammatory factors comparable to sulfasalazin were observed with these compounds, revealing the effectiveness of our artificial intelligence strategy in discovering bioactive constituents from TCM prescriptions.
Zhixiang Ren, Qi Shu, Huijuan Ma
IEEE Trans. Comput. Biol. Bioinform.1
2024 Self-supervised learning on millions of primary RNA sequences from 72 vertebrates improves sequence-based RNA splicing prediction
abstract
Language models pretrained by self-supervised learning (SSL) have been widely utilized to study protein sequences, while few models were developed for genomic sequences and were limited to single species. Due to the lack of genomes from different species, these models cannot effectively leverage evolutionary information. In this study, we have developed SpliceBERT, a language model pretrained on primary ribonucleic acids (RNA) sequences from 72 vertebrates by masked language modeling, and applied it to sequence-based modeling of RNA splicing. Pretraining SpliceBERT on diverse species enables effective identification of evolutionarily conserved elements. Meanwhile, the learned hidden states and attention weights can characterize the biological properties of splice sites. As a result, SpliceBERT was shown effective on several downstream tasks: zero-shot prediction of variant effects on splicing, prediction of branchpoints in humans, and cross-species prediction of splice sites. Our study highlighted the importance of pretraining genomic language models on a diverse range of species and suggested that SSL is a promising approach to enhance our understanding of the regulatory logic underlying genomic sequences.
Ken Chen 0006, Yue Zhou 0014, Maolin Ding, Yu Wang 0008, Zhixiang Ren, Yuedong Yang
Briefings Bioinform.5
2024 CELA-MFP: a contrast-enhanced and label-adaptive framework for multi-functional therapeutic peptides prediction
abstract
Functional peptides play crucial roles in various biological processes and hold significant potential in many fields such as drug discovery and biotechnology. Accurately predicting the functions of peptides is essential for understanding their diverse effects and designing peptide-based therapeutics. Here, we propose CELA-MFP, a deep learning framework that incorporates feature Contrastive Enhancement and Label Adaptation for predicting Multi-Functional therapeutic Peptides. CELA-MFP utilizes a protein language model (pLM) to extract features from peptide sequences, which are then fed into a Transformer decoder for function prediction, effectively modeling correlations between different functions. To enhance the representation of each peptide sequence, contrastive learning is employed during training. Experimental results demonstrate that CELA-MFP outperforms state-of-the-art methods on most evaluation metrics for two widely used datasets, MFBP and MFTP. The interpretability of CELA-MFP is demonstrated by visualizing attention patterns in pLM and Transformer decoder. Finally, a user-friendly online server for predicting multi-functional peptides is established as the implementation of the proposed CELA-MFP and can be freely accessed at http://dreamai.cmii.online/CELA-MFP.
Yitian Fang, Mingshuang Luo, Zhixiang Ren, Leyi Wei
Briefings Bioinform.3
2024 3D-Mol: A Novel Contrastive Learning Framework for Molecular Property Prediction with 3D Information
Taojie Kuang, Zhixiang Ren
Pattern Anal. Appl.3
2023 DeepProSite: structure-aware protein binding site prediction using ESMFold and pretrained language model
abstract
MOTIVATION: Identifying the functional sites of a protein, such as the binding sites of proteins, peptides, or other biological components, is crucial for understanding related biological processes and drug design. However, existing sequence-based methods have limited predictive accuracy, as they only consider sequence-adjacent contextual features and lack structural information. RESULTS: In this study, DeepProSite is presented as a new framework for identifying protein binding site that utilizes protein structure and sequence information. DeepProSite first generates protein structures from ESMFold and sequence representations from pretrained language models. It then uses Graph Transformer and formulates binding site predictions as graph node classifications. In predicting protein-protein/peptide binding sites, DeepProSite outperforms state-of-the-art sequence- and structure-based methods on most metrics. Moreover, DeepProSite maintains its performance when predicting unbound structures, in contrast to competing structure-based prediction methods. DeepProSite is also extended to the prediction of binding sites for nucleic acids and other ligands, verifying its generalization capability. Finally, an online server for predicting multiple types of residue is established as the implementation of the proposed DeepProSite. AVAILABILITY AND IMPLEMENTATION: The datasets and source codes can be accessed at https://github.com/WeiLab-Biology/DeepProSite. The proposed DeepProSite can be accessed at https://inner.wei-group.net/DeepProSite/.
Yitian Fang, Leyi Wei, Qin Ma 0003, Zhixiang Ren, Qianmu Yuan
Bioinform.5
2014 Region-Based Saliency Detection and Its Application in Object Recognition
abstract
The objective of this paper is twofold. First, we introduce an effective region-based solution for saliency detection. Then, we apply the achieved saliency map to better encode the image features for solving object recognition task. To find the perceptually and semantically meaningful salient regions, we extract superpixels based on an adaptive mean shift algorithm as the basic elements for saliency detection. The saliency of each superpixel is measured by using its spatial compactness, which is calculated according to the results of Gaussian mixture model (GMM) clustering. To propagate saliency between similar clusters, we adopt a modified PageRank algorithm to refine the saliency map. Our method not only improves saliency detection through large salient region detection and noise tolerance in messy background, but also generates saliency maps with a well-defined object shape. Experimental results demonstrate the effectiveness of our method. Since the objects usually correspond to salient regions, and these regions usually play more important roles for object recognition than background, we apply our achieved saliency map for object recognition by incorporating a saliency map into sparse coding-based spatial pyramid matching (ScSPM) image representation. To learn a more discriminative codebook and better encode the features corresponding to the patches of the objects, we propose a weighted sparse coding for feature coding. Moreover, we also propose a saliency weighted max pooling to further emphasize the importance of those salient regions in feature pooling module. Experimental results on several datasets illustrate that our weighted ScSPM framework greatly outperforms ScSPM framework, and achieves excellent performance for object recognition.
Zhixiang Ren, Shenghua Gao, Liang-Tien Chia, Ivor W. Tsang
IEEE Trans. Circuits Syst. Video Technol.1
2014 Concurrent Single-Label Image Classification and Annotation via Efficient Multi-Layer Group Sparse Coding
abstract
We present a multi-layer group sparse coding framework for concurrent single-label image classification and annotation. By leveraging the dependency between image class label and tags, we introduce a multi-layer group sparse structure of the reconstruction coefficients. Such structure fully encodes the mutual dependency between the class label, which describes image content as a whole, and tags, which describe the components of the image content. Therefore we propose a multi-layer group based tag propagation method, which combines the class label and subgroups of instances with similar tag distribution to annotate test images. To make our model more suitable for nonlinear separable features, we also extend our multi-layer group sparse coding in the Reproducing Kernel Hilbert Space (RKHS), which further improves performances of image classification and annotation. Moreover, we also integrate our multi-layer group sparse coding with kNN strategy, which greatly improves the computational efficiency. Experimental results on the LabelMe, UIUC-Sports and NUS-WIDE-Object databases show that our method outperforms the baseline methods, and achieves excellent performances in both image classification and annotation tasks.
Shenghua Gao, Liang-Tien Chia, Ivor W. Tsang, Zhixiang Ren
IEEE Trans. Multim.4
2013 Background subtraction via coherent trajectory decomposition
abstract
Background subtraction, the task to detect moving objects in a scene, is an important step in video analysis. In this paper, we propose an efficient background subtraction method based on coherent trajectory decomposition. We assume that the trajectories from background lie in a low-rank subspace, and foreground trajectories are sparse outliers in this background subspace. Meanwhile, the Markov Random Field (MRF) is used to encode the spatial coherency and trajectory consistency. With the low-rank decomposition and the MRF, our method can better handle videos with moving camera and obtain coherent foreground. Experimental results on a video dataset show our method achieves very competitive performance.
Zhixiang Ren, Liang-Tien Chia, Deepu Rajan, Shenghua Gao
ACM Multimedia1
2013 Regularized Feature Reconstruction for Spatio-Temporal Saliency Detection
abstract
Multimedia applications such as image or video retrieval, copy detection, and so forth can benefit from saliency detection, which is essentially a method to identify areas in images and videos that capture the attention of the human visual system. In this paper, we propose a new spatio-temporal saliency detection framework on the basis of regularized feature reconstruction. Specifically, for video saliency detection, both the temporal and spatial saliency detection are considered. For temporal saliency, we model the movement of the target patch as a reconstruction process using the patches in neighboring frames. A Laplacian smoothing term is introduced to model the coherent motion trajectories. With psychological findings that abrupt stimulus could cause a rapid and involuntary deployment of attention, our temporal model combines the reconstruction error, regularizer, and local trajectory contrast to measure the temporal saliency. For spatial saliency, a similar sparse reconstruction process is adopted to capture the regions with high center-surround contrast. Finally, the temporal saliency and spatial saliency are combined together to favor salient regions with high confidence for video saliency detection. We also apply the spatial saliency part of the spatio-temporal model to image saliency detection. Experimental results on a human fixation video dataset and an image saliency detection dataset show that our method achieves the best performance over several state-of-the-art approaches.
Zhixiang Ren, Shenghua Gao, Liang-Tien Chia, Deepu Rajan
IEEE Trans. Image Process.1
2012 Spatiotemporal Saliency Detection via Sparse Representation
abstract
Multimedia applications like retrieval, copy detection etc. can gain from saliency detection, which is essentially a method to identify areas in images and videos that capture the attention of the human visual system. In this paper, we propose a new spatiotemporal saliency framework for videos based on sparse representation. For temporal saliency, we model the movement of the target patch as a reconstruction process, and the overlapping patches in neighboring frames are used to reconstruct the target patch. The learned coefficients encode the positions of the matched patches, which are able to represent the motion trajectory of the target patch. We also introduce a smoothing term into our sparse coding framework to learn coherent motion trajectories. Based on the psychological findings that abrupt stimulus could cause a rapid and involuntary deployment of attention, our temporal model combines the reconstruction error, sparsity regularizer, and local trajectory contrast to measure the motion saliency. For spatial saliency, a similar sparse reconstruction process is adopted to capture the regions with high center-surround contrast. Finally, the temporal saliency and spatial saliency are combined by agreement to favor the salient regions with high confidence. Experimental results on a human fixation video dataset show our method achieved the best performance over five state-of-the-art approaches.
Zhixiang Ren, Shenghua Gao, Deepu Rajan, Liang-Tien Chia
ICME1
2012 Salient Object Detection through Over-Segmentation
abstract
In this paper we present a salient object detection model from an over-segmented image. The input image is initially segmented by the mean-shift segmentation algorithm and then over-segmented by a quad mesh to even smaller segments. Such segmented regions overcome the disadvantage of using patches or single pixels to compute saliency. Segments that are similar and spread over the image receive low saliency and a segment which is distinct in the whole image or in a local region receives high saliency. We express this as a color compactness measure which is used to derive saliency level directly. Our method is shown to outperform six existing methods in the literature using a saliency detection database containing images with human-labeled object contour ground truth. The proposed saliency model has been shown to be useful for an image retargeting application.
Zhixiang Ren, Deepu Rajan, Yiqun Hu
ICME2
2012 Video saliency detection with robust temporal alignment and local-global spatial contrast
abstract
Video saliency detection, the task to detect attractive content in a video, has broad applications in multimedia understanding and retrieval. In this paper, we propose a new framework for spatiotemporal saliency detection. To better estimate the salient motion in temporal domain, we take advantage of robust alignment by sparse and low-rank decomposition to jointly estimate the salient foreground motion and the camera motion. Consecutive frames are transformed and aligned, and then decomposed to a low-rank matrix representing the background and a sparse matrix indicating the objects with salient motion. In the spatial domain, we address several problems of local center-surround contrast based model, and demonstrate how to utilize global information and prior knowledge to improve spatial saliency detection. Individual component evaluation demonstrates the effectiveness of our temporal and spatial methods. Final experimental results show that the combination of our spatial and temporal saliency maps achieve the best overall performance compared to several state-of-the-art methods.
Zhixiang Ren, Liang-Tien Chia, Deepu Rajan
ICMR1
2010 Salient Region Detection by Jointly Modeling Distinctness and Redundancy of Image Content
Yiqun Hu, Zhixiang Ren, Deepu Rajan, Liang-Tien Chia
ACCV (2)2
2010 Improved saliency detection based on superpixel clustering and saliency propagation
abstract
Saliency detection is useful for high level applications such as adaptive compression, image retargeting, object recognition, etc. In this paper, we introduce an effective region-based solution for saliency detection. We first use the adaptive mean shift algorithm to extract superpixels from the input image, then apply Gaussian Mixture Model (GMM) to cluster superpixels based on their color similarity, and finally calculate the saliency value for each cluster using compactness metric together with modified PageRank propagation. This solution is able to represent the image in a perceptually meaningful way and is robust to over-segmentation. It highlights salient regions with full resolution, well-defined boundary. Experimental results show that both the adaptive mean shift and the modified PageRank algorithm contribute substantially to the saliency detection result. In addition, the ROC analysis demonstrates that our approach significantly outperforms five existing popular methods.
Zhixiang Ren, Yiqun Hu, Liang-Tien Chia, Deepu Rajan
ACM Multimedia1