Ruochi Zhang

dblp:37/3381 · DBLP profile ↗
← Back
19ranked-venue papers
11as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 PepHarmony: a multi-view contrastive learning framework for integrated sequence and structure-based peptide representation
Ruochi Zhang, Chang Liu 0082, Huaping Li, Yuqian Wu, Fengfeng Zhou, Xin Gao 0001
Neural Networks1
2026 DeepSelective: Interpretable prognosis prediction via feature selection and compression in EHR data
Ruochi Zhang, Xiaoyang Wang 0009, Qiong Zhou, Ziqi Deng, Yueying Wang, Yusi Fan, Jiale Zhang 0002, Lan Huang 0002, Chang Liu 0082, Fengfeng Zhou
Pattern Recognit.1
2025 DeepMolTex: Deep Alignment of Molecular Graphs with Large Language Models via Mixture of Modality Experts
abstract
Recent advances in Molecular Graph-Language Models (MGLMs) have demonstrated promising capabilities in molecular understanding tasks. However, existing approaches face critical limitations: (1) shallow alignment methods which employ identical processing modules for both modalities, resulting in compromised expressiveness and catastrophic forgetting of pre-trained language capabilities; and (2) over-reliance on high-level molecular representations that inadequately capture fine-grained structural information essential for comprehensive molecular understanding. To address these challenges, we present DeepMolTex, a novel framework for Deep fusion of Molecular structure and Textual representations across multiple scales. Our approach introduces a Mixture of Modality Experts (MoME) architecture that facilitates deep alignment between molecular graph features and large language models while preserving language capabilities, and a multi-scale graph projector that extracts and aligns molecular features at atom, motif, and molecule levels. Experimental results demonstrate that DeepMolTex significantly outperforms existing methods on fundamental molecular understanding tasks, including molecule description generation and IUPAC name prediction, while effectively preserving the language capabilities of the pre-trained LLM.
Mingliang Yan, Yanhua Yu, Ruochi Zhang, Zhiyuan Liu 0010, Ruicheng Zhang, Yimeng Ren 0001, Kangkang Lu 0002, Zhiyong Huang 0010, Feng Luo 0004
ACM Multimedia3
2025 PepLand: a large-scale pre-trained peptide representation model for a comprehensive landscape of both canonical and non-canonical amino acids
abstract
The recent interest in peptides incorporating non-canonical amino acids has surged within the scientific community, driven by their enhanced stability and resistance to proteolytic degradation. These so-called non-canonical peptides offer significant potential for modifying biological, pharmacological, and physiochemical characteristics in both native and synthetic contexts. Despite their advantages, there remains a notable gap in the availability of an efficient pre-trained model capable of effectively capturing feature representations from such intricate peptide sequences. This study herein introduces PepLand, a novel pre-training framework designed for the comprehensive representation and analysis of peptides, encompassing both canonical and non-canonical amino acids. PepLand leverages a general-purpose multi-view heterogeneous graph neural network to unveil the subtle structural representations of peptides. Our empirical evaluations demonstrate PepLand's proficiency in a range of peptide property prediction tasks, including cell penetrability, solubility, and protein-peptide binding affinity. These rigorous assessments affirm PepLand's superior capability in discerning critical representations of peptides with both canonical and non-canonical amino acids, and provide a robust foundation for transformative advances in peptide-focused pharmaceutical research. We have made the entire source code and datasets available at http://www.healthinformaticslab.org/supp/resources.php or https://github.com/zhangruochi/PepLand.
Ruochi Zhang, Chang Liu 0082, Yuting Xiu, Ningning Chen, Yu Wang 0225, Yan Wang 0028, Xin Gao 0001, Fengfeng Zhou
Briefings Bioinform.1
2025 TaiChiNet: PCA-based Ying-Yang dilution of inter- and intra-BERT layers to represent anti-coronavirus peptides
Shiying Ding, Yusi Fan, Yannan Sun, Gongyou Zhang, Ruochi Zhang, Lan Huang 0002, Fengfeng Zhou
Expert Syst. Appl.8
2024 HELM-GPT: de novo macrocyclic peptide design using generative pre-trained transformer
abstract
MOTIVATION: Macrocyclic peptides hold great promise as therapeutics targeting intracellular proteins. This stems from their remarkable ability to bind flat protein surfaces with high affinity and specificity while potentially traversing the cell membrane. Research has already explored their use in developing inhibitors for intracellular proteins, such as KRAS, a well-known driver in various cancers. However, computational approaches for de novo macrocyclic peptide design remain largely unexplored. RESULTS: Here, we introduce HELM-GPT, a novel method that combines the strength of the hierarchical editing language for macromolecules (HELM) representation and generative pre-trained transformer (GPT) for de novo macrocyclic peptide design. Through reinforcement learning (RL), our experiments demonstrate that HELM-GPT has the ability to generate valid macrocyclic peptides and optimize their properties. Furthermore, we introduce a contrastive preference loss during the RL process, further enhanced the optimization performance. Finally, to co-optimize peptide permeability and KRAS binding affinity, we propose a step-by-step optimization strategy, demonstrating its effectiveness in generating molecules fulfilling both criteria. In conclusion, the HELM-GPT method can be used to identify novel macrocyclic peptides to target intracellular proteins. AVAILABILITY AND IMPLEMENTATION: The code and data of HELM-GPT are freely available on GitHub (https://github.com/charlesxu90/helm-gpt).
Xiaopeng Xu, Chencheng Xu, Lesong Wei, Haoyang Li 0011, Juexiao Zhou, Ruochi Zhang, Yu Wang 0225, Yuanpeng Xiong, Xin Gao 0001
Bioinform.7
2024 MolFeSCue: enhancing molecular property prediction in data-limited and imbalanced contexts using few-shot and contrastive learning
abstract
MOTIVATION: Predicting molecular properties is a pivotal task in various scientific domains, including drug discovery, material science, and computational chemistry. This problem is often hindered by the lack of annotated data and imbalanced class distributions, which pose significant challenges in developing accurate and robust predictive models. RESULTS: This study tackles these issues by employing pretrained molecular models within a few-shot learning framework. A novel dynamic contrastive loss function is utilized to further improve model performance in the situation of class imbalance. The proposed MolFeSCue framework not only facilitates rapid generalization from minimal samples, but also employs a contrastive loss function to extract meaningful molecular representations from imbalanced datasets. Extensive evaluations and comparisons of MolFeSCue and state-of-the-art algorithms have been conducted on multiple benchmark datasets, and the experimental data demonstrate our algorithm's effectiveness in molecular representations and its broad applicability across various pretrained models. Our findings underscore MolFeSCues potential to accelerate advancements in drug discovery. AVAILABILITY AND IMPLEMENTATION: We have made all the source code utilized in this study publicly accessible via GitHub at http://www.healthinformaticslab.org/supp/ or https://github.com/zhangruochi/MolFeSCue. The code (MolFeSCue-v1-00) is also available as the supplementary file of this paper.
Ruochi Zhang, Chang Liu 0082, Yan Wang 0028, Lan Huang 0002, Fengfeng Zhou
Bioinform.1
2024 FairCare: Adversarial training of a heterogeneous graph neural network with attention mechanism to learn fair representations of electronic health records
Yan Wang 0028, Ruochi Zhang, Qiong Zhou, Shengde Zhang, Yusi Fan, Lan Huang 0002, Fengfeng Zhou
Inf. Process. Manag.2
2023 Orchestrating information across tissues via a novel multitask GAT framework to improve quantitative gene regulation relation modeling for survival analysis
abstract
Survival analysis is critical to cancer prognosis estimation. High-throughput technologies facilitate the increase in the dimension of genic features, but the number of clinical samples in cohorts is relatively small due to various reasons, including difficulties in participant recruitment and high data-generation costs. Transcriptome is one of the most abundantly available OMIC (referring to the high-throughput data, including genomic, transcriptomic, proteomic and epigenomic) data types. This study introduced a multitask graph attention network (GAT) framework DQSurv for the survival analysis task. We first used a large dataset of healthy tissue samples to pretrain the GAT-based HealthModel for the quantitative measurement of the gene regulatory relations. The multitask survival analysis framework DQSurv used the idea of transfer learning to initiate the GAT model with the pretrained HealthModel and further fine-tuned this model using two tasks i.e. the main task of survival analysis and the auxiliary task of gene expression prediction. This refined GAT was denoted as DiseaseModel. We fused the original transcriptomic features with the difference vector between the latent features encoded by the HealthModel and DiseaseModel for the final task of survival analysis. The proposed DQSurv model stably outperformed the existing models for the survival analysis of 10 benchmark cancer types and an independent dataset. The ablation study also supported the necessity of the main modules. We released the codes and the pretrained HealthModel to facilitate the feature encodings and survival analysis of transcriptome-based future studies, especially on small datasets. The model and the code are available at http://www.healthinformaticslab.org/supp/.
Meiyu Duan, Yueying Wang, Gongyou Zhang, Haotian Zhang 0018, Lan Huang 0002, Ruochi Zhang, Fengfeng Zhou
Briefings Bioinform.9
2022 Ultrafast and Interpretable Single-Cell 3D Genome Analysis with Fast-Higashi
Ruochi Zhang, Tianming Zhou, Jian Ma 0004
RECOMB1
2020 Hyper-SAGNN: a self-attention based graph neural network for hypergraphs
Ruochi Zhang, Yuesong Zou, Jian Ma 0004
ICLR1
2020 Probing Multi-way Chromatin Interaction with Hypergraph Representation Learning
Ruochi Zhang, Jian Ma 0004
RECOMB1
2018 Mutual-Information-Private Online Gradient Descent Algorithm
abstract
A user implemented privacy preservation mechanism is proposed for the online gradient descent (OGD) algorithm. Privacy is measured through the information leakage as quantified by the mutual information between the users outputs and learners inputs. The input perturbation mechanism proposed can be implemented by individual users with a space and time complexity that is independent of the horizon T. For the proposed mechanism, the information leakage is shown to be bounded by the Gaussian channel capacity in the full information setting. The regret bound of the privacy preserving learning mechanism is identical to the non private OGD with only differing in constant factors.
Ruochi Zhang, Parv Venkitasubramaniam
ICASSP1
2018 Predicting CTCF-mediated chromatin loops using CTCF-MP
abstract
Motivation: The three dimensional organization of chromosomes within the cell nucleus is highly regulated. It is known that CCCTC-binding factor (CTCF) is an important architectural protein to mediate long-range chromatin loops. Recent studies have shown that the majority of CTCF binding motif pairs at chromatin loop anchor regions are in convergent orientation. However, it remains unknown whether the genomic context at the sequence level can determine if a convergent CTCF motif pair is able to form a chromatin loop. Results: In this article, we directly ask whether and what sequence-based features (other than the motif itself) may be important to establish CTCF-mediated chromatin loops. We found that motif conservation measured by 'branch-of-origin' that accounts for motif turn-over in evolution is an important feature. We developed a new machine learning algorithm called CTCF-MP based on word2vec to demonstrate that sequence-based features alone have the capability to predict if a pair of convergent CTCF motifs would form a loop. Together with functional genomic signals from CTCF ChIP-seq and DNase-seq, CTCF-MP is able to make highly accurate predictions on whether a convergent CTCF motif pair would form a loop in a single cell type and also across different cell types. Our work represents an important step further to understand the sequence determinants that may guide the formation of complex chromatin architectures. Availability and implementation: The source code of CTCF-MP can be accessed at: https://github.com/ma-compbio/CTCF-MP. Supplementary information: Supplementary data are available at Bioinformatics online.
Ruochi Zhang, Yang Yang 0094, Yang Zhang 0042, Jian Ma 0004
Bioinform.1
2018 pyHIVE, a health-related image visualization and engineering system using Python
abstract
BACKGROUND: Imaging is one of the major biomedical technologies to investigate the status of a living object. But the biomedical image based data mining problem requires extensive knowledge across multiple disciplinaries, e.g. biology, mathematics and computer science, etc. RESULTS: pyHIVE (a Health-related Image Visualization and Engineering system using Python) was implemented as an image processing system, providing five widely used image feature engineering algorithms. A standard binary classification pipeline was also provided to help researchers build data models immediately after the data is collected. pyHIVE may calculate five widely-used image feature engineering algorithms efficiently using multiple computing cores, and also featured the modules of Principal Component Analysis (PCA) based preprocessing and normalization. CONCLUSIONS: The demonstrative example shows that the image features generated by pyHIVE achieved very good classification performances based on the gastrointestinal endoscopic images. This system pyHIVE and the demonstrative example are freely available and maintained at http://www.healthinformaticslab.org/supp/resources.php .
Ruochi Zhang, Ruixue Zhao, Xin Feng 0004, Fengfeng Zhou
BMC Bioinform.1
2017 Exploiting sequence-based features for predicting enhancer-promoter interactions
abstract
MOTIVATION: A large number of distal enhancers and proximal promoters form enhancer-promoter interactions to regulate target genes in the human genome. Although recent high-throughput genome-wide mapping approaches have allowed us to more comprehensively recognize potential enhancer-promoter interactions, it is still largely unknown whether sequence-based features alone are sufficient to predict such interactions. RESULTS: Here, we develop a new computational method (named PEP) to predict enhancer-promoter interactions based on sequence-based features only, when the locations of putative enhancers and promoters in a particular cell type are given. The two modules in PEP (PEP-Motif and PEP-Word) use different but complementary feature extraction strategies to exploit sequence-based information. The results across six different cell types demonstrate that our method is effective in predicting enhancer-promoter interactions as compared to the state-of-the-art methods that use functional genomic signals. Our work demonstrates that sequence-based features alone can reliably predict enhancer-promoter interactions genome-wide, which could potentially facilitate the discovery of important sequence determinants for long-range gene regulation. AVAILABILITY AND IMPLEMENTATION: The source code of PEP is available at: https://github.com/ma-compbio/PEP . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yang Yang 0094, Ruochi Zhang, Shashank Singh 0005, Jian Ma 0004
Bioinform.2
2017 Stealthy Control Signal Attacks in Linear Quadratic Gaussian Control Systems: Detectability Reward Tradeoff
abstract
The problem of false data injection through compromised cyber links to a physical control system modeled by linear quadratic Gaussian dynamics is studied in this paper. The control input stream is compromised by an attacker who modifies the (cyber) control signals transmitted with the objective of increasing the quadratic cost incurred by the (physical) controller whilst maintaining a degree of stealthiness. The tradeoff between the increase in quadratic cost and the stealthiness (or detectability), are measured by the Kullback-Leibler distance between legitimate and falsified state dynamics is characterized analytically. It is shown that the optimal adversarial strategy is a sequence of independent Gaussian noise signals with carefully chosen variances whose eigenvalues align with those of the legitimate noise covariance with the scaling reflecting the desired quadratic cost increase. As the stealthiness decreases, the optimal tradeoff is shown to be linear with slope inversely proportional to the maximal of maximal eigenvalue of modified reward matrices. Numerical simulations are presented that showcase the optimal tradeoff and the comparison of the legitimate and falsified dynamics under different requirements on detectability.
Ruochi Zhang, Parv Venkitasubramaniam
IEEE Trans. Inf. Forensics Secur.1
2011 Select-the-Best-Ones: A new way to judge relative relevance
Ruihua Song, Qingwei Guo, Ruochi Zhang, Guomao Xin, Ji-Rong Wen, Yong Yu 0001, Hsiao-Wuen Hon
Inf. Process. Manag.3
2008 GeoLife: Managing and Understanding Your Past Life over Maps
abstract
The increasing popularity of GPS device has boosted many applications where more and more GPS logs have been accumulating continuously. Managing and understanding the collected GPS data are two important issues for these applications. On one hand, by indexing the increasing GPS data, we can provide effective retrieval method for users to find the corresponding GPS data interests them. On the other hand, by understanding user's GPS data, we are more likely to enable novel services which would stimulate people's passion on contributing GPS data in turn. However, so far, GPS data are still used directly without much understanding. In our project, referred to as GeoLife, we focus on visualization, organization, fast retrieval, and effective understanding of GPS track logs for both personal and public use. It not only provides a powerful platform for people to effectively manage their GPS data but also help them well understand a person's past experience from GPS data.
Yu Zheng 0004, Longhao Wang, Ruochi Zhang, Xing Xie 0001, Wei-Ying Ma
MDM3