VLDB 2026 Research / reviewers in the wild / expert
Ruhong Zhou
dblp:04/5523
· DBLP profile ↗
12ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0001-8624-5591ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decoding TCR recognition via geometric deep learning of immunological fingerprintsabstractT cell receptor (TCR) recognition of peptide-major histocompatibility complex (pMHC) molecules is the critical first step in adaptive immune activation, shaping immunity against pathogens and tumors, as well as tolerance to self. Despite extensive structural characterization of TCR-pMHC complexes, the molecular principles underlying this process remain incompletely understood, hindered by the inherent duality of TCR specificity and cross-reactivity. Traditional structural analyses often fall short in capturing the multidimensional features that govern TCR-pMHC engagement. Here, we introduce a multimodal geometric deep learning framework that systematically extracts and learns various physicochemical and spatial features from pMHC interfaces, which encode key immunological cues for TCR recognition. Applied to a curated dataset of human leukocyte antigens HLA-A*02-peptide-TCR crystal structures, our model robustly predicts TCR binding preferences and uncovers interfacial "immunological fingerprints" that inform receptor engagement. Through an integrated explainability module, we identify critical contact residues and interaction motifs, thus providing interpretable insights into the determinants of TCR specificity. We further demonstrate the model's generalizability by analyzing HLA-B*27-peptide complexes, revealing potential TCR cross-reactivity between self-derived and bacterial peptides-highlighting its utility in probing molecular mimicry. This work establishes a scalable, structure-based approach for decoding T cell recognition and offers a powerful tool for guiding antigen design, vaccine development, and TCR-based immunotherapies. Chun Shang, Kevin C. Chan, Ruhong Zhou |
Briefings Bioinform. | 3 |
| 2025 | A free energy perturbation-assisted machine learning strategy for mimotope screening in neoantigen-based vaccine designabstractNeoantigen-based immunotherapy has emerged as a promising approach for cancer treatment. One key strategy in neoantigen-based vaccine design is to alter known neoantigens into enhanced mimotopes that elicit more robust immune responses. However, screening mimotopes presents challenges in both diversity and precision. While machine learning (ML) models facilitate high-throughput screening of immunogenic candidates, they struggle to distinguish mimotopes from original neoantigens (i.e. identify mimotopes with higher binding affinities, rather than solely distinguish between binding and nonbinding peptides). In contrast, alchemical methods such as free energy perturbation (FEP) provide quantitative binding free-energy differences between mimotopes and neoantigens but are computationally intensive. To leverage the strengths of both approaches, we propose an FEP-assisted ML (FEPaML) strategy that employs Bayesian optimization to iteratively refine knowledge-based predictions with physics-based evaluations, thereby progressively achieving locally optimized, precise, and robust outcomes. Our FEPaML strategy is then applied to screen mimotopes for several representative neoantigens. It has demonstrated excellent predictive precisions (exceeding 0.9) with a relatively small number of FEP samplings, significantly outperforming existing ML models. Qinglu Zhong, Kevin C. Chan, Ruhong Zhou |
Briefings Bioinform. | 4 |
| 2024 | IsRNAcirc: 3D structure prediction of circular RNAs based on coarse-grained molecular dynamics simulationabstractAs an emerging class of RNA molecules, circular RNAs play pivotal roles in various biological processes, thereby determining their three-dimensional (3D) structure is crucial for a deep understanding of their biological significances. Similar to linear RNAs, the development of computational methods for circular RNA 3D structure prediction is challenging, especially considering the inherent flexibility and potentially long length of circular RNAs. Here, we introduce an extension of our previous IsRNA2 model, named IsRNAcirc, to enable circular RNA 3D structure predictions through coarse-grained molecular dynamics simulations. The workflow of IsRNAcirc consists of four main steps, including input preparation, end closure, structure prediction, and model refinement. Our results demonstrate that IsRNAcirc can provide reasonable 3D structure predictions for circular RNAs, which significantly reduce the locally irrational elements contained in the initial input. Moreover, for a validation test set comprising 34 circular RNAs, our IsRNAcirc can generate 3D models with better scores than the template-based 3dRNA method. These findings demonstrate that our IsRNAcirc method is a promising tool to explore the structural details along with intricate interactions of circular RNAs. Haolin Jiang, Yulian Xu, Yunguang Tong, Dong Zhang 0020, Ruhong Zhou |
PLoS Comput. Biol. | 5 |
| 2023 | OPUS-Fold3: a gradient-based protein all-atom folding and docking framework on TensorFlowabstractFor refining and designing protein structures, it is essential to have an efficient protein folding and docking framework that generates a protein 3D structure based on given constraints. In this study, we introduce OPUS-Fold3 as a gradient-based, all-atom protein folding and docking framework, which accurately generates 3D protein structures in compliance with specified constraints, such as a potential function as long as it can be expressed as a function of positions of heavy atoms. Our tests show that, for example, OPUS-Fold3 achieves performance comparable to pyRosetta in backbone folding and significantly better in side-chain modeling. Developed using Python and TensorFlow 2.4, OPUS-Fold3 is user-friendly for any source-code level modifications and can be seamlessly combined with other deep learning models, thus facilitating collaboration between the biology and AI communities. The source code of OPUS-Fold3 can be downloaded from http://github.com/OPUS-MaLab/opus_fold3. It is freely available for academic usage. Zhenwei Luo, Ruhong Zhou, Jianpeng Ma 0004 |
Briefings Bioinform. | 3 |
| 2021 | CASTELO: clustered atom subtypes aided lead optimization - a combined machine learning and molecular modeling methodabstractBACKGROUND: Drug discovery is a multi-stage process that comprises two costly major steps: pre-clinical research and clinical trials. Among its stages, lead optimization easily consumes more than half of the pre-clinical budget. We propose a combined machine learning and molecular modeling approach that partially automates lead optimization workflow in silico, providing suggestions for modification hot spots. RESULTS: The initial data collection is achieved with physics-based molecular dynamics simulation. Contact matrices are calculated as the preliminary features extracted from the simulations. To take advantage of the temporal information from the simulations, we enhanced contact matrices data with temporal dynamism representation, which are then modeled with unsupervised convolutional variational autoencoder (CVAE). Finally, conventional and CVAE-based clustering methods are compared with metrics to rank the submolecular structures and propose potential candidates for lead optimization. CONCLUSION: With no need for extensive structure-activity data, our method provides new hints for drug modification hotspots which can be used to improve drug potency and reduce the lead optimization time. It can potentially become a valuable tool for medicinal chemists. Leili Zhang, Giacomo Domeniconi, Chih-Chieh Yang, Seung-gu Kang, Ruhong Zhou, Guojing Cong |
BMC Bioinform. | 5 |
| 2012 | Multiscale modeling of macromolecular biosystemsabstractIn this article, we review the recent progress in multiresolution modeling of structure and dynamics of protein, RNA and their complexes. Many approaches using both physics-based and knowledge-based potentials have been developed at multiple granularities to model both protein and RNA. Coarse graining can be achieved not only in the length, but also in the time domain using discrete time and discrete state kinetic network models. Models with different resolutions can be combined either in a sequential or parallel fashion. Similarly, the modeling of assemblies is also often achieved using multiple granularities. The progress shows that a multiresolution approach has considerable potential to continue extending the length and time scales of macromolecular modeling. Samuel Flores, Julie Bernauer, Seokmin Shin, Ruhong Zhou, Xuhui Huang |
Briefings Bioinform. | 4 |
| 2009 | Using a mutual information-based site transition network to map the genetic evolution of influenza A/H3N2 virusabstractMOTIVATION: Mapping the antigenic and genetic evolution pathways of influenza A is of critical importance in the vaccine development and drug design of influenza virus. In this article, we have analyzed more than 4000 A/H3N2 hemagglutinin (HA) sequences from 1968 to 2008 to model the evolutionary path of the influenza virus, which allows us to predict its future potential drifts with specific mutations. RESULTS: The mutual information (MI) method was used to design a site transition network (STN) for each amino acid site in the A/H3N2 HA sequence. The STN network indicates that most of the dynamic interactions are positioned around the epitopes and the receptor binding domain regions, with strong preferences in both the mutation sites and amino acid types being mutated to. The network also shows that antigenic changes accumulate over time, with occasional large changes due to multiple co-occurring mutations at antigenic sites. Furthermore, the cluster analysis by subdividing the STN into several subnetworks reveals a more detailed view about the features of the antigenic change: the characteristic inner sites and the connecting inter-subnetwork sites are both responsible for the drifts. A novel five-step prediction algorithm based on the STN shows a reasonable accuracy in reproducing historical HA mutations. For example, our method can reproduce the 2003-2004 A/H3N2 mutations with approximately 70% accuracy. The method also predicts seven possible mutations for the next antigenic drift in the coming 2009-2010 season. The STN approach also agrees well with the phylogenetic tree and antigenic maps based on HA inhibition assays. AVAILABILITY: All code and data are available at http://ibi.zju.edu.cn/birdflu/. Zhen Xia, Gulei Jin, Ruhong Zhou |
Bioinform. | 4 |
| 2007 | PROTERAN: animated terrain evolution for visual analysis of patterns in protein folding trajectoryabstractThe mechanism of protein folding remains largely a mystery in molecular biology, despite the enormous effort from many groups in the past decades. Currently, the protein folding mechanism is often characterized by calculating the free energy landscape versus various reaction coordinates such as the fraction of native contacts, the radius of gyration and so on. In this paper, we present an integrated approach towards understanding the folding process via visual analysis of patterns of these reaction coordinates. The three disparate processes (1) protein folding simulation, (2) pattern elicitation and (3) visualization of patterns, work in tandem. Thus as the protein folds, the changing landscape in the pattern space can be viewed via the visualization tool, PROTERAN, a program we developed for this purpose. We first present an incremental (on-line) trie-based pattern discovery algorithm to elicit the patterns and then describe the terrain metaphor based visualization tool. Using two example small proteins, a beta-hairpin and a designed protein Trp-cage, we next demonstrate that this combined pattern discovery and visualization approach extracts crucial information about protein folding intermediates and mechanism. Ruhong Zhou, Laxmi Parida, Kush Kapila, Sudhir P. Mudur |
Bioinform. | 1 |
| 2006 | Parallel implementation of the replica exchange molecular dynamics algorithm on Blue Gene/LabstractThe replica exchange method is a popular approach for studying the folding thermodynamics of small to modest size proteins in explicit solvent, since it is easily parallelized. However, replica exchange can become computationally expensive for large-scale studies, due to the number of replicas needed as well as interprocess or communication requirements both between and within replicas. In this paper, we discuss an implementation of replica exchange molecular dynamics on Blue Gene/L for performing large scale simulation studies of systems of biological interest. The algorithm is tuned with an awareness of the physical network topology and hardware performance features of the Blue Gene/L architecture. Performance measurements for replica exchange using the blue matter molecular dynamics application are presented on Blue Gene/L hardware with up to 256 replicas simulated on 8,192 compute nodes. Both scalability and performance are achieved with this implementation. Maria Eleftheriou, Aleksandr Rayshubskiy, Jed W. Pitera, Blake G. Fitch, Ruhong Zhou, Robert S. Germain |
IPDPS | 5 |
| 2005 | Protein folding trajectory analysis using patterned clusters
Laxmi Parida, Ruhong Zhou |
APBC | 3 |
| 2005 | Combinatorial Pattern Discovery Approach for the Folding Trajectory Analysis of a β-HairpinabstractThe study of protein folding mechanisms continues to be one of the most challenging problems in computational biology. Currently, the protein folding mechanism is often characterized by calculating the free energy landscape versus various reaction coordinates, such as the fraction of native contacts, the radius of gyration, RMSD from the native structure, and so on. In this paper, we present a combinatorial pattern discovery approach toward understanding the global state changes during the folding process. This is a first step toward an unsupervised (and perhaps eventually automated) approach toward identification of global states. The approach is based on computing biclusters (or patterned clusters)-each cluster is a combination of various reaction coordinates, and its signature pattern facilitates the computation of the Z-score for the cluster. For this discovery process, we present an algorithm of time complexity c in RO((N + nm) log n), where N is the size of the output patterns and (n x m) is the size of the input with n time frames and m reaction coordinates. To date, this is the best time complexity for this problem. We next apply this to a beta-hairpin folding trajectory and demonstrate that this approach extracts crucial information about protein folding intermediate states and mechanism. We make three observations about the approach: (1) The method recovers states previously obtained by visually analyzing free energy surfaces. (2) It also succeeds in extracting meaningful patterns and structures that had been overlooked in previous works, which provides a better understanding of the folding mechanism of the beta-hairpin. These new patterns also interconnect various states in existing free energy surfaces versus different reaction coordinates. (3) The approach does not require calculating the free energy values, yet it offers an analysis comparable to, and sometimes better than, the methods that use free energy landscapes, thus validating the choice of reaction coordinates. (An abstract version of this work was presented at the 2005 Asia Pacific Bioinformatics Conference [1].). Laxmi Parida, Ruhong Zhou |
PLoS Comput. Biol. | 2 |
| 2003 | Blue Matter, an application framework for molecular simulation on Blue Gene
Blake G. Fitch, Robert S. Germain, Mark P. Mendell, Jed W. Pitera, Michael Pitman, Aleksandr Rayshubskiy, Yuk Yin Sham, Frank Suits, William C. Swope, T. J. Christopher Ward, Yuriy Zhestkov, Ruhong Zhou |
J. Parallel Distributed Comput. | 12 |