VLDB 2026 Research / reviewers in the wild / expert
Dong Si
dblp:63/10786
· DBLP profile ↗
15ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0001-7039-2589ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 9 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GOBoost: leveraging long-tail gene ontology terms for accurate protein function predictionabstractMOTIVATION: With the advancement of deep learning, researchers have increasingly proposed computational methods based on deep learning techniques to predict protein function. However, many of these methods treat protein function prediction as a multi-label classification problem, often overlooking the long-tail distribution of functional labels (i.e., Gene Ontology Terms) in datasets. To address this issue, we propose the GOBoost method, which incorporates the proposed long-tail optimization ensemble strategy. Besides, GOBoost introduces the proposed global-local label graph module and multi-granularity focal loss function to enhance long-tail functional information, mitigate the long-tail phenomenon, and improve overall prediction accuracy. RESULTS: We evaluate GOBoost and other state-of-the-art (SOTA) protein function prediction methods on the PDB and AF2 datasets. The GOBoost outperformed SOTA methods on all evaluation metrics for both datasets. Notably, in the AUPR evaluation on the PDB test set, GOBoost improved by 10.71%, 35.91%, and 22.71% compared to the SOTA HEAL method in the MF, BP, and CC functions. The experimental results show the necessity and superiority of designing models from the label long-tail distribution perspective. AVAILABILITY AND IMPLEMENTATION: The source code of GOBoost is available at https://github.com/Cao-Labs/GOBoost. Lei Zhang 0060, Jie Hou 0001, Dong Si, Bo Jiang 0002, Hailey Ledenko, Renzhi Cao |
Bioinform. | 5 |
| 2025 | AnglesRefine: Refinement of 3D Protein Structures Using Transformer Based on Torsion AnglesabstractThe goal of protein structure refinement is to enhance the precision of predicted protein models, particularly at the residue level of the local structure. Existing refinement approaches primarily rely on physics, whereas molecular simulation methods are resource-intensive and time-consuming. In this study, we employ deep learning methods to extract structural constraints from protein structure residues to assist in protein structure refinement. We introduce a novel method, AnglesRefine, which focuses on a protein's secondary structure and employs transformer to refine various protein structure angles (psi, phi, omega, CA_C_N_angle, C_N_CA_angle, N_CA_C_angle), ultimately generating a superior protein model based on the refined angles. We evaluate our approach against other cutting-edge methods using the CASP11-14 and CASP15 datasets. Experimental outcomes indicate that our method generally surpasses other techniques on the CASP11-14 test dataset, while performing comparably or marginally better on the CASP15 test dataset. Our method consistently demonstrates the least likelihood of model quality degradation, e.g., the degradation percentage of our method is less than 10%, while other methods are about 50%. Furthermore, as our approach eliminates the need for conformational search and sampling, it significantly reduces computational time compared to existing refinement methods. Lei Zhang 0060, Jun-Yong Zhu, Jie Hou 0001, Dong Si, Renzhi Cao |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | Enhancing cryo-EM structure prediction with DeepTracer and AlphaFold2 integrationabstractUnderstanding the protein structures is invaluable in various biomedical applications, such as vaccine development. Protein structure model building from experimental electron density maps is a time-consuming and labor-intensive task. To address the challenge, machine learning approaches have been proposed to automate this process. Currently, the majority of the experimental maps in the database lack atomic resolution features, making it challenging for machine learning-based methods to precisely determine protein structures from cryogenic electron microscopy density maps. On the other hand, protein structure prediction methods, such as AlphaFold2, leverage evolutionary information from protein sequences and have recently achieved groundbreaking accuracy. However, these methods often require manual refinement, which is labor intensive and time consuming. In this study, we present DeepTracer-Refine, an automated method that refines AlphaFold predicted structures by aligning them to DeepTracers modeled structure. Our method was evaluated on 39 multi-domain proteins and we improved the average residue coverage from 78.2 to 90.0% and average local Distance Difference Test score from 0.67 to 0.71. We also compared DeepTracer-Refine with Phenixs AlphaFold refinement and demonstrated that our method not only performs better when the initial AlphaFold model is less precise but also surpasses Phenix in run-time performance. Ayisha Zia, Albert Luo, Hanze Meng, Fengbin Wang, Jie Hou 0001, Renzhi Cao, Dong Si |
Briefings Bioinform. | 8 |
| 2023 | Fast and automated protein-DNA/RNA macromolecular complex modeling from cryo-EM mapsabstractCryo-electron microscopy (cryo-EM) allows a macromolecular structure such as protein-DNA/RNA complexes to be reconstructed in a three-dimensional coulomb potential map. The structural information of these macromolecular complexes forms the foundation for understanding the molecular mechanism including many human diseases. However, the model building of large macromolecular complexes is often difficult and time-consuming. We recently developed DeepTracer-2.0, an artificial-intelligence-based pipeline that can build amino acid and nucleic acid backbones from a single cryo-EM map, and even predict the best-fitting residues according to the density of side chains. The experiments showed improved accuracy and efficiency when benchmarking the performance on independent experimental maps of protein-DNA/RNA complexes and demonstrated the promising future of macromolecular modeling from cryo-EM maps. Our method and pipeline could benefit researchers worldwide who work in molecular biomedicine and drug discovery, and substantially increase the throughput of the cryo-EM model building. The pipeline has been integrated into the web portal https://deeptracer.uw.edu/. Andrew Nakamura, Hanze Meng, Minglei Zhao, Fengbin Wang, Jie Hou 0001, Renzhi Cao, Dong Si |
Briefings Bioinform. | 7 |
| 2023 | ComplexQA: a deep graph learning approach for protein complex structure assessmentabstractMOTIVATION: In recent years, the end-to-end deep learning method for single-chain protein structure prediction has achieved high accuracy. For example, the state-of-the-art method AlphaFold, developed by Google, has largely increased the accuracy of protein structure predictions to near experimental accuracy in some of the cases. At the same time, there are few methods that can evaluate the quality of protein complexes at the residue level. In particular, evaluating the quality of residues at the interface of protein complexes can lead to a wide range of applications, such as protein function analysis and drug design. In this paper, we introduce a new deep graph neural network-based method ComplexQA, to evaluate the local quality of interfaces for protein complexes by utilizing the residue-level structural information in 3D space and the sequence-level constraints. RESULTS: We benchmark our method to other state-of-the-art quality assessment approaches on the HAF2 and DBM55-AF2 datasets (high-quality structural models predicted by AlphaFold-Multimer), and the BM5 docking dataset. The experimental results show that our proposed method achieves better or similar performance compared with other state-of-the-art methods, especially on difficult targets which only contain a few acceptable models. Our method is able to suggest a score for each interfac e residue, which demonstrates a powerful assessment tool for the ever-increasing number of protein complexes. AVAILABILITY: https://github.com/Cao-Labs/ComplexQA.git. Contact: [email protected]. Lei Zhang 0060, Jie Hou 0001, Dong Si, Jun-Yong Zhu, Renzhi Cao |
Briefings Bioinform. | 4 |
| 2022 | DeepTracer-Denoising: Deep Learning for 3D Electron Density Map DenoisingabstractCryo-electron microscopy (Cryo-EM) is widely used in molecular structure determination and drug discovery. Experimental cryo-EM images suffer from the noises introduced by electron beam dose and sample preparation. Although many approaches have been proposed to improve the signal-to-noise ratio (SNR) for cryo-EM image denoising, the noises are still presented after 3D reconstruction and can obstruct the analysis and visualization of the 3D density map. Here we present DeepTracer-Denoising, a method for 3D electron density map denoising. We employ a 3D Neural Network to learn the pattern of noises and the biological structure from density maps. Our method is designed to work on medium to high-resolution maps ranging from 2.5 A to 10.0A. It is configurated with two modes to tackle both background noise and structural noise in a 3D density map. Our method can correctly identify 97.70% background noise while preserving 96.46% density of the native structure. For the maps that contain structural noise, DeepTracer-Denoising achieves an overall accuracy of 98.95%. Haowen Guan, Dong Si |
BIBM | 2 |
| 2022 | Psychosis iREACH: Reach for Psychosis Treatment using Artificial IntelligenceabstractPsychosis iREACH aims to optimize the delivery of evidence-based cognitive behavioral therapy to family caregivers who have a loved one with psychosis. It is an accessible digital platform that can utilize the user’s intent and entities to determine the appropriate response. The platform is implemented based on an artificial intelligence and natural language understanding (NLU) framework, RASA. We developed the web application of the platform, and the chatbot has been integrated into the platform to collect data and evaluate the performance. The results showed that the NLU model’s accuracy, precision, recall, and F1- score of the intent prediction are 88.31%, 89.80%, 88.22%, and 88.65% respectively. The link to the website is https://psychosisireach.uw.edu/. Sarah Kopelovich, Sunny Chieh Cheng, Dong Si |
BIBM | 4 |
| 2022 | ZoomQA: residue-level protein model accuracy estimation with machine learning on sequential and 3D structural featuresabstractMOTIVATION: The Estimation of Model Accuracy problem is a cornerstone problem in the field of Bioinformatics. As of CASP14, there are 79 global QA methods, and a minority of 39 residue-level QA methods with very few of them working on protein complexes. Here, we introduce ZoomQA, a novel, single-model method for assessing the accuracy of a tertiary protein structure/complex prediction at residue level, which have many applications such as drug discovery. ZoomQA differs from others by considering the change in chemical and physical features of a fragment structure (a portion of a protein within a radius $r$ of the target amino acid) as the radius of contact increases. Fourteen physical and chemical properties of amino acids are used to build a comprehensive representation of every residue within a protein and grade their placement within the protein as a whole. Moreover, we have shown the potential of ZoomQA to identify problematic regions of the SARS-CoV-2 protein complex. RESULTS: We benchmark ZoomQA on CASP14, and it outperforms other state-of-the-art local QA methods and rivals state of the art QA methods in global prediction metrics. Our experiment shows the efficacy of these new features and shows that our method is able to match the performance of other state-of-the-art methods without the use of homology searching against databases or PSSM matrices. AVAILABILITY: http://zoomQA.renzhitech.com. Kyle Hippe, Cade Lilley, Joshua William Berkenpas, Ciri Chandana Pocha, Kiyomi Kishaba, Jie Hou 0001, Dong Si, Renzhi Cao |
Briefings Bioinform. | 8 |
| 2022 | Neural representations of cryo-EM maps and a graph-based interpretationabstractBACKGROUND: Advances in imagery at atomic and near-atomic resolution, such as cryogenic electron microscopy (cryo-EM), have led to an influx of high resolution images of proteins and other macromolecular structures to data banks worldwide. Producing a protein structure from the discrete voxel grid data of cryo-EM maps involves interpolation into the continuous spatial domain. We present a novel data format called the neural cryo-EM map, which is formed from a set of neural networks that accurately parameterize cryo-EM maps and provide native, spatially continuous data for density and gradient. As a case study of this data format, we create graph-based interpretations of high resolution experimental cryo-EM maps. RESULTS: Normalized cryo-EM map values interpolated using the non-linear neural cryo-EM format are more accurate, consistently scoring less than 0.01 mean absolute error, than a conventional tri-linear interpolation, which scores up to 0.12 mean absolute error. Our graph-based interpretations of 115 experimental cryo-EM maps from 1.15 to 4.0 Å resolution provide high coverage of the underlying amino acid residue locations, while accuracy of nodes is correlated with resolution. The nodes of graphs created from atomic resolution maps (higher than 1.6 Å) provide greater than 99% residue coverage as well as 85% full atomic coverage with a mean of 0.19 Å root mean squared deviation. Other graphs have a mean 84% residue coverage with less specificity of the nodes due to experimental noise and differences of density context at lower resolutions. CONCLUSIONS: The fully continuous and differentiable nature of the neural cryo-EM map enables the adaptation of the voxel data to alternative data formats, such as a graph that characterizes the atomic locations of the underlying protein or macromolecular structure. Graphs created from atomic resolution maps are superior in finding atom locations and may serve as input to predictive residue classification and structure segmentation methods. This work may be generalized to transform any 3D grid-based data format into non-linear, continuous, and differentiable format for downstream geometric deep learning applications. Nathan Ranno, Dong Si |
BMC Bioinform. | 2 |
| 2019 | Scaling up Prediction of Psychosis by Natural Language ProcessingabstractMental health professionals currently diagnose and treat mental disorders, such as schizophrenia, mainly by analyzing the language and speech of their patients, a method that maybe improved with the usage of artificial intelligence. This study aims to use machine learning to distinguish between the speech of patients who suffer from mental disorders which cause psychosis from that of healthy individuals to improve early detection of schizophrenia. We analyzed forty interview transcripts from patients who have been diagnosed with first episode psychosis. Word embeddings and convolutional neural network were utilized for the classification of patients from healthy individuals. The preliminary test results achieved a prediction rate of 99%, which indicated that our speech classifier was able to discriminate speech in patients from healthy individuals' daily conversations. This suggested that machine learning models can learn and train upon features of natural languages to predict whether or not an individual is beginning to show the first signs of early psychosis based on their speech. This line of inquiry will contribute to the improved identification of individuals at risk for psychiatric symptoms and lead to the development of targeted therapies. Source code and data of this work have been made public on https://github.com/DrDongSi/Psychosis_NLP. Dong Si, Sunny Chieh Cheng, Ruiwen Xing, Hoi Yan Wu |
ICTAI | 1 |
| 2018 | GPU Accelerated Ray Tracing for the Beta-Barrel Detection from Three-Dimensional Cryo-EM Maps
Albert Ng, Adedayo Odesile, Dong Si |
ISBRA | 3 |
| 2017 | Genetic Algorithm Based Beta-Barrel Detection for Medium Resolution Cryo-EM Density Maps
Albert Ng, Dong Si |
ISBRA | 2 |
| 2016 | Deep convolutional neural networks for detecting secondary structures in protein density maps from cryo-electron microscopyabstractThe detection of secondary structure of proteins using three dimensional (3D) cryo-electron microscopy (cryo-EM) images is still a challenging task when the spatial resolution of cryo-EM images is at medium level (5-10Å). Prior researches focused on the usage of local features that may not capture the global information of image objects. In this study, we propose to use deep learning methods to extract high representative global features and then automatically detect secondary structures of proteins. In particular, we build a convolutional neural network (CNN) classifier that predicts the probability of label for every individual voxel in 3D cryo-EM image with respect to the secondary structure elements of proteins such as α-helix, β-sheet and background. To effectively incorporate the 3D spatial information in protein structures, we propose to perform 3D convolutions in the convolutional layers of CNNs. We show that the proposed CNN classifier can outperform existing SVM method on identifying the secondary structure elements of proteins from 3D cryo-EM medium resolution images. Rongjian Li, Dong Si, Shuiwang Ji, Jing He 0002 |
BIBM | 2 |
| 2015 | Detection of Secondary Structures from 3D Protein Images of Medium Resolutions and its Challenges
Jing He 0002, Dong Si, Maryam Arab |
ICIG (2) | 2 |
| 2011 | A Constraint Dynamic Graph Approach to Identify the Secondary Structure Topology from cryoEM Density Data in Presence of ErrorsabstractThe determination of the secondary structure topology is a critical step in deriving the atomic structure from the protein density map obtained from electron cryo-microscopy technique. This step often relies on the matching of two sources of information. One source comes from the secondary structures detected from the protein density map at the medium resolution, such as 5-10 A. The other source comes from the predicted secondary structures from the amino acid sequence. Due to the uncertainty in either source of information, a pool of possible secondary structure positions has to be sampled in order to include the true answer. A naive way to find the native topology is to exhaustively map the pool of possible secondary structures detected in the density map with the pool of the secondary structures predicted from the sequence and search for the topology with the lowest cost. This paper studies the question that is how to reduce the computation of the mapping when the uncertainty of the secondary structure predictions is considered. We present a method that combines the concept of dynamic graph with our previous work of using constrained shortest path to identify the topology of the secondary structures. We show a reduction of about 34.55% time as comparison to the naive way of handling the inaccuracies. To our knowledge, this is the Is computationally effective exact algorithm to identify the optimal topology of the secondary structures when the inaccuracy of the predicted data is considered. Abhishek Biswas, Dong Si, Kamal Al-Nasr, Desh Ranjan, Mohammad Zubair, Jing He 0002 |
BIBM | 2 |