EDBT 2026 Demo / reviewers in the wild / expert
Zhenghe Yang
dblp:137/1031
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Boosting Deep Learning-based Docking with Cross-attention and Centrality EmbeddingabstractDocking is a classic computational biology problem that is widely used to predict binding conformations and to virtually screen binding molecules. Recently, a deep learning-based method, DeepDock was proposed to address the docking problem. The method shows great performance on conformation prediction. One major limitation of the method is that the interaction of ligands and targets is too simple. Here, we introduce caDeepDock, which is a geometric deep learning model for protein-ligand binding conformation prediction. Inspired by DeepDock, caDeepDock has two major advantages over DeepDock. First, cross-attention is employed to enable communications between the molecule and the protein binding pocket. Second, a positional embedding based on node degrees is used to incorporate both data-dependent and position-dependent communications. Experiments on the CASF-2016 benchmark have shown that the potential function learned by caDeepDock is able to pick a near-native conformation (with RMSD$\leqslant$2Å) from a set of decoy conformations with a success rate of 91.6%. This result outperforms not only DeepDock by 4.6% but also AutoDock by 1.4%. Consequently, the potential function learned by caDeepDock yields 5.96% more near-native conformations in conformation optimization experiments. Zongzhao Qiu, Zhenghe Yang, Xuefeng Cui |
BIBM | 4 |
| 2022 | Graph-based Reaction Classification by Contrasting between Precursors and ProductsabstractThe classification of organic reactions is a complex and tedious process, and it often requires certain domain knowledge to understand the classification rules. To reduce such requirements on domain knowledge, BERT-based deep learning methods have been applied to classify reactions based on SMILES strings. However, the same reaction can be represented by different but equivalent SMILES strings, and it can be observed that BERT-based methods are highly sensitive to the choice of SMILES strings. Logically, GNN-based methods are robust to equivalent SMILES strings. Here, we propose a graph isomorphism network (GIN) with contrastive learning, called ContraGIN, to encourage feature fusion learning between precursors and products of the reaction. Indeed, experiments have shown that the new method focuses on learning the atoms near the bond changes before and after the reaction, and consequently achieves a classification accuracy of 99.30% on the USPTO 1k TPL dataset. Additional experiments have also shown that ContraGIN is more robust and much faster for lager reactions (with more atoms) and complicated reactions (with rings). Yuxiao Wang 0002, Zhaoxu Meng, Zhenghe Yang, Xuefeng Cui |
BIBM | 5 |
| 2022 | CoAtGIN: Marrying Convolution and Attention for Graph-based Molecule Property PredictionabstractMolecule property prediction based on computational strategies plays a key role in the process of drug discovery and design processes, such as DFT. However, these traditional methods are time-consuming, labor-intensive, and cannot satisfy the need for biomedicine. Owning to the development of deep learning, there are many variants of Graph Neural Networks (GNN) for molecular representation learning. However, the existing well-performing graph-based methods that have a number of parameters or light models cannot achieve good grades on various tasks. To manage the trade-off between efficiency and performance, we propose a novel model architecture, CoAtGIN, using both Convolution and Attention. At the local level, k-hop convolution is designed to capture long-range neighbor information. At the global level, in addition to using the virtual node to pass identical messages, we utilize linear attention to the aggregate global graph representation according to the importance of each node and edge. In the recent Open Graph Benchmark (OGB) Large-Scale Benchmark, CoAtGIN achieves the 0.0901 Mean Absolute Error (MAE) on the large-scale dataset PCQM4Mv2 with only 6.4 M model parameters. Moreover, using the linear attention block improves the performance, which helps to capture the global representation. Zhaoxu Meng, Zhenghe Yang, Haitao Jiang 0005, Xuefeng Cui |
BIBM | 4 |
| 2021 | Jointly Learning to Align and Aggregate with Cross Attention Pooling for Peptide-MHC Class I Binding PredictionabstractPredicting binding affinities of peptide antigens presented on major histocompatibility complex (MHC) is of great importance in T-cell immune response research. Accurate prediction of peptide-MHC binding affinities is essential for vaccine design and disease treatment. Recent deep learning-based prediction methods have shown that effective sequence embedding is critical to accurately predict binding affinities. One common neural network layer shared by these methods is the global average pooling layer that aggregates features. However, can we design a better global pooling layer? Here, we introduce a novel cross attention pooling (caPool) layer to aggregate features. As our initial application of caPool, a novel end-to-end transformer model, called capTransformer, is proposed for peptide-MHC class I binding prediction. In our model, caPool jointly aligns peptide-MHC residual pairs and aggregates residual features. Thus, instead of treating all residues equally and independently, caPool focuses more on correlated residue pairs that are potentially contact pairs contributing major forces to stabilize the complex structure. Using a five-fold cross-validation experiment, we found that caPool achieved the highest PCC value of 0.845, which was 0.139 higher than a global average pooling. Here, the global pooling layer was the only difference between the two tested models, and this observation indicated that global average pooling was not always the best choice. Importantly, our capTransformer model achieved a SRCC value of 0.614 (i.e., 6.4% higher than the best-performing method) when applied to the IEDB dataset. Cheng Chen 0051, Zongzhao Qiu, Zhenghe Yang, Bin Yu 0007, Xuefeng Cui |
BIBM | 3 |
| 2021 | Edge-Gated Graph Neural Network for Predicting Protein-Ligand Binding AffinitiesabstractPredicting Protein-ligand binding affinities using Deep Learning can significantly shorten the drug development cycle. Recently, Graph neural network models have been developed, and are successfully used to accelerate the development of potential drugs. One major limitation of these GNN models is that they focus on node features (i.e., atom features), as these nodes carry the most important information of molecules. However, atoms are connected via different bonds in molecules, and we argue that such chemical bonds carry critical information for assessing how atomic features should be aggregated. To overcome the lack of bond-related information in earlier models, we here proposed a novel edge-gated graph neural network (egGNN) that predict the binding affinities between proteins and ligands. Specifically, our model treats chemical bonds as gates that control how information is extracted between atoms, this modification enables our model to extract more accurate information for different bonds. We tested our model using the CASF-2016 dataset, and found that the Pearson’s correlation coefficient (R) of egGNN is three percent higher than that of the best tested method (i.e., 0.86 vs 0.83), and the Root Mean Square Error (RMSE) is significantly lower than that of the best tested method (i.e., 1.12 vs 1.23) when compared to state-of-the-art-approaches. In ablation experiments, we demonstrate that our edge-gated feature extraction (EGFE) consoderably improves the performance of GNNs. These results indicate that egGNN represents a promising tool applicable for virtual screening, and should greatly assist in accelerating drug development. Qihong Jiao, Zongzhao Qiu, Yuxiao Wang 0002, Zhenghe Yang, Xuefeng Cui |
BIBM | 5 |
| 2021 | PG-TFNet: Transformer-based Fusion Network Integrating Pathological Images and Genomic Data for Cancer Survival AnalysisabstractSurvival analysis is crucial to the evaluation of cancer treatment options and deep learning-based methods integrating pathological images and genomic data have been used for prognosis prediction. However, the most methods are based on the analysis of pathological image patches, thus ignoring the morphological structure information at larger field-of-view and intrinsic relationships between patches. Meanwhile, the existing models fail to exploit the powerful representation learning capabilities of the neural networks for effective multimodal feature fusion of pathological images and genomic data. In this paper, we propose a novel transformer-based fusion network integrating pathological images and genomic data (PGTFNet) for cancer survival analysis. Specifically, we present a transformer-based feature fusion module for multi-scale pathological slides to fully exploit the intra-modality relationships between image patches at various fields of view. Moreover, in order to make effective inter-modality feature fusion of pathological images and genomic data, we introduce a cross-attention transformer module that can exchange feature representations of different modalities between two transformers branches. The PG-TFNet is performed on the colorectal cancer dataset from the Cancer Genome Atlas (TCGA), which contains paired whole-slide images and genomic data with ground truth survival data. The experimental results from a 10-fold cross validation demonstrate that the proposed PG-TFNet facilitates the prognosis prediction of colorectal cancer and shows superiority over the existing methods. Zhilong Lv, Yuexiao Lin, Rui Yan 0009, Zhenghe Yang, Ying Wang 0043, Fa Zhang 0001 |
BIBM | 4 |
| 2021 | TransPicker: a Transformer-based Framework for Particle Picking in cryoEM MicrographsabstractSingle-particle cryo-electron microscopy (cryoEM) methods are powerful for solving high-resolution structures of biological macromolecules. Locating numerous particles from micrographs is essential for three-dimensional reconstruction but challenging due to the extremely low signal-to-noise ratio and various particle shapes in micrographs. In this study, we devise the TransPicker, a two-dimensional particle picking framework based on a novel end-to-end transformer-based detective method named crDETR (cryoEM DEtection TRansformer). crDETR applies an improved deformable Transformer to perform inference in parallel on particle relocation and global context, without hand-crafted components like anchors, non-maximum suppression procedure, or sliding windows. Also, it uses a combined loss function to guarantee fast convergence. It uses divide-and-conquer to overcome the limitations of the object query number. Moreover, we develop a series of optimizations, including denoising, enhancing, bad particle filtering, adding masks on carbon areas and ice contaminants to decrease the false-positive ratio and improve the accuracy of particle picking. Experimental results on various datasets demonstrate that TransPicker can select particles with more accuracy, especially in high noise compared with other methods. To our knowledge, TransPicker is the first application of the transformer technique in cryoEM particle picking. Chi Zhang 0110, Hongjia Li 0001, Zhenghe Yang, Jieqing Feng, Fa Zhang 0001 |
BIBM | 5 |