Yingfan Ma

dblp:287/6312 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing Protein Structure Learning Using a Size-Guided Conditional Mixture-of-Experts
abstract
In recent years, deep learning on protein structures has attracted widespread attention, as structures determine proteins' function. A series of structure-based protein property prediction methods have been proposed, achieving remarkable performance. However, these methods often neglect the importance of the protein size and fail to fully leverage it, leading to biases toward certain sizes and suboptimal overall performance. To address this issue, we propose a protein size-guided conditional mixture-of-experts for improving deep learning on protein structures. It can adaptively activate the sub-networks with the guidance of protein sizes and network features. Its flexible combinations of sub-networks help mitigate biases toward certain protein sizes, while the deliberate incorporation of protein size guidance enables the network to effectively capture both universal and size-specific characteristics, resulting in more accurate predictive performance. Based on it, we propose a framework for protein property prediction and benchmark it on eight tasks with two representation forms of proteins and three different dataset splits, a total of forty-eight tests. Experiments show that our method can be seamlessly integrated into numerous existing models and achieve performance improvement across tasks under almost all settings. More importantly, our experiments reveal that although often overlooked, protein size serves as an important prior knowledge in deep learning on protein structures.
Mingzhi Yuan, Siqi Yin, Yingfan Ma, Manning Wang
IEEE J. Biomed. Health Informatics4
2025 Flow-MIL: Constructing Highly-expressive Latent Feature Space for Whole Slide Image Classification using Normalizing Flow
Yingfan Ma, Bohan An, Mingzhi Yuan, Minghong Duan, Manning Wang
ICCV1
2025 Drug-TTA: Test-Time Adaptation for Drug Virtual Screening via Multi-task Meta-Auxiliary Learning
abstract
Virtual screening is a critical step in drug discovery, aiming at identifying potential drugs that bind to a specific protein pocket from a large database of molecules. Traditional docking methods are time-consuming, while learning-based approaches supervised by high-precision conformational or affinity labels are limited by the scarcity of training data. Recently, a paradigm of feature alignment through contrastive learning has gained widespread attention. This method does not require explicit binding affinity scores, but it suffers from the issue of overly simplistic construction of negative samples, which limits their generalization to more difficult test cases. In this paper, we propose Drug-TTA, which leverages a large number of self-supervised auxiliary tasks to adapt the model to each test instance. Specifically, we incorporate the auxiliary tasks into both the training and the inference process via meta-learning to improve the performance of the primary task of virtual screening. Additionally, we design a multi-scale feature based Auxiliary Loss Balance Module (ALBM) to balance the auxiliary tasks to improve their efficiency. Extensive experiments demonstrate that Drug-TTA achieves state-of-the-art (SOTA) performance in all five virtual screening tasks under a zero-shot setting, showing an average improvement of 9.86% in AUROC metric compared to the baseline without test-time adaptation.
Mingzhi Yuan, Yingfan Ma, Manning Wang
ICML3
2025 Knowledge-Guided Multi-scale Graph Mamba for Whole Slide Image Classification
Minghong Duan, Yingfan Ma, Manning Wang, Zhijian Song
MICCAI (12)3
2025 ProteinF3S: boosting enzyme function prediction by fusing protein sequence, structure, and surface
abstract
Proteins can be represented in different data forms, including sequence, structure, and surface, each of which has unique advantages and certain limitations. It is promising to fuse the complementary information among them. In this work, we propose a framework called ProteinF3S for enzyme function prediction that fuses the complementary information across protein sequence, structure, and surface. To achieve more effective fusion, we propose a multi-scale bidirectional fusion strategy between protein structure and surface, in which the hierarchical features of a surface encoder and a structure encoder interact with each other bidirectionally. Based on these interactions, more distinctive features can be obtained. After that, we achieve further fusion by concatenating the sequence features with the features containing structure and surface information, so that better performance can be achieved. To validate our method, we conduct extensive experiments on tasks including enzyme reaction classification and enzyme commission number prediction. Our method achieves new state-of-the-art performance and shows that fusing different forms of data is effective in enzyme function prediction.
Mingzhi Yuan, Yingfan Ma, Bohan An, Manning Wang
Briefings Bioinform.3
2024 Transformer-Based Video-Structure Multi-Instance Learning for Whole Slide Image Classification
abstract
Pathological images play a vital role in clinical cancer diagnosis. Computer-aided diagnosis utilized on digital Whole Slide Images (WSIs) has been widely studied. The major challenge of using deep learning models for WSI analysis is the huge size of WSI images and existing methods struggle between end-to-end learning and proper modeling of contextual information. Most state-of-the-art methods utilize a two-stage strategy, in which they use a pre-trained model to extract features of small patches cut from a WSI and then input these features into a classification model. These methods can not perform end-to-end learning and consider contextual information at the same time. To solve this problem, we propose a framework that models a WSI as a pathologist's observing video and utilizes Transformer to process video clips with a divide-and-conquer strategy, which helps achieve both context-awareness and end-to-end learning. Extensive experiments on three public WSI datasets show that our proposed method outperforms existing SOTA methods in both WSI classification and positive region detection.
Yingfan Ma, Xiaoyuan Luo, Kexue Fu 0001, Manning Wang
AAAI1
2024 Complementary multi-modality molecular self-supervised learning via non-overlapping masking for property prediction
abstract
Self-supervised learning plays an important role in molecular representation learning because labeled molecular data are usually limited in many tasks, such as chemical property prediction and virtual screening. However, most existing molecular pre-training methods focus on one modality of molecular data, and the complementary information of two important modalities, SMILES and graph, is not fully explored. In this study, we propose an effective multi-modality self-supervised learning framework for molecular SMILES and graph. Specifically, SMILES data and graph data are first tokenized so that they can be processed by a unified Transformer-based backbone network, which is trained by a masked reconstruction strategy. In addition, we introduce a specialized non-overlapping masking strategy to encourage fine-grained interaction between these two modalities. Experimental results show that our framework achieves state-of-the-art performance in a series of molecular property prediction tasks, and a detailed ablation study demonstrates efficacy of the multi-modality framework and the masking strategy.
Mingzhi Yuan, Yingfan Ma, Manning Wang
Briefings Bioinform.3
2024 PGBind: pocket-guided explicit attention learning for protein-ligand docking
abstract
As more and more protein structures are discovered, blind protein-ligand docking will play an important role in drug discovery because it can predict protein-ligand complex conformation without pocket information on the target proteins. Recently, deep learning-based methods have made significant advancements in blind protein-ligand docking, but their protein features are suboptimal because they do not fully consider the difference between potential pocket regions and non-pocket regions in protein feature extraction. In this work, we propose a pocket-guided strategy for guiding the ligand to dock to potential docking regions on a protein. To this end, we design a plug-and-play module to enhance the protein features, which can be directly incorporated into existing deep learning-based blind docking methods. The proposed module first estimates potential pocket regions on the target protein and then leverages a pocket-guided attention mechanism to enhance the protein features. Experiments are conducted on integrating our method with EquiBind and FABind, and the results show that their blind-docking performances are both significantly improved and new start-of-the-art performance is achieved by integration with FABind.
Mingzhi Yuan, Yingfan Ma, Manning Wang
Briefings Bioinform.3
2024 SS-Pro: a simplified Siamese contrastive learning approach for protein surface representation
Mingzhi Yuan, Yingfan Ma, Manning Wang
Frontiers Comput. Sci.3
2024 Rethinking Multiple Instance Learning for Whole Slide Image Classification: A Good Instance Classifier Is All You Need
abstract
Weakly supervised whole slide image classification is usually formulated as a multiple instance learning (MIL) problem, where each slide is treated as a bag, and the patches cut out of it are treated as instances. Existing methods either train an instance classifier through pseudo-labeling or aggregate instance features into a bag feature through attention mechanisms and then train a bag classifier, where the attention scores can be used for instance-level classification. However, the pseudo instance labels constructed by the former usually contain a lot of noise, and the attention scores constructed by the latter are not accurate enough, both of which affect their performance. In this paper, we propose an instance-level MIL framework based on contrastive learning and prototype learning to effectively accomplish both instance classification and bag classification tasks. To this end, we propose an instance-level weakly supervised contrastive learning algorithm for the first time under the MIL setting to effectively learn instance feature representation. We also propose an accurate pseudo label generation method through prototype learning. We then develop a joint training strategy for weakly supervised contrastive learning, prototype learning, and instance classifier training. Extensive experiments and visualizations on four datasets demonstrate the powerful performance of our method. Codes will be available.
Linhao Qu, Yingfan Ma, Xiaoyuan Luo, Qinhao Guo, Manning Wang, Zhijian Song
IEEE Trans. Circuits Syst. Video Technol.2
2023 Boosting Whole Slide Image Classification from the Perspectives of Distribution, Correlation and Magnification
abstract
Bag-based multiple instance learning (MIL) methods have become the mainstream for Whole Slide Image (WSI) classification. However, there are still three important issues that have not been fully addressed: (1) positive bags with a low positive instance ratio are prone to the influence of a large number of negative instances; (2) the correlation between local and global features of pathology images has not been fully modeled; and (3) there is a lack of effective information interaction between different magnifications. In this paper, we propose MILBooster, a powerful dual-scale multi-stage MIL framework to address these issues from the perspectives of distribution, correlation, and magnification. Specifically, to address issue (1), we propose a plug-and-play bag filter that effectively increases the positive instance ratio of positive bags. For issue (2), we propose a novel window-based Transformer architecture called PiceBlock to model the correlation between local and global features of pathology images. For issue (3), we propose a dual-branch architecture to process different magnifications and design an information interaction module called Scale Mixer for efficient information interaction between them. We conducted extensive experiments on four clinical WSI classification tasks using three datasets. MILBooster achieved new state-of-the-art performance on all these tasks. Codes will be available at https://github.com/miccaiif/MILBooster.
Linhao Qu, Minghong Duan, Yingfan Ma, Shuo Wang 0011, Manning Wang, Zhijian Song
ICCV4
2023 OpenAL: An Efficient Deep Active Learning Framework for Open-Set Pathology Image Classification
Linhao Qu, Yingfan Ma, Manning Wang, Zhijian Song
MICCAI (2)2
2023 ProteinMAE: masked autoencoder for protein surface self-supervised learning
abstract
SUMMARY: The biological functions of proteins are determined by the chemical and geometric properties of their surfaces. Recently, with the booming progress of deep learning, a series of learning-based surface descriptors have been proposed and achieved inspirational performance in many tasks such as protein design, protein-protein interaction prediction, etc. However, they are still limited by the problem of label scarcity, since the labels are typically obtained through wet experiments. Inspired by the great success of self-supervised learning in natural language processing and computer vision, we introduce ProteinMAE, a self-supervised framework specifically designed for protein surface representation to mitigate label scarcity. Specifically, we propose an efficient network and utilize a large number of accessible unlabeled protein data to pretrain it by self-supervised learning. Then we use the pretrained weights as initialization and fine-tune the network on downstream tasks. To demonstrate the effectiveness of our method, we conduct experiments on three different downstream tasks including binding site identification in protein surface, ligand-binding protein pocket classification, and protein-protein interaction prediction. The extensive experiments show that our method not only successfully improves the network's performance on all downstream tasks, but also achieves competitive performance with state-of-the-art methods. Moreover, our proposed network also exhibits significant advantages in terms of computational cost, which only requires less than a tenth of memory cost of previous methods. AVAILABILITY AND IMPLEMENTATION: https://github.com/phdymz/ProteinMAE.
Mingzhi Yuan, Kexue Fu 0001, Jiaming Guan, Yingfan Ma, Qin Qiao, Manning Wang
Bioinform.5
2021 Mining frequent pyramid patterns from time series transaction data with custom constraints
Guodong Xin, Yingfan Ma, Bailing Wang
Comput. Secur.5