EDBT 2026 Demo / reviewers in the wild / expert
Mingzhi Yuan
dblp:315/2046
· DBLP profile ↗
19ranked-venue papers
8as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Protein Structure Learning Using a Size-Guided Conditional Mixture-of-ExpertsabstractIn recent years, deep learning on protein structures has attracted widespread attention, as structures determine proteins' function. A series of structure-based protein property prediction methods have been proposed, achieving remarkable performance. However, these methods often neglect the importance of the protein size and fail to fully leverage it, leading to biases toward certain sizes and suboptimal overall performance. To address this issue, we propose a protein size-guided conditional mixture-of-experts for improving deep learning on protein structures. It can adaptively activate the sub-networks with the guidance of protein sizes and network features. Its flexible combinations of sub-networks help mitigate biases toward certain protein sizes, while the deliberate incorporation of protein size guidance enables the network to effectively capture both universal and size-specific characteristics, resulting in more accurate predictive performance. Based on it, we propose a framework for protein property prediction and benchmark it on eight tasks with two representation forms of proteins and three different dataset splits, a total of forty-eight tests. Experiments show that our method can be seamlessly integrated into numerous existing models and achieve performance improvement across tasks under almost all settings. More importantly, our experiments reveal that although often overlooked, protein size serves as an important prior knowledge in deep learning on protein structures. Mingzhi Yuan, Siqi Yin, Yingfan Ma, Manning Wang |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Dual Focus-Attention Transformer for Robust Point Cloud RegistrationabstractRecently, coarse-to-fine methods for point cloud registration have achieved great success, but few works deeply explore the impact of feature interaction at both coarse and fine scales. By visualizing attention scores and correspondences, we find that existing methods fail to achieve effective feature aggregation at the two scales during the feature interaction. To tackle this issue, we propose a Dual Focus-Attention Transformer framework, which only focuses on points relevant to the current point for feature interaction, avoiding interactions with irrelevant points. For the coarse scale, we design a superpoint focus-attention transformer guided by sparse keypoints, which are selected from the neighborhood of superpoints. For the fine scale, we only perform feature interaction between the point sets that belong to the same superpoint. Experiments show that our method achieve the state-of-the-art performance on three standard benchmarks. The code and pre-trained models are available at https://github.com/fukexue/DFAT.git. Kexue Fu 0001, Mingzhi Yuan, Changwei Wang 0001, Weiguang Pang, Jing Chi, Manning Wang, Longxiang Gao |
CVPR | 2 |
| 2025 | Flow-MIL: Constructing Highly-expressive Latent Feature Space for Whole Slide Image Classification using Normalizing Flow
Yingfan Ma, Bohan An, Mingzhi Yuan, Minghong Duan, Manning Wang |
ICCV | 4 |
| 2025 | Drug-TTA: Test-Time Adaptation for Drug Virtual Screening via Multi-task Meta-Auxiliary LearningabstractVirtual screening is a critical step in drug discovery, aiming at identifying potential drugs that bind to a specific protein pocket from a large database of molecules. Traditional docking methods are time-consuming, while learning-based approaches supervised by high-precision conformational or affinity labels are limited by the scarcity of training data. Recently, a paradigm of feature alignment through contrastive learning has gained widespread attention. This method does not require explicit binding affinity scores, but it suffers from the issue of overly simplistic construction of negative samples, which limits their generalization to more difficult test cases. In this paper, we propose Drug-TTA, which leverages a large number of self-supervised auxiliary tasks to adapt the model to each test instance. Specifically, we incorporate the auxiliary tasks into both the training and the inference process via meta-learning to improve the performance of the primary task of virtual screening. Additionally, we design a multi-scale feature based Auxiliary Loss Balance Module (ALBM) to balance the auxiliary tasks to improve their efficiency. Extensive experiments demonstrate that Drug-TTA achieves state-of-the-art (SOTA) performance in all five virtual screening tasks under a zero-shot setting, showing an average improvement of 9.86% in AUROC metric compared to the baseline without test-time adaptation. Mingzhi Yuan, Yingfan Ma, Manning Wang |
ICML | 2 |
| 2025 | ProteinF3S: boosting enzyme function prediction by fusing protein sequence, structure, and surfaceabstractProteins can be represented in different data forms, including sequence, structure, and surface, each of which has unique advantages and certain limitations. It is promising to fuse the complementary information among them. In this work, we propose a framework called ProteinF3S for enzyme function prediction that fuses the complementary information across protein sequence, structure, and surface. To achieve more effective fusion, we propose a multi-scale bidirectional fusion strategy between protein structure and surface, in which the hierarchical features of a surface encoder and a structure encoder interact with each other bidirectionally. Based on these interactions, more distinctive features can be obtained. After that, we achieve further fusion by concatenating the sequence features with the features containing structure and surface information, so that better performance can be achieved. To validate our method, we conduct extensive experiments on tasks including enzyme reaction classification and enzyme commission number prediction. Our method achieves new state-of-the-art performance and shows that fusing different forms of data is effective in enzyme function prediction. Mingzhi Yuan, Yingfan Ma, Bohan An, Manning Wang |
Briefings Bioinform. | 1 |
| 2024 | Complementary multi-modality molecular self-supervised learning via non-overlapping masking for property predictionabstractSelf-supervised learning plays an important role in molecular representation learning because labeled molecular data are usually limited in many tasks, such as chemical property prediction and virtual screening. However, most existing molecular pre-training methods focus on one modality of molecular data, and the complementary information of two important modalities, SMILES and graph, is not fully explored. In this study, we propose an effective multi-modality self-supervised learning framework for molecular SMILES and graph. Specifically, SMILES data and graph data are first tokenized so that they can be processed by a unified Transformer-based backbone network, which is trained by a masked reconstruction strategy. In addition, we introduce a specialized non-overlapping masking strategy to encourage fine-grained interaction between these two modalities. Experimental results show that our framework achieves state-of-the-art performance in a series of molecular property prediction tasks, and a detailed ablation study demonstrates efficacy of the multi-modality framework and the masking strategy. Mingzhi Yuan, Yingfan Ma, Manning Wang |
Briefings Bioinform. | 2 |
| 2024 | PGBind: pocket-guided explicit attention learning for protein-ligand dockingabstractAs more and more protein structures are discovered, blind protein-ligand docking will play an important role in drug discovery because it can predict protein-ligand complex conformation without pocket information on the target proteins. Recently, deep learning-based methods have made significant advancements in blind protein-ligand docking, but their protein features are suboptimal because they do not fully consider the difference between potential pocket regions and non-pocket regions in protein feature extraction. In this work, we propose a pocket-guided strategy for guiding the ligand to dock to potential docking regions on a protein. To this end, we design a plug-and-play module to enhance the protein features, which can be directly incorporated into existing deep learning-based blind docking methods. The proposed module first estimates potential pocket regions on the target protein and then leverages a pocket-guided attention mechanism to enhance the protein features. Experiments are conducted on integrating our method with EquiBind and FABind, and the results show that their blind-docking performances are both significantly improved and new start-of-the-art performance is achieved by integration with FABind. Mingzhi Yuan, Yingfan Ma, Manning Wang |
Briefings Bioinform. | 2 |
| 2024 | SS-Pro: a simplified Siamese contrastive learning approach for protein surface representation
Mingzhi Yuan, Yingfan Ma, Manning Wang |
Frontiers Comput. Sci. | 2 |
| 2024 | Decoupled deep hough voting for point cloud registration
Mingzhi Yuan, Kexue Fu 0001, Manning Wang |
Frontiers Comput. Sci. | 1 |
| 2024 | Boosting Point-BERT by Multi-Choice TokensabstractMasked language modeling (MLM) has become one of the most successful self-supervised pre-training task. Inspired by its success, Point-BERT, as a pioneer work in point cloud, proposed masked point modeling (MPM) to pre-train point transformer on large scale unanotated dataset. Despite its great performance, we find the inherent difference between language and point cloud tends to cause ambiguous tokenization for point cloud, and no gold standard is available for point cloud tokenization. Point-BERT uses a discrete Variational AutoEncoder (dVAE) as tokenizer, but it might generate different token ids for semantically-similar patches and the same token ids for semantically-dissimilar patches. To tackle the above problems, we propose our McP-BERT, a pre-training framework with multi-choice tokens. Specifically, we ease the previous single-choice constraint on patch token ids in Point-BERT, and provide multi-choice token ids for each patch as supervision. Moreover, we utilitze the high-level semantics learned by transformer to further refine our supervision signals. Extensive experiments on point cloud classification, few-shot classification and part segmentation tasks demonstrate the superiority of our method, e.g., the pre-trained transformer achieves 94.1% accuracy on ModelNet40, 84.28% accuracy on the hardest setting of ScanObjectNN and new state-of-the-art performance on few-shot learning. Our method improves the performance of Point-BERT on all downstream tasks without extra computational overhead. Kexue Fu 0001, Mingzhi Yuan, Shaolei Liu, Manning Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Robust Point Cloud Registration via Random Network Co-EnsembleabstractLearning-based point cloud registration has achieved great success in recent years but is still limited by its generalization. The performance of these methods declines when they are extended to unseen datasets that have inconsistent distributions with the training set. In this paper, we propose a novel random network-based method, which does not require training. Our approach utilizes multiple randomly initialized networks for feature extraction and correspondence building. Furthermore, we also introduce a co-ensemble strategy to prune the outliers in correspondences built upon random networks, which leverages spatial consistency. Through our co-ensemble pruning, a large proportion of outliers can be removed, thereby achieving robust registration in affordable RANSAC iterations. Extensive experiments on 3DMatch and KITTI demonstrate that our method outperforms not only the traditional methods but also the learning-based methods trained on datasets inconsistent with the test set. The code will be released at https://github.com/phdymz/RandPCR. Mingzhi Yuan, Kexue Fu 0001, Yucong Meng, Manning Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | PointMBF: A Multi-scale Bidirectional Fusion Network for Unsupervised RGB-D Point Cloud RegistrationabstractPoint cloud registration is a task to estimate the rigid transformation between two unaligned scans, which plays an important role in many computer vision applications. Previous learning-based works commonly focus on supervised registration, which have limitations in practice. Recently, with the advance of inexpensive RGB-D sensors, several learning-based works utilize RGB-D data to achieve unsupervised registration. However, most of existing unsupervised methods follow a cascaded design or fuse RGB-D data in a unidirectional manner, which do not fully exploit the complementary information in the RGB-D data. To leverage the complementary information more effectively, we propose a network implementing multi-scale bidirectional fusion between RGB images and point clouds generated from depth images. By bidirectionally fusing visual and geometric features in multi-scales, more distinctive deep features for correspondence estimation can be obtained, making our registration more accurate. Extensive experiments on ScanNet and 3DMatch demonstrate that our method achieves new state-of-the-art performance. Code will be released at https://github.com/phdymz/PointMBF. Mingzhi Yuan, Kexue Fu 0001, Yucong Meng, Manning Wang |
ICCV | 1 |
| 2023 | Boosting 3D Point Cloud Registration by Transferring Multi-modality KnowledgeabstractThe recent multi-modality models have achieved great performance in many vision tasks because the extracted features contain the multi-modality knowledge. However, most of the current registration descriptors have only concentrated on local geometric structures. This paper proposes a method to boost point cloud registration accuracy by transferring the multi-modality knowledge of pre-trained multi-modality model to a new descriptor neural network. Different to the previous multi-modality methods that requires both modalities, the proposed method only requires point clouds during inference. Specifically, we propose an ensemble descriptor neural network combining pre-trained sparse convolution branch and a new point-based convolution branch. By fine-tuning on a single modality data, the proposed method achieves new state-of-the-art results on 3DMatch and competitive accuracy on 3DLoMatch and KITTI. The code and the trained model will be released at https://github.com/phdymz/DBENet.git. Mingzhi Yuan, Xiaoshui Huang, Kexue Fu 0001, Manning Wang |
ICRA | 1 |
| 2023 | ProteinMAE: masked autoencoder for protein surface self-supervised learningabstractSUMMARY: The biological functions of proteins are determined by the chemical and geometric properties of their surfaces. Recently, with the booming progress of deep learning, a series of learning-based surface descriptors have been proposed and achieved inspirational performance in many tasks such as protein design, protein-protein interaction prediction, etc. However, they are still limited by the problem of label scarcity, since the labels are typically obtained through wet experiments. Inspired by the great success of self-supervised learning in natural language processing and computer vision, we introduce ProteinMAE, a self-supervised framework specifically designed for protein surface representation to mitigate label scarcity. Specifically, we propose an efficient network and utilize a large number of accessible unlabeled protein data to pretrain it by self-supervised learning. Then we use the pretrained weights as initialization and fine-tune the network on downstream tasks. To demonstrate the effectiveness of our method, we conduct experiments on three different downstream tasks including binding site identification in protein surface, ligand-binding protein pocket classification, and protein-protein interaction prediction. The extensive experiments show that our method not only successfully improves the network's performance on all downstream tasks, but also achieves competitive performance with state-of-the-art methods. Moreover, our proposed network also exhibits significant advantages in terms of computational cost, which only requires less than a tenth of memory cost of previous methods. AVAILABILITY AND IMPLEMENTATION: https://github.com/phdymz/ProteinMAE. Mingzhi Yuan, Kexue Fu 0001, Jiaming Guan, Yingfan Ma, Qin Qiao, Manning Wang |
Bioinform. | 1 |
| 2023 | Density-based one-shot active learning for image segmentation
Qiuye Jin, Shiman Li, Xiaofei Du 0002, Mingzhi Yuan, Manning Wang, Zhijian Song |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | PointCLM: A Contrastive Learning-based Framework for Multi-instance Point Cloud Registration
Mingzhi Yuan, Qiuye Jin, Xinrong Chen, Manning Wang |
ECCV (9) | 1 |
| 2022 | Cold-start active learning for image classification
Qiuye Jin, Mingzhi Yuan, Shiman Li, Haoran Wang 0009, Manning Wang, Zhijian Song |
Inf. Sci. | 2 |
| 2022 | One-shot active learning for image segmentation via contrastive learning and diversity-based sampling
Qiuye Jin, Mingzhi Yuan, Qin Qiao, Zhijian Song |
Knowl. Based Syst. | 2 |
| 2022 | Deep active learning models for imbalanced image classification
Qiuye Jin, Mingzhi Yuan, Haoran Wang 0009, Manning Wang, Zhijian Song |
Knowl. Based Syst. | 2 |