VLDB 2026 Research / reviewers in the wild / expert
Manning Wang
dblp:23/5931
· DBLP profile ↗
55ranked-venue papers
1as first author
50since 2021 · last 2026
0000-0002-9255-3897ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 1 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 18 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 16 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Protein Structure Learning Using a Size-Guided Conditional Mixture-of-ExpertsabstractIn recent years, deep learning on protein structures has attracted widespread attention, as structures determine proteins' function. A series of structure-based protein property prediction methods have been proposed, achieving remarkable performance. However, these methods often neglect the importance of the protein size and fail to fully leverage it, leading to biases toward certain sizes and suboptimal overall performance. To address this issue, we propose a protein size-guided conditional mixture-of-experts for improving deep learning on protein structures. It can adaptively activate the sub-networks with the guidance of protein sizes and network features. Its flexible combinations of sub-networks help mitigate biases toward certain protein sizes, while the deliberate incorporation of protein size guidance enables the network to effectively capture both universal and size-specific characteristics, resulting in more accurate predictive performance. Based on it, we propose a framework for protein property prediction and benchmark it on eight tasks with two representation forms of proteins and three different dataset splits, a total of forty-eight tests. Experiments show that our method can be seamlessly integrated into numerous existing models and achieve performance improvement across tasks under almost all settings. More importantly, our experiments reveal that although often overlooked, protein size serves as an important prior knowledge in deep learning on protein structures. Mingzhi Yuan, Siqi Yin, Yingfan Ma, Manning Wang |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Dual Focus-Attention Transformer for Robust Point Cloud RegistrationabstractRecently, coarse-to-fine methods for point cloud registration have achieved great success, but few works deeply explore the impact of feature interaction at both coarse and fine scales. By visualizing attention scores and correspondences, we find that existing methods fail to achieve effective feature aggregation at the two scales during the feature interaction. To tackle this issue, we propose a Dual Focus-Attention Transformer framework, which only focuses on points relevant to the current point for feature interaction, avoiding interactions with irrelevant points. For the coarse scale, we design a superpoint focus-attention transformer guided by sparse keypoints, which are selected from the neighborhood of superpoints. For the fine scale, we only perform feature interaction between the point sets that belong to the same superpoint. Experiments show that our method achieve the state-of-the-art performance on three standard benchmarks. The code and pre-trained models are available at https://github.com/fukexue/DFAT.git. Kexue Fu 0001, Mingzhi Yuan, Changwei Wang 0001, Weiguang Pang, Jing Chi, Manning Wang, Longxiang Gao |
CVPR | 6 |
| 2025 | Flow-MIL: Constructing Highly-expressive Latent Feature Space for Whole Slide Image Classification using Normalizing Flow
Yingfan Ma, Bohan An, Mingzhi Yuan, Minghong Duan, Manning Wang |
ICCV | 6 |
| 2025 | Drug-TTA: Test-Time Adaptation for Drug Virtual Screening via Multi-task Meta-Auxiliary LearningabstractVirtual screening is a critical step in drug discovery, aiming at identifying potential drugs that bind to a specific protein pocket from a large database of molecules. Traditional docking methods are time-consuming, while learning-based approaches supervised by high-precision conformational or affinity labels are limited by the scarcity of training data. Recently, a paradigm of feature alignment through contrastive learning has gained widespread attention. This method does not require explicit binding affinity scores, but it suffers from the issue of overly simplistic construction of negative samples, which limits their generalization to more difficult test cases. In this paper, we propose Drug-TTA, which leverages a large number of self-supervised auxiliary tasks to adapt the model to each test instance. Specifically, we incorporate the auxiliary tasks into both the training and the inference process via meta-learning to improve the performance of the primary task of virtual screening. Additionally, we design a multi-scale feature based Auxiliary Loss Balance Module (ALBM) to balance the auxiliary tasks to improve their efficiency. Extensive experiments demonstrate that Drug-TTA achieves state-of-the-art (SOTA) performance in all five virtual screening tasks under a zero-shot setting, showing an average improvement of 9.86% in AUROC metric compared to the baseline without test-time adaptation. Mingzhi Yuan, Yingfan Ma, Manning Wang |
ICML | 6 |
| 2025 | Vector-Quantization-Driven Active Learning for Efficient Multi-modal Medical Segmentation with Cross-Modal Assistance
Xiaofei Du 0002, Haoran Wang 0009, Manning Wang, Zhijian Song |
MICCAI (7) | 3 |
| 2025 | Knowledge-Guided Multi-scale Graph Mamba for Whole Slide Image Classification
Minghong Duan, Yingfan Ma, Manning Wang, Zhijian Song |
MICCAI (12) | 4 |
| 2025 | ProteinF3S: boosting enzyme function prediction by fusing protein sequence, structure, and surfaceabstractProteins can be represented in different data forms, including sequence, structure, and surface, each of which has unique advantages and certain limitations. It is promising to fuse the complementary information among them. In this work, we propose a framework called ProteinF3S for enzyme function prediction that fuses the complementary information across protein sequence, structure, and surface. To achieve more effective fusion, we propose a multi-scale bidirectional fusion strategy between protein structure and surface, in which the hierarchical features of a surface encoder and a structure encoder interact with each other bidirectionally. Based on these interactions, more distinctive features can be obtained. After that, we achieve further fusion by concatenating the sequence features with the features containing structure and surface information, so that better performance can be achieved. To validate our method, we conduct extensive experiments on tasks including enzyme reaction classification and enzyme commission number prediction. Our method achieves new state-of-the-art performance and shows that fusing different forms of data is effective in enzyme function prediction. Mingzhi Yuan, Yingfan Ma, Bohan An, Manning Wang |
Briefings Bioinform. | 6 |
| 2025 | DDFP: Data-dependent frequency prompt for source free domain adaptation of medical image segmentation
Siqi Yin, Shaolei Liu, Manning Wang |
Knowl. Based Syst. | 3 |
| 2024 | Transformer-Based Video-Structure Multi-Instance Learning for Whole Slide Image ClassificationabstractPathological images play a vital role in clinical cancer diagnosis. Computer-aided diagnosis utilized on digital Whole Slide Images (WSIs) has been widely studied. The major challenge of using deep learning models for WSI analysis is the huge size of WSI images and existing methods struggle between end-to-end learning and proper modeling of contextual information. Most state-of-the-art methods utilize a two-stage strategy, in which they use a pre-trained model to extract features of small patches cut from a WSI and then input these features into a classification model. These methods can not perform end-to-end learning and consider contextual information at the same time. To solve this problem, we propose a framework that models a WSI as a pathologist's observing video and utilizes Transformer to process video clips with a divide-and-conquer strategy, which helps achieve both context-awareness and end-to-end learning. Extensive experiments on three public WSI datasets show that our proposed method outperforms existing SOTA methods in both WSI classification and positive region detection. Yingfan Ma, Xiaoyuan Luo, Kexue Fu 0001, Manning Wang |
AAAI | 4 |
| 2024 | Local Implicit Wavelet Transformer for Arbitrary-Scale Super-Resolution
Minghong Duan, Linhao Qu, Shaolei Liu, Manning Wang |
BMVC | 4 |
| 2024 | FAST: A Dual-tier Few-Shot Learning Paradigm for Whole Slide Image ClassificationabstractThe expensive fine-grained annotation and data scarcity
have become the primary obstacles for the widespread adoption of deep learning-based Whole Slide Images (WSI) classification algorithms in clinical practice. Unlike few-shot learning methods in natural images that can leverage the labels of each image, existing few-shot WSI classification methods only utilize a small number of fine-grained labels or weakly supervised slide labels for training in order to avoid expensive fine-grained annotation. They lack sufficient mining of available WSIs, severely limiting WSI classification performance. To address the above issues, we propose a novel and efficient dual-tier few-shot learning paradigm for WSI classification, named FAST. FAST consists of a dual-level annotation strategy and a dual-branch classification framework. Firstly, to avoid expensive fine-grained annotation, we collect a very small number of WSIs at the slide level, and annotate an extremely small number of patches. Then, to fully mining the available WSIs, we use all the patches and available patch labels to build a cache branch, which utilizes the labeled patches to learn the labels of unlabeled patches and through knowledge retrieval for patch classification. In addition to the cache branch, we also construct a prior branch that includes learnable prompt vectors, using the text encoder of visual-language models for patch classification. Finally, we integrate the results from both branches to achieve WSI classification. Extensive experiments on binary and multi-class datasets demonstrate that our proposed method significantly surpasses existing few-shot classification methods and approaches the accuracy of fully supervised methods with only 0.22% annotation costs. All codes and models will be publicly available on https://github.com/fukexue/FAST. Kexue Fu 0001, Xiaoyuan Luo, Linhao Qu, Shuo Wang 0011, Ilias Maglogiannis, Longxiang Gao, Manning Wang |
NeurIPS | 8 |
| 2024 | Complementary multi-modality molecular self-supervised learning via non-overlapping masking for property predictionabstractSelf-supervised learning plays an important role in molecular representation learning because labeled molecular data are usually limited in many tasks, such as chemical property prediction and virtual screening. However, most existing molecular pre-training methods focus on one modality of molecular data, and the complementary information of two important modalities, SMILES and graph, is not fully explored. In this study, we propose an effective multi-modality self-supervised learning framework for molecular SMILES and graph. Specifically, SMILES data and graph data are first tokenized so that they can be processed by a unified Transformer-based backbone network, which is trained by a masked reconstruction strategy. In addition, we introduce a specialized non-overlapping masking strategy to encourage fine-grained interaction between these two modalities. Experimental results show that our framework achieves state-of-the-art performance in a series of molecular property prediction tasks, and a detailed ablation study demonstrates efficacy of the multi-modality framework and the masking strategy. Mingzhi Yuan, Yingfan Ma, Manning Wang |
Briefings Bioinform. | 5 |
| 2024 | PGBind: pocket-guided explicit attention learning for protein-ligand dockingabstractAs more and more protein structures are discovered, blind protein-ligand docking will play an important role in drug discovery because it can predict protein-ligand complex conformation without pocket information on the target proteins. Recently, deep learning-based methods have made significant advancements in blind protein-ligand docking, but their protein features are suboptimal because they do not fully consider the difference between potential pocket regions and non-pocket regions in protein feature extraction. In this work, we propose a pocket-guided strategy for guiding the ligand to dock to potential docking regions on a protein. To this end, we design a plug-and-play module to enhance the protein features, which can be directly incorporated into existing deep learning-based blind docking methods. The proposed module first estimates potential pocket regions on the target protein and then leverages a pocket-guided attention mechanism to enhance the protein features. Experiments are conducted on integrating our method with EquiBind and FABind, and the results show that their blind-docking performances are both significantly improved and new start-of-the-art performance is achieved by integration with FABind. Mingzhi Yuan, Yingfan Ma, Manning Wang |
Briefings Bioinform. | 5 |
| 2024 | POS-BERT: Point cloud one-stage BERT pre-training
Kexue Fu 0001, Peng Gao 0007, Shaolei Liu, Linhao Qu, Longxiang Gao, Manning Wang |
Expert Syst. Appl. | 6 |
| 2024 | Trans2Fuse: Empowering image fusion through self-supervised learning and multi-modal transformations via transformer networks
Linhao Qu, Shaolei Liu, Manning Wang, Shiman Li, Siqi Yin, Zhijian Song |
Expert Syst. Appl. | 3 |
| 2024 | SS-Pro: a simplified Siamese contrastive learning approach for protein surface representation
Mingzhi Yuan, Yingfan Ma, Manning Wang |
Frontiers Comput. Sci. | 4 |
| 2024 | Decoupled deep hough voting for point cloud registration
Mingzhi Yuan, Kexue Fu 0001, Manning Wang |
Frontiers Comput. Sci. | 4 |
| 2024 | A comprehensive survey on deep active learning in medical image analysis
Haoran Wang 0009, Qiuye Jin, Shiman Li, Manning Wang, Zhijian Song |
Medical Image Anal. | 5 |
| 2024 | Wavelet-based spectrum transfer with collaborative learning for unsupervised bidirectional cross-modality domain adaptation on medical image segmentation
Shaolei Liu, Linhao Qu, Siqi Yin, Manning Wang, Zhijian Song |
Neural Comput. Appl. | 4 |
| 2024 | Boosting Point-BERT by Multi-Choice TokensabstractMasked language modeling (MLM) has become one of the most successful self-supervised pre-training task. Inspired by its success, Point-BERT, as a pioneer work in point cloud, proposed masked point modeling (MPM) to pre-train point transformer on large scale unanotated dataset. Despite its great performance, we find the inherent difference between language and point cloud tends to cause ambiguous tokenization for point cloud, and no gold standard is available for point cloud tokenization. Point-BERT uses a discrete Variational AutoEncoder (dVAE) as tokenizer, but it might generate different token ids for semantically-similar patches and the same token ids for semantically-dissimilar patches. To tackle the above problems, we propose our McP-BERT, a pre-training framework with multi-choice tokens. Specifically, we ease the previous single-choice constraint on patch token ids in Point-BERT, and provide multi-choice token ids for each patch as supervision. Moreover, we utilitze the high-level semantics learned by transformer to further refine our supervision signals. Extensive experiments on point cloud classification, few-shot classification and part segmentation tasks demonstrate the superiority of our method, e.g., the pre-trained transformer achieves 94.1% accuracy on ModelNet40, 84.28% accuracy on the hardest setting of ScanObjectNN and new state-of-the-art performance on few-shot learning. Our method improves the performance of Point-BERT on all downstream tasks without extra computational overhead. Kexue Fu 0001, Mingzhi Yuan, Shaolei Liu, Manning Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Rethinking Multiple Instance Learning for Whole Slide Image Classification: A Good Instance Classifier Is All You NeedabstractWeakly supervised whole slide image classification is usually formulated as a multiple instance learning (MIL) problem, where each slide is treated as a bag, and the patches cut out of it are treated as instances. Existing methods either train an instance classifier through pseudo-labeling or aggregate instance features into a bag feature through attention mechanisms and then train a bag classifier, where the attention scores can be used for instance-level classification. However, the pseudo instance labels constructed by the former usually contain a lot of noise, and the attention scores constructed by the latter are not accurate enough, both of which affect their performance. In this paper, we propose an instance-level MIL framework based on contrastive learning and prototype learning to effectively accomplish both instance classification and bag classification tasks. To this end, we propose an instance-level weakly supervised contrastive learning algorithm for the first time under the MIL setting to effectively learn instance feature representation. We also propose an accurate pseudo label generation method through prototype learning. We then develop a joint training strategy for weakly supervised contrastive learning, prototype learning, and instance classifier training. Extensive experiments and visualizations on four datasets demonstrate the powerful performance of our method. Codes will be available. Linhao Qu, Yingfan Ma, Xiaoyuan Luo, Qinhao Guo, Manning Wang, Zhijian Song |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Robust Point Cloud Registration via Random Network Co-EnsembleabstractLearning-based point cloud registration has achieved great success in recent years but is still limited by its generalization. The performance of these methods declines when they are extended to unseen datasets that have inconsistent distributions with the training set. In this paper, we propose a novel random network-based method, which does not require training. Our approach utilizes multiple randomly initialized networks for feature extraction and correspondence building. Furthermore, we also introduce a co-ensemble strategy to prune the outliers in correspondences built upon random networks, which leverages spatial consistency. Through our co-ensemble pruning, a large proportion of outliers can be removed, thereby achieving robust registration in affordable RANSAC iterations. Extensive experiments on 3DMatch and KITTI demonstrate that our method outperforms not only the traditional methods but also the learning-based methods trained on datasets inconsistent with the test set. The code will be released at https://github.com/phdymz/RandPCR. Mingzhi Yuan, Kexue Fu 0001, Yucong Meng, Manning Wang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Negative Instance Guided Self-Distillation Framework for Whole Slide Image AnalysisabstractHistopathology image classification is an important clinical task, and current deep learning-based whole-slide image (WSI) classification methods typically cut WSIs into small patches and cast the problem as multi-instance learning. The mainstream approach is to train a bag-level classifier, but their performance on both slide classification and positive patch localization is limited because the instance-level information is not fully explored. In this article, we propose a negative instance-guided, self-distillation framework to directly train an instance-level classifier end-to-end. Instead of depending only on the self-supervised training of the teacher and the student classifiers in a typical self-distillation framework, we input the true negative instances into the student classifier to guide the classifier to better distinguish positive and negative instances. In addition, we propose a prediction bank to constrain the distribution of pseudo instance labels generated by the teacher classifier to prevent the self-distillation from falling into the degeneration of classifying all instances as negative. We conduct extensive experiments and analysis on three publicly available pathological datasets: CAMELYON16, PANDA, and TCGA, as well as an in-house pathological dataset for cervical cancer lymph node metastasis prediction. The results show that our method outperforms existing methods by a large margin. Code will be publicly available. Xiaoyuan Luo, Linhao Qu, Qinhao Guo, Zhijian Song, Manning Wang |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Reducing Domain Gap in Frequency and Spatial Domain for Cross-Modality Domain Adaptation on Medical Image SegmentationabstractUnsupervised domain adaptation (UDA) aims to learn a model trained on source domain and performs well on unlabeled target domain. In medical image segmentation field, most existing UDA methods depend on adversarial learning to address the domain gap between different image modalities, which is ineffective due to its complicated training process. In this paper, we propose a simple yet effective UDA method based on frequency and spatial domain transfer under multi-teacher distillation framework. In the frequency domain, we first introduce non-subsampled contourlet transform for identifying domain-invariant and domain-variant frequency components (DIFs and DVFs), and then keep the DIFs unchanged while replacing the DVFs of the source domain images with that of the target domain images to narrow the domain gap. In the spatial domain, we propose a batch momentum update-based histogram matching strategy to reduce the domain-variant image style bias. Experiments on two commonly used cross-modality medical image segmentation datasets show that our proposed method achieves superior performance compared to state-of-the-art methods. Shaolei Liu, Siqi Yin, Linhao Qu, Manning Wang |
AAAI | 4 |
| 2023 | Boosting Whole Slide Image Classification from the Perspectives of Distribution, Correlation and MagnificationabstractBag-based multiple instance learning (MIL) methods have become the mainstream for Whole Slide Image (WSI) classification. However, there are still three important issues that have not been fully addressed: (1) positive bags with a low positive instance ratio are prone to the influence of a large number of negative instances; (2) the correlation between local and global features of pathology images has not been fully modeled; and (3) there is a lack of effective information interaction between different magnifications. In this paper, we propose MILBooster, a powerful dual-scale multi-stage MIL framework to address these issues from the perspectives of distribution, correlation, and magnification. Specifically, to address issue (1), we propose a plug-and-play bag filter that effectively increases the positive instance ratio of positive bags. For issue (2), we propose a novel window-based Transformer architecture called PiceBlock to model the correlation between local and global features of pathology images. For issue (3), we propose a dual-branch architecture to process different magnifications and design an information interaction module called Scale Mixer for efficient information interaction between them. We conducted extensive experiments on four clinical WSI classification tasks using three datasets. MILBooster achieved new state-of-the-art performance on all these tasks. Codes will be available at https://github.com/miccaiif/MILBooster. Linhao Qu, Minghong Duan, Yingfan Ma, Shuo Wang 0011, Manning Wang, Zhijian Song |
ICCV | 6 |
| 2023 | PointMBF: A Multi-scale Bidirectional Fusion Network for Unsupervised RGB-D Point Cloud RegistrationabstractPoint cloud registration is a task to estimate the rigid transformation between two unaligned scans, which plays an important role in many computer vision applications. Previous learning-based works commonly focus on supervised registration, which have limitations in practice. Recently, with the advance of inexpensive RGB-D sensors, several learning-based works utilize RGB-D data to achieve unsupervised registration. However, most of existing unsupervised methods follow a cascaded design or fuse RGB-D data in a unidirectional manner, which do not fully exploit the complementary information in the RGB-D data. To leverage the complementary information more effectively, we propose a network implementing multi-scale bidirectional fusion between RGB images and point clouds generated from depth images. By bidirectionally fusing visual and geometric features in multi-scales, more distinctive deep features for correspondence estimation can be obtained, making our registration more accurate. Extensive experiments on ScanNet and 3DMatch demonstrate that our method achieves new state-of-the-art performance. Code will be released at https://github.com/phdymz/PointMBF. Mingzhi Yuan, Kexue Fu 0001, Yucong Meng, Manning Wang |
ICCV | 5 |
| 2023 | Boosting 3D Point Cloud Registration by Transferring Multi-modality KnowledgeabstractThe recent multi-modality models have achieved great performance in many vision tasks because the extracted features contain the multi-modality knowledge. However, most of the current registration descriptors have only concentrated on local geometric structures. This paper proposes a method to boost point cloud registration accuracy by transferring the multi-modality knowledge of pre-trained multi-modality model to a new descriptor neural network. Different to the previous multi-modality methods that requires both modalities, the proposed method only requires point clouds during inference. Specifically, we propose an ensemble descriptor neural network combining pre-trained sparse convolution branch and a new point-based convolution branch. By fine-tuning on a single modality data, the proposed method achieves new state-of-the-art results on 3DMatch and competitive accuracy on 3DLoMatch and KITTI. The code and the trained model will be released at https://github.com/phdymz/DBENet.git. Mingzhi Yuan, Xiaoshui Huang, Kexue Fu 0001, Manning Wang |
ICRA | 5 |
| 2023 | OpenAL: An Efficient Deep Active Learning Framework for Open-Set Pathology Image Classification
Linhao Qu, Yingfan Ma, Manning Wang, Zhijian Song |
MICCAI (2) | 4 |
| 2023 | PI-NeRF: A Partial-Invertible Neural Radiance Fields for Pose EstimationabstractIn recent years, Neural Radiance Fields (NeRF) have been used as a map of 3D scene to estimate the 6-DoF pose of new observed images - given an image, estimate the relative rotation and translation of a camera using a trained NeRF. However, existing NeRF-based pose estimation methods have a small convergence region and need to be optimized iteratively over a given initial pose, which makes them slow and sensitive to the initial pose. In this paper, we propose PI-NeRF that directly outputs the pose of a given image without pose initialization and iterative optimization. This is achieved by integrating NeRF with invertible neural network (INN). Our method employs INNs to establish a bijective mapping between the rays and pixel features, which allows us to directly estimate the ray corresponding to each image pixel using the feature map extracted by an image encoder. Based on these rays, we can directly estimate the pose of the image using the PnP algorithm. Experiments conducted on both synthetic and real-world datasets demonstrate that our method is two orders of magnitude faster than existing NeRF-based methods, while the accuracy is competitive without initial pose. The accuracy of our method also outperforms NeRF-free absolute pose regression methods by a large margin. Kexue Fu 0001, Haoran Wang 0009, Manning Wang |
ACM Multimedia | 4 |
| 2023 | The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image ClassificationabstractThis paper introduces the novel concept of few-shot weakly supervised learning for pathology Whole Slide Image (WSI) classification, denoted as FSWC. A solution is proposed based on prompt learning and the utilization of a large language model, GPT-4. Since a WSI is too large and needs to be divided into patches for processing, WSI classification is commonly approached as a Multiple Instance Learning (MIL) problem. In this context, each WSI is considered a bag, and the obtained patches are treated as instances. The objective of FSWC is to classify both bags and instances with only a limited number of labeled bags. Unlike conventional few-shot learning problems, FSWC poses additional challenges due to its weak bag labels within the MIL framework. Drawing inspiration from the recent achievements of vision-language models (V-L models) in downstream few-shot classification tasks, we propose a two-level prompt learning MIL framework tailored for pathology, incorporating language prior knowledge. Specifically, we leverage CLIP to extract instance features for each patch, and introduce a prompt-guided pooling strategy to aggregate these instance features into a bag feature. Subsequently, we employ a small number of labeled bags to facilitate few-shot prompt learning based on the bag features. Our approach incorporates the utilization of GPT-4 in a question-and-answer mode to obtain language prior knowledge at both the instance and bag levels, which are then integrated into the instance and bag level language prompts. Additionally, a learnable component of the language prompts is trained using the available few-shot labeled data. We conduct extensive experiments on three real WSI datasets encompassing breast cancer, lung cancer, and cervical cancer, demonstrating the notable performance of the proposed method in bag and instance classification. All codes will be made publicly accessible. Linhao Qu, Xiaoyuan Luo, Kexue Fu 0001, Manning Wang, Zhijian Song |
NeurIPS | 4 |
| 2023 | ProteinMAE: masked autoencoder for protein surface self-supervised learningabstractSUMMARY: The biological functions of proteins are determined by the chemical and geometric properties of their surfaces. Recently, with the booming progress of deep learning, a series of learning-based surface descriptors have been proposed and achieved inspirational performance in many tasks such as protein design, protein-protein interaction prediction, etc. However, they are still limited by the problem of label scarcity, since the labels are typically obtained through wet experiments. Inspired by the great success of self-supervised learning in natural language processing and computer vision, we introduce ProteinMAE, a self-supervised framework specifically designed for protein surface representation to mitigate label scarcity. Specifically, we propose an efficient network and utilize a large number of accessible unlabeled protein data to pretrain it by self-supervised learning. Then we use the pretrained weights as initialization and fine-tune the network on downstream tasks. To demonstrate the effectiveness of our method, we conduct experiments on three different downstream tasks including binding site identification in protein surface, ligand-binding protein pocket classification, and protein-protein interaction prediction. The extensive experiments show that our method not only successfully improves the network's performance on all downstream tasks, but also achieves competitive performance with state-of-the-art methods. Moreover, our proposed network also exhibits significant advantages in terms of computational cost, which only requires less than a tenth of memory cost of previous methods. AVAILABILITY AND IMPLEMENTATION: https://github.com/phdymz/ProteinMAE. Mingzhi Yuan, Kexue Fu 0001, Jiaming Guan, Yingfan Ma, Qin Qiao, Manning Wang |
Bioinform. | 7 |
| 2023 | Density-based one-shot active learning for image segmentation
Qiuye Jin, Shiman Li, Xiaofei Du 0002, Mingzhi Yuan, Manning Wang, Zhijian Song |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | AIM-MEF: Multi-exposure image fusion based on adaptive information mining in both spatial and frequency domains
Linhao Qu, Siqi Yin, Shaolei Liu, Manning Wang, Zhijian Song |
Expert Syst. Appl. | 5 |
| 2023 | A learnable self-supervised task for unsupervised domain adaptation on point cloud classification and segmentation
Shaolei Liu, Xiaoyuan Luo, Kexue Fu 0001, Manning Wang, Zhijian Song |
Frontiers Comput. Sci. | 4 |
| 2023 | Robust Point Cloud Registration Framework Based on Deep Graph Matchingabstract3D point cloud registration is a fundamental problem in computer vision and robotics. Recently, learning-based point cloud registration methods have made great progress. However, these methods are sensitive to outliers, which lead to more incorrect correspondences. In this paper, we propose a novel deep graph matching-based framework for point cloud registration. Specifically, we first transform point clouds into graphs and extract deep features for each point. Then, we develop a module based on deep graph matching to calculate a soft correspondence matrix. By using graph matching, not only the local geometry of each point but also its structure and topology in a larger range are considered in establishing correspondences, so that more correct correspondences are found. We train the network with a loss directly defined on the correspondences, and in the test stage the soft correspondences are transformed into hard one-to-one correspondences so that registration can be performed by a correspondence-based solver. Furthermore, we introduce a transformer-based method to generate edges for graph construction, which further improves the quality of the correspondences. Extensive experiments on object-level and scene-level benchmark datasets show that the proposed method achieves state-of-the-art performance. Kexue Fu 0001, Jiazheng Luo, Xiaoyuan Luo, Shaolei Liu, Chenxi Zhang 0004, Manning Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Dual-Branch Deep Point Cloud Registration Framework for Unconstrained RotationabstractLearning-based rigid point cloud registration (RPCR) studies have made great progress recently but most existing methods have a small convergence region and can only be used to solve the registration problem with a small rotation angle, which is usually constrained within$[0, 45^\circ ]$. However, the relative rotation between point clouds is usually unconstrained in practice. To address this challenging problem, we propose a new RPCR network and integrate it into a new dual-branch registration framework for unconstrained rotation point cloud registration. The dual-branch framework consists of a large-rotation branch and a small-rotation branch, which are used to accurately register point clouds with large and small relative rotations, respectively. In addition, we propose a multiview intersection over the union module to select a better registration result from the output of the two branches. Extensive experiments on both ModelNet40 and MVP-RG datasets demonstrate that our proposed method outperforms existing state-of-the-art techniques by a large margin. Kexue Fu 0001, Mingye Xu, Xiaoyuan Luo, Manning Wang |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | A Structure-Aware Framework of Unsupervised Cross-Modality Domain Adaptation via Frequency and Spatial Knowledge DistillationabstractUnsupervised domain adaptation (UDA) aims to train a model on a labeled source domain and adapt it to an unlabeled target domain. In medical image segmentation field, most existing UDA methods rely on adversarial learning to address the domain gap between different image modalities. However, this process is complicated and inefficient. In this paper, we propose a simple yet effective UDA method based on both frequency and spatial domain transfer under a multi-teacher distillation framework. In the frequency domain, we introduce non-subsampled contourlet transform for identifying domain-invariant and domain-variant frequency components (DIFs and DVFs) and replace the DVFs of the source domain images with those of the target domain images while keeping the DIFs unchanged to narrow the domain gap. In the spatial domain, we propose a batch momentum update-based histogram matching strategy to minimize the domain-variant image style bias. Additionally, we further propose a dual contrastive learning module at both image and pixel levels to learn structure-related information. Our proposed method outperforms state-of-the-art methods on two cross-modality medical image segmentation datasets (cardiac and abdominal). Codes are avaliable at https://github.com/slliuEric/FSUDA. Shaolei Liu, Siqi Yin, Linhao Qu, Manning Wang, Zhijian Song |
IEEE Trans. Medical Imaging | 4 |
| 2022 | TransMEF: A Transformer-Based Multi-Exposure Image Fusion Framework Using Self-Supervised Multi-Task LearningabstractIn this paper, we propose TransMEF, a transformer-based multi-exposure image fusion framework that uses self-supervised multi-task learning. The framework is based on an encoder-decoder network, which can be trained on large natural image datasets and does not require ground truth fusion images. We design three self-supervised reconstruction tasks according to the characteristics of multi-exposure images and conduct these tasks simultaneously using multi-task learning; through this process, the network can learn the characteristics of multi-exposure images and extract more generalized features. In addition, to compensate for the defect in establishing long-range dependencies in CNN-based architectures, we design an encoder that combines a CNN module with a transformer module. This combination enables the network to focus on both local and global information. We evaluated our method and compared it to 11 competitive traditional and deep learning-based methods on the latest released multi-exposure image fusion benchmark dataset, and our method achieved the best performance in both subjective and objective evaluations. Code will be available at https://github.com/miccaiif/TransMEF. Linhao Qu, Shaolei Liu, Manning Wang, Zhijian Song |
AAAI | 3 |
| 2022 | PointCLM: A Contrastive Learning-based Framework for Multi-instance Point Cloud Registration
Mingzhi Yuan, Qiuye Jin, Xinrong Chen, Manning Wang |
ECCV (9) | 5 |
| 2022 | DGMIL: Distribution Guided Multiple Instance Learning for Whole Slide Image Classification
Linhao Qu, Xiaoyuan Luo, Shaolei Liu, Manning Wang, Zhijian Song |
MICCAI (2) | 4 |
| 2022 | Bi-directional Weakly Supervised Knowledge Distillation for Whole Slide Image ClassificationabstractComputer-aided pathology diagnosis based on the classification of Whole Slide Image (WSI) plays an important role in clinical practice, and it is often formulated as a weakly-supervised Multiple Instance Learning (MIL) problem. Existing methods solve this problem from either a bag classification or an instance classification perspective. In this paper, we propose an end-to-end weakly supervised knowledge distillation framework (WENO) for WSI classification, which integrates a bag classifier and an instance classifier in a knowledge distillation framework to mutually improve the performance of both classifiers. Specifically, an attention-based bag classifier is used as the teacher network, which is trained with weak bag labels, and an instance classifier is used as the student network, which is trained using the normalized attention scores obtained from the teacher network as soft pseudo labels for the instances in positive bags. An instance feature extractor is shared between the teacher and the student to further enhance the knowledge exchange between them. In addition, we propose a hard positive instance mining strategy based on the output of the student network to force the teacher network to keep mining hard positive instances. WENO is a plug-and-play framework that can be easily applied to any existing attention-based bag classification methods. Extensive experiments on five datasets demonstrate the efficiency of WENO. Code is available at https://github.com/miccaiif/WENO. Linhao Qu, Xiaoyuan Luo, Manning Wang, Zhijian Song |
NeurIPS | 3 |
| 2022 | Globally Optimal Linear Model Fitting with Unit-Norm Constraint
Yinlong Liu, Manning Wang, Guang Chen 0001, Alois C. Knoll, Zhijian Song |
Int. J. Comput. Vis. | 3 |
| 2022 | Self-supervised learning and semi-supervised learning for multi-sequence medical image classification
Danjun Song, Shengxiang Rao, Manning Wang |
Neurocomputing | 6 |
| 2022 | Cold-start active learning for image classification
Qiuye Jin, Mingzhi Yuan, Shiman Li, Haoran Wang 0009, Manning Wang, Zhijian Song |
Inf. Sci. | 5 |
| 2022 | Deep active learning models for imbalanced image classification
Qiuye Jin, Mingzhi Yuan, Haoran Wang 0009, Manning Wang, Zhijian Song |
Knowl. Based Syst. | 4 |
| 2022 | Wavelet-based self-supervised learning for multi-scene image fusion
Shaolei Liu, Linhao Qu, Qin Qiao, Manning Wang, Zhijian Song |
Neural Comput. Appl. | 4 |
| 2022 | Efficient and Outlier-Robust Simultaneous Pose and Correspondence Determination by Branch-and-Bound and Transformation DecompositionabstractEstimating the pose of a calibrated camera relative to a 3D point set from one image is an important task in computer vision. Perspective-n-Point algorithms are often used if perfect 2D-3D correspondences are known. However, it is difficult to determine 2D-3D correspondences perfectly, and then the simultaneous pose and correspondence determination problem is needed to be solved. Early methods aimed to solve this problem by local optimization. Recently, several new methods are proposed to globally solve this problem by using branch-and-bound (BnB) method, but they tend to be slow because the time complexity of the BnB-based methods is exponential to the dimensionality of the parameter space, and they directly search the 6D parameter space. In this paper, we propose to decompose the joint searching into two separate searching processes by introducing a rotation invariant feature (RIF). Specifically, we construct RIFs from the original 3D and 2D point sets and search for the globally optimal translation to match these two RIFs first. Then, the original 3D point set is translated and matched with the 2D point set to find a globally optimal rotation. Experiments on challenging data show that the proposed method outperforms state-of-the-art methods in terms of both speed and accuracy. Chen Wang 0025, Yinlong Liu, Xuechen Li 0002, Manning Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Robust Point Cloud Registration Framework Based on Deep Graph Matchingabstract3D point cloud registration is a fundamental problem in computer vision and robotics. Recently, learning-based point cloud registration methods have made great progress. However, these methods are sensitive to outliers, which lead to more incorrect correspondences. In this paper, we propose a novel deep graph matching-based framework for point cloud registration. Specifically, we first transform point clouds into graphs and extract deep features for each point. Then, we develop a module based on deep graph matching to calculate a soft correspondence matrix. By using graph matching, not only the local geometry of each point but also its structure and topology in a larger range are considered in establishing correspondences, so that more correct correspondences are found. We train the network with a loss directly defined on the correspondences, and in the test stage the soft correspondences are transformed into hard one-to-one correspondences so that registration can be performed by singular value decomposition. Furthermore, we introduce a transformer-based method to generate edges for graph construction, which further improves the quality of the correspondences. Extensive experiments on registering clean, noisy, partial-to-partial and unseen category point clouds show that the proposed method achieves state-of-the-art performance. The code will be made publicly available at https://github.com/fukexue/RGM. Kexue Fu 0001, Shaolei Liu, Xiaoyuan Luo, Manning Wang |
CVPR | 4 |
| 2021 | WaveFuse: A Unified Unsupervised Framework for Image Fusion with Discrete Wavelet Transform
Shaolei Liu, Manning Wang, Zhijian Song |
ICONIP (4) | 2 |
| 2021 | Practical globally optimal consensus maximization by Branch-and-bound based on interval arithmetic
Yinlong Liu, Xuechen Li 0002, Chen Wang 0025, Manning Wang, Zhijian Song |
Pattern Recognit. | 5 |
| 2020 | Fast correspondence-based point cloud registration by pair-wise inlier checking and transformation decomposition
Chen Wang 0025, Yuxi Jiang, Manning Wang |
Pattern Recognit. Lett. | 3 |
| 2020 | GORFLM: Globally Optimal Robust Fitting for Linear Model
Yinlong Liu, Xuechen Li 0002, Chen Wang 0025, Manning Wang, Zhijian Song |
Signal Process. Image Commun. | 5 |
| 2019 | 2D-3D Point Set Registration Based on Global Rotation SearchabstractSimultaneously determining the relative pose and correspondence between a set of 3D points and its 2D projection is a fundamental problem in computer vision, and the problem becomes more difficult when the point sets are contaminated by noise and outliers. Traditionally, this problem is solved by local optimization methods, which usually start from an initial guess of the pose and alternately optimize the pose and the correspondence. In this paper, we formulate the problem as optimizing the pose of the 3D points in the SE(3) space to make its 2D projection best align with the 2D point set, which is measured by the cardinality of the inlier set on the 2D projection plane. We propose four geometric bounds for the position of the projection of a 3D point on the 2D projection plane and solve the 2D-3D point set registration problem by combining a global optimal rotation search and a grid search of translation. Compared with existing global optimization approaches, the proposed method utilizes a different problem formulation and more efficiently searches the translation space, which improves the registration speed. Experiments with synthetic and real data showed that the proposed approach significantly outperformed state-of-the-art local and global methods. Yinlong Liu, Zhijian Song, Manning Wang |
IEEE Trans. Image Process. | 4 |
| 2018 | Efficient Global Point Cloud Registration by Matching Rotation Invariant Features Through Translation Search
Yinlong Liu, Chen Wang 0025, Zhijian Song, Manning Wang |
ECCV (12) | 4 |
| 2009 | Automatic localization of the center of fiducial markers in 3D CT/MRI images for image-guided neurosurgery
Manning Wang, Zhijian Song |
Pattern Recognit. Lett. | 1 |