EDBT 2026 Demo / reviewers in the wild / expert
Kexue Fu 0001
dblp:287/4634-1
· DBLP profile ↗
33ranked-venue papers
7as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Image Content Matters: An Image Content Aware State Space Model for Accelerated MRI ReconstructionabstractThe challenge of accelerated MRI reconstruction lies in recovering high-quality images from undersampled k-space. Recently, the selective state space model (Mamba) has shown promising results in various tasks with balanced global receptive field and computational efficiency, shedding new light on MRI reconstruction. However, existing approaches directly flatten 2D images based on spatial positions and apply Mamba to vision tasks, failing to preserve and explore the content properties. In this paper, we posit that the key to unlocking Mamba's full potential for MRI reconstruction lies in content-aware sequence modeling. We investigate two fundamental challenges: (1) how to reasonably preserve semantic information when converting 2D images into 1D sequences, and (2) how to effectively identify and recover the crucial high-frequency textures. To this end, we introduce CAM, a novel framework that shifts Mamba-based MRI reconstruction from position-based to content-aware sequence modeling. Specifically, we introduce three modules: (1) the Semantic Preservation Scanning Module (SPSM) introduces learnable clustering centers to group similar pixels, establishing the semantic preserved sequence. (2) The Texture Extraction Scanning Module (TESM) acts as a differentiable local texture descriptor to estimate crucial high-frequency information, forming the texture emphasized sequence. (3) The Texture Enhancement Mamba Module (TEMM) further modulates the semantic sequence with texture-informed system matrices derived from the texture sequence, yielding both context- and texture-aware sequential representations. With these enhancements, CAM significantly outperforms existing methods across various datasets and under-sampling masks. Yucong Meng, Kexue Fu 0001, Zhijian Song, Yonghong Shi |
AAAI | 3 |
| 2026 | TransHER2: Prediction of HER2 Expression Status in Breast Ultrasound Videos Based on Transformer Spatiotemporal Interactive Feature Fusion
Xuejing Li, Longxiang Gao, Lei Cui 0006, Kexue Fu 0001 |
ICIC (3) | 5 |
| 2026 | DH-Mamba: Exploring Dual-Domain Hierarchical State Space Models for MRI Reconstruction
Yucong Meng, Kexue Fu 0001, Zhijian Song, Yonghong Shi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | DiCLIP: Diffusion Model Enhances CLIP's Dense Knowledge for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels typically leverages Class Activation Maps (CAMs) to achieve pixel-level predictions. Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced to generate CAMs in WSSS. However, previous WSSS methods solely adopt CLIP's vision-language paired property for dense localization, neglecting its inherently limited dense knowledge across both visual and text modalities, which renders CAM generation suboptimal. In this work, we propose DiCLIP, a novel WSSS framework that leverages the generative diffusion model to enhance CLIP's dense knowledge across two modalities. Specifically, Visual Correlation Enhancement (VCE) and Text Semantic Augmentation (TSA) modules are proposed for dense prediction enhancement. To improve the spatial awareness of visual features, our VCE module utilizes diffusion's reliable spatial consistency to mitigate the over-smoothing issue in CLIP's attention. It designs the Attention Clustering Refinement (ACR) module to reliably extract diverse correlation maps from the diffusion model. The correlation maps act as a diversity bias for CLIP's self-attention, recursively pushing its visual features towards a more discriminative dense distribution. To augment the semantics of text embeddings, our TSA module argues that a single text modality is insufficient to encompass the variability of visual categories. Thus, we leverage diffusion's generative power to maintain a dynamic key-value cache model, shifting CAM generation from a patch-text matching mechanism to a novel visual knowledge retrieval paradigm. With these enhancements, DiCLIP not only outperforms state-of-the-art methods on PASCAL VOC and MS COCO but also significantly reduces training costs. Code is publicly available at https://github.com/zwyang6/DiCLIP. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
IEEE Trans. Image Process. | 4 |
| 2025 | Focus on Local: Finding Reliable Discriminative Regions for Visual Place RecognitionabstractVisual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual aliasing because of easy overlap. However, existing methods lack precisely modeling and full exploitation of these discriminative regions. In addition, the lack of pixel-level correspondence supervision in the VPR dataset hinders further improvement of the local feature matching capability in the re-ranking stage. In this paper, we propose the Focus on Local (FoL) approach to stimulate the performance of image retrieval and re-ranking in VPR simultaneously by mining and exploiting reliable discriminative local regions in images and introducing pseudo-correlation supervision. First, we design two losses, Extraction-Aggregation Spatial Alignment Loss (SAL) and Foreground-Background Contrast Enhancement Loss (CEL), to explicitly model reliable discriminative local regions and use them to guide the generation of global representations and efficient re-ranking. Second, we introduce a weakly-supervised local feature training strategy based on pseudo-correspondences obtained from aggregating global features to alleviate the lack of local correspondences ground truth for the VPR task. Third, we suggest an efficient re-ranking pipeline that is efficiently and precisely based on discriminative region guidance. Finally, experimental results show that our FoL achieves the state-of-the-art on multiple VPR benchmarks in both image retrieval and re-ranking stages and also significantly outperforms existing two-stage VPR methods in terms of computational efficiency. Changwei Wang 0001, Shunpeng Chen, Rongtao Xu, Jiguang Zhang, Haoran Yang 0003, Yu Zhang 0133, Kexue Fu 0001, Shide Du, Zhiwei Xu 0005, Longxiang Gao, Li Guo 0004, Shibiao Xu |
AAAI | 9 |
| 2025 | MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels typically uses Class Activation Maps (CAM) to achieve dense predictions. Recently, Vision Transformer (ViT) has provided an alternative to generate localization maps from class-patch attention. However, due to insufficient constraints on modeling such attention, we observe that the Localization Attention Maps (LAM) often struggle with the artifact issue, i.e., patch regions with minimal semantic relevance are falsely activated by class tokens. In this work, we propose MoRe to address this issue and further explore the potential of LAM. Our findings suggest that imposing additional regularization on class-patch attention is necessary. To this end, we first view the attention as a novel directed graph and propose the Graph Category Representation module to implicitly regularize the interaction among class-patch entities. It ensures that class tokens dynamically condense the related patch information and suppress unrelated artifacts at a graph level. Second, motivated by the observation that CAM from classification weights maintains smooth localization of objects, we devise the Localization-informed Regularization module to explicitly regularize the class-patch attention. It directly mines the token relations from CAM and further supervises the consistency between class and patch tokens in a learnable manner. Extensive experiments on PASCAL VOC and MS COCO validate that MoRe effectively addresses the artifact issue and achieves state-of-the-art performance, surpassing recent single-stage and even multi-stage methods. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
AAAI | 3 |
| 2025 | Dual Focus-Attention Transformer for Robust Point Cloud RegistrationabstractRecently, coarse-to-fine methods for point cloud registration have achieved great success, but few works deeply explore the impact of feature interaction at both coarse and fine scales. By visualizing attention scores and correspondences, we find that existing methods fail to achieve effective feature aggregation at the two scales during the feature interaction. To tackle this issue, we propose a Dual Focus-Attention Transformer framework, which only focuses on points relevant to the current point for feature interaction, avoiding interactions with irrelevant points. For the coarse scale, we design a superpoint focus-attention transformer guided by sparse keypoints, which are selected from the neighborhood of superpoints. For the fine scale, we only perform feature interaction between the point sets that belong to the same superpoint. Experiments show that our method achieve the state-of-the-art performance on three standard benchmarks. The code and pre-trained models are available at https://github.com/fukexue/DFAT.git. Kexue Fu 0001, Mingzhi Yuan, Changwei Wang 0001, Weiguang Pang, Jing Chi, Manning Wang, Longxiang Gao |
CVPR | 1 |
| 2025 | Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced in WSSS. However, recent methods primarily focus on image-text alignment for CAM generation, while CLIP’s potential in patch-text alignment remains unexplored. In this work, we propose ExCEL to explore CLIP’s dense knowledge via a novel patch-text alignment paradigm for WSSS. Specifically, we propose Text Semantic Enrichment (TSE) and Visual Calibration (VC) modules to improve the dense alignment across both text and vision modalities. To make text embeddings semantically informative, our TSE module applies Large Language Models (LLMs) to build a dataset-wide knowledge base and enriches the text representations with an implicit attribute-hunting process. To mine fine-grained knowledge from visual features, our VC module first proposes Static Visual Calibration (SVC) to propagate fine-grained knowledge in a non-parametric manner. Then Learnable Visual Calibration (LVC) is further proposed to dynamically shift the frozen features towards distributions with diverse semantics. With these enhancements, ExCEL not only retains CLIP’s training-free advantages but also significantly outperforms other state-of-the-art methods with much less training cost on PASCAL VOC and MS COCO. Code is available at https://github.com/zwyang6/ExCEL. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
CVPR | 3 |
| 2025 | AASD: Accelerate Inference by Aligning Speculative Decoding in Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLMs) have achieved notable success in visual instruction tuning, yet their inference is time-consuming due to the auto-regressive decoding of Large Language Model (LLM) backbone. Traditional methods for accelerating inference, including model compression and migration from language model acceleration, often compromise output quality or face challenges in effectively integrating multimodal features. To address these issues, we propose AASD, a novel framework for Accelerating inference with refined KV Cache and Aligning speculative decoding in MLLMs. Our approach leverages the target model’s cached KeyValue (KV) pairs to extract vital information for generating draft tokens, enabling efficient speculative decoding. To reduce the computational burden associated with long multimodal token sequences, we introduce a KV Projector to compress the KV Cache while maintaining representational fidelity. Additionally, we design a Target-Draft Attention mechanism that optimizes the alignment between the draft model and the target model, achieving the benefits of real inference scenarios with minimal computational overhead. Extensive experiments on mainstream MLLMs demonstrate that our method achieves up to a $2 \times$ inference speedup without sacrificing accuracy. This study not only provides an effective and lightweight solution for accelerating MLLM inference but also introduces a novel alignment strategy for speculative decoding in multimodal contexts, laying a strong foundation for future research in efficient MLLMs. Code is availiable at https://github.com/transcend-0/ASD Muyang Zhang, Weiguang Pang, Yuzhi Chen, Rongtao Xu, Kexue Fu 0001, Changwei Wang 0001, Longxiang Gao |
DAC | 7 |
| 2025 | Directed Spatial Consistency-Based Partial-to-Partial Point Cloud Registration with Deep Graph Matchingabstract3D point cloud registration is an essential problem in computer vision, robotics, surgical navigation and augmented reality. Accurate registration of partially overlapped intraoperative point clouds (e.g., femoral reconstruction) remains critical yet challenging in orthopedic navigation due to incomplete overlap and dynamic noise. In this study, we propose a partial-to-partial point cloud registration framework based on directional spatial consistency. First, we extract overlapped areas from partially overlapping point clouds and leverage the point registration graph matching module to calculate the hard point matching matrix. Second, we sample nodes from the source point cloud and generate translation-invariant edge vectors (direction/scale-preserving) via their k-nearest neighbors, guided by predicted point correspondences. This bypasses translation ambiguities by encoding spatial consistency through edges, reducing pose estimation to 3DoF alignment (rotation). The loss explicitly couples point-level matches with edge-level geometric constraints for dual optimization. Building upon this framework, we extract reliable overlapping edge representations and prune their similarity matrix by thresholding low-confidence scores, effectively suppressing spurious matches. The proposed edge-aware matching mechanism further exploits the translation invariance of local structures to refine point correspondences with enhanced accuracy. Finally, we introduce a bidirectional registration mechanism to reinforce optimization stability, achieving state-of-the-art performance across benchmarks. Extensive experiments on ModelNet40, ShapeNet, and MedShapeNet validate our method under diverse scenarios: partial-to-partial, unseen categories, partial-to-full, and cross-dataset generalization, surpassing existing methods in registration accuracy. The codes are available at https://github.com/pidan0824/DSCGM. Kexue Fu 0001, Xinzhe Du, Rui Song 0002, Max Q.-H. Meng, Zhe Min |
IROS | 2 |
| 2025 | Collaboration Wins More: Dual-Modal Collaborative Attention Reinforcement for Mitigating Large Vision Language Models HallucinationabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual-language understanding for downstream multimodal tasks. However, these models often generate descriptions containing objects or details not present in the input image, a phenomenon commonly referred to as ''hallucination''. Existing methods focus solely on single-side hallucination mitigation: Intra-modal-only reinforcement (e.g. visual attention enhancement) ignores prompt-based guidance; Inter-modal-only correlation correction may introduce low-information visual tokens to mislead reasoning. To tackle this challenge, we propose Dual-Modal Collaborative Attention Reinforcement (DuCAR). Specifically, DuCAR is equipped with intra-visual CLS-driven sampling and cross-modal dynamic sampling, extracting important visual tokens guided by intra- and inter-modal joint information. During the multimodal fusion stage, DuCAR adaptively enhances the attention weights of these visual tokens. Our sampling and enhancement strategies in DuCAR simultaneously reinforces informative visual tokens, and suppresses attention dispersion towards question-irrelevant visual information. We conduct extensive experiments on the POPE and CHAIR hallucination benchmarks, demonstrating that our method outperforms existing state-of-the-art mitigation baselines and effectively reduces hallucinations in text generated by LVLMs. The code is available in the https://github.com/xjy2020/DuCAR. Jiye Xie, Liangliang You, Zhiqiang Kou, Kexue Fu 0001, Youyang Qu, Wenjie Yang 0005, Jianwei Guo 0003, Weiliang Meng, Longxiang Gao, Haoran Yang 0003, Changwei Wang 0001, Yu Zhang 0133 |
ACM Multimedia | 7 |
| 2025 | PointMM: A Hybrid Mamba-Transformer Framework for Point Cloud Analysis with Morton Reordering Strategy
Changwei Wang 0001, Shujun Gu, Chuanfu Wu, Longxiang Gao, Kexue Fu 0001, Youyang Qu |
PRCV (10) | 7 |
| 2025 | A Mamba-KAN Joint UNet Framework for Medical Image Segmentation
Haoyu Zhou, Changwei Wang 0001, Weiguang Pang, Lei Cui 0006, Shujun Gu, Longxiang Gao, Kexue Fu 0001, Youyang Qu |
PRCV (3) | 7 |
| 2025 | C2Fi-NeRF: Coarse to fine inversion NeRF for 6D pose estimation
Jiguang Zhang, Zhaohui Zhang 0002, Xuxiang Feng, Shibiao Xu, Rongtao Xu, Changwei Wang 0001, Kexue Fu 0001, Jiaxi Sun, Weilong Ding 0001 |
Expert Syst. Appl. | 7 |
| 2025 | Ddog: optimizing multi-hop inference via dual-driven retrieval and reasoning path
Bruce Gu, Longxiang Gao, Kexue Fu 0001, Youyang Qu, Lei Cui 0006 |
Mach. Learn. | 4 |
| 2025 | Tackling Ambiguity From Perspectives of Uncertainty Inference and Affinity Diversification for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) with image-level labels aims to achieve dense predictions without laborious annotations. However, due to the ambiguous contexts and fuzzy regions, the performance of WSSS, particularly during the stages of generating Class Activation Maps (CAMs) and refining pseudo masks, is widely hindered by ambiguity. Despite this, this issue has received little attention in previous literature. In this work, we propose UniA, a unified single-staged WSSS framework, to efficiently tackle this issue from the perspectives of uncertainty inference and affinity diversification. When activating class objects, we argue that the false activation stems from the bias to ambiguous regions during the feature extraction. Therefore, we formulate a robust feature representation with a Gaussian distribution and introduce the uncertainty estimation to avoid the bias. A distribution loss is proposed to supervise the process, which effectively captures the ambiguity and models the complex dependencies among features. When refining pseudo labels, we observe that the affinity from the prevailing refinement methods intends to be overly similar among ambiguities. To this end, we design an affinity diversification module to promote diversity among semantics. A mutual complementing refinement is first proposed to statically rectify the ambiguous affinity with multiple inferred pseudo labels. Then a contrastive affinity loss is further designed to dynamically diversify the relations among unrelated semantics. It stably propagates the diversity into the feature representation and helps generate better pseudo masks. Extensive experiments are conducted on PASCAL VOC, MS COCO, and medical ACDC datasets, which validate the efficiency of UniA tackling ambiguity and its superiority over recent single-staged or even most multi-staged competitors. Code is publicly available athttps://github.com/zwyang6/UniA. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
IEEE Trans. Multim. | 3 |
| 2024 | Transformer-Based Video-Structure Multi-Instance Learning for Whole Slide Image ClassificationabstractPathological images play a vital role in clinical cancer diagnosis. Computer-aided diagnosis utilized on digital Whole Slide Images (WSIs) has been widely studied. The major challenge of using deep learning models for WSI analysis is the huge size of WSI images and existing methods struggle between end-to-end learning and proper modeling of contextual information. Most state-of-the-art methods utilize a two-stage strategy, in which they use a pre-trained model to extract features of small patches cut from a WSI and then input these features into a classification model. These methods can not perform end-to-end learning and consider contextual information at the same time. To solve this problem, we propose a framework that models a WSI as a pathologist's observing video and utilizes Transformer to process video clips with a divide-and-conquer strategy, which helps achieve both context-awareness and end-to-end learning. Extensive experiments on three public WSI datasets show that our proposed method outperforms existing SOTA methods in both WSI classification and positive region detection. Yingfan Ma, Xiaoyuan Luo, Kexue Fu 0001, Manning Wang |
AAAI | 3 |
| 2024 | Separate and Conquer: Decoupling Co-occurrence via Decomposition and Representation for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) with image-level labels aims to achieve segmentation tasks with-out dense annotations. However, attributed to the frequent coupling of co-occurring objects and the limited supervision from image-level labels, the challenging co-occurrence problem is widely present and leads to false activation of objects in WSSS. In this work, we devise a ‘Separate and Conquer’ scheme SeCo to tackle this issue from di-mensions of image space and feature space. In the im-age space, we propose to ‘separate’ the co-occurring ob-jects with image decomposition by subdividing images into patches. Importantly, we assign each patch a category tag from Class Activation Maps (CAMs), which spatially helps remove the co-context bias and guide the subsequent rep-resentation. In the feature space, we propose to ‘conquer’ the false activation by enhancing semantic representation with multi-granularity knowledge contrast. To this end, a dual-teacher-single-student architecture is designed and tag-guided contrast is conducted, which guarantee the cor-rectness of knowledge and further facilitate the discrepancy among co-contexts. We streamline the multi-staged WSSS pipeline end-to-end and tackle this issue without external supervision. Extensive experiments are conducted, validating the efficiency of our method and the superiority over previous single-staged and even multi-staged competitors on PASCAL VOC and MS COCO. Code is available here. Kexue Fu 0001, Minghong Duan, Linhao Qu, Shuo Wang 0011, Zhijian Song |
CVPR | 2 |
| 2024 | Control Flow Divergence Optimization by Exploiting Tensor CoresabstractKernels are scheduled on Graphics Processing Units (GPUs) in the granularity of GPU warp, which is a bunch of threads that must be scheduled together. When executing kernels with conditional branches, the threads within a warp may execute different branches sequentially, resulting in a considerable utilization loss and unpredictable execution time. This problem is known as the control flow divergence. In this work, we propose a novel method to predict threads' execution path before the launch of the kernel by deploying a branch prediction network on the GPU's tensor cores, which can efficiently parallel run with the kernels on CUDA cores, so that the divergence problem can be eased in a large extent with the lowest overhead. Combined with a well-designed thread data reorganization algorithm, this solution can better mitigate GPUs' control flow divergence problem. Weiguang Pang, Xu Jiang 0004, Songran Liu, Lei Qiao 0002, Kexue Fu 0001, Longxiang Gao, Wang Yi 0001 |
DAC | 5 |
| 2024 | FAST: A Dual-tier Few-Shot Learning Paradigm for Whole Slide Image ClassificationabstractThe expensive fine-grained annotation and data scarcity
have become the primary obstacles for the widespread adoption of deep learning-based Whole Slide Images (WSI) classification algorithms in clinical practice. Unlike few-shot learning methods in natural images that can leverage the labels of each image, existing few-shot WSI classification methods only utilize a small number of fine-grained labels or weakly supervised slide labels for training in order to avoid expensive fine-grained annotation. They lack sufficient mining of available WSIs, severely limiting WSI classification performance. To address the above issues, we propose a novel and efficient dual-tier few-shot learning paradigm for WSI classification, named FAST. FAST consists of a dual-level annotation strategy and a dual-branch classification framework. Firstly, to avoid expensive fine-grained annotation, we collect a very small number of WSIs at the slide level, and annotate an extremely small number of patches. Then, to fully mining the available WSIs, we use all the patches and available patch labels to build a cache branch, which utilizes the labeled patches to learn the labels of unlabeled patches and through knowledge retrieval for patch classification. In addition to the cache branch, we also construct a prior branch that includes learnable prompt vectors, using the text encoder of visual-language models for patch classification. Finally, we integrate the results from both branches to achieve WSI classification. Extensive experiments on binary and multi-class datasets demonstrate that our proposed method significantly surpasses existing few-shot classification methods and approaches the accuracy of fully supervised methods with only 0.22% annotation costs. All codes and models will be publicly available on https://github.com/fukexue/FAST. Kexue Fu 0001, Xiaoyuan Luo, Linhao Qu, Shuo Wang 0011, Ilias Maglogiannis, Longxiang Gao, Manning Wang |
NeurIPS | 1 |
| 2024 | POS-BERT: Point cloud one-stage BERT pre-training
Kexue Fu 0001, Peng Gao 0007, Shaolei Liu, Linhao Qu, Longxiang Gao, Manning Wang |
Expert Syst. Appl. | 1 |
| 2024 | Decoupled deep hough voting for point cloud registration
Mingzhi Yuan, Kexue Fu 0001, Manning Wang |
Frontiers Comput. Sci. | 2 |
| 2024 | Boosting Point-BERT by Multi-Choice TokensabstractMasked language modeling (MLM) has become one of the most successful self-supervised pre-training task. Inspired by its success, Point-BERT, as a pioneer work in point cloud, proposed masked point modeling (MPM) to pre-train point transformer on large scale unanotated dataset. Despite its great performance, we find the inherent difference between language and point cloud tends to cause ambiguous tokenization for point cloud, and no gold standard is available for point cloud tokenization. Point-BERT uses a discrete Variational AutoEncoder (dVAE) as tokenizer, but it might generate different token ids for semantically-similar patches and the same token ids for semantically-dissimilar patches. To tackle the above problems, we propose our McP-BERT, a pre-training framework with multi-choice tokens. Specifically, we ease the previous single-choice constraint on patch token ids in Point-BERT, and provide multi-choice token ids for each patch as supervision. Moreover, we utilitze the high-level semantics learned by transformer to further refine our supervision signals. Extensive experiments on point cloud classification, few-shot classification and part segmentation tasks demonstrate the superiority of our method, e.g., the pre-trained transformer achieves 94.1% accuracy on ModelNet40, 84.28% accuracy on the hardest setting of ScanObjectNN and new state-of-the-art performance on few-shot learning. Our method improves the performance of Point-BERT on all downstream tasks without extra computational overhead. Kexue Fu 0001, Mingzhi Yuan, Shaolei Liu, Manning Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Robust Point Cloud Registration via Random Network Co-EnsembleabstractLearning-based point cloud registration has achieved great success in recent years but is still limited by its generalization. The performance of these methods declines when they are extended to unseen datasets that have inconsistent distributions with the training set. In this paper, we propose a novel random network-based method, which does not require training. Our approach utilizes multiple randomly initialized networks for feature extraction and correspondence building. Furthermore, we also introduce a co-ensemble strategy to prune the outliers in correspondences built upon random networks, which leverages spatial consistency. Through our co-ensemble pruning, a large proportion of outliers can be removed, thereby achieving robust registration in affordable RANSAC iterations. Extensive experiments on 3DMatch and KITTI demonstrate that our method outperforms not only the traditional methods but also the learning-based methods trained on datasets inconsistent with the test set. The code will be released at https://github.com/phdymz/RandPCR. Mingzhi Yuan, Kexue Fu 0001, Yucong Meng, Manning Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | PointMBF: A Multi-scale Bidirectional Fusion Network for Unsupervised RGB-D Point Cloud RegistrationabstractPoint cloud registration is a task to estimate the rigid transformation between two unaligned scans, which plays an important role in many computer vision applications. Previous learning-based works commonly focus on supervised registration, which have limitations in practice. Recently, with the advance of inexpensive RGB-D sensors, several learning-based works utilize RGB-D data to achieve unsupervised registration. However, most of existing unsupervised methods follow a cascaded design or fuse RGB-D data in a unidirectional manner, which do not fully exploit the complementary information in the RGB-D data. To leverage the complementary information more effectively, we propose a network implementing multi-scale bidirectional fusion between RGB images and point clouds generated from depth images. By bidirectionally fusing visual and geometric features in multi-scales, more distinctive deep features for correspondence estimation can be obtained, making our registration more accurate. Extensive experiments on ScanNet and 3DMatch demonstrate that our method achieves new state-of-the-art performance. Code will be released at https://github.com/phdymz/PointMBF. Mingzhi Yuan, Kexue Fu 0001, Yucong Meng, Manning Wang |
ICCV | 2 |
| 2023 | Boosting 3D Point Cloud Registration by Transferring Multi-modality KnowledgeabstractThe recent multi-modality models have achieved great performance in many vision tasks because the extracted features contain the multi-modality knowledge. However, most of the current registration descriptors have only concentrated on local geometric structures. This paper proposes a method to boost point cloud registration accuracy by transferring the multi-modality knowledge of pre-trained multi-modality model to a new descriptor neural network. Different to the previous multi-modality methods that requires both modalities, the proposed method only requires point clouds during inference. Specifically, we propose an ensemble descriptor neural network combining pre-trained sparse convolution branch and a new point-based convolution branch. By fine-tuning on a single modality data, the proposed method achieves new state-of-the-art results on 3DMatch and competitive accuracy on 3DLoMatch and KITTI. The code and the trained model will be released at https://github.com/phdymz/DBENet.git. Mingzhi Yuan, Xiaoshui Huang, Kexue Fu 0001, Manning Wang |
ICRA | 3 |
| 2023 | PI-NeRF: A Partial-Invertible Neural Radiance Fields for Pose EstimationabstractIn recent years, Neural Radiance Fields (NeRF) have been used as a map of 3D scene to estimate the 6-DoF pose of new observed images - given an image, estimate the relative rotation and translation of a camera using a trained NeRF. However, existing NeRF-based pose estimation methods have a small convergence region and need to be optimized iteratively over a given initial pose, which makes them slow and sensitive to the initial pose. In this paper, we propose PI-NeRF that directly outputs the pose of a given image without pose initialization and iterative optimization. This is achieved by integrating NeRF with invertible neural network (INN). Our method employs INNs to establish a bijective mapping between the rays and pixel features, which allows us to directly estimate the ray corresponding to each image pixel using the feature map extracted by an image encoder. Based on these rays, we can directly estimate the pose of the image using the PnP algorithm. Experiments conducted on both synthetic and real-world datasets demonstrate that our method is two orders of magnitude faster than existing NeRF-based methods, while the accuracy is competitive without initial pose. The accuracy of our method also outperforms NeRF-free absolute pose regression methods by a large margin. Kexue Fu 0001, Haoran Wang 0009, Manning Wang |
ACM Multimedia | 2 |
| 2023 | The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image ClassificationabstractThis paper introduces the novel concept of few-shot weakly supervised learning for pathology Whole Slide Image (WSI) classification, denoted as FSWC. A solution is proposed based on prompt learning and the utilization of a large language model, GPT-4. Since a WSI is too large and needs to be divided into patches for processing, WSI classification is commonly approached as a Multiple Instance Learning (MIL) problem. In this context, each WSI is considered a bag, and the obtained patches are treated as instances. The objective of FSWC is to classify both bags and instances with only a limited number of labeled bags. Unlike conventional few-shot learning problems, FSWC poses additional challenges due to its weak bag labels within the MIL framework. Drawing inspiration from the recent achievements of vision-language models (V-L models) in downstream few-shot classification tasks, we propose a two-level prompt learning MIL framework tailored for pathology, incorporating language prior knowledge. Specifically, we leverage CLIP to extract instance features for each patch, and introduce a prompt-guided pooling strategy to aggregate these instance features into a bag feature. Subsequently, we employ a small number of labeled bags to facilitate few-shot prompt learning based on the bag features. Our approach incorporates the utilization of GPT-4 in a question-and-answer mode to obtain language prior knowledge at both the instance and bag levels, which are then integrated into the instance and bag level language prompts. Additionally, a learnable component of the language prompts is trained using the available few-shot labeled data. We conduct extensive experiments on three real WSI datasets encompassing breast cancer, lung cancer, and cervical cancer, demonstrating the notable performance of the proposed method in bag and instance classification. All codes will be made publicly accessible. Linhao Qu, Xiaoyuan Luo, Kexue Fu 0001, Manning Wang, Zhijian Song |
NeurIPS | 3 |
| 2023 | ProteinMAE: masked autoencoder for protein surface self-supervised learningabstractSUMMARY: The biological functions of proteins are determined by the chemical and geometric properties of their surfaces. Recently, with the booming progress of deep learning, a series of learning-based surface descriptors have been proposed and achieved inspirational performance in many tasks such as protein design, protein-protein interaction prediction, etc. However, they are still limited by the problem of label scarcity, since the labels are typically obtained through wet experiments. Inspired by the great success of self-supervised learning in natural language processing and computer vision, we introduce ProteinMAE, a self-supervised framework specifically designed for protein surface representation to mitigate label scarcity. Specifically, we propose an efficient network and utilize a large number of accessible unlabeled protein data to pretrain it by self-supervised learning. Then we use the pretrained weights as initialization and fine-tune the network on downstream tasks. To demonstrate the effectiveness of our method, we conduct experiments on three different downstream tasks including binding site identification in protein surface, ligand-binding protein pocket classification, and protein-protein interaction prediction. The extensive experiments show that our method not only successfully improves the network's performance on all downstream tasks, but also achieves competitive performance with state-of-the-art methods. Moreover, our proposed network also exhibits significant advantages in terms of computational cost, which only requires less than a tenth of memory cost of previous methods. AVAILABILITY AND IMPLEMENTATION: https://github.com/phdymz/ProteinMAE. Mingzhi Yuan, Kexue Fu 0001, Jiaming Guan, Yingfan Ma, Qin Qiao, Manning Wang |
Bioinform. | 3 |
| 2023 | A learnable self-supervised task for unsupervised domain adaptation on point cloud classification and segmentation
Shaolei Liu, Xiaoyuan Luo, Kexue Fu 0001, Manning Wang, Zhijian Song |
Frontiers Comput. Sci. | 3 |
| 2023 | Robust Point Cloud Registration Framework Based on Deep Graph Matchingabstract3D point cloud registration is a fundamental problem in computer vision and robotics. Recently, learning-based point cloud registration methods have made great progress. However, these methods are sensitive to outliers, which lead to more incorrect correspondences. In this paper, we propose a novel deep graph matching-based framework for point cloud registration. Specifically, we first transform point clouds into graphs and extract deep features for each point. Then, we develop a module based on deep graph matching to calculate a soft correspondence matrix. By using graph matching, not only the local geometry of each point but also its structure and topology in a larger range are considered in establishing correspondences, so that more correct correspondences are found. We train the network with a loss directly defined on the correspondences, and in the test stage the soft correspondences are transformed into hard one-to-one correspondences so that registration can be performed by a correspondence-based solver. Furthermore, we introduce a transformer-based method to generate edges for graph construction, which further improves the quality of the correspondences. Extensive experiments on object-level and scene-level benchmark datasets show that the proposed method achieves state-of-the-art performance. Kexue Fu 0001, Jiazheng Luo, Xiaoyuan Luo, Shaolei Liu, Chenxi Zhang 0004, Manning Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Dual-Branch Deep Point Cloud Registration Framework for Unconstrained RotationabstractLearning-based rigid point cloud registration (RPCR) studies have made great progress recently but most existing methods have a small convergence region and can only be used to solve the registration problem with a small rotation angle, which is usually constrained within$[0, 45^\circ ]$. However, the relative rotation between point clouds is usually unconstrained in practice. To address this challenging problem, we propose a new RPCR network and integrate it into a new dual-branch registration framework for unconstrained rotation point cloud registration. The dual-branch framework consists of a large-rotation branch and a small-rotation branch, which are used to accurately register point clouds with large and small relative rotations, respectively. In addition, we propose a multiview intersection over the union module to select a better registration result from the output of the two branches. Extensive experiments on both ModelNet40 and MVP-RG datasets demonstrate that our proposed method outperforms existing state-of-the-art techniques by a large margin. Kexue Fu 0001, Mingye Xu, Xiaoyuan Luo, Manning Wang |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Robust Point Cloud Registration Framework Based on Deep Graph Matchingabstract3D point cloud registration is a fundamental problem in computer vision and robotics. Recently, learning-based point cloud registration methods have made great progress. However, these methods are sensitive to outliers, which lead to more incorrect correspondences. In this paper, we propose a novel deep graph matching-based framework for point cloud registration. Specifically, we first transform point clouds into graphs and extract deep features for each point. Then, we develop a module based on deep graph matching to calculate a soft correspondence matrix. By using graph matching, not only the local geometry of each point but also its structure and topology in a larger range are considered in establishing correspondences, so that more correct correspondences are found. We train the network with a loss directly defined on the correspondences, and in the test stage the soft correspondences are transformed into hard one-to-one correspondences so that registration can be performed by singular value decomposition. Furthermore, we introduce a transformer-based method to generate edges for graph construction, which further improves the quality of the correspondences. Extensive experiments on registering clean, noisy, partial-to-partial and unseen category point clouds show that the proposed method achieves state-of-the-art performance. The code will be made publicly available at https://github.com/fukexue/RGM. Kexue Fu 0001, Shaolei Liu, Xiaoyuan Luo, Manning Wang |
CVPR | 1 |