Xinwei He 0001

dblp:392/9058 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
18since 2021 · last 2025
0000-0001-7267-0394ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Tetrahedron-Net for Medical Image Registration
abstract
Medical image registration plays a vital role in medical image processing. Extracting expressive representations for medical images is crucial for improving the registration quality. One common practice for this end is constructing a convolutional backbone to enable interactions with skip connections among feature extraction layers. The de facto structure, U-Net-like networks, has attempted to design skip connections such as nested or full-scale ones to connect one single encoder and one single decoder to improve its representation capacity. Despite being effective, it still does not fully explore interactions with a single encoder and decoder architectures. In this paper, we embrace this observation and introduce a simple yet effective alternative strategy to enhance the representations for registrations by appending one additional decoder. The new decoder is designed to interact with both the original encoder and decoder. In this way, it not only reuses feature presentation from corresponding layers in the encoder but also interacts with the original decoder to corporately give more accurate registration results. The new architecture is concise yet generalized, with only one encoder and two decoders forming a “Tetrahedron” structure, thereby dubbed Tetrahedron-Net. Three instantiations of Tetrahedron-Net are further constructed regarding the different structures of the appended decoder. Our extensive experiments prove that superior performance can be obtained on several representative benchmarks of medical image registration. Finally, such a “Tetrahedron” design can also be easily integrated into popular U-Net-like architectures including VoxelMorph, ViT-V-Net, and TransMorph, leading to consistent performance gains.
Jinhai Xiang, Dantong Shi, Xinwei He 0001
BIBM6
2025 Graph-Level Anomaly Detection of Brain Connectivity with Structural Interpretation and Knowledge Distillation
abstract
Graph-level anomaly detection is a critical yet un-derexplored task, especially in the neuroimaging domain where early identification of abnormal brain patterns is vital. In this paper, we propose KDGAE, a generalizable graph autoencoder framework based on knowledge distillation. It employs a teacher-student architecture for graph reconstruction and latent representation distillation. Anomaly scores are then derived from both reconstruction errors and embedding discrepancies. To improve interpretability, post-hoc explanation tools such as feature masking and GNNExplainer are integrated. KDGAE is evaluated on the UB-GOLD benchmark and the Autism Brain Imaging Data Exchange (ABIDE) dataset for autism detection, achieving competitive performance against state-of-the-art methods under limited supervision.
Javeed Muhammad Ahmad, Xinwei He 0001, Jingbo Xia
BIBM4
2025 WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation
abstract
Weakly supervised referring expression comprehension (WREC) and segmentation (WRES) aim to learn object grounding based on a given expression using weak super-vision signals like image-text pairs. While these tasks have traditionally been modeled separately, we argue that they can benefit from joint learning in a multi-task framework. To this end, we propose WeakMCN, a novel multi-task collaborative network that effectively combines WREC and WRES with a dual-branch architecture. Specifically, the WREC branch is formulated as anchor-based contrastive learning, which also acts as a teacher to supervise the WRES branch. In WeakMCN, we propose two innovative designs to facilitate multi-task collaboration, namely Dynamic Visual Feature Enhancement (DVFE) and Collaborative Consistency Module (CCM). DVFE dynamically combines various pre-trained visual knowledge to meet different task requirements, while CCM promotes cross-task consistency from the perspective of optimization. Extensive experimental results on three popular REC and RES benchmarks, i.e., RefCOCO, RefCOCO+, and RefCOCOg, consistently demonstrate performance gains of WeakMCN over state-of-the-art single-task alternatives, e.g., up to 3.91% and 13.11% on RefCOCO for WREC and WRES tasks, respectively. Furthermore, experiments also validate the strong generalization ability of WeakMCN in both semi-supervised REC and RES settings against existing methods, e.g., +8.94% for semi-REC and +7.71% for semi-RES on 1% RefCOCO. The code is publicly available at https://github.com/MRUIL/WeakMCN.
Silin Cheng 0001, Yang Liu 0271, Xinwei He 0001, Sébastien Ourselin, Gen Luo
CVPR3
2025 LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
abstract
The success of Large Language Models (LLMs) has inspired the development of Multimodal Large Language Models (MLLMs) for unified understanding of vision and language. However, the increasing model size and computational complexity of large-scale MLLMs (l-MLLMs) limit their use in resource-constrained scenarios. Although small-scale MLLMs (s-MLLMs) are designed to reduce computational costs, they typically suffer from performance degradation. To mitigate this limitation, we propose a novel LLaVA-KD framework to transfer knowledge from l-MLLMs to s-MLLMs. Specifically, we introduce Multimodal Distillation (MDist) to transfer teacher model's robust representations across both visual and linguistic modalities, and Relation Distillation (RDist) to transfer teacher model's ability to capture visual token relationships. Additionally, we propose a three-stage training scheme to fully exploit the potential of the proposed distillation strategy: 1) Distilled Pre-Training to strengthen the alignment between visual-linguistic representations in s-MLLMs, 2) Supervised Fine-Tuning to equip the s-MLLMs with multimodal understanding capacity, and 3) Distilled Fine-Tuning to refine s-MLLM's knowledge. Our approach significantly improves s-MLLMs performance without altering the model architecture. Extensive experiments and ablation studies validate the effectiveness of each proposed component. Code will be available at https://github.com/Fantasyele/LLaVA-KD.
Jiangning Zhang, Haoyang He, Xinwei He 0001, Ao Tong, Zhenye Gan, Chengjie Wang 0001, Zhucun Xue, Yong Liu 0007, Xiang Bai
ICCV4
2025 Describe, Adapt and Combine: Empowering CLIP Encoders for Open-Set 3D Object Retrieval
Yang Zhou 0007, Zhe Liu 0033, Rui Yu 0002, Song Bai 0001, Yulong Wang 0002, Xinwei He 0001, Xiang Bai
ICCV7
2025 From Individual to Universal: Regularized Multi-view Joint Representation for Multi-view Subspace-Preserving Recovery
abstract
Recent years have witnessed an explosion of Multi- view Subspace Classification (MSCla) and Multi-view Subspace Clustering (MSClu) methods for various applications. However, their theoretical foundation have not been well explored and understood. In this paper, we investigate the multi-view subspace-preserving recovery theory, which is the theoretical underpinnings for MSCla and MSClu methods. Specifically, we derive novel geometrically interpretable conditions for the success of multi-view subspace-preserving recovery. Compared with prior related works, we make the following innovations: First, our theory does not require the equality constraint, which is a common requirement in prior theoretical works and may be too restrictive in reality. Second, we provide both Individual Theoretical Guarantee (ITG) and Universal Theoretical Guarantee (UTG) for multi-view subspace-preserving recovery while prior works only give the UTG. Third, we also apply the proposed theory to establish theoretical guarantees for MSCla and MSClu, respectively. Numerical results validate the proposed theory for multi-view subspace-preserving recovery.
Yulong Wang 0002, Xinwei He 0001, Qiwei Xie, Kit Ian Kou, Yuan Yan Tang
IJCAI3
2025 TeDA: Boosting Vision-Lanuage Models for Zero-Shot 3D Object Retrieval via Testing-time Distribution Alignment
abstract
Learning discriminative 3D representations that generalize well to unknown testing categories is an emerging requirement for many real-world 3D applications. Existing well-established methods often struggle to attain this goal due to insufficient 3D training data from broader concepts. Meanwhile, pre-trained large vision-language models (e.g., CLIP) have shown remarkable zero-shot generalization capabilities. Yet, they are limited in extracting suitable 3D representations due to substantial gaps between their 2D training and 3D testing distributions. To address these challenges, we propose Testing-time Distribution Alignment (TeDA), a novel framework that adapts a pretrained 2D vision-language model CLIP for unknown 3D object retrieval at test time. To our knowledge, it is the first work that studies the test-time adaptation of a vision-language model for 3D feature learning. TeDA projects 3D objects into multi-view images, extracts features using CLIP, and refines 3D query embeddings with an iterative optimization strategy by confident query-target sample pairs in a self-boosting manner. Additionally, TeDA integrates textual descriptions generated by a multimodal language model (InternVL) to enhance 3D object understanding, leveraging CLIP's aligned feature space to fuse visual and textual cues. Extensive experiments on four open-set 3D object retrieval benchmarks demonstrate that TeDA greatly outperforms state-of-the-art methods, even those requiring extensive training. We also experimented with depth maps on Objaverse-LVIS, further validating its effectiveness. Code is available at https://github.com/wangzhichuan123/TeDA.
Yang Zhou 0007, Jinhai Xiang, Yulong Wang 0002, Xinwei He 0001
ICMR5
2025 CLIP-AdaM: Adapting Multi-view CLIP for Open-set 3D Object Retrieval
abstract
Open-set 3D object retrieval (3DOR) aims to learn discriminative and generalizable embeddings for unseen categories of 3D objects. However, attaining this objective typically requires the costly acquisition of large-scale 3D object datasets and associated resources for model training. Building upon the strong open-world representation capabilities of CLIP, we introduce CLIP-AdaM, which, to our knowledge, represents the first attempt to adapt a CLIP model for open-set 3DOR with minimal effort. We first find that a pretrained CLIP already delivers a surprisingly acceptable performance on multi-view images. To further unleash its potential, we design a customized adapter for learning to aggregate and adapt its pretrained features towards better 3D embeddings. For aggregation, it learns two sets of view scores to weigh the contributions of view images for fusion. One is learned by a tiny view-score network at the instance level, and the other is learned implicitly at the dataset level, aiding generalization to unseen categories. The adaptation component comprises only a basic linear layer yet yields superior results. During training, the adapter with such a small amount of parameters can be efficiently fine-tuned with limited 3D closed-set data, effectively mitigating the overfitting issue while harnessing the prior knowledge from pretrained models. Without bells and whistles, CLIP-AdaM attains state-of-the-art performance on four open-set 3DOR benchmarks. Additionally, it demonstrates strong extensibility to broader scenarios, including zero-shot, few-shot, and seen/unseen 3D representation learning.
Xinwei He 0001, Yuxuan Cheng, Yulong Wang 0002, Yang Zhou 0007, Xiang Bai
SIGIR1
2025 Dynamic Expert Routing for Unsupervised Continual Anomaly Detection
abstract
Unsupervised continual anomaly detection (UCAD) aims to develop a model that can continuously learn in dynamic scenarios while avoiding forgetting previously acquired knowledge from past tasks, enabling effective anomaly detection in both past and present tasks. Recently, some works have attempted to alleviate forgetting by exploiting knowledge banks to preserve past knowledge. However, they are proven to be time-consuming, limiting their practical applicability. In this article, we propose a new framework calledDynamicExpertRouting (DER) for UCAD. The key idea is to construct a task-specific expert for each anomaly detection task during training, and dynamically select the appropriate expert during inference. To build experts for anomaly detection, DER first employs a shared pre-trained encoder for feature extraction across all tasks. On top of that, an adaptor and a learnable prompt are then jointly introduced for each task. Meanwhile, in UCAD, the task identity of the sample is unknown during inference, making it challenging to select the appropriate expert. To address this issue, we propose an Adaptive Selection Module that dynamically determines the task identity based on the sample’s semantic representation. DER achieves encouraging UCAD performance in terms of accuracy and inference speed. For instance, it advances the state-of-the-art under both the MVTec-MCCL and MVTec-KSDD settings. Meanwhile, it performs 3× faster than the previous art with comparable parameters.
Xinwei He 0001, Ao Tong, Xiang Bai
IEEE Trans. Ind. Informatics2
2024 BcMatch: Semi-supervised Medical Image Segmentation with Bias Correction
abstract
Semi-supervised semantic segmentation (SSS), which allows for learning a better model with a small fraction of labeled samples and a large amount of unlabeled ones, is valuable yet challenging in medical image analysis. Recent works (e.g., UniMatch) have found that weak-to-strong consistency via augmentation is especially conducive to SSS training. However, they inadvertently introduce cognitive biases for unlabeled images, making it difficult to segment accurately near the target’s edge regions. In this paper we present a lightweight bias-correct module to self-correct these mistakes between the strong perturbations. Based on it, we design a new framework named BcMatch by plugging it into UniMatch to reduce such cognitive biases stemming from incorrect pseudo-labels for unlabeled images. Moreover, we also introduce a bias correction loss, which works in tandem with the consistency loss to guide the model learning, focusing more on the edge regions of the targets. Experiments on the representative semi-supervised segmentation dataset, ACDC, demonstrate our BcMatch surpasses UniMatch by a large margin, attaining new state-of-the-art performance. The code is at https://github.com/zhangyan498/BcMatch.
Jinhai Xiang, Jiakun Yu, Xinwei He 0001
BIBM4
2024 SimpleFusion: A Simple Fusion Framework for Infrared and Visible Images
Yuxuan Cheng, Xinwei He 0001, Yan Aze, Jinhai Xiang
PRCV (8)3
2024 Point-StyleGAN: Multi-scale point cloud synthesis with style modulation
Yang Zhou 0007, Zhiqiang Lin 0004, Xinwei He 0001, Hui Huang 0004
Comput. Aided Geom. Des.4
2024 LATFormer: Locality-Aware Point-View Fusion Transformer for 3D shape recognition
Xinwei He 0001, Silin Cheng 0001, Dingkang Liang, Song Bai 0001, Xi Wang 0044, Yingying Zhu 0005
Pattern Recognit.1
2024 A Discrepancy Aware Framework for Robust Anomaly Detection
abstract
Defect detection is a critical research area in artificial intelligence. Recently, synthetic data-based self-supervised learning has shown great potential on this task. Although many sophisticated synthesizing strategies exist, little research has been done to investigate the robustness of models when faced with different strategies. In this article, we focus on this issue and find that existing methods are highly sensitive to them. To alleviate this issue, we present a discrepancy aware framework (DAF), which demonstrates robust performance consistently with simple and cheap strategies across different anomaly detection benchmarks. We hypothesize that the high sensitivity to synthetic data of existing self-supervised methods arises from their heavy reliance on the visual appearance of synthetic data during decoding. In contrast, our method leverages an appearance-agnostic cue to guide the decoder in identifying defects, thereby alleviating its reliance on synthetic appearance. To this end, inspired by existing knowledge distillation methods, we employ a teacher-student network, which is trained based on synthesized outliers, to compute the discrepancy map as the cue. Extensive experiments on two challenging datasets prove the robustness of our method. Under the simple synthesis strategies, it outperforms existing methods by a large margin. Furthermore, it also achieves the state-of-the-art localization performance.
Dingkang Liang, Dongliang Luo, Xinwei He 0001, Xin Yang 0008, Xiang Bai
IEEE Trans. Ind. Informatics4
2024 Multi-Modal 3D Object Detection by Box Matching
abstract
Multi-modal 3D object detection has received growing attention as the information from different sensors like LiDAR and cameras are complementary. Most fusion methods for 3D detection rely on an accurate alignment and calibration between 3D point clouds and RGB images. However, such an assumption is not reliable in a real-world self-driving system, as the alignment between different modalities is easily affected by asynchronous sensors and disturbed sensor placement. We propose a novel Fusion network by Box Matching (FBMNet) for multi-modal 3D detection, which provides an alternative way for cross-modal feature alignment by learning the correspondence at the bounding box level to free up the dependency of calibration during inference. With the learned assignments between 3D and 2D object proposals, the fusion for detection can be effectively performed by combining their ROI features. Extensive experiments on the nuScenes dataset demonstrate that our method is much more robust in dealing with challenging cases such as asynchronous sensors, misaligned sensor placement, and degenerated camera images than existing fusion methods. We hope that our FBMNet could provide an available solution to dealing with these challenging cases for safety in real autonomous driving scenarios.
Zhe Liu 0033, Xiaoqing Ye, Zhikang Zou, Xinwei He 0001, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001, Xiang Bai
IEEE Trans. Intell. Transp. Syst.4
2023 AiA-UNet: Attention in Attention for Medical Image Segmentation
abstract
In medical image segmentation, U-Net has consistently played a vital role. Recently, the U-Net networks based on the Vision Transformer (ViT) architecture have become more and more popular. ViT exhibits superior capabilities in handling long-range dependencies and capturing global contextual information. However, it requires significant computational cost, and does not explore the optimal matching and the potential dependencies between different patches. To address the aforementioned issues, we propose a novel network framework, called AiA-UNet, for medical image segmentation. The AiA-UNet makes two main contributions. A convolutional self-attention mechanism is proposed to replace the self-attention module in ViT, effectively reducing computational complexity. Moreover, an Attention in Attention module (AiA) is applied within the ViT block. Experimental results on the Synapse multi-organ segmentation dataset demonstrate that AiA-UNet outperforms Trans-UNet by 5.40% and Swin-UNet by 3.75%. Code and models are available at https://github.com/xiaoqin1998/AiA-UNet.
Jianfeng Qin, Xinwei He 0001, Jiakun Yu, Jinhai Xiang, Lulu Wu
BIBM2
2023 Trans-UNeter: A new Decoder of TransUNet for Medical Image Segmentation
abstract
Recently, how to integrate convolutional neural networks and transformers into a U-Net-like encoder-decoder structure has drawn growing interest in medical image segmentation, as transformer is more efficient in capturing longrange relations. Following this line of research, TransUNet is one representative work. However, it still insufficiently explores the rich relations of features from the encoder layer and the decoder layer with just a simple concatenation, which weakens their effectiveness to some extent. To address this issue, we propose two important design improvements to strengthen TransUNet: 1) a novel skip connection module, which upsamples the high-level semantic features and fuse it with low-level features, producing comprehensive semantic-aware features for the decoder. 2) an improved decoder network cascades reverse attention and spatial attention to adaptively combines features from the corresponding encoder layer and the previously decoded outputs.The results of the abdominal multi-organ segmentation experiment on the Synapse multi-organ segmentation dataset indicated that Trans-UNeter improved the mean similarity coefficient(DSC) by 3.71% compared to TransUNet. Code and models are available at https://github.com/iaoqin/Trans-UNeter.
Jiakun Yu, Xinwei He 0001, Jianfeng Qin, Jinhai Xiang, Weiming Zhao
BIBM2
2021 PRA-Net: Point Relation-Aware Network for 3D Point Cloud Analysis
abstract
Learning intra-region contexts and inter-region relations are two effective strategies to strengthen feature representations for point cloud analysis. However, unifying the two strategies for point cloud representation is not fully emphasized in existing methods. To this end, we propose a novel framework named Point Relation-Aware Network (PRA-Net), which is composed of an Intra-region Structure Learning (ISL) module and an Inter-region Relation Learning (IRL) module. The ISL module can dynamically integrate the local structural information into the point features, while the IRL module captures inter-region relations adaptively and efficiently via a differentiable region partition scheme and a representative point-based strategy. Extensive experiments on several 3D benchmarks covering shape classification, keypoint estimation, and part segmentation have verified the effectiveness and the generalization ability of PRA-Net. Code will be available at https://github.com/XiwuChen/PRA-Net.
Silin Cheng 0001, Xiwu Chen, Xinwei He 0001, Zhe Liu 0033, Xiang Bai
IEEE Trans. Image Process.3
2020 A comparison of methods for 3D scene shape retrieval
Juefei Yuan, Hameed Abdul-Rashid, Bo Li 0013, Yijuan Lu, Tobias Schreck, Song Bai 0001, Xiang Bai, Ngoc-Minh Bui, Minh N. Do, Trong-Le Do, Anh Duc Duong, Xinwei He 0001, Mike Holenderski, Dmitri Jarnikov, Tu-Khiem Le, Wenhui Li 0001, Anan Liu
Comput. Vis. Image Underst.13
2020 An Improved Multi-View Convolutional Neural Network for 3D Object Retrieval
abstract
Learning robust and discriminative representations is essential for 3D object retrieval. In this paper, we present an improved Multi-view Convolutional Neural Network (MVCNN) for view-based 3D object representation learning. Our technical contributions are divided into two aspects. First, we propose to employ Group-view Similarity Learning (GSL) over the multi-view representations before the aggregation operation (i.e., max-pooling in MVCNN). We assume that the similarity information among the view groups of different 3D objects can provide an important cue but has been neglected more or less by previous methods. To enhance it, we add a branch to the original MVCNN architecture and learn to maintain such group-view similarity relationships. Second, we utilize an end-to-end metric learning loss function to improve the representation learning process. In particular, we propose an improved Triplet-Center Loss (TCL) named Adaptive Margin based Triplet-Center Loss (AMTCL). The original TCL assumes a fixed and common margin to control the relative distance relationship between a sample to its corresponding class center and to the nearest negative center. Though TCL has demonstrated its great capacity on the 3D object retrieval task, however, when considering the distinguishability between samples of one class and samples of another class, we assume that it would be more appropriate that the margin takes different values based on the distinguishability of samples of different classes. Therefore we propose to adaptively and dynamically adjust the margin hyperparameter based on the normalized confusion matrix which is obtained on the training set during the training process. Extensive experiments on several public 3D shape benchmarks show that our method, GSL + AMTCL, can learn more suitable representations for 3D object retrieval, obtaining superior performance against state-of-the-art methods.
Xinwei He 0001, Song Bai 0001, Jiajia Chu, Xiang Bai
IEEE Trans. Image Process.1
2019 View N-Gram Network for 3D Object Retrieval
abstract
How to aggregate multi-view representations of a 3D object into an informative and discriminative one remains a key challenge for multi-view 3D object retrieval. Existing methods either use view-wise pooling strategies which neglect the spatial information across different views or employ recurrent neural networks which may face the efficiency problem. To address these issues, we propose an effective and efficient framework called View N-gram Network (VNN). Inspired by n-gram models in natural language processing, VNN divides the view sequence into a set of visual n-grams, which involve overlapping consecutive view sub-sequences. By doing so, spatial information across multiple views is captured, which helps to learn a discriminative global embedding for each 3D object. Experiments on 3D shape retrieval benchmarks, including ModelNet10, ModelNet40 and ShapeNetCore55 datasets, demonstrate the superiority of our proposed method.
Xinwei He 0001, Tengteng Huang, Song Bai 0001, Xiang Bai
ICCV1
2019 VD-SAN: Visual-Densely Semantic Attention Network for Image Caption Generation
Xinwei He 0001, Baoguang Shi, Xiang Bai
Neurocomputing1
2019 Image Caption Generation with Part of Speech Guidance
Xinwei He 0001, Baoguang Shi, Xiang Bai, Gui-Song Xia, Zhaoxiang Zhang 0001, Weisheng Dong
Pattern Recognit. Lett.1
2018 Triplet-Center Loss for Multi-View 3D Object Retrieval
abstract
Most existing 3D object recognition algorithms focus on leveraging the strong discriminative power of deep learning models with softmax loss for the classification of 3D data, while learning discriminative features with deep metric learning for 3D object retrieval is more or less neglected. In the paper, we study variants of deep metric learning losses for 3D object retrieval, which did not receive enough attention from this area. First, two kinds of representative losses, triplet loss and center loss, are introduced which could learn more discriminative features than traditional classification loss. Then, we propose a novel loss named triplet-center loss, which can further enhance the discriminative power of the features. The proposed triplet-center loss learns a center for each class and requires that the distances between samples and centers from the same class are closer than those from different classes. Extensive experimental results on two popular 3D object retrieval benchmarks and two widely-adopted sketch-based 3D shape retrieval benchmarks consistently demonstrate the effectiveness of our proposed loss, and significant improvements have been achieved compared with the state-of-the-arts.
Xinwei He 0001, Yang Zhou 0007, Song Bai 0001, Xiang Bai
CVPR1