Weijie Wang 0002

dblp:18/689-2 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-1168-3527ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Disentangled Instrumentalization Learning for Dialogue Emotion Detection
abstract
Emotion detection is a pivotal research area spanning diverse fields. Despite advancements in existing methods enhancing detection accuracy and robustness through various research perspectives, the complexity of dialogue content and variability in dialogue conditions introduce significant challenges. These challenges often obscure the relationship between utterances and true emotions, leading to spurious correlations. To address these challenges, we proposeDisentangled Instrumentalization Learning (DIL), a disentangled instrumentalization learning model based on instrumental variables. Leveraging pre-extracted topic information from GPT-4 as an observable confounder, we disentangle it to learn the representation of required conditional instrumental variables. Our approach employs prior and posterior networks for covariate distribution inference, optimized via reconstruction techniques. Additionally, we introduce an instrumentalization learning method for conditional variables to effectively represent instrumental variables for emotion detection. We illustrate the variable relationships in our model using a causal Directed Acyclic Graph (DAG). Experimental results demonstrate that our model achieves a weighted-average F1-score of 69.0% with 69.4% accuracy on IEMOCAP, and a weighted-average F1-score of 67.6% on MELD, showing consistent improvements over strong baseline methods.
Weizhi Nie, Zixuan Peng, Weijie Wang 0002, Anan Liu
IEEE Trans. Affect. Comput.4
2025 Fully-Geometric Cross-Attention for Point Cloud Registration
abstract
Point cloud registration approaches often fail when the overlap between point clouds is low due to noisy point correspondences. This work introduces a novel cross-attention mechanism tailored for Transformer-based architectures that tackles this problem, by fusing information from coordinates and features at the super-point level between point clouds. This formulation has remained unexplored primarily because it must guarantee rotation and translation invariance since point clouds reside in different and independent reference frames. We integrate the Gromov-Wasserstein distance into the cross-attention formulation to jointly compute distances between points across different point clouds and account for their geometric structure. By doing so, points from two distinct point clouds can attend to each other under arbitrary rigid transformations. At the point level, we also devise a self-attention mechanism that aggregates the local geometric structure information into point features for fine matching. Our formulation boosts the number of inlier correspondences, thereby yielding more precise registration results compared to state-of-the-art approaches. We have conducted an extensive evaluation on 3DMatch, 3DLoMatch, KITTI, and 3DCSR datasets. Project page: https://github.com/twowwj/FLAT.
Weijie Wang 0002, Guofeng Mei, Jian Zhang 0002, Nicu Sebe, Bruno Lepri, Fabio Poiesi
3DV1
2025 FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
Weijie Wang 0002, Qiang Li 0048, Nicu Sebe, Bruno Lepri, Weizhi Nie
ACM Multimedia2
2025 MBT-Polyp: A new Multi-Branch Memory-augmented Transformer for polyp segmentation
Tao Wang 0180, Weijie Wang 0002, Fausto Giunchiglia, Fengzhi Zhao, Ye Zhang 0014, Duo Yu, Guixia Liu
Image Vis. Comput.2
2025 T2TD: Text-3D Generation Model Based on Prior Knowledge Guidance
abstract
In recent years, 3D models have been utilized in many applications, such as auto-drivers, 3D reconstruction, VR, and AR. However, the scarcity of 3D model data does not meet its practical demands. Thus, generating high-quality 3D models efficiently from textual descriptions is a promising but challenging way to solve this problem. In this paper, inspired by the creative mechanisms of human imagination, which concretely supplement the target model from ambiguous descriptions built upon human experiential knowledge, we propose a novel text-3D generation model (T2TD). T2TD aims to generate the target model based on the textual description with the aid of experiential knowledge. Its target creation process simulates the imaginative mechanisms of human beings. In this process, we first introduce the text-3D knowledge graph to preserve the relationship between 3D models and textual semantic information, which provides related shapes like humans' experiential information. Second, we propose an effective causal inference model to select useful feature information from these related shapes, which can remove the unrelated structure information and only retain solely the feature information strongly related to the textual description. Third, we adopt a novel multi-layer transformer structure to progressively fuse this strongly related structure information and textual information, compensating for the lack of structural information, and enhancing the final performance of the 3D generation model. The final experimental results demonstrate that our approach significantly improves 3D model generation quality and outperforms the SOTA methods on the text2shape datasets.
Weizhi Nie, Rui-dong Chen, Weijie Wang 0002, Bruno Lepri, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Structure Causal Models and LLMs Integration in Medical Visual Question Answering
abstract
Medical Visual Question Answering (MedVQA) aims to answer medical questions according to medical images. However, the complexity of medical data leads to confounders that are difficult to observe, so bias between images and questions is inevitable. Such cross-modal bias makes it challenging to infer medically meaningful answers. In this work, we propose a causal inference framework for the MedVQA task, which effectively eliminates the relative confounding effect between the image and the question to ensure the precision of the question-answering (QA) session. We are the first to introduce a novel causal graph structure that represents the interaction between visual and textual elements, explicitly capturing how different questions influence visual features. During optimization, we apply the mutual information to discover spurious correlations and propose a multi-variable resampling front-door adjustment method to eliminate the relative confounding effect, which aims to align features based on their true causal relevance to the question-answering task. In addition, we also introduce a prompt strategy that combines multiple prompt forms to improve the model's ability to understand complex medical data and answer accurately. Extensive experiments on three MedVQA datasets demonstrate that 1) our method significantly improves the accuracy of MedVQA, and 2) our method achieves true causal correlations in the face of complex medical data.
Zibo Xu, Qiang Li 0048, Weizhi Nie, Weijie Wang 0002, Anan Liu
IEEE Trans. Medical Imaging4
2024 Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-supervised Learning
Bin Ren 0005, Guofeng Mei, Danda Pani Paudel, Weijie Wang 0002, Yawei Li 0001, Mengyuan Liu 0001, Rita Cucchiara, Luc Van Gool, Nicu Sebe
ACCV (7)4
2024 UVMap-ID: A Controllable and Personalized UV Map Generative Model
abstract
Recently, diffusion models have made significant strides in synthesizing realistic 2D human images based on provided text prompts. Building upon this, researchers have extended 2D text-to-image diffusion models into the 3D domain for generating human textures (UV Maps). However, some important problems about UV Map Generative models are still not solved, i.e., how to generate personalized texture maps for any given face image, and how to define and evaluate the quality of these generated texture maps. To solve the above problems, we introduce a novel method, UVMap-ID, which is a controllable and personalized UV Map generative model. Unlike traditional large-scale training methods in 2D, we propose to fine-tune a pre-trained text-to-image diffusion model which is integrated with a face fusion module for achieving ID-driven customized generation. To support the finetuning strategy, we introduce a small-scale attribute-balanced training dataset, including high-quality textures with labeled text and Face ID. Additionally, we introduce some metrics to evaluate the multiple aspects of the textures. Finally, both quantitative and qualitative analyses demonstrate the effectiveness of our method in controllable and personalized UV Map generation.
Weijie Wang 0002, Jichao Zhang, Chang Liu 0030, Xia Li 0005, Xingqian Xu, Humphrey Shi, Nicu Sebe, Bruno Lepri
ACM Multimedia1
2024 Learning spatial-spectral dual adaptive graph embedding for multispectral and hyperspectral image fusion
Xuquan Wang, Feng Zhang 0028, Kai Zhang 0010, Weijie Wang 0002, Xiong Dun, Jiande Sun 0001
Pattern Recognit.4
2023 Unsupervised Deep Probabilistic Approach for Partial Point Cloud Registration
abstract
Deep point cloud registration methods face challenges to partial overlaps and rely on labeled data. To address these issues, we propose UDPReg, an unsupervised deep probabilistic registration framework for point clouds with partial overlaps. Specifically, we first adopt a network to learn posterior probability distributions of Gaussian mixture models (GMMs) from point clouds. To handle partial point cloud registration, we apply the Sinkhorn algorithm to predict the distribution-level correspondences under the constraint of the mixing weights of GMMs. To enable unsupervised learning, we design three distribution consistency-based losses: self-consistency, cross-consistency, and local contrastive. The self-consistency loss is formulated by encouraging GMMs in Euclidean and feature spaces to share identical posterior distributions. The cross-consistency loss derives from the fact that the points of two partially overlapping point clouds belonging to the same clusters share the cluster centroids. The cross-consistency loss allows the network to flexibly learn a transformation-invariant posterior distribution of two aligned point clouds. The local contrastive loss facilitates the network to extract discriminative local features. Our UDPReg achieves competitive performance on the 3DMatch/3DLoMatch and ModelNet/ModelLoNet benchmarks.
Guofeng Mei, Hao Tang 0005, Xiaoshui Huang, Weijie Wang 0002, Juan Liu 0006, Jian Zhang 0002, Luc Van Gool, Qiang Wu 0001
CVPR4
2023 Dynamically Instance-Guided Adaptation: A Backward-free Approach for Test-Time Domain Adaptive Semantic Segmentation
abstract
In this paper, we study the application of Test-time domain adaptation in semantic segmentation (TTDA-Seg) where both efficiency and effectiveness are crucial. Existing methods either have low efficiency (e.g., backward optimization) or ignore semantic adaptation (e.g., distribution alignment). Besides, they would suffer from the accumulated errors caused by unstable optimization and abnormal distributions. To solve these problems, we propose a novel backward-free approach for TTDA-Seg, called Dynamically Instance-Guided Adaptation (DIGA). Our principle is utilizing each instance to dynamically guide its own adaptation in a non-parametric way, which avoids the error accumulation issue and expensive optimizing cost. Specifically, DIGA is composed of a distribution adaptation module (DAM) and a semantic adaptation module (SAM), enabling us to jointly adapt the model in two indispensable aspects. DAM mixes the instance and source BN statistics to encourage the model to capture robust representation. SAM combines the historical prototypes with instance-level prototypes to adjust semantic predictions, which can be associated with the parametric classifier to mutually benefit the final results. Extensive experiments evaluated on five target domains demonstrate the effectiveness and efficiency of the proposed method. Our DIGA establishes new state-of-the-art performance in TTDA-Seg. Source code is available at: https://github.com/Waybaba/DIGA.
Zhun Zhong, Weijie Wang 0002, Charles Ling 0001, Boyu Wang 0004, Nicu Sebe
CVPR3
2020 HGAN: Holistic Generative Adversarial Networks for Two-dimensional Image-based Three-dimensional Object Retrieval
abstract
In this article, we propose a novel method to address the two-dimensional (2D) image-based 3D object retrieval problem. First, we extract a set of virtual views to represent each 3D object. Then, a soft-attention model is utilized to find the weight of each view to select one characteristic view for each 3D object. Second, we propose a novel Holistic Generative Adversarial Network (HGAN) to solve the cross-domain feature representation problem and make the feature space of virtual characteristic view more inclined to the feature space of the real picture. This will effectively mitigate the distribution discrepancies across the 2D image domains and 3D object domains. Finally, we utilize the generative model of the HGAN to obtain the “virtual real image” of each 3D object and make the characteristic view of the 3D object and real picture possess the same feature space for retrieval. To demonstrate the performance of our approach, We established a new dataset that includes pairs of 2D images and 3D objects, where the 3D objects are based on the ModelNet40 dataset. The experimental results demonstrate the superiority of our proposed method over the state-of-the-art methods.
Weizhi Nie, Weijie Wang 0002, Anan Liu, Jie Nie, Yuting Su 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2019 Characteristic Views Extraction Modal Based-on Deep Reinforcement Learning for 3D Model Retrieval
abstract
In this paper, we propose an effective framework for view-based 3D model retrieval based on reinforcement learning, which aims to extract the most important views from the view set of 3D models for model representation. Since the choice of the representative views can be treated as a Markov decision process (MDP), we model the view extraction as a progressive process through deep reinforcement learning. First, the visual tool developed by OpenGL is utilized to extract 12 rendered pictures in order from a 3D model. Second, these 12 views are fed into the network as a view-based sequence, and then we utilize characteristic views extraction model to extract the most informative views of the 3D model. Finally, a similarity measure between 3D models is performed to complete the last step pf the model retrieval problem. In the experiment section, the classic ModelNet40 is utilized to evaluate the performance of the proposed method. Extensive experiments and corresponding experimental results have demonstrated the superiority of our approach.
Weizhi Nie, Weijie Wang 0002, Anan Liu
ICIP2