Changshuo Wang 0001

dblp:297/9002 · DBLP profile ↗
← Back
30ranked-venue papers
11as first author
30since 2021 · last 2026
0000-0002-4056-4922ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 7 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization
abstract
Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness against adversarial attacks—particularly those exploiting both modalities—remains underexplored, posing risks to critical applications like autonomous driving and content moderation. Existing attacks focus on single modalities or require impractical white-box access, limiting their real-world relevance. In this paper, we introduce Multi-Modal Adversarial Synergy (MMAS), a groundbreaking framework that crafts universal, black-box multi-modal attacks against LVLMs. MMAS simultaneously generates a texture scale-constrained Universal Adversarial Perturbation (UAP) for images and a learnable prompt perturbation for text, optimized jointly using only model queries. The image perturbation, bounded by an L∞-norm, leverages wavelet-based texture constraints to ensure imperceptibility and robustness across diverse visual inputs. The text perturbation, constrained by an L2-norm in the embedding space, maintains semantic coherence while steering outputs toward a target. A novel cross-modal regularization term aligns the perturbations’ gradient directions, enhancing their synergistic impact and transferability across tasks and models. Extensive experiments are conducted to verify the strong universal adversarial capabilities of our proposed attack with prevalent LVLMs, spanning a spectrum of tasks on various datasets, all achieved without delving into the details of the model structures.
Wanlong Fang, Changshuo Wang 0001
AAAI3
2026 Rethinking Video-Language Model from the Language Input Perspective
abstract
Driven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between videos and texts. Although previous VLM works have made significant progress, almost all of them implicitly assume that all the texts are predefined by the specific template. In real-world applications, such a strict assumption is impossible to satisfy since 1) predefining all the texts is extremely time-consuming and labor-intensive. 2) these predefined text inputs are too restrictive and user-unfriendly, limiting their applications. It is observed that given a video input, texts with similar semantics but different templates lead to various performances. To this end, in this paper, we propose a novel plug-and-play framework for various VLM-based methods to fully bridge videos and texts. Specifically, we first generate positive and negative texts from the original ones to target specific text components. Then, we propose an attribute-based text reasoning strategy to mine fine-grained textual semantics of generated texts. Finally, we utilize videos as guidance to conduct cross-modal bridging by designing a self-weighted loss. Extensive experiments show that the proposed method can serve as the plug-and-play module to effectively improve the performance of state-of-the-art VLMs.
Wanlong Fang, Changshuo Wang 0001, Xiaoye Qu, Daizong Liu
AAAI3
2026 Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs
abstract
Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these VLMs are task-specific and assume that both video and language inputs are complete. However, real-world VLM applications might face challenges due to deactivated sensors (e.g., cameras are unavailable due to data privacy), yielding modality-incomplete data and leading to inconsistency between training and testing data. While straightforward incomplete input can boast training generalization-ability and lead to training failure, its potential risks to VLMs regarding safety and trustworthiness have been largely neglected. To this end, we make the first attempt to propose a unified incomplete video-language model to process the incomplete multi-modal inputs. Extensive experimental results show that our method can serve as a plug-and-play module for previous works to improve their performance in various multi-modal tasks.
Wanlong Fang, Changshuo Wang 0001, Keke Tang, Daizong Liu, Wei Ji 0008
AAAI3
2026 R²D-LPCC: Relevance-Ranking Guided Region-Adaptive Dynamic LiDAR Point Cloud Compression
abstract
Dynamic LiDAR point cloud compression (LPCC) is crucial for the efficient transmission and storage of large-scale three-dimensional data in applications such as autonomous driving. However, many existing methods, which primarily focus on compressing geometric or motion information, face a fundamental limitation: they treat all points as equally important. This approach neglects the semantic priorities of a scene, resulting in inefficient bit allocation and particularly compromising the reconstruction quality of safety-critical regions, such as pedestrians and vehicles, which are vital to downstream perception tasks. To address these limitations, we propose R²D-LPCC, a relevance-ranking framework for region adaptive LPCC that prioritizes fidelity in semantically important regions. Central to our approach is the Adaptive Relevance Learning (ARL) module, which integrates semantic context with uncertainty to evaluate regional significance and guide compression. We also introduce a Multi-scale Region-Adaptive Transform (MRAT) module to enhance semantic feature modeling and preserve fine-grained details in key areas. Additionally, we develop an adaptive multi-modal motion estimation module to improve motion prediction in complex three-dimensional environments. Extensive experiments conducted on the SemanticKITTI benchmark demonstrate that R²D-LPCC significantly surpasses ten recent state-of-the-art methods, achieving a 45.48% BD-rate gain over the previous leading method, Unicorn, and a 98.58% gain over the GPCC standard, while ensuring superior reconstruction quality in semantically important regions.
Fangzhe Nan, Frederick W. B. Li, Gary K. L. Tam, Zhaoyi Jiang, Bailin Yang, Jingke Cui, Changshuo Wang 0001
AAAI7
2026 Biologically-Inspired Evolutionary Domain Symbiosis for Few-shot and Zero-shot Point Cloud Semantic Segmentation
abstract
Few-shot and zero-shot point cloud semantic segmentation aim to accurately segment novel categories using limited or no labeled samples, respectively. However, existing methods face significant challenges including domain shifts between support and query sets and the inability to handle both few-shot and zero-shot scenarios within a unified framework. To address these issues, we propose a biologically-inspired Evolutionary Domain Symbiosis Network EDS-Net for unified few-shot and zero-shot point cloud semantic segmentation. Specifically, inspired by natural symbiotic evolution, we propose a Symbiotic Evolution Module (SEM) that models co-adaptation between support and query features through self-correlation and cross-correlation mechanisms. Second, motivated by genetic crossover mechanisms, we introduce a Vision-Semantic Bridging Module (VSBM) that treats visual prototypes and semantic prototypes as two “parent” individuals, creating fused offspring prototypes through adaptive crossover operations and mutation strategies for zero-shot scenarios. Third, we develop a multi-generational evolutionary optimization framework employing an adaptive gating network to learn optimal fusion weights across different evolutionary stages. Extensive experiments demonstrate that EDS-Net with biological interpretability achieves state-of-the-art performance on both few-shot and zero-shot settings.
Changshuo Wang 0001, Zhijian Hu, Zaiyang Yu, Yibin Wu, Mingkun Xu, Yusong Wang 0003, Xingyu Gao 0001, Prayag Tiwari
AAAI1
2026 MMPG: MoE-based Adaptive Multi-Perspective Graph Fusion for Protein Representation Learning
abstract
Graph Neural Networks (GNNs) have been widely adopted for Protein Representation Learning (PRL), as residue interaction networks can be naturally represented as graphs. Current GNN-based PRL methods typically rely on single-perspective graph construction strategies, which capture partial properties of residue interactions, resulting in incomplete protein representations. To address this limitation, we propose MMPG, a framework that constructs protein graphs from multiple perspectives and adaptively fuses them via Mixture of Experts (MoE) for PRL. MMPG constructs graphs from physical, chemical, and geometric perspectives to characterize different properties of residue interactions. To capture both perspective-specific features and their synergies, we develop an MoE module, which dynamically routes perspectives to specialized experts, where experts learn intrinsic features and cross-perspective interactions. We quantitatively verify that MoE automatically specializes experts in modeling distinct levels of interaction—from individual representations, to pairwise inter-perspective synergies, and ultimately to a global consensus across all perspectives. Through integrating this multi-level information, MMPG produces superior protein representations and achieves advanced performance on four different downstream protein tasks.
Yusong Wang 0003, Jialun Shen, Shiyin Tan, Mingkun Xu, Changshuo Wang 0001, Zixing Song, Prayag Tiwari
AAAI7
2026 FantasyStyle: Controllable Stylized Distillation for 3D Gaussian Splatting
abstract
The success of 3DGS in generative and editing applications has sparked growing interest in 3DGS-based style transfer. However, current methods still face two major challenges: (1) multi-view inconsistency often leads to style conflicts, resulting in appearance smoothing and distortion; and (2) heavy reliance on VGG features, which struggle to disentangle style and content from style images, often causing content leakage and excessive stylization. To tackle these issues, we introduce FantasyStyle, a 3DGS-based style transfer framework, and the first to rely entirely on diffusion model distillation. It comprises two key components: (1) Multi-View Frequency Consistency. We enhance cross-view consistency by applying a 3D filter to multi-view noisy latent, selectively reducing low-frequency components to mitigate stylized prior conflicts. (2) Controllable Stylized Distillation. To suppress content leakage from style images, we introduce negative guidance to exclude undesired content. In addition, we identify the limitations of Score Distillation Sampling and Delta Denoising Score in 3D style transfer and remove the reconstruction term accordingly. Building on these insights, we propose a controllable stylized distillation that leverages negative guidance to more effectively optimize the 3D Gaussians. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art approaches, achieving higher stylization quality and visual realism across various scenes and styles.
Yitong Yang, Changshuo Wang 0001, Huajie Wang, Shuting He
AAAI3
2026 Efficient large-scale road surface reconstruction via curvature-based frame selection
Hongjia Xing, Changshuo Wang 0001, Zaiyang Yu, Tingran Wang, Yuanjin Fang, Liping Zhang 0014, Xin Ning 0001
Eng. Appl. Artif. Intell.3
2025 Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer Network
abstract
Given some video-query pairs with untrimmed videos and sentence queries, temporal sentence grounding (TSG) aims to locate query-relevant segments in these videos. Although previous respectable TSG methods have achieved remarkable success, they train each video-query pair separately and ignore the relationship between different pairs. To this end, in this paper, we pose a brand-new setting: Multi-Pair TSG, which aims to co-train these pairs. We propose a novel video-query co-training approach, Multi-Thread Knowledge Transfer Network, to locate a variety of video-query pairs effectively and efficiently. Firstly, we mine the spatial and temporal semantics across different queries to cooperate with each other. To learn intra- and inter-modal representations simultaneously, we design a cross-modal contrast module to explore the semantic consistency by a self-supervised strategy. To fully align visual and textual representations between different pairs, we design a prototype alignment strategy to 1) match object prototypes and phrase prototypes for spatial alignment, and 2) align activity prototypes and sentence prototypes for temporal alignment. Finally, we develop an adaptive negative selection module to adaptively generate a threshold for cross-modal matching. Extensive experiments show the effectiveness and efficiency of our proposed method.
Wanlong Fang, Changshuo Wang 0001, Daizong Liu, Keke Tang, Jianfeng Dong, Pan Zhou 0001, Beibei Li 0002
AAAI3
2025 Taylor Series-Inspired Local Structure Fitting Network for Few-shot Point Cloud Semantic Segmentation
abstract
Few-shot point cloud semantic segmentation aims to accurately segment "unseen" new categories in point cloud scenes using limited labeled data. However, pretraining-based methods not only introduce excessive time overhead but also overlook the local structure representation among irregular point clouds. To address these issues, we propose a pretraining-free local structure fitting network for few-shot point cloud semantic segmentation, named TaylorSeg. Specifically, inspired by Taylor series, we treat the local structure representation of irregular point clouds as a polynomial fitting problem and propose a novel local structure fitting convolution, called TaylorConv. This convolution learns the low-order basic information and high-order refined information of point clouds from explicit encoding of local geometric structures. Then, using TaylorConv as the basic component, we construct two variant of TaylorSeg: a non-parametric TaylorSeg-NN and a parametric TaylorSeg-PN. The former can achieve performance comparable to existing parametric models without pretraining. For the latter, we equip it with an Adaptive Push-Pull (APP) module to mitigate the feature distribution differences between the query set and the support set. Extensive experiments validate the effectiveness of the proposed method. Notably, under the 2-way 1-shot setting, TaylorSeg-PN achieves improvements of +2.28% and +4.37% mIoU on the S3DIS and ScanNet datasets respectively, compared to the previous state-of-the-art methods.
Changshuo Wang 0001, Shuting He, Meiqing Wu, Siew-Kei Lam, Prayag Tiwari
AAAI1
2025 Point Clouds Meets Physics: Dynamic Acoustic Field Fitting Network for Point Cloud Understanding
abstract
While existing pre-training-based methods have enhanced point cloud model performance, they have not fundamentally resolved the challenge of local structure representation in point clouds. The limited representational capacity of pure point cloud models continues to constrain the potential of cross-modal fusion methods and performance across various tasks. To address this challenge, we propose a Dynamic Acoustic Field Fitting Network (DAF-Net), inspired by physical acoustic principles. Specifically, we represent local point clouds as acoustic fields and introduce a novel Acoustic Field Convolution (AF-Conv), which treats local aggregation as an acoustic energy field modeling problem and captures fine-grained local shape awareness by dividing the local area into near field and far field. Furthermore, drawing inspiration from multi-frequency wave phenomena and dynamic convolution, we develop the Dynamic Acoustic Field Convolution (DAF-Conv) based on AF-Conv. DAF-Conv dynamically generates multiple weights based on local geometric priors, effectively enhancing adaptability to diverse geometric features. Additionally, we design a Global Shape-Aware (GSA) layer incorporating EdgeConv and multi-head attention mechanisms, which combines with DAF-Conv to form the DAF Block. These blocks are then stacked to create a hierarchical DAFNet architecture. Extensive experiments demonstrate that DAFNet significantly outperforms existing methods across multiple tasks.
Changshuo Wang 0001, Shuting He, Jiawei Han 0008, Zhonghang Liu, Xin Ning 0001, Weijun Li 0002, Prayag Tiwari
CVPR1
2025 FlexUOD: The Answer to Real-world Unsupervised Image Outlier Detection
abstract
How many outliers are within an unlabeled and contaminated dataset? Despite a series of unsupervised outlier detection (UOD) approaches have been proposed, they cannot correctly answer this critical question, resulting in their performance instability across various real-world (varying contamination factor) scenarios. To address this problem, we propose FlexUOD, with a novel contamination factor estimation perspective. FlexUOD not only achieves its remarkable robustness but also is a general and plug-and-play framework, which can significantly improve the performance of existing UOD methods. Extensive experiments demonstrate that FlexUOD achieves state-of-the-art results as well as high efficacy on diverse evaluation benchmarks.
Zhonghang Liu, Kun Zhou 0001, Changshuo Wang 0001, Wen-Yan Lin, Jiangbo Lu
CVPR3
2025 DyPolySeg: Taylor Series-Inspired Dynamic Polynomial Fitting Network for Few-shot Point Cloud Semantic Segmentation
abstract
Few-shot point cloud semantic segmentation effectively addresses data scarcity by identifying unlabeled query samples through semantic prototypes generated from a small set of labeled support samples. However, pre-training-based methods suffer from domain shifts and increased training time. Additionally, existing methods using DGCNN as the backbone have limited geometric structure modeling capabilities and struggle to bridge the categorical information gap between query and support sets. To address these challenges, we propose DyPolySeg, a pre-training-free Dynamic Polynomial fitting network for few-shot point cloud semantic segmentation. Specifically, we design a unified Dynamic Polynomial Convolution (DyPolyConv) that extracts flat and detailed features of local geometry through Low-order Convolution (LoConv) and Dynamic High-order Convolution (DyHoConv), complemented by Mamba Block for capturing global context information. Furthermore, we propose a lightweight Prototype Completion Module (PCM) that reduces structural differences through self-enhancement and interactive enhancement between query and support sets. Experiments demonstrate that DyPolySeg achieves state-of-the-art performance on S3DIS and ScanNet datasets.
Changshuo Wang 0001, Prayag Tiwari
ICML1
2025 ReferSplat: Referring Segmentation in 3D Gaussian Splatting
abstract
We introduce Referring 3D Gaussian Splatting Segmentation (R3DGS), a new task that aims to segment target objects in a 3D Gaussian scene based on natural language descriptions, which often contain spatial relationships or object attributes. This task requires the model to identify newly described objects that may be occluded or not directly visible in a novel view, posing a significant challenge for 3D multi-modal understanding. Developing this capability is crucial for advancing embodied AI. To support research in this area, we construct the first R3DGS dataset, Ref-LERF. Our analysis reveals that 3D multi-modal understanding and spatial relationship modeling are key challenges for R3DGS. To address these challenges, we propose ReferSplat, a framework that explicitly models 3D Gaussian points with natural language expressions in a spatially aware paradigm. ReferSplat achieves state-of-the-art performance on both the newly proposed R3DGS task and 3D open-vocabulary segmentation benchmarks. Dataset and code are available at https://github.com/heshuting555/ReferSplat.
Shuting He, Guangquan Jie, Changshuo Wang 0001, Shuming Hu, Guanbin Li, Henghui Ding
ICML3
2025 Seeing the Overlooked: Bio-Visual Inspired Weak Saliency Feedback Transformer for Person Re-identification
abstract
The domain gap between pretraining data (e.g., ImageNet, LUPerson) and downstream ReID datasets often leads to suboptimal performance when directly fine-tuning pretrained models. While existing methods attempt to bridge this gap by incorporating additional modalities (e.g., text, 3D data) or visual cues (e.g., pose, body masks), these approaches introduce two key limitations: (1) they may distract the model with irrelevant factors like background clutter or clothing variations, and (2) they inevitably increase computational overhead during inference. To address these issues, we propose the Weak Saliency Feedback Transformer (WSFFormer), inspired by the feedback mechanisms in biological visual systems. Unlike traditional one-way feature propagation, WSFFormer employs an adaptive feedback loop during training to enhance low-response regions, enabling the model to capture richer and more discriminative features. The WSFFormer introduces three key components: (1) The Lateral Feedback Module (LFM) mimics retinal lateral inhibition by adaptively suppressing high-response regions and amplifying weak discriminative features, forcing attention on subtle details; (2) The Progressive Feedback Module (PFM) refines feedback through deep-to-shallow closed-loop propagation, blending high-level semantics with spatial details; (3) The Feedback Sensitive Entropy Loss (FSE Loss) optimizes target-domain adaptation by quantifying divergence between forward and feedback-corrected features. Experiments on holistic/occluded ReID benchmarks show WSFFormer outperforms ViT/Swin-based SOTA methods without extra inference cost.
Changshuo Wang 0001, Shuting He, Fangzhe Nan, Prayag Tiwari
ACM Multimedia1
2025 Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation
abstract
Vision-Language Navigation in Continuous Environments (VLN-CE) poses a formidable challenge for autonomous agents, requiring seamless integration of natural language instructions and visual observations to navigate complex 3D indoor spaces. Existing approaches often falter in long-horizon tasks due to limited scene understanding, inefficient planning, and lack of robust decision-making frameworks. We introduce the \textbf{Hierarchical Semantic-Augmented Navigation (HSAN)} framework, a groundbreaking approach that redefines VLN-CE through three synergistic innovations. First, HSAN constructs a dynamic hierarchical semantic scene graph, leveraging vision-language models to capture multi-level environmental representations—from objects to regions to zones—enabling nuanced spatial reasoning. Second, it employs an optimal transport-based topological planner, grounded in Kantorovich's duality, to select long-term goals by balancing semantic relevance and spatial accessibility with theoretical guarantees of optimality. Third, a graph-aware reinforcement learning policy ensures precise low-level control, navigating subgoals while robustly avoiding obstacles. By integrating spectral graph theory, optimal transport, and advanced multi-modal learning, HSAN addresses the shortcomings of static maps and heuristic planners prevalent in prior work. Extensive experiments on multiple challenging VLN-CE datasets demonstrate that HSAN achieves state-of-the-art performance, with significant improvements in navigation success and generalization to unseen environments.
Wanlong Fang, Changshuo Wang 0001
NeurIPS3
2025 Reasoning Beyond Points: A Visual Introspective Approach for Few-Shot 3D Segmentation
abstract
Point Cloud Few-Shot Semantic Segmentation (PC-FSS) aims to segment unknown categories in query samples using only a small number of annotated support samples. However, scene complexity and insufficient representation of local geometric structures pose significant challenges to PC-FSS. To address these issues, we propose a novel pre-training-free Visual Introspective Prototype Segmentation network (VIP-Seg). Specifically, we design a Visual Introspective Prototype (VIP) module that employs a multi-step reasoning approach to tackle intra-class diversity and domain gaps between support and query sets. The VIP module consists of a Prototype Enhancement Module (PEM) and a Prototype Difference Module (PDM), which work alternately to progressively refine prototypes. The PEM enhances prototype discriminability and reduces intra-class diversity, while the PDM learns common representations from the differences between query and support features, effectively eliminating semantic inconsistencies caused by domain gaps. To further reduce intra-class diversity and enhance point discriminative ability, we propose a Dynamic Power Convolution (DyPowerConv) that leverages learnable power functions to effectively capture local geometric structures and detailed features of point clouds. Extensive experiments on S3DIS and ScanNet demonstrate that our proposed VIP-Seg significantly outperforms current state-of-the-art methods, proving its effectiveness in PC-FSS tasks. Our code will be available at https://github.com/changshuowang/VIP-Seg .
Changshuo Wang 0001, Shuting He, Zhijian Hu, Jia-Hong Huang, Yixian Shen, Prayag Tiwari
NeurIPS1
2025 Dual adversity training for domain generalization
Youjia Shao, Changshuo Wang 0001, Wenna Liu, Wencang Zhao
Knowl. Based Syst.3
2025 Learning discriminative topological structure information representation for 2D shape and social network classification via persistent homology
Changshuo Wang 0001, Rongsheng Cao, Ruiping Wang 0005
Knowl. Based Syst.1
2025 Comprehensive disentanglement with fine-grained feature mitigation for domain generalization
Youjia Shao, Changshuo Wang 0001, Qihang Jia, Wencang Zhao
Neural Networks2
2025 Looking Clearer With Text: A Hierarchical Context Blending Network for Occluded Person Re-Identification
abstract
Existing occluded person re-identification (re-ID) methods mainly learn limited visual information for occluded pedestrians from images. However, textual information, which can describe various human appearance attributes, is rarely fully utilized in the task. To address this issue, we propose a Text-guided Hierarchical Context Blending Network ( THCB-Net) for occluded person re-ID. Specifically, at the data level, informative multi-modal inputs are first generated to make full use of the auxiliary role of textual information and make image data have a strong inductive bias for occluded environments. At the feature expression level, we design a novel Hierarchical Context Blending (HCB) module that can adaptively integrate shallow appearance features obtained by CNNs and multi-scale semantic features from visual transformer encoder. At the model optimization level, a Multi-modal Feature Interaction (MFI) module is proposed to learn the multi-modal information of pedestrians from texts and images, then guide the visual transformer encoder and HCB module to further learn discriminative identity information for occluded pedestrians through Image-Multimodal Contrastive (IMC) learning. Extensive experiments on standard occluded person re-ID benchmarks demonstrate that the proposed THCB-Net outperforms state-of-the-art methods.
Changshuo Wang 0001, Xingyu Gao 0001, Meiqing Wu, Siew-Kei Lam, Shuting He, Prayag Tiwari
IEEE Trans. Inf. Forensics Secur.1
2025 EDS-Depth: Enhancing Self-Supervised Monocular Depth Estimation in Dynamic Scenes
abstract
Self-supervised monocular depth estimation usually assumes that training samples contain only static objects, which leads to poor performance in real-world environments. The presence of dynamic objects incurs camera motion estimation errors, motion blur, and occlusions, which induce significant challenges for network training. To address these issues, we introduce EDS-Depth, a self-supervised learning framework, that improves monocular depth estimation in dynamic scenes. Firstly, we propose a novel TCE (Temporal Continuity Enhancement) strategy to reduce camera motion estimation errors and motion blur caused by dynamic objects. Video frames are interpolated to generate more continuous frames in order to smooth dynamic changes and enrich motion details. Secondly, we design a novel IPDM (Iterative Pseudo Depth Masking) module to address inaccurate object motion and occlusions in dynamic scenes. The module integrates multiple optical flows from different frames for triangulation, generating optimal depth as pseudo-supervision labels in dynamic regions. Extensive experiments on Cityscapes and KITTI datasets demonstrate the effectiveness of EDS-Depth, which surpasses state-of-the-art self-supervised monocular depth estimation methods, particularly in dynamic scenes.
Shangshu Yu, Meiqing Wu, Siew-Kei Lam, Changshuo Wang 0001, Ruiping Wang 0005
IEEE Trans. Intell. Transp. Syst.4
2024 GPSFormer: A Global Perception and Local Structure Fitting-Based Transformer for Point Cloud Understanding
Changshuo Wang 0001, Meiqing Wu, Siew-Kei Lam, Xin Ning 0001, Shangshu Yu, Ruiping Wang 0005, Weijun Li 0002, Thambipillai Srikanthan
ECCV (8)1
2024 Occluded person re-identification with deep learning: A survey and perspectives
Enhao Ning, Changshuo Wang 0001, Xin Ning 0001, Prayag Tiwari
Expert Syst. Appl.2
2024 Enhancement, integration, expansion: Activating representation of detailed features for occluded person re-identification
Enhao Ning, Changshuo Wang 0001, Xin Ning 0001
Neural Networks3
2024 Deformation depth decoupling network for point cloud domain adaptation
Xin Ning 0001, Changshuo Wang 0001, Enhao Ning, Lusi Li
Neural Networks3
2024 3D Person Re-Identification Based on Global Semantic Guidance and Local Feature Aggregation
abstract
Person re-identification (Re-ID) has played an extremely crucial role in ensuring social safety and has attracted considerable research attention. 3D shape information is an important clue to understand the posture and shape of pedestrians. However, most existing person Re-ID methods learn pedestrian feature representations from images, ignoring the real 3D human body structure and the spatial relationship between the pedestrians and interferents. To address this problem, our devise a new point cloud Re-ID network (PointReIDNet), designed to obtain 3D shape representations of pedestrians from point clouds of 3D scenes. The model consists of modules, namely global semantic guidance module and local feature extraction module. The global semantic guidance module is designed by enhancing the point cloud feature representation in similar feature neighborhoods and to reduce the interference caused by 3D shape reconstruction or noise. Further, to provide an efficient representation of point clouds, we propose space cover convolution (SC-Conv), which efficiently encodes information on human shapes in local point clouds by constructing anisotropic geometries in the coordinate neighborhoods. Extensive experiments are conducted on four holistic person Re-ID datasets, one occlusion person Re-ID dataset and one point cloud classification dataset. The results exhibit significant improvements over point-cloud-based person Re-ID methods. In particular, the proposed efficient PointReIDNet decreases the number of parameters from 2.30M to 0.35M with an insignificant drop in performance. The source code is available at: https://github.com/changshuowang/PointReIDNet.
Changshuo Wang 0001, Xin Ning 0001, Weijun Li 0002, Xiao Bai 0001, Xingyu Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Pedestrian 3D Shape Understanding for Person Re-Identification via Multi-View Learning
abstract
Recent development in computing power has resulted in performance improvements on holistic(none-occluded) person Re-Identification (ReID) tasks. Nevertheless, the precision of the recent research will diminish when a pedestrian is obstructed by obstacles. Within the realm of 2D space, the loss of information from obstructed objects continues to pose significant challenges in the context of person ReID. Person is a 3D non-grid object, and thus semantic representation learning in only 2D space limits the understanding of occluded person. In the present work, we propose a network based on 3D multi-view learning, allowing it to acquire geometric and shape details of an occluded pedestrian from 3D space. Simultaneously, it capitalizes on advancements in 2D-based networks to extract semantic representations from 3D multi-views. Specifically, the surface random selection strategy is proposed to convert images of 2D RGB into 3D multi-views. Using this strategy, we build four extensive 3D multi-view data collections for person ReID. After that, Pedestrian 3D Shape Understanding for Person Re-Identification via Multi-View Learning(MV-3DSReID), is proposed for identifying the person by learning person geometry and structure representation from the groups of multi-view images. In comparison to alternative data formats (e.g., 2D RGB, 3D point cloud), multi-view images complement each other’s detailed features of the 3D object by adjusting rendering viewpoints, thus facilitating a more comprehensive understanding of the person for both holistic and occluded ReID situations. Experiments on occluded and holistic ReID tasks demonstrate performance levels comparable to state-of-the-art methods, validating the effectiveness of our proposed approach in tackling challenges related to occlusion. The code is available at https://github.com/hangjiaqi1/MV-TransReID.
Zaiyang Yu, Lusi Li, Jinlong Xie, Changshuo Wang 0001, Weijun Li 0002, Xin Ning 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 PointGT: A Method for Point-Cloud Classification and Segmentation Based on Local Geometric Transformation
abstract
Recently, three-dimensional (3D) point-cloud analysis has been extensively utilized in the domain of machine vision, encompassing tasks include shape classification and segmentation. However the inherent disorder in point clouds poses a challenge in capturing relationships among points, particularly when dealing with mutilated and occluded data. To this end, We propose the Point Geometry Transformation (PointGT) method for 3D point-cloud classification and part segmentation, by exploring the underlying geometric structure in the local and global of points. Specifically, the efficacy of PointGT arises from the integration of a local abstraction (LA) module and an optimization strategy. The LA module is tailored to address the localized features inherent to point clouds. This module encapsulates the multidimensional attributes of local edge and inside points. The bi-directional cross-attention mechanism amalgamates these two constituents into the native channel with the primary objective of optimizing the exploitation of edge and inside delineations, thereby judiciously mitigating noise artifacts. Ultimately, the channel residual connections disseminate the postdownsampling point attributes, thereby inheriting the edge and inside delineations gleaned via post bi-directional attention. The effectiveness of the proposed method was verified through the validation of point-cloud classification and segmentation datasets. The empirical findings confirmed the efficacy of PointGT; accuracies of 93.2% and 87.8% were achieved for the ModelNet40 and ScanObjectNN datasets, respectively.
Changshuo Wang 0001, Long Yu 0001, Shengwei Tian, Xin Ning 0001, Joel J. P. C. Rodrigues
IEEE Trans. Multim.2
2022 Learning Discriminative Features by Covering Local Geometric Space for Point Cloud Analysis
abstract
At present, effectively aggregating and transferring the local features of point cloud is still an unresolved technological conundrum. In this study, we propose a new space-cover convolutional neural network (SC-CNN) for tasks such as point cloud classification and segmentation. The core of this network is space-cover convolution (SC-Conv), which implements depthwise separable convolution on the point cloud. In addition, a newly designed space-cover operator (SCOP) replaces depthwise convolution. The key to SC-Conv is constructing anisotropic spatial geometry in the local point cloud. The SCOP achieves this by utilizing the positional and feature relationships to learn the high-order relationship expression between points. First, data-driven adaptive learning from the 3-D coordinate relationship between the local points is used to determine the weight of the SCOP. Then, the edge feature of the neighboring point relative to the sampling point is used as the input of the SCOP. Finally, a deformable spatial geometry is constructed in the feature space between local points to aggregate the local high-order features. By stacking SC-Conv to construct SC-CNN with a hierarchical network structure for point cloud analysis, we can better perceive the shape information of point cloud and improve network robustness. Finally, we provide numerous experiments to verify that SC-CNN parallels or even outperforms advanced methods in shape classification, part segmentation, and large-scale indoor scene segmentation tasks. The open-source code was published athttps://github.com/changshuowang/SC-CNN.
Changshuo Wang 0001, Xin Ning 0001, Linjun Sun, Liping Zhang 0014, Weijun Li 0002, Xiao Bai 0001
IEEE Trans. Geosci. Remote. Sens.1