VLDB 2026 Research / reviewers in the wild / expert
Wenfei Yang
dblp:211/5775
· DBLP profile ↗
50ranked-venue papers
9as first author
48since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-author · 34 since 2021Artificial intelligence and machine learning · 28 · 4 first-author · 28 since 2021Computer networks · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Near-Field Propagation and Spatial Non-Stationarity Channel Model for 6-24 GHz (FR3) Extremely Large-Scale MIMO: Adopted by 3GPP for 6GabstractNext generation cellular deployments are expected to exploit the 6–24 GHz frequency range 3 (FR3) and extremely large-scale multiple-input multiple-output (XL-MIMO) to enable ultra-high data rates and reliability. However, the significantly enlarged antenna apertures and higher carrier frequencies make the far-field and spatial stationarity assumptions in the existing 3rd generation partnership project (3GPP) channel models no longer valid, giving rise to new features such as near-field propagation and spatial non-stationarity (SNS). Despite extensive prior research, incorporating these new features within the standardized channel modeling framework remains an open issue. To address this, this paper presents a channel modeling framework for XL-MIMO systems that incorporates both near-field and SNS features, adopted by 3GPP. For the near-field propagation feature, the framework models the distances from the base station (BS) and user equipment to the spherical-wave sources associated with clusters. These distances are used to characterize element-wise variations of path parameters, such as nonlinear changes in phase and angle. To capture the effect of SNS at the BS side, a stochastic-based approach is proposed to model SNS caused by incomplete scattering, by establishing power attenuation factors from visibility probability and visibility region to characterize antenna element-wise path power variation. In addition, a physical blocker-based approach is introduced to model SNS effects caused by partial blockage. The near-field and SNS channel modeling approaches are validated against ray-tracing simulations. Finally, a simulation framework for near-field and SNS is developed based on the existing 3GPP channel model. Performance evaluations demonstrate that the near-field model captures higher channel capacity potential compablack to the far-field model. Coupling loss results indicate that SNS leads to more pronounced propagation fading relative to the spatial stationary model. Huixin Xu, Jianhua Zhang 0001, Hongbo Xing, Haiyang Miao, Wenfei Yang, Zhening Zhang, Afshin Haghighat, Qixing Wang, Guangyi Liu 0001 |
IEEE J. Sel. Areas Commun. | 9 |
| 2026 | UniSOT: A Unified Framework for Multi-Modality Single Object TrackingabstractSingle object tracking aims to localize target object with specific reference modalities (bounding box, natural language or both) in a sequence of specific video modalities (RGB, RGB+Depth, RGB+Thermal or RGB+Event.). Different reference modalities enable various human-machine interactions, and different video modalities are demanded in complex scenarios to enhance tracking robustness. Existing trackers are designed for single or several video modalities with single or several reference modalities, which leads to separate model designs and limits practical applications. Practically, a unified tracker is needed to handle various requirements. To the best of our knowledge, there is still no tracker that can perform tracking with these above reference modalities across these video modalities simultaneously. Thus, in this paper, we present a unified tracker, UniSOT, for different combinations of three reference modalities and four video modalities with uniform parameters. Extensive experimental results on 18 visual tracking, vision-language tracking and RGB+X tracking benchmarks demonstrate that UniSOT shows superior performance against modality-specific counterparts. Notably, UniSOT outperforms previous counterparts by over 3.0% AUC on TNL2K across all three reference modalities and outperforms Un-Track by over 2.0% main metric across all three RGB+X video modalities. Yinchao Ma, Yuyang Tang 0001, Wenfei Yang, Tianzhu Zhang 0001, Feng Wu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Uncertainty-Aware Projection Network for Multi-View 3D Object DetectionabstractIn recent years, multi-view 3D object detection has drawn a lot of interest in autonomous driving. The projection from 2D image point to 3D world space is the core step of multi-view 3D object detection. However, the uncertainty of depth estimation and camera parameters can make the projected 3D position inaccurate, leading to a performance drop. To this end, we propose an Uncertainty-aware Projection Network (UPN), including a depth uncertainty-guided 3D point generation module and a camera-aware 3D position encoder. For the depth uncertainty-guided 3D point generation module, we predict the depth as a distribution to generate depth uncertainty-aware 3D points. For the camera-aware 3D position encoder, we alleviate the uncertainty of camera parameters by employing adaptive receptive fields to refine the position embedding. Based on these two modules, the proposed UPN can effectively deal with the performance drop caused by inaccurate 2D-3D projection. To the best of our knowledge, this is the first work to model the 2D-3D projection uncertainty for multi-view 3D object detection. The proposed UPN achieves 56.9% mAP and 65.4 % NDS on the nuScenes test dataset, demonstrating the superiority of our method. Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | SMTrack: State-Aware Mamba for Efficient Temporal Modeling in Visual TrackingabstractVisual tracking aims to automatically estimate the state of a target object in a video sequence, which is challenging especially in dynamic scenarios. Thus, numerous methods are proposed to introduce temporal cues to enhance tracking robustness. However, conventional CNN and Transformer architectures exhibit inherent limitations in modeling long-range temporal dependencies in visual tracking, often necessitating either complex customized modules or substantial computational costs to integrate temporal cues. Inspired by the success of the state space model, we propose a novel temporal modeling paradigm for visual tracking, termed State-aware Mamba Tracker (SMTrack), providing a neat pipeline for training and tracking without needing customized modules or substantial computational costs to build long-range temporal dependencies. It enjoys several merits. First, we propose a novel selective state-aware space model with state-wise parameters to capture more diverse temporal cues for robust tracking. Second, SMTrack facilitates long-range temporal interactions with linear computational complexity during training. Third, SMTrack enables each frame to interact with previously tracked frames via hidden state propagation and updating, which releases computational costs of handling temporal cues during tracking. Extensive experimental results demonstrate that SMTrack achieves promising performance with low computational costs. Yinchao Ma, Dengqing Yang, Zhangyu He, Wenfei Yang, Tianzhu Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | DiffCorr: Conditional Diffusion Model with Reliable Pseudo-Label Guidance for Unsupervised Point Cloud Shape CorrespondenceabstractUnsupervised point cloud shape correspondence aims to establish dense correspondences between source and target point clouds. Existing methods universally follow a one-step paradigm to obtain shape correspondence directly, but it often fails in large-scale motions of humans and animals. To address this challenge, we propose a conditional Diffusion model with reliable pseudo-label guidance for unsupervised point cloud shape Correspondence (DiffCorr), including a transformer-based conditional diffusion model and a reliable pseudo-label generator. The proposed DiffCorr enjoys several merits. Firstly, the transformer-based conditional diffusion model implements a coarse-to-fine optimization for coarse correspondences. Secondly, we design a reliable pseudo-label generator to provide high-quality pseudo-labels for training. Extensive experiments on four human and animal datasets demonstrate that DiffCorr surpasses state-of-the-art methods and exhibits favorable generalization capabilities. Jiacheng Deng 0002, Jiahao Lu 0001, Zhixin Cheng, Wenfei Yang |
AAAI | 4 |
| 2025 | Pamba: Enhancing Global Interaction in Point Clouds via State Space ModelabstractTransformers have demonstrated impressive results for 3D point cloud semantic segmentation. However, the quadratic complexity of transformer makes computation costs high, limiting the number of points that can be processed simultaneously and impeding the modeling of long-range dependencies between objects in a single scene. Drawing inspiration from the great potential of recent state space models (SSM) for long sequence modeling, we introduce Mamba, an SSM-based architecture, to the point cloud domain and propose Pamba, a novel architecture with strong global modeling capability under linear complexity. Specifically, to make the disorderness of point clouds fit in with the causal nature of Mamba, we propose a multi-path serialization strategy applicable to point clouds. Besides, we propose the ConvMamba block to compensate for the shortcomings of Mamba in modeling local geometries and in unidirectional modeling. Pamba obtains state-of-the-art results on several 3D point cloud segmentation tasks, including ScanNet v2, ScanNet200, S3DIS and nuScenes, while its effectiveness is validated by extensive experiments. Yubo Ai, Jiahao Lu 0001, Chuxin Wang, Jiacheng Deng 0002, Hanzhi Chang, Yanzhe Liang, Wenfei Yang, Tianzhu Zhang 0001 |
AAAI | 8 |
| 2025 | Structure-Aware Correspondence Learning for Relative Pose EstimationabstractRelative pose estimation provides a promising way for achieving object-agnostic pose estimation. Despite the success of existing 3D correspondence-based methods, the reliance on explicit feature matching suffers from small overlaps in visible regions and unreliable feature estimation for invisible regions. Inspired by humans’ ability to assemble two object parts that have small or no overlapping regions by considering object structure, we propose a novel Structure-Aware Correspondence Learning method for Relative Pose Estimation, which consists of two key modules. First, a structure-aware keypoint extraction module is designed to locate a set of kepoints that can represent the structure of objects with different shapes and appearance, under the guidance of a keypoint based image reconstruction loss. Second, a structure-aware correspondence estimation module is designed to model the intra-image and inter-image relationships between keypoints to extract structure-aware features for correspondence estimation. By jointly leveraging these two modules, the proposed method can naturally estimate 3D-3D correspondences for unseen objects without explicit feature matching for precise relative pose estimation. Experimental results on the CO3D, Objaverse and LineMOD datasets demonstrate that the proposed method significantly outperforms prior methods, i.e., with 5.7° reduction in mean angular error on the CO3D dataset. Wenfei Yang, Tianzhu Zhang 0001, Feng Wu 0001 |
CVPR | 2 |
| 2025 | Implicit Correspondence Learning for Image-to-Point Cloud RegistrationabstractImage-to-point cloud registration aims to estimate the camera pose of a given image within a 3D scene point cloud. In this area, matching-based methods have achieved leading performance by first detecting the overlapping region, then matching point and pixel features learned by neural networks and finally using the PnP-RANSAC algorithm to estimate camera pose. However, achieving accurate image-to-point cloud registration remains challenging because the overlapping region detection is unreliable merely relying on point-wise classification, direct alignment of cross-modal data is difficult and indirect optimization objective leads to unstable registration results. To address these challenges, we propose a novel implicit correspondence learning method, including a Geometric Prior-guided overlapping region Detection Module (GPDM), an Implicit Correspondence Learning Module (ICLM), and a Pose Regression Module (PRM). The proposed method enjoys several merits. First, the proposed GPDM can precisely detect the overlapping region. Second, the ICLM can generate robust cross-modality correspondences. Third, the PRM can enable end-to-end optimization. Extensive experimental results on KITTI and nuScenes datasets demonstrate that the proposed model sets a new state-of-the-art performance in registration accuracy. Xinjun Li, Wenfei Yang, Jiacheng Deng 0002, Zhixin Cheng, Tianzhu Zhang 0001 |
CVPR | 2 |
| 2025 | Rethinking Correspondence-based Category-Level Object Pose EstimationabstractCategory-level object pose estimation aims to determine the pose and size of arbitrary objects within given categories. Existing two-stage correspondence-based methods first establish correspondences between camera and object coordinates, and then acquire the object pose using a pose fitting algorithm. In this paper, we conduct a comprehensive analysis of this paradigm and introduce two crucial essentials: 1) shape-sensitive and pose-invariant feature extraction for accurate correspondence prediction, and 2) outlier correspondence removal for robust pose fitting. Based on these insights, we propose a simple yet effective correspondence-based method called SpotPose, which includes two stages. During the correspondence prediction stage, pose-invariant geometric structure of objects is thoroughly exploited to facilitate shape-sensitive holistic interaction among keypoint-wise features. During the pose fitting stage, outlier scores of correspondences are explicitly predicted to facilitate efficient identification and removal of outliers. Experimental results on CAMERA25, REAL275 and HouseCat6D benchmarks demonstrate that the proposed SpotPose outperforms state-of-the-art approaches by a large margin. Wenfei Yang, Tianzhu Zhang 0001 |
CVPR | 2 |
| 2025 | Learning Neural Scene Representation from iToF Imaging
Wenjie Chang, Hanzhi Chang, Wenfei Yang, Tianzhu Zhang 0001 |
ICCV | 4 |
| 2025 | CA-I2P: Channel-Adaptive Registration Network with Global Optimal SelectionabstractDetection-free methods typically follow a coarse-to-fine pipeline, extracting image and point cloud features for patch-level matching and refining dense pixel-to-point correspondences. However, differences in feature channel attention between images and point clouds may lead to degraded matching results, ultimately impairing registration accuracy. Furthermore, similar structures in the scene could lead to redundant correspondences in cross-modal matching. To address these issues, we propose Channel Adaptive Adjustment Module (CAA) and Global Optimal Selection Module (GOS). CAA enhances intra-modal features and suppresses cross-modal sensitivity, while GOS replaces local selection with global optimization. Experiments on RGB-D Scenes V2 and 7-Scenes demonstrate the superiority of our method, achieving state-of-the-art performance in image-to-point cloud registration. Zhixin Cheng, Jiacheng Deng 0002, Xinjun Li, Xiaotian Yin, Bohao Liao, Baoqun Yin, Wenfei Yang, Tianzhu Zhang 0001 |
ICCV | 7 |
| 2025 | Diffusion-Based Source-Biased Model for Single Domain Generalized Object Detection
Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001 |
ICCV | 2 |
| 2025 | StruMamba3D: Exploring Structural Mamba for Self-Supervised Point Cloud Representation LearningabstractRecently, Mamba-based methods have demonstrated impressive performance in point cloud representation learning by leveraging State Space Model (SSM) with the efficient context modeling ability and linear complexity. However, these methods still face two key issues that limit the potential of SSM: Destroying the adjacency of 3D points during SSM processing and failing to retain long-sequence memory as the input length increases in downstream tasks. To address these issues, we propose StruMamba3D, a novel paradigm for self-supervised point cloud representation learning. It enjoys several merits. First, we design spatial states and use them as proxies to preserve spatial dependencies among points. Second, we enhance the SSM with a state-wise update strategy and incorporate a lightweight convolution to facilitate interactions between spatial states for efficient structure modeling. Third, our method reduces the sensitivity of pre-trained Mamba-based models to varying input lengths by introducing a sequence length-adaptive strategy. Experimental results across four downstream tasks showcase the superior performance of our method. In addition, our method attains the SOTA 95.1% accuracy on ModelNet40 and 92.75% accuracy on the most challenging split of ScanObjectNN without voting strategy. Chuxin Wang, Yixin Zha, Wenfei Yang, Tianzhu Zhang 0001 |
ICCV | 3 |
| 2025 | Learning Shape-Independent Transformation via Spherical Representations for Category-Level Object Pose EstimationabstractCategory-level object pose estimation aims to determine the pose and size of novel objects in specific categories. Existing correspondence-based approaches typically adopt point-based representations to establish the correspondences between primitive observed points and normalized object coordinates. However, due to the inherent shape-dependence of canonical coordinates, these methods suffer from semantic incoherence across diverse object shapes. To resolve this issue, we innovatively leverage the sphere as a shared proxy shape of objects to learn shape-independent transformation via spherical representations. Based on this insight, we introduce a novel architecture called SpherePose, which yields precise correspondence prediction through three core designs. Firstly, We endow the point-wise feature extraction with SO(3)-invariance, which facilitates robust mapping between camera coordinate space and object coordinate space regardless of rotation transformation. Secondly, the spherical attention mechanism is designed to propagate and integrate features among spherical anchors from a comprehensive perspective, thus mitigating the interference of noise and incomplete point cloud. Lastly, a hyperbolic correspondence loss function is designed to distinguish subtle distinctions, which can promote the precision of correspondence prediction. Experimental results on CAMERA25, REAL275 and HouseCat6D benchmarks demonstrate the superior performance of our method, verifying the effectiveness of spherical representations and architectural innovations. Wenfei Yang, Xiang Liu 0020, Tianzhu Zhang 0001 |
ICLR | 2 |
| 2025 | State Space Model Meets Transformer: A New Paradigm for 3D Object DetectionabstractDETR-based methods, which use multi-layer transformer decoders to refine object queries iteratively, have shown promising performance in 3D indoor object detection. However, the scene point features in the transformer decoder remain fixed, leading to minimal contributions from later decoder layers, thereby limiting performance improvement. Recently, State Space Models (SSM) have shown efficient context modeling ability with linear complexity through iterative interactions between system states and inputs. Inspired by SSMs, we propose a new 3D object DEtection paradigm with an interactive STate space model (DEST). In the interactive SSM, we design a novel state-dependent SSM parameterization method that enables system states to effectively serve as queries in 3D indoor detection tasks. In addition, we introduce four key designs tailored to the characteristics of point cloud and SSM: The serialization and bidirectional scanning strategies enable bidirectional feature interaction among scene points within the SSM. The inter-state attention mechanism models the relationships between state points, while the gated feed-forward network enhances inter-channel correlations. To the best of our knowledge, this is the first method to model queries as system states and scene points as system inputs, which can simultaneously update scene point features and query features with linear complexity. Extensive experiments on two challenging datasets demonstrate the effectiveness of our DEST-based method. Our method improves the GroupFree baseline in terms of $\text{AP}_{50}$ on ScanNet V2 (+5.3) and SUN RGB-D (+3.2) datasets. Based on the VDETR baseline, Our method sets a new state-of-the-art on the ScanNetV2 and SUN RGB-D datasets. Chuxin Wang, Wenfei Yang, Xiang Liu 0020, Tianzhu Zhang 0001 |
ICLR | 2 |
| 2025 | Prototype Optimal Transport for Box-Supervised 3D Instance Segmentationabstract3D Instance segmentation (3DIS) on point clouds is a fundamental task in the field of 3D scene understanding. Existing fully-supervised networks have achieved promising results but remain heavily reliant on point-wise annotated data. Using the instance bounding boxes as annotations for weakly-supervised learning is a feasible way to solve the label-efficiency problem. In this paper, we propose POTNet, a novel training paradigm designed to generate point-wise pseudo-labels using only bounding box annotations, which effectively considering both local and global information. We leverage prototype learning method to extract local features from the non-overlapping regions indicated by the bounding boxes as instance prototypes. We employ an optimal transport algorithm to assign points in overlapping regions to their corresponding instances based on the similarity matrix between prototypes and these points. Our approach enables the generation of point-wise pseudo-labels that fully account for local and global correlations. We demonstrate the effectiveness of our method by achieving performance comparable to state-of-the-art approaches on multiple datasets, without requiring any additional supplementary data or retraining processes. Wenfei Yang, Tianzhu Zhang 0001, Xiang Liu 0020 |
ICME | 2 |
| 2025 | Exploring Vision Semantic Prompt for Efficient Point Cloud UnderstandingabstractA series of pre-trained models have demonstrated promising results in point cloud understanding tasks and are widely applied to downstream tasks through fine-tuning. However, full fine-tuning leads to the forgetting of pretrained knowledge and substantial storage costs on edge devices. To address these issues, Parameter-Efficient Transfer Learning (PETL) methods have been proposed. According to our analysis, we find that existing 3D PETL methods cannot adequately align with semantic relationships of features required by downstream tasks, resulting in suboptimal performance. To ensure parameter efficiency while introducing rich semantic cues, we propose a novel fine-tuning paradigm for 3D pre-trained models. We utilize frozen 2D pre-trained models to provide vision semantic prompts and design a new Hybrid Attention Adapter to efficiently fuse 2D semantic cues into 3D representations with minimal trainable parameters(1.8M). Extensive experiments conducted on datasets including ScanObjectNN, ModelNet40, and ShapeNetPart demonstrate the effectiveness of our proposed paradigm. In particular, our method achieves 95.6% accuracy on ModelNet40 and attains 90.09% performance on the most challenging classification split ScanObjectNN(PB-T50-RS). Yixin Zha, Chuxin Wang, Wenfei Yang, Tianzhu Zhang 0001, Feng Wu 0001 |
ICML | 3 |
| 2025 | Exploring Semantic Masked Autoencoder for Self-supervised Point Cloud UnderstandingabstractPoint cloud understanding aims to acquire robust and general feature representations from unlabeled data. Masked point modeling-based methods have recently shown significant performance across various downstream tasks. These pre-training methods rely on random masking strategies to establish the perception of point clouds by restoring corrupted point cloud inputs, which leads to the failure of capturing reasonable semantic relationships by the self-supervised models. To address this issue, we propose Semantic Masked Autoencoder, which comprises two main components: a prototype-based component semantic modeling module and a component semantic-enhanced masking strategy. Specifically, in the component semantic modeling module, we design a component semantic guidance mechanism to direct a set of learnable prototypes in capturing the semantics of different components from objects. Leveraging these prototypes, we develop a component semantic-enhanced masking strategy that addresses the limitations of random masking in effectively covering complete component structures. Furthermore, we introduce a component semantic-enhanced prompt-tuning strategy, which further leverages these prototypes to improve the performance of pre-trained models in downstream tasks. Extensive experiments conducted on datasets such as ScanObjectNN, ModelNet40, and ShapeNetPart demonstrate the effectiveness of our proposed modules. Yixin Zha, Chuxin Wang, Wenfei Yang, Tianzhu Zhang 0001 |
IJCAI | 3 |
| 2025 | EF-3DGS: Event-Aided Free-Trajectory 3D Gaussian SplattingabstractScene reconstruction from casually captured videos has wide real-world applications. Despite recent progress, existing methods relying on traditional cameras tend to fail in high-speed scenarios due to insufficient observations and inaccurate pose estimation. Event cameras, inspired by biological vision, record pixel-wise intensity changes asynchronously with high temporal resolution and low latency, providing valuable scene and motion information in blind inter-frame intervals. In this paper, we introduce the event cameras to aid scene construction from a casually captured video for the first time, and propose Event-Aided Free-Trajectory 3DGS, called EF-3DGS, which seamlessly integrates the advantages of event cameras into 3DGS through three key components. First, we leverage the Event Generation Model (EGM) to fuse events and frames, enabling continuous supervision between discrete frames. Second, we extract motion information through Contrast Maximization (CMax) of warped events, which calibrates camera poses and provides gradient-domain constraints for 3DGS. Third, to address the absence of color information in events, we combine photometric bundle adjustment (PBA) with a Fixed-GS training strategy that separates structure and color optimization, effectively ensuring color consistency across different views. We evaluate our method on the public Tanks and Temples benchmark and a newly collected real-world dataset, RealEv-DAVIS. Our method achieves up to 3dB higher PSNR and 40% lower Absolute Trajectory Error (ATE) compared to state-of-the-art methods under challenging high-speed scenarios. Bohao Liao, Wei Zhai, Zengyu Wan, Zhixin Cheng, Wenfei Yang, Yang Cao 0010, Tianzhu Zhang 0001, Zhengjun Zha |
NeurIPS | 5 |
| 2025 | Learning Discriminative Features for Visual Tracking via Scenario Decoupling
Yinchao Ma, Qianjin Yu, Wenfei Yang, Tianzhu Zhang 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | Cross-Task Relation-Aware Consistency for Weakly Supervised Temporal Action DetectionabstractTemporal action detection aims to predict temporal boundaries and category labels of actions in untrimmed videos. In the past years, many weakly supervised temporal action detection methods have been proposed to relieve the annotation cost of fully supervised methods. Due to the discrepancy between action localization and action classification, the two-branch structure is widely adopted by existing weakly supervised methods, where the classification branch is used to predict category-wise score and the localization branch is used to predict foreground score for each segment. Under the weakly supervised setting, the model training is mainly guided by the video-level or sparse segment-level annotations. As a result, the classification branch tends to focus on the most discriminative segments while ignore less discriminative ones so as to minimize the classification cost, and the localization branch may assign high foreground scores for some negative segments. This phenomenon can severely damage the action detection performance, because the foreground scores and classification scores are combined together in the testing stage for action detection. To deal with this problem, several methods have been proposed to encourage the consistency between the classification branch and localization branch. However, these methods only consider the video-level or segment-level consistency, without considering the relation among different segments to be consistent. In this paper, we propose a Cross-Task Relation-Aware Consistency (CRC) strategy for weakly supervised temporal action detection, including an intra-video consistency module and an inter-video consistency module. The intra-video consistency module can well guarantee the relationship among segments from the same video to be consistent, and the inter-video consistency module guarantees the relationship among segments from different videos to be consistent. These two modules are complementary to each other by combining both intra-video and inter-video consistency. Experimental results show that the proposed CRC strategy can consistently improve the performance of existing weakly supervised methods, including click-level supervised methods (e.g., LACP Lee et al., 2021), video-level supervised methods (e.g., DELU Chen et al., 2022) and unsupervised methods (e.g., BaS-Net Lee et al., 2020), verifying the generality and effectiveness of the proposed method. Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001, Feng Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | SA3Det++: Side-Aware Quality Estimation for Semi-Supervised 3D Object DetectionabstractSemi-supervised 3D object detection from point cloud aims to train a detector with a small number of labeled data and a large number of unlabeled data. Among existing methods, the pseudo-label based methods have achieved superior performance, and the core lies in how to select high-quality pseudo-labels with the designed quality evaluation criterion. Despite the success of these methods, they all consider the localization and classification quality estimation from a global perspective. For localization quality, they use a global score threshold to filter out low-quality pseudo-labels and assign equal importance to each side during training, ignoring the fact that sides with different localization quality should not be treat equally. Besides, a large number of pseudo-labels are discarded due to the high global threshold, which may also contain some correctly predicted sides that are helpful for model training. For the classification quality, they usually combine the objectness score and classification confidence score to filter out pseudo-labels. The main focus of them is designing effective classification confidence evaluation metrics, neglecting the importance of predicting better objectness score. In this paper, we propose SA3Det++, a side-aware quality estimation method for semi-supervised object detection, which consists of a probabilistic side localization strategy, a side-aware quality estimation strategy, and a soft pseudo-label selection strategy. Extensive results demonstrate that the proposed method consistently outperforms the baseline methods under different scenes and evaluation criterions. Wenfei Yang, Chuxin Wang, Tianzhu Zhang 0001, Yongdong Zhang 0001, Feng Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Plane2Depth: Hierarchical Adaptive Plane Guidance for Monocular Depth EstimationabstractMonocular depth estimation aims to infer a dense depth map from a single image, which is a fundamental and prevalent task in computer vision. Many previous works have shown impressive depth estimation results through carefully designed network structures, but they usually ignore the planar information and therefore perform poorly in low-texture areas of indoor scenes. In this paper, we propose Plane2Depth, which adaptively utilizes plane information to improve depth prediction within a hierarchical framework. Specifically, in the proposed plane guided depth generator (PGDG), we design a set of plane queries as prototypes to softly model planes in the scene and predict per-pixel plane coefficients. Then the predicted plane coefficients can be converted into metric depth values with the pinhole camera model. In the proposed adaptive plane query aggregation (APGA) module, we introduce a novel feature interaction approach to improve the aggregation of multi-scale plane features in a top-down manner. Extensive experiments show that our method can achieve outstanding performance, especially in low-texture or repetitive areas. Furthermore, under the same backbone network, our method outperforms the state-of-the-art methods on the NYU-Depth-v2 dataset, achieves competitive results with state-of-the-art methods KITTI dataset and can be generalized to unseen scenes effectively. Li Liu 0067, Ruijie Zhu 0002, Jiacheng Deng 0002, Ziyang Song 0001, Wenfei Yang, Tianzhu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Purify Then Guide: A Bi-Directional Bridge Network for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation (OVSS) aims to segment an image into regions of corresponding semantic vocabularies, without being limited to a predefined set of object categories. Existing works mainly utilize large-scale vision-language models (e.g., CLIP) to leverage their superior open-vocabulary classification abilities in a two-stage manner. However, their heavy reliance on the first-stage segmentation network leaves the full potential of CLIP untapped, creating an unresolved gap between the rich pre-training knowledge and the challenging per-pixel classification task. Although the recent one-stage paradigm has further leveraged pre-trained vision knowledge from CLIP, it fails to effectively utilize text information due to the inclusion of numerous unrelated semantics in the vocabulary list. How to avoid noise interference in text information and utilize language guidance remains a Gordian knot. In this paper, we propose a bi-directional bridge network (BBN) to bridge the gap between upstream pre-trained models and downstream segmentation tasks. It first purifies the noisy text embedding and then guides semantics-vision aggregation with the purified information in a purification-then-guidance manner, thereby facilitating effective semantic utilization. Specifically, we design an optimal purification modulator to purify noisy text information via the optimal transport algorithm, and a reliable guidance modulator to integrate proper textual information into vision embedding via the designed reliable attention in an adaptive manner. Extensive experimental results on five challenging benchmarks demonstrate that our BBN performs favorably against state-of-the-art open-vocabulary semantic segmentation methods. Yuwen Pan, Rui Sun 0006, Yuan Wang 0064, Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Learning Adaptive Conceptual Prototypes for 3D Single Object Trackingabstract3D single object tracking (3D SOT) in LiDAR point clouds plays a crucial role in autonomous driving. It remains a challenging problem due to the incompleteness and the sparsity of points caused by occlusion and limited sensor capabilities. Previous methods design various modules to propagate target perceptual cues to the current frame for target localization. However, perceptual cues may contain less information for occluded or distant objects, which brings great challenges to estimating the target state accurately. To address the above limitations, we propose a novel 3D SOT framework based on the adaptive conceptual prototypes named ACPTrack, which first learns the conceptual prototype from the prior knowledge of the category structure, and then associates weak perceptual cues with the learned conceptual prototypes to improve tracking performance. The proposed ACPTrack enjoys several merits. First, we propose a universal learning method of adaptive conceptual prototype, which can quickly adapt to target-specific structure with given perceptual cues. Second, we design two modules based on the conceptual prototype for structure completion and positioning refinement, which can exploit the rich structure information of the conceptual prototype to deal with sparse and incomplete targets for robust tracking. Third, our framework is generic and compatible with various 3D trackers and brings consistent performance gains. Extensive experiments validate that our method achieves competitive performance on three large-scale datasets. Yinchao Ma, Wenfei Yang, Tianzhu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Spatio-Temporal Pyramid Keypoint Detection With Event CamerasabstractEvent cameras are bio-inspired sensors with diverse advantages, including high temporal resolution and minimal power consumption. Therefore, event cameras enjoy a wide range of applications in computer vision, among which event keypoint detection plays a vital role. However, repeatable event keypoint detection remains challenging because the lack of temporal interframe interaction leads to descriptors with limited temporal consistency, which restricts the ability to perceive keypoint motion. Besides, detectors learned at single scale features are not suitable for event keypoints with significant motion speed differences in high-speed scenarios. To deal with these problems, we propose a novel Spatio-Temporal Pyramid Keypoint Detection Network (STPNet) for event cameras via a temporally consistent descriptor learning (TCL) module and a spatially diverse detector learning (SDL) module. The proposed STPNet enjoys several merits. First, the TCL module generates temporally consistent descriptors for specific keypoint motion patterns. Second, the SDL module produces spatially diverse detectors for applications in high-speed motion scenarios. Extensive experimental results on three challenging benchmarks show that our method notably outperforms state-of-the-art event keypoint detection methods. Specifically, our STPNet can outperform the best event keypoint detection method by 0.21px in reprj. error on Event-Camera, 4% in IoU on N-Caltech101, 0.13px in reprj. error on HVGA ATIS Corner and 5.94% in matching accuracy on DSEC. Yuan Gao 0015, Tianle Ding, Xiang Liu 0020, Wenfei Yang, Tianzhu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Adaptive Prototype Learning for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised Temporal Action Localization (WTAL) aims to localize action instances with only video-level labels during training, where two primary issues are localization incompleteness and background interference. To relieve these two issues, recent methods adopt an attention mechanism to activate action instances and simultaneously suppress background ones, which have achieved remarkable progress. Nevertheless, we argue that these two issues have not been well resolved yet. On the one hand, the attention mechanism adopts fixed weights for different videos, which are incapable of handling the diversity of different videos, thus deficient in addressing the problem of localization incompleteness. On the other hand, previous methods only focus on learning the foreground attention and the attention weights usually suffer from ambiguity, resulting in difficulty of suppressing background interference. To deal with the above issues, in this paper we propose an Adaptive Prototype Learning (APL) method for WTAL, which includes two key designs: 1) an Adaptive Transformer Network (ATN) to explicitly model background and learn video-adaptive prototypes for each specific video; 2) an OT-based Collaborative (OTC) training strategy to guide the learning of prototypes and remove the ambiguity of the foreground-background separation by introducing an Optimal Transport (OT) algorithm into the collaborative training scheme between RGB and FLOW streams. These two key designs can work together to learn video-adaptive prototypes and solve the above two issues, achieving robust localization. Extensive experimental results on two standard benchmarks (THUMOS14 and ActivityNet) demonstrate that our proposed APL performs favorably against state-of-the-art methods. Tianzhu Zhang 0001, Wenfei Yang, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Rethinking Masked Representation Learning for 3D Point Cloud UnderstandingabstractSelf-supervised point cloud representation learning aims to acquire robust and general feature representations from unlabeled data. Recently, masked point modeling-based methods have shown significant performance improvements for point cloud understanding, yet these methods rely on overlapping grouping strategies (k-nearest neighbor algorithm) resulting in early leakage of structural information of mask groups, and overlook the semantic modeling of object components resulting in parts with the same semantics having obvious feature differences due to position differences. In this work, we rethink grouping strategies and pretext tasks that are more suitable for self-supervised point cloud representation learning and propose a novel hierarchical masked representation learning method, including an optimal transport-based hierarchical grouping strategy, a prototype-based part modeling module, and a hierarchical attention encoder. The proposed method enjoys several merits. First, the proposed grouping strategy partitions the point cloud into non-overlapping groups, eliminating the early leakage of structural information in the masked groups. Second, the proposed prototype-based part modeling module dynamically models different object components, ensuring feature consistency on parts with the same semantics. Extensive experiments on four downstream tasks demonstrate that our method surpasses state-of-the-art 3D representation learning methods. Furthermore, Comprehensive ablation studies and visualizations demonstrate the effectiveness of the proposed modules. Chuxin Wang, Yixin Zha, Wenfei Yang, Tianzhu Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | ER-Depth: Enhancing the Robustness of Self-Supervised Monocular Depth Estimation in Challenging ScenesabstractSelf-supervised monocular depth estimation holds significant importance in the fields of autonomous driving and robotics. However, existing methods are typically trained and evaluated on clear, sunny datasets, overlooking the impact of various adverse conditions commonly encountered in real-world applications, such as rainy weather, low visibility, and motion blur. As a result, they often struggle in challenging scenarios and produce artifacts. To address this issue, we propose ER-Depth, a novel two-stage self-supervised framework designed for robust depth estimation. In the first stage, we propose perturbation-invariant depth consistency regularization to propagate reliable supervision from standard to challenging scenes. In the second stage, we adopt the Mean Teacher paradigm for self-distillation and present a novel consistency-based pseudo-label filtering strategy to improve the quality of pseudo-labels. Extensive experiments demonstrate that our method exhibits exceptional robustness in challenging scenarios while maintaining high performance in standard scenes, significantly outperforming existing state-of-the-art methods on challenging KITTI-C, DrivingStereo, and NuScenes-Night benchmarks. Project page: https://ruijiezhu94.github.io/ERDepth_page . Ziyang Song 0001, Ruijie Zhu 0002, Chuxin Wang, Jiacheng Deng 0002, Wenfei Yang, Tianzhu Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2024 | Unifying Visual and Vision-Language Tracking via Contrastive LearningabstractSingle object tracking aims to locate the target object in a video sequence according to the state specified by different modal references, including the initial bounding box (BBOX), natural language (NL), or both (NL+BBOX). Due to the gap between different modalities, most existing trackers are designed for single or partial of these reference settings and overspecialize on the specific modality. Differently, we present a unified tracker called UVLTrack, which can simultaneously handle all three reference settings (BBOX, NL, NL+BBOX) with the same parameters. The proposed UVLTrack enjoys several merits. First, we design a modality-unified feature extractor for joint visual and language feature learning and propose a multi-modal contrastive loss to align the visual and language features into a unified semantic space. Second, a modality-adaptive box head is proposed, which makes full use of the target reference to mine ever-changing scenario features dynamically from video contexts and distinguish the target in a contrastive way, enabling robust performance in different reference settings. Extensive experimental results demonstrate that UVLTrack achieves promising performance on seven visual tracking datasets, three vision-language tracking datasets, and three visual grounding datasets. Codes and models will be open-sourced at https://github.com/OpenSpaceAI/UVLTrack. Yinchao Ma, Yuyang Tang 0001, Wenfei Yang, Tianzhu Zhang 0001, Mengxue Kang |
AAAI | 3 |
| 2024 | Instance-Adaptive and Geometric-Aware Keypoint Learning for Category-Level 6D Object Pose EstimationabstractCategory-level 6D object pose estimation aims to estimate the rotation, translation and size of unseen instances within specific categories. In this area, dense correspondence-based methods have achieved leading performance. However, they do not explicitly consider the local and global geometric information of different instances, resulting in poor generalization ability to unseen instances with significant shape variations. To deal with this problem, we propose a novel Instance-Adaptive and Geometric-Aware Keypoint Learning method for category-level 6D object pose estimation (AG-Pose), which includes two key designs: (1) The first design is an Instance-Adaptive Keypoint Detection module, which can adaptively detect a set of sparse keypoints for various instances to represent their geometric structures. (2) The second design is a Geometric-Aware Feature Aggregation module, which can efficiently integrate the local and global geometric information into keypoint features. These two modules can work together to establish robust keypoint-level correspondences for unseen instances, thus enhancing the generalization ability of the model.Experimental results on CAMERA25 and REAL275 datasets show that the proposed AG-Pose outperforms state-of-the-art methods by a large margin without category-specific shape priors. Wenfei Yang, Yuan Gao 0015, Tianzhu Zhang 0001 |
CVPR | 2 |
| 2024 | Aggregation and Purification: Dual Enhancement Network for Point Cloud Few-shot Segmentation
Guoxin Xiong, Yuan Wang 0064, Zhaoyang Li 0010, Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001 |
IJCAI | 4 |
| 2024 | MotionGS: Exploring Explicit Motion Guidance for Deformable 3D Gaussian SplattingabstractDynamic scene reconstruction is a long-term challenge in the field of 3D vision. Recently, the emergence of 3D Gaussian Splatting has provided new insights into this problem. Although subsequent efforts rapidly extend static 3D Gaussian to dynamic scenes, they often lack explicit constraints on object motion, leading to optimization difficulties and performance degradation. To address the above issues, we propose a novel deformable 3D Gaussian splatting framework called MotionGS, which explores explicit motion priors to guide the deformation of 3D Gaussians. Specifically, we first introduce an optical flow decoupling module that decouples optical flow into camera flow and motion flow, corresponding to camera movement and object motion respectively. Then the motion flow can effectively constrain the deformation of 3D Gaussians, thus simulating the motion of dynamic objects. Additionally, a camera pose refinement module is proposed to alternately optimize 3D Gaussians and camera poses, mitigating the impact of inaccurate camera poses. Extensive experiments in the monocular dynamic scenes validate that MotionGS surpasses state-of-the-art methods and exhibits significant superiority in both qualitative and quantitative results. Project page: https://ruijiezhu94.github.io/MotionGS_page. Ruijie Zhu 0002, Yanzhe Liang, Hanzhi Chang, Jiacheng Deng 0002, Jiahao Lu 0001, Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001 |
NeurIPS | 6 |
| 2024 | DN-4DGS: Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene RenderingabstractDynamic scenes rendering is an intriguing yet challenging problem. Although current methods based on NeRF have achieved satisfactory performance, they still can not reach real-time levels. Recently, 3D Gaussian Splatting (3DGS) has garnered researchers' attention due to their outstanding rendering quality and real-time speed. Therefore, a new paradigm has been proposed: defining a canonical 3D gaussians and deforming it to individual frames in deformable fields. However, since the coordinates of canonical 3D gaussians are filled with noise, which can transfer noise into the deformable fields, and there is currently no method that adequately considers the aggregation of 4D information. Therefore, we propose Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene Rendering (DN-4DGS). Specifically, a Noise Suppression Strategy is introduced to change the distribution of the coordinates of the canonical 3D gaussians and suppress noise. Additionally, a Decoupled Temporal-Spatial Aggregation Module is designed to aggregate information from adjacent points and frames. Extensive experiments on various real-world datasets demonstrate that our method achieves state-of-the-art rendering quality under a real-time level. Code is available at https://github.com/peoplelu/DN-4DGS. Jiahao Lu 0001, Jiacheng Deng 0002, Ruijie Zhu 0002, Yanzhe Liang, Wenfei Yang, Tianzhu Zhang 0001 |
NeurIPS | 5 |
| 2024 | Reliable Phrase Feature Mining for Hierarchical Video-Text RetrievalabstractVideo-Text Retrieval is a fundamental task in multi-modal understanding and has attracted increasing attention from both academia and industry communities in recent years. Generally, video inherently contains multi-grained semantic and each video corresponds to several different texts, which is challenging. Previous best-performing methods adopt video-sentence, phrase-phrase, and frame-word interactions simultaneously. Different from word/frame features that can be obtained directly, phrase features need to be adaptively aggregated from correlative word/frame features, which makes it very demanding. However, existing method utilizes simple intra-modal self-attention to generate phrase features without considering the following three aspects: cross-modality semantic correlation, phrase generation noise and diversity. In this paper, we propose a novel Reliable Phrase Mining model (RPM) to construct reliable phrase features and conduct hierarchical cross-modal interactions for video-text retrieval. The proposed RPM model enjoys several merits. Firstly, to guarantee the semantic consistency between video phrases and text phrases, we propose a set of modality-shared prototypes as the joint query to aggregate the semantically related frame/word features into adaptive-grained phrase features. Secondly, to deal with the phrase generation noise, the proposed denoised decoder module is responsible for obtaining more reliable similarity between prototypes and frame/word features. Specifically, not only the correlation between frame/word features and prototypes, but also the correlation among prototypes, should be taken into account when calculating the similarity. Furthermore, to encourage different prototypes to focus on different semantic information, we design a prototype contrastive loss whose core idea is enabling phrases produced by the same prototype to be more similar than those produced by different prototypes. Extensive experiment results demonstrate that the proposed method performs favorably on three benchmark datasets, including MSR-VTT, MSVD, and LSMDC. Huakai Lai, Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Multi-Modal Attribute Prompting for Vision-Language ModelsabstractPre-trained Vision-Language Models (VLMs), like CLIP, exhibit strong generalization ability to downstream tasks but struggle in few-shot scenarios. Existing prompting techniques primarily focus on global text and image representations, yet overlooking multi-modal attribute characteristics. This limitation hinders the model’s ability to perceive fine-grained visual details and restricts its generalization ability to a broader range of unseen classes. To address this issue, we propose a Multi-modal Attribute Prompting method (MAP) by jointly exploring textual attribute prompting, visual attribute prompting, and attribute-level alignment. The proposed MAP enjoys several merits. First, we introduce learnable visual attribute prompts enhanced by textual attribute semantics to adaptively capture visual attributes for images from unknown categories, boosting fine-grained visual perception capabilities for CLIP. Second, the proposed attribute-level alignment complements the global alignment to enhance the robustness of cross-modal alignment for open-vocabulary objects. To our knowledge, this is the first work to establish cross-modal attribute-level alignment for CLIP-based few-shot adaptation. Extensive experimental results on 11 datasets demonstrate that our method performs favorably against state-of-the-art approaches. Xin Liu 0089, Wenfei Yang, Tianzhu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Hierarchy-Aware Interactive Prompt Learning for Few-Shot ClassificationabstractFew-Shot Learning (FSL) leverages prior knowledge and generalization strategies to quickly adapt to new tasks or recognize new objects with minimal input. Recently, CLIP-based methods, aided by contrastive language-image pre-training, have demonstrated impressive few-shot performance. However, these methods solely employ fixed-length uni-modal prompts at the initial encoder layer, neglecting the multi-level adaptation and cross-modal interaction for the intermediate features. To address this issue, we propose Hierarchy-Aware Interactive Prompt Learning (HIPL), by jointly exploring hierarchical prompt learning and cross-modal prompt interaction for CLIP-based FSC. The proposed HIPL enjoys several merits. First, we design a hierarchical prompt aggregation module to progressively generate higher-level prompts via the attention mechanisms, equipping the CLIP with hierarchical adaptation capability. Second, a cross-modal prompt interaction module is proposed to facilitate deep interaction between stage-wise prompts, ensuring mutual synergy between vision and textual features. To the best of our knowledge, this is the first work to learn multi-level prompts by progressive aggregation. Our extensive experiments demonstrate that HIPL outperforms previous methods in few-shot classification and base-to-new generalization. Our code is available athttps://github.com/Yxt1212/HIPL Xiaotian Yin, Wenfei Yang, Tianzhu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization aims to localize and recognize actions in untrimmed videos with only video-level category labels during training. Without instance-level annotations, most existing methods follow the Segment-based Multiple Instance Learning (S-MIL) framework, where the predictions of segments are supervised by the labels of videos. However, the objective for acquiring segment-level scores during training is not consistent with the target for acquiring proposal-level scores during testing, leading to suboptimal results. To deal with this problem, we propose a novel Proposal-based Multiple Instance Learning (P-MIL) framework that directly classifies the candidate proposals in both the training and testing stages, which includes three key designs: 1) a surrounding contrastive feature extraction module to suppress the discriminative short proposals by considering the surrounding contrastive information, 2) a proposal completeness evaluation module to inhibit the low-quality proposals with the guidance of the completeness pseudo labels, and 3) an instance-level rank consistency loss to achieve robust detection by leveraging the complementarity of RGB and FLOW modalities. Extensive experimental results on two challenging benchmarks including THUMOS14 and ActivityNet demonstrate the superior performance of our method. Our code is available at github.com/RenHuan1999/CVPR2023_P-MIL. Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001 |
CVPR | 2 |
| 2023 | Channel Measurements at 140 and 220 GHz in an Outdoor Street Canyon EnvironmentabstractTerahertz (THz) communication is considered as one of the potential candidate technologies in the sixth generation (6G) wireless systems. This paper introduces channel measure-ments in two sub-THz bands, i.e., 140- and 220-GHz bands, in an outdoor street canyon environment for both line-of-sight (LoS) and non-line-of-sight (NLoS) scenarios with a frequency-domain vector network analyzer (VNA)-based sounder. Based on the measurement results, we computed and analyzed the statistical features of wireless propagation channels, including path loss, root mean square (RMS) delay spreads (DSs), azimuth spreads of arrival (ASA), and elevation spreads of arrival (ESA). Moreover, we observe the birth and death of clusters over a straight trajectory under the NLoS condition. A high-resolution param-eter estimation algorithm, i.e., the space-alternating generalized expectation-maximization (SAGE) algorithm, was employed to eliminate the effects of antenna patterns and the density-based spatial clustering of applications with noise (DBSCAN) algorithm was used to clustered the multipath components (MPCs). The obtained statistical properties of measured channels and obser-vations on channel evolution can be employed in THz channel modeling and system design. Wenfei Yang, Ziming Yu, Yi Chen 0013, Mate Boban, Tommaso Zugno, Jian Li 0058 |
GLOBECOM | 1 |
| 2023 | Integrating the Sensing and Radio Communications Channel Modelling From Radar Mutual InterferenceabstractThe current growing interest in integrated sensing and communications (ISAC) for the next generation of radio access networks towards 6G is opening new challenges on the channel estimation and modelling. New frequency bands and novel techniques for joining the sensing and the transmission of information over the same waveform are becoming a hot topic for research. This paper proposes a generalist dual channel model description for ISAC and a technique to relate the sensing and the communications channel, which is described analytically and verified with measurements. The technique is based on exploiting the mutual interference between radar sensors to estimate the multipath communications channel between them. Narcís Cardona, J. Samuel Romero, Wenfei Yang |
ICASSP | 3 |
| 2023 | Not Every Side Is Equal: Localization Uncertainty Estimation for Semi-Supervised 3D Object DetectionabstractSemi-supervised 3D object detection from point cloud aims to train a detector with a small number of labeled data and a large number of unlabeled data. The core of existing methods lies in how to select high-quality pseudo-labels using the designed quality evaluation criterion. However, these methods treat each pseudo bounding box as a whole and assign equal importance to each side during training, which is detrimental to model performance due to many sides having poor localization quality. Besides, existing methods filter out a large number of low-quality pseudo-labels, which also contain some correct regression values that can help with model training. To address the above issues, we propose a side-aware framework for semi-supervised 3D object detection consisting of three key designs: a 3D bounding box parameterization method, an uncertainty estimation module, and a pseudo-label selection strategy. These modules work together to explicitly estimate the localization quality of each side and assign different levels of importance during the training phase. Extensive experiment results demonstrate that the proposed method can consistently outperform baseline models under different scenes and evaluation criteria. Moreover, our method achieves state-of-the-art performance on three datasets with different labeled ratios. Chuxin Wang, Wenfei Yang, Tianzhu Zhang 0001 |
ICCV | 2 |
| 2023 | Joint Estimation on the Reflector Velocity and Normal Direction Through NLOS Echo SignalsabstractThis paper studies sensing through the non-line-of-sight (NLOS) paths to obtain the channel information. We assume that the reflector is in the near-field of the base station (BS), while the user equipment (UE) is in the far-field. We demonstrate that the normal direction of the reflector is necessary for the localization of the UE and the estimation of the Downlink (DL) angle of arrival (AOA). Moreover, the movement of the reflectors could cause the channel information to be fast varying. To solve these problems, we propose an algorithm to jointly estimate the magnitude of velocity and the normal direction of the reflector. We first derive the expression of the NLOS echo signal considering the dual Doppler effect caused by the transmitter, the UE and the reflector. We then propose an optimization problem to estimate the parameters, and propose to use the Levenberg-Marquardt (LM) method to solve the problem. We derive the Cramer-Rao lower bound (CRLB) for parameter estimations of the proposed algorithm. Simulation results validates the performance of our proposed algorithm for varying signal-to-noise-ratios (SNRs) and snapshot durations. Tianxiao Zhao, Wenfei Yang |
VTC2023-Spring | 3 |
| 2023 | Uncertainty Guided Collaborative Training for Weakly Supervised and Unsupervised Temporal Action LocalizationabstractIn weakly supervised (WSAL) and unsupervised temporal action localization (UAL), the target is to simultaneously localize temporal boundaries and identify category labels of actions with only video-level category labels (WSAL) or category numbers in a dataset (UAL) during training. Among existing methods, attention based methods have achieved superior performance in both tasks by highlighting action segments with foreground attention weights. However, without the segment-level supervision on the attention weight learning, the quality of the attention weight hinders the performance of these methods. In this paper, we propose a novel Uncertainty Guided Collaborative Training (UGCT) strategy to alleviate this problem, which mainly includes two key designs: (1) The first design is an online pseudo label generation module, in which the RGB and FLOW streams work collaboratively to learn from each other. (2) The second design is an uncertainty aware learning module, which can mitigate the noise in the generated pseudo labels. These two designs work together to promote the model performance effectively and efficiently by exchanging information between RGB and FLOW streams. Extensive experimental results on two benchmark datasets with three attention based methods demonstrate the effectiveness of the proposed method, e.g, more than 7.0% performance gain for mAP@IoU=0.5 on THUMOS14 dataset. Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001, Feng Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Diverse Complementary Part Mining for Weakly Supervised Object LocalizationabstractWeakly Supervised Object Localization (WSOL) aims to localize objects with only image-level labels, which has better scalability and practicability than fully supervised methods in the actual deployment. However, a common limitation for available techniques based on classification networks is that they only highlight the most discriminative part of the object, not the entire object. To alleviate this problem, we propose a novel end-to-end part discovery model (PDM) to learn multiple discriminative object parts in a unified network for accurate object localization and classification. The proposed PDM enjoys several merits. First, to the best of our knowledge, it is the first work to directly model diverse and robust object parts by exploiting part diversity, compactness, and importance jointly for WSOL. Second, three effective mechanisms including diversity, compactness, and importance learning mechanisms are designed to learn robust object parts. Therefore, our model can exploit complementary spatial information and local details from the learned object parts, which help to produce precise bounding boxes and discriminate different object categories. Extensive experiments on two standard benchmarks demonstrate that our PDM performs favorably against state-of-the-art WSOL approaches. Tianzhu Zhang 0001, Wenfei Yang, Jian Zhao 0006, Yongdong Zhang 0001, Feng Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Action Unit Memory Network for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization aims to detect and localize actions in untrimmed videos with only video-level labels during training. However, without frame-level annotations, it is challenging to achieve localization completeness and relieve background interference. In this paper, we present an Action Unit Memory Network (AUMN) for weakly supervised temporal action localization, which can mitigate the above two challenges by learning an action unit memory bank. In the proposed AUMN, two attention modules are designed to update the memory bank adaptively and learn action units specific classifiers. Furthermore, three effective mechanisms (diversity, homogeneity and sparsity) are designed to guide the updating of the memory network. To the best of our knowledge, this is the first work to explicitly model the action units with a memory network. Extensive experimental results on two standard benchmarks (THUMOS14 and ActivityNet) demonstrate that our AUMN performs favorably against state-of-the-art methods. Specifically, the average mAP of IoU thresholds from 0.1 to 0.5 on the THUMOS14 dataset is significantly improved from 47.0% to 52.1%. Tianzhu Zhang 0001, Wenfei Yang, Jingen Liu, Tao Mei 0001, Feng Wu 0001, Yongdong Zhang 0001 |
CVPR | 3 |
| 2021 | Uncertainty Guided Collaborative Training for Weakly Supervised Temporal Action DetectionabstractWeakly supervised temporal action detection aims to localize temporal boundaries of actions and identify their categories simultaneously with only video-level category labels during training. Among existing methods, attention based methods have achieved superior performance by separating action and non-action segments. However, without the segment-level ground-truth supervision, the quality of the attention weight hinders the performance of these methods. To alleviate this problem, we propose a novel Uncertainty Guided Collaborative Training (UGCT) strategy, which mainly includes two key designs: (1) The first design is an online pseudo label generation module, in which the RGB and FLOW streams work collaboratively to learn from each other. (2) The second design is an uncertainty aware learning module, which can mitigate the noise in the generated pseudo labels. These two designs work together to promote the model performance effectively and efficiently by imposing pseudo label supervision on attention weight learning. Experimental results on three state-of-the-art attention based methods demonstrate that the proposed training strategy can significantly improve the performance of these methods, e.g., more than 4% for all three methods in terms of mAP@IoU=0.5 on the THUMOS14 dataset. Wenfei Yang, Tianzhu Zhang 0001, Xiaoyuan Yu, Qi Tian 0001, Yongdong Zhang 0001, Feng Wu 0001 |
CVPR | 1 |
| 2021 | Multi-Scale Structure-Aware Network for Weakly Supervised Temporal Action DetectionabstractWeakly supervised temporal action detection has better scalability and practicability than fully supervised action detection in reality deployment. However, it is difficult to learn a robust model without temporal action boundary annotations. In this paper, we propose an en-to-end Multi-Scale Structure-Aware Network (MSA-Net) for weakly supervised temporal action detection by exploring both the global structure information of a video and the local structure information of actions. The proposed SA-Net enjoys several merits. First, to localize actions with different durations, each video is encoded into feature representations with different temporal scales. Second, based on the multi-scale feature representation, the proposed model has designed two effective structure modeling mechanisms including global structure modeling and local structure modeling, which can effectively learn discriminative structure aware representations for robust and complete action detection. To the best of our knowledge, this is the first work to fully explore the global and local structure information in a unified deep model for weakly supervised action detection. And extensive experimental results on two benchmark datasets demonstrate that the proposed MSA-Net performs favorably against state-of-the-art methods. Wenfei Yang, Tianzhu Zhang 0001, Zhendong Mao 0001, Yongdong Zhang 0001, Qi Tian 0001, Feng Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Local Correspondence Network for Weakly Supervised Temporal Sentence GroundingabstractWeakly supervised temporal sentence grounding has better scalability and practicability than fully supervised methods in real-world application scenarios. However, most of existing methods cannot model the fine-grained video-text local correspondences well and do not have effective supervision information for correspondence learning, thus yielding unsatisfying performance. To address the above issues, we propose an end-to-end Local Correspondence Network (LCNet) for weakly supervised temporal sentence grounding. The proposed LCNet enjoys several merits. First, we represent video and text features in a hierarchical manner to model the fine-grained video-text correspondences. Second, we design a self-supervised cycle-consistent loss as a learning guidance for video and text matching. To the best of our knowledge, this is the first work to fully explore the fine-grained correspondences between video and text for temporal sentence grounding by using self-supervised learning. Extensive experimental results on two benchmark datasets demonstrate that the proposed LCNet significantly outperforms existing weakly supervised methods. Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001, Feng Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Tracking Assisted Faster Video Object DetectionabstractRecent approaches have achieved great success on still image object detection. Despite the high accuracy, directly applying image object detectors for video object detection is rather slow. Inspired from the fact that object tracking is much more efficient than object detection, we propose to combine object detection and tracking for fast video object detection. Computational expensive detection network is applied on sparsely arranged key frames, while proposals of non-key frames are obtained through tracking and regression of previous frame's proposals. Assisted with an adaptive key-frame arrangement module, our method can adaptively decide whether to track or to detect based on tracking quality. Extensive experiments show that the proposed method can significantly boost detection speed with a rather small drop in detection accuracy. Wenfei Yang, Bin Liu 0016, Weihai Li, Nenghai Yu |
ICME | 1 |
| 2017 | Key-Region Representation Learning for Anomaly Detection
Wenfei Yang, Bin Liu 0016, Nenghai Yu |
ICIG (1) | 1 |