Wei Tang 0016

dblp:58/1874-16 · DBLP profile ↗
← Back
66ranked-venue papers
6as first author
48since 2021 · last 2026
0000-0001-8879-8325ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 4 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 4 first-author · 33 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author
YearPublicationVenuePosition
2026 FairScene: Learning Class-Disentangled 2D/3D Representations for Semantic Scene Completion
abstract
Semantic Scene Completion (SSC) aims to predict the semantic occupancy of each voxel within a 3D scene using sensor data, a critical task for autonomous driving and robotics. Despite recent progress, camera-based SSC remains challenging due to various difficulties, including voxel class imbalance, occlusion, and depth ambiguity. This paper introduces FairScene, a novel approach that learns class-disentangled 2D/3D representations to improve SSC. By ensuring balanced representations across classes, FairScene mitigates the dominance of majority classes and promotes fairer voxel categorization. Additionally, FairScene explicitly models spatial dependencies between different classes through a novel inter-class occupancy reasoning mechanism. Such explicit modeling helps alleviate occlusion and depth ambiguities in SSC. To address the scarcity of SSC training data, we propose OccMix, a novel augmentation strategy that generalizes MixUp from 2D to 2.5D and 3D metric spaces while maintaining geometric consistency. Extensive quantitative and qualitative experiments demonstrate that FairScene out-performs prior methods on both the SemanticKITTI and SSCBench-KITTI-360 benchmarks. The code is available at https://github.com/DianJJ/FairScene.
Dian Jia, Pei Yu, Wei Tang 0016
WACV3
2026 Modeling and Learning Multiple Hypotheses for Monocular 3D Object Detection
abstract
Detecting objects in 3D space using a monocular image is inherently a highly ill-posed problem: multiple plausible 3D bounding boxes can explain the same 2D observation of an object. Existing approaches typically follow a single-point prediction paradigm, failing to capture this multimodal nature and often regressing to an implausible mean solution. This paper introduces MonoMH, a novel multi-hypothesis framework for monocular 3D object detection. By explicitly modeling and learning the multimodal distribution of plausible 3D object configurations, MonoMH not only significantly improves detection performance but also provides richer information to support downstream decision-making. MonoMH introduces three key innovations: (1) a novel multi-hypothesis predictor that leverages spatially diverse features across different windows within an RoI to generate a rich variety of hypotheses without increasing model complexity; (2) a new multi-hypothesis learning approach that derives diverse and relevant hypotheses from single-modal ground truth by integrating uncertainty modeling with Best-of-Many learning; and (3) a hypothesis filtering mechanism that enhances detection capability by dynamically retaining a variable number of plausible hypotheses based on each object’s uncertainty. Experimental results demonstrate the effectiveness of our approach. Notably, MonoMH achieves 29.12/20.88/17.93 Car AP3D(easy/mod./hard) on the KITTI test set, significantly outperforming previous state-of-the-art methods. The code can be found at https://github.com/HyeonjeongPark37/MonoMH.
Hyeonjeong Park, Peixi Xiong, Pei Yu, Wei Tang 0016
WACV4
2026 Sparse Trajectory Prediction
abstract
Pedestrian trajectory prediction is crucial for ensuring safe decision-making in intelligent robotic systems. While this task demands real-time performance, previous works have primarily focused on improving prediction accuracy, often neglecting efficiency. Dense predictions with time-consuming post-clustering steps and global interactions with quadratic computational complexity result in a trade-off between accuracy and speed. In this paper, we propose a novel Sparse Trajectory Prediction (STP) model that aims to achieve both high accuracy and real-time speed by following an efficient principle: leveraging sparse structures to achieve global effects. STP instantiates this principle within a transformer-style encoder-decoder framework. In the encoder, STP introduces irregular interaction, which builds sparse interactions with dynamic interactive positions, reducing computational complexity to linearithmic/linear while maintaining global interaction. In the decoder, STP applies an early-sparsity strategy to generate sparse motion modes that represent global motion behaviors. These modes are shared across all predictions, eliminating redundant computations. By harnessing the expressive power of transformers, STP maps these sparse motion modes into multimodal future trajectories, significantly improving prediction speed while ensuring accuracy. Experimental results on four commonly used datasets demonstrate that STP maximizes both accuracy and prediction speed, achieving state-of-the-art performance and significantly improving prediction speed by about $100 \times$100× - $150 \times$150× to satisfy the real-time demand.
Liushuai Shi, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 RAMP: Iterative Refinement and Adaptive Multi-granularity Perception for embodied dialog localization
Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016
Pattern Recognit.6
2026 RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation
abstract
Recent image-goal navigation (ImageNav) methods learn a perception-action policy by separately capturing semantic features of the goal and egocentric images, then passing them to a policy network. However, challenges remain: (1) Semantic features often fail to provide accurate directional information, leading to superfluous actions, and (2) performance drops significantly when viewpoint inconsistencies arise between training and application. To address these challenges, we propose RSRNav, a simple yet effective method that reasons spatial relationships between the goal and current observations as navigation guidance. Specifically, we model the spatial relationship by constructing correlations between the goal and current observations, which are then passed to the policy network for action prediction. These correlations are progressively refined using fine-grained cross-correlation and direction-aware correlation for more precise navigation. Extensive evaluation of RSRNav on three benchmark datasets demonstrates superior navigation performance, particularly in the "user-matched goal" setting, highlighting its potential for real-world applications. Code: https://github.com/ qinzheng2000/RSRNav.git.
Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016
IEEE Trans. Circuits Syst. Video Technol.6
2026 MoDe-Track: Robust Multi-Object Tracking With Motion Decoupling in UAV Videos
abstract
Multi-Object Tracking (MOT) in Unmanned Aerial Vehicle (UAV) scenarios is characterized by frequent and abrupt camera motion, which presents two unique challenges: nonlinear motion and appearance degradation. Traditional motion models, designed for smooth and consistent motion, struggle to capture the complex background motion patterns caused by UAV movement; while appearance-based methods are easily disrupted by occlusion and blur, leading to unreliable associations. Even though dense optical flow is widely utilized to model complex motion patterns, the entanglement of background and object motion often introduces interference, limiting its effectiveness. To this end, we propose MoDe-Track, a unified framework that explicitly decouples scene motion into background and object components, and serves as an elegant integration of three robust components. Specifically, the Scene Motion Decomposition (SMD) module decouples the motion into the background and object components based on robust principal component analysis, serving as the foundation for motion compensation and feature propagation. Afterwards, the Background Motion Compensation (BMC) uses the decomposed background flow to estimate and compensate for camera motion, mitigating the effects of nonlinear motion. Finally, the Foreground-guided Feature Propagation (FFP) module uses the decoupled object flow to guide feature propagation across frames, achieving temporal consistency and enhancing robustness against occlusion and motion blur. Extensive experimental results on two benchmarks, VisDrone2019 and UAVDT, demonstrate that MoDe-Track consistently outperforms current multi-object tracking methods. We achieve 56.0% MOTA on VisDrone2019 and 56.2% MOTA on UAVDT, reaching the state-of-the-art among existing methods.
Zixuan Song, Sanping Zhou, Wei Tang 0016, Le Wang 0003
IEEE Trans. Multim.4
2025 Learning Partonomic 3D Reconstruction from Image Collections
abstract
Reconstructing the 3D shape of an object from a single-view image is a fundamental task in computer vision. Recent advances in differentiable rendering have enabled 3D reconstruction from image collections using only 2D annotations. However, these methods mainly focus on whole-object reconstruction and overlook object partonomy, which is essential for intelligent agents interacting with physical environments. This paper aims at learning partonomic 3D reconstruction from collections of images with only 2D annotations. Our goal is not only to reconstruct the shape of an object from a single-view image but also to decompose the shape into meaningful semantic parts. To handle the expanded solution space and frequent part occlusions in single-view images, we introduce a novel approach that represents, parses, and learns the structural compositionality of 3D objects. This approach comprises: (1) a compact and expressive compositional representation of object geometry, achieved through disentangled modeling of large shape variations, constituent parts, and detailed part deformations as multi-granularity neural fields; (2) a part transformer that recovers precise partonomic geometry and handles occlusions, through effective part-to-pixel grounding and part-to-part relational modeling; and (3) a 2D-supervised learning method that jointly learns the compositional representation and part transformer, by bridging object shape and parts, image synthesis, and differentiable rendering. Extensive experiments on ShapeNetPart, PartNet, and CUB-200-2011 demonstrate the effectiveness of our approach on both overall and partonomic reconstruction. Code, models, and data are avaliable at https://github.com/XiaoqianRuan1/Partonomic_Reconstruction.
Xiaoqian Ruan, Pei Yu, Dian Jia, Hyeonjeong Park, Peixi Xiong, Wei Tang 0016
CVPR6
2025 PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic Manipulation
abstract
Robotic manipulation based on visual observations and natural language instructions is a long-standing challenge in robotics. Yet prevailing approaches model action distribution by adopting explicit or implicit representations, which often struggle to achieve a trade-off between accuracy and efficiency. In response, we propose PDFactor, a novel framework that models action distribution with a hybrid triplane representation. In particular, PDFactor decomposes 3D point cloud into three orthogonal feature planes and leverages a tri-perspective view transformer to produce dense cubic features as a latent diffusion field aligned with observation space representing 6-DoF action probability distribution at an arbitrary location. We employ a small denoising network conceptually as both a parameterized loss function measuring the quality of the learned latent features and an action gradient decoder to sample actions from the latent diffusion field during inference. This design enables our PDFactor to benefit from spatial awareness of explicit representation and arbitrary resolution of implicit representation, rendering it with manipulation accuracy, inference efficiency, and model scalability. Experiments demonstrate that PDFactor outperforms state-of-the-art approaches across a diverse range of manipulation tasks in RLBench simulation. Moreover, PDFactor can effectively learn multi-task policies from a limited number of human demonstrations, achieving promising accuracy in a variety of real-world manipulation tasks.
Jingyi Tian, Le Wang 0003, Sanping Zhou, Haowen Sun 0003, Wei Tang 0016
CVPR7
2025 FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation
abstract
Robotic manipulation in high-precision tasks is essential for numerous industrial and real-world applications where accuracy and speed are required. Yet current diffusion-based policy learning methods generally suffer from low computational efficiency due to the iterative denoising process during inference. Moreover, these methods do not fully explore the potential of generative models for enhancing information exploration in 3D environments. In response, we propose FlowRAM, a novel framework that leverages generative models to achieve region-aware perception, enabling efficient multimodal information processing. Specifically, we devise a Dynamic Radius Schedule, which allows adaptive perception, facilitating transitions from global scene comprehension to fine-grained geometric details. Furthermore, we integrate state space models to integrate multimodal information, while preserving linear computational complexity. In addition, we employ conditional flow matching to learn action poses by regressing deterministic vector fields, simplifying the learning process while maintaining performance. We verify the effectiveness of the FlowRAM in the RLBench, an established manipulation benchmark, and achieve state-of-the-art performance. The results demonstrate that FlowRAM achieves a remarkable improvement, particularly in high-precision tasks, where it outperforms previous methods by 12.0% in average success rate. Additionally, FlowRAM is able to generate physically plausible actions for a variety of real-world tasks in less than 4 time steps, significantly increasing inference speed.
Le Wang 0003, Sanping Zhou, Jingyi Tian, Haowen Sun 0003, Wei Tang 0016
CVPR7
2025 Towards Precise Embodied Dialogue Localization via Causality Guided Diffusion
abstract
Embodied localization based on vision and natural language dialogues presents a persistent challenge in embodied intelligence. Existing methods often approach this task as an image translation problem, leveraging encoder-decoder architectures to predict heatmaps. However, these methods frequently experience a deficiency in accuracy, largely due to their heavy reliance on resolution. To address this issue, we introduce CGD, a novel framework that utilizes causality guided diffusion model to directly model coordinate distributions. Specifically, CGD employs a denoising network to regress coordinates, while integrating causal learning modules, namely back-door adjustment (BDA) and front-door adjustment (FDA) to mitigate confounders during the diffusion process. This approach reduces the dependency on high resolution for improving accuracy, while effectively minimizing spurious correlations, thereby promoting unbiased learning. By guiding the denoising process with causal adjustments, CGD offers flexible control over intensity, ensuring seamless integration with diffusion models. Experimental results demonstrate that CGD outperforms state-of-the-art methods across all metrics. Additionally, we also evaluate CGD in a multi-shot setting, achieving consistently high accuracy.
Le Wang 0003, Sanping Zhou, Jingyi Tian, Gang Hua 0001, Wei Tang 0016
CVPR8
2025 SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World Models
abstract
World models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions due to disrupted spatial structure, inefficient decoding, and inadequate motion modeling. In response, we propose Scale-wise Autoregression with Motion PrOmpt (SAMPO), a hybrid framework that combines visual autoregressive modeling for intra-frame generation with causal modeling for next-frame generation. Specifically, SAMPO integrates temporal causal decoding with bidirectional spatial attention, which preserves spatial locality and supports parallel decoding within each scale. This design significantly enhances both temporal consistency and rollout efficiency. To further improve dynamic scene understanding, we devise an asymmetric multi-scale tokenizer that preserves spatial details in observed frames and extracts compact dynamic representations for future frames, optimizing both memory usage and model performance. Additionally, we introduce a trajectory-aware motion prompt module that injects spatiotemporal cues about object and robot trajectories, focusing attention on dynamic regions and improving temporal consistency and physical realism. Extensive experiments show that SAMPO achieves competitive performance in action-conditioned video prediction and model-based control, improving generation quality with 4.4× faster inference. We also evaluate SAMPO's zero-shot generalization and scaling behavior, demonstrating its ability to generalize to unseen tasks and benefit from larger model sizes.
Jingyi Tian, Le Wang 0003, Zhimin Liao, Huaiyi Dong, Sanping Zhou, Wei Tang 0016, Gang Hua 0001
NeurIPS9
2025 AFC-RNN: Adaptive Forgetting-Controlled Recurrent Neural Network for Pedestrian Trajectory Prediction
abstract
Pedestrian trajectory prediction plays a crucial and fundamental role in many computer vision tasks. Most existing works utilize recurrent neural networks to extract temporal features from trajectories because their recursive structure is inherently well-suited for time series data. However, previous methods overlook the forgetting characteristics of pedestrians when modeling historical trajectories, which may cause the model to focus on the wrong positions of historical information. In this paper, we propose a simple yet effective Adaptive Forgetting-Controlled Recurrent Neural Network (AFC-RNN) for pedestrian trajectory prediction. The core idea of AFC-RNN is a novel Adaptive Forgetting Controller (AFC), which controls the forgetting degree of the historical information at each time step explicitly and adaptively. Specifically, AFC first learns memory factors for each time step based on the temporal correlation of observed trajectories using the self-attention mechanism. Then, AFC-RNN applies these memory factors to regulate the forgetting degree of observed features at each time step from RNN. Extensive experiments and ablation studies on ETH, UCY, SDD, and NBA datasets demonstrate that our method outperforms existing state-of-the-art approaches. Additionally, we provide a mathematical analysis to demonstrate the superiority of our adaptive forgetting strategy in the AFC-RNN over traditional RNNs for trajectory forgetting modeling.
Yonghao Dong, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001, Changyin Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Towards unsupervised learning of joint facial landmark detection and head pose estimation
Zhiming Zou, Dian Jia, Wei Tang 0016
Pattern Recognit.3
2025 Meta Pairwise Relationship Distillation for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) is challenging due to the lack of ground-truth labels. Most existing methods rely on pseudo labels estimated via iterative clustering and thus are highly susceptible to performance penalties incurred by the inaccurate estimated number of clusters. Alternatively, we utilize the sample pairs with pairwise pseudo labels to guide the feature learning to avoid the dilemma of determining cluster numbers. In this article, we propose a meta pairwise relationship distillation (MPRD) method that incorporates a graph convolutional network (GCN) to provide high-fidelity pairwise relationships to supervise the model training. A small amount of metadata with very-confidence pairwise relationships and the unlabeled pairs with the provided pseudo pairwise relationships participate in the GCN training. Besides, we introduce a hard sample deduction (HSD) module to timely mine the sample pairs with error-prone pairwise pseudo labels to mitigate the misled optimization by noisy labels. Furthermore, since the features of each positive pair represent the same person, we design a positive pair alignment (PPA) module to reduce the redundant information in each feature, which is achieved by minimizing the difference between each positive pair's feature distributions. Extensive experiments on the Market-1501, DukeMTMC-reID, and MSMT17 datasets show that our method outperforms the state-of-the-art unsupervised methods.
Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Towards Generalizable Multi-Object Tracking
abstract
Multi-Object Tracking (MOT) encompasses various tracking scenarios, each characterized by unique traits. Ef-fective trackers should demonstrate a high degree of gen-eralizability across diverse scenarios. However, existing trackers struggle to accommodate all aspects or necessi-tate hypothesis and experimentation to customize the asso-ciation information (motion and/or appearance) for a given scenario, leading to narrowly tailored solutions with limited generalizability. In this paper, we investigate the factors that influence trackers' generalization to different scenar-ios and concretize them into a set of tracking scenario at-tributes to guide the design of more generalizable trackers. Furthermore, we propose a “point-wise to instance-wise relation” framework for MOT, i.e., GeneralTrack, which can generalize across diverse scenarios while eliminating the need to balance motion and appearance. Thanks to its supe-rior generalizability, our proposed GeneralTrack achieves state-of-the-art performance on multiple benchmarks and demonstrates the potential for domain generalization.
Le Wang 0003, Sanping Zhou, Panpan Fu, Gang Hua 0001, Wei Tang 0016
CVPR6
2024 Analysis-by-Synthesis Transformer for Single-View 3D Reconstruction
Dian Jia, Xiaoqian Ruan, Zhiming Zou, Le Wang 0003, Wei Tang 0016
ECCV (21)6
2024 Learning Anomalies with Normality Prior for Unsupervised Video Anomaly Detection
Haoyue Shi 0002, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016
ECCV (6)5
2024 Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
abstract
Cross-lingual cross-modal retrieval (CCR) aims to retrieve visually relevant content based on non-English queries, without relying on human-labeled cross-modal data pairs during training. One popular approach involves utilizing machine translation (MT) to create pseudo-parallel data pairs, establishing correspondence between visual and non-English textual data. However, aligning their representations poses challenges due to the significant semantic gap between vision and text, as well as the lower quality of non-English representations caused by pre-trained encoders and data noise. To overcome these challenges, we propose LECCR, a novel solution that incorporates the multi-modal large language model (MLLM) to improve the alignment between visual and non-English representations. Specifically, we first employ MLLM to generate detailed visual content descriptions and aggregate them into multi-view semantic slots that encapsulate different semantics. Then, we take these semantic slots as internal features and leverage them to interact with the visual features. By doing so, we enhance the semantic information within the visual features, narrowing the semantic gap between modalities and generating local visual semantics for subsequent multi-level matching. Additionally, to further enhance the alignment between visual and non-English features, we introduce softened matching under English guidance. This approach provides more comprehensive and reliable inter-modal correspondences between visual and non-English features. Extensive experiments on four CCR benchmarks, i.e., Multi30K, MSCOCO, VATEX, and MSR-VTT-CN, demonstrate the effectiveness of our proposed method. Code: https://github.com/LiJiaBei-7/leccr.
Le Wang 0003, Qiang Zhou 0001, Zhibin Wang 0004, Hao Li 0030, Gang Hua 0001, Wei Tang 0016
ACM Multimedia7
2024 Transfer easy to hard: Adversarial contrastive feature learning for unsupervised person re-identification
Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
Pattern Recognit.4
2024 Disentangled Sample Guidance Learning for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) is challenging due to the lack of ground truth labels. Most existing methods employ iterative clustering to generate pseudo labels for unlabeled training data to guide the learning process. However, how to select samples that are both associated with high-confidence pseudo labels and hard (discriminative) enough remains a critical problem. To address this issue, a disentangled sample guidance learning (DSGL) method is proposed for unsupervised Re-ID. The method consists of disentangled sample mining (DSM) and discriminative feature learning (DFL). DSM disentangles (unlabeled) person images into identity-relevant and identity-irrelevant factors, which are used to construct disentangled positive/negative groups that contain discriminative enough information. DFL incorporates the mined disentangled sample groups into model training by a surrogate disentangled learning loss and a disentangled second-order similarity regularization, to help the model better distinguish the characteristics of different persons. By using the DSGL training strategy, the mAP on Market-1501 and MSMT17 increases by 6.6% and 10.1% when applying the ResNet50 framework, and by 0.6% and 6.9% with the vision transformer (VIT) framework, respectively, validating the effectiveness of the DSGL method. Moreover, DSGL surpasses previous state-of-the-art methods by achieving higher Top-1 accuracy and mAP on the Market-1501, MSMT17, PersonX, and VeRi-776 datasets. The source code for this paper is available at https://github.com/jihaoxuanye/DiseSGL.
Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001
IEEE Trans. Image Process.4
2024 Abnormal Ratios Guided Multi-Phase Self-Training for Weakly-Supervised Video Anomaly Detection
abstract
Weakly-supervised Video Anomaly Detection (W-VAD) aims to detect abnormal events in videos given only video-level labels for training. Recent methods relying on multiple instance learning (MIL) and self-training achieve good performance, but they tend to focus on learning easy abnormal patterns while ignoring hard ones, e.g., unusual driving trajectory or over-speeding driving. How to detect hard anomalies is a critical but largely ignored problem in W-VAD. To tackle this challenge, we propose a novel framework, termed Abnormal Ratios guided Multi-phase Self-training (ARMS), for W-VAD. It includes a new abnormal ratio-based MIL (AR-MIL) loss and a new multi-phase self-training paradigm. The AR-MIL loss guides the learning of hard anomalies by enforcing a minimum ratio of abnormal snippets in an abnormal video and no abnormal snippets in a normal video. Our multi-phase self-training paradigm sequentially performs bootstrapping, hard anomalies mining, and adaptive self-training so as to address pseudo labeling on easy anomalies, detect hard anomalies, and setting adaptive abnormal ratios for different videos in a unified framework. Experimental results on three benchmark datasets, i.e., ShanghaiTech, UCF-Crime, and XD-Violence, show that ARMS outperforms all previous state-of-the-art methods and has a great advantage in detecting hard anomalies.
Haoyue Shi 0002, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016
IEEE Trans. Multim.5
2023 Multi-Stream Representation Learning for Pedestrian Trajectory Prediction
abstract
Forecasting the future trajectory of pedestrians is an important task in computer vision with a range of applications, from security cameras to autonomous driving. It is very challenging because pedestrians not only move individually across time but also interact spatially, and the spatial and temporal information is deeply coupled with one another in a multi-agent scenario. Learning such complex spatio-temporal correlation is a fundamental issue in pedestrian trajectory prediction. Inspired by the procedure that the hippocampus processes and integrates spatio-temporal information to form memories, we propose a novel multi-stream representation learning module to learn complex spatio-temporal features of pedestrian trajectory. Specifically, we learn temporal, spatial and cross spatio-temporal correlation features in three respective pathways and then adaptively integrate these features with learnable weights by a gated network. Besides, we leverage the sparse attention gate to select informative interactions and correlations brought by complex spatio-temporal modeling and reduce complexity of our model. We evaluate our proposed method on two commonly used datasets, i.e. ETH-UCY and SDD, and the experimental results demonstrate our method achieves the state-of-the-art performance. Code: https://github.com/YuxuanIAIR/MSRL-master
Le Wang 0003, Sanping Zhou, Jinghai Duan, Gang Hua 0001, Wei Tang 0016
AAAI6
2023 MotionTrack: Learning Robust Short-Term and Long-Term Motions for Multi-Object Tracking
abstract
The main challenge of Multi-Object Tracking (MOT) lies in maintaining a continuous trajectory for each target. Existing methods often learn reliable motion patterns to match the same target between adjacent frames and discriminative appearance features to re-identify the lost targets after a long period. However, the reliability of motion prediction and the discriminability of appearances can be easily hurt by dense crowds and extreme occlusions in the tracking process. In this paper, we propose a simple yet effective multi-object tracker, i.e., MotionTrack, which learns robust short-term and long-term motions in a unified framework to associate trajectories from a short to long range. For dense crowds, we design a novel Interaction Module to learn interaction-aware motions from short-term trajectories, which can estimate the complex movement of each target. For extreme occlusions, we build a novel Refind Module to learn reliable long-term motions from the target's history trajectory, which can link the interrupted trajectory with its corresponding detection. Our Interaction Module and Refind Module are embedded in the well-known tracking-by-detection paradigm, which can work in tandem to maintain superior performance. Extensive experimental results on MOT17 and MOT20 datasets demonstrate the superiority of our approach in challenging scenarios, and it achieves state-of-the-art performances at various MOT metrics. Code is available at https://github.com/qwomeng/MotionTrack.
Sanping Zhou, Le Wang 0003, Jinghai Duan, Gang Hua 0001, Wei Tang 0016
CVPR6
2023 Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action Localization
abstract
Semi-Supervised Temporal Action Localization (SS-TAL) aims to improve the generalization ability of action detectors with large-scale unlabeled videos. Albeit the recent advancement, one of the major challenges still remains: noisy pseudo labels hinder efficient learning on abundant unlabeled videos, embodied as location biases and category errors. In this paper, we dive deep into such an important but understudied dilemma. To this end, we propose a unified framework, termed Noisy Pseudo-Label Learning, to handle both location biases and category errors. Specifically, our method is featured with (1) Noisy Label Ranking to rank pseudo labels based on the semantic confidence and boundary reliability, (2) Noisy Label Filtering to address the class-imbalance problem of pseudo labels caused by category errors, (3) Noisy Label Learning to penalize in-consistent boundary predictions to achieve noise-tolerant learning for heavy location biases. As a result, our method could effectively handle the label noise problem and improve the utilization of a large amount of unlabeled videos. Extensive experiments on THUMOS14 and ActivityNet v1.3 demonstrate the effectiveness of our method. The code is available at github.com/kunnxia/NPL.
Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016
ICCV5
2023 Representing Multimodal Behaviors With Mean Location for Pedestrian Trajectory Prediction
abstract
Representing multimodal behaviors is a critical challenge for pedestrian trajectory prediction. Previous methods commonly represent this multimodality with multiple latent variables repeatedly sampled from a latent space, encountering difficulties in interpretable trajectory prediction. Moreover, the latent space is usually built by encoding global interaction into future trajectory, which inevitably introduces superfluous interactions and thus leads to performance reduction. To tackle these issues, we propose a novel Interpretable Multimodality Predictor (IMP) for pedestrian trajectory prediction, whose core is to represent a specific mode by its mean location. We model the distribution of mean location as a Gaussian Mixture Model (GMM) conditioned on sparse spatio-temporal features, and sample multiple mean locations from the decoupled components of GMM to encourage multimodality. Our IMP brings four-fold benefits: 1) Interpretable prediction to provide semantics about the motion behavior of a specific mode; 2) Friendly visualization to present multimodal behaviors; 3) Well theoretical feasibility to estimate the distribution of mean locations supported by the central-limit theorem; 4) Effective sparse spatio-temporal features to reduce superfluous interactions and model temporal continuity of interaction. Extensive experiments validate that our IMP not only outperforms state-of-the-art methods but also can achieve a controllable prediction by customizing the corresponding mean location.
Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Adaptive Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (W-TAL) aims to classify and localize all action instances in untrimmed videos under only video-level supervision. Without frame-level annotations, it is challenging for W-TAL methods to clearly distinguish actions and background, which severely degrades the action boundary localization and action proposal scoring. In this paper, we present an adaptive two-stream consensus network (A-TSCN) to address this problem. Our A-TSCN features an iterative refinement training scheme: a frame-level pseudo ground truth is generated and iteratively updated from a late-fusion activation sequence, and used to provide frame-level supervision for improved model training. Besides, we introduce an adaptive attention normalization loss, which adaptively selects action and background snippets according to video attention distribution. By differentiating the attention values of the selected action snippets and background snippets, it forces the predicted attention to act as a binary selection and promotes the precise localization of action boundaries. Furthermore, we propose a video-level and a snippet-level uncertainty estimator, and they can mitigate the adverse effect caused by learning from noisy pseudo ground truth. Experiments conducted on the THUMOS14, ActivityNet v1.2, ActivityNet v1.3, and HACS datasets show that our A-TSCN outperforms current state-of-the-art methods, and even achieves comparable performance with several fully-supervised methods.
Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Nanning Zheng 0001, David S. Doermann, Junsong Yuan 0001, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 ContextLoc++: A Unified Context Model for Temporal Action Localization
abstract
Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by enriching the local, global and multi-scale contexts in the popular two-stage temporal localization framework. Our proposed model, dubbed ContextLoc++, can be divided into three sub-networks: L-Net, G-Net, and M-Net. L-Net enriches the local context via fine-grained modeling of snippet-level features, which is formulated as a query-and-retrieval process. Furthermore, the spatial and temporal snippet-level features, functioning as keys and values, are fused by temporal gating. G-Net enriches the global context via higher-level modeling of the video-level representation. In addition, we introduce a novel context adaptation module to adapt the global context to different proposals. M-Net further fuses the local and global contexts with multi-scale proposal features. Specially, proposal-level features from multi-scale video snippets can focus on different action characteristics. Short-term snippets with fewer frames pay attention to the action details while long-term snippets with more frames focus on the action variations. Experiments on the THUMOS14 and ActivityNet v1.3 datasets validate the efficacy of our method against existing state-of-the-art TAL algorithms.
Zixin Zhu, Le Wang 0003, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Instance Motion Tendency Learning for Video Panoptic Segmentation
abstract
Video panoptic segmentation is an important but challenging task in computer vision. It not only performs panoptic segmentation of each frame, but also associates the same instance across adjacent frames. Due to the lack of temporal coherence modeling, most existing approaches often generate identity switches during instance association, and they cannot handle ambiguous segmentation boundaries caused by motion blur. To address these difficult issues, we introduce a simple yet effective Instance Motion Tendency Network (IMTNet) for video panoptic segmentation. It learns a global motion tendency map for instance association, and a hierarchical classifier for motion boundary refinement. Specifically, a Global Motion Tendency Module (GMTM) is designed to learn robust motion features from optical flows, which can directly associate each instance in the previous frame to the corresponding instance in the current frame. In addition, we propose a Motion Boundary Refinement Module (MBRM) to learn a hierarchical classifier to handle the boundary pixels of moving targets, which can effectively revise the inaccurate segmentation predictions. Experimental results on both Cityscapes and Cityscapes-VPS datasets show that our IMTNet outperforms most state-of-the-art approaches.
Le Wang 0003, Hongzhen Liu, Sanping Zhou, Wei Tang 0016, Gang Hua 0001
IEEE Trans. Image Process.4
2023 A Reconstruction-Based Visual-Acoustic-Semantic Embedding Method for Speech-Image Retrieval
abstract
Speech-image retrieval aims at learning the relevance between image and speech.Prior approaches are mainly based on bi-modal contrastive learning, which can not alleviate the cross-modal heterogeneous issue between visual and acoustic modalities well. To address this issue, we propose a visual-acoustic-semantic embedding (VASE) method. First, we propose a tri-modal ranking loss by taking advantage of semantic information corresponding to the acoustic data, which introduces the auxiliary alignment to enhance the alignment between image and speech. Second, we introduce a cycle-consistency loss based on feature reconstruction. It can further alleviate the heterogeneous issue between different data modalities (e.g., visual-acoustic, visual-textual and acoustic-textual). Extensive experiments have demonstrated the effectiveness of our proposed method. In addition, our VASE model achieves state-of-the-art performance on the speech-image retrieval task on the Flickr8K [Harwath and Glass, 2015]s and Places [Harwathet al., 2018] datasets.
Wei Tang 0016, Yan Huang 0008, Yiwen Luo, Liang Wang 0001
IEEE Trans. Multim.2
2023 Exploring Action Centers for Temporal Action Localization
abstract
Temporal action localization aims at detecting the temporal intervals of human actions in untrimmed videos. Most previous methods rely on locating and matching the start and end times of actions. However, action boundaries are ambiguous and uncertain in nature, which leads to inaccurate action localization and a lot of false positives. In this paper, we introduce a new framework for temporal action localization. It explicitly models temporal action centers to reduce unreliable action detection results caused by ambiguous action boundaries. Since action centers are highly related to semantic actions, they can be detected more reliably than the conventional action boundaries. As a result, our framework can exclude false positives and promote high-quality proposals. Based on action centers, we propose a triplet feature fusion mechanism. It performs neural message passing among the boundaries and the center as well as contextual regions outside of the proposal to enrich its representation. In addition, we introduce a centerness scoring method to suppress proposals deviating from the centers of action instances. Consequently, our network can retrieve high-quality action proposals and locate actions more precisely. Experimental results show our method outperforms state-of-the-art methods on the THUMOS14 and ActivityNet v1.3 datasets.
Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016
IEEE Trans. Multim.6
2022 Learning Disentangled Classification and Localization Representations for Temporal Action Localization
abstract
A common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that this shared representation focuses on the most discriminative frames for classification, e.g., ``take-offs" rather than ``run-ups" in distinguishing ``high jump" and ``long jump", while frames most relevant to localization, such as the start and end frames of an action, are largely ignored. In other words, such a shared representation can not simultaneously handle both classification and localization tasks well, and it makes precise TAL difficult. To address this challenge, this paper disentangles the shared representation into classification and localization representations. The disentangled classification representation focuses on the most discriminative frames, and the disentangled localization representation focuses on the action phase as well as the action start and end. Our model could be divided into two sub-networks, i.e., the disentanglement network and the context-based aggregation network. The disentanglement network is an autoencoder to learn orthogonal hidden variables of classification and localization. The context-based aggregation network aggregates the classification and localization representations by modeling local and global contexts. We evaluate our proposed method on two popular benchmarks for TAL, which outperforms all state-of-the-art methods.
Zixin Zhu, Le Wang 0003, Wei Tang 0016, Ziyi Liu 0001, Nanning Zheng 0001, Gang Hua 0001
AAAI3
2022 Learning to Refactor Action and Co-occurrence Features for Temporal Action Localization
abstract
The main challenge of Temporal Action Localization is to retrieve subtle human actions from various co-occurring ingredients, e.g., context and background, in an untrimmed video. While prior approaches have achieved substantial progress through devising advanced action detectors, they still suffer from these co-occurring ingredients which often dominate the actual action content in videos. In this paper, we explore two orthogonal but complementary aspects of a video snippet, i.e., the action features and the co-occurrence features. Especially, we develop a novel auxiliary task by decoupling these two types of features within a video snippet and recombining them to generate a new feature representation with more salient action information for accurate action localization. We term our method RefactorNet, which first explicitly factorizes the action content and regularizes its co-occurrence features, and then synthesizes a new action-dominated video representation. Extensive experimental results and ablation studies on THUMOS14 and ActivityNet v 1.3 demonstrate that our new representation, combined with a simple action detector, can significantly improve the action localization performance.
Le Wang 0003, Sanping Zhou, Nanning Zheng 0001, Wei Tang 0016
CVPR5
2022 Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation Networks
abstract
Given only video-level action categorical labels during training, weakly-supervised temporal action localization (WS-TAL) learns to detect action instances and locates their temporal boundaries in untrimmed videos. Compared to its fully supervised counterpart, WS-TAL is more cost-effective in data labeling and thus favorable in practical applications. However, the coarse video-level supervision inevitably incurs ambiguities in action localization, especially in untrimmed videos containing multiple action instances. To overcome this challenge, we observe that significant temporal contrasts among video snippets, e.g., caused by temporal discontinuities and sudden changes, often occur around true action boundaries. This motivates us to introduce a Contrast-based Localization EvaluAtioN Network (CleanNet), whose core is a new temporal action proposal evaluator, which provides fine-grained pseudo supervision by leveraging the temporal contrasts among snippet-level classification predictions. As a result, the uncertainty in locating action instances can be resolved via evaluating their temporal contrast scores. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Besides, we also explore the usage of temporal contrast on temporal action proposal (TAP) generation task, which we believe is the first attempt with the weak supervision setting. Experiments on the THUMOS14, ActivityNet v1.2 and v1.3 datasets validate the efficacy of our method against existing state-of-the-art WS-TAL algorithms.
Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Exploiting appearance transfer and multi-scale context for efficient person image generation
Chengkang Shen, Peiyan Wang, Wei Tang 0016
Pattern Recognit.3
2022 Loss functions for pose guided person image generation
Haoyue Shi 0002, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001, Wei Tang 0016
Pattern Recognit.5
2022 Dual relation network for temporal action localization
Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016
Pattern Recognit.5
2022 Density-Aware Haze Image Synthesis by Self-Supervised Content-Style Disentanglement
abstract
The key procedure of haze image synthesis with adversarial training lies in the disentanglement of the feature involved only in haze synthesis, i.e.,the style feature, from the feature representing the invariant semantic content, i.e.,the content feature. Previous methods introduced a binary classifier to constrain the domain membership from being distinguished through the learned content feature during the training stage, thereby the style information is separated from the content feature. However, we find that these methods cannot achieve complete content-style disentanglement. The entanglement of the flawed style feature with content information inevitably leads to the inferior rendering of haze images. To address this issue, we propose a self-supervised style regression model with stochastic linear interpolation that can suppress the content information in the style feature. Ablative experiments demonstrate the disentangling completeness and its superiority in density-aware haze image synthesis. Moreover, the synthesized haze data are applied to test the generalization ability of vehicle detectors. Further study on the relation between haze density and detection performance shows that haze has an obvious impact on the generalization ability of vehicle detectors and that the degree of performance degradation is linearly correlated to the haze density, which in turn validates the effectiveness of the proposed method.
Chi Zhang 0020, Zihang Lin, Liheng Xu, Zongliang Li, Wei Tang 0016, Yuehu Liu, Gaofeng Meng, Le Wang 0003, Li Li 0013
IEEE Trans. Circuits Syst. Video Technol.5
2022 Action Coherence Network for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised Temporal Action Localization (W-TAL) aims at simultaneously classifying and locating all action instances with only video-level supervision. However, current W-TAL methods have two limitations. First, they ignore the difference in video representations between an action instance and its surrounding background when generating and scoring action proposals. Second, the unique characteristics of the RGB frames and optical flow are largely ignored when fusing these two modalities. To address these problems, an Action Coherence Network (ACN) is proposed in this paper. Its core is a new coherence loss which exploits both classification predictions and video content representations to supervise action boundary regression and thus leads to more accurate action localization results. Besides, the proposed ACN explicitly takes into account the specific characteristics of RGB frames and optical flow by training two separate sub-networks, each of which is able to generate modality-specific action proposals independently. Finally, to take advantage of the complementary action proposals generated by two streams, a novel fusion module is introduced to reconcile them and obtain the final action localization results. Experiments on the THUMOS14 and ActivityNet datasets show that our ACN outperforms the state-of-the-art W-TAL methods, and is even comparable to some recent fully-supervised methods. Particularly, ACN achieves a mean average precision of 26.4% on the THUMOS14 dataset under the IoU threshold 0.5.
Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Multim.3
2021 Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and Context
abstract
Weakly-supervised Temporal Action Localization (WS-TAL) methods learn to localize temporal starts and ends of action instances in a video under only video-level supervision. Existing WS-TAL methods rely on deep features learned for action recognition. However, due to the mismatch between classification and localization, these features cannot distinguish the frequently co-occurring contextual background, i.e., the context, and the actual action instances. We term this challenge action-context confusion, and it will adversely affect the action localization accuracy. To address this challenge, we introduce a framework that learns two feature subspaces respectively for actions and their context. By explicitly accounting for action visual elements, the action instances can be localized more precisely without the distraction from the context. To facilitate the learning of these two feature subspaces with only video-level categorical labels, we leverage the predictions from both spatial and temporal streams for snippets grouping. In addition, an unsupervised learning task is introduced to make the proposed module focus on mining temporal information. The proposed approach outperforms state-of-the-art WS-TAL methods on three benchmarks, i.e., THUMOS14, ActivityNet v1.2 and v1.3 datasets.
Ziyi Liu 0001, Le Wang 0003, Wei Tang 0016, Junsong Yuan 0001, Nanning Zheng 0001, Gang Hua 0001
AAAI3
2021 ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action Localization
abstract
The object of Weakly-supervised Temporal Action Localization (WS-TAL) is to localize all action instances in an untrimmed video with only video-level supervision. Due to the lack of frame-level annotations during training, current WS-TAL methods rely on attention mechanisms to localize the foreground snippets or frames that contribute to the video-level classification task. This strategy frequently confuse context with the actual action, in the localization result. Separating action and context is a core problem for precise WS-TAL, but it is very challenging and has been largely ignored in the literature. In this paper, we introduce an Action-Context Separation Network (ACSNet) that explicitly takes into account context for accurate action localization. It consists of two branches (i.e., the Foreground-Background branch and the Action-Context branch). The Foreground-Background branch first distinguishes foreground from background within the entire video while the Action-Context branch further separates the foreground as action and context. We associate video snippets with two latent components (i.e., a positive component and a negative component), and their different combinations can effectively characterize foreground, action and context. Furthermore, we introduce extended labels with auxiliary context categories to facilitate the learning of action-context separation. Experiments on THUMOS14 and ActivityNet v1.2/v1.3 datasets demonstrate the ACSNet outperforms existing state-of-the-art WS-TAL methods by a large margin.
Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Junsong Yuan 0001, Nanning Zheng 0001, Gang Hua 0001
AAAI4
2021 Exploit Visual Dependency Relations for Semantic Segmentation
abstract
Dependency relations among visual entities are ubiquity because both objects and scenes are highly structured. They provide prior knowledge about the real world that can help improve the generalization ability of deep learning approaches. Different from contextual reasoning which focuses on feature aggregation in the spatial domain, visual dependency reasoning explicitly models the dependency relations among visual entities. In this paper, we introduce a novel network architecture, termed the dependency network or DependencyNet, for semantic segmentation. It unifies dependency reasoning at three semantic levels. Intra-class reasoning decouples the representations of different object categories and updates them separately based on the internal object structures. Inter-class reasoning then performs spatial and semantic reasoning based on the dependency relations among different object categories. We will have an in-depth investigation on how to discover the dependency graph from the training annotations. Global dependency reasoning further refines the representations of each object category based on the global scene information. Extensive ablative studies with a controlled model size and the same network depth show that each individual dependency reasoning component benefits semantic segmentation and they together significantly improve the base network. Experimental results on two benchmark datasets show the DependencyNet achieves comparable performance to the recent states of the art.
Mingyuan Liu 0002, Dan Schonfeld, Wei Tang 0016
CVPR3
2021 Compositional Graph Convolutional Networks for 3D Human Pose Estimation
abstract
3D human pose estimation (HPE) from a single image can be decomposed into two subtasks, i.e., 2D HPE followed by 2D-to-3D pose lifting. Despite recent success in 2D HPE, 3D pose regression from 2D detections remains challenging due to the substantial depth ambiguity. Recently, graph convolutional networks (GCNs) have been exploited to model the relationships among body joints and demonstrate promising results. In this paper, we go one step further along this direction and propose a novel framework, termed Compositional GCN, for 3D HPE. It learns compositional relationships among body parts of different semantic levels and then exploits multilevel structural reasoning to reduce the depth uncertainty. Furthermore, we introduce a novel part-aware graph convolution. It not only disentangles self and neighbor transformations but also captures different relational patterns between each part and their respective neighbors. Experimental results demonstrate the effectiveness of the proposed approach.
Zhiming Zou, Dapeng Oliver Wu, Wei Tang 0016
FG4
2021 Meta Pairwise Relationship Distillation for Unsupervised Person Re-identification
abstract
Unsupervised person re-identification (Re-ID) remains challenging due to the lack of ground-truth labels. Existing methods often rely on estimated pseudo labels via iterative clustering and classification, and they are unfortunately highly susceptible to performance penalties incurred by the inaccurate estimated number of clusters. Alternatively, we propose the Meta Pairwise Relationship Distillation (MPRD) method to estimate the pseudo labels of sample pairs for unsupervised person Re-ID. Specifically, it consists of a Convolutional Neural Network (CNN) and Graph Convolutional Network (GCN), in which the GCN estimates the pseudo labels of sample pairs based on the current features extracted by CNN, and the CNN learns better features by involving high-fidelity positive and negative sample pairs imposed by GCN. To achieve this goal, a small amount of labeled samples are used to guide GCN training, which can distill meta knowledge to judge the difference in the neighborhood structure between positive and negative sample pairs. Extensive experiments on Market-1501, DukeMTMC-reID and MSMT17 datasets show that our method outperforms the state-of-the-art approaches.
Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
ICCV4
2021 Unlimited Neighborhood Interaction for Heterogeneous Trajectory Prediction
abstract
Understanding complex social interactions among agents is a key challenge for trajectory prediction. Most existing methods consider the interactions between pairwise traffic agents or in a local area, while the nature of interactions is unlimited, involving an uncertain number of agents and non-local areas simultaneously. Besides, they treat heterogeneous traffic agents the same, namely those among agents of different categories, while neglecting people’s diverse reaction patterns toward traffic agents in different categories. To address these problems, we propose a simple yet effective Unlimited Neighborhood Interaction Network (UNIN), which predicts trajectories of heterogeneous agents in multiple categories. Specifically, the proposed unlimited neighborhood interaction module generates the fused-features of all agents involved in an interaction simultaneously, which is adaptive to any number of agents and any range of interaction area. Meanwhile, a hierarchical graph attention module is proposed to obtain category-to-category interaction and agent-to-agent interaction. Finally, parameters of a Gaussian Mixture Model are estimated for generating the future trajectories. Extensive experimental results on benchmark datasets demonstrate a significant performance improvement of our method over the state-of-the-art methods.
Fang Zheng 0009, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001
ICCV4
2021 Enriching Local and Global Contexts for Temporal Action Localization
abstract
Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by enriching both the local and global contexts in the popular two-stage temporal localization framework, where action proposals are first generated followed by action classification and temporal boundary regression. Our proposed model, dubbed ContextLoc, can be divided into three sub-networks: L-Net, G-Net and P-Net. L-Net enriches the local context via fine-grained modeling of snippet-level features, which is formulated as a query-and-retrieval process. G-Net enriches the global context via higher-level modeling of the video-level representation. In addition, we introduce a novel context adaptation module to adapt the global context to different proposals. P-Net further models the context-aware inter-proposal relations. We explore two existing models to be the P-Net in our experiments. The efficacy of our proposed method is validated by experimental results on the THUMOS14 (54.3% at [email protected]) and ActivityNet v1.3 (56.01% at [email protected]) datasets, which outperforms recent states of the art. Code is available at https://github.com/buxiangzhiren/ContextLoc.
Zixin Zhu, Wei Tang 0016, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001
ICCV2
2021 Modulated Graph Convolutional Network for 3D Human Pose Estimation
abstract
The graph convolutional network (GCN) has recently achieved promising performance of 3D human pose estimation (HPE) by modeling the relationship among body parts. However, most prior GCN approaches suffer from two main drawbacks. First, they share a feature transformation for each node within a graph convolution layer. This prevents them from learning different relations between different body joints. Second, the graph is usually defined according to the human skeleton and is suboptimal because human activities often exhibit motion patterns beyond the natural connections of body joints. To address these limitations, we introduce a novel Modulated GCN for 3D HPE. It consists of two main components: weight modulation and affinity modulation. Weight modulation learns different modulation vectors for different nodes so that the feature transformations of different nodes are disentangled while retaining a small model size. Affinity modulation adjusts the graph structure in a GCN so that it can model additional edges beyond the human skeleton. We investigate several affinity modulation methods as well as the impact of regularizations. Rigorous ablation study indicates both types of modulation improve performance with negligible overhead. Compared with state-of-the-art GCNs for 3D HPE, our approach either significantly reduces the estimation errors, e.g., by around 10%, while retaining a small model size or drastically reduces the model size, e.g., from 4.22M to 0.29M (a 14.5× reduction), while achieving comparable performance. Results on two benchmarks show our Modulated GCN outperforms some recent states of the art. Our code is available at https://github.com/ZhimingZo/Modulated-GCN.
Zhiming Zou, Wei Tang 0016
ICCV2
2021 Graph-based temporal action co-localization from an untrimmed video
Le Wang 0003, Changbo Zhai, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
Neurocomputing4
2021 Giant Panda Identification
abstract
The lack of automatic tools to identify giant panda makes it hard to keep track of and manage giant pandas in wildlife conservation missions. In this paper, we introduce a new Giant Panda Identification (GPID) task, which aims to identify each individual panda based on an image. Though related to the human re-identification and animal classification problem, GPID is extraordinarily challenging due to subtle visual differences between pandas and cluttered global information. In this paper, we propose a new benchmark dataset iPanda-50 for GPID. The iPanda-50 consists of 6, 874 images from 50 giant panda individuals, and is collected from panda streaming videos. We also introduce a new Feature-Fusion Network with Patch Detector (FFN-PD) for GPID. The proposed FFN-PD exploits the patch detector to detect discriminative local patches without using any part annotations or extra location sub-networks, and builds a hierarchical representation by fusing both global and local features to enhance the inter-layer patch feature interactions. Specifically, an attentional cross-channel pooling is embedded in the proposed FFN-PD to improve the identify-specific patch detectors. Experiments performed on the iPanda-50 datasets demonstrate the proposed FFN-PD significantly outperforms competing methods. Besides, experiments on other fine-grained recognition datasets (i.e., CUB-200-2011, Stanford Cars, and FGVC-Aircraft) demonstrate that the proposed FFN-PD outperforms existing state-of-the-art methods.
Le Wang 0003, Rizhi Ding, Yuanhao Zhai 0001, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
IEEE Trans. Image Process.5
2020 Addressing Class Imbalance in Scene Graph Parsing by Learning to Contrast and Score
He Huang 0008, Shunta Saito, Yuta Kikuchi, Eiichi Matsumoto, Wei Tang 0016, Philip S. Yu
ACCV (6)5
2020 Learning Global Pose Features in Graph Convolutional Networks for 3D Human Pose Estimation
Kenkun Liu, Zhiming Zou, Wei Tang 0016
ACCV (1)3
2020 Multi-label Zero-shot Classification by Learning to Transfer from External Knowledge
He Huang 0008, Wei Tang 0016, Philip S. Yu, Yuanwei Chen, Wenhao Zheng 0001
BMVC2
2020 Loss Functions for Person Image Generation
Haoyue Shi 0002, Le Wang 0003, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001
BMVC3
2020 High-order Graph Convolutional Networks for 3D Human Pose Estimation
Zhiming Zou, Kenkun Liu, Le Wang 0003, Wei Tang 0016
BMVC4
2020 A Comprehensive Study of Weight Sharing in Graph Networks for 3D Human Pose Estimation
Kenkun Liu, Rongqi Ding, Zhiming Zou, Le Wang 0003, Wei Tang 0016
ECCV (10)5
2020 Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization
Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Junsong Yuan 0001, Gang Hua 0001
ECCV (6)3
2019 Does Learning Specific Features for Related Parts Help Human Pose Estimation?
abstract
Human pose estimation (HPE) is inherently a homogeneous multi-task learning problem, with the localization of each body part as a different task. Recent HPE approaches universally learn a shared representation for all parts, from which their locations are linearly regressed. However, our statistical analysis indicates not all parts are related to each other. As a result, such a sharing mechanism can lead to negative transfer and deteriorate the performance. This potential issue drives us to raise an interesting question. Can we identify related parts and learn specific features for them to improve pose estimation? Since unrelated tasks no longer share a high-level representation, we expect to avoid the adverse effect of negative transfer. In addition, more explicit structural knowledge, e.g., ankles and knees are highly related, is incorporated into the model, which helps resolve ambiguities in HPE. To answer this question, we first propose a data-driven approach to group related parts based on how much information they share. Then a part-based branching network (PBN) is introduced to learn representations specific to each part group. We further present a multi-stage version of this network to repeatedly refine intermediate features and pose estimates. Ablation experiments indicate learning specific features significantly improves the localization of occluded parts and thus benefits HPE. Our approach also outperforms all state-of-the-art methods on two benchmark datasets, with an outstanding advantage when occlusion occurs.
Wei Tang 0016, Ying Wu 0001
CVPR1
2018 Deeply Learned Compositional Models for Human Pose Estimation
Wei Tang 0016, Pei Yu, Ying Wu 0001
ECCV (3)1
2017 How to Train a Compact Binary Neural Network with High Accuracy?
abstract
How to train a binary neural network (BinaryNet) with both high compression rate and high accuracy on large scale dataset? We answer this question through a careful analysis of previous work on BinaryNets, in terms of training strategies, regularization, and activation approximation. Our findings first reveal that a low learning rate is highly preferred to avoid frequent sign changes of the weights, which often makes the learning of BinaryNets unstable. Secondly, we propose to use PReLU instead of ReLU in a BinaryNet to conveniently absorb the scale factor for weights to the activation function, which enjoys high computation efficiency for binarized layers while maintains high approximation accuracy. Thirdly, we reveal that instead of imposing L2 regularization, driving all weights to zero which contradicts with the setting of BinaryNets, we introduce a regularization term that encourages the weights to be bipolar. Fourthly, we discover that the failure of binarizing the last layer, which is essential for high compression rate, is due to the improper output range. We propose to use a scale layer to bring it to normal. Last but not least, we propose multiple binarizations to improve the approximation of the activations. The composition of all these enables us to train BinaryNets with both high compression rate and high accuracy, which is strongly supported by our extensive empirical study.
Wei Tang 0016, Gang Hua 0001, Liang Wang 0001
AAAI1
2017 Towards a Unified Compositional Model for Visual Pattern Modeling
abstract
Compositional models represent visual patterns as hierarchies of meaningful and reusable parts. They are attractive to vision modeling due to their ability to decompose complex patterns into simpler ones and resolve the lowlevel ambiguities in high-level image interpretations. However, current compositional models separate structure and part discovery from parameter estimation, which generally leads to suboptimal learning and fitting of the model. Moreover, the commonly adopted latent structural learning is not scalable for deep architectures. To address these difficult issues for compositional models, this paper quests for a unified framework for compositional pattern modeling, inference and learning. Represented by And-or graphs (AOGs), it jointly models the compositional structure, parts, features, and composition/sub-configuration relationships. We show that the inference algorithm of the proposed framework is equivalent to a feed-forward network. Thus, all the parameters can be learned efficiently via the highly-scalable back-propagation (BP) in an end-to-end fashion. We validate the model via the task of handwritten digit recognition. By visualizing the processes of bottom-up composition and top-down parsing, we show that our model is fully interpretable, being able to learn the hierarchical compositions from visual primitives to visual patterns at increasingly higher levels. We apply this new compositional model to natural scene character recognition and generic object detection. Experimental results have demonstrated its effectiveness.
Wei Tang 0016, Pei Yu, Jiahuan Zhou, Ying Wu 0001
ICCV1
2017 Efficient Online Local Metric Adaptation via Negative Samples for Person Re-identification
abstract
Many existing person re-identification (PRID) methods typically attempt to train a faithful global metric offline to cover the enormous visual appearance variations, so as to directly use it online on various probes for identity matching. However, their need for a huge set of positive training pairs is very demanding in practice. In contrast to these methods, this paper advocates a different paradigm: part of the learning can be performed online but with nominal costs, so as to achieve online metric adaptation for different input probes. A major challenge here is that no positive training pairs are available for the probe anymore. By only exploiting easily-available negative samples, we propose a novel solution to achieve local metric adaptation effectively and efficiently. For each probe at the test time, it learns a strictly positive semi-definite dedicated local metric. Comparing to offline global metric learning, its computational cost is negligible. The insight of this new method is that the local hard negative samples can actually provide tight constraints to fine tune the metric locally. This new local metric adaptation method is generally applicable, as it can be used on top of any global metric to enhance its performance. In addition, this paper gives in-depth theoretical analysis and justification of the new method. We prove that our new method guarantees the reduction of the classification error asymptotically, and prove that it actually learns the optimal local metric to best approximate the asymptotic case by a finite number of training data. Extensive experiments and comparative studies on almost all major benchmarks (VIPeR, QMUL GRID, CUHK Campus, CUHK03 and Market-1501) have confirmed the effectiveness and superiority of our method.
Jiahuan Zhou, Pei Yu, Wei Tang 0016, Ying Wu 0001
ICCV3
2016 Collaborative sparse unmixing of hyperspectral data using L2, P norm
abstract
Sparse unmixing is a popular method in remotely sensed hyperspectral imagery interpretation. Recently, the collaborative sparse unmixing model has shown its advantage over the traditional single channel sparse unmixing method since it can utilize the subspace nature of the high-dimensional hyperspectral data to alleviate the difficulty caused by the usually high mutual coherence of the spectral library. However, the existing collaborative sparse unmixing model is constructed in the convex ℓ1norm framework while it is known that using the ℓp(02;p(02;p(0 <; p <; 1) norm collaborative sparse unmixing model.
Dan Wang 0005, Zhenwei Shi 0001, Wei Tang 0016
IGARSS3
2015 Sparse Unmixing of Hyperspectral Data Using Spectral A Priori Information
abstract
Given a spectral library, sparse unmixing aims at finding the optimal subset of endmembers from it to model each pixel in the hyperspectral scene. However, sparse unmixing still remains a challenging task due to the usually high mutual coherence of the spectral library. In this paper, we exploit the spectral a priori information in the hyperspectral image to alleviate this difficulty. It assumes that some materials in the spectral library are known to exist in the scene. Such information can be obtained via field investigation or hyperspectral data analysis. Then, we propose a novel model to incorporate the spectral a priori information into sparse unmixing. Based on the alternating direction method of multipliers, we present a new algorithm, which is termed sparse unmixing using spectral a priori information (SUnSPI), to solve the model. Experimental results on both synthetic and real data demonstrate that the spectral a priori information is beneficial to sparse unmixing and that SUnSPI can exploit this information effectively to improve the abundance estimation.
Wei Tang 0016, Zhenwei Shi 0001, Ying Wu 0001, Changshui Zhang
IEEE Trans. Geosci. Remote. Sens.1
2015 Robust Hyperspectral Image Target Detection Using an Inequality Constraint
abstract
In real hyperspectral images, there exist variations within spectra of materials. The inherent spectral variability is one of the major obstacles for the successful hyperspectral image target detection. Although several hyperspectral image target detection algorithms have been proposed, there are few algorithms considering the spectral variability. Under such circumstances, in this paper, we propose a hyperspectral image target detection algorithm that is robust to the target spectral variability. The proposed algorithm utilizes an inequality constraint to guarantee that the outputs of target spectra, which vary in a certain set, are larger than one, so that these target spectra could be detected. The proposed algorithm transforms the target detection to a convex optimization problem and uses a kind of interior point method named barrier method to solve the formulated optimization problem effectively. Two synthetic hyperspectral images and two real hyperspectral images are used to conduct experiments. The experimental results demonstrate the proposed algorithm is robust to the target spectral variability and performs better than other classical algorithms.
Zhenwei Shi 0001, Wei Tang 0016
IEEE Trans. Geosci. Remote. Sens.3
2014 Single Remote Sensing Image Dehazing
abstract
Remote sensing images are widely used in various fields. However, they usually suffer from the poor contrast caused by haze. In this letter, we propose a simple, but effective, way to eliminate the haze effect on remote sensing images. Our work is based on the dark channel prior and a common haze imaging model. In order to eliminate halo artifacts, we use a low-pass Gaussian filter to refine the coarse estimated atmospheric veil. We then redefine the transmission, with the aim of preventing the color distortion of the recovered images. The main advantage of the proposed algorithm is its fast speed, while it can also achieve good results. The experimental results demonstrate that our algorithm produces visually appealing dehazing images and retains the very fine details. Moreover, for images containing partly clear and partly hazy areas, our algorithm can also achieve good results.
Jiao Long, Zhenwei Shi 0001, Wei Tang 0016, Changshui Zhang
IEEE Geosci. Remote. Sens. Lett.3
2014 Subspace Matching Pursuit for Sparse Unmixing of Hyperspectral Data
abstract
Sparse unmixing assumes that each mixed pixel in the hyperspectral image can be expressed as a linear combination of only a few spectra (endmembers) in a spectral library, known a priori. It then aims at estimating the fractional abundances of these endmembers in the scene. Unfortunately, because of the usually high correlation of the spectral library, the sparse unmixing problem still remains a great challenge. Moreover, most related work focuses on the l1convex relaxation methods, and little attention has been paid to the use of simultaneous sparse representation via greedy algorithms (GAs) (SGA) for sparse unmixing. SGA has advantages such as that it can get an approximate solution for the l0problem directly without smoothing the penalty term in a low computational complexity as well as exploit the spatial information of the hyperspectral data. Thus, it is necessary to explore the potential of using such algorithms for sparse unmixing. Inspired by the existing SGA methods, this paper presents a novel GA termed subspace matching pursuit (SMP) for sparse unmixing of hyperspectral data. SMP makes use of the low-degree mixed pixels in the hyperspectral image to iteratively find a subspace to reconstruct the hyperspectral data. It is proved that, under certain conditions, SMP can recover the optimal endmembers from the spectral library. Moreover, SMP can serve as a dictionary pruning algorithm. Thus, it can boost other sparse unmixing algorithms, making them more accurate and time efficient. Experimental results on both synthetic and real data demonstrate the efficacy of the proposed algorithm.
Zhenwei Shi 0001, Wei Tang 0016, Zhana Duren, Zhiguo Jiang 0001
IEEE Trans. Geosci. Remote. Sens.2
2014 Regularized Simultaneous Forward-Backward Greedy Algorithm for Sparse Unmixing of Hyperspectral Data
abstract
Sparse unmixing assumes that each observed signature of a hyperspectral image is a linear combination of only a few spectra (endmembers) in an available spectral library. It then estimates the fractional abundances of these endmembers in the scene. The sparse unmixing problem still remains a great difficulty due to the usually high correlation of the spectral library. Under such circumstances, this paper presents a novel algorithm termed as the regularized simultaneous forward-backward greedy algorithm (RSFoBa) for sparse unmixing of hyperspectral data. The RSFoBa has low computational complexity of getting an approximate solution for the l0problem directly and can exploit the joint sparsity among all the pixels in the hyperspectral data. In addition, the combination of the forward greedy step and the backward greedy step makes the RSFoBa more stable and less likely to be trapped into the local optimum than the conventional greedy algorithms. Furthermore, when updating the solution in each iteration, a regularizer that enforces the spatial-contextual coherence within the hyperspectral image is considered to make the algorithm more effective. We also show that the sublibrary obtained by the RSFoBa can serve as input for any other sparse unmixing algorithms to make them more accurate and time efficient. Experimental results on both synthetic and real data demonstrate the effectiveness of the proposed algorithm.
Wei Tang 0016, Zhenwei Shi 0001, Ying Wu 0001
IEEE Trans. Geosci. Remote. Sens.1