EDBT 2026 Demo / reviewers in the wild / expert
Le Wang 0003
dblp:79/652-3
· DBLP profile ↗
123ranked-venue papers
13as first author
88since 2021 · last 2026
0000-0001-6636-6396ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 87 · 6 first-author · 62 since 2021Graphics, computer vision, multimedia, augmented reality and games · 85 · 9 first-author · 57 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMsabstractWhile Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the understanding of complex human intentions and the provision of empathetic, context-aware responses. Here we introduce HumanSense, a comprehensive benchmark designed to evaluate the human-centered perception and interaction capabilities of MLLMs, with a particular focus on deep understanding of extended multimodal contexts and the formulation of rational feedback. Our evaluation reveals that leading MLLMs still have considerable room for improvement, particularly for advanced interaction-oriented tasks. Supplementing visual input with audio and text information yields substantial improvements, and Omni-modal models show advantages on these tasks.Furthermore, grounded in the observation that appropriate feedback stems from a contextual analysis of the interlocutor's needs and emotions, we posit that reasoning ability serves as the key to unlocking it. We devise a multi-stage, modality-progressive reinforcement learning approach, resulting in HumanSense-Omni-Reasoning, which substantially enhances performance on higher-level understanding and interactive tasks. Additionally, we observe that successful reasoning processes appear to exhibit consistent thought patterns. By designing corresponding prompts, we also enhance the performance of non-reasoning models in a training-free manner. Ruobing Zheng, Jingdong Chen, Le Wang 0003 |
AAAI | 7 |
| 2026 | Sparse Trajectory PredictionabstractPedestrian trajectory prediction is crucial for ensuring safe decision-making in intelligent robotic systems. While this task demands real-time performance, previous works have primarily focused on improving prediction accuracy, often neglecting efficiency. Dense predictions with time-consuming post-clustering steps and global interactions with quadratic computational complexity result in a trade-off between accuracy and speed. In this paper, we propose a novel Sparse Trajectory Prediction (STP) model that aims to achieve both high accuracy and real-time speed by following an efficient principle: leveraging sparse structures to achieve global effects. STP instantiates this principle within a transformer-style encoder-decoder framework. In the encoder, STP introduces irregular interaction, which builds sparse interactions with dynamic interactive positions, reducing computational complexity to linearithmic/linear while maintaining global interaction. In the decoder, STP applies an early-sparsity strategy to generate sparse motion modes that represent global motion behaviors. These modes are shared across all predictions, eliminating redundant computations. By harnessing the expressive power of transformers, STP maps these sparse motion modes into multimodal future trajectories, significantly improving prediction speed while ensuring accuracy. Experimental results on four commonly used datasets demonstrate that STP maximizes both accuracy and prediction speed, achieving state-of-the-art performance and significantly improving prediction speed by about $100 \times$100× - $150 \times$150× to satisfy the real-time demand. Liushuai Shi, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Data-Driven Bidirectional Spatial-Adaptive Network for Weakly Supervised Object Detection in Remote Sensing ImagesabstractWeakly-supervised object detection (WSOD) learns detectors with only image-level classification annotations. Without precise instance-level labels, most previous WSOD methods in remote sensing images (RSIs) select the highest-scoring proposals as the final detection results, which are confronted by two major challenges: (1) instances with small scale or rare poses are easily neglected; (2) optimizing network by the top-scoring region inevitably overlooks many valuable candidate proposals. To mitigate the above-mentioned challenges, we propose a data-driven bidirectional spatial-adaptive network (BSANet). It contains a forward-reverse spatial dropout (FRSD) module to reduce instance ambiguity induced from extreme scales and poses, as well as crowded scene, and to better excavate the entire instances. From attention learning perspective, the proposed FRSD is conceptually similar to a data-driven hard attention mechanism, which adaptively samples and reconstructs the spatially related regions for mining more latent feature responses. Meanwhile, our FRSD effectively alleviates the inherent problem that non-parametric hard attention learning fashion cannot adapt to different datasets. In addition, we build a soft attention branch to simultaneously model soft pixel-level and hard region-level attention information for exploring the complementary benefit between soft and hard attention learning. We evaluate our BSANet on the challenging NWPU VHR-10.v2 and DIOR datasets. Experimental results demonstrate that our method sets a new state-of-the-art. Zebin Wu 0001, Shangdong Zheng, Yang Xu 0006, Le Wang 0003, Zhihui Wei, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | PR-DETR: Injecting position and relation prior for dense video captioning
Sanping Zhou, Le Wang 0003 |
Pattern Recognit. | 4 |
| 2026 | Action hints: Semantic typicality and context uniqueness for generalizable skeleton-based video anomaly detection
Canhui Tang, Sanping Zhou, Haoyue Shi 0002, Le Wang 0003 |
Pattern Recognit. | 4 |
| 2026 | RAMP: Iterative Refinement and Adaptive Multi-granularity Perception for embodied dialog localization
Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
Pattern Recognit. | 2 |
| 2026 | Pedestrian Trajectory Prediction via Hierarchical Dynamics DecompositionabstractPredicting human future trajectories is crucial for various intelligent systems and applications. Previous approaches typically adopt a direct prediction strategy, which decodes trajectory features directly into future coordinates. However, they overlook different hierarchical high-order velocities, which have stronger representational abilities in dynamics. In this paper, we introduce HDDNet, a novel trajectory prediction framework that follows dynamical principles and employs a hierarchical dynamics decomposition strategy. Specifically, HDDNet models future trajectories by progressively transferring trajectory coordinates into velocity, acceleration, and jerk, up to the highest-order velocity, which sequentially represent a broader receptive field and a more compact representation of motion dynamics. Furthermore, we design a hierarchical dynamics decomposition decoder with a corresponding dynamics loss, which predicts future trajectories by sequentially refining human motions from the highest-order velocity down to the final coordinates. Compared to the traditional direct prediction strategy, our approach makes better use of dynamic information at different levels. Extensive experiments and ablation studies on the ETH-UCY, SDD and GigaTraj datasets demonstrate that our method outperforms existing state-of-the-art approaches. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | RSRNav: Reasoning Spatial Relationship for Image-Goal NavigationabstractRecent image-goal navigation (ImageNav) methods learn a perception-action policy by separately capturing semantic features of the goal and egocentric images, then passing them to a policy network. However, challenges remain: (1) Semantic features often fail to provide accurate directional information, leading to superfluous actions, and (2) performance drops significantly when viewpoint inconsistencies arise between training and application. To address these challenges, we propose RSRNav, a simple yet effective method that reasons spatial relationships between the goal and current observations as navigation guidance. Specifically, we model the spatial relationship by constructing correlations between the goal and current observations, which are then passed to the policy network for action prediction. These correlations are progressively refined using fine-grained cross-correlation and direction-aware correlation for more precise navigation. Extensive evaluation of RSRNav on three benchmark datasets demonstrates superior navigation performance, particularly in the "user-matched goal" setting, highlighting its potential for real-world applications. Code: https://github.com/ qinzheng2000/RSRNav.git. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Representation Sampling and Hybrid Transformer Network for Image Compressed Sensing
Heping Song, Jingyao Gong, Hongjie Jia, Xiangjun Shen, Jianping Gou, Hongying Meng, Le Wang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Advancing Pre-Trained Teacher: Towards Robust Feature Discrepancy for Anomaly DetectionabstractWith the wide application of knowledge distillation between an ImageNet pre-trained teacher model and a learnable student model, unsupervised anomaly detection has witnessed a significant achievement in the past few years. The success of this framework mainly relies on how to keep the feature discrepancy between the teacher and student model, in which it has two underlying sub-assumptions: (1) The teacher model can represent two separable distributions for the normal and abnormal patterns, while (2) the student model can only reconstruct the normal distribution. However, it still remains a challenging issue to maintain these ideal assumptions in practice. In this paper, we propose a simple yet effective two-stage industrial anomaly detection framework, termed AAND, which sequentially performs Anomaly Amplification and Normality Distillation to enhance the two assumptions. In the first anomaly amplification stage, we propose a novel Residual Anomaly Amplification (RAA) module to advance the pre-trained teacher encoder with synthetic anomalies. It generates adaptive residuals to amplify anomalies while maintaining the feature integrity of pre-trained model. It mainly comprises a Matching-guided Residual Gate and an Attribute-scaling Residual Generator, which can determine the residuals' proportion and characteristic, respectively. In the second normality distillation stage, we further employ a reverse distillation paradigm to train a student decoder, in which a novel Hard Knowledge Distillation (HKD) loss is built to better facilitate the reconstruction of normal patterns. Comprehensive experiments on the MvTecAD, VisA, and MvTec3D-RGB datasets show that our method achieves state-of-the-art performance. Our code is available at https://github.com/Hui-design/AAND. Canhui Tang, Sanping Zhou, Yonghao Dong, Le Wang 0003 |
IEEE Trans. Image Process. | 5 |
| 2026 | MoDe-Track: Robust Multi-Object Tracking With Motion Decoupling in UAV VideosabstractMulti-Object Tracking (MOT) in Unmanned Aerial Vehicle (UAV) scenarios is characterized by frequent and abrupt camera motion, which presents two unique challenges: nonlinear motion and appearance degradation. Traditional motion models, designed for smooth and consistent motion, struggle to capture the complex background motion patterns caused by UAV movement; while appearance-based methods are easily disrupted by occlusion and blur, leading to unreliable associations. Even though dense optical flow is widely utilized to model complex motion patterns, the entanglement of background and object motion often introduces interference, limiting its effectiveness. To this end, we propose MoDe-Track, a unified framework that explicitly decouples scene motion into background and object components, and serves as an elegant integration of three robust components. Specifically, the Scene Motion Decomposition (SMD) module decouples the motion into the background and object components based on robust principal component analysis, serving as the foundation for motion compensation and feature propagation. Afterwards, the Background Motion Compensation (BMC) uses the decomposed background flow to estimate and compensate for camera motion, mitigating the effects of nonlinear motion. Finally, the Foreground-guided Feature Propagation (FFP) module uses the decoupled object flow to guide feature propagation across frames, achieving temporal consistency and enhancing robustness against occlusion and motion blur. Extensive experimental results on two benchmarks, VisDrone2019 and UAVDT, demonstrate that MoDe-Track consistently outperforms current multi-object tracking methods. We achieve 56.0% MOTA on VisDrone2019 and 56.2% MOTA on UAVDT, reaching the state-of-the-art among existing methods. Zixuan Song, Sanping Zhou, Wei Tang 0016, Le Wang 0003 |
IEEE Trans. Multim. | 5 |
| 2025 | Diversifying Query: Region-Guided Transformer for Temporal Sentence GroundingabstractTemporal sentence grounding is a challenging task that aims to localize the moment spans relevant to a language description. Although recent DETR-based models have achieved notable progress by leveraging multiple learnable moment queries, they suffer from overlapped and redundant proposals, leading to inaccurate predictions. We attribute this limitation to the lack of task-related guidance for the learnable queries to serve a specific mode. Furthermore, the complex solution space generated by variable and open-vocabulary language descriptions complicates optimization, making it harder for learnable queries to adaptively distinguish each other, leading to more severe overlapped proposals. To address this limitation, we present the Region-Guided TRansformer (RGTR) for temporal sentence grounding, which introduces regional guidance to increase query diversity and eliminate overlapped proposals. Instead of using learnable queries, RGTR adopts a set of anchor pairs as moment queries to introduce explicit regional guidance. Each moment query takes charge of moment prediction for a specific temporal region, which reduces the optimization difficulty and ensures the diversity of the proposals. In addition, we design an IoU-aware scoring head to improve proposal quality. Extensive experiments demonstrate the effectiveness of RGTR, outperforming state-of-the-art methods on three public benchmarks and exhibiting good generalization and robustness on out-of-distribution splits. Xiaolong Sun, Liushuai Shi, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
AAAI | 3 |
| 2025 | RefDetector: A Simple Yet Effective Matching-based Method for Referring Expression ComprehensionabstractDespite the rapid and substantial advancements in object detection, it continues to face limitations imposed by pre-defined category sets. Current methods for visual grounding primarily focus on how to better leverage the visual backbone to generate text-tailored visual features, which may require adjusting the parameters of the entire model. Besides, some early methods, \ie, matching-based method, build upon and extend the functionality of existing object detectors by enabling them to localize an object based on free-form linguistic expressions, which have good application potential. However, the untapped potential of the matching-based approach has not been fully realized due to inadequate exploration. In this paper, we first analyze the limitations that exist in the current matching-based method (\ie, mismatch problem and complicated fusion mechanisms), and then present a simple yet effective matching-based method, namely RefDetector. To tackle the above issues, we devise a simple heuristic rule to generate proposals with improved referent recall. Additionally, we introduce a straightforward vision-language interaction module that eliminates the need for intricate manually-designed mechanisms. Moreover, we have explored the visual grounding based on the modern detector DETR, and achieved significant performance improvement. Extensive experiments on three REC benchmark datasets, \ie, RefCOCO, RefCOCO+, and RefCOCOg validate the effectiveness of the proposed method. Zhuotao Tian, Sanping Zhou, Le Wang 0003 |
AAAI | 5 |
| 2025 | Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal TransportabstractPoint-supervised Temporal Action Localization poses significant challenges due to the difficulty of identifying complete actions with a single-point annotation per action. Existing methods typically employ Multiple Instance Learning, which struggles to capture global temporal context and requires heuristic post-processing. In research on fully-supervised tasks, DETR-based structures have effectively addressed these limitations. However, it is nontrivial to merely adapt DETR to this task, encountering two major bottlenecks. (1) How to integrate point label information into the model and (2) How to select optimal decoder proposals for training in the absence of complete action segment annotations. To address this issue, we introduce an end-to-end framework by integrating Query Reformation and Optimal Transport (QROT). Specifically, we encode point labels through a set of semantic consensus queries, enabling effective focus on action-relevant snippets. Furthermore, we integrate an optimal transport mechanism to generate high-quality pseudo labels. These pseudo-labels facilitate precise proposals selection based on the Hungarian algorithm, significantly enhancing localization accuracy in point-supervised settings. Extensive experiments on the THUMOS14 and ActivityNet-v1.3 datasets demonstrate that our method outperforms existing MIL-based approaches, offering more stable and accurate temporal action localization in point-level supervision. Mengnan Liu 0001, Le Wang 0003, Sanping Zhou, Xiaolong Sun, Gang Hua 0001 |
CVPR | 2 |
| 2025 | PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic ManipulationabstractRobotic manipulation based on visual observations and natural language instructions is a long-standing challenge in robotics. Yet prevailing approaches model action distribution by adopting explicit or implicit representations, which often struggle to achieve a trade-off between accuracy and efficiency. In response, we propose PDFactor, a novel framework that models action distribution with a hybrid triplane representation. In particular, PDFactor decomposes 3D point cloud into three orthogonal feature planes and leverages a tri-perspective view transformer to produce dense cubic features as a latent diffusion field aligned with observation space representing 6-DoF action probability distribution at an arbitrary location. We employ a small denoising network conceptually as both a parameterized loss function measuring the quality of the learned latent features and an action gradient decoder to sample actions from the latent diffusion field during inference. This design enables our PDFactor to benefit from spatial awareness of explicit representation and arbitrary resolution of implicit representation, rendering it with manipulation accuracy, inference efficiency, and model scalability. Experiments demonstrate that PDFactor outperforms state-of-the-art approaches across a diverse range of manipulation tasks in RLBench simulation. Moreover, PDFactor can effectively learn multi-task policies from a limited number of human demonstrations, achieving promising accuracy in a variety of real-world manipulation tasks. Jingyi Tian, Le Wang 0003, Sanping Zhou, Haowen Sun 0003, Wei Tang 0016 |
CVPR | 2 |
| 2025 | FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic ManipulationabstractRobotic manipulation in high-precision tasks is essential for numerous industrial and real-world applications where accuracy and speed are required. Yet current diffusion-based policy learning methods generally suffer from low computational efficiency due to the iterative denoising process during inference. Moreover, these methods do not fully explore the potential of generative models for enhancing information exploration in 3D environments. In response, we propose FlowRAM, a novel framework that leverages generative models to achieve region-aware perception, enabling efficient multimodal information processing. Specifically, we devise a Dynamic Radius Schedule, which allows adaptive perception, facilitating transitions from global scene comprehension to fine-grained geometric details. Furthermore, we integrate state space models to integrate multimodal information, while preserving linear computational complexity. In addition, we employ conditional flow matching to learn action poses by regressing deterministic vector fields, simplifying the learning process while maintaining performance. We verify the effectiveness of the FlowRAM in the RLBench, an established manipulation benchmark, and achieve state-of-the-art performance. The results demonstrate that FlowRAM achieves a remarkable improvement, particularly in high-precision tasks, where it outperforms previous methods by 12.0% in average success rate. Additionally, FlowRAM is able to generate physically plausible actions for a variety of real-world tasks in less than 4 time steps, significantly increasing inference speed. Le Wang 0003, Sanping Zhou, Jingyi Tian, Haowen Sun 0003, Wei Tang 0016 |
CVPR | 2 |
| 2025 | Towards Precise Embodied Dialogue Localization via Causality Guided DiffusionabstractEmbodied localization based on vision and natural language dialogues presents a persistent challenge in embodied intelligence. Existing methods often approach this task as an image translation problem, leveraging encoder-decoder architectures to predict heatmaps. However, these methods frequently experience a deficiency in accuracy, largely due to their heavy reliance on resolution. To address this issue, we introduce CGD, a novel framework that utilizes causality guided diffusion model to directly model coordinate distributions. Specifically, CGD employs a denoising network to regress coordinates, while integrating causal learning modules, namely back-door adjustment (BDA) and front-door adjustment (FDA) to mitigate confounders during the diffusion process. This approach reduces the dependency on high resolution for improving accuracy, while effectively minimizing spurious correlations, thereby promoting unbiased learning. By guiding the denoising process with causal adjustments, CGD offers flexible control over intensity, ensuring seamless integration with diffusion models. Experimental results demonstrate that CGD outperforms state-of-the-art methods across all metrics. Additionally, we also evaluate CGD in a multi-shot setting, achieving consistently high accuracy. Le Wang 0003, Sanping Zhou, Jingyi Tian, Gang Hua 0001, Wei Tang 0016 |
CVPR | 2 |
| 2025 | Semantic Graph Embedded Energy Minimization Learning for Scene Graph GenerationabstractThe performance of current scene graph generation models is affected by training with cross-entropy loss, exacerbating the problem of prediction bias stemming from biased training data. Energy-based model adopts a learning method for joint image and scene graph to alleviate this challenge. However, this method only focuses on the visual features of images, neglecting the rich relation information contained in the semantic space. To address this issue, we innovatively employ powerful pre-trained large models to realize a simple yet effective semantic graph embedded energy minimization framework for the SGG task. Specifically, we use large models to generate image descriptions and extract relation triplets, which are then transformed into semantic graphs with entities as nodes and relations as edges. Moreover, by mapping these graphs into the same space using GNN to learn the minimal energy value, our approach enables SGG model to learn structural information in both semantic and visual spaces. We validate the effectiveness and efficiency of our method on the SGG benchmark Visual Genome dataset. Compared with prevailing models and EBM, we achieve a significant performance improvement of up to 2.99% and 2.24%, respectively. Jinghang Chen, Chi Zhang 0020, Yuehu Liu, Le Wang 0003 |
ICASSP | 4 |
| 2025 | Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense PredictionabstractSufficient cross-task interaction is crucial for success in multi-task dense prediction. However, sufficient interaction often results in high computational complexity, forcing existing methods to face the trade-off between interaction completeness and computational efficiency. To address this limitation, this work proposes a Bidirectional Interaction Mamba (BIM), which incorporates novel scanning mechanisms to adapt the Mamba modeling approach for multi-task dense prediction. On the one hand, we introduce a novel Bidirectional Interaction Scan (BI-Scan) mechanism, which constructs task-specific representations as bidirectional sequences during interaction. By integrating task-first and position-first scanning modes within a unified linear complexity architecture, BI-Scan efficiently preserves critical cross-task information. On the other hand, we employ a Multi-Scale Scan~(MS-Scan) mechanism to achieve multi-granularity scene modeling. This design not only meets the diverse granularity requirements of various tasks but also enhances nuanced cross-task feature interactions. Extensive experiments on two challenging benchmarks, \emph{i.e.}, NYUD-V2 and PASCAL-Context, show the superiority of our BIM vs its state-of-the-art competitors. Mang Cao, Sanping Zhou, Ye Deng 0005, Wenli Huang 0004, Le Wang 0003 |
ICCV | 6 |
| 2025 | Moment Quantization for Video Temporal GroundingabstractVideo temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishing relevant and irrelevant moments. Previous methods focused on learning continuous features exhibit weak differentiation between foreground and background features. In this paper, we propose a novel Moment-Quantization based Video Temporal Grounding method (MQVTG), which quantizes the input video into various discrete vectors to enhance the discrimination between relevant and irrelevant moments. Specifically, MQVTG maintains a learnable moment codebook, where each video moment matches a codeword. Considering the visual diversity, i.e., various visual expressions for the same moment, MQVTG treats moment-codeword matching as a clustering process without using discrete vectors, avoiding the loss of useful information from direct hard quantization. Additionally, we employ effective prior-initialization and joint-projection strategies to enhance the maintained moment codebook. With its simple implementation, the proposed method can be integrated into existing temporal grounding models as a plug-and-play component. Extensive experiments on six popular benchmarks demonstrate the effectiveness and generalizability of MQVTG, significantly outperforming state-of-the-art methods. Further qualitative analysis shows that our method effectively groups relevant features and separates irrelevant ones, aligning with our goal of enhancing discrimination. Xiaolong Sun, Le Wang 0003, Sanping Zhou, Liushuai Shi, Mengnan Liu 0001, Gang Hua 0001 |
ICCV | 2 |
| 2025 | Versatile Multimodal Controls for Expressive Talking Human Animation
Ruobing Zheng, Zixin Zhu, Sanping Zhou, Ming Yang 0007, Le Wang 0003 |
ACM Multimedia | 8 |
| 2025 | DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic ManipulationabstractLearning generalizable robotic manipulation policies remains a key challenge due to the scarcity of diverse real-world training data. While recent approaches have attempted to mitigate this through self-supervised representation learning, most either rely on 2D vision pretraining paradigms such as masked image modeling, which primarily focus on static semantics or scene geometry, or utilize large-scale video prediction models that emphasize 2D dynamics, thus failing to jointly learn the geometry, semantics, and dynamics required for effective manipulation. In this paper, we present DynaRend, a representation learning framework that learns 3D-aware and dynamics-informed triplane features via masked reconstruction and future prediction using differentiable volumetric rendering. By pretraining on multi-view RGB-D video data, DynaRend jointly captures spatial geometry, future dynamics, and task semantics in a unified triplane representation. The learned representations can be effectively transferred to downstream robotic manipulation tasks via action value map prediction. We evaluate DynaRend on two challenging benchmarks, RLBench and Colosseum, as well as in real-world robotic experiments, demonstrating substantial improvements in policy success rate, generalization to environmental perturbations, and real-world applicability across diverse manipulation tasks. Jingyi Tian, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
NeurIPS | 2 |
| 2025 | SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World ModelsabstractWorld models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions due to disrupted spatial structure, inefficient decoding, and inadequate motion modeling. In response, we propose Scale-wise Autoregression with Motion PrOmpt (SAMPO), a hybrid framework that combines visual autoregressive modeling for intra-frame generation with causal modeling for next-frame generation. Specifically, SAMPO integrates temporal causal decoding with bidirectional spatial attention, which preserves spatial locality and supports parallel decoding within each scale. This design significantly enhances both temporal consistency and rollout efficiency. To further improve dynamic scene understanding, we devise an asymmetric multi-scale tokenizer that preserves spatial details in observed frames and extracts compact dynamic representations for future frames, optimizing both memory usage and model performance. Additionally, we introduce a trajectory-aware motion prompt module that injects spatiotemporal cues about object and robot trajectories, focusing attention on dynamic regions and improving temporal consistency and physical realism. Extensive experiments show that SAMPO achieves competitive performance in action-conditioned video prediction and model-based control, improving generation quality with 4.4× faster inference. We also evaluate SAMPO's zero-shot generalization and scaling behavior, demonstrating its ability to generalize to unseen tasks and benefit from larger model sizes. Jingyi Tian, Le Wang 0003, Zhimin Liao, Huaiyi Dong, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
NeurIPS | 3 |
| 2025 | UniCuboid: Cuboid-based dense shape supervision for monocular 3D object detection
Yuanqi Su, Haoyue Shi 0002, Haoang Lu, Yuehu Liu, Le Wang 0003 |
Neurocomputing | 6 |
| 2025 | AFC-RNN: Adaptive Forgetting-Controlled Recurrent Neural Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction plays a crucial and fundamental role in many computer vision tasks. Most existing works utilize recurrent neural networks to extract temporal features from trajectories because their recursive structure is inherently well-suited for time series data. However, previous methods overlook the forgetting characteristics of pedestrians when modeling historical trajectories, which may cause the model to focus on the wrong positions of historical information. In this paper, we propose a simple yet effective Adaptive Forgetting-Controlled Recurrent Neural Network (AFC-RNN) for pedestrian trajectory prediction. The core idea of AFC-RNN is a novel Adaptive Forgetting Controller (AFC), which controls the forgetting degree of the historical information at each time step explicitly and adaptively. Specifically, AFC first learns memory factors for each time step based on the temporal correlation of observed trajectories using the self-attention mechanism. Then, AFC-RNN applies these memory factors to regulate the forgetting degree of observed features at each time step from RNN. Extensive experiments and ablation studies on ETH, UCY, SDD, and NBA datasets demonstrate that our method outperforms existing state-of-the-art approaches. Additionally, we provide a mathematical analysis to demonstrate the superiority of our adaptive forgetting strategy in the AFC-RNN over traditional RNNs for trajectory forgetting modeling. Yonghao Dong, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Long and Short-Term Collaborative Decision-Making Transformer for Online Action Detection and Anticipation
Chi Zhang 0020, Le Wang 0003, Yuehu Liu |
Pattern Recognit. | 3 |
| 2025 | Recurrent Aligned Network for Generalized Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a crucial component in computer vision and robotics, but remains challenging due to the domain shift problem. Previous studies have tried to tackle this problem by leveraging a portion of trajectory data from the target domain to fine-tune the model. However, such domain adaptation methods are impractical in real-world scenarios, as it is infeasible to collect trajectory data from all potential target domains. In this paper, we study a new task named generalized pedestrian trajectory prediction, with the aim of generalizing the model to unseen domains without accessing their trajectories. To tackle this task, we further introduce a Recurrent Aligned Network (RAN) to minimize the domain gap through domain alignment. Specifically, we devise a recurrent alignment module to effectively align the trajectory feature spaces at both time-state and time-sequence levels by the recurrent alignment strategy. Furthermore, we introduce a pre-aligned representation module to combine social interactions with the recurrent alignment strategy, which aims to consider social interactions during the alignment process instead of just target trajectories. We extensively evaluate our method and compare it with state-of-the-art methods on three widely used benchmarks. The experimental results demonstrate the superior generalization capability of our method. Our work not only fills the gap in the generalization setting for practical pedestrian trajectory prediction, but also sets strong baselines in this field. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Robust Noisy Label Learning via Two-Stream Sample DistillationabstractNoisy label learning aims to learn robust networks under the supervision of noisy labels, which plays a critical role in deep learning. Existing work either conducts sample selection or label correction to deal with noisy labels during the model training process. In this paper, we design a simple yet effective sample selection framework, termed Two-Stream Sample Distillation (TSSD), for noisy label learning, which can extract more high-quality samples with clean labels to improve the robustness of network training. Firstly, a novel Parallel Sample Division (PSD) module is designed to generate acertaintraining set with sufficient reliable positive and negative samples by jointly considering the sample structure in feature space and the human prior in loss space. Secondly, a novel Meta Sample Purification (MSP) module is further designed to mine adequate semi-hard samples from the remaininguncertaintraining set by learning a strong meta classifier with extra golden data. As a result, more and more high-quality samples will be distilled from the noisy training set to train networks robustly in every iteration. Extensive experiments on four benchmark datasets, including CIFAR-10, CIFAR-100, Tiny-ImageNet and Clothing-1M, show that our method has achieved state-of-the-art results over its competitors. Sihan Bai, Sanping Zhou, Le Wang 0003, Nanning Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Visual-Linguistic Feature Alignment With Semantic and Kinematic Guidance for Referring Multi-Object TrackingabstractReferring Multi-Object Tracking (RMOT) aims to dynamically track an arbitrary number of referred targets in a video sequence according to the language expression. Previous methods mainly focus on cross-modal fusion at the feature level with designed structures. However, the insufficient visual-linguistic alignment is prone to causing visual-linguistic mismatches, leading to some targets being tracked but not correctly referred especially when facing the language expression with complex semantics or motion descriptions. To this end, we propose to conduct visual-linguistic alignment with semantic and kinematic guidance to effectively align the visual features with more diverse language expressions. In this paper, we put forward a novel end-to-end RMOT framework SKTrack, which follows the transformer-based architecture with a Language-Guided Decoder (LGD) and a Motion-Aware Aggregator (MAA). In particular, the LGD performs deep semantic interaction layer-by-layer in a single frame to enhance the alignment ability of the model, while the MAA conducts temporal feature fusion and alignment across multiple frames to enable the alignment between visual targets and language expression with motion descriptions. Extensive experiments on the Refer-KITTI and Refer-KITTI-v2 demonstrate that SKTrack achieves state-of-the-art performance and verify the effectiveness of our framework and its components. Sanping Zhou, Le Wang 0003 |
IEEE Trans. Multim. | 4 |
| 2025 | Meta Pairwise Relationship Distillation for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) is challenging due to the lack of ground-truth labels. Most existing methods rely on pseudo labels estimated via iterative clustering and thus are highly susceptible to performance penalties incurred by the inaccurate estimated number of clusters. Alternatively, we utilize the sample pairs with pairwise pseudo labels to guide the feature learning to avoid the dilemma of determining cluster numbers. In this article, we propose a meta pairwise relationship distillation (MPRD) method that incorporates a graph convolutional network (GCN) to provide high-fidelity pairwise relationships to supervise the model training. A small amount of metadata with very-confidence pairwise relationships and the unlabeled pairs with the provided pseudo pairwise relationships participate in the GCN training. Besides, we introduce a hard sample deduction (HSD) module to timely mine the sample pairs with error-prone pairwise pseudo labels to mitigate the misled optimization by noisy labels. Furthermore, since the features of each positive pair represent the same person, we design a positive pair alignment (PPA) module to reduce the redundant information in each feature, which is achieved by minimizing the difference between each positive pair's feature distributions. Extensive experiments on the Market-1501, DukeMTMC-reID, and MSMT17 datasets show that our method outperforms the state-of-the-art unsupervised methods. Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Temporal Correlation Vision Transformer for Video Person Re-IdentificationabstractVideo Person Re-Identification (Re-ID) is a task of retrieving persons from multi-camera surveillance systems. Despite the progress made in leveraging spatio-temporal information in videos, occlusion in dense crowds still hinders further progress. To address this issue, we propose a Temporal Correlation Vision Transformer (TCViT) for video person Re-ID. TCViT consists of a Temporal Correlation Attention (TCA) module and a Learnable Temporal Aggregation (LTA) module. The TCA module is designed to reduce the impact of non-target persons by relative state, while the LTA module is used to aggregate frame-level features based on their completeness. Specifically, TCA is a parameter-free module that first aligns frame-level features to restore semantic coherence in videos and then enhances the features of the target person according to temporal correlation. Additionally, unlike previous methods that treat each frame equally with a pooling layer, LTA introduces a lightweight learnable module to weigh and aggregate frame-level features under the guidance of a classification score. Extensive experiments on four prevalent benchmarks demonstrate that our method achieves state-of-the-art performance in video Re-ID. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
AAAI | 2 |
| 2024 | Towards Generalizable Multi-Object TrackingabstractMulti-Object Tracking (MOT) encompasses various tracking scenarios, each characterized by unique traits. Ef-fective trackers should demonstrate a high degree of gen-eralizability across diverse scenarios. However, existing trackers struggle to accommodate all aspects or necessi-tate hypothesis and experimentation to customize the asso-ciation information (motion and/or appearance) for a given scenario, leading to narrowly tailored solutions with limited generalizability. In this paper, we investigate the factors that influence trackers' generalization to different scenar-ios and concretize them into a set of tracking scenario at-tributes to guide the design of more generalizable trackers. Furthermore, we propose a “point-wise to instance-wise relation” framework for MOT, i.e., GeneralTrack, which can generalize across diverse scenarios while eliminating the need to balance motion and appearance. Thanks to its supe-rior generalizability, our proposed GeneralTrack achieves state-of-the-art performance on multiple benchmarks and demonstrates the potential for domain generalization. Le Wang 0003, Sanping Zhou, Panpan Fu, Gang Hua 0001, Wei Tang 0016 |
CVPR | 2 |
| 2024 | PMT: Progressive Mean Teacher via Exploring Temporal Consistency for Semi-Supervised Medical Image Segmentation
Sanping Zhou, Le Wang 0003, Nanning Zheng 0001 |
ECCV (68) | 3 |
| 2024 | Analysis-by-Synthesis Transformer for Single-View 3D Reconstruction
Dian Jia, Xiaoqian Ruan, Zhiming Zou, Le Wang 0003, Wei Tang 0016 |
ECCV (21) | 5 |
| 2024 | Stepwise Multi-grained Boundary Detector for Point-Supervised Temporal Action Localization
Mengnan Liu 0001, Le Wang 0003, Sanping Zhou, Qilin Zhang 0004, Gang Hua 0001 |
ECCV (7) | 2 |
| 2024 | Learning Anomalies with Normality Prior for Unsupervised Video Anomaly Detection
Haoyue Shi 0002, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
ECCV (6) | 2 |
| 2024 | Multimodal LLM Enhanced Cross-lingual Cross-modal RetrievalabstractCross-lingual cross-modal retrieval (CCR) aims to retrieve visually relevant content based on non-English queries, without relying on human-labeled cross-modal data pairs during training. One popular approach involves utilizing machine translation (MT) to create pseudo-parallel data pairs, establishing correspondence between visual and non-English textual data. However, aligning their representations poses challenges due to the significant semantic gap between vision and text, as well as the lower quality of non-English representations caused by pre-trained encoders and data noise. To overcome these challenges, we propose LECCR, a novel solution that incorporates the multi-modal large language model (MLLM) to improve the alignment between visual and non-English representations. Specifically, we first employ MLLM to generate detailed visual content descriptions and aggregate them into multi-view semantic slots that encapsulate different semantics. Then, we take these semantic slots as internal features and leverage them to interact with the visual features. By doing so, we enhance the semantic information within the visual features, narrowing the semantic gap between modalities and generating local visual semantics for subsequent multi-level matching. Additionally, to further enhance the alignment between visual and non-English features, we introduce softened matching under English guidance. This approach provides more comprehensive and reliable inter-modal correspondences between visual and non-English features. Extensive experiments on four CCR benchmarks, i.e., Multi30K, MSCOCO, VATEX, and MSR-VTT-CN, demonstrate the effectiveness of our proposed method. Code: https://github.com/LiJiaBei-7/leccr. Le Wang 0003, Qiang Zhou 0001, Zhibin Wang 0004, Hao Li 0030, Gang Hua 0001, Wei Tang 0016 |
ACM Multimedia | 2 |
| 2024 | Referencing Where to Focus: Improving Visual Grounding with Referential QueryabstractVisual Grounding aims to localize the referring object in an image given a natural language expression. Recent advancements in DETR-based visual grounding methods have attracted considerable attention, as they directly predict the coordinates of the target object without relying on additional efforts, such as pre-generated proposal candidates or pre-defined anchor boxes. However, existing research primarily focuses on designing stronger multi-modal decoder, which typically generates learnable queries by random initialization or by using linguistic embeddings. This vanilla query generation approach inevitably increases the learning difficulty for the model, as it does not involve any target-related information at the beginning of decoding. Furthermore, they only use the deepest image feature during the query learning process, overlooking the importance of features from other levels. To address these issues, we propose a novel approach, called RefFormer. It consists of the query adaption module that can be seamlessly integrated into CLIP and generate the referential query to provide the prior context for decoder, along with a task-specific decoder. By incorporating the referential query into the decoder, we can effectively mitigate the learning difficulty of the decoder, and accurately concentrate on the target object. Additionally, our proposed query adaption module can also act as an adapter, preserving the rich knowledge within CLIP without the need to tune the parameters of the backbone network. Extensive experiments demonstrate the effectiveness and efficiency of our proposed method, outperforming state-of-the-art approaches on five visual grounding benchmarks. Zhuotao Tian, Qingpei Guo, Sanping Zhou, Ming Yang 0007, Le Wang 0003 |
NeurIPS | 7 |
| 2024 | End-to-end pedestrian trajectory prediction via Efficient Multi-modal Predictors
Sanping Zhou, Le Wang 0003, Liushuai Shi, Yonghao Dong, Gang Hua 0001 |
Comput. Vis. Image Underst. | 3 |
| 2024 | Residual feature learning with hierarchical calibration for gaze estimation
Zhengdan Yin, Sanping Zhou, Le Wang 0003, Gang Hua 0001, Nanning Zheng 0001 |
Mach. Vis. Appl. | 3 |
| 2024 | Adversarial Attack and Defense in Deep RankingabstractDeep Neural Network classifiers are vulnerable to adversarial attacks, where an imperceptible perturbation could result in misclassification. However, the vulnerability of DNN-based image ranking systems remains under-explored. In this paper, we propose two attacks against deep ranking systems, i.e., Candidate Attack and Query Attack, that can raise or lower the rank of chosen candidates by adversarial perturbations. Specifically, the expected ranking order is first represented as a set of inequalities. Then a triplet-like objective function is designed to obtain the optimal perturbation. Conversely, an anti-collapse triplet defense is proposed to improve the ranking model robustness against all proposed attacks, where the model learns to prevent the adversarial attack from pulling the positive and negative samples close to each other. To comprehensively measure the empirical adversarial robustness of a ranking model with our defense, we propose an empirical robustness score, which involves a set of representative attacks against ranking models. Our adversarial ranking attacks and defenses are evaluated on MNIST, Fashion-MNIST, CUB200-2011, CARS196, and Stanford Online Products datasets. Experimental results demonstrate that our attacks can effectively compromise a typical deep ranking system. Nevertheless, our defense can significantly improve the ranking system's robustness and simultaneously mitigate a wide range of attacks. Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Transfer easy to hard: Adversarial contrastive feature learning for unsupervised person re-identification
Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
Pattern Recognit. | 2 |
| 2024 | Bidirectional feature learning network for RGB-D salient object detectionabstractRGB-D salient object detection aims to perform the pixel-wise localization of salient objects from both RGB and depth images, whose challenge mainly comes from how to learn complementary features from each modality. Existing works often use increasingly large models for performance enhancement, which need large memory and time consumption in practice. In this paper, we propose a simple yet effective B idirectional F eature L earning Net work (BFLNet) for RGB-D salient object detection under limited memory and time conditions. To achieve accurate performance with lightweight backbone networks , an effective B idirectional F eature F usion (BFF) module is designed to merge features from both RGB and depth streams, in which the cross-modal fusions and cross-scale fusions are jointly conducted to fuse the immediate features in multiple scales and multiple modals. What is more, a simple D ual C onsistency L oss (DCL) function is designed to prompt cross-modal fusion by keeping the consistency between cross-modal target predictions. Extensive experiments on four benchmark datasets demonstrate that our method has achieved the state-of-the-art performance with high efficiency in RGB-D salient object detection. Code will be available at https://github.com/nightsky-nostar/BFLNet . Ye Niu, Sanping Zhou, Yonghao Dong, Le Wang 0003, Jinjun Wang, Nanning Zheng 0001 |
Pattern Recognit. | 4 |
| 2024 | Disentangled Sample Guidance Learning for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) is challenging due to the lack of ground truth labels. Most existing methods employ iterative clustering to generate pseudo labels for unlabeled training data to guide the learning process. However, how to select samples that are both associated with high-confidence pseudo labels and hard (discriminative) enough remains a critical problem. To address this issue, a disentangled sample guidance learning (DSGL) method is proposed for unsupervised Re-ID. The method consists of disentangled sample mining (DSM) and discriminative feature learning (DFL). DSM disentangles (unlabeled) person images into identity-relevant and identity-irrelevant factors, which are used to construct disentangled positive/negative groups that contain discriminative enough information. DFL incorporates the mined disentangled sample groups into model training by a surrogate disentangled learning loss and a disentangled second-order similarity regularization, to help the model better distinguish the characteristics of different persons. By using the DSGL training strategy, the mAP on Market-1501 and MSMT17 increases by 6.6% and 10.1% when applying the ResNet50 framework, and by 0.6% and 6.9% with the vision transformer (VIT) framework, respectively, validating the effectiveness of the DSGL method. Moreover, DSGL surpasses previous state-of-the-art methods by achieving higher Top-1 accuracy and mAP on the Market-1501, MSMT17, PersonX, and VeRi-776 datasets. The source code for this paper is available at https://github.com/jihaoxuanye/DiseSGL. Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Worst Perception Scenario Search via Recurrent Neural Controller and K-Reciprocal Re-RankingabstractAchieving excellent generalization on perceiving real traffic scenarios with diversity is the long-term goal for building robust autonomous driving systems. A recent theoretical study shows that the generalization on the worst-group of test samples is far more difficult than others. Therefore, we propose to discover potential shortness of certain perception module by analyzing its worst-scenario performance. However, with the benchmark datasets growing huge and tremendous, exhaustive searching for the worst perception scenario (WPS) seems to be time consuming and unnecessary. To address this, we present an automatic searching scheme empowered by reinforcement learning. In this case, worst scenario mining is formulated as the discrete search on the Visual Operation Design Domain (ODD), namely scenario representation, by optimizing LSTM-RNN controller with the worst-performance reward. Moreover, a time-efficient K-reciprocal re-ranking technique is utilized to match the predicted scenario parameters with existing test data. The proposed method has been validated by finding the most challenging scenarios for various vehicle detectors on KITTI, BDD100k and our own benchmark set EVB. Furthermore, searching performances w.r.t different Visual ODDs are investigated and it is found that visual representations through generative adversarial network contribute to a better performance. Chi Zhang 0020, Xiaoning Ma, Liheng Xu, Haoang Lu, Le Wang 0003, Yuanqi Su, Yuehu Liu, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Sparse Pedestrian Character Learning for Trajectory PredictionabstractPedestrian trajectory prediction in a first-person view has recently attracted much attention due to its importance in autonomous driving. Recent work utilizes pedestrian character information, i.e., action and appearance, to improve the learned trajectory embedding and achieves state-of-the-art performance. However, it neglects the invalid and negative pedestrian character information, which is harmful to trajectory representation and thus leads to performance degradation. To address this issue, we present a two-stream sparse-character-based network (TSNet) for pedestrian trajectory prediction. Specifically, TSNet learns the negative-removed characters in the sparse character representation stream to improve the trajectory embedding obtained in the trajectory representation stream. Moreover, to model the negative-removed characters, we propose a novel sparse character graph, including the sparse category and sparse temporal character graphs, to learn the different effects of various characters in category and temporal dimensions, respectively. Extensive experiments on two first-person view datasets, PIE and JAAD, show that our method outperforms existing state-of-the-art methods. In addition, ablation studies demonstrate different effects of various characters and prove that TSNet outperforms approaches without eliminating negative characters. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Single-Shot and Multi-Shot Feature Learning for Multi-Object TrackingabstractMulti-Object Tracking (MOT) remains a vital component of intelligent video analysis, which aims to locate targets and maintain a consistent identity for each target throughout a video sequence. Existing works usually learn a discriminative feature representation, such as motion and appearance, to associate the detections across frames, which are easily affected by mutual occlusion and background clutter in practice. In this paper, we propose a simple yet effective two-stage feature learning paradigm to jointly learn single-shot and multi-shot features for different targets, so as to achieve robust data association in the tracking process. For the detections without being associated, we design a novel single-shot feature learning module to extract discriminative features of each detection, which can efficiently associate targets between adjacent frames. For the tracklets being lost several frames, we design a novel multi-shot feature learning module to extract discriminative features of each tracklet, which can accurately refind these lost targets after a long period. Once equipped with a simple data association logic, the resulting VisualTracker can perform robust MOT based on the single-shot and multi-shot feature representations. Extensive experimental results demonstrate that our method has achieved significant improvements on MOT17 and MOT20 datasets while reaching state-of-the-art performance on DanceTrack dataset. Sanping Zhou, Le Wang 0003, Jinjun Wang, Nanning Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Abnormal Ratios Guided Multi-Phase Self-Training for Weakly-Supervised Video Anomaly DetectionabstractWeakly-supervised Video Anomaly Detection (W-VAD) aims to detect abnormal events in videos given only video-level labels for training. Recent methods relying on multiple instance learning (MIL) and self-training achieve good performance, but they tend to focus on learning easy abnormal patterns while ignoring hard ones, e.g., unusual driving trajectory or over-speeding driving. How to detect hard anomalies is a critical but largely ignored problem in W-VAD. To tackle this challenge, we propose a novel framework, termed Abnormal Ratios guided Multi-phase Self-training (ARMS), for W-VAD. It includes a new abnormal ratio-based MIL (AR-MIL) loss and a new multi-phase self-training paradigm. The AR-MIL loss guides the learning of hard anomalies by enforcing a minimum ratio of abnormal snippets in an abnormal video and no abnormal snippets in a normal video. Our multi-phase self-training paradigm sequentially performs bootstrapping, hard anomalies mining, and adaptive self-training so as to address pseudo labeling on easy anomalies, detect hard anomalies, and setting adaptive abnormal ratios for different videos in a unified framework. Experimental results on three benchmark datasets, i.e., ShanghaiTech, UCF-Crime, and XD-Violence, show that ARMS outperforms all previous state-of-the-art methods and has a great advantage in detecting hard anomalies. Haoyue Shi 0002, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
IEEE Trans. Multim. | 2 |
| 2024 | Inverse Adversarial Diversity Learning for Network EnsembleabstractNetwork ensemble aims to obtain better results by aggregating the predictions of multiple weak networks, in which how to keep the diversity of different networks plays a critical role in the training process. Many existing approaches keep this kind of diversity either by simply using different network initializations or data partitions, which often requires repeated attempts to pursue a relatively high performance. In this article, we propose a novel inverse adversarial diversity learning (IADL) method to learn a simple yet effective ensemble regime, which can be easily implemented in the following two steps. First, we take each weak network as a generator and design a discriminator to judge the difference between the features extracted by different weak networks. Second, we present an inverse adversarial diversity constraint to push the discriminator to cheat generators that all the resulting features of the same image are too similar to distinguish each other. As a result, diverse features will be extracted by these weak networks through a min-max optimization. What is more, our method can be applied to a variety of tasks, such as image classification and image retrieval, by applying a multitask learning objective function to train all these weak networks in an end-to-end manner. We conduct extensive experiments on the CIFAR-10, CIFAR-100, CUB200-2011, and CARS196 datasets, in which the results show that our method significantly outperforms most of the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Le Wang 0003, Xingyu Wan, Siqi Hui, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Multi-Stream Representation Learning for Pedestrian Trajectory PredictionabstractForecasting the future trajectory of pedestrians is an important task in computer vision with a range of applications, from security cameras to autonomous driving. It is very challenging because pedestrians not only move individually across time but also interact spatially, and the spatial and temporal information is deeply coupled with one another in a multi-agent scenario. Learning such complex spatio-temporal correlation is a fundamental issue in pedestrian trajectory prediction. Inspired by the procedure that the hippocampus processes and integrates spatio-temporal information to form memories, we propose a novel multi-stream representation learning module to learn complex spatio-temporal features of pedestrian trajectory. Specifically, we learn temporal, spatial and cross spatio-temporal correlation features in three respective pathways and then adaptively integrate these features with learnable weights by a gated network. Besides, we leverage the sparse attention gate to select informative interactions and correlations brought by complex spatio-temporal modeling and reduce complexity of our model. We evaluate our proposed method on two commonly used datasets, i.e. ETH-UCY and SDD, and the experimental results demonstrate our method achieves the state-of-the-art performance. Code: https://github.com/YuxuanIAIR/MSRL-master Le Wang 0003, Sanping Zhou, Jinghai Duan, Gang Hua 0001, Wei Tang 0016 |
AAAI | 2 |
| 2023 | Progressive Backdoor Erasing via connecting Backdoor and Adversarial AttacksabstractDeep neural networks (DNNs) are known to be vulnera-ble to both backdoor attacks as well as adversarial attacks. In the literature, these two types of attacks are commonly treated as distinct problems and solved separately, since they belong to training-time and inference-time attacks respectively. However, in this paper we find an intriguing connection between them: for a model planted with backdoors, we observe that its adversarial examples have similar behaviors as its triggered images, i.e., both activate the same subset of DNN neurons. It indicates that planting a back-door into a model will significantly affect the model's adversarial examples. Based on these observations, a novel Progressive Backdoor Erasing (PBE) algorithm is proposed to progressively purify the infected model by leveraging un-targeted adversarial attacks. Different from previous back-door defense methods, one significant advantage of our approach is that it can erase backdoor even when the clean extra dataset is unavailable. We empirically show that, against 5 state-of-the-art backdoor attacks, our PBE can effectively erase the backdoor without obvious performance degradation on clean samples and outperforms existing de-fense methods. Bingxu Mu, Zhenxing Niu, Le Wang 0003, Xue Wang 0010, Qiguang Miao, Rong Jin 0001, Gang Hua 0001 |
CVPR | 3 |
| 2023 | MotionTrack: Learning Robust Short-Term and Long-Term Motions for Multi-Object TrackingabstractThe main challenge of Multi-Object Tracking (MOT) lies in maintaining a continuous trajectory for each target. Existing methods often learn reliable motion patterns to match the same target between adjacent frames and discriminative appearance features to re-identify the lost targets after a long period. However, the reliability of motion prediction and the discriminability of appearances can be easily hurt by dense crowds and extreme occlusions in the tracking process. In this paper, we propose a simple yet effective multi-object tracker, i.e., MotionTrack, which learns robust short-term and long-term motions in a unified framework to associate trajectories from a short to long range. For dense crowds, we design a novel Interaction Module to learn interaction-aware motions from short-term trajectories, which can estimate the complex movement of each target. For extreme occlusions, we build a novel Refind Module to learn reliable long-term motions from the target's history trajectory, which can link the interrupted trajectory with its corresponding detection. Our Interaction Module and Refind Module are embedded in the well-known tracking-by-detection paradigm, which can work in tandem to maintain superior performance. Extensive experimental results on MOT17 and MOT20 datasets demonstrate the superiority of our approach in challenging scenarios, and it achieves state-of-the-art performances at various MOT metrics. Code is available at https://github.com/qwomeng/MotionTrack. Sanping Zhou, Le Wang 0003, Jinghai Duan, Gang Hua 0001, Wei Tang 0016 |
CVPR | 3 |
| 2023 | Sparse Instance Conditioned Multimodal Trajectory PredictionabstractPedestrian trajectory prediction is critical in many vision tasks but challenging due to the multimodality of the future trajectory. Most existing methods predict multi-modal trajectories conditioned by goals (future endpoints) or instances (all future points). However, goal-conditioned methods ignore the intermediate process and instance-conditioned methods ignore the stochasticity of pedestrian motions. In this paper, we propose a simple yet effective Sparse Instance Conditioned Network (SICNet), which gives a balanced solution between goal-conditioned and instance-conditioned methods. Specifically, SICNet learns comprehensive sparse instances, i.e., representative points of the future trajectory, through a mask generated by a long short-term memory encoder and uses the memory mechanism to store and retrieve such sparse instances. Hence SICNet can decode the observed trajectory into the future prediction conditioned on the stored sparse instance. Moreover, we design a memory refinement module that refines the retrieved sparse instances from the memory to reduce memory recall errors. Extensive experiments on ETH-UCY and SDD datasets show that our method outperforms existing state-of-the-art methods. In addition, ablation studies demonstrate the superiority of our method compared with goal-conditioned and instance-conditioned approaches. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
ICCV | 2 |
| 2023 | Parallel Attention Interaction Network for Few-Shot Skeleton-based Action RecognitionabstractLearning discriminative features from very few labeled samples to identify novel classes has received increasing attention in skeleton-based action recognition. Existing works aim to learn action-specific embeddings by exploiting either intra-skeleton or inter-skeleton spatial associations, which may lead to less discriminative representations. To address these issues, we propose a novel Parallel Attention Interaction Network (PAINet) that incorporates two complementary branches to strengthen the match by inter-skeleton and intraskeleton correlation. Specifically, a topology encoding module utilizing topology and physical information is proposed to enhance the modeling of interactive parts and joint pairs in both branches. In the Cross Spatial Alignment branch, we employ a spatial cross-attention module to establish joint associations across sequences, and a directional Average Symmetric Surface Metric is introduced to locate the closest temporal similarity. In parallel, the Cross Temporal Alignment branch incorporates a spatial self-attention module to aggregate spatial context within sequences as well as applies the temporal cross-attention network to correct misalignment temporally and calculate similarity. Extensive experiments on three skeleton benchmarks, namely NTU-T, NTU-S, and Kinetics, demonstrate the superiority of our framework and consistently outperform state-of-the-art methods. Sanping Zhou, Le Wang 0003, Gang Hua 0001 |
ICCV | 3 |
| 2023 | Trajectory Unified Transformer for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is an essential link to understanding human behavior. Recent work achieves state-of-the-art performance gained from hand-designed post-processing, e.g., clustering. However, this post-processing suffers from expensive inference time and neglects the probability that the predicted trajectory disturbs downstream safety decisions. In this paper, we present Trajectory Unified TRansformer, called TUTR, which unifies the trajectory prediction components, social interaction, and multimodal trajectory prediction, into a transformer encoder-decoder architecture to effectively remove the need for post-processing. Specifically, TUTR parses the relationships across various motion modes using an explicit global prediction and an implicit mode-level transformer encoder. Then, TUTR attends to the social interactions with neighbors by a social-level transformer decoder. Finally, a dual prediction forecasts diverse trajectories and corresponding probabilities in parallel without post-processing. TUTR achieves state-of-the-art accuracy performance and improvements in inference speed of about 10× - 40× compared to previous well-tuned state-of-the-art methods using post-processing. Liushuai Shi, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
ICCV | 2 |
| 2023 | Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action LocalizationabstractSemi-Supervised Temporal Action Localization (SS-TAL) aims to improve the generalization ability of action detectors with large-scale unlabeled videos. Albeit the recent advancement, one of the major challenges still remains: noisy pseudo labels hinder efficient learning on abundant unlabeled videos, embodied as location biases and category errors. In this paper, we dive deep into such an important but understudied dilemma. To this end, we propose a unified framework, termed Noisy Pseudo-Label Learning, to handle both location biases and category errors. Specifically, our method is featured with (1) Noisy Label Ranking to rank pseudo labels based on the semantic confidence and boundary reliability, (2) Noisy Label Filtering to address the class-imbalance problem of pseudo labels caused by category errors, (3) Noisy Label Learning to penalize in-consistent boundary predictions to achieve noise-tolerant learning for heavy location biases. As a result, our method could effectively handle the label noise problem and improve the utilization of a large amount of unlabeled videos. Extensive experiments on THUMOS14 and ActivityNet v1.3 demonstrate the effectiveness of our method. The code is available at github.com/kunnxia/NPL. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
ICCV | 2 |
| 2023 | Representing Multimodal Behaviors With Mean Location for Pedestrian Trajectory PredictionabstractRepresenting multimodal behaviors is a critical challenge for pedestrian trajectory prediction. Previous methods commonly represent this multimodality with multiple latent variables repeatedly sampled from a latent space, encountering difficulties in interpretable trajectory prediction. Moreover, the latent space is usually built by encoding global interaction into future trajectory, which inevitably introduces superfluous interactions and thus leads to performance reduction. To tackle these issues, we propose a novel Interpretable Multimodality Predictor (IMP) for pedestrian trajectory prediction, whose core is to represent a specific mode by its mean location. We model the distribution of mean location as a Gaussian Mixture Model (GMM) conditioned on sparse spatio-temporal features, and sample multiple mean locations from the decoupled components of GMM to encourage multimodality. Our IMP brings four-fold benefits: 1) Interpretable prediction to provide semantics about the motion behavior of a specific mode; 2) Friendly visualization to present multimodal behaviors; 3) Well theoretical feasibility to estimate the distribution of mean locations supported by the central-limit theorem; 4) Effective sparse spatio-temporal features to reduce superfluous interactions and model temporal continuity of interaction. Extensive experiments validate that our IMP not only outperforms state-of-the-art methods but also can achieve a controllable prediction by customizing the corresponding mean location. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Adaptive Two-Stream Consensus Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (W-TAL) aims to classify and localize all action instances in untrimmed videos under only video-level supervision. Without frame-level annotations, it is challenging for W-TAL methods to clearly distinguish actions and background, which severely degrades the action boundary localization and action proposal scoring. In this paper, we present an adaptive two-stream consensus network (A-TSCN) to address this problem. Our A-TSCN features an iterative refinement training scheme: a frame-level pseudo ground truth is generated and iteratively updated from a late-fusion activation sequence, and used to provide frame-level supervision for improved model training. Besides, we introduce an adaptive attention normalization loss, which adaptively selects action and background snippets according to video attention distribution. By differentiating the attention values of the selected action snippets and background snippets, it forces the predicted attention to act as a binary selection and promotes the precise localization of action boundaries. Furthermore, we propose a video-level and a snippet-level uncertainty estimator, and they can mitigate the adverse effect caused by learning from noisy pseudo ground truth. Experiments conducted on the THUMOS14, ActivityNet v1.2, ActivityNet v1.3, and HACS datasets show that our A-TSCN outperforms current state-of-the-art methods, and even achieves comparable performance with several fully-supervised methods. Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Nanning Zheng 0001, David S. Doermann, Junsong Yuan 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | ContextLoc++: A Unified Context Model for Temporal Action LocalizationabstractEffectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by enriching the local, global and multi-scale contexts in the popular two-stage temporal localization framework. Our proposed model, dubbed ContextLoc++, can be divided into three sub-networks: L-Net, G-Net, and M-Net. L-Net enriches the local context via fine-grained modeling of snippet-level features, which is formulated as a query-and-retrieval process. Furthermore, the spatial and temporal snippet-level features, functioning as keys and values, are fused by temporal gating. G-Net enriches the global context via higher-level modeling of the video-level representation. In addition, we introduce a novel context adaptation module to adapt the global context to different proposals. M-Net further fuses the local and global contexts with multi-scale proposal features. Specially, proposal-level features from multi-scale video snippets can focus on different action characteristics. Short-term snippets with fewer frames pay attention to the action details while long-term snippets with more frames focus on the action variations. Experiments on the THUMOS14 and ActivityNet v1.3 datasets validate the efficacy of our method against existing state-of-the-art TAL algorithms. Zixin Zhu, Le Wang 0003, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Memory-augmented appearance-motion network for video anomaly detection
Le Wang 0003, Junwen Tian, Sanping Zhou, Haoyue Shi 0002, Gang Hua 0001 |
Pattern Recognit. | 1 |
| 2023 | Instance Motion Tendency Learning for Video Panoptic SegmentationabstractVideo panoptic segmentation is an important but challenging task in computer vision. It not only performs panoptic segmentation of each frame, but also associates the same instance across adjacent frames. Due to the lack of temporal coherence modeling, most existing approaches often generate identity switches during instance association, and they cannot handle ambiguous segmentation boundaries caused by motion blur. To address these difficult issues, we introduce a simple yet effective Instance Motion Tendency Network (IMTNet) for video panoptic segmentation. It learns a global motion tendency map for instance association, and a hierarchical classifier for motion boundary refinement. Specifically, a Global Motion Tendency Module (GMTM) is designed to learn robust motion features from optical flows, which can directly associate each instance in the previous frame to the corresponding instance in the current frame. In addition, we propose a Motion Boundary Refinement Module (MBRM) to learn a hierarchical classifier to handle the boundary pixels of moving targets, which can effectively revise the inaccurate segmentation predictions. Experimental results on both Cityscapes and Cityscapes-VPS datasets show that our IMTNet outperforms most state-of-the-art approaches. Le Wang 0003, Hongzhen Liu, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Multi-Panda TrackingabstractMulti-Panda Tracking (MPT) is a video-based tracking task for panda individuals, which is conducive to the observation and measurement of distribution and status of pandas. Different from tracking general objects such as pedestrians and vehicles, MPT is extremely challenging due to the indistinguishable appearances and diversified postures of pandas. In this case, existing tracking methods cannot appropriately tackle with the excessive occlusion between different panda individuals, hence suffering from identity switch, missing and inaccurate detections. To address these problems, we propose a simple yet effective MPT framework in the tracking-by-detection paradigm, which is benefited both from a short-term prediction filtering module and a discriminative feature learning network. In particular, the short-term prediction filtering module introduces similarity learning to enhance the temporal consistency among detections, which is capable of supplementing the missing detections and discarding false positive detections. Besides, the discriminative feature learning network leverages a two-branch network to learn both local and global discriminative features, so as to distinguish different panda individuals with a very similar appearance with a subtle difference. To evaluate the proposed method, we annotate a large-scale MPT dataset, named PANDA2021, which is particularly challenging due to the similar appearance and dramatic occlusion between panda individuals. Experiments on PANDA2021 demonstrate that the proposed MPT method significantly outperforms the competing methods. Moreover, experimental results on pedestrian tracking dataset MOT16 further demonstrate that the proposed MPT method achieves comparative performance with competing methods. Le Wang 0003, Sanping Zhou, Nanning Zheng 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Adaptive Ladder Loss for Learning Coherent Visual-Semantic EmbeddingabstractFor visual-semantic embedding, the existing methods normally treat the relevance between queries and candidates in a bipolar way – relevant or irrelevant, and all “irrelevant” candidates are uniformly pushed away from the query by an equal margin in the embedding space, regardless of their various proximity to the query. This practice disregards relatively discriminative information and could lead to suboptimal ranking in the retrieval results and poorer user experience, especially in the long-tail query scenario where a matching candidate may not necessarily exist. In this paper, we introduce a continuous variable to model the relevance degree between queries and multiple candidates, and propose to learn a coherent embedding space, where candidates with higher relevance degrees are mapped closer to the query than those with lower relevance degrees. In particular, the new ladder loss is proposed by extending the triplet loss inequality to a more general inequality chain, which implements variable push-away margins according to respective relevance degrees. To adapt to the varying mini-batch statistics and improve the efficiency of the ladder loss, we also propose a Silhouette score-based method to adaptively decide the ladder level and hence the underlying inequality chain. In addition, a proper Coherent Score metric is proposed to better measure the ranking results including those “irrelevant” candidates. Extensive experiments on multiple datasets validate the efficacy of our proposed method, which achieves significant improvement over existing state-of-the-art methods. Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Exploring Action Centers for Temporal Action LocalizationabstractTemporal action localization aims at detecting the temporal intervals of human actions in untrimmed videos. Most previous methods rely on locating and matching the start and end times of actions. However, action boundaries are ambiguous and uncertain in nature, which leads to inaccurate action localization and a lot of false positives. In this paper, we introduce a new framework for temporal action localization. It explicitly models temporal action centers to reduce unreliable action detection results caused by ambiguous action boundaries. Since action centers are highly related to semantic actions, they can be detected more reliably than the conventional action boundaries. As a result, our framework can exclude false positives and promote high-quality proposals. Based on action centers, we propose a triplet feature fusion mechanism. It performs neural message passing among the boundaries and the center as well as contextual regions outside of the proposal to enrich its representation. In addition, we introduce a centerness scoring method to suppress proposals deviating from the centers of action instances. Consequently, our network can retrieve high-quality action proposals and locate actions more precisely. Experimental results show our method outperforms state-of-the-art methods on the THUMOS14 and ActivityNet v1.3 datasets. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
IEEE Trans. Multim. | 2 |
| 2022 | Complementary Attention Gated Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is crucial in many practical applications due to the diversity of pedestrian movements, such as social interactions and individual motion behaviors. With similar observable trajectories and social environments, different pedestrians may make completely different future decisions. However, most existing methods only focus on the frequent modal of the trajectory and thus are difficult to generalize to the peculiar scenario, which leads to the decline of the multimodal fitting ability when facing similar scenarios. In this paper, we propose a complementary attention gated network (CAGN) for pedestrian trajectory prediction, in which a dual-path architecture including normal and inverse attention is proposed to capture both frequent and peculiar modals in spatial and temporal patterns, respectively. Specifically, a complementary block is proposed to guide normal and inverse attention, which are then be summed with learnable weights to get attention features by a gated network. Finally, multiple trajectory distributions are estimated based on the fused spatio-temporal attention features due to the multimodality of future trajectory. Experimental results on benchmark datasets, i.e., the ETH, and the UCY, demonstrate that our method outperforms state-of-the-art methods by 13.8% in Average Displacement Error (ADE) and 10.4% in Final Displacement Error (FDE). Code will be available at https://github.com/jinghaiD/CAGN Jinghai Duan, Le Wang 0003, Chengjiang Long, Sanping Zhou, Fang Zheng 0009, Liushuai Shi, Gang Hua 0001 |
AAAI | 2 |
| 2022 | Social Interpretable Tree for Pedestrian Trajectory PredictionabstractUnderstanding the multiple socially-acceptable future behaviors is an essential task for many vision applications. In this paper, we propose a tree-based method, termed as Social Interpretable Tree (SIT), to address this multi-modal prediction task, where a hand-crafted tree is built depending on the prior information of observed trajectory to model multiple future trajectories. Specifically, a path in the tree from the root to leaf represents an individual possible future trajectory. SIT employs a coarse-to-fine optimization strategy, in which the tree is first built by high-order velocity to balance the complexity and coverage of the tree and then optimized greedily to encourage multimodality. Finally, a teacher-forcing refining operation is used to predict the final fine trajectory. Compared with prior methods which leverage implicit latent variables to represent possible future trajectories, the path in the tree can explicitly explain the rough moving behaviors (e.g., go straight and then turn right), and thus provides better interpretability. Despite the hand-crafted tree, the experimental results on ETH-UCY and Stanford Drone datasets demonstrate that our method is capable of matching or exceeding the performance of state-of-the-art methods. Interestingly, the experiments show that the raw built tree without training outperforms many prior deep neural network based approaches. Meanwhile, our method presents sufficient flexibility in long-term prediction and different best-of-K predictions. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Fang Zheng 0009, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 2 |
| 2022 | Learning Disentangled Classification and Localization Representations for Temporal Action LocalizationabstractA common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that this shared representation focuses on the most discriminative frames for classification, e.g., ``take-offs" rather than ``run-ups" in distinguishing ``high jump" and ``long jump", while frames most relevant to localization, such as the start and end frames of an action, are largely ignored. In other words, such a shared representation can not simultaneously handle both classification and localization tasks well, and it makes precise TAL difficult. To address this challenge, this paper disentangles the shared representation into classification and localization representations. The disentangled classification representation focuses on the most discriminative frames, and the disentangled localization representation focuses on the action phase as well as the action start and end. Our model could be divided into two sub-networks, i.e., the disentanglement network and the context-based aggregation network. The disentanglement network is an autoencoder to learn orthogonal hidden variables of classification and localization. The context-based aggregation network aggregates the classification and localization representations by modeling local and global contexts. We evaluate our proposed method on two popular benchmarks for TAL, which outperforms all state-of-the-art methods. Zixin Zhu, Le Wang 0003, Wei Tang 0016, Ziyi Liu 0001, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 2 |
| 2022 | Learning to Refactor Action and Co-occurrence Features for Temporal Action LocalizationabstractThe main challenge of Temporal Action Localization is to retrieve subtle human actions from various co-occurring ingredients, e.g., context and background, in an untrimmed video. While prior approaches have achieved substantial progress through devising advanced action detectors, they still suffer from these co-occurring ingredients which often dominate the actual action content in videos. In this paper, we explore two orthogonal but complementary aspects of a video snippet, i.e., the action features and the co-occurrence features. Especially, we develop a novel auxiliary task by decoupling these two types of features within a video snippet and recombining them to generate a new feature representation with more salient action information for accurate action localization. We term our method RefactorNet, which first explicitly factorizes the action content and regularizes its co-occurrence features, and then synthesizes a new action-dominated video representation. Extensive experimental results and ablation studies on THUMOS14 and ActivityNet v 1.3 demonstrate that our new representation, combined with a simple action detector, can significantly improve the action localization performance. Le Wang 0003, Sanping Zhou, Nanning Zheng 0001, Wei Tang 0016 |
CVPR | 2 |
| 2022 | Switching: understanding the class-reversed sampling in tail sample memorization
Chi Zhang 0020, Benyi Hu, Yuhang Liuzhang, Le Wang 0003, Yuehu Liu |
Mach. Learn. | 4 |
| 2022 | Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksabstractGiven only video-level action categorical labels during training, weakly-supervised temporal action localization (WS-TAL) learns to detect action instances and locates their temporal boundaries in untrimmed videos. Compared to its fully supervised counterpart, WS-TAL is more cost-effective in data labeling and thus favorable in practical applications. However, the coarse video-level supervision inevitably incurs ambiguities in action localization, especially in untrimmed videos containing multiple action instances. To overcome this challenge, we observe that significant temporal contrasts among video snippets, e.g., caused by temporal discontinuities and sudden changes, often occur around true action boundaries. This motivates us to introduce a Contrast-based Localization EvaluAtioN Network (CleanNet), whose core is a new temporal action proposal evaluator, which provides fine-grained pseudo supervision by leveraging the temporal contrasts among snippet-level classification predictions. As a result, the uncertainty in locating action instances can be resolved via evaluating their temporal contrast scores. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Besides, we also explore the usage of temporal contrast on temporal action proposal (TAP) generation task, which we believe is the first attempt with the weak supervision setting. Experiments on the THUMOS14, ActivityNet v1.2 and v1.3 datasets validate the efficacy of our method against existing state-of-the-art WS-TAL algorithms. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Loss functions for pose guided person image generation
Haoyue Shi 0002, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001, Wei Tang 0016 |
Pattern Recognit. | 2 |
| 2022 | Dual relation network for temporal action localization
Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
Pattern Recognit. | 2 |
| 2022 | Local to Global Feature Learning for Salient Object Detection
Xuelu Feng, Sanping Zhou, Zixin Zhu, Le Wang 0003, Gang Hua 0001 |
Pattern Recognit. Lett. | 4 |
| 2022 | Density-Aware Haze Image Synthesis by Self-Supervised Content-Style DisentanglementabstractThe key procedure of haze image synthesis with adversarial training lies in the disentanglement of the feature involved only in haze synthesis, i.e.,the style feature, from the feature representing the invariant semantic content, i.e.,the content feature. Previous methods introduced a binary classifier to constrain the domain membership from being distinguished through the learned content feature during the training stage, thereby the style information is separated from the content feature. However, we find that these methods cannot achieve complete content-style disentanglement. The entanglement of the flawed style feature with content information inevitably leads to the inferior rendering of haze images. To address this issue, we propose a self-supervised style regression model with stochastic linear interpolation that can suppress the content information in the style feature. Ablative experiments demonstrate the disentangling completeness and its superiority in density-aware haze image synthesis. Moreover, the synthesized haze data are applied to test the generalization ability of vehicle detectors. Further study on the relation between haze density and detection performance shows that haze has an obvious impact on the generalization ability of vehicle detectors and that the degree of performance degradation is linearly correlated to the haze density, which in turn validates the effectiveness of the proposed method. Chi Zhang 0020, Zihang Lin, Liheng Xu, Zongliang Li, Wei Tang 0016, Yuehu Liu, Gaofeng Meng, Le Wang 0003, Li Li 0013 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2022 | Action Coherence Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised Temporal Action Localization (W-TAL) aims at simultaneously classifying and locating all action instances with only video-level supervision. However, current W-TAL methods have two limitations. First, they ignore the difference in video representations between an action instance and its surrounding background when generating and scoring action proposals. Second, the unique characteristics of the RGB frames and optical flow are largely ignored when fusing these two modalities. To address these problems, an Action Coherence Network (ACN) is proposed in this paper. Its core is a new coherence loss which exploits both classification predictions and video content representations to supervise action boundary regression and thus leads to more accurate action localization results. Besides, the proposed ACN explicitly takes into account the specific characteristics of RGB frames and optical flow by training two separate sub-networks, each of which is able to generate modality-specific action proposals independently. Finally, to take advantage of the complementary action proposals generated by two streams, a novel fusion module is introduced to reconcile them and obtain the final action localization results. Experiments on the THUMOS14 and ActivityNet datasets show that our ACN outperforms the state-of-the-art W-TAL methods, and is even comparable to some recent fully-supervised methods. Particularly, ACN achieves a mean average precision of 26.4% on the THUMOS14 dataset under the IoU threshold 0.5. Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Multinetwork Collaborative Feature Learning for Semisupervised Person ReidentificationabstractPerson reidentification (Re-ID) aims at matching images of the same identity captured from the disjoint camera views, which remains a very challenging problem due to the large cross-view appearance variations. In practice, the mainstream methods usually learn a discriminative feature representation using a deep neural network, which needs a large number of labeled samples in the training process. In this article, we design a simple yet effective multinetwork collaborative feature learning (MCFL) framework to alleviate the data annotation requirement for person Re-ID, which can confidently estimate the pseudolabels of unlabeled sample pairs and consistently learn the discriminative features of input images. To keep the precision of pseudolabels, we further build a novel self-paced collaborative regularizer to extensively exchange the weight information of unlabeled sample pairs between different networks. Once the pseudolabels are correctly estimated, we take the corresponding sample pairs into the training process, which is beneficial to learn more discriminative features for person Re-ID. Extensive experimental results on the Market1501, DukeMTMC, and CUHK03 data sets have shown that our method outperforms most of the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Deyu Meng, Le Wang 0003, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and ContextabstractWeakly-supervised Temporal Action Localization (WS-TAL) methods learn to localize temporal starts and ends of action instances in a video under only video-level supervision. Existing WS-TAL methods rely on deep features learned for action recognition. However, due to the mismatch between classification and localization, these features cannot distinguish the frequently co-occurring contextual background, i.e., the context, and the actual action instances. We term this challenge action-context confusion, and it will adversely affect the action localization accuracy. To address this challenge, we introduce a framework that learns two feature subspaces respectively for actions and their context. By explicitly accounting for action visual elements, the action instances can be localized more precisely without the distraction from the context. To facilitate the learning of these two feature subspaces with only video-level categorical labels, we leverage the predictions from both spatial and temporal streams for snippets grouping. In addition, an unsupervised learning task is introduced to make the proposed module focus on mining temporal information. The proposed approach outperforms state-of-the-art WS-TAL methods on three benchmarks, i.e., THUMOS14, ActivityNet v1.2 and v1.3 datasets. Ziyi Liu 0001, Le Wang 0003, Wei Tang 0016, Junsong Yuan 0001, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 2 |
| 2021 | ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action LocalizationabstractThe object of Weakly-supervised Temporal Action Localization (WS-TAL) is to localize all action instances in an untrimmed video with only video-level supervision. Due to the lack of frame-level annotations during training, current WS-TAL methods rely on attention mechanisms to localize the foreground snippets or frames that contribute to the video-level classification task. This strategy frequently confuse context with the actual action, in the localization result. Separating action and context is a core problem for precise WS-TAL, but it is very challenging and has been largely ignored in the literature. In this paper, we introduce an Action-Context Separation Network (ACSNet) that explicitly takes into account context for accurate action localization. It consists of two branches (i.e., the Foreground-Background branch and the Action-Context branch). The Foreground-Background branch first distinguishes foreground from background within the entire video while the Action-Context branch further separates the foreground as action and context. We associate video snippets with two latent components (i.e., a positive component and a negative component), and their different combinations can effectively characterize foreground, action and context. Furthermore, we introduce extended labels with auxiliary context categories to facilitate the learning of action-context separation. Experiments on THUMOS14 and ActivityNet v1.2/v1.3 datasets demonstrate the ACSNet outperforms existing state-of-the-art WS-TAL methods by a large margin. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Junsong Yuan 0001, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 2 |
| 2021 | SGCN: Sparse Graph Convolution Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a key technology in autopilot, which remains to be very challenging due to complex interactions between pedestrians. However, previous works based on dense undirected interaction suffer from modeling superfluous interactions and neglect of trajectory motion tendency, and thus inevitably result in a considerable deviance from the reality. To cope with these issues, we present a Sparse Graph Convolution Network (SGCN) for pedestrian trajectory prediction. Specifically, the SGCN explicitly models the sparse directed interaction with a sparse directed spatial graph to capture adaptive interaction pedestrians. Meanwhile, we use a sparse directed temporal graph to model the motion tendency, thus to facilitate the prediction based on the observed direction. Finally, parameters of a bi-Gaussian distribution for trajectory prediction are estimated by fusing the above two sparse graphs. We evaluate our proposed method on the ETH and UCY datasets, and the experimental results show our method outperforms comparative state-of-the-art methods by 9% in Average Displacement Error (ADE) and 13% in Final Displacement Error (FDE). Notably, visualizations indicate that our method can capture adaptive interactions between pedestrians and their effective motion tendencies. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Zhenxing Niu, Gang Hua 0001 |
CVPR | 2 |
| 2021 | Meta Pairwise Relationship Distillation for Unsupervised Person Re-identificationabstractUnsupervised person re-identification (Re-ID) remains challenging due to the lack of ground-truth labels. Existing methods often rely on estimated pseudo labels via iterative clustering and classification, and they are unfortunately highly susceptible to performance penalties incurred by the inaccurate estimated number of clusters. Alternatively, we propose the Meta Pairwise Relationship Distillation (MPRD) method to estimate the pseudo labels of sample pairs for unsupervised person Re-ID. Specifically, it consists of a Convolutional Neural Network (CNN) and Graph Convolutional Network (GCN), in which the GCN estimates the pseudo labels of sample pairs based on the current features extracted by CNN, and the CNN learns better features by involving high-fidelity positive and negative sample pairs imposed by GCN. To achieve this goal, a small amount of labeled samples are used to guide GCN training, which can distill meta knowledge to judge the difference in the neighborhood structure between positive and negative sample pairs. Extensive experiments on Market-1501, DukeMTMC-reID and MSMT17 datasets show that our method outperforms the state-of-the-art approaches. Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 2 |
| 2021 | Unlimited Neighborhood Interaction for Heterogeneous Trajectory PredictionabstractUnderstanding complex social interactions among agents is a key challenge for trajectory prediction. Most existing methods consider the interactions between pairwise traffic agents or in a local area, while the nature of interactions is unlimited, involving an uncertain number of agents and non-local areas simultaneously. Besides, they treat heterogeneous traffic agents the same, namely those among agents of different categories, while neglecting people’s diverse reaction patterns toward traffic agents in different categories. To address these problems, we propose a simple yet effective Unlimited Neighborhood Interaction Network (UNIN), which predicts trajectories of heterogeneous agents in multiple categories. Specifically, the proposed unlimited neighborhood interaction module generates the fused-features of all agents involved in an interaction simultaneously, which is adaptive to any number of agents and any range of interaction area. Meanwhile, a hierarchical graph attention module is proposed to obtain category-to-category interaction and agent-to-agent interaction. Finally, parameters of a Gaussian Mixture Model are estimated for generating the future trajectories. Extensive experimental results on benchmark datasets demonstrate a significant performance improvement of our method over the state-of-the-art methods. Fang Zheng 0009, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 2 |
| 2021 | Practical Relative Order Attack in Deep RankingabstractRecent studies unveil the vulnerabilities of deep ranking models, where an imperceptible perturbation can trigger dramatic changes in the ranking result. While previous attempts focus on manipulating absolute ranks of certain candidates, the possibility of adjusting their relative order remains under-explored. In this paper, we formulate a new adversarial attack against deep ranking systems, i.e., the Order Attack, which covertly alters the relative order among a selected set of candidates according to an attacker-specified permutation, with limited interference to other unrelated candidates. Specifically, it is formulated as a triplet-style loss imposing an inequality chain reflecting the specified permutation. However, direct optimization of such white-box objective is infeasible in a real-world attack scenario due to various black-box limitations. To cope with them, we propose a Short-range Ranking Correlation metric as a surrogate objective for black-box Order Attack to approximate the white-box method. The Order Attack is evaluated on the Fashion-MNIST and Stanford-Online-Products datasets under both white-box and black-box threat models. The black-box attack is also successfully implemented on a major e-commerce platform. Comprehensive experimental evaluations demonstrate the effectiveness of the proposed methods, revealing a new type of ranking model vulnerability. Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 2 |
| 2021 | Enriching Local and Global Contexts for Temporal Action LocalizationabstractEffectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by enriching both the local and global contexts in the popular two-stage temporal localization framework, where action proposals are first generated followed by action classification and temporal boundary regression. Our proposed model, dubbed ContextLoc, can be divided into three sub-networks: L-Net, G-Net and P-Net. L-Net enriches the local context via fine-grained modeling of snippet-level features, which is formulated as a query-and-retrieval process. G-Net enriches the global context via higher-level modeling of the video-level representation. In addition, we introduce a novel context adaptation module to adapt the global context to different proposals. P-Net further models the context-aware inter-proposal relations. We explore two existing models to be the P-Net in our experiments. The efficacy of our proposed method is validated by experimental results on the THUMOS14 (54.3% at [email protected]) and ActivityNet v1.3 (56.01% at [email protected]) datasets, which outperforms recent states of the art. Code is available at https://github.com/buxiangzhiren/ContextLoc. Zixin Zhu, Wei Tang 0016, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 3 |
| 2021 | Graph-based temporal action co-localization from an untrimmed video
Le Wang 0003, Changbo Zhai, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
Neurocomputing | 1 |
| 2021 | Optimal Schedule of Secure Transmissions for Remote State Estimation Against EavesdroppingabstractIn this article, we investigate the privacy issue of the remote state estimation problem in cyber-physical systems. Specifically, in the presence of an eavesdropper, a sensor observes a discrete linear time-invariant process and then sends the measurements to a remote state estimator with arbitrary finite kinds of transmission options through an unreliable wireless channel. The transmission options of the sensor are in silence state or transmitting aided by injection noise with different energy levels. The eavesdropper wiretaps the channel when the sensor transmits packets to the estimator. Aiming at minimizing the remote estimation error and the cost of the sensors transmission energy while maximizing the eavesdropper state estimation error, we theoretically prove that there exist some structural properties for the optimal transmission schedule for both the known and the unknown eavesdropper's estimation errors. Numerical simulation results are provided to validate the theoretical analysis. Le Wang 0003, Xianghui Cao, Heng Zhang 0001, Changyin Sun 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Giant Panda IdentificationabstractThe lack of automatic tools to identify giant panda makes it hard to keep track of and manage giant pandas in wildlife conservation missions. In this paper, we introduce a new Giant Panda Identification (GPID) task, which aims to identify each individual panda based on an image. Though related to the human re-identification and animal classification problem, GPID is extraordinarily challenging due to subtle visual differences between pandas and cluttered global information. In this paper, we propose a new benchmark dataset iPanda-50 for GPID. The iPanda-50 consists of 6, 874 images from 50 giant panda individuals, and is collected from panda streaming videos. We also introduce a new Feature-Fusion Network with Patch Detector (FFN-PD) for GPID. The proposed FFN-PD exploits the patch detector to detect discriminative local patches without using any part annotations or extra location sub-networks, and builds a hierarchical representation by fusing both global and local features to enhance the inter-layer patch feature interactions. Specifically, an attentional cross-channel pooling is embedded in the proposed FFN-PD to improve the identify-specific patch detectors. Experiments performed on the iPanda-50 datasets demonstrate the proposed FFN-PD significantly outperforms competing methods. Besides, experiments on other fine-grained recognition datasets (i.e., CUB-200-2011, Stanford Cars, and FGVC-Aircraft) demonstrate that the proposed FFN-PD outperforms existing state-of-the-art methods. Le Wang 0003, Rizhi Ding, Yuanhao Zhai 0001, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Hierarchical and Interactive Refinement Network for Edge-Preserving Salient Object DetectionabstractSalient object detection has undergone a very rapid development with the blooming of Deep Neural Network (DNN), which is usually taken as an important preprocessing procedure in various computer vision tasks. However, the down-sampling operations, such as pooling and striding, always make the final predictions blurred at edges, which has seriously degenerated the performance of salient object detection. In this paper, we propose a simple yet effective approach, i.e., Hierarchical and Interactive Refinement Network (HIRN), to preserve the edge structures in detecting salient objects. In particular, a novel multi-stage and dual-path network structure is designed to estimate the salient edges and regions from the low-level and high-level feature maps, respectively. As a result, the predicted regions will become more accurate by enhancing the weak responses at edges, while the predicted edges will become more semantic by suppressing the false positives in background. Once the salient maps of edges and regions are obtained at the output layers, a novel edge-guided inference algorithm is introduced to further filter the resulting regions along the predicted edges. Extensive experiments on several benchmark datasets have been conducted, in which the results show that our method significantly outperforms a variety of state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Le Wang 0003, Jimuyang Zhang, Fei Wang 0037, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Object Cosegmentation in Noisy Videos With Multilevel HypergraphabstractWith the target of simultaneously segmenting semantically related videos to identify the common objects, video object cosegmentation has attracted the attention of researchers in recent years. Existing methods are primarily based on pair-wise relations between adjacent pixels and regions, which are susceptible to performance degradation from object entries/exists or occlusions. Specifically, we refer these video frames without the common objects present as the “empty” frames. In this paper, we propose a multilevel hypergraph-based full Video object CoSegmentation (VCS) method, which incorporates high-level semantics and low-level appearance/motion/saliency to construct the hyperedge among multiple spatially and temporally adjacent regions. Specifically, the high-level semantic model fuses multiple object proposals from each frame instead of relying on a single object proposal per frame. A hypergraph cut is subsequently utilized to calculate the object cosegmentation. Experiments on four video object segmentation/cosegmentation datasets against state-of-the-art methods with both objective and subjective results manifest the effectiveness of the proposed VCS method, including the SegTrack and VCoSeg datasets without “empty” frames, the XJTU-Stevens dataset with 3.7% “empty” frames, and the Noisy-ViCoSeg dataset proposed together with our method with 30.3% “empty” frames. Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Ladder Loss for Coherent Visual-Semantic EmbeddingabstractFor visual-semantic embedding, the existing methods normally treat the relevance between queries and candidates in a bipolar way – relevant or irrelevant, and all “irrelevant” candidates are uniformly pushed away from the query by an equal margin in the embedding space, regardless of their various proximity to the query. This practice disregards relatively discriminative information and could lead to suboptimal ranking in the retrieval results and poorer user experience, especially in the long-tail query scenario where a matching candidate may not necessarily exist. In this paper, we introduce a continuous variable to model the relevance degree between queries and multiple candidates, and propose to learn a coherent embedding space, where candidates with higher relevance degrees are mapped closer to the query than those with lower relevance degrees. In particular, the new ladder loss is proposed by extending the triplet loss inequality to a more general inequality chain, which implements variable push-away margins according to respective relevance degrees. In addition, a proper Coherent Score metric is proposed to better measure the ranking results including those “irrelevant” candidates. Extensive experiments on multiple datasets validate the efficacy of our proposed method, which achieves significant improvement over existing state-of-the-art methods. Zhenxing Niu, Le Wang 0003, Zhanning Gao, Qilin Zhang 0004, Gang Hua 0001 |
AAAI | 3 |
| 2020 | Multi-label X-Ray Imagery Classification via Bottom-Up Attention and Meta Fusion
Benyi Hu, Chi Zhang 0020, Le Wang 0003, Qilin Zhang 0004, Yuehu Liu |
ACCV (6) | 3 |
| 2020 | Loss Functions for Person Image Generation
Haoyue Shi 0002, Le Wang 0003, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
BMVC | 2 |
| 2020 | High-order Graph Convolutional Networks for 3D Human Pose Estimation
Zhiming Zou, Kenkun Liu, Le Wang 0003, Wei Tang 0016 |
BMVC | 3 |
| 2020 | A Comprehensive Study of Weight Sharing in Graph Networks for 3D Human Pose Estimation
Kenkun Liu, Rongqi Ding, Zhiming Zou, Le Wang 0003, Wei Tang 0016 |
ECCV (10) | 4 |
| 2020 | Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization
Yuanhao Zhai 0001, Le Wang 0003, Wei Tang 0016, Qilin Zhang 0004, Junsong Yuan 0001, Gang Hua 0001 |
ECCV (6) | 2 |
| 2020 | Adversarial Ranking Attack and Defense
Zhenxing Niu, Le Wang 0003, Qilin Zhang 0004, Gang Hua 0001 |
ECCV (14) | 3 |
| 2020 | Fine-Grained Giant Panda IdentificationabstractThe image-based fine-grained identification of individual giant pandas (Ailuropoda melanoleuca) is an emerging technology, and it is extraordinarily challenging due to the extremely subtle visual differences between individual giant pandas and limited annotated training data. To address these challenges, we propose the Feature-Fusion Convolutional Neural Network with Patch Detector (FFCNN-PD) algorithm, which exploits the discriminative local patches and builds a hierarchical representation generated by fusing both global and local features. Specifically, an attentional cross-channel pooling is embedded in the FFCNN-PD to improve the class- specific patch detectors. In addition, we propose a new giant panda identification dataset (iPanda-30) to establish a benchmark. Experiments on the proposed iPanda-30 dataset and other fine-grained recognition datasets demonstrate the effectiveness of the FFCNN-PD algorithm against the existing state-of-the-arts. Rizhi Ding, Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICASSP | 2 |
| 2020 | Caption Generation from Road Images for Traffic Scene ConstructionabstractIn this paper, an image captioning network is proposed for traffic scene modeling, which incorporates element attention into the encoder-decoder mechanism to generate more reasonable scene captions. Firstly, the traffic scene elements are detected and segmented according to their clustered locations. Then, the image captioning network is applied to generate the corresponding caption of each subregion. The static and dynamic traffic elements are appropriately organized to construct a 3D corridor scene model. The semantic relationships between the traffic elements are specified according to the captions. The constructed 3D scene model can be utilized for the offline test of unmanned vehicles. The evaluations and comparisons based on the TSD-max and COCO datasets prove the effectiveness of the proposed framework. Yaochen Li, Le Wang 0003, Yuehu Liu |
IV | 4 |
| 2020 | Worst Perception Scenario Search for Autonomous DrivingabstractAchieving excellent generalization on perceiving real traffic scenarios with diversity is the long-term goal for building robust autonomous driving systems. In this paper, we propose to discover potential shortness of certain perception module by analyzing its worst-scenario performance. However, with the benchmark datasets growing huge and tremendous, exhaustive searching for the worst perception scenario (WPS) seems to be time consuming and unnecessary. To address, we present an automatic searching scheme empowered by reinforcement learning. In this case, worst scenario mining is formulated as a discrete search problem. A single layer recurrent neural network with LSTM neurons is employed to predict WPS according to the searching reward, which is optimized by a vanilla policy gradient method. Moreover, to deal with the imbalanced distribution of real traffic scenarios, a KNN-like retrieval is utilized for searching the closest scenario samples. Effective yet efficient, the proposed method has been validated by finding the most challenging scenarios for various vehicle detectors on KITTI, BDD100k and our own benchmark set EVB. Further experiments reveal that detection networks with structural similarity share the similar WPS. Liheng Xu, Chi Zhang 0020, Yuehu Liu, Le Wang 0003, Li Li 0013 |
IV | 4 |
| 2020 | Meta Corrupted Pixels Mining for Medical Image Segmentation
Sanping Zhou, Chaowei Fang, Le Wang 0003, Jinjun Wang |
MICCAI (1) | 4 |
| 2020 | Action Co-localization in an Untrimmed Video by Graph Neural Networks
Changbo Zhai, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
MMM (1) | 2 |
| 2020 | Hierarchical U-Shape Attention Network for Salient Object DetectionabstractSalient object detection aims at locating the most conspicuous objects in natural images, which usually acts as a very important pre-processing procedure in many computer vision tasks. In this paper, we propose a simple yet effective Hierarchical U-shape Attention Network (HUAN) to learn a robust mapping function for salient object detection. Firstly, a novel attention mechanism is formulated to improve the well-known U-shape network [1], in which the memory consumption can be extensively reduced and the mask quality can be significantly improved by the resulting U-shape Attention Network (UAN). Secondly, a novel hierarchical structure is constructed to well bridge the low-level and high-level feature representations between different UANs, in which both the intra-network and inter-network connections are considered to explore the salient patterns from a local to global view. Thirdly, a novel Mask Fusion Network (MFN) is designed to fuse the intermediate prediction results, so as to generate a salient mask which is in higher-quality than any of those inputs. Our HUAN can be trained together with any backbone network in an end-to-end manner, and high-quality masks can be finally learned to represent the salient objects. Extensive experimental results on several benchmark datasets show that our method significantly outperforms most of the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Jimuyang Zhang, Le Wang 0003, Shaoyi Du, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Video Imprint Segmentation for Temporal Action Detection in Untrimmed VideosabstractWe propose a temporal action detection by spatial segmentation framework, which simultaneously categorize actions and temporally localize action instances in untrimmed videos. The core idea is the conversion of temporal detection task into a spatial semantic segmentation task. Firstly, the video imprint representation is employed to capture the spatial/temporal interdependences within/among frames and represent them as spatial proximity in a feature space. Subsequently, the obtained imprint representation is spatially segmented by a fully convolutional network. With such segmentation labels projected back to the video space, both temporal action boundary localization and per-frame spatial annotation can be obtained simultaneously. The proposed framework is robust to variable lengths of untrimmed videos, due to the underlying fixed-size imprint representations. The efficacy of the framework is validated in two public action detection datasets. Zhanning Gao, Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 2 |
| 2019 | Object Affordances Graph Network for Action Recognition
Haoliang Tan, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Nanning Zheng 0001, Gang Hua 0001 |
BMVC | 2 |
| 2019 | Compressing Unknown Images With Product Quantizer for Efficient Zero-Shot ClassificationabstractFor Zero-Shot Learning (ZSL), the Nearest Neighbor (NN) search is generally conducted for classification, which may cause unacceptable computational complexity for large-scale datasets. To compress zero-shot classes by the trained quantizer for efficient search, it tends to induce large quantization error because distributions between seen and unseen classes are different. However, as semantic attributes of classes are available in ZSL, both seen and unseen classes have the same distribution for one specific property, e.g., animals have or not have spots. Based on this intuition, a Product Quantization Zero-Shot Learning (PQZSL) method is proposed to learn embeddings as well as quantizers to compress visual features into compact codes for Approximate NN (ANN) search. Particularly, visual features are projected into an orthogonal semantic space, and then the Product Quantization (PQ) is utilized to quantize individual properties. Experimental results on five benchmark datasets demonstrate that unseen classes are represented by the Cartesian product of quantized properties with little quantization error. As classes in orthogonal common space are more discriminative, the classification based on PQZSL achieves state-of-the-art performance in Generalized Zero-Shot Learning (GZSL) task, meanwhile, the speed of ANN search is 10-100 times higher than traditional NN search. Jin Li 0011, Xuguang Lan, Yang Liu 0069, Le Wang 0003, Nanning Zheng 0001 |
CVPR | 4 |
| 2019 | Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksabstractWeakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotations in training data, WS-TAL could possibly exploit automatically retrieved video tags as video-level labels. However, such coarse video-level supervision inevitably incurs confusions, especially in untrimmed videos containing multiple action instances. To address this challenge, we propose the Contrast-based Localization EvaluAtioN Network (CleanNet) with our new action proposal evaluator, which provides pseudo-supervision by leveraging the temporal contrast in snippet-level action classification predictions. Essentially, the new action proposal evaluator enforces an additional temporal contrast constraint so that high-evaluation-score action proposals are more likely to coincide with true action instances. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Experiments on THUMOS14 and ActivityNet datasets validate the efficacy of CleanNet against existing state-ofthe- art WS-TAL algorithms. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 2 |
| 2019 | Action Coherence Network for Weakly Supervised Temporal Action LocalizationabstractMost prominent temporal action localization methods are of the fully-supervised type, which rely heavily on frame-level labels, which could be prohibitively expensive to annotate. Thanks to recent developments on the Weakly-supervised Temporal Action Localization (W-TAL), this alternative paradigm requires only video-level labels in training, alleviating such annotation efforts. Specifically, we present Action Coherence Network (ACN) for W-TAL, which features a new coherence loss that better supervises action boundary learning and facilitate proposal regression. In addition, a purpose-built fusion module is proposed for localization inference based on features extracted by two streams of convolutional neural network. Overall, the proposed ACN achieves state-of-the-art W-TAL performance on two challenging datasets (THU-MOS14 and ActivityNet1.2, particularly ACN attains mAP of 24.2% on THUMOS14 under IoU threshold 0.5), which is approaching some recent fully-supervised TAL methods. Yuanhao Zhai 0001, Le Wang 0003, Ziyi Liu 0001, Qilin Zhang 0004, Gang Hua 0001, Nanning Zheng 0001 |
ICIP | 2 |
| 2019 | Jointly Detecting and Retrieving Vehicles from Road Image Sequences based on CNNabstractIn this paper, a CNN-based vehicle detection and retrieval framework is proposed for the intelligent transportation system. Firstly, the vehicle target is detected from the traffic scene. The proposed object detection method uses a fully convolutional neural network (CNN) based on SqueezeNet, which has the characteristics of real-time, high accuracy and has small model size. Secondly, an intra-class image retrieval method is presented to search vehicles which are similar to the target vehicle in the dataset. The image retrieval results can be used for traffic scenes simulation and modeling. The experiments and comparisons prove the effectiveness of our framework. Yaochen Li, Yuehu Liu, Shanmin Pang, Le Wang 0003, Huihui Huo |
IV | 5 |
| 2019 | Video ImprintabstractA new unified video analytics framework (ER3) is proposed for complex event retrieval, recognition and recounting, based on the proposed video imprint representation, which exploits temporal correlations among image features across video frames. With the video imprint representation, it is convenient to reverse map back to both temporal and spatial locations in video frames, allowing for both key frame identification and key areas localization within each frame. In the proposed framework, a dedicated feature alignment module is incorporated for redundancy removal across frames to produce the tensor representation, i.e., the video imprint. Subsequently, the video imprint is individually fed into both a reasoning network and a feature aggregation module, for event recognition/recounting and event retrieval tasks, respectively. Thanks to its attention mechanism inspired by the memory networks used in language modeling, the proposed reasoning network is capable of simultaneous event category recognition and localization of the key pieces of evidence for event recounting. In addition, the latent structure in our reasoning network highlights the areas of the video imprint, which can be directly used for event recounting. With the event retrieval task, the compact video representation aggregated from the video imprint contributes to better retrieval results than existing state-of-the-art methods. Zhanning Gao, Le Wang 0003, Nebojsa Jojic, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Joint Spatio-Temporal Action Localization in Untrimmed Videos with Per-Frame SegmentationabstractInspired by the recent spatio-temporal action localization efforts with tubelets (sequences of bounding boxes), we present a new spatio-temporal action detector Segment-tube, which consists of sequences of per-frame segmentation masks. The proposed Segment-tube detector can temporally pinpoint the starting/ending frame of each action class in the presence of preceding/subsequent interference actions in untrimmed videos. Simultaneously, the Segment-tube detector produces per-frame segmentation masks instead of bounding boxes, offering superior spatial accuracy to tubelets. This is achieved by alternating iterative optimization between temporal action localization and spatial action segmentation. Experimental results on multiple datasets validate the efficacy of the proposed detector. Xuhuan Duan, Le Wang 0003, Changbo Zhai, Nanning Zheng 0001, Qilin Zhang 0004, Zhenxing Niu, Gang Hua 0001 |
ICIP | 2 |
| 2018 | Video Object Co-Segmentation from Noisy Videos by a Multi-Level Hypergraph ModelabstractDefined as simultaneously segmenting a set of related videos to identify the common objects, video co-segmentation has attracted the attention of researchers in recent years. Existing methods are primarily based on pair-wise relations between adjacent pixels/regions, which are susceptible to performance degradation from “empty” video frames (e.g., due to transient/intermittent common objects). In this paper, a new multilevel hypergraph based method, termed the full Video object Co-Segmentation method (VCS), is proposed, which incorporates both a high-level semantics object model and a low-level appearance/motion/saliency object model to construct the hyperedge among multiple spatially and temporally adjacent regions. Specifically, the high-level semantic model fuses multiple object proposals from each frame instead of relying on a single object proposal per frame. A hypergraph cut is subsequently utilized to calculate the object co-segmentation. Experiments on three datasets demonstrate the efficacy of the proposed VCS method. Le Wang 0003, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
ICIP | 2 |
| 2018 | A Graded Offline Evaluation Framework for Intelligent Vehicle's Cognitive AbilityabstractCognitive ability evaluation in intelligent vehicles is conventionally evaluated by classical autonomous driving dataset, which lacks comprehensive annotations of driving difficulty. Realistically, different driving conditions require vast different level of cognitive ability, e.g., driving in highly congested traffic is much more challenging than driving on limited access highway; driving in a blizzard/hurricane requires much more robust environmental cognition abilities than driving under ordinary conditions. Different datasets contain different proportions of various driving conditions, rendering intelligent vehicle evaluation susceptible to dataset variations. To overcome such limitations, we propose to first benchmark the driving difficulty with the proposed “Cascaded Tanks Model” and obtain a fine-grained per-segment difficulty rating based on our proposed Semantic Descriptor. With the proposed Graded Offline Evaluation (GOE) framework, it is demonstrated that offline validation of the cognitive abilities in Intelligent Vehicles (IV) is more consistent regardless of dataset choice. Chi Zhang 0020, Yuehu Liu, Qilin Zhang 0004, Le Wang 0003 |
Intelligent Vehicles Symposium | 4 |
| 2018 | Joint Video Object Discovery and Segmentation by Coupled Dynamic Markov NetworksabstractIt is a challenging task to extract segmentation mask of a target from a single noisy video, which involves object discovery coupled with segmentation. To solve this challenge, we present a method to jointly discover and segment an object from a noisy video, where the target disappears intermittently throughout the video. Previous methods either only fulfill video object discovery, or video object segmentation presuming the existence of the object in each frame. We argue that jointly conducting the two tasks in a unified way will be beneficial. In other words, video object discovery and video object segmentation tasks can facilitate each other. To validate this hypothesis, we propose a principled probabilistic model, where two dynamic Markov networks are coupled-one for discovery and the other for segmentation. When conducting the Bayesian inference on this model using belief propagation, the bi-directional message passing reveals a clear collaboration between these two inference tasks. We validated our proposed method in five data sets. The first three video data sets, i.e., the SegTrack data set, the YouTube-objects data set, and the Davis data set, are not noisy, where all video frames contain the objects. The two noisy data sets, i.e., the XJTU-Stevens data set, and the Noisy-ViDiSeg data set, newly introduced in this paper, both have many frames that do not contain the objects. When compared with state of the art, it is shown that although our method produces inferior results on video data sets without noisy frames, we are able to obtain better results on video data sets with noisy frames. Ziyi Liu 0001, Le Wang 0003, Gang Hua 0001, Qilin Zhang 0004, Zhenxing Niu, Ying Wu 0001, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Knowledge-Based Topic Model for Unsupervised Object Discovery and LocalizationabstractUnsupervised object discovery and localization is to discover some dominant object classes and localize all of object instances from a given image collection without any supervision. Previous work has attempted to tackle this problem with vanilla topic models, such as latent Dirichlet allocation (LDA). However, in those methods no prior knowledge for the given image collection is exploited to facilitate object discovery. On the other hand, the topic models used in those methods suffer from the topic coherence issue-some inferred topics do not have clear meaning, which limits the final performance of object discovery. In this paper, prior knowledge in terms of the so-called must-links are exploited from Web images on the Internet. Furthermore, a novel knowledge-based topic model, called LDA with mixture of Dirichlet trees, is proposed to incorporate the must-links into topic modeling for object discovery. In particular, to better deal with the polysemy phenomenon of visual words, the must-link is re-defined as that one must-link only constrains one or some topic(s) instead of all topics, which leads to significantly improved topic coherence. Moreover, the must-links are built and grouped with respect to specific object classes, thus the must-links in our approach are semantic-specific, which allows to more efficiently exploit discriminative prior knowledge from Web images. Extensive experiments validated the efficiency of our proposed approach on several data sets. It is shown that our method significantly improves topic coherence and outperforms the unsupervised methods for object discovery and localization. In addition, compared with discriminative methods, the naturally existing object classes in the given image collection can be subtly discovered, which makes our approach well suited for realistic applications of unsupervised object discovery. Zhenxing Niu, Gang Hua 0001, Le Wang 0003, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | ER3: A Unified Framework for Event Retrieval, Recognition and RecountingabstractWe develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alignment procedure identifies and removes the feature redundancies across frames and outputs an intermediate tensor representation we call video imprint. The video imprint is then fed into a reasoning network, whose attention mechanism parallels that of memory networks used in language modeling. The reasoning network simultaneously recognizes the event category and locates the key pieces of evidence for event recounting. In event retrieval tasks, we show that the compact video representation aggregated from the video imprint achieves significantly better retrieval accuracy compared with existing methods. We also set new state of the art results in event recognition tasks with an additional benefit: The latent structure in our reasoning network highlights the areas of the video imprint and can be directly used for event recounting. As video imprint maps back to locations in the video frames, the network allows not only the identification of key frames but also specific areas inside each frame which are most influential to the decision process. Zhanning Gao, Gang Hua 0001, Dongqing Zhang, Nebojsa Jojic, Le Wang 0003, Jianru Xue, Nanning Zheng 0001 |
CVPR | 5 |
| 2017 | Hierarchical Multimodal LSTM for Dense Visual-Semantic EmbeddingabstractWe address the problem of dense visual-semantic embedding that maps not only full sentences and whole images but also phrases within sentences and salient regions within images into a multimodal embedding space. Such dense embeddings, when applied to the task of image captioning, enable us to produce several region-oriented and detailed phrases rather than just an overview sentence to describe an image. Specifically, we present a hierarchical structured recurrent neural network (RNN), namely Hierarchical Multimodal LSTM (HM-LSTM). Compared with chain structured RNN, our proposed model exploits the hierarchical relations between sentences and phrases, and between whole images and image regions, to jointly establish their representations. Without the need of any supervised labels, our proposed model automatically learns the fine-grained correspondences between phrases and image regions towards the dense embedding. Extensive experiments on several datasets validate the efficacy of our method, which compares favorably with the state-of-the-art methods. Zhenxing Niu, Le Wang 0003, Xinbo Gao 0001, Gang Hua 0001 |
ICCV | 3 |
| 2017 | Video Object Discovery and Co-Segmentation with Extremely Weak SupervisionabstractWe present a spatio-temporal energy minimization formulation for simultaneous video object discovery and co-segmentation across multiple videos containing irrelevant frames. Our approach overcomes a limitation that most existing video co-segmentation methods possess, i.e., they perform poorly when dealing with practical videos in which the target objects are not present in many frames. Our formulation incorporates a spatio-temporal auto-context model, which is combined with appearance modeling for superpixel labeling. The superpixel-level labels are propagated to the frame level through a multiple instance boosting algorithm with spatial reasoning, based on which frames containing the target object are identified. Our method only needs to be bootstrapped with the frame-level labels for a few video frames (e.g., usually 1 to 3) to indicate if they contain the target objects or not. Extensive experiments on four datasets validate the efficacy of our proposed method: 1) object segmentation from a single video on the SegTrack dataset, 2) object co-segmentation from multiple videos on a video co-segmentation dataset, and 3) joint object discovery and co-segmentation from multiple videos containing irrelevant frames on the MOViCS dataset and XJTU-Stevens, a new dataset that we introduce in this paper. The proposed method compares favorably with the state-of-the-art in all of these experiments. Le Wang 0003, Gang Hua 0001, Rahul Sukthankar, Jianru Xue, Zhenxing Niu, Nanning Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Ordinal Regression with Multiple Output CNN for Age EstimationabstractTo address the non-stationary property of aging patterns, age estimation can be cast as an ordinal regression problem. However, the processes of extracting features and learning a regression model are often separated and optimized independently in previous work. In this paper, we propose an End-to-End learning approach to address ordinal regression problems using deep Convolutional Neural Network, which could simultaneously conduct feature learning and regression modeling. In particular, an ordinal regression problem is transformed into a series of binary classification sub-problems. And we propose a multiple output CNN learning algorithm to collectively solve these classification sub-problems, so that the correlation between these tasks could be explored. In addition, we publish an Asian Face Age Dataset (AFAD) containing more than 160K facial images with precise age ground-truths, which is the largest public age dataset to date. To the best of our knowledge, this is the first work to address ordinal regression problems by using CNN, and achieves the state-of-the-art performance on both the MORPH and AFAD datasets. Zhenxing Niu, Le Wang 0003, Xinbo Gao 0001, Gang Hua 0001 |
CVPR | 3 |
| 2014 | Video Object Discovery and Co-segmentation with Extremely Weak Supervision
Le Wang 0003, Gang Hua 0001, Rahul Sukthankar, Jianru Xue, Nanning Zheng 0001 |
ECCV (4) | 1 |
| 2014 | Joint Segmentation and Recognition of Categorized Objects From Noisy Web Image CollectionabstractThe segmentation of categorized objects addresses the problem of joint segmentation of a single category of object across a collection of images, where categorized objects are referred to objects in the same category. Most existing methods of segmentation of categorized objects made the assumption that all images in the given image collection contain the target object. In other words, the given image collection is noise free. Therefore, they may not work well when there are some noisy images which are not in the same category, such as those image collections gathered by a text query from modern image search engines. To overcome this limitation, we propose a method for automatic segmentation and recognition of categorized objects from noisy Web image collections. This is achieved by cotraining an automatic object segmentation algorithm that operates directly on a collection of images, and an object category recognition algorithm that identifies which images contain the target object. The object segmentation algorithm is trained on a subset of images from the given image collection which are recognized to contain the target object with high confidence, while training the object category recognition model is guided by the intermediate segmentation results obtained from the object segmentation algorithm. This way, our co-training algorithm automatically identifies the set of true positives in the noisy Web image collection, and simultaneously extracts the target objects from all the identified images. Extensive experiments validated the efficacy of our proposed approach on four datasets: 1) the Weizmann horse dataset, 2) the MSRC object category dataset, 3) the iCoseg dataset, and 4) a new 30-categories dataset including 15,634 Web images with both hand-annotated category labels and ground truth segmentation labels. It is shown that our method compares favorably with the state-of-the-art, and has the ability to deal with noisy image collections. Le Wang 0003, Gang Hua 0001, Jianru Xue, Zhanning Gao, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Automatic salient object extraction with contextual cue and its applications to recognition and alpha matting
Jianru Xue, Le Wang 0003, Nanning Zheng 0001, Gang Hua 0001 |
Pattern Recognit. | 2 |
| 2012 | Large-Scale Bundle Adjustment by Parameter Vector Partition
Shanmin Pang, Jianru Xue, Le Wang 0003, Nanning Zheng 0001 |
ACCV (4) | 3 |
| 2012 | Concurrent segmentation of categorized objects from an image collection
Le Wang 0003, Jianru Xue, Nanning Zheng 0001, Gang Hua 0001 |
ICPR | 1 |
| 2011 | Automatic salient object extraction with contextual cueabstractWe present a method for automatically extracting salient object from a single image, which is cast in an energy minimization framework. Unlike most previous methods that only leverage appearance cues, we employ an auto-context cue as a complementary data term. Benefitting from a generic saliency model for bootstrapping, the segmentation of the salient object and the learning of the auto-context model are iteratively performed without any user intervention. Upon convergence, we obtain not only a clear separation of the salient object, but also an auto-context classifier which can be used to recognize the same type of object in other images. Our experiments on four benchmarks demonstrated the efficacy of the added contextual cue. It is shown that our method compares favorably with the state-of-the-art, some of which even embraced user interactions. Le Wang 0003, Jianru Xue, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 1 |