VLDB 2026 Research / reviewers in the wild / expert
Sanping Zhou
dblp:179/0508
· DBLP profile ↗
121ranked-venue papers
11as first author
102since 2021 · last 2026
0000-0002-2946-6395ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 87 · 7 first-author · 72 since 2021Graphics, computer vision, multimedia, augmented reality and games · 73 · 6 first-author · 64 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TSPO: Temporal Sampling Policy Optimization for Long-form Video Language UnderstandingabstractMultimodal Large Language Models (MLLMs) have demonstrated significant progress in vision-language tasks, yet they still face challenges when processing long-duration video inputs. The limitation arises from MLLMs' context limit and training costs, necessitating sparse frame sampling before feeding videos into MLLMs. However, building a trainable sampling method remains challenging due to the unsupervised and non-differentiable nature of sparse frame sampling in Video-MLLMs. To address these problems, we propose Temporal Sampling Policy Optimization (**TSPO**), advancing MLLMs' long-form video-language understanding via reinforcement learning. Specifically, we first propose a trainable event-aware temporal agent, which captures event-query correlation for performing probabilistic keyframe selection. Then, we propose the TSPO reinforcement learning paradigm, which models keyframe selection and language generation as a joint decision-making process, enabling end-to-end group relative optimization for the temporal sampling policy. Furthermore, we propose a dual-style long video training data construction pipeline, balancing comprehensive temporal understanding and key segment localization. Finally, we incorporate rule-based answering accuracy and temporal locating reward mechanisms to optimize the temporal sampling policy. Comprehensive experiments show that our TSPO achieves state-of-the-art performance across multiple long video understanding benchmarks, and shows transferable ability across different cutting-edge Video-MLLMs. Canhui Tang, Zifan Han, Sanping Zhou, Xuchong Zhang, Jinglin Xu |
AAAI | 4 |
| 2026 | Sparse Trajectory PredictionabstractPedestrian trajectory prediction is crucial for ensuring safe decision-making in intelligent robotic systems. While this task demands real-time performance, previous works have primarily focused on improving prediction accuracy, often neglecting efficiency. Dense predictions with time-consuming post-clustering steps and global interactions with quadratic computational complexity result in a trade-off between accuracy and speed. In this paper, we propose a novel Sparse Trajectory Prediction (STP) model that aims to achieve both high accuracy and real-time speed by following an efficient principle: leveraging sparse structures to achieve global effects. STP instantiates this principle within a transformer-style encoder-decoder framework. In the encoder, STP introduces irregular interaction, which builds sparse interactions with dynamic interactive positions, reducing computational complexity to linearithmic/linear while maintaining global interaction. In the decoder, STP applies an early-sparsity strategy to generate sparse motion modes that represent global motion behaviors. These modes are shared across all predictions, eliminating redundant computations. By harnessing the expressive power of transformers, STP maps these sparse motion modes into multimodal future trajectories, significantly improving prediction speed while ensuring accuracy. Experimental results on four commonly used datasets demonstrate that STP maximizes both accuracy and prediction speed, achieving state-of-the-art performance and significantly improving prediction speed by about $100 \times$100× - $150 \times$150× to satisfy the real-time demand. Liushuai Shi, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | MonoA2: Adaptive depth with augmented head for monocular 3D object detection
Jinpeng Dong, Sanping Zhou, Jingjing Jiang, Weiliang Zuo, Shi-tao Chen, Nanning Zheng 0001 |
Pattern Recognit. | 2 |
| 2026 | DictCR-former: Content-aware dictionary transformer for cloud removal
Wenli Huang 0004, Yang Wu 0001, Sanping Zhou, Xiaomeng Xin, Xiaobo Jia, Ye Deng 0005 |
Pattern Recognit. | 3 |
| 2026 | Frequency-guided generalizable representation learning for cross-domain few-shot learning
Siqi Hui, Sanping Zhou, Ye Deng 0005, Wenli Huang 0004, Yang Wu 0001, Jinjun Wang |
Pattern Recognit. | 2 |
| 2026 | Query-enhanced motion transformer with dilated static query and bridged dynamic query
Miao Kang, Liushuai Shi, Ke Ye, Sanping Zhou, Nanning Zheng 0001 |
Pattern Recognit. | 4 |
| 2026 | PR-DETR: Injecting position and relation prior for dense video captioning
Sanping Zhou, Le Wang 0003 |
Pattern Recognit. | 2 |
| 2026 | Time-Unified Diffusion Policy with action discrimination for robotic manipulation
Ye Niu, Sanping Zhou |
Pattern Recognit. | 2 |
| 2026 | Action hints: Semantic typicality and context uniqueness for generalizable skeleton-based video anomaly detection
Canhui Tang, Sanping Zhou, Haoyue Shi 0002, Le Wang 0003 |
Pattern Recognit. | 2 |
| 2026 | RAMP: Iterative Refinement and Adaptive Multi-granularity Perception for embodied dialog localization
Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
Pattern Recognit. | 3 |
| 2026 | Pedestrian Trajectory Prediction via Hierarchical Dynamics DecompositionabstractPredicting human future trajectories is crucial for various intelligent systems and applications. Previous approaches typically adopt a direct prediction strategy, which decodes trajectory features directly into future coordinates. However, they overlook different hierarchical high-order velocities, which have stronger representational abilities in dynamics. In this paper, we introduce HDDNet, a novel trajectory prediction framework that follows dynamical principles and employs a hierarchical dynamics decomposition strategy. Specifically, HDDNet models future trajectories by progressively transferring trajectory coordinates into velocity, acceleration, and jerk, up to the highest-order velocity, which sequentially represent a broader receptive field and a more compact representation of motion dynamics. Furthermore, we design a hierarchical dynamics decomposition decoder with a corresponding dynamics loss, which predicts future trajectories by sequentially refining human motions from the highest-order velocity down to the final coordinates. Compared to the traditional direct prediction strategy, our approach makes better use of dynamic information at different levels. Extensive experiments and ablation studies on the ETH-UCY, SDD and GigaTraj datasets demonstrate that our method outperforms existing state-of-the-art approaches. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | RSRNav: Reasoning Spatial Relationship for Image-Goal NavigationabstractRecent image-goal navigation (ImageNav) methods learn a perception-action policy by separately capturing semantic features of the goal and egocentric images, then passing them to a policy network. However, challenges remain: (1) Semantic features often fail to provide accurate directional information, leading to superfluous actions, and (2) performance drops significantly when viewpoint inconsistencies arise between training and application. To address these challenges, we propose RSRNav, a simple yet effective method that reasons spatial relationships between the goal and current observations as navigation guidance. Specifically, we model the spatial relationship by constructing correlations between the goal and current observations, which are then passed to the policy network for action prediction. These correlations are progressively refined using fine-grained cross-correlation and direction-aware correlation for more precise navigation. Extensive evaluation of RSRNav on three benchmark datasets demonstrates superior navigation performance, particularly in the "user-matched goal" setting, highlighting its potential for real-world applications. Code: https://github.com/ qinzheng2000/RSRNav.git. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Advancing Pre-Trained Teacher: Towards Robust Feature Discrepancy for Anomaly DetectionabstractWith the wide application of knowledge distillation between an ImageNet pre-trained teacher model and a learnable student model, unsupervised anomaly detection has witnessed a significant achievement in the past few years. The success of this framework mainly relies on how to keep the feature discrepancy between the teacher and student model, in which it has two underlying sub-assumptions: (1) The teacher model can represent two separable distributions for the normal and abnormal patterns, while (2) the student model can only reconstruct the normal distribution. However, it still remains a challenging issue to maintain these ideal assumptions in practice. In this paper, we propose a simple yet effective two-stage industrial anomaly detection framework, termed AAND, which sequentially performs Anomaly Amplification and Normality Distillation to enhance the two assumptions. In the first anomaly amplification stage, we propose a novel Residual Anomaly Amplification (RAA) module to advance the pre-trained teacher encoder with synthetic anomalies. It generates adaptive residuals to amplify anomalies while maintaining the feature integrity of pre-trained model. It mainly comprises a Matching-guided Residual Gate and an Attribute-scaling Residual Generator, which can determine the residuals' proportion and characteristic, respectively. In the second normality distillation stage, we further employ a reverse distillation paradigm to train a student decoder, in which a novel Hard Knowledge Distillation (HKD) loss is built to better facilitate the reconstruction of normal patterns. Comprehensive experiments on the MvTecAD, VisA, and MvTec3D-RGB datasets show that our method achieves state-of-the-art performance. Our code is available at https://github.com/Hui-design/AAND. Canhui Tang, Sanping Zhou, Yonghao Dong, Le Wang 0003 |
IEEE Trans. Image Process. | 2 |
| 2026 | MoDe-Track: Robust Multi-Object Tracking With Motion Decoupling in UAV VideosabstractMulti-Object Tracking (MOT) in Unmanned Aerial Vehicle (UAV) scenarios is characterized by frequent and abrupt camera motion, which presents two unique challenges: nonlinear motion and appearance degradation. Traditional motion models, designed for smooth and consistent motion, struggle to capture the complex background motion patterns caused by UAV movement; while appearance-based methods are easily disrupted by occlusion and blur, leading to unreliable associations. Even though dense optical flow is widely utilized to model complex motion patterns, the entanglement of background and object motion often introduces interference, limiting its effectiveness. To this end, we propose MoDe-Track, a unified framework that explicitly decouples scene motion into background and object components, and serves as an elegant integration of three robust components. Specifically, the Scene Motion Decomposition (SMD) module decouples the motion into the background and object components based on robust principal component analysis, serving as the foundation for motion compensation and feature propagation. Afterwards, the Background Motion Compensation (BMC) uses the decomposed background flow to estimate and compensate for camera motion, mitigating the effects of nonlinear motion. Finally, the Foreground-guided Feature Propagation (FFP) module uses the decoupled object flow to guide feature propagation across frames, achieving temporal consistency and enhancing robustness against occlusion and motion blur. Extensive experimental results on two benchmarks, VisDrone2019 and UAVDT, demonstrate that MoDe-Track consistently outperforms current multi-object tracking methods. We achieve 56.0% MOTA on VisDrone2019 and 56.2% MOTA on UAVDT, reaching the state-of-the-art among existing methods. Zixuan Song, Sanping Zhou, Wei Tang 0016, Le Wang 0003 |
IEEE Trans. Multim. | 3 |
| 2025 | REGNav: Room Expert Guided Image-Goal NavigationabstractImage-goal navigation aims to steer an agent towards the goal location specified by an image. Most prior methods tackle this task by learning a navigation policy, which extracts visual features of goal and observation images, compares their similarity and predicts actions. However, if the agent is in a different room from the goal image, it's extremely challenging to identify their similarity and infer the likely goal location, which may result in the agent wandering around. Intuitively, when humans carry out this task, they may roughly compare the current observation with the goal image, having an approximate concept of whether they are in the same room before executing the actions. Inspired by this intuition, we try to imitate human behaviour and propose a Room Expert Guided Image-Goal Navigation model~(REGNav) to equip the agent with the ability to analyze whether goal and observation images are taken in the same room. Specifically, we first pre-train a room expert with an unsupervised learning technique on the self-collected unlabelled room images. The expert can extract the hidden room style information of goal and observation images and predict their relationship about whether they belong to the same room. In addition, two different fusion approaches are explored to efficiently guide the agent navigation with the room relation knowledge. Extensive experiments show that our REGNav surpasses prior state-of-the-art works on three popular benchmarks. Pengna Li, Kangyi Wu, Jingwen Fu, Sanping Zhou |
AAAI | 4 |
| 2025 | Diversifying Query: Region-Guided Transformer for Temporal Sentence GroundingabstractTemporal sentence grounding is a challenging task that aims to localize the moment spans relevant to a language description. Although recent DETR-based models have achieved notable progress by leveraging multiple learnable moment queries, they suffer from overlapped and redundant proposals, leading to inaccurate predictions. We attribute this limitation to the lack of task-related guidance for the learnable queries to serve a specific mode. Furthermore, the complex solution space generated by variable and open-vocabulary language descriptions complicates optimization, making it harder for learnable queries to adaptively distinguish each other, leading to more severe overlapped proposals. To address this limitation, we present the Region-Guided TRansformer (RGTR) for temporal sentence grounding, which introduces regional guidance to increase query diversity and eliminate overlapped proposals. Instead of using learnable queries, RGTR adopts a set of anchor pairs as moment queries to introduce explicit regional guidance. Each moment query takes charge of moment prediction for a specific temporal region, which reduces the optimization difficulty and ensures the diversity of the proposals. In addition, we design an IoU-aware scoring head to improve proposal quality. Extensive experiments demonstrate the effectiveness of RGTR, outperforming state-of-the-art methods on three public benchmarks and exhibiting good generalization and robustness on out-of-distribution splits. Xiaolong Sun, Liushuai Shi, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
AAAI | 4 |
| 2025 | RefDetector: A Simple Yet Effective Matching-based Method for Referring Expression ComprehensionabstractDespite the rapid and substantial advancements in object detection, it continues to face limitations imposed by pre-defined category sets. Current methods for visual grounding primarily focus on how to better leverage the visual backbone to generate text-tailored visual features, which may require adjusting the parameters of the entire model. Besides, some early methods, \ie, matching-based method, build upon and extend the functionality of existing object detectors by enabling them to localize an object based on free-form linguistic expressions, which have good application potential. However, the untapped potential of the matching-based approach has not been fully realized due to inadequate exploration. In this paper, we first analyze the limitations that exist in the current matching-based method (\ie, mismatch problem and complicated fusion mechanisms), and then present a simple yet effective matching-based method, namely RefDetector. To tackle the above issues, we devise a simple heuristic rule to generate proposals with improved referent recall. Additionally, we introduce a straightforward vision-language interaction module that eliminates the need for intricate manually-designed mechanisms. Moreover, we have explored the visual grounding based on the modern detector DETR, and achieved significant performance improvement. Extensive experiments on three REC benchmark datasets, \ie, RefCOCO, RefCOCO+, and RefCOCOg validate the effectiveness of the proposed method. Zhuotao Tian, Sanping Zhou, Le Wang 0003 |
AAAI | 4 |
| 2025 | Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal TransportabstractPoint-supervised Temporal Action Localization poses significant challenges due to the difficulty of identifying complete actions with a single-point annotation per action. Existing methods typically employ Multiple Instance Learning, which struggles to capture global temporal context and requires heuristic post-processing. In research on fully-supervised tasks, DETR-based structures have effectively addressed these limitations. However, it is nontrivial to merely adapt DETR to this task, encountering two major bottlenecks. (1) How to integrate point label information into the model and (2) How to select optimal decoder proposals for training in the absence of complete action segment annotations. To address this issue, we introduce an end-to-end framework by integrating Query Reformation and Optimal Transport (QROT). Specifically, we encode point labels through a set of semantic consensus queries, enabling effective focus on action-relevant snippets. Furthermore, we integrate an optimal transport mechanism to generate high-quality pseudo labels. These pseudo-labels facilitate precise proposals selection based on the Hungarian algorithm, significantly enhancing localization accuracy in point-supervised settings. Extensive experiments on the THUMOS14 and ActivityNet-v1.3 datasets demonstrate that our method outperforms existing MIL-based approaches, offering more stable and accurate temporal action localization in point-level supervision. Mengnan Liu 0001, Le Wang 0003, Sanping Zhou, Xiaolong Sun, Gang Hua 0001 |
CVPR | 3 |
| 2025 | PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic ManipulationabstractRobotic manipulation based on visual observations and natural language instructions is a long-standing challenge in robotics. Yet prevailing approaches model action distribution by adopting explicit or implicit representations, which often struggle to achieve a trade-off between accuracy and efficiency. In response, we propose PDFactor, a novel framework that models action distribution with a hybrid triplane representation. In particular, PDFactor decomposes 3D point cloud into three orthogonal feature planes and leverages a tri-perspective view transformer to produce dense cubic features as a latent diffusion field aligned with observation space representing 6-DoF action probability distribution at an arbitrary location. We employ a small denoising network conceptually as both a parameterized loss function measuring the quality of the learned latent features and an action gradient decoder to sample actions from the latent diffusion field during inference. This design enables our PDFactor to benefit from spatial awareness of explicit representation and arbitrary resolution of implicit representation, rendering it with manipulation accuracy, inference efficiency, and model scalability. Experiments demonstrate that PDFactor outperforms state-of-the-art approaches across a diverse range of manipulation tasks in RLBench simulation. Moreover, PDFactor can effectively learn multi-task policies from a limited number of human demonstrations, achieving promising accuracy in a variety of real-world manipulation tasks. Jingyi Tian, Le Wang 0003, Sanping Zhou, Haowen Sun 0003, Wei Tang 0016 |
CVPR | 3 |
| 2025 | FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic ManipulationabstractRobotic manipulation in high-precision tasks is essential for numerous industrial and real-world applications where accuracy and speed are required. Yet current diffusion-based policy learning methods generally suffer from low computational efficiency due to the iterative denoising process during inference. Moreover, these methods do not fully explore the potential of generative models for enhancing information exploration in 3D environments. In response, we propose FlowRAM, a novel framework that leverages generative models to achieve region-aware perception, enabling efficient multimodal information processing. Specifically, we devise a Dynamic Radius Schedule, which allows adaptive perception, facilitating transitions from global scene comprehension to fine-grained geometric details. Furthermore, we integrate state space models to integrate multimodal information, while preserving linear computational complexity. In addition, we employ conditional flow matching to learn action poses by regressing deterministic vector fields, simplifying the learning process while maintaining performance. We verify the effectiveness of the FlowRAM in the RLBench, an established manipulation benchmark, and achieve state-of-the-art performance. The results demonstrate that FlowRAM achieves a remarkable improvement, particularly in high-precision tasks, where it outperforms previous methods by 12.0% in average success rate. Additionally, FlowRAM is able to generate physically plausible actions for a variety of real-world tasks in less than 4 time steps, significantly increasing inference speed. Le Wang 0003, Sanping Zhou, Jingyi Tian, Haowen Sun 0003, Wei Tang 0016 |
CVPR | 3 |
| 2025 | Towards Precise Embodied Dialogue Localization via Causality Guided DiffusionabstractEmbodied localization based on vision and natural language dialogues presents a persistent challenge in embodied intelligence. Existing methods often approach this task as an image translation problem, leveraging encoder-decoder architectures to predict heatmaps. However, these methods frequently experience a deficiency in accuracy, largely due to their heavy reliance on resolution. To address this issue, we introduce CGD, a novel framework that utilizes causality guided diffusion model to directly model coordinate distributions. Specifically, CGD employs a denoising network to regress coordinates, while integrating causal learning modules, namely back-door adjustment (BDA) and front-door adjustment (FDA) to mitigate confounders during the diffusion process. This approach reduces the dependency on high resolution for improving accuracy, while effectively minimizing spurious correlations, thereby promoting unbiased learning. By guiding the denoising process with causal adjustments, CGD offers flexible control over intensity, ensuring seamless integration with diffusion models. Experimental results demonstrate that CGD outperforms state-of-the-art methods across all metrics. Additionally, we also evaluate CGD in a multi-shot setting, achieving consistently high accuracy. Le Wang 0003, Sanping Zhou, Jingyi Tian, Gang Hua 0001, Wei Tang 0016 |
CVPR | 3 |
| 2025 | Event-Equalized Dense Video CaptioningabstractDense video captioning aims to localize and caption all events in arbitrary untrimmed videos. Although previous methods have achieved appealing results, they still face the issue of temporal bias, i.e, models tend to focus more on events with certain temporal characteristics. Specifically, 1) the temporal distribution of events in training datasets is uneven. Models trained on these datasets will pay less attention to out-of-distribution events. 2) long-duration events have more frame features than short ones and will attract more attention. To address this, we argue that events, with varying temporal characteristics, should be treated equally when it comes to dense video captioning. Intuitively, different events tend to have distinct visual differences due to varied camera views, backgrounds, or subjects. Inspired by that, we intend to utilize visual features to have an approximate perception of possible events and pay equal attention to them. In this paper, we introduce a simple but effective framework, called Event-Equalized Dense Video Captioning (E2DVC) to overcome the temporal bias and treat all possible events equally. Experimental results on ActivityNet Captions and YouCook2 dataset validate the effectiveness of the proposed methods and show State-of-the-art (SOTA) performance on dense video captioning. Kangyi Wu, Pengna Li, Jingwen Fu, Yang Wu 0001, Yuhan Liu 0006, Jinjun Wang, Sanping Zhou |
CVPR | 8 |
| 2025 | Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense PredictionabstractSufficient cross-task interaction is crucial for success in multi-task dense prediction. However, sufficient interaction often results in high computational complexity, forcing existing methods to face the trade-off between interaction completeness and computational efficiency. To address this limitation, this work proposes a Bidirectional Interaction Mamba (BIM), which incorporates novel scanning mechanisms to adapt the Mamba modeling approach for multi-task dense prediction. On the one hand, we introduce a novel Bidirectional Interaction Scan (BI-Scan) mechanism, which constructs task-specific representations as bidirectional sequences during interaction. By integrating task-first and position-first scanning modes within a unified linear complexity architecture, BI-Scan efficiently preserves critical cross-task information. On the other hand, we employ a Multi-Scale Scan~(MS-Scan) mechanism to achieve multi-granularity scene modeling. This design not only meets the diverse granularity requirements of various tasks but also enhances nuanced cross-task feature interactions. Extensive experiments on two challenging benchmarks, \emph{i.e.}, NYUD-V2 and PASCAL-Context, show the superiority of our BIM vs its state-of-the-art competitors. Mang Cao, Sanping Zhou, Ye Deng 0005, Wenli Huang 0004, Le Wang 0003 |
ICCV | 2 |
| 2025 | DAMap: Distance-Aware MapNet for High Quality HD Map Construction
Jinpeng Dong, Yutong Lin, Jingwen Fu, Sanping Zhou, Nanning Zheng 0001 |
ICCV | 5 |
| 2025 | Mind the Gap: Aligning Vision Foundation Models to Image Feature MatchingabstractLeveraging the vision foundation models has emerged as a mainstream paradigm that improves the performance of image feature matching. However, previous works have ignored the misalignment when introducing the foundation models into feature matching. The misalignment arises from the discrepancy between the foundation models focusing on single-image understanding and the cross-image understanding requirement of feature matching. Specifically, 1) the embeddings derived from commonly used foundation models exhibit discrepancies with the optimal embeddings required for feature matching; 2) lacking an effective mechanism to leverage the single-image understanding ability into cross-image understanding. A significant consequence of the misalignment is they struggle when addressing multi-instance feature matching problems. To address this, we introduce a simple but effective framework, called IMD (Image feature Matching with a pre-trained Diffusion model) with two parts: 1) Unlike the dominant solutions employing contrastive-learning based foundation models that emphasize global semantics, we integrate the generative-based diffusion models to effectively capture instance-level details. 2) We leverage the prompt mechanism in generative model as a natural tunnel, propose a novel cross-image interaction prompting module to facilitate bidirectional information interaction between image pairs. To more accurately measure the misalignment, we propose a new benchmark called IMIM, which focuses on multi-instance scenarios. Our proposed IMD establishes a new state-of-the-art in commonly evaluated benchmarks, and the superior improvement 12% in IMIM indicates our method efficiently mitigates the misalignment. Yuhan Liu 0006, Jingwen Fu, Yang Wu 0001, Kangyi Wu, Pengna Li, Jiayi Wu 0002, Sanping Zhou, Jingmin Xin |
ICCV | 7 |
| 2025 | FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection TransformersabstractDetecting 3D objects accurately from multi-view 2D images is a challenging yet essential task in the field of autonomous driving. Current methods resort to integrating depth prediction to recover the spatial information for object query decoding, which necessitates explicit supervision from LiDAR points during the training phase. However, the predicted depth quality is still unsatisfactory such as depth discontinuity of object boundaries and indistinction of small objects, which are mainly caused by the sparse supervision of projected points and the use of high-level image features for depth prediction. Besides, cross-view consistency and scale invariance are also overlooked in previous methods. In this paper, we introduce Frequency-aware Positional Depth Embedding (FreqPDE) to equip 2D image features with spatial information for 3D detection transformer decoder, which can be obtained through three main modules. Specifically, the Frequency-aware Spatial Pyramid Encoder (FSPE) constructs a feature pyramid by combining high-frequency edge clues and low-frequency semantics from different levels respectively. Then the Cross-view Scale-invariant Depth Predictor (CSDP) estimates the pixel-level depth distribution with cross-view and efficient channel attention mechanism. Finally, the Positional Depth Encoder (PDE) combines the 2D image features and 3D position embeddings to generate the 3D depth-aware features for query decoding. Additionally, hybrid depth supervision is adopted for complementary depth learning from both metric and distribution aspects. Extensive experiments conducted on the nuScenes dataset demonstrate the effectiveness and superiority of our proposed method. Haisheng Su, Feixiang Song, Sanping Zhou, Wei Wu 0021, Junchi Yan, Nanning Zheng 0001 |
ICCV | 4 |
| 2025 | Moment Quantization for Video Temporal GroundingabstractVideo temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishing relevant and irrelevant moments. Previous methods focused on learning continuous features exhibit weak differentiation between foreground and background features. In this paper, we propose a novel Moment-Quantization based Video Temporal Grounding method (MQVTG), which quantizes the input video into various discrete vectors to enhance the discrimination between relevant and irrelevant moments. Specifically, MQVTG maintains a learnable moment codebook, where each video moment matches a codeword. Considering the visual diversity, i.e., various visual expressions for the same moment, MQVTG treats moment-codeword matching as a clustering process without using discrete vectors, avoiding the loss of useful information from direct hard quantization. Additionally, we employ effective prior-initialization and joint-projection strategies to enhance the maintained moment codebook. With its simple implementation, the proposed method can be integrated into existing temporal grounding models as a plug-and-play component. Extensive experiments on six popular benchmarks demonstrate the effectiveness and generalizability of MQVTG, significantly outperforming state-of-the-art methods. Further qualitative analysis shows that our method effectively groups relevant features and separates irrelevant ones, aligning with our goal of enhancing discrimination. Xiaolong Sun, Le Wang 0003, Sanping Zhou, Liushuai Shi, Mengnan Liu 0001, Gang Hua 0001 |
ICCV | 3 |
| 2025 | Versatile Multimodal Controls for Expressive Talking Human Animation
Ruobing Zheng, Zixin Zhu, Sanping Zhou, Ming Yang 0007, Le Wang 0003 |
ACM Multimedia | 6 |
| 2025 | DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic ManipulationabstractLearning generalizable robotic manipulation policies remains a key challenge due to the scarcity of diverse real-world training data. While recent approaches have attempted to mitigate this through self-supervised representation learning, most either rely on 2D vision pretraining paradigms such as masked image modeling, which primarily focus on static semantics or scene geometry, or utilize large-scale video prediction models that emphasize 2D dynamics, thus failing to jointly learn the geometry, semantics, and dynamics required for effective manipulation. In this paper, we present DynaRend, a representation learning framework that learns 3D-aware and dynamics-informed triplane features via masked reconstruction and future prediction using differentiable volumetric rendering. By pretraining on multi-view RGB-D video data, DynaRend jointly captures spatial geometry, future dynamics, and task semantics in a unified triplane representation. The learned representations can be effectively transferred to downstream robotic manipulation tasks via action value map prediction. We evaluate DynaRend on two challenging benchmarks, RLBench and Colosseum, as well as in real-world robotic experiments, demonstrating substantial improvements in policy success rate, generalization to environmental perturbations, and real-world applicability across diverse manipulation tasks. Jingyi Tian, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
NeurIPS | 3 |
| 2025 | SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World ModelsabstractWorld models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions due to disrupted spatial structure, inefficient decoding, and inadequate motion modeling. In response, we propose Scale-wise Autoregression with Motion PrOmpt (SAMPO), a hybrid framework that combines visual autoregressive modeling for intra-frame generation with causal modeling for next-frame generation. Specifically, SAMPO integrates temporal causal decoding with bidirectional spatial attention, which preserves spatial locality and supports parallel decoding within each scale. This design significantly enhances both temporal consistency and rollout efficiency. To further improve dynamic scene understanding, we devise an asymmetric multi-scale tokenizer that preserves spatial details in observed frames and extracts compact dynamic representations for future frames, optimizing both memory usage and model performance. Additionally, we introduce a trajectory-aware motion prompt module that injects spatiotemporal cues about object and robot trajectories, focusing attention on dynamic regions and improving temporal consistency and physical realism. Extensive experiments show that SAMPO achieves competitive performance in action-conditioned video prediction and model-based control, improving generation quality with 4.4× faster inference. We also evaluate SAMPO's zero-shot generalization and scaling behavior, demonstrating its ability to generalize to unseen tasks and benefit from larger model sizes. Jingyi Tian, Le Wang 0003, Zhimin Liao, Huaiyi Dong, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
NeurIPS | 8 |
| 2025 | Meta channel masking for cross-domain few-shot image classificationabstractCross-domain Few-shot Learning (CD-FSL) aims to address the challenges of FSL where significant domain gaps exist between source and target image datasets. Unlike many existing CD-FSL methods that utilize an auxiliary target dataset with a few labeled target images to enhance model generalization, our approach directly tackles the limitations imposed by the reliance on source-specific knowledge. We observe that models trained on unbalanced datasets tend to overfit to source-specific features, which, while effective in the source domain, generalize poorly to the target image domain. To address this, we introduce a novel dropout-based framework named Meta Channel Masking (MCM). This framework attenuates the learning of model channels on the source domain by dynamically masking source feature channels during training. In contrast to traditional dropout techniques that manually set masking probabilities based on statistical assumptions about the source data, our MCM framework employs a meta-learning process that automatically adjusts channel mask probabilities. This adjustment is informed by auxiliary target data, effectively minimizing few-shot loss on the auxiliary target dataset and thereby enhancing the model’s generalization capabilities in the target domain. Our extensive experiments across various image classification benchmark datasets demonstrate that our framework outperforms state-of-the-art methods. Siqi Hui, Sanping Zhou, Ye Deng 0005, Pengna Li, Jinjun Wang |
Neurocomputing | 2 |
| 2025 | AFC-RNN: Adaptive Forgetting-Controlled Recurrent Neural Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction plays a crucial and fundamental role in many computer vision tasks. Most existing works utilize recurrent neural networks to extract temporal features from trajectories because their recursive structure is inherently well-suited for time series data. However, previous methods overlook the forgetting characteristics of pedestrians when modeling historical trajectories, which may cause the model to focus on the wrong positions of historical information. In this paper, we propose a simple yet effective Adaptive Forgetting-Controlled Recurrent Neural Network (AFC-RNN) for pedestrian trajectory prediction. The core idea of AFC-RNN is a novel Adaptive Forgetting Controller (AFC), which controls the forgetting degree of the historical information at each time step explicitly and adaptively. Specifically, AFC first learns memory factors for each time step based on the temporal correlation of observed trajectories using the self-attention mechanism. Then, AFC-RNN applies these memory factors to regulate the forgetting degree of observed features at each time step from RNN. Extensive experiments and ablation studies on ETH, UCY, SDD, and NBA datasets demonstrate that our method outperforms existing state-of-the-art approaches. Additionally, we provide a mathematical analysis to demonstrate the superiority of our adaptive forgetting strategy in the AFC-RNN over traditional RNNs for trajectory forgetting modeling. Yonghao Dong, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | StructVPR++: Distill Structural and Semantic Knowledge With Weighting Samples for Visual Place RecognitionabstractVisual place recognition is a challenging task for autonomous driving and robotics, which is usually considered as an image retrieval problem. A commonly used two-stage strategy involves global retrieval followed by re-ranking using patch-level descriptors. Most deep learning-based methods in an end-to-end manner cannot extract global features with sufficient semantic information from RGB images. In contrast, re-ranking can utilize more explicit structural and semantic information in one-to-one matching process, but it is time-consuming. To bridge the gap between global retrieval and re-ranking and achieve a good trade-off between accuracy and efficiency, we propose StructVPR++, a framework that embeds structural and semantic knowledge into RGB global representations via segmentation-guided distillation. Our key innovation lies in decoupling label-specific features from global descriptors, enabling explicit semantic alignment between image pairs without requiring segmentation during deployment. Furthermore, we introduce a sample-wise weighted distillation strategy that prioritizes reliable training pairs while suppressing noisy ones. Experiments on four benchmarks demonstrate that StructVPR++ surpasses state-of-the-art global methods by 5-23% in Recall@1 and even outperforms many two-stage approaches, achieving real-time efficiency with a single RGB input. Yanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang 0005, Shi-tao Chen, Nanning Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Recurrent Aligned Network for Generalized Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a crucial component in computer vision and robotics, but remains challenging due to the domain shift problem. Previous studies have tried to tackle this problem by leveraging a portion of trajectory data from the target domain to fine-tune the model. However, such domain adaptation methods are impractical in real-world scenarios, as it is infeasible to collect trajectory data from all potential target domains. In this paper, we study a new task named generalized pedestrian trajectory prediction, with the aim of generalizing the model to unseen domains without accessing their trajectories. To tackle this task, we further introduce a Recurrent Aligned Network (RAN) to minimize the domain gap through domain alignment. Specifically, we devise a recurrent alignment module to effectively align the trajectory feature spaces at both time-state and time-sequence levels by the recurrent alignment strategy. Furthermore, we introduce a pre-aligned representation module to combine social interactions with the recurrent alignment strategy, which aims to consider social interactions during the alignment process instead of just target trajectories. We extensively evaluate our method and compare it with state-of-the-art methods on three widely used benchmarks. The experimental results demonstrate the superior generalization capability of our method. Our work not only fills the gap in the generalization setting for practical pedestrian trajectory prediction, but also sets strong baselines in this field. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | A Novel Dense Object Detector With Scale Balanced Sample Assignment and RefinementabstractScale variation of objects remains one of the crucial challenges in object detection. Currently, conventional dense detectors with fixed receptive fields and label weights are not conducive to the detection of multi-scale objects. However, the design limitations of unbalanced label weights and fixed refinement for multi-scale objects and multi-tasks in these studies make it difficult to achieve better detection performance. In this paper, we propose a novel dense detector named Balanced FCOS which consists of two components: Balanced Label Assignment (BLA) and Flexible Shape-based Refinement (FSR). The BLA implements scale-balanced sample assignment by introducing reweighting factors consisting of localization and classification scores into the label assignment. Low-quality but high-weight samples can be weakened by the BLA. Furthermore, we design a cross-reweighting mechanism in the BLA to ensure score consistency between classification and localization. The FSR implements scale-balanced sample refinement by learning flexible sample points’ offsets for multi-scale objects and multi-tasks based on objects’ coarse features to get more discriminative features with appropriate receptive field. In addition, better features obtained by FSR are beneficial to get better classification and localization scores, which can be used by BLA to produce accurate label weights. Only equipped with the BLA, we can achieve 41.7/46.6 AP under R50/R101-FCOS without any additional parameters. When combining the BLA with the FSR, our Balanced FCOS achieves SOTA results among dense detectors on the COCO test-dev set. Experiments conducted on other heads (T-Head, DyHead), detectors (DINO), and datasets (AI-TOD) further demonstrate the effectiveness of our method. Jinpeng Dong, Dingyi Yao, Sanping Zhou, Nanning Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Robust Noisy Label Learning via Two-Stream Sample DistillationabstractNoisy label learning aims to learn robust networks under the supervision of noisy labels, which plays a critical role in deep learning. Existing work either conducts sample selection or label correction to deal with noisy labels during the model training process. In this paper, we design a simple yet effective sample selection framework, termed Two-Stream Sample Distillation (TSSD), for noisy label learning, which can extract more high-quality samples with clean labels to improve the robustness of network training. Firstly, a novel Parallel Sample Division (PSD) module is designed to generate acertaintraining set with sufficient reliable positive and negative samples by jointly considering the sample structure in feature space and the human prior in loss space. Secondly, a novel Meta Sample Purification (MSP) module is further designed to mine adequate semi-hard samples from the remaininguncertaintraining set by learning a strong meta classifier with extra golden data. As a result, more and more high-quality samples will be distilled from the noisy training set to train networks robustly in every iteration. Extensive experiments on four benchmark datasets, including CIFAR-10, CIFAR-100, Tiny-ImageNet and Clothing-1M, show that our method has achieved state-of-the-art results over its competitors. Sihan Bai, Sanping Zhou, Le Wang 0003, Nanning Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Visual-Linguistic Feature Alignment With Semantic and Kinematic Guidance for Referring Multi-Object TrackingabstractReferring Multi-Object Tracking (RMOT) aims to dynamically track an arbitrary number of referred targets in a video sequence according to the language expression. Previous methods mainly focus on cross-modal fusion at the feature level with designed structures. However, the insufficient visual-linguistic alignment is prone to causing visual-linguistic mismatches, leading to some targets being tracked but not correctly referred especially when facing the language expression with complex semantics or motion descriptions. To this end, we propose to conduct visual-linguistic alignment with semantic and kinematic guidance to effectively align the visual features with more diverse language expressions. In this paper, we put forward a novel end-to-end RMOT framework SKTrack, which follows the transformer-based architecture with a Language-Guided Decoder (LGD) and a Motion-Aware Aggregator (MAA). In particular, the LGD performs deep semantic interaction layer-by-layer in a single frame to enhance the alignment ability of the model, while the MAA conducts temporal feature fusion and alignment across multiple frames to enable the alignment between visual targets and language expression with motion descriptions. Extensive experiments on the Refer-KITTI and Refer-KITTI-v2 demonstrate that SKTrack achieves state-of-the-art performance and verify the effectiveness of our framework and its components. Sanping Zhou, Le Wang 0003 |
IEEE Trans. Multim. | 2 |
| 2025 | Meta Pairwise Relationship Distillation for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) is challenging due to the lack of ground-truth labels. Most existing methods rely on pseudo labels estimated via iterative clustering and thus are highly susceptible to performance penalties incurred by the inaccurate estimated number of clusters. Alternatively, we utilize the sample pairs with pairwise pseudo labels to guide the feature learning to avoid the dilemma of determining cluster numbers. In this article, we propose a meta pairwise relationship distillation (MPRD) method that incorporates a graph convolutional network (GCN) to provide high-fidelity pairwise relationships to supervise the model training. A small amount of metadata with very-confidence pairwise relationships and the unlabeled pairs with the provided pseudo pairwise relationships participate in the GCN training. Besides, we introduce a hard sample deduction (HSD) module to timely mine the sample pairs with error-prone pairwise pseudo labels to mitigate the misled optimization by noisy labels. Furthermore, since the features of each positive pair represent the same person, we design a positive pair alignment (PPA) module to reduce the redundant information in each feature, which is achieved by minimizing the difference between each positive pair's feature distributions. Extensive experiments on the Market-1501, DukeMTMC-reID, and MSMT17 datasets show that our method outperforms the state-of-the-art unsupervised methods. Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Voxel or Pillar: Exploring Efficient Point Cloud Representation for 3D Object DetectionabstractEfficient representation of point clouds is fundamental for LiDAR-based 3D object detection. While recent grid-based detectors often encode point clouds into either voxels or pillars, the distinctions between these approaches remain underexplored. In this paper, we quantify the differences between the current encoding paradigms and highlight the limited vertical learning within. To tackle these limitations, we propose a hybrid detection framework named Voxel-Pillar Fusion (VPF), which synergistically combines the unique strengths of both voxels and pillars. To be concrete, we first develop a sparse voxel-pillar encoder that encodes point clouds into voxel and pillar features through 3D and 2D sparse convolutions respectively, and then introduce the Sparse Fusion Layer (SFL), facilitating bidirectional interaction between sparse voxel and pillar features. Our computationally efficient, fully sparse method can be seamlessly integrated into both dense and sparse detectors. Leveraging this powerful yet straightforward representation, VPF delivers competitive performance, achieving real-time inference speeds on the nuScenes and Waymo Open Dataset. Sanping Zhou, Jinpeng Dong, Nanning Zheng 0001 |
AAAI | 2 |
| 2024 | Temporal Correlation Vision Transformer for Video Person Re-IdentificationabstractVideo Person Re-Identification (Re-ID) is a task of retrieving persons from multi-camera surveillance systems. Despite the progress made in leveraging spatio-temporal information in videos, occlusion in dense crowds still hinders further progress. To address this issue, we propose a Temporal Correlation Vision Transformer (TCViT) for video person Re-ID. TCViT consists of a Temporal Correlation Attention (TCA) module and a Learnable Temporal Aggregation (LTA) module. The TCA module is designed to reduce the impact of non-target persons by relative state, while the LTA module is used to aggregate frame-level features based on their completeness. Specifically, TCA is a parameter-free module that first aligns frame-level features to restore semantic coherence in videos and then enhances the features of the target person according to temporal correlation. Additionally, unlike previous methods that treat each frame equally with a pooling layer, LTA introduces a lightweight learnable module to weigh and aggregate frame-level features under the guidance of a classification score. Extensive experiments on four prevalent benchmarks demonstrate that our method achieves state-of-the-art performance in video Re-ID. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
AAAI | 3 |
| 2024 | Towards Generalizable Multi-Object TrackingabstractMulti-Object Tracking (MOT) encompasses various tracking scenarios, each characterized by unique traits. Ef-fective trackers should demonstrate a high degree of gen-eralizability across diverse scenarios. However, existing trackers struggle to accommodate all aspects or necessi-tate hypothesis and experimentation to customize the asso-ciation information (motion and/or appearance) for a given scenario, leading to narrowly tailored solutions with limited generalizability. In this paper, we investigate the factors that influence trackers' generalization to different scenar-ios and concretize them into a set of tracking scenario at-tributes to guide the design of more generalizable trackers. Furthermore, we propose a “point-wise to instance-wise relation” framework for MOT, i.e., GeneralTrack, which can generalize across diverse scenarios while eliminating the need to balance motion and appearance. Thanks to its supe-rior generalizability, our proposed GeneralTrack achieves state-of-the-art performance on multiple benchmarks and demonstrates the potential for domain generalization. Le Wang 0003, Sanping Zhou, Panpan Fu, Gang Hua 0001, Wei Tang 0016 |
CVPR | 3 |
| 2024 | AugDETR: Improving Multi-scale Learning for Detection Transformer
Jinpeng Dong, Yutong Lin, Sanping Zhou, Nanning Zheng 0001 |
ECCV (24) | 4 |
| 2024 | PMT: Progressive Mean Teacher via Exploring Temporal Consistency for Semi-Supervised Medical Image Segmentation
Sanping Zhou, Le Wang 0003, Nanning Zheng 0001 |
ECCV (68) | 2 |
| 2024 | Stepwise Multi-grained Boundary Detector for Point-Supervised Temporal Action Localization
Mengnan Liu 0001, Le Wang 0003, Sanping Zhou, Qilin Zhang 0004, Gang Hua 0001 |
ECCV (7) | 3 |
| 2024 | Learning Anomalies with Normality Prior for Unsupervised Video Anomaly Detection
Haoyue Shi 0002, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
ECCV (6) | 3 |
| 2024 | Vehicle Trajectory Prediction with Soft Behavior ConstraintsabstractTrajectory prediction plays a crucial role in autonomous driving, but it is challenging due to the multi-modal nature of future trajectories. Behavior information is frequently employed to capture more diverse modalities of future trajectories. Traditional behavior information is typically hard-encoded, which is often inaccurate and inadequate for reflecting future multimodality. Therefore, we introduce the concept of soft vehicle behavior, which is represented as a probability distribution over a predefined comprehensive set of behaviors. This approach allows for a more rational depiction of vehicle behavior and captures potential future driving modalities. Based on it, we propose a new soft-behavior-constrained vehicle trajectory prediction framework. The framework consists of a backbone and a lightweight and plug-and-play behavior prediction module, which is used to imbue soft behavior constraints to assist in representation learning. We integrated the behavior prediction module into five representative trajectory predictors and achieved improvements of at least 4.2% in minFDE(K=5) on the nuScenes dataset and 0.5% in minFDE(K=6) on the Argoverse 1 motion forecasting dataset. These universal increments prove the effectiveness and generalizability of soft behavior constraints in vehicle trajectory prediction. Ke Ye, Sanping Zhou, Miao Kang, Jingwen Fu, Nanning Zheng 0001 |
IROS | 2 |
| 2024 | Semantic-aware Representation Learning for Homography EstimationabstractHomography estimation is the task of determining the transformation from an image pair. Our approach focuses on employing detector-free feature matching methods to address this issue. Previous work has underscored the importance of incorporating semantic information, however there still lacks an efficient way to utilize semantic information. Previous methods suffer from treating the semantics as a pre-processing, causing the utilization of semantics overly coarse-grained and lack adaptability when dealing with different tasks. In our work, we seek another way to use the semantic information, that is semantic-aware feature representation learning framework. Based on this, we propose SRMatcher, a new detector-free feature matching method, which encourages the network to learn integrated semantic feature representation. Specifically, to capture precise and rich semantics, we leverage the capabilities of recently popularized vision foundation models (VFMs) trained on extensive datasets. Then, a cross-images Semantic-aware Fusion Block (SFB) is proposed to integrate its fine-grained semantic features into the feature representation space. In this way, by reducing errors stemming from semantic inconsistencies in matching pairs, our proposed SRMatcher is able to deliver more accurate and realistic outcomes. Extensive experiments show that SRMatcher surpasses solid baselines and attains SOTA results on multiple real-world datasets. Compared to the previous SOTA approach GeoFormer, SRMatcher increases the area under the cumulative curve (AUC) by about 11% on HPatches. Additionally, the SRMatcher could serve as a plug-and-play framework for other matching methods like LoFTR, yielding substantial precision improvement. Yuhan Liu 0006, Qianxin Huang, Siqi Hui, Jingwen Fu, Sanping Zhou, Kangyi Wu, Pengna Li, Jinjun Wang |
ACM Multimedia | 5 |
| 2024 | Molecule Design by Latent Prompt TransformerabstractThis work explores the challenging problem of molecule design by framing it as a conditional generative modeling task, where target biological properties or desired chemical constraints serve as conditioning variables.
We propose the Latent Prompt Transformer (LPT), a novel generative model comprising three components: (1) a latent vector with a learnable prior distribution modeled by a neural transformation of Gaussian white noise; (2) a molecule generation model based on a causal Transformer, which uses the latent vector as a prompt; and (3) a property prediction model that predicts a molecule's target properties and/or constraint values using the latent prompt. LPT can be learned by maximum likelihood estimation on molecule-property pairs. During property optimization, the latent prompt is inferred from target properties and constraints through posterior sampling and then used to guide the autoregressive molecule generation.
After initial training on existing molecules and their properties, we adopt an online learning algorithm to progressively shift the model distribution towards regions that support desired target properties. Experiments demonstrate that LPT not only effectively discovers useful molecules across single-objective, multi-objective, and structure-constrained optimization tasks, but also exhibits strong sample efficiency. Deqian Kong, Jianwen Xie, Edouardo Honig, Shuanghong Xue, Pei Lin, Sanping Zhou, Nanning Zheng 0001, Ying Nian Wu |
NeurIPS | 8 |
| 2024 | Referencing Where to Focus: Improving Visual Grounding with Referential QueryabstractVisual Grounding aims to localize the referring object in an image given a natural language expression. Recent advancements in DETR-based visual grounding methods have attracted considerable attention, as they directly predict the coordinates of the target object without relying on additional efforts, such as pre-generated proposal candidates or pre-defined anchor boxes. However, existing research primarily focuses on designing stronger multi-modal decoder, which typically generates learnable queries by random initialization or by using linguistic embeddings. This vanilla query generation approach inevitably increases the learning difficulty for the model, as it does not involve any target-related information at the beginning of decoding. Furthermore, they only use the deepest image feature during the query learning process, overlooking the importance of features from other levels. To address these issues, we propose a novel approach, called RefFormer. It consists of the query adaption module that can be seamlessly integrated into CLIP and generate the referential query to provide the prior context for decoder, along with a task-specific decoder. By incorporating the referential query into the decoder, we can effectively mitigate the learning difficulty of the decoder, and accurately concentrate on the target object. Additionally, our proposed query adaption module can also act as an adapter, preserving the rich knowledge within CLIP without the need to tune the parameters of the backbone network. Extensive experiments demonstrate the effectiveness and efficiency of our proposed method, outperforming state-of-the-art approaches on five visual grounding benchmarks. Zhuotao Tian, Qingpei Guo, Sanping Zhou, Ming Yang 0007, Le Wang 0003 |
NeurIPS | 5 |
| 2024 | End-to-end pedestrian trajectory prediction via Efficient Multi-modal Predictors
Sanping Zhou, Le Wang 0003, Liushuai Shi, Yonghao Dong, Gang Hua 0001 |
Comput. Vis. Image Underst. | 2 |
| 2024 | KPDet: Keypoint-based 3D object detection with Parametric Radius Learning
Sanping Zhou, Xinrui Yan, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2024 | Gradient-guided channel masking for cross-domain few-shot learning
Siqi Hui, Sanping Zhou, Ye Deng 0005, Yang Wu 0001, Jinjun Wang |
Knowl. Based Syst. | 2 |
| 2024 | Residual feature learning with hierarchical calibration for gaze estimation
Zhengdan Yin, Sanping Zhou, Le Wang 0003, Gang Hua 0001, Nanning Zheng 0001 |
Mach. Vis. Appl. | 2 |
| 2024 | Sparse self-attention transformer for image inpainting
Wenli Huang 0004, Ye Deng 0005, Siqi Hui, Yang Wu 0001, Sanping Zhou, Jinjun Wang |
Pattern Recognit. | 5 |
| 2024 | Transfer easy to hard: Adversarial contrastive feature learning for unsupervised person re-identification
Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
Pattern Recognit. | 3 |
| 2024 | Bidirectional feature learning network for RGB-D salient object detectionabstractRGB-D salient object detection aims to perform the pixel-wise localization of salient objects from both RGB and depth images, whose challenge mainly comes from how to learn complementary features from each modality. Existing works often use increasingly large models for performance enhancement, which need large memory and time consumption in practice. In this paper, we propose a simple yet effective B idirectional F eature L earning Net work (BFLNet) for RGB-D salient object detection under limited memory and time conditions. To achieve accurate performance with lightweight backbone networks , an effective B idirectional F eature F usion (BFF) module is designed to merge features from both RGB and depth streams, in which the cross-modal fusions and cross-scale fusions are jointly conducted to fuse the immediate features in multiple scales and multiple modals. What is more, a simple D ual C onsistency L oss (DCL) function is designed to prompt cross-modal fusion by keeping the consistency between cross-modal target predictions. Extensive experiments on four benchmark datasets demonstrate that our method has achieved the state-of-the-art performance with high efficiency in RGB-D salient object detection. Code will be available at https://github.com/nightsky-nostar/BFLNet . Ye Niu, Sanping Zhou, Yonghao Dong, Le Wang 0003, Jinjun Wang, Nanning Zheng 0001 |
Pattern Recognit. | 2 |
| 2024 | CR-former: Single-Image Cloud Removal With Focused Taylor AttentionabstractCloud removal aims to restore high-quality images from cloud-contaminated captures, which is essential in remote sensing applications. Effectively modeling the long-range relationships between image features is key to achieving high-quality cloud-free images. While self-attention mechanisms excel at modeling long-distance relationships, their computational complexity scales quadratically with image resolution, limiting their applicability to high-resolution remote sensing images. Current cloud removal methods have mitigated this issue by restricting the global receptive field to smaller regions or adopting channel attention to model long-range relationships. However, these methods either compromise pixel-level long-range dependencies or lose spatial information, potentially leading to structural inconsistencies in restored images. In this work, we propose the focused Taylor attention (FT-Attention), which captures pixel-level long-range relationships without limiting the spatial extent of attention and achieves the$\mathcal {O}(N)$computational complexity, where N represents the image resolution. Specifically, we utilize Taylor series expansions to reduce the computational complexity of the attention mechanism from$\mathcal {O}(N^{2})$to$\mathcal {O}(N)$, enabling efficient capture of pixel relationships directly in high-resolution images. Additionally, to fully leverage the informative pixel, we develop a new normalization function for the query and key, which produces more distinguishable attention weights, enhancing focus on important features. Building on FT-Attention, we design a U-net style network, termed the CR-former, specifically for cloud removal. Extensive experimental results on representative cloud removal datasets demonstrate the superior performance of our CR-former. The code is available athttps://github.com/wuyang2691/CR-former. Yang Wu 0001, Ye Deng 0005, Sanping Zhou, Yuhan Liu 0006, Wenli Huang 0004, Jinjun Wang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Disentangled Sample Guidance Learning for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) is challenging due to the lack of ground truth labels. Most existing methods employ iterative clustering to generate pseudo labels for unlabeled training data to guide the learning process. However, how to select samples that are both associated with high-confidence pseudo labels and hard (discriminative) enough remains a critical problem. To address this issue, a disentangled sample guidance learning (DSGL) method is proposed for unsupervised Re-ID. The method consists of disentangled sample mining (DSM) and discriminative feature learning (DFL). DSM disentangles (unlabeled) person images into identity-relevant and identity-irrelevant factors, which are used to construct disentangled positive/negative groups that contain discriminative enough information. DFL incorporates the mined disentangled sample groups into model training by a surrogate disentangled learning loss and a disentangled second-order similarity regularization, to help the model better distinguish the characteristics of different persons. By using the DSGL training strategy, the mAP on Market-1501 and MSMT17 increases by 6.6% and 10.1% when applying the ResNet50 framework, and by 0.6% and 6.9% with the vision transformer (VIT) framework, respectively, validating the effectiveness of the DSGL method. Moreover, DSGL surpasses previous state-of-the-art methods by achieving higher Top-1 accuracy and mAP on the Market-1501, MSMT17, PersonX, and VeRi-776 datasets. The source code for this paper is available at https://github.com/jihaoxuanye/DiseSGL. Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | FFINet: Future Feedback Interaction Network for Motion ForecastingabstractMotion forecasting plays a crucial role in autonomous driving, with the aim of predicting the future reasonable motions of traffic agents. Most existing methods mainly model the historical interactions between agents and the environment, and predict multi-modal trajectories in a feedforward process, ignoring potential trajectory changes caused by future interactions between agents. In this paper, we propose a novel Future Feedback Interaction Network (FFINet) to aggregate the current, observations and potential future interaction features for trajectory prediction. Firstly, we employ different spatial-temporal encoders to embed the decomposed position vectors and the current position of each scene, providing rich features for the subsequent cross-temporal aggregation. Secondly, the relative interaction and cross-temporal aggregation strategies are sequentially adopted to integrate features in the current fusion module, observation interaction module, future feedback module and global fusion module, in which the future feedback module can enable the understanding of pre-action by feeding the influence of preview information to feedforward prediction. Thirdly, the comprehensive interaction features are further fed into final predictor to generate the joint predicted trajectories of multiple agents. Extensive experimental results show that our FFINet achieves the state-of-the-art performance on Argoverse 1 and Argoverse 2 motion forecasting benchmarks. Miao Kang, Shengqi Wang, Sanping Zhou, Ke Ye, Jingjing Jiang, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Sparse Pedestrian Character Learning for Trajectory PredictionabstractPedestrian trajectory prediction in a first-person view has recently attracted much attention due to its importance in autonomous driving. Recent work utilizes pedestrian character information, i.e., action and appearance, to improve the learned trajectory embedding and achieves state-of-the-art performance. However, it neglects the invalid and negative pedestrian character information, which is harmful to trajectory representation and thus leads to performance degradation. To address this issue, we present a two-stream sparse-character-based network (TSNet) for pedestrian trajectory prediction. Specifically, TSNet learns the negative-removed characters in the sparse character representation stream to improve the trajectory embedding obtained in the trajectory representation stream. Moreover, to model the negative-removed characters, we propose a novel sparse character graph, including the sparse category and sparse temporal character graphs, to learn the different effects of various characters in category and temporal dimensions, respectively. Extensive experiments on two first-person view datasets, PIE and JAAD, show that our method outperforms existing state-of-the-art methods. In addition, ablation studies demonstrate different effects of various characters and prove that TSNet outperforms approaches without eliminating negative characters. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Single-Shot and Multi-Shot Feature Learning for Multi-Object TrackingabstractMulti-Object Tracking (MOT) remains a vital component of intelligent video analysis, which aims to locate targets and maintain a consistent identity for each target throughout a video sequence. Existing works usually learn a discriminative feature representation, such as motion and appearance, to associate the detections across frames, which are easily affected by mutual occlusion and background clutter in practice. In this paper, we propose a simple yet effective two-stage feature learning paradigm to jointly learn single-shot and multi-shot features for different targets, so as to achieve robust data association in the tracking process. For the detections without being associated, we design a novel single-shot feature learning module to extract discriminative features of each detection, which can efficiently associate targets between adjacent frames. For the tracklets being lost several frames, we design a novel multi-shot feature learning module to extract discriminative features of each tracklet, which can accurately refind these lost targets after a long period. Once equipped with a simple data association logic, the resulting VisualTracker can perform robust MOT based on the single-shot and multi-shot feature representations. Extensive experimental results demonstrate that our method has achieved significant improvements on MOT17 and MOT20 datasets while reaching state-of-the-art performance on DanceTrack dataset. Sanping Zhou, Le Wang 0003, Jinjun Wang, Nanning Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Abnormal Ratios Guided Multi-Phase Self-Training for Weakly-Supervised Video Anomaly DetectionabstractWeakly-supervised Video Anomaly Detection (W-VAD) aims to detect abnormal events in videos given only video-level labels for training. Recent methods relying on multiple instance learning (MIL) and self-training achieve good performance, but they tend to focus on learning easy abnormal patterns while ignoring hard ones, e.g., unusual driving trajectory or over-speeding driving. How to detect hard anomalies is a critical but largely ignored problem in W-VAD. To tackle this challenge, we propose a novel framework, termed Abnormal Ratios guided Multi-phase Self-training (ARMS), for W-VAD. It includes a new abnormal ratio-based MIL (AR-MIL) loss and a new multi-phase self-training paradigm. The AR-MIL loss guides the learning of hard anomalies by enforcing a minimum ratio of abnormal snippets in an abnormal video and no abnormal snippets in a normal video. Our multi-phase self-training paradigm sequentially performs bootstrapping, hard anomalies mining, and adaptive self-training so as to address pseudo labeling on easy anomalies, detect hard anomalies, and setting adaptive abnormal ratios for different videos in a unified framework. Experimental results on three benchmark datasets, i.e., ShanghaiTech, UCF-Crime, and XD-Violence, show that ARMS outperforms all previous state-of-the-art methods and has a great advantage in detecting hard anomalies. Haoyue Shi 0002, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
IEEE Trans. Multim. | 3 |
| 2024 | Inverse Adversarial Diversity Learning for Network EnsembleabstractNetwork ensemble aims to obtain better results by aggregating the predictions of multiple weak networks, in which how to keep the diversity of different networks plays a critical role in the training process. Many existing approaches keep this kind of diversity either by simply using different network initializations or data partitions, which often requires repeated attempts to pursue a relatively high performance. In this article, we propose a novel inverse adversarial diversity learning (IADL) method to learn a simple yet effective ensemble regime, which can be easily implemented in the following two steps. First, we take each weak network as a generator and design a discriminator to judge the difference between the features extracted by different weak networks. Second, we present an inverse adversarial diversity constraint to push the discriminator to cheat generators that all the resulting features of the same image are too similar to distinguish each other. As a result, diverse features will be extracted by these weak networks through a min-max optimization. What is more, our method can be applied to a variety of tasks, such as image classification and image retrieval, by applying a multitask learning objective function to train all these weak networks in an end-to-end manner. We conduct extensive experiments on the CIFAR-10, CIFAR-100, CUB200-2011, and CARS196 datasets, in which the results show that our method significantly outperforms most of the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Le Wang 0003, Xingyu Wan, Siqi Hui, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Multi-Stream Representation Learning for Pedestrian Trajectory PredictionabstractForecasting the future trajectory of pedestrians is an important task in computer vision with a range of applications, from security cameras to autonomous driving. It is very challenging because pedestrians not only move individually across time but also interact spatially, and the spatial and temporal information is deeply coupled with one another in a multi-agent scenario. Learning such complex spatio-temporal correlation is a fundamental issue in pedestrian trajectory prediction. Inspired by the procedure that the hippocampus processes and integrates spatio-temporal information to form memories, we propose a novel multi-stream representation learning module to learn complex spatio-temporal features of pedestrian trajectory. Specifically, we learn temporal, spatial and cross spatio-temporal correlation features in three respective pathways and then adaptively integrate these features with learnable weights by a gated network. Besides, we leverage the sparse attention gate to select informative interactions and correlations brought by complex spatio-temporal modeling and reduce complexity of our model. We evaluate our proposed method on two commonly used datasets, i.e. ETH-UCY and SDD, and the experimental results demonstrate our method achieves the state-of-the-art performance. Code: https://github.com/YuxuanIAIR/MSRL-master Le Wang 0003, Sanping Zhou, Jinghai Duan, Gang Hua 0001, Wei Tang 0016 |
AAAI | 3 |
| 2023 | MotionTrack: Learning Robust Short-Term and Long-Term Motions for Multi-Object TrackingabstractThe main challenge of Multi-Object Tracking (MOT) lies in maintaining a continuous trajectory for each target. Existing methods often learn reliable motion patterns to match the same target between adjacent frames and discriminative appearance features to re-identify the lost targets after a long period. However, the reliability of motion prediction and the discriminability of appearances can be easily hurt by dense crowds and extreme occlusions in the tracking process. In this paper, we propose a simple yet effective multi-object tracker, i.e., MotionTrack, which learns robust short-term and long-term motions in a unified framework to associate trajectories from a short to long range. For dense crowds, we design a novel Interaction Module to learn interaction-aware motions from short-term trajectories, which can estimate the complex movement of each target. For extreme occlusions, we build a novel Refind Module to learn reliable long-term motions from the target's history trajectory, which can link the interrupted trajectory with its corresponding detection. Our Interaction Module and Refind Module are embedded in the well-known tracking-by-detection paradigm, which can work in tandem to maintain superior performance. Extensive experimental results on MOT17 and MOT20 datasets demonstrate the superiority of our approach in challenging scenarios, and it achieves state-of-the-art performances at various MOT metrics. Code is available at https://github.com/qwomeng/MotionTrack. Sanping Zhou, Le Wang 0003, Jinghai Duan, Gang Hua 0001, Wei Tang 0016 |
CVPR | 2 |
| 2023 | StructVPR: Distill Structural Knowledge with Weighting Samples for Visual Place RecognitionabstractVisual place recognition (VPR) is usually considered as a specific image retrieval problem. Limited by existing training frameworks, most deep learning-based works cannot extract sufficiently stable global features from RGB images and rely on a time-consuming re-ranking step to exploit spatial structural information for better performance. In this paper, we propose StructVPR, a novel training architecture for VPR, to enhance structural knowledge in RGB global features and thus improve feature stability in a constantly changing environment. Specifically, StructVPR uses segmentation images as a more definitive source of structural knowledge input into a CNN network and applies knowledge distillation to avoid online segmentation and inference of seg-branch in testing. Considering that not all samples contain high-quality and helpful knowledge, and some even hurt the performance of distillation, we partition samples and weigh each sample's distillation loss to enhance the expected knowledge precisely. Finally, StructVPR achieves impressive performance on several benchmarks using only global retrieval and even outperforms many two-stage approaches by a large margin. After adding additional re-ranking, ours achieves state-of-the-art performance while maintaining a low computational cost. Yanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang 0005, Shi-tao Chen, Nanning Zheng 0001 |
CVPR | 2 |
| 2023 | MLF-DET: Multi-Level Fusion for Cross-Modal 3D Object Detection
Zewei Lin, Yanqing Shen, Sanping Zhou, Shi-tao Chen, Nanning Zheng 0001 |
ICANN (7) | 3 |
| 2023 | Sparse Instance Conditioned Multimodal Trajectory PredictionabstractPedestrian trajectory prediction is critical in many vision tasks but challenging due to the multimodality of the future trajectory. Most existing methods predict multi-modal trajectories conditioned by goals (future endpoints) or instances (all future points). However, goal-conditioned methods ignore the intermediate process and instance-conditioned methods ignore the stochasticity of pedestrian motions. In this paper, we propose a simple yet effective Sparse Instance Conditioned Network (SICNet), which gives a balanced solution between goal-conditioned and instance-conditioned methods. Specifically, SICNet learns comprehensive sparse instances, i.e., representative points of the future trajectory, through a mask generated by a long short-term memory encoder and uses the memory mechanism to store and retrieve such sparse instances. Hence SICNet can decode the observed trajectory into the future prediction conditioned on the stored sparse instance. Moreover, we design a memory refinement module that refines the retrieved sparse instances from the memory to reduce memory recall errors. Extensive experiments on ETH-UCY and SDD datasets show that our method outperforms existing state-of-the-art methods. In addition, ablation studies demonstrate the superiority of our method compared with goal-conditioned and instance-conditioned approaches. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
ICCV | 3 |
| 2023 | Parallel Attention Interaction Network for Few-Shot Skeleton-based Action RecognitionabstractLearning discriminative features from very few labeled samples to identify novel classes has received increasing attention in skeleton-based action recognition. Existing works aim to learn action-specific embeddings by exploiting either intra-skeleton or inter-skeleton spatial associations, which may lead to less discriminative representations. To address these issues, we propose a novel Parallel Attention Interaction Network (PAINet) that incorporates two complementary branches to strengthen the match by inter-skeleton and intraskeleton correlation. Specifically, a topology encoding module utilizing topology and physical information is proposed to enhance the modeling of interactive parts and joint pairs in both branches. In the Cross Spatial Alignment branch, we employ a spatial cross-attention module to establish joint associations across sequences, and a directional Average Symmetric Surface Metric is introduced to locate the closest temporal similarity. In parallel, the Cross Temporal Alignment branch incorporates a spatial self-attention module to aggregate spatial context within sequences as well as applies the temporal cross-attention network to correct misalignment temporally and calculate similarity. Extensive experiments on three skeleton benchmarks, namely NTU-T, NTU-S, and Kinetics, demonstrate the superiority of our framework and consistently outperform state-of-the-art methods. Sanping Zhou, Le Wang 0003, Gang Hua 0001 |
ICCV | 2 |
| 2023 | Trajectory Unified Transformer for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is an essential link to understanding human behavior. Recent work achieves state-of-the-art performance gained from hand-designed post-processing, e.g., clustering. However, this post-processing suffers from expensive inference time and neglects the probability that the predicted trajectory disturbs downstream safety decisions. In this paper, we present Trajectory Unified TRansformer, called TUTR, which unifies the trajectory prediction components, social interaction, and multimodal trajectory prediction, into a transformer encoder-decoder architecture to effectively remove the need for post-processing. Specifically, TUTR parses the relationships across various motion modes using an explicit global prediction and an implicit mode-level transformer encoder. Then, TUTR attends to the social interactions with neighbors by a social-level transformer decoder. Finally, a dual prediction forecasts diverse trajectories and corresponding probabilities in parallel without post-processing. TUTR achieves state-of-the-art accuracy performance and improvements in inference speed of about 10× - 40× compared to previous well-tuned state-of-the-art methods using post-processing. Liushuai Shi, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
ICCV | 3 |
| 2023 | Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action LocalizationabstractSemi-Supervised Temporal Action Localization (SS-TAL) aims to improve the generalization ability of action detectors with large-scale unlabeled videos. Albeit the recent advancement, one of the major challenges still remains: noisy pseudo labels hinder efficient learning on abundant unlabeled videos, embodied as location biases and category errors. In this paper, we dive deep into such an important but understudied dilemma. To this end, we propose a unified framework, termed Noisy Pseudo-Label Learning, to handle both location biases and category errors. Specifically, our method is featured with (1) Noisy Label Ranking to rank pseudo labels based on the semantic confidence and boundary reliability, (2) Noisy Label Filtering to address the class-imbalance problem of pseudo labels caused by category errors, (3) Noisy Label Learning to penalize in-consistent boundary predictions to achieve noise-tolerant learning for heavy location biases. As a result, our method could effectively handle the label noise problem and improve the utilization of a large amount of unlabeled videos. Extensive experiments on THUMOS14 and ActivityNet v1.3 demonstrate the effectiveness of our method. The code is available at github.com/kunnxia/NPL. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
ICCV | 3 |
| 2023 | Pseudo Labels Refinement with Intra-Camera Similarity for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) aims to retrieve person images across cameras without any identity labels. Most clustering-based methods roughly divide image features into clusters and neglect the feature distribution noise caused by domain shifts among different cameras, leading to inevitable performance degradation. To address this challenge, we propose a novel label refinement framework with clustering intra-camera similarity. Intra-camera feature distribution pays more attention to the appearance of pedestrians and labels are more reliable. We conduct intra-camera training to get local clusters in each camera, respectively, and refine inter-camera clusters with local results. We hence train the Re-ID model with refined reliable pseudo labels in a self-paced way. Extensive experiments demonstrate that the proposed method surpasses state-of-the-art performance. Code is available at https://github.com/leeBooMla/ICSR. Pengna Li, Kangyi Wu, Sanping Zhou, Qianxin Huang, Jinjun Wang |
ICIP | 3 |
| 2023 | TEMI-MOT: Towards Efficient Multi-Modality Instance-Aware Feature Learning for 3D Multi-Object Trackingabstract3D multi-object tracking is one of the key technologies of autonomous driving, which aims to ensure that autonomous driving vehicles accurately perceive the movements and intentions of surrounding traffic participants. In recent years, some 3D multi-object tracking methods based on multimodality have been proposed. Although these methods improve the accuracy of object association in the tracking process, these methods are still difficult to effectively deal with the problems of feature ambiguity due to occlusion, incorrect feature alignment between different modalities, and confusion of adjacent target features caused by coarse-grained feature maps. To address these problems, we propose a new multi-modality feature learning method for 3D multi-object tracking, named TEMI-MOT, which is composed of three modules in series: the point-guided image feature sampler, the instance-aware feature encoder, and the tracking pipeline. The point-guided feature sampler realizes the alignment between the point cloud and image features, the instance-aware feature encoder fuses the aligned image features with each object's points to generate the discriminative instance-aware features, and the tracking pipeline finally outputs the results based on instance-aware features and G-IoU geometric similarities. Our approach achieves state-of-the-art results on the nuScenes dataset among the methods using CenterPoint detections. The experimental results show that the proposed method has better robustness and effectiveness for 3D multi-object tracking. Sanping Zhou, Jinpeng Dong, Nanning Zheng 0001 |
IJCNN | 2 |
| 2023 | Representing Multimodal Behaviors With Mean Location for Pedestrian Trajectory PredictionabstractRepresenting multimodal behaviors is a critical challenge for pedestrian trajectory prediction. Previous methods commonly represent this multimodality with multiple latent variables repeatedly sampled from a latent space, encountering difficulties in interpretable trajectory prediction. Moreover, the latent space is usually built by encoding global interaction into future trajectory, which inevitably introduces superfluous interactions and thus leads to performance reduction. To tackle these issues, we propose a novel Interpretable Multimodality Predictor (IMP) for pedestrian trajectory prediction, whose core is to represent a specific mode by its mean location. We model the distribution of mean location as a Gaussian Mixture Model (GMM) conditioned on sparse spatio-temporal features, and sample multiple mean locations from the decoupled components of GMM to encourage multimodality. Our IMP brings four-fold benefits: 1) Interpretable prediction to provide semantics about the motion behavior of a specific mode; 2) Friendly visualization to present multimodal behaviors; 3) Well theoretical feasibility to estimate the distribution of mean locations supported by the central-limit theorem; 4) Effective sparse spatio-temporal features to reduce superfluous interactions and model temporal continuity of interaction. Extensive experiments validate that our IMP not only outperforms state-of-the-art methods but also can achieve a controllable prediction by customizing the corresponding mean location. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Memory-augmented appearance-motion network for video anomaly detection
Le Wang 0003, Junwen Tian, Sanping Zhou, Haoyue Shi 0002, Gang Hua 0001 |
Pattern Recognit. | 3 |
| 2023 | Context Adaptive Network for Image InpaintingabstractIn a typical image inpainting task, the location and shape of the damaged or masked area is often random and irregular. The vanilla convolutions widely used in learning-based inpainting models treat all spatial features as valid and share parameters across regions, making it difficult for them to cope with those irregular damages, and models tend to produce inpainting results with color discrepancy and blurriness. In this paper, we propose a novel Context Adaptive Network (CANet) to address this issue. The main idea of the proposed CANet is able to generate different weights depending on the miscellaneous input, which may help to complement images with multiple broken forms in a flexible way. Specifically, the proposed CANet has two novel context adaptive modules, namely, the context adaptive block (CAB) and the cross-scale contextual attention (CSCA), which utilize attention mechanisms to cope with diverse content breakdowns. The proposed CAB, during the forward propagation, uses an adaptive term to determine the importance between adaptive term and convolution kernel, so as to dynamically balance features based on the degree of breakage (confidence level or soft mask), and the overall calculation is formulated as a classic convolution implementation with an additional attention term to describe local structure. Besides, the proposed CSCA, not only takes advantage of the contextual attention module, but also considers cross-scale information transfer to generate reasonable features for damaged areas, thus alleviating the inefficiency of the long-range modeling capability of convolutional neural networks. Qualitative and quantitative experiments show that our method performs better than state-of-the-arts, producing clearer, more coherent and visually plausible inpainting results. The code can be found at github.com/dengyecode/CANet_image_inpainting. Ye Deng 0005, Siqi Hui, Sanping Zhou, Wenli Huang 0004, Jinjun Wang |
IEEE Trans. Image Process. | 3 |
| 2023 | Instance Motion Tendency Learning for Video Panoptic SegmentationabstractVideo panoptic segmentation is an important but challenging task in computer vision. It not only performs panoptic segmentation of each frame, but also associates the same instance across adjacent frames. Due to the lack of temporal coherence modeling, most existing approaches often generate identity switches during instance association, and they cannot handle ambiguous segmentation boundaries caused by motion blur. To address these difficult issues, we introduce a simple yet effective Instance Motion Tendency Network (IMTNet) for video panoptic segmentation. It learns a global motion tendency map for instance association, and a hierarchical classifier for motion boundary refinement. Specifically, a Global Motion Tendency Module (GMTM) is designed to learn robust motion features from optical flows, which can directly associate each instance in the previous frame to the corresponding instance in the current frame. In addition, we propose a Motion Boundary Refinement Module (MBRM) to learn a hierarchical classifier to handle the boundary pixels of moving targets, which can effectively revise the inaccurate segmentation predictions. Experimental results on both Cityscapes and Cityscapes-VPS datasets show that our IMTNet outperforms most state-of-the-art approaches. Le Wang 0003, Hongzhen Liu, Sanping Zhou, Wei Tang 0016, Gang Hua 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Multi-Panda TrackingabstractMulti-Panda Tracking (MPT) is a video-based tracking task for panda individuals, which is conducive to the observation and measurement of distribution and status of pandas. Different from tracking general objects such as pedestrians and vehicles, MPT is extremely challenging due to the indistinguishable appearances and diversified postures of pandas. In this case, existing tracking methods cannot appropriately tackle with the excessive occlusion between different panda individuals, hence suffering from identity switch, missing and inaccurate detections. To address these problems, we propose a simple yet effective MPT framework in the tracking-by-detection paradigm, which is benefited both from a short-term prediction filtering module and a discriminative feature learning network. In particular, the short-term prediction filtering module introduces similarity learning to enhance the temporal consistency among detections, which is capable of supplementing the missing detections and discarding false positive detections. Besides, the discriminative feature learning network leverages a two-branch network to learn both local and global discriminative features, so as to distinguish different panda individuals with a very similar appearance with a subtle difference. To evaluate the proposed method, we annotate a large-scale MPT dataset, named PANDA2021, which is particularly challenging due to the similar appearance and dramatic occlusion between panda individuals. Experiments on PANDA2021 demonstrate that the proposed MPT method significantly outperforms the competing methods. Moreover, experimental results on pedestrian tracking dataset MOT16 further demonstrate that the proposed MPT method achieves comparative performance with competing methods. Le Wang 0003, Sanping Zhou, Nanning Zheng 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Exploring Action Centers for Temporal Action LocalizationabstractTemporal action localization aims at detecting the temporal intervals of human actions in untrimmed videos. Most previous methods rely on locating and matching the start and end times of actions. However, action boundaries are ambiguous and uncertain in nature, which leads to inaccurate action localization and a lot of false positives. In this paper, we introduce a new framework for temporal action localization. It explicitly models temporal action centers to reduce unreliable action detection results caused by ambiguous action boundaries. Since action centers are highly related to semantic actions, they can be detected more reliably than the conventional action boundaries. As a result, our framework can exclude false positives and promote high-quality proposals. Based on action centers, we propose a triplet feature fusion mechanism. It performs neural message passing among the boundaries and the center as well as contextual regions outside of the proposal to enrich its representation. In addition, we introduce a centerness scoring method to suppress proposals deviating from the centers of action instances. Consequently, our network can retrieve high-quality action proposals and locate actions more precisely. Experimental results show our method outperforms state-of-the-art methods on the THUMOS14 and ActivityNet v1.3 datasets. Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
IEEE Trans. Multim. | 4 |
| 2022 | Complementary Attention Gated Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is crucial in many practical applications due to the diversity of pedestrian movements, such as social interactions and individual motion behaviors. With similar observable trajectories and social environments, different pedestrians may make completely different future decisions. However, most existing methods only focus on the frequent modal of the trajectory and thus are difficult to generalize to the peculiar scenario, which leads to the decline of the multimodal fitting ability when facing similar scenarios. In this paper, we propose a complementary attention gated network (CAGN) for pedestrian trajectory prediction, in which a dual-path architecture including normal and inverse attention is proposed to capture both frequent and peculiar modals in spatial and temporal patterns, respectively. Specifically, a complementary block is proposed to guide normal and inverse attention, which are then be summed with learnable weights to get attention features by a gated network. Finally, multiple trajectory distributions are estimated based on the fused spatio-temporal attention features due to the multimodality of future trajectory. Experimental results on benchmark datasets, i.e., the ETH, and the UCY, demonstrate that our method outperforms state-of-the-art methods by 13.8% in Average Displacement Error (ADE) and 10.4% in Final Displacement Error (FDE). Code will be available at https://github.com/jinghaiD/CAGN Jinghai Duan, Le Wang 0003, Chengjiang Long, Sanping Zhou, Fang Zheng 0009, Liushuai Shi, Gang Hua 0001 |
AAAI | 4 |
| 2022 | Social Interpretable Tree for Pedestrian Trajectory PredictionabstractUnderstanding the multiple socially-acceptable future behaviors is an essential task for many vision applications. In this paper, we propose a tree-based method, termed as Social Interpretable Tree (SIT), to address this multi-modal prediction task, where a hand-crafted tree is built depending on the prior information of observed trajectory to model multiple future trajectories. Specifically, a path in the tree from the root to leaf represents an individual possible future trajectory. SIT employs a coarse-to-fine optimization strategy, in which the tree is first built by high-order velocity to balance the complexity and coverage of the tree and then optimized greedily to encourage multimodality. Finally, a teacher-forcing refining operation is used to predict the final fine trajectory. Compared with prior methods which leverage implicit latent variables to represent possible future trajectories, the path in the tree can explicitly explain the rough moving behaviors (e.g., go straight and then turn right), and thus provides better interpretability. Despite the hand-crafted tree, the experimental results on ETH-UCY and Stanford Drone datasets demonstrate that our method is capable of matching or exceeding the performance of state-of-the-art methods. Interestingly, the experiments show that the raw built tree without training outperforms many prior deep neural network based approaches. Meanwhile, our method presents sufficient flexibility in long-term prediction and different best-of-K predictions. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Fang Zheng 0009, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 4 |
| 2022 | TransVPR: Transformer-Based Place Recognition with Multi-Level Attention AggregationabstractVisual place recognition is a challenging task for applications such as autonomous driving navigation and mobile robot localization. Distracting elements presenting in complex scenes often lead to deviations in the perception of visual place. To address this problem, it is crucial to integrate information from only task-relevant regions into image representations. In this paper, we introduce a novel holistic place recognition model, TransVPR, based on vision Transformers. It benefits from the desirable property of the self-attention operation in Transformers which can naturally aggregate task-relevant features. Attentions from multiple levels of the Transformer, which focus on different regions of interest, are further combined to generate a global image representation. In addition, the output tokens from Transformer layers filtered by the fused attention mask are considered as key-patch descriptors, which are used to perform spatial matching to re-rank the candidates retrieved by the global image features. The whole model allows end-to-end training with a single objective and image-level supervision. TransVPR achieves state-of-the-art performance on several real-world benchmarks while maintaining low computational time and storage requirements. Ruotong Wang 0005, Yanqing Shen, Weiliang Zuo, Sanping Zhou, Nanning Zheng 0001 |
CVPR | 4 |
| 2022 | Learning to Refactor Action and Co-occurrence Features for Temporal Action LocalizationabstractThe main challenge of Temporal Action Localization is to retrieve subtle human actions from various co-occurring ingredients, e.g., context and background, in an untrimmed video. While prior approaches have achieved substantial progress through devising advanced action detectors, they still suffer from these co-occurring ingredients which often dominate the actual action content in videos. In this paper, we explore two orthogonal but complementary aspects of a video snippet, i.e., the action features and the co-occurrence features. Especially, we develop a novel auxiliary task by decoupling these two types of features within a video snippet and recombining them to generate a new feature representation with more salient action information for accurate action localization. We term our method RefactorNet, which first explicitly factorizes the action content and regularizes its co-occurrence features, and then synthesizes a new action-dominated video representation. Extensive experimental results and ablation studies on THUMOS14 and ActivityNet v 1.3 demonstrate that our new representation, combined with a simple action detector, can significantly improve the action localization performance. Le Wang 0003, Sanping Zhou, Nanning Zheng 0001, Wei Tang 0016 |
CVPR | 3 |
| 2022 | Hourglass Attention Network for Image Inpainting
Ye Deng 0005, Siqi Hui, Rongye Meng, Sanping Zhou, Jinjun Wang |
ECCV (18) | 4 |
| 2022 | Pedestrian Intention Prediction Based on Traffic-Aware Scene Graph ModelabstractAnticipating the future behavior of pedestrians is a crucial part of deploying Automated Driving Systems (ADS) in urban traffic scenarios. Most recent works utilize a convolutional neural network (CNN) to extract visual information, which is then input to a recurrent neural network (RNN) along with pedestrian-specific features like location and speed to obtain temporal features. However, the majority of these approaches lack the ability to parse the relationships of the related objects in the specific traffic scene, which leads to omitting the interactions between the pedestrians and the interactions between the pedestrians and the traffic. For this purpose, we propose a graph-structured model which can dig out pedestrians' dynamic constraints by constructing a traffic-aware scene graph within each frame. In addition, to capture pedestrian movement more effectively, we also introduce a temporal feature representation model, which first uses inter-frame and intra-frame GRU (II-GRU) to mine inter-frame information and intra-frame information together, and then employs a novel attention mechanism to adaptively generate attention weights. Extensive experiments on the JAAD and PIE datasets prove that our proposed model is effective in reaching and enhancing the state-of-the-art performance. Xingchen Song, Miao Kang, Sanping Zhou, Jianji Wang 0001, Yishu Mao 0003, Nanning Zheng 0001 |
IROS | 3 |
| 2022 | T-former: An Efficient Transformer for Image InpaintingabstractBenefiting from powerful convolutional neural networks (CNNs), learning-based image inpainting methods have made significant breakthroughs over the years. However, some nature of CNNs (e.g. local prior, spatially shared parameters) limit the performance in the face of broken images with diverse and complex forms. Recently, a class of attention-based network architectures, called transformer, has shown significant performance on natural language processing fields and high-level vision tasks. Compared with CNNs, attention operators are better at long-range modeling and have dynamic weights, but their computational complexity is quadratic in spatial resolution, and thus less suitable for applications involving higher resolution images, such as image inpainting. In this paper, we design a novel attention linearly related to the resolution according to Taylor expansion. And based on this attention, a network called T-former is designed for image inpainting. Experiments on several benchmark datasets demonstrate that our proposed method achieves state-of-the-art accuracy while maintaining a relatively low number of parameters and computational complexity. Ye Deng 0005, Siqi Hui, Sanping Zhou, Deyu Meng, Jinjun Wang |
ACM Multimedia | 3 |
| 2022 | Learning to predict diverse trajectory from human motion patterns
Miao Kang, Jingwen Fu, Sanping Zhou, Songyi Zhang, Nanning Zheng 0001 |
Neurocomputing | 3 |
| 2022 | Dual relation network for temporal action localization
Le Wang 0003, Sanping Zhou, Gang Hua 0001, Wei Tang 0016 |
Pattern Recognit. | 3 |
| 2022 | Local to Global Feature Learning for Salient Object Detection
Xuelu Feng, Sanping Zhou, Zixin Zhu, Le Wang 0003, Gang Hua 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | A Novel Hybrid Level Set Model for Non-Rigid Object Contour TrackingabstractMost existing trackers use bounding boxes for object tracking. However, the background contained in the bounding box inevitably decreases the accuracy of the target model, which affects the performance of the tracker and is particularly pronounced for non-rigid objects. To address the above issue, this paper proposes a novel hybrid level set model, which can robustly address the issue of topology changing, occlusions and abrupt motion in non-rigid object tracking by accurately tracking the object contour. In particular, an appearance model is first obtained by repeatedly training and relabeling the initial labeled frame using competing one-class SVMs. Then, by integrating the trained appearance model, an edge detector and image spatial information into the level set model, a new hybrid level set model is presented, which accurately locates the object contour and feeds back to the competing one-class SVMs to update the appearance model of the next frame. In addition, a motion model is defined to predict the accurate location of the object when occlusion and abrupt motion occur in the next frame. Finally, the experimental results on state-of-the-art benchmarks demonstrate the feasibility and effectiveness of the proposed model and the superiority of the proposed method over existing trackers in terms of accuracy and robustness. Yiming Qian, Sanping Zhou, Jinjun Wang, Yee-Hong Yang |
IEEE Trans. Image Process. | 4 |
| 2022 | AVLSM: Adaptive Variational Level Set Model for Image Segmentation in the Presence of Severe Intensity Inhomogeneity and High NoiseabstractIntensity inhomogeneity and noise are two common issues in images but inevitably lead to significant challenges for image segmentation and is particularly pronounced when the two issues simultaneously appear in one image. As a result, most existing level set models yield poor performance when applied to this images. To this end, this paper proposes a novel hybrid level set model, named adaptive variational level set model (AVLSM) by integrating an adaptive scale bias field correction term and a denoising term into one level set framework, which can simultaneously correct the severe inhomogeneous intensity and denoise in segmentation. Specifically, an adaptive scale bias field correction term is first defined to correct the severe inhomogeneous intensity by adaptively adjusting the scale according to the degree of intensity inhomogeneity while segmentation. More importantly, the proposed adaptive scale truncation function in the term is model-agnostic, which can be applied to most off-the-shelf models and improves their performance for image segmentation with severe intensity inhomogeneity. Then, a denoising energy term is constructed based on the variational model, which can remove not only common additive noise but also multiplicative noise often occurred in medical image during segmentation. Finally, by integrating the two proposed energy terms into a variational level set framework, the AVLSM is proposed. The experimental results on synthetic and real images demonstrate the superiority of AVLSM over most state-of-the-art level set models in terms of accuracy, robustness and running time. Yiming Qian, Sanping Zhou, Jinxing Li 0003, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Multinetwork Collaborative Feature Learning for Semisupervised Person ReidentificationabstractPerson reidentification (Re-ID) aims at matching images of the same identity captured from the disjoint camera views, which remains a very challenging problem due to the large cross-view appearance variations. In practice, the mainstream methods usually learn a discriminative feature representation using a deep neural network, which needs a large number of labeled samples in the training process. In this article, we design a simple yet effective multinetwork collaborative feature learning (MCFL) framework to alleviate the data annotation requirement for person Re-ID, which can confidently estimate the pseudolabels of unlabeled sample pairs and consistently learn the discriminative features of input images. To keep the precision of pseudolabels, we further build a novel self-paced collaborative regularizer to extensively exchange the weight information of unlabeled sample pairs between different networks. Once the pseudolabels are correctly estimated, we take the corresponding sample pairs into the training process, which is beneficial to learn more discriminative features for person Re-ID. Extensive experimental results on the Market1501, DukeMTMC, and CUHK03 data sets have shown that our method outperforms most of the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Deyu Meng, Le Wang 0003, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | SGCN: Sparse Graph Convolution Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a key technology in autopilot, which remains to be very challenging due to complex interactions between pedestrians. However, previous works based on dense undirected interaction suffer from modeling superfluous interactions and neglect of trajectory motion tendency, and thus inevitably result in a considerable deviance from the reality. To cope with these issues, we present a Sparse Graph Convolution Network (SGCN) for pedestrian trajectory prediction. Specifically, the SGCN explicitly models the sparse directed interaction with a sparse directed spatial graph to capture adaptive interaction pedestrians. Meanwhile, we use a sparse directed temporal graph to model the motion tendency, thus to facilitate the prediction based on the observed direction. Finally, parameters of a bi-Gaussian distribution for trajectory prediction are estimated by fusing the above two sparse graphs. We evaluate our proposed method on the ETH and UCY datasets, and the experimental results show our method outperforms comparative state-of-the-art methods by 9% in Average Displacement Error (ADE) and 13% in Final Displacement Error (FDE). Notably, visualizations indicate that our method can capture adaptive interactions between pedestrians and their effective motion tendencies. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Zhenxing Niu, Gang Hua 0001 |
CVPR | 4 |
| 2021 | Meta Pairwise Relationship Distillation for Unsupervised Person Re-identificationabstractUnsupervised person re-identification (Re-ID) remains challenging due to the lack of ground-truth labels. Existing methods often rely on estimated pseudo labels via iterative clustering and classification, and they are unfortunately highly susceptible to performance penalties incurred by the inaccurate estimated number of clusters. Alternatively, we propose the Meta Pairwise Relationship Distillation (MPRD) method to estimate the pseudo labels of sample pairs for unsupervised person Re-ID. Specifically, it consists of a Convolutional Neural Network (CNN) and Graph Convolutional Network (GCN), in which the GCN estimates the pseudo labels of sample pairs based on the current features extracted by CNN, and the CNN learns better features by involving high-fidelity positive and negative sample pairs imposed by GCN. To achieve this goal, a small amount of labeled samples are used to guide GCN training, which can distill meta knowledge to judge the difference in the neighborhood structure between positive and negative sample pairs. Extensive experiments on Market-1501, DukeMTMC-reID and MSMT17 datasets show that our method outperforms the state-of-the-art approaches. Haoxuanye Ji, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 3 |
| 2021 | Unlimited Neighborhood Interaction for Heterogeneous Trajectory PredictionabstractUnderstanding complex social interactions among agents is a key challenge for trajectory prediction. Most existing methods consider the interactions between pairwise traffic agents or in a local area, while the nature of interactions is unlimited, involving an uncertain number of agents and non-local areas simultaneously. Besides, they treat heterogeneous traffic agents the same, namely those among agents of different categories, while neglecting people’s diverse reaction patterns toward traffic agents in different categories. To address these problems, we propose a simple yet effective Unlimited Neighborhood Interaction Network (UNIN), which predicts trajectories of heterogeneous agents in multiple categories. Specifically, the proposed unlimited neighborhood interaction module generates the fused-features of all agents involved in an interaction simultaneously, which is adaptive to any number of agents and any range of interaction area. Meanwhile, a hierarchical graph attention module is proposed to obtain category-to-category interaction and agent-to-agent interaction. Finally, parameters of a Gaussian Mixture Model are estimated for generating the future trajectories. Extensive experimental results on benchmark datasets demonstrate a significant performance improvement of our method over the state-of-the-art methods. Fang Zheng 0009, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 3 |
| 2021 | Learning Generic Feature Representations with Adversarial Regularization for Person Re-IdentificationabstractMany existing person re-identification (Re-ID) methods can achieve human-level accuracy on a single dataset, while most of them can be poorly generalized to other datasets. This is mainly caused by different data distributions between different domains. In this paper, we propose a novel adversarial regularization method to address this issue. Specifically, the features extracted from different datasets will be constrained and focused to follow a more similar distribution during the training process. As a result, our method can learn a feature representation with better inter-domain invariance, which will improve the generalization ability of the resulting model. Besides, our method is flexible and can be combined with any feature learning network. Extensive experiments on both Market1501 and DukeMTMC-reID datasets have demonstrated the effectiveness of our method. Qindong Zhang, Sanping Zhou, Jinjun Wang |
ICIP | 2 |
| 2021 | Learning Contextual Transformer Network for Image InpaintingabstractFully Convolutional Networks with attention modules have been proven effective for learning-based image inpainting. While many existing approaches could produce visually reasonable results, the generated images often show blurry textures or distorted structures around corrupted areas. The main reason is due to the fact that convolutional neural networks have limited capacity for modeling contextual information with long range dependencies. Although the attention mechanism can alleviate this problem to some extent, existing attention modules tend to emphasize similarities between the corrupted and the uncorrupted regions while ignoring the dependencies from within each of them. Hence, this paper proposes the Contextual Transformer Network (CTN) which not only learns relationships between the corrupted and the uncorrupted regions but also exploits their respective internal closeness. Besides, instead of a fully convolutional network, in our CTN, we stack several transformer blocks to replace convolution layers to better model the long range dependencies. Finally, by dividing the image into patches of different sizes, we propose a multi-scale multi-head attention module to better model the affinity among various image regions. Experiments on several benchmark datasets demonstrate superior performance by our proposed approach. Ye Deng 0005, Siqi Hui, Sanping Zhou, Deyu Meng, Jinjun Wang |
ACM Multimedia | 3 |
| 2021 | Multiple Object Tracking by Trajectory Map Regression with Temporal Priors EmbeddingabstractPrevailing Multiple Object Tracking (MOT) works following the Tracking-by-Detection (TBD) paradigm pay most attention to either object detection in a first step or data association in a second step. In this paper, we approach the MOT problem from a different perspective by directly obtaining the embedded spatial-temporal information of trajectories from raw video data. For the purpose we propose a joint trajectory locating and attributes encoding framework for real-time, on-line MOT. We firstly introduce a trajectory attribute representation scheme designed for each tracked target (instead of object) where the extracted Trajectory Map (TM) encodes the spatial-temporal attributes of a trajectory across a window of consecutive video frames. Next we present a Temporal Priors Embedding (TPE) methodology to infer these attributes with a logical reasoning strategy based on long-term feature dynamics. The proposed MOT framework projects multiple attributes of tracked targets, e.g., presence, enter/exit, location, scale, motion, etc. into a continuous TM to perform one-shot regression for real-time MOT. Experimental results show that, our proposed video-based method runs at 33 FPS and is more accurate and robust as compared to the detection-based tracking methods and a few other State-of-the- Art (SOTA) approaches on MOT16/17/20 benchmarks. Xingyu Wan, Sanping Zhou, Jinjun Wang, Rongye Meng |
ACM Multimedia | 2 |
| 2021 | Single-Image super-resolution - When model adaptation matters
Yudong Liang, Radu Timofte, Jinjun Wang, Sanping Zhou, Yihong Gong, Nanning Zheng 0001 |
Pattern Recognit. | 4 |
| 2021 | Predicting Task-Driven Attention via Integrating Bottom-Up Stimulus and Top-Down GuidanceabstractTask-free attention has gained intensive interest in the computer vision community while relatively few works focus on task-driven attention (TDAttention). Thus this paper handles the problem of TDAttention prediction in daily scenarios where a human is doing a task. Motivated by the cognition mechanism that human attention allocation is jointly controlled by the top-down guidance and bottom-up stimulus, this paper proposes a cognitively-explanatory deep neural network model to predict TDAttention. Given an image sequence, bottom-up features, such as human pose and motion, are firstly extracted. At the same time, the coarse-grained task information and fine-grained task information are embedded as a top-down feature. The bottom-up features are then fused with the top-down feature to guide the model to predict TDAttention. Two public datasets are re-annotated to make them qualified for TDAttention prediction, and our model is widely compared with other models on the two datasets. In addition, some ablation studies are conducted to evaluate the individual modules in our model. Experiment results demonstrate the effectiveness of our model. Zhixiong Nan, Jingjing Jiang, Xiaofeng Gao 0002, Sanping Zhou, Weiliang Zuo, Ping Wei 0001, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Tracking Beyond Detection: Learning a Global Response Map for End-to-End Multi-Object TrackingabstractMost of the existing Multi-Object Tracking (MOT) approaches follow the Tracking-by-Detection and Data Association paradigm, in which objects are firstly detected and then associated in the tracking process. In recent years, deep neural network has been utilized to obtain more discriminative appearance features for cross-frame association, and noticeable performance improvement has been reported. On the other hand, the Tracking-by-Detection framework is yet not completely end-to-end, which leads to huge computation and limited performance especially in the inference (tracking) process. To address this problem, we present an effective end-to-end deep learning framework which can directly take image-sequence/video as input and output the located and tracked objects of learned types. Specifically, a novel global response network is learned to project multiple objects in the image-sequence/video into a continuous response map, and the trajectory of each tracked object can then be easily picked out. The overall process is similar to how a detector inputs an image and outputs the bounding boxes of each detected object. Experimental results based on the MOT16 and MOT17 benchmarks show that our proposed on-line tracker achieves state-of-the-art performance on several tracking metrics. Xingyu Wan, Jiakai Cao, Sanping Zhou, Jinjun Wang, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Hierarchical and Interactive Refinement Network for Edge-Preserving Salient Object DetectionabstractSalient object detection has undergone a very rapid development with the blooming of Deep Neural Network (DNN), which is usually taken as an important preprocessing procedure in various computer vision tasks. However, the down-sampling operations, such as pooling and striding, always make the final predictions blurred at edges, which has seriously degenerated the performance of salient object detection. In this paper, we propose a simple yet effective approach, i.e., Hierarchical and Interactive Refinement Network (HIRN), to preserve the edge structures in detecting salient objects. In particular, a novel multi-stage and dual-path network structure is designed to estimate the salient edges and regions from the low-level and high-level feature maps, respectively. As a result, the predicted regions will become more accurate by enhancing the weak responses at edges, while the predicted edges will become more semantic by suppressing the false positives in background. Once the salient maps of edges and regions are obtained at the output layers, a novel edge-guided inference algorithm is introduced to further filter the resulting regions along the predicted edges. Extensive experiments on several benchmark datasets have been conducted, in which the results show that our method significantly outperforms a variety of state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Le Wang 0003, Jimuyang Zhang, Fei Wang 0037, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Meta Corrupted Pixels Mining for Medical Image Segmentation
Sanping Zhou, Chaowei Fang, Le Wang 0003, Jinjun Wang |
MICCAI (1) | 2 |
| 2020 | Hierarchical U-Shape Attention Network for Salient Object DetectionabstractSalient object detection aims at locating the most conspicuous objects in natural images, which usually acts as a very important pre-processing procedure in many computer vision tasks. In this paper, we propose a simple yet effective Hierarchical U-shape Attention Network (HUAN) to learn a robust mapping function for salient object detection. Firstly, a novel attention mechanism is formulated to improve the well-known U-shape network [1], in which the memory consumption can be extensively reduced and the mask quality can be significantly improved by the resulting U-shape Attention Network (UAN). Secondly, a novel hierarchical structure is constructed to well bridge the low-level and high-level feature representations between different UANs, in which both the intra-network and inter-network connections are considered to explore the salient patterns from a local to global view. Thirdly, a novel Mask Fusion Network (MFN) is designed to fuse the intermediate prediction results, so as to generate a salient mask which is in higher-quality than any of those inputs. Our HUAN can be trained together with any backbone network in an end-to-end manner, and high-quality masks can be finally learned to represent the salient objects. Extensive experimental results on several benchmark datasets show that our method significantly outperforms most of the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Jimuyang Zhang, Le Wang 0003, Shaoyi Du, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Person-in-WiFi: Fine-Grained Person Perception Using WiFiabstractFine-grained person perception such as body segmentation and pose estimation has been achieved with many 2D and 3D sensors such as RGB/depth cameras, radars (e.g. RF-Pose), and LiDARs. These solutions require 2D images, depth maps or 3D point clouds of person bodies as input. In this paper, we take one step forward to show that fine-grained person perception is possible even with 1D sensors: WiFi antennas. Specifically, we used two sets of WiFi antennas to acquire signals, i.e., one transmitter set and one receiver set. Each set contains three antennas horizontally lined-up as a regular household WiFi router. The WiFi signal generated by a transmitter antenna, penetrates through and reflects on human bodies, furniture, and walls, and then superposes at a receiver antenna as 1D signal samples. We developed a deep learning approach that uses annotations on 2D images, takes the received 1D WiFi signals as input, and performs body segmentation and pose estimation in an end-to-end manner. To our knowledge, our solution is the first work based on off-the-shelf WiFi antennas and standard IEEE 802.11n WiFi signals. Demonstrating comparable results to image-based solutions, our WiFi-based person perception solution is cheaper and more ubiquitous than radars and LiDARs, while invariant to illumination and has little privacy concern comparing to cameras. Fei Wang 0037, Sanping Zhou, Stanislav Panev, Jinsong Han |
ICCV | 2 |
| 2019 | Discriminative Feature Learning With Consistent Attention Regularization for Person Re-IdentificationabstractPerson re-identification (Re-ID) has undergone a rapid development with the blooming of deep neural network. Most methods are very easily affected by target misalignment and background clutter in the training process. In this paper, we propose a simple yet effective feedforward attention network to address the two mentioned problems, in which a novel consistent attention regularizer and an improved triplet loss are designed to learn foreground attentive features for person Re-ID. Specifically, the consistent attention regularizer aims to keep the deduced foreground masks similar from the low-level, mid-level and high-level feature maps. As a result, the network will focus on the foreground regions at the lower layers, which is benefit to learn discriminative features from the foreground regions at the higher layers. Last but not least, the improved triplet loss is introduced to enhance the feature learning capability, which can jointly minimize the intra-class distance and maximize the inter-class distance in each triplet unit. Experimental results on the Market1501, DukeMTMC-reID and CUHK03 datasets have shown that our method outperforms most of the state-of-the-art approaches. Sanping Zhou, Fei Wang 0037, Zeyi Huang, Jinjun Wang |
ICCV | 1 |
| 2019 | Meta-Weight-Net: Learning an Explicit Mapping For Sample WeightingabstractCurrent deep neural networks(DNNs) can easily overfit to biased training data with corrupted labels or class imbalance. Sample re-weighting strategy is commonly used to alleviate this issue by designing a weighting function mapping from training loss to sample weight, and then iterating between weight recalculating and classifier updating. Current approaches, however, need manually pre-specify the weighting function as well as its additional hyper-parameters. It makes them fairly hard to be generally applied in practice due to the significant variation of proper weighting schemes relying on the investigated problem and training data. To address this issue, we propose a method capable of adaptively learning an explicit weighting function directly from data. The weighting function is an MLP with one hidden layer, constituting a universal approximator to almost any continuous functions, making the method able to fit a wide range of weighting function forms including those assumed in conventional research. Guided by a small amount of unbiased meta-data, the parameters of the weighting function can be finely updated simultaneously with the learning process of the classifiers. Synthetic and real experiments substantiate the capability of our method for achieving proper weighting functions in class imbalance and noisy label cases, fully complying with the common settings in traditional methods, and more complicated scenarios beyond conventional cases. This naturally leads to its better accuracy than other state-of-the-art methods. Qi Xie 0002, Lixuan Yi, Qian Zhao 0002, Sanping Zhou, Zongben Xu, Deyu Meng |
NeurIPS | 5 |
| 2019 | Saliency-guided level set model for automatic object segmentation
Yiming Qian, Sanping Zhou, Xiaojun Duan, Yee-Hong Yang |
Pattern Recognit. | 4 |
| 2019 | Semi-supervised person re-identification using multi-view clustering
Xiaomeng Xin, Jinjun Wang, Ruji Xie, Sanping Zhou, Wenli Huang 0004, Nanning Zheng 0001 |
Pattern Recognit. | 4 |
| 2019 | Discriminative Feature Learning With Foreground Attention for Person Re-IdentificationabstractThe performance of person re-identification (Re-ID) has been seriously affected by the large cross-view appearance variations caused by mutual occlusions and background clutter. Hence, learning a feature representation that can adaptively emphasize the foreground persons becomes very critical to solve the person Re-ID problem. In this paper, we propose a simple yet effective foreground attentive neural network (FANN) to learn a discriminative feature representation for person Re-ID, which can adaptively enhance the positive side of foreground and weaken the negative side of background. Specifically, a novel foreground attentive subnetwork is designed to drive the network’s attention, in which a decoder network is used to reconstruct the binary mask by using a novel local regression loss function, and an encoder network is regularized by the decoder network to focus its attention on the foreground persons. The resulting feature maps of encoder network are further fed into the body part subnetwork and feature fusion subnetwork to learn discriminative features. Besides, a novel symmetric triplet loss function is introduced to supervise feature learning, in which the intra-class distance is minimized and the inter-class distance is maximized in each triplet unit, simultaneously. Training our FANN in a multi-task learning framework, a discriminative feature representation can be learned to find out the matched reference to each probe among various candidates in the gallery. Extensive experimental results on several public benchmark datasets are evaluated, which have shown clear improvements of our method over the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Deyu Meng, Yudong Liang, Yihong Gong, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Continuous Action Recognition and Segmentation in Untrimmed VideosabstractRecognizing continuous human action is a fundamental task in many real-world computer vision applications including video surveillance, video retrieval, and human-computer interaction, etc. It requires to recognize each action performed as well as their segmentation boundaries in a continuous sequence. In previous works, great progress has been reported for single action recognition, by using deep convolutional networks. In order to further improve the performance for continuous action recognition, in this paper, we introduce a discriminative approach consisting of three modules. The first feature extraction module uses a two stream Convolutional Neural Network to capture the appearance and the short-term motion information from the raw video input. Based on the obtained features, the second classification module performs spatial and temporal recognition and then fuses the two scores from respective feature stream. In the final segmentation module, a semi-Markov Conditional Field model, capable of handling long-term action interactions, is built to partition the action sequence. As can be seen in the experimental results, our approach obtains state-of-the-art performance on public datasets including 50Salads, Breakfast, and MERL Shopping. We have also visualized the continuous actions segmentation results for more insightful discussion in the paper. Ruibin Bai, Sanping Zhou, Xueji Zhao, Jinjun Wang |
ICPR | 3 |
| 2018 | An adaptive-scale active contour model for inhomogeneous image segmentation and bias field estimation
Sanping Zhou, Jingfeng Sun |
Pattern Recognit. | 3 |
| 2018 | Face alignment recurrent network
Qiqi Hou, Jinjun Wang, Ruibin Bai, Sanping Zhou, Yihong Gong |
Pattern Recognit. | 4 |
| 2018 | Deep ranking model by large adaptive margin learning for person re-identification
Sanping Zhou, Jinjun Wang, Qiqi Hou |
Pattern Recognit. | 2 |
| 2018 | Deep self-paced learning for person re-identification
Sanping Zhou, Jinjun Wang, Deyu Meng, Xiaomeng Xin, Yihong Gong, Nanning Zheng 0001 |
Pattern Recognit. | 1 |
| 2018 | Large Margin Learning in Set-to-Set Similarity Comparison for Person ReidentificationabstractPerson reidentification aims at matching images of the same person across disjoint camera views, which is a challenging problem in multimedia analysis, multimedia editing, and content-based media retrieval communities. The major challenge lies in how to preserve similarity of the same person across video footages with large appearance variations, while discriminating different individuals. To address this problem, conventional methods usually consider the pairwise similarity between persons by only measuring the point-to-point distance. In this paper, we propose using a deep learning technique to model a novel set-to-set (S2S) distance, in which the underline objective focuses on preserving the compactness of intraclass samples for each camera view, while maximizing the margin between the intraclass set and interclass set. The S2S distance metric consists of three terms, namely, the class-identity term, the relative distance term, and the regularization term. The class-identity term keeps the intraclass samples within each camera view gathering together, the relative distance term maximizes the distance between the intraclass class set and interclass set across different camera views, and the regularization term smoothes the parameters of the deep convolutional neural network. As a result, the final learned deep model can effectively find out the matched target to the probe object among various candidates in the video gallery by learning discriminative and stable feature representations. Using the CUHK01, CUHK03, PRID2011, and Market1501 benchmark datasets, we extensively conducted comparative evaluations to demonstrate the advantages of our method over the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Qiqi Hou, Yihong Gong, Nanning Zheng 0001 |
IEEE Trans. Multim. | 1 |
| 2017 | Point to Set Similarity Based Deep Feature Learning for Person Re-IdentificationabstractPerson re-identification (Re-ID) remains a challenging problem due to significant appearance changes caused by variations in view angle, background clutter, illumination condition and mutual occlusion. To address these issues, conventional methods usually focus on proposing robust feature representation or learning metric transformation based on pairwise similarity, using Fisher-type criterion. The recent development in deep learning based approaches address the two processes in a joint fashion and have achieved promising progress. One of the key issues for deep learning based person Re-ID is the selection of proper similarity comparison criteria, and the performance of learned features using existing criterion based on pairwise similarity is still limited, because only P2P distances are mostly considered. In this paper, we present a novel person Re-ID method based on P2S similarity comparison. The P2S metric can jointly minimize the intra-class distance and maximize the inter-class distance, while back-propagating the gradient to optimize parameters of the deep model. By utilizing our proposed P2S metric, the learned deep model can effectively distinguish different persons by learning discriminative and stable feature representations. Comprehensive experimental evaluations on 3DPeS, CUHK01, PRID2011 and Market1501 datasets demonstrate the advantages of our method over the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Yihong Gong, Nanning Zheng 0001 |
CVPR | 1 |
| 2017 | Correntropy-based level set method for medical image segmentation and bias correction
Sanping Zhou, Jinjun Wang, Yihong Gong |
Neurocomputing | 1 |
| 2016 | Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss FunctionabstractPerson re-identification across cameras remains a very challenging problem, especially when there are no overlapping fields of view between cameras. In this paper, we present a novel multi-channel parts-based convolutional neural network (CNN) model under the triplet framework for person re-identification. Specifically, the proposed CNN model consists of multiple channels to jointly learn both the global full-body and local body-parts features of the input persons. The CNN model is trained by an improved triplet loss function that serves to pull the instances of the same person closer, and at the same time push the instances belonging to different persons farther from each other in the learned feature space. Extensive comparative evaluations demonstrate that our proposed method significantly outperforms many state-of-the-art approaches, including both traditional and deep network-based ones, on the challenging i-LIDS, VIPeR, PRID2011 and CUHK01 datasets. De Cheng, Yihong Gong, Sanping Zhou, Jinjun Wang, Nanning Zheng 0001 |
CVPR | 3 |
| 2016 | Incorporating image priors with deep convolutional neural networks for image super-resolution
Yudong Liang, Jinjun Wang, Sanping Zhou, Yihong Gong, Nanning Zheng 0001 |
Neurocomputing | 3 |
| 2016 | Active contour model based on local and global intensity information for medical image segmentation
Sanping Zhou, Jinjun Wang, Yudong Liang, Yihong Gong |
Neurocomputing | 1 |