VLDB 2026 Research / reviewers in the wild / expert
Dong Wang 0004
dblp:40/3934-4
· DBLP profile ↗
141ranked-venue papers
26as first author
69since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 94 · 16 first-author · 43 since 2021Artificial intelligence and machine learning · 85 · 12 first-author · 48 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingabstractRGB-Thermal (RGBT) tracking aims to exploit visible and thermal infrared modalities for robust all-weather object tracking. However, existing RGBT trackers struggle to resolve modality discrepancies, which poses great challenges for robust feature representation. This limitation hinders effective cross-modal information propagation and fusion, which significantly reduces the tracking accuracy. To address this limitation, we propose a novel Contextual Aggregation with Deformable Alignment framework called CADTrack for RGBT Tracking. To be specific, we first deploy the Mamba-based Feature Interaction (MFI) that establishes efficient feature interaction via state space models. This interaction module can operate with linear complexity, reducing computational cost and improving feature discrimination. Then, we propose the Contextual Aggregation Module (CAM) that dynamically activates backbone layers through sparse gating based on the Mixture-of-Experts (MoE). This module can encode complementary contextual information from cross-layer features. Finally, we propose the Deformable Alignment Module (DAM) to integrate deformable sampling and temporal propagation, mitigating spatial misalignment and localization drift. With the above components, our CADTrack achieves robust and accurate tracking in complex scenarios. Extensive experiments on five RGBT tracking benchmarks verify the effectiveness of our proposed method. Hao Li 0101, Xiantao Hu, Wenning Hao, Dong Wang 0004, Huchuan Lu |
AAAI | 6 |
| 2026 | One Aligned LLM to Serve Them All: A Transfer Recipe for Training VLMs without Visual-Language Re-Alignment
Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Wei Zhou 0021, Dong Wang 0004, Huchuan Lu, You He 0002 |
Int. J. Comput. Vis. | 6 |
| 2026 | What Makes You Unique? Attribute Prompt Composition for Object Re-Identification
Yingquan Wang, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | DiMuS: Disentangled Multi-Signal Learning for Weakly Supervised Point-Based 3D Object DetectionabstractWeakly supervised 3D object detection has emerged as a promising paradigm to reduce the reliance on costly 3D annotations. Existing methods often rely on 2D projection constraints or heuristic priors to supervise 3D box regression with inexpensive 2D labels. However, they still suffer from projection ambiguity and geometry inconsistency due to the entangled optimization of 3D parameters. In this paper, we propose DiMuS, a Disentangled Multi- $\boldsymbol {S}$ ignal learning framework that integrates complementary supervision from 2D boxes, LLM-derived semantic prior, and 3D geometric alignment to enhance distinct 3D properties of position, dimension, and orientation, respectively. Specifically, DiMuS incorporates three key components: (i) a Centerness-enhanced Projection Constraint (CPC) that improves position estimation through a centerness weighting strategy, (ii) a Semantic Prior Anchoring (SPA) module that leverages LLM-derived category-specific priors for robust dimension decoding, and (iii) a Rotation-aware Consistency Regularization (RCR) mechanism that enforces orientation consistency through synthetic rotations and self-supervised invariance learning. Additionally, an Adversarial Geometric Alignment (AGA) module is proposed to build attraction/repulsion forces between LiDAR points and box edges for dynamic boundary refinement. Extensive experiments on the KITTI dataset demonstrate that DiMuS outperforms previous weakly supervised methods, achieving 96.82% of fully supervised performance on car detection while maintaining robustness across different categories. Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Image Process. | 5 |
| 2025 | SUTrack: Towards Simple and Unified Single Object TrackingabstractIn this paper, we propose a simple yet unified single object tracking (SOT) framework, dubbed SUTrack. It consolidates five SOT tasks (RGB-based, RGB-Depth, RGB-Thermal, RGB-Event, RGB-Language Tracking) into a unified model trained in a single session. Due to the distinct nature of the data, current methods typically design individual architectures and train separate models for each task. This fragmentation results in redundant training processes, repetitive technological innovations, and limited cross-modal knowledge sharing. In contrast, SUTrack demonstrates that a single model with a unified input representation can effectively handle various SOT tasks, eliminating the need for task-specific designs and separate training sessions. Additionally, we introduce a task-recognition training strategy and a soft token type embedding to further enhance SUTrack's performance with minimal overhead. Experiments show that SUTrack outperforms previous task-specific counterparts across 11 datasets spanning five SOT tasks. Moreover, we provide a range of models catering edge devices as well as high-performance GPUs, striking a good trade-off between speed and accuracy. We hope SUTrack could serve as a strong foundation for further compelling research into unified tracking models. Xin Chen 0032, Ben Kang, Wanting Geng, Jiawen Zhu 0003, Dong Wang 0004, Huchuan Lu |
AAAI | 6 |
| 2025 | Knowledge Graph Completion with Relation-Aware Anchor EnhancementabstractText-based knowledge graph completion methods take advantage of pre-trained language models (PLM) to enhance intrinsic semantic connections of raw triplets with detailed text descriptions. Typical methods in this branch map an input query (textual descriptions associated with an entity and a relation) and its candidate entities into feature vectors, respectively, and then maximize the probability of valid triples. These methods are gaining promising performance and increasing attention for the rapid development of large language models. According to the property of the language models, the more related and specific context information the input query provides, the more discriminative the resultant embedding will be. In this paper, through observation and validation, we find a neglected fact that the relation-aware neighbors of the head entities in queries could act as effective contexts for more precise link prediction. Driven by this finding, we propose a relation-aware anchor enhanced knowledge graph completion method (RAA-KGC). Specifically, in our method, to provide a reference of what might the target entity be like, we first generate anchor entities within the relation-aware neighborhood of the head entity. Then, by pulling the query embedding towards the neighborhoods of the anchors, it is tuned to be more discriminative for target entity matching. The results of our extensive experiments not only validate the efficacy of RAA-KGC but also reveal that by integrating our relation-aware anchor enhancement strategy, the performance of current leading methods can be notably enhanced without substantial modifications. Duanyang Yuan, Sihang Zhou 0001, Xiaoshu Chen, Dong Wang 0004, Ke Liang 0006, Xinwang Liu 0002, Jian Huang 0010 |
AAAI | 4 |
| 2025 | Two-stream Beats One-stream: Asymmetric Siamese Network for Efficient Visual TrackingabstractEfficient tracking has garnered attention for its ability to operate on resource-constrained platforms for real-world deployment beyond desktop GPUs. Current efficient trackers mainly follow precision-oriented trackers, adopting a one-stream framework with lightweight modules. However, blindly adhering to the one-stream paradigm may not be optimal, as incorporating template computation in every frame leads to redundancy, and pervasive semantic interaction between template and search region places stress on edge devices. In this work, we propose a novel asymmetric Siamese tracker named AsymTrack for efficient tracking. AsymTrack disentangles template and search streams into separate branches, with template computing only once during initialization to generate modulation signals. Building on this architecture, we devise an efficient template modulation mechanism to unidirectional inject crucial cues into the search features, and design an object perception enhancement module that integrates abstract semantics and local details to overcome the limited representation in lightweight tracker. Extensive experiments demonstrate that AsymTrack offers superior speed-precision trade-offs across different platforms compared to the current state-of-the-arts. For instance, AsymTrack-T achieves 60.8% AUC on LaSOT and 224/81/84 FPS on GPU/CPU/AGX, surpassing HiT-Tiny by 6.0% AUC with higher speeds. Jiawen Zhu 0003, Huayi Tang, Xin Chen 0032, Xinying Wang 0005, Dong Wang 0004, Huchuan Lu |
AAAI | 5 |
| 2025 | CAT: A Unified Click-and-Track Framework for Realistic Tracking
Yongsheng Yuan, Jie Zhao 0014, Dong Wang 0004, Huchuan Lu |
ICCV | 3 |
| 2025 | Efficient Motion Prompt Learning for Robust Visual TrackingabstractDue to the challenges of processing temporal information, most trackers depend solely on visual discriminability and overlook the unique temporal coherence of video data. In this paper, we propose a lightweight and plug-and-play motion prompt tracking method. It can be easily integrated into existing vision-based trackers to build a joint tracking framework leveraging both motion and vision cues, thereby achieving robust tracking through efficient prompt learning. A motion encoder with three different positional encodings is proposed to encode the long-term motion trajectory into the visual embedding space, while a fusion decoder and an adaptive weight mechanism are designed to dynamically fuse visual and motion features. We integrate our motion module into three different trackers with five models in total. Experiments on seven challenging tracking benchmarks demonstrate that the proposed motion module significantly improves the robustness of vision-based trackers, with minimal training costs and negligible speed sacrifice. Code is available at https://github.com/zj5559/Motion-Prompt-Tracking. Jie Zhao 0014, Xin Chen 0032, Yongsheng Yuan, Michael Felsberg, Dong Wang 0004, Huchuan Lu |
ICML | 5 |
| 2025 | GFM-Planner: Perception-Aware Trajectory Planning with Geometric Feature MetricabstractLike humans who rely on landmarks for orientation, autonomous robots depend on feature-rich environments for accurate localization. In this paper, we propose the GFM-Planner, a perception-aware trajectory planning framework based on the geometric feature metric, which enhances LiDAR localization accuracy by guiding the robot to avoid degraded areas. First, we derive the Geometric Feature Metric (GFM) from the fundamental LiDAR localization problem. Next, we design a 2D grid-based Metric Encoding Map (MEM) to efficiently store GFM values across the environment. A constant-time decoding algorithm is further proposed to retrieve GFM values for arbitrary poses from the MEM. Finally, we develop a perception-aware trajectory planning algorithm that improves LiDAR localization capabilities by guiding the robot in selecting trajectories through feature-rich areas. Both simulation and real-world experiments demonstrate that our approach enables the robot to actively select trajectories that significantly enhance LiDAR localization accuracy. Dong Wang 0004, Huchuan Lu |
IROS | 4 |
| 2025 | Equilibrium Policy Generalization: A Reinforcement Learning Framework for Cross-Graph Zero-Shot Generalization in Pursuit-Evasion GamesabstractEquilibrium learning in adversarial games is an important topic widely examined in the fields of game theory and reinforcement learning (RL). Pursuit-evasion game (PEG), as an important class of real-world games from the fields of robotics and security, requires exponential time to be accurately solved. When the underlying graph structure varies, even the state-of-the-art RL methods require recomputation or at least fine-tuning, which can be time-consuming and impair real-time applicability. This paper proposes an Equilibrium Policy Generalization (EPG) framework to effectively learn a generalized policy with robust cross-graph zero-shot performance. In the context of PEGs, our framework is generally applicable to both pursuer and evader sides in both no-exit and multi-exit scenarios. These two generalizability properties, to our knowledge, are the first to appear in this domain. The core idea of the EPG framework is to train an RL policy across different graph structures against the equilibrium policy for each single graph. To construct an equilibrium oracle for single-graph policies, we present a dynamic programming (DP) algorithm that provably generates pure-strategy Nash equilibrium with near-optimal time complexity. To guarantee scalability with respect to pursuer number, we further extend DP and RL by designing a grouping mechanism and a sequence model for joint policy decomposition, respectively. Experimental results show that, using equilibrium guidance and a distance feature proposed for cross-graph PEG training, the EPG framework guarantees desirable zero-shot performance in various unseen real-world graphs. Besides, when trained under an equilibrium heuristic proposed for the graphs with exits, our generalized pursuer policy can even match the performance of the fine-tuned policies from the state-of-the-art PEG methods. Runyu Lu, Peng Zhang 0127, Ruochuan Shi, Yuanheng Zhu, Dongbin Zhao, Yang Liu 0066, Dong Wang 0004, Cesare Alippi |
NeurIPS | 7 |
| 2025 | Exploring a Hierarchical Cross-Attention Transformer for High-Speed Tracking
Xin Chen 0032, Ben Kang, Jiawen Zhu 0003, Dongdong Li 0004, Chunjuan Bo, Dong Wang 0004 |
Comput. Vis. Media | 6 |
| 2025 | Complex knowledge base question answering with difficulty-aware active data augmentation
Dong Wang 0004, Sihang Zhou 0001, Ke Liang 0006, Chuanli Wang, Huang Jian |
Expert Syst. Appl. | 1 |
| 2025 | Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual TrackingabstractAbstract Transformer-based visual trackers have demonstrated significant advancements due to their powerful modeling capabilities. However, their practicality is limited on resource-constrained devices because of their slow processing speeds. To address this challenge, we present HiT, a novel family of efficient tracking models that achieve high performance while maintaining fast operation across various devices. The core innovation of HiT lies in its Bridge Module, which connects lightweight transformers to the tracking framework, enhancing feature representation quality. Additionally, we introduce a dual-image position encoding approach to effectively encode spatial information. HiT achieves an impressive speed of 61 frames per second (fps) on the NVIDIA Jetson AGX platform, alongside a competitive AUC of 64.6% on the LaSOT benchmark, outperforming all previous efficient trackers. Building on HiT, we propose DyHiT, an efficient dynamic tracker that flexibly adapts to scene complexity by selecting routes with varying computational requirements. DyHiT uses search area features extracted by the backbone network and inputs them into an efficient dynamic router to classify tracking scenarios. Based on the classification, DyHiT applies a divide-and-conquer strategy, selecting appropriate routes to achieve a superior trade-off between accuracy and speed. The fastest version of DyHiT achieves 111 fps on NVIDIA Jetson AGX while maintaining an AUC of 62.4% on LaSOT. Furthermore, we introduce a training-free acceleration method based on the dynamic routing architecture of DyHiT. This method significantly improves the execution speed of various high-performance trackers without sacrificing accuracy. For instance, our acceleration method enables the state-of-the-art tracker SeqTrack-B256 to achieve a $$2.68\times $$ 2.68 × speedup on an NVIDIA GeForce RTX 2080 Ti GPU while maintaining the same AUC of 69.9% on the LaSOT. Codes, models, and results are available at https://github.com/kangben258/HiT . Ben Kang, Xin Chen 0032, Jie Zhao 0014, Chunjuan Bo, Dong Wang 0004, Huchuan Lu |
Int. J. Comput. Vis. | 5 |
| 2025 | MoE-Adapters++: Toward More Efficient Continual Learning of Vision-Language Models Via Dynamic Mixture-of-Experts AdaptersabstractIn this paper, we first propose MoE-Adapters, a parameter-efficient training framework to alleviate long-term forgetting issues in incremental learning with Vision-Language Models (VLM). Our MoE-Adapters leverages incrementally added routers to activate and integrate exclusive expert adapters from a pre-defined static expert set, enabling the pre-trained CLIP to efficiently adapt to new tasks. To preserve the zero-shot capability of VLM, a Distribution Discriminative Auto-Selector (DDAS) is introduced that automatically routes in-distribution and out-of-distribution inputs to the MoE-Adapters and the original CLIP, respectively. However, relying on a static expert set and a separate distribution selector can lead to parameter redundancy and increased training complexity. In response, we further extend an MoE-Adapters++ framework by introducing dynamic MoE-adapters, which allows experts to be adaptively involved during the continual learning process. Additionally, a Latent Embedding Auto-Selector (LEAS) is proposed that incorporates distribution selection within CLIP to create a more unified architecture. Extensive experiments across diverse settings demonstrate that the proposed method consistently surpasses previous state-of-the-art approaches while concurrently improving training efficiency. Jiazuo Yu 0001, Zichen Huang 0004, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | A Workload-Aware Encrypted Index for Efficient Privacy-Preserving Range Queries
Dong Wang 0004, Ningning Cui, Jianxin Li 0001, Jianzhong Qi 0001, Jianliang Xu |
Proc. VLDB Endow. | 1 |
| 2025 | MambaVT: Spatio-Temporal Contextual Modeling for Robust RGB-T Tracking
Simiao Lai, Chang Liu 0071, Jiawen Zhu 0003, Ben Kang, Yang Liu 0066, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | EMTrack: Efficient Multimodal Object TrackingabstractMulti-modal object tracking has received increasing attention, given the limitations the representation ability in certain challenging scenarios of single RGB modality. Recent prompt tuning techniques enable multimodal tracking to effectively inherit knowledge from foundation models trained with a large amount of RGB tracking data and achieve parameter-efficient training. However, few works focus on the efficient inference of multimodal tracking handling multiple RGB-X (RGB-Thermal, RGB-Depth, RGB-Event, etc.) tracking tasks simultaneously, especially on resource-limited devices such as CPU. In this work, we propose an efficient multimodal tracker named EMTrack. EMTrack follows a concise and unified multimodal tracking framework with simple knowledge distillation. RGB modality and auxiliary modality are added after patch-embedding layer for fusion, reducing the computational complexity of multimodal tracking compared with that of single modality. Before fusion operation, we introduce a modal-specific spatial modulation module to exploit and realize adaptive spatial adjustment of different modality features. Multiple modal-specific experts are adopted to capture specific information for different RGB-X tracking tasks, which assists in handling such tasks in a unified model with joint training. EMTrack achieves competitive performance on various RGB-X tracking benchmarks while reaching a good balance of performance and speed on different platforms. Especially on an Intel Core i9-10850K CPU device, EMTrack achieves 29.1 fps, a real-time speed, with only 2.0G MAC computation. Chang Liu 0071, Ziqi Guan, Simiao Lai, Yang Liu 0066, Huchuan Lu, Dong Wang 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Learning Language Prompt for Vision-Language TrackingabstractVision-language object tracking integrates advanced linguistic information, enhancing its robustness and accuracy in complex scenarios. Nevertheless, current methods are constrained by a lack of sufficient vision-language data, making it challenging for the model to learn generalized knowledge. To alleviate this issue, we propose a new prompt-based framework for vision-language tracking, named ProVLT. This framework casts language information as a prompt for pretrained visionbased tracking models, thereby leveraging the knowledge from extensive tracking data. Experiments demonstrate that ProVLT achieves competitive performance while training only a fraction of parameters (approximately 29% of modal parameters). For instance, ProVLT achieves competitive performance, attaining AUC of 59.8% on TNL2K benchmark. Furthermore, we augment five mainstream vision-only tracking benchmarks with language annotations, and find that the inclusion of linguistic information consistently improves tracking performance. On these benchmarks, the linguistic information improves the performance by an average of 2.9% compared with the vision-based tracker. We will release the code, models, and benchmarks for the community. ChengAo Zong, Jie Zhao 0014, Xin Chen 0032, Huchuan Lu, Dong Wang 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Self-Adaptive Vision-Language Tracking With Context PromptingabstractDue to the substantial gap between vision and language modalities, along with the mismatch problem between fixed language descriptions and dynamic visual information, existing vision-language tracking methods exhibit performance on par with or slightly worse than vision-only tracking. Effectively exploiting the rich semantics of language to enhance tracking robustness remains an open challenge. To address these issues, we propose a self-adaptive vision-language tracking framework that leverages the pre-trained multi-modal CLIP model to obtain well-aligned visual-language representations. A novel context-aware prompting mechanism is introduced to dynamically adapt linguistic cues based on the evolving visual context during tracking. Specifically, our context prompter extracts dynamic visual features from the current search image and integrates them into the text encoding process, enabling self-updating language embeddings. Furthermore, our framework employs a unified one-stream Transformer architecture, supporting joint training for both vision-only and vision-language tracking scenarios. Our method not only bridges the modality gap but also enhances robustness by allowing language features to evolve with visual context. Extensive experiments on four vision-language tracking benchmarks demonstrate that our method effectively leverages the advantages of language to enhance visual tracking. Our large model can obtain 55.0% AUC on $\text {LaSOT}_{\text {EXT}}$ and 69.0% AUC on TNL2K. Additionally, our language-only tracking model achieves performance comparable to that of state-of-the-art vision-only tracking methods on TNL2K. Code is available at https://github.com/zj5559/SAVLT. Jie Zhao 0014, Xin Chen 0032, Shengming Li, Chunjuan Bo, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Image Process. | 5 |
| 2025 | Enhancing the Two-Stream Framework for Efficient Visual TrackingabstractPractical deployments, especially on resource-limited edge devices, necessitate high speed for visual object trackers. To meet this demand, we introduce a new efficient tracker with a Two-Stream architecture, named ToS. While the recent one-stream tracking framework, employing a unified backbone for simultaneous processing of both the template and search region, has demonstrated exceptional efficacy, we find the conventional two-stream tracking framework, which employs two separate backbones for the template and search region, offers inherent advantages. The two-stream tracking framework is more compatible with advanced lightweight backbones and can efficiently utilize benefits from large templates. We demonstrate that the two-stream setup can exceed the one-stream tracking model in both speed and accuracy through strategic designs. Our methodology rejuvenates the two-stream tracking paradigm with lightweight pre-trained backbones and the proposed three efficient strategies: 1) A feature-aggregation module that improves the representation capability of the backbone, 2) A channel-wise approach for feature fusion, presenting a more effective and lighter alternative to spatial concatenation techniques, and 3) An expanded template strategy to boost tracking accuracy with negligible additional computational cost. Extensive evaluations across multiple tracking benchmarks demonstrate that the proposed method sets a new state-of-the-art performance in efficient visual tracking. ChengAo Zong, Xin Chen 0032, Jie Zhao 0014, Yang Liu 0066, Huchuan Lu, Dong Wang 0004 |
IEEE Trans. Image Process. | 6 |
| 2025 | Refocus the Attention for Parameter-Efficient Thermal Infrared Object TrackingabstractIntroducing deep trackers to thermal infrared (TIR) tracking is hampered by the scarcity of large training datasets. To alleviate the predicament, a common approach is full fine-tuning (FFT) based on pretrained RGB parameters. Nevertheless, due to its inefficient training pattern and representation collapse risk, some parameter-efficient fine-tuning (PEFT) alternatives have been promoted recently. However, the existing PEFT algorithms typically follow a bottom-up way, where their attention solely relies on the input and lacks the capability of task-guided top-down attention, which provides the task-relevant representation such as the human visual perception system. In this article, we introduce ReFocus, a new PEFT method that adapts the pretrained RGB foundation tracking model to the downstream TIR tracking task through the guidance of high-level task-specific signals in a top-down attention manner. By freezing the entire foundation model and only training query-guided feature selection and top-down blocks, ReFocus achieves state-of-the-art (SOTA) TIR tracking performance while keeping training efficiency. Extensive experiments on five TIR tracking benchmarks demonstrate that ReFocus significantly improves the performance of the foundation tracker. Besides, further ablation studies show the effectiveness and flexible adaptability of the proposed method to lighter foundation models and different tracking frameworks. Compared to FFT and other bottom-up PEFT paradigms, such as head probe, low-rank adaptation (LoRA), and adapter, our method achieves comparable or superior performance with fewer training parameters and reveals the advantage of learning stability. Simiao Lai, Chang Liu 0071, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Event-Assisted Recurrent Network for Arbitrary-Temporal-Scale Blurry Image UnfoldingabstractRecovering a sequence of latent sharp frames from a motion-blurred image is a challenging task. The bio-inspired event camera, which produces an event stream with high temporal resolution, has been exploited to promote the recovery performance. However, recovering sharp sequences with arbitrary temporal scales has been ignored for a long time. Existing works can only recover a fixed number of latent frames from a blurry image once they are trained. In this work, we propose an event-assisted blurry image unfolding framework that can work across arbitrary temporal scales. A bi-directional recurrent network is employed to encode events corresponding to each latent frame, which gathers information over all events in the exposure time. Features of both the blurry image and events are fused together and fed to a bi-directional latent sequence decoder (BiLSD) to produce a sequence of latent sharp frames. Extensive experiments show that the proposed method not only performs favorably against state-of-the-art methods in recovering a fixed number of frames from a blurry image but can be well generalized to arbitrary-temporal-scale blurry image unfolding. Hao Ju 0004, Weihua He, Yaoyuan Wang, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Exploring Dynamic Transformer for Efficient Object TrackingabstractThe speed-precision tradeoff is a critical problem in visual object tracking, as it typically requires low latency and is deployed on resource-constrained platforms. Existing solutions for efficient tracking primarily focus on lightweight backbones or modules, which, however, come at a sacrifice in precision. In this article, inspired by dynamic network routing, we propose DyTrack, a dynamic transformer framework for efficient tracking. Real-world tracking scenarios exhibit varying levels of complexity. We argue that a simple network is sufficient for easy video frames, while more computational resources should be assigned to difficult ones. DyTrack automatically learns to configure proper reasoning routes for different inputs, thereby improving the utilization of the available computational budget and achieving higher performance at the same running speed. We formulate instance-specific tracking as a sequential decision problem and incorporate terminating branches to intermediate layers of the model. Furthermore, we propose a feature recycling mechanism to maximize computational efficiency by reusing the outputs of predecessors. Additionally, a target-aware self-distillation strategy is designed to enhance the discriminating capabilities of early-stage predictions by mimicking the representation patterns of the deep model. Extensive experiments demonstrate that DyTrack achieves promising speed-precision tradeoffs with only a single model. For instance, DyTrack obtains 64.9% area under the curve (AUC) on LaSOT with a speed of 256 fps. Jiawen Zhu 0003, Xin Chen 0032, Haiwen Diao, Shuai Li 0014, Jun-Yan He, Chenyang Li 0007, Bin Luo 0008, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | Hybrid-SORT: Weak Cues Matter for Online Multi-Object TrackingabstractMulti-Object Tracking (MOT) aims to detect and associate all desired objects across frames. Most methods accomplish the task by explicitly or implicitly leveraging strong cues (i.e., spatial and appearance information), which exhibit powerful instance-level discrimination. However, when object occlusion and clustering occur, spatial and appearance information will become ambiguous simultaneously due to the high overlap among objects. In this paper, we demonstrate this long-standing challenge in MOT can be efficiently and effectively resolved by incorporating weak cues to compensate for strong cues. Along with velocity direction, we introduce the confidence and height state as potential weak cues. With superior performance, our method still maintains Simple, Online and Real-Time (SORT) characteristics. Also, our method shows strong generalization for diverse trackers and scenarios in a plug-and-play and training-free manner. Significant and consistent improvements are observed when applying our method to 5 different representative trackers. Further, with both strong and weak cues, our method Hybrid-SORT achieves superior performance on diverse benchmarks, including MOT17, MOT20, and especially DanceTrack where interaction and severe occlusion frequently happen with complex motions. The code and models are available at https://github.com/ymzis69/HybridSORT. Mingzhan Yang, Guangxin Han, Bin Yan 0004, Jinqing Qi, Huchuan Lu, Dong Wang 0004 |
AAAI | 7 |
| 2024 | Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts AdaptersabstractContinual learning can empower vision-language models to continuously acquire new knowledge, without the need for access to the entire historical dataset. However, mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout life-long learning and (ii) significant computational burdens associated with full-model tuning. In this work, we present a parameter-efficient continual learning framework to alleviate long-term forgetting in incremental learning with vision-language models. Our approach involves the dynamic expansion of a pre-trained CLIP model, through the integration of Mixture-of-Experts (MoE) adapters in response to new tasks. To preserve the zero-shot recognition capability of vision-language models, we further introduce a Distribution Discriminative Auto-Selector (DDAS) that automatically routes in-distribution and out-of-distribution inputs to the MoE Adapter and the original CLIP, respectively. Through extensive experiments across various settings, our proposed method consistently outperforms previous state-of-the-art approaches while concurrently reducing parameter training burdens by 60%. Our code locates at https://github.com/JiazuoYu/MoE-Adapters4CL Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002 |
CVPR | 5 |
| 2024 | EvSign: Sign Language Recognition and Translation with Streaming Events
Zeren Wang, Wenyue Chen, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012 |
ECCV (5) | 6 |
| 2024 | DepthRefiner: Adapting RGB Trackers to RGBD Scenes via Depth-Fused RefinementabstractThe increasing availability of depth sensors has facilitated the acquisition of depth images, thereby driving advancements in RGBD tracking. However, compared to RGB benchmarks, deficient data hampers the sufficient learning of RGBD trackers. In this paper, instead of developing a new RGBD tracker from scratch, we aim to learn a depth-fused refinement module that enables existing RGB trackers to adapt to RGBD scenes. Specifically, we introduce a compact yet effective module, named DepthRefiner (DR), based on multi-head self-attention, a simple bimodal fusion technique, and the center-based head. This approach leverages the learned prior representations of RGB trackers from large-scale RGB data and can be flexibly integrated into various off-the-shelf trackers without modifying original pipelines. Comprehensive experiments on CDTB, DepthTrack, VOT-RGBD2022, and RGBD1K benchmarks with multiple base trackers validate that our approach significantly improves the base tracker’s performance while adding minimal computational overhead. Simiao Lai, Dong Wang 0004, Huchuan Lu |
ICME | 2 |
| 2024 | Multi-Stage Fusion for Event-based Multimodal TrackerabstractEvent cameras are bio-inspired sensors with high dynamic range and time resolution, which are favorable properties for visual object tracking. There are already some methods that fuse the event modality and RGB modality with cross-domain feature integrator to achieve improved tracking performance. Researchers have developed some architectures for event modality processing or fusion, successfully boosting the tracking performance. In this work, we design a RGB-E tracker with multi-stage fusion. In the early stage, frames are enhanced with aid of events to mitigate blur or under/over-exposure degradation. During the middle stage, we utilize a fusion module for feature-level integration. At the late stage, we carry out decision-level fusion by predicting tracking boxes based on frame features, event features, and fused features, and the one with highest score is taken as the final estimation. Our design thoroughly integrate information from various levels, allowing each modality to contribute to the tracking process as much as possible. Extensive experiments demonstrate that the proposed method performs favorably against state-of-the-art RGB-E trackers in both accuracy and efficiency. Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Wenyue Chen, Dong Wang 0004, Shengming Li, Huchuan Lu |
ICME | 5 |
| 2024 | Safety-First Tracker: A Trajectory Planning Framework for Omnidirectional Robot TrackingabstractThis paper introduces a Safety-First Tracker (SF-Tracker) designed for omnidirectional autonomous tracking robots. The position and orientation of omnidirectional robots are decoupled for stepwise planning to ensure trajectory safety and maintain target visibility. SF-Tracker puts the trajectory safety in the first place. First, a collision-free and occlusion-free reference path is efficiently initialized by constructing a directed weighted graph. By building upon this path, safe trajectory optimization is implemented to ensure safe movement. Finally, an orientation planner is developed to achieve target visibility based on the safe trajectory. Extensive experimental evaluations in simulated environments and the real world demonstrate that the SF-Tracker outperforms state-of-the-art methods in terms trajectory safety and target visibility. Ablation experiments further demonstrate the significance of each step of the SF-Tracker. The source code and demonstration video can be found at https://github.com/Yue-0/SF-Tracker. Yang Liu 0003, Xin Chen 0032, Dong Wang 0004, Huchuan Lu |
IROS | 5 |
| 2024 | LLMs Can Evolve Continually on Modality for X-Modal Reasoning
Jiazuo Yu 0001, Haomiao Xiong, Lu Zhang 0053, Haiwen Diao, Yunzhi Zhuge, Lanqing Hong, Dong Wang 0004, Huchuan Lu, You He 0002, Long Chen 0016 |
NeurIPS | 7 |
| 2024 | Leveraging the Power of Data Augmentation for Transformer-based TrackingabstractDue to long-distance correlation and powerful pretrained models, transformer-based methods have initiated a breakthrough in visual object tracking performance. Previous works focus on designing effective architectures suited for tracking, but ignore that data augmentation is equally crucial for training a well-performing model. In this paper, we first explore the impact of general data augmentations on transformer-based trackers via systematic experiments, and reveal the limited effectiveness of these common strategies. Motivated by experimental observations, we then propose two data augmentation methods customized for tracking. First, we optimize existing random cropping via a dynamic search radius mechanism and simulation for boundary samples. Second, we propose a token-level feature mixing augmentation strategy, which enables the model against challenges like background interference. Extensive experiments on two transformer-based trackers and six benchmarks demonstrate the effectiveness and data efficiency of our methods, especially under challenging settings, like one-shot tracking and small image resolutions. Code is available at https://github.com/zj5559/DATr. Jie Zhao 0014, Johan Edstedt, Michael Felsberg, Dong Wang 0004, Huchuan Lu |
WACV | 4 |
| 2024 | Other tokens matter: Exploring global and local features of Vision Transformers for Object Re-Identification
Yingquan Wang, Dong Wang 0004, Huchuan Lu |
Comput. Vis. Image Underst. | 3 |
| 2024 | Neural image re-exposure
Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Dong Wang 0004, Lihe Zhang, Bolun Zheng, Wei Zhou 0021, Huchuan Lu |
Comput. Vis. Image Underst. | 4 |
| 2024 | Multi-modal visual tracking: Review and experimental comparisonabstractVisual object tracking has been drawing increasing attention in recent years, as a fundamental task in computer vision. To extend the range of tracking applications, researchers have been introducing information from multiple modalities to handle specific scenes, with promising research prospects for emerging methods and benchmarks. To provide a thorough review of multi-modal tracking, different aspects of multi-modal tracking algorithms are summarized under a unified taxonomy, with specific focus on visible-depth (RGB-D) and visible-thermal (RGB-T) tracking. Subsequently, a detailed description of the related benchmarks and challenges is provided. Extensive experiments were conducted to analyze the effectiveness of trackers on five datasets: PTB, VOT19-RGBD, GTOT, RGBT234, and VOT19-RGBT. Finally, various future directions, including model design and dataset construction, are discussed from different perspectives for further research. Dong Wang 0004, Huchuan Lu |
Comput. Vis. Media | 2 |
| 2024 | Triple alignment-enhanced complex question answering over knowledge bases
Dong Wang 0004, Sihang Zhou 0001, Jian Huang 0010, Xiangrong Ni |
Neurocomputing | 1 |
| 2024 | LGTrack: Exploiting Local and Global Properties for Robust Visual TrackingabstractRe-detection is a necessary capability for long-term tracking. Target candidate proposals in the whole image can provide a chance of tracking reset when tracking fails due to tracking drift or target invisibility. In this paper, we propose a unified local-global tracker based on the same transformer architecture sharing weights, which can not only search in a continuous local region but also provide target candidates of the global image in every frame. The requirements of both long-term and short-term scenarios can be addressed using a unified model. A simple proposal selection scheme is adopted to properly select the candidate proposals of re-detection, to assist tracking and obtain better performance. The scheme performs reevaluation of all high-quality proposals based on a transformer-based embedding network, once the predicted state of the local tracking is not sufficient to be accurate. To capture appearance variations brought by online updates in minimum risks, a long-term-friendly dynamic template update scheme is also designed. Extensive experiments are conducted to demonstrate the effectiveness of our proposed tracker, including three short-term tracking benchmarks and six long-term benchmarks. Our tracker can achieve results comparable to that of the state-of-the-art. The proposed tracker can also work well in balancing the performance and speed, achieving an average speed of approximately 25 fps tested on LaSOT testing set. Chang Liu 0071, Jie Zhao 0014, Chunjuan Bo, Shengming Li, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | SRRT: Exploring Search Region Regulation for Visual Object TrackingabstractThe dominant trackers generate a fixed-size rectangular region based on the previous prediction or initial bounding box as the model input, i.e., search region. While this manner obtains promising tracking efficiency, a fixed-size search region lacks flexibility and is likely to fail in some cases, e.g., fast motion and distractor interference. Trackers tend to lose the target object due to the limited search region or experience interference from distractors due to the excessive search region. Drawing inspiration from the pattern humans track an object, we propose a novel tracking paradigm, called Search Region Regulation Tracking (SRRT) that applies a small eyereach when the target is captured and zooms out the search field when the target is about to be lost. SRRT applies a proposed search region regulator to estimate an optimal search region dynamically for each frame, by which the tracker can flexibly respond to transient changes in the location of object occurrences. To adapt the object’s appearance variation during online tracking, we further propose a locking-state determined updating strategy for reference frame updating. The proposed SRRT is concise without bells and whistles, yet achieves evident improvements and competitive results with other state-of-the-art trackers on eight benchmarks. On the large-scale LaSOT benchmark, SRRT improves SiamRPN++ and TransT with absolute gains of 4.6% and 3.1% in terms of AUC. The code and models will be released. Jiawen Zhu 0003, Xin Chen 0032, Xinying Wang 0005, Dong Wang 0004, Wenda Zhao 0003, Huchuan Lu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Event-Assisted Blurriness Representation Learning for Blurry Image UnfoldingabstractThe goal of blurry image deblurring and unfolding task is to recover a single sharp frame or a sequence from a blurry one. Recently, its performance is greatly improved with introduction of a bio-inspired visual sensor, event camera. Most existing event-assisted deblurring methods focus on the design of powerful network architectures and effective training strategy, while ignoring the role of blur modeling in removing various blur in dynamic scenes. In this work, we propose to implicitly model blur in an image by computing blurriness representation with an event-assisted blurriness encoder. The learning of blurriness representation is formulated as a ranking problem based on specially synthesized pairs. Blurriness-aware image unfolding is achieved by integrating blur relevant information contained in the representation into a base unfolding network. The integration is mainly realized by the proposed blurriness-guided modulation and multi-scale aggregation modules. Experiments on GOPRO and HQF datasets show favorable performance of the proposed method against state-of-the-art approaches. More results on real-world data validate its effectiveness in recovering a sequence of latent sharp frames from a blurry image. Hao Ju 0004, Lei Yu 0006, Weihua He, Yaoyuan Wang, Qi Xu 0008, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012 |
IEEE Trans. Image Process. | 9 |
| 2023 | Dual Memory Aggregation Network for Event-Based Object Detection with Learnable RepresentationabstractEvent-based cameras are bio-inspired sensors that capture brightness change of every pixel in an asynchronous manner. Compared with frame-based sensors, event cameras have microsecond-level latency and high dynamic range, hence showing great potential for object detection under high-speed motion and poor illumination conditions. Due to sparsity and asynchronism nature with event streams, most of existing approaches resort to hand-crafted methods to convert event data into 2D grid representation. However, they are sub-optimal in aggregating information from event stream for object detection. In this work, we propose to learn an event representation optimized for event-based object detection. Specifically, event streams are divided into grids in the x-y-t coordinates for both positive and negative polarity, producing a set of pillars as 3D tensor representation. To fully exploit information with event streams to detect objects, a dual-memory aggregation network (DMANet) is proposed to leverage both long and short memory along event streams to aggregate effective information for object detection. Long memory is encoded in the hidden state of adaptive convLSTMs while short memory is modeled by computing spatial-temporal correlation between event pillars at neighboring time intervals. Extensive experiments on the recently released event-based automotive detection dataset demonstrate the effectiveness of the proposed method. Xu Jia 0012, Xinyu Zhang 0017, Yaoyuan Wang, Dong Wang 0004, Huchuan Lu |
AAAI | 7 |
| 2023 | Universal Instance Perception as Object Discovery and RetrievalabstractAll instance perception tasks aim at finding certain objects specified by some queries such as category names, language expressions, and target annotations, but this complete field has been split into multiple independent sub-tasks. In this work, we present a universal instance perception model of the next generation, termed UNINEXT. UNINEXT reformulates diverse instance perception tasks into a unified object discovery and retrieval paradigm and can flexibly perceive different types of objects by simply changing the input prompts. This unified formulation brings the following benefits: (1) enormous data from different tasks and label vocabularies can be exploited for jointly training general instance-level representations, which is especially beneficial for tasks lacking in training data. (2) the unified model is parameter-efficient and can save redundant computation when handling multiple tasks simultaneously. UNINEXT shows superior performance on 20 challenging benchmarks from 10 instance-level tasks including classical image-level tasks (object detection and instance segmentation), vision-and-language tasks (referring expression comprehension and segmentation), and six video-level object tracking tasks. Code is available at https://github.com/MasterBin-IIAU/UNINEXT. Bin Yan 0004, Yi Jiang 0009, Jiannan Wu, Dong Wang 0004, Ping Luo 0002, Zehuan Yuan, Huchuan Lu |
CVPR | 4 |
| 2023 | SeqTrack: Sequence to Sequence Learning for Visual Object TrackingabstractIn this paper, we present a new sequence-to-sequence learning framework for visual tracking, dubbed SeqTrack. It casts visual tracking as a sequence generation problem, which predicts object bounding boxes in an autoregressive fashion. This is different from prior Siamese trackers and transformer trackers, which rely on designing complicated head networks, such as classification and regression heads. SeqTrack only adopts a simple encoder-decoder transformer architecture. The encoder extracts visual features with a bidirectional transformer, while the decoder generates a sequence of bounding box values autoregressively with a causal transformer. The loss function is a plain cross-entropy. Such a sequence learning paradigm not only simplifies tracking framework, but also achieves competitive performance on benchmarks. For instance, SeqTrack gets 72.5% AUC on LaSOT, establishing a new state-of-the-art performance. Code and models are available at https://github.com/microsoft/VideoX. Xin Chen 0032, Houwen Peng, Dong Wang 0004, Huchuan Lu, Han Hu 0001 |
CVPR | 3 |
| 2023 | Representation Learning for Visual Object Tracking by Masked Appearance TransferabstractVisual representation plays an important role in visual object tracking. However, few works study the tracking-specified representation learning method. Most trackers directly use ImageNet pre-trained representations. In this paper, we propose masked appearance transfer, a simple but effective representation learning method for tracking, based on an encoder-decoder architecture. First, we encode the visual appearances of the template and search region jointly, and then we decode them separately. During decoding, the original search region image is reconstructed. However, for the template, we make the decoder reconstruct the target appearance within the search region. By this target appearance transfer, the tracking-specified representations are learned. We randomly mask out the inputs, thereby making the learned representations more discriminative. For sufficient evaluation, we design a simple and lightweight tracker that can evaluate the representation for both target localization and box regression. Extensive experiments show that the proposed method is effective, and the learned representations can enable the simple tracker to obtain state-of-the-art performance on six datasets. https://github.com/difhnp/MAT Haojie Zhao, Dong Wang 0004, Huchuan Lu |
CVPR | 2 |
| 2023 | Visual Prompt Multi-Modal TrackingabstractVisible-modal object tracking gives rise to a series of downstream multi-modal tracking tributaries. To inherit the powerful representations of the foundation model, a natural modus operandi for multi-modal tracking is full fine-tuning on the RGB-based parameters. Albeit effective, this manner is not optimal due to the scarcity of downstream data and poor transferability, etc. In this paper, inspired by the recent success of the prompt learning in language models, we develop Visual Prompt multi-modal Tracking (ViPT), which learns the modal-relevant prompts to adapt the frozen pre-trained foundation model to various downstream multi-modal tracking tasks. ViPT finds a better way to stimulate the knowledge of the RGB-based model that is pre-trained at scale, meanwhile only introducing a few trainable parameters (less than 1% of model parameters). ViPT outperforms the full fine-tuning paradigm on multiple downstream tracking tasks including RGB+Depth, RGB+Thermal, and RGB+Event tracking. Extensive experiments show the potential of visual prompt learning for multi-modal tracking, and ViPT can achieve state-of-the-art performance while satisfying parameter efficiency. Code and models are available at https://github.com/jiawen-zhu/ViPT. Jiawen Zhu 0003, Simiao Lai, Xin Chen 0032, Dong Wang 0004, Huchuan Lu |
CVPR | 4 |
| 2023 | Efficient Siamese Network for UAV TrackingabstractIn this work, we propose an efficient Siamese-based tracker (ESTrack) for aerial visual object tracking using dual global correlation and accurate center localization. The dual correlation module embeds task-specific global similarity information for target classification and state estimation. The center localization highlights the classification score where the target exists, which improves the accuracy of the classification and reduces complex hyperparameter tuning. Extensive experiments on five UAV benchmarks show that our ESTrack-50 performs favorably against many state-of-the-art Siamese- based trackers with a speed of 110 fps. Meanwhile, ESTrack-18 and ESTrack-AO achieve 180 fps with comparable performance to most UAV-based trackers. Dong Wang 0004 |
ICASSP | 2 |
| 2023 | Exploring Lightweight Hierarchical Vision Transformers for Efficient Visual TrackingabstractTransformer-based visual trackers have demonstrated significant progress owing to their superior modeling capabilities. However, existing trackers are hampered by low speed, limiting their applicability on devices with limited computational power. To alleviate this problem, we propose HiT, a new family of efficient tracking models that can run at high speed on different devices while retaining high performance. The central idea of HiT is the Bridge Module, which bridges the gap between modern lightweight transformers and the tracking framework. The Bridge Module incorporates the high-level information of deep features into the shallow large-resolution features. In this way, it produces better features for the tracking head. We also propose a novel dual-image position encoding technique that simultaneously encodes the position information of both the search region and template images. The HiT model achieves promising speed with competitive performance. For instance, it runs at 61 frames per second (fps) on the Nvidia Jetson AGX edge device. Furthermore, HiT attains 64.6% AUC on the LaSOT benchmark, surpassing all previous efficient trackers. Code and models are available at https://github.com/kangben258/HiT. Ben Kang, Xin Chen 0032, Dong Wang 0004, Houwen Peng, Huchuan Lu |
ICCV | 3 |
| 2023 | Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior RefinementabstractThe popularity of Contrastive Language-Image Pretraining (CLIP) has propelled its application to diverse downstream vision tasks. To improve its capacity on downstream tasks, few-shot learning has become a widelya-dopted technique. However, existing methods either exhibit limited performance or suffer from excessive learnable parameters. In this paper, we propose APE, an Adaptive Prior rEfinement method for CLIP’s pre-trained knowledge, which achieves superior accuracy with high computational efficiency. Via a prior refinement module, we analyze the inter-class disparity in the downstream data and decouple the domain-specific knowledge from the CLIP-extracted cache model. On top of that, we introduce two model variants, a training-free APE and a training-required APE-T. We explore the trilateral affinities between the test image, prior cache model, and textual representations, and only enable a lightweight category-residual module to be trained. For the average accuracy over 11 benchmarks, both APE and APE-T attain state-of-the-art and respectively outperform the second-best by +1.59% and +1.99% under 16 shots with ×30 less learnable parameters. Code is available at https://github.com/yangyangyang127/APE. Renrui Zhang, Bowei He, Aojun Zhou, Dong Wang 0004, Bin Zhao 0001, Peng Gao 0007 |
ICCV | 5 |
| 2023 | High-Performance Transformer TrackingabstractCorrelation has a critical role in the tracking field, especially in recent popular Siamese-based trackers. The correlation operation is a simple fusion method that considers the similarity between the template and the search region. However, the correlation operation is a local linear matching process, losing semantic information and easily falling into a local optimum, which may be the bottleneck in designing high-accuracy tracking algorithms. In this work, to determine whether a better feature fusion method exists than correlation, a novel attention-based feature fusion network, inspired by the transformer, is presented. This network effectively combines the template and search region features using attention mechanism. Specifically, the proposed method includes an ego-context augment module based on self-attention and a cross-feature augment module based on cross-attention. First, we present a transformer tracking (named TransT) method based on the Siamese-like feature extraction backbone, the designed attention-based fusion mechanism, and the classification and regression heads. Based on the TransT baseline, we also design a segmentation branch to generate the accurate mask. Finally, we propose a stronger version of TransT by extending it with a multi-template scheme and an IoU prediction head, named TransT-M. Experiments show that our TransT and TransT-M methods achieve promising results on seven popular benchmarks. Code and models are available at https://github.com/chenxin-dlut/TransT-M. Xin Chen 0032, Bin Yan 0004, Jiawen Zhu 0003, Huchuan Lu, Xiang Ruan, Dong Wang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Effective Local and Global Search for Fast Long-Term TrackingabstractCompared with short-term tracking, long-term tracking remains a challenging task that usually requires the tracking algorithm to track targets within a local region and re-detect targets over the entire image. However, few works have been done and their performances have also been limited. In this paper, we present a novel robust and real-time long-term tracking framework based on the proposed local search module and re-detection module. The local search module consists of an effective bounding box regressor to generate a series of candidate proposals and a target verifier to infer the optimal candidate with its confidence score. For local search, we design a long short-term updated scheme to improve the target verifier. The verification capability of the tracker can be improved by using several templates updated at different times. Based on the verification scores, our tracker determines whether the tracked object is present or absent and then chooses the tracking strategies of local or global search, respectively, in the next frame. For global re-detection, we develop a novel re-detection module that can estimate the target position and target size for a given base tracker. We conduct a series of experiments to demonstrate that this module can be flexibly integrated into many other tracking algorithms for long-term tracking and that it can improve long-term tracking performance effectively. Numerous experiments and discussions are conducted on several popular tracking datasets, including VOT, OxUvA, TLP, and LaSOT. The experimental results demonstrate that the proposed tracker achieves satisfactory performance with a real-time speed. Code is available at https://github.com/difhnp/ELGLT. Haojie Zhao, Bin Yan 0004, Dong Wang 0004, Xuesheng Qian, Xiaoyun Yang, Huchuan Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Robust Online Tracking With Meta-UpdaterabstractIn a sequence, the appearance of both the target and background often changes dramatically. Offline-trained models may not handle huge appearance variations well, causing tracking failures. Most discriminative trackers address this issue by introducing an online update scheme, making the model dynamically adapt the changes of the target and background. Although the online update scheme plays an important role in improving the tracker's accuracy, it inevitably pollutes the model with noisy observation samples. It is necessary to reduce the risk of the online update scheme for better tracking. In this work, we propose a novel offline-trained Meta-Updater to address an important but unsolved problem: Is the tracker ready for updating in the current frame? The proposed module can effectively integrate geometric, discriminative, and appearance cues in a sequential manner, and then mine the sequential information with a designed cascaded LSTM module. Moreover, we strengthen the effect of appearance information on the module, i.e., the additional local outlier factor is introduced to integrate into a newly designed network. We integrate our meta-updater into eight different types of online update trackers. Extensive experiments on four long-term and two short-term tracking benchmarks demonstrate that our meta-updater is effective and has strong generalization ability. Jie Zhao 0014, Kenan Dai, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Transformer vision-language tracking via proxy token guided cross-modal fusion
Haojie Zhao, Xiao Wang 0014, Dong Wang 0004, Huchuan Lu, Xiang Ruan |
Pattern Recognit. Lett. | 3 |
| 2023 | Multiscale Latent-Guided Entropy Model for LiDAR Point Cloud CompressionabstractThe non-uniform distribution and extremely sparse nature of the LiDAR point cloud (LPC) bring significant challenges to its high-efficient compression. This paper proposes a novel end-to-end, fully-factorized deep framework that represents the original LiDAR point cloud into an octree structure and hierarchically constructs the octree entropy model in layers. The proposed framework utilizes a hierarchical latent variable as side information to encapsulate the sibling and ancestor dependence, which provides sufficient context information for the modeling of point cloud distribution while enabling the parallel encoding and decoding of octree nodes in the same layer. Besides, we propose a residual coding framework for the compression of the latent variable, which explores the spatial correlation of each layer by progressive downsampling, and model the corresponding residual with a fully-factorized entropy model. Furthermore, we propose soft addition and subtraction for residual coding to improve network flexibility. The comprehensive experiment results on the LiDAR benchmark SemanticKITTI and MPEG-specified dataset Ford demonstrate that our proposed framework achieves state-of-the-art performance among all the previous LPC frameworks. Besides, our end-to-end, fully-factorized framework is proved by experiment to be high-parallelized and time-efficient, which saves more than 99.8% of decoding time compared to previous state-of-the-art methods on LPC compression. Tingyu Fan, Linyao Gao, Yiling Xu, Dong Wang 0004, Zhu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Visible-Thermal UAV Tracking: A Large-Scale Benchmark and New BaselineabstractWith the popularity of multi-modal sensors, visible-thermal (RGB-T) object tracking is to achieve robust performance and wider application scenarios with the guidance of objects' temperature information. However, the lack of paired training samples is the main bottleneck for unlocking the power of RGB-T tracking. Since it is laborious to collect high-quality RGB-T sequences, recent benchmarks only provide test sequences. In this paper, we construct a large-scale benchmark with high diversity for visible-thermal UAV tracking (VTUAV), including 500 sequences with 1.7 million high-resolution (1920* 1080 pixels) frame pairs. In addition, comprehensive applications (short-term tracking, long-term tracking and segmentation mask prediction) with diverse categories and scenes are considered for exhaustive evaluation. Moreover, we provide a coarse-to-fine attribute annotation, where frame-level attributes are provided to exploit the potential of challenge-specific trackers. In addition, we design a new RGB-T baseline, named Hierarchical Multi-modal Fusion Tracker (HMFT), which fuses RGB-T data in various levels. Numerous experiments on several datasets are conducted to reveal the effectiveness of HMFT and the complement of different fusion types. The project is available at here. Jie Zhao 0014, Dong Wang 0004, Huchuan Lu, Xiang Ruan |
CVPR | 3 |
| 2022 | Towards Grand Unification of Object Tracking
Bin Yan 0004, Yi Jiang 0009, Peize Sun, Dong Wang 0004, Zehuan Yuan, Ping Luo 0002, Huchuan Lu |
ECCV (21) | 4 |
| 2022 | PWPROP: A Progressive Weighted Adaptive Method for Training Deep Neural NetworksabstractIn recent years, adaptive optimization methods for deep learning have attracted considerable attention. AMSGRAD indicates that the adaptive methods may be hard to converge to optimal solutions of some convex problems due to the divergence of its adaptive learning rate as in ADAM. However, we find that AMSGRAD may generalize worse than ADAM for some deep learning tasks. We first show that AMSGRAD may not find a flat minimum. So how can we design an optimization method to find a flat minimum with low training loss? Few works focus on this important problem. We propose a novel progressive weighted adaptive optimization algorithm, called PWPROP, with fewer hyperparameters than its counterparts such as ADAM. By intuitively constructing a “sharp-flat minima” model, we show that how different second-order estimates affect the ability to escape a sharp minimum. Moreover, we also prove that PWPROP can address the non-convergence issue of ADAM and has a sublinear convergence rate for non-convex problems. Extensive experimental results show that PWPROP is effective and suitable for various deep learning architectures such as Transformer, and achieves state-of-the-art results. Dong Wang 0004, Huatian Zhang 0001, Fanhua Shang, Hongying Liu 0001, Yuanyuan Liu 0001, Shengmei Shen |
ICTAI | 1 |
| 2022 | D-DPCC: Deep Dynamic Point Cloud Compression via 3D Motion PredictionabstractThe non-uniformly distributed nature of the 3D Dynamic Point Cloud (DPC) brings significant challenges to its high-efficient inter-frame compression. This paper proposes a novel 3D sparse convolution-based Deep Dynamic Point Cloud Compression (D-DPCC) network to compensate and compress the DPC geometry with 3D motion estimation and motion compensation in the feature space. In the proposed D-DPCC network, we design a Multi-scale Motion Fusion (MMF) module to accurately estimate the 3D optical flow between the feature representations of adjacent point cloud frames. Specifically, we utilize a 3D sparse convolution-based encoder to obtain the latent representation for motion estimation in the feature space and introduce the proposed MMF module for fused 3D motion embedding. Besides, for motion compensation, we propose a 3D Adaptively Weighted Interpolation (3DAWI) algorithm with a penalty coefficient to adaptively decrease the impact of distant neighbours. We compress the motion embedding and the residual with a lossy autoencoder-based network. To our knowledge, this paper is the first work proposing an end-to-end deep dynamic point cloud compression framework. The experimental result shows that the proposed D-DPCC framework achieves an average 76% BD-Rate (Bjontegaard Delta Rate) gains against state-of-the-art Video-based Point Cloud Compression (V-PCC) v13 in inter mode. Tingyu Fan, Linyao Gao, Yiling Xu, Zhu Li 0001, Dong Wang 0004 |
IJCAI | 5 |
| 2022 | Balanced Gradient Penalty Improves Deep Long-Tailed LearningabstractIn recent years, deep learning has achieved a great success in various image recognition tasks. However, the long-tailed setting over a semantic class plays a leading role in real-world applications. Common methods focus on optimization on balanced distribution or naive models. Few works explore long-tailed learning from a deep learning-based generalization perspective. The loss landscape on long-tailed learning is first investigated in this work. Empirical results show that sharpness-aware optimizers work not well on long-tailed learning. Because they do not take class priors into consideration, and they fail to improve performance of few-shot classes. To better guide the network and explicitly alleviate sharpness without extra computational burden, we develop a universal Balanced Gradient Penalty (BGP) method. Surprisingly, our BGP method does not need the detailed class priors and preserves privacy. Our new algorithm BGP, as a regularization loss, can achieve the state-of-the-art results on various image datasets (i.e., CIFAR-LT, ImageNet-LT and iNaturalist-2018) in the settings of different imbalance ratios. Dong Wang 0004, Liangji Fang, Fanhua Shang, Yuanyuan Liu 0001, Hongying Liu 0001 |
ACM Multimedia | 1 |
| 2022 | Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingabstractMasked Autoencoders (MAE) have shown great potentials in self-supervised pre-training for language and 2D image transformers. However, it still remains an open question on how to exploit masked autoencoding for learning 3D representations of irregular point clouds. In this paper, we propose Point-M2AE, a strong Multi-scale MAE pre-training framework for hierarchical self-supervised learning of 3D point clouds. Unlike the standard transformer in MAE, we modify the encoder and decoder into pyramid architectures to progressively model spatial geometries and capture both fine-grained and high-level semantics of 3D shapes. For the encoder that downsamples point tokens by stages, we design a multi-scale masking strategy to generate consistent visible regions across scales, and adopt a local spatial self-attention mechanism during fine-tuning to focus on neighboring patterns. By multi-scale token propagation, the lightweight decoder gradually upsamples point tokens with complementary skip connections from the encoder, which further promotes the reconstruction from a global-to-local perspective. Extensive experiments demonstrate the state-of-the-art performance of Point-M2AE for 3D representation learning. With a frozen encoder after pre-training, Point-M2AE achieves 92.9% accuracy for linear SVM on ModelNet40, even surpassing some fully trained methods. By fine-tuning on downstream tasks, Point-M2AE achieves 86.43% accuracy on ScanObjectNN, +3.36% to the second-best, and largely benefits the few-shot classification, part segmentation and 3D object detection with the hierarchical pre-training scheme. Code is available at https://github.com/ZrrSkywalker/Point-M2AE. Renrui Zhang, Peng Gao 0007, Rongyao Fang, Bin Zhao 0001, Dong Wang 0004, Yu Qiao 0001, Hongsheng Li 0001 |
NeurIPS | 6 |
| 2022 | Vision-Based Anti-UAV Detection and TrackingabstractUnmanned aerial vehicles (UAV) have been widely used in various fields, and their invasion of security and privacy has aroused social concern. Several detection and tracking systems for UAVs have been introduced in recent years, but most of them are based on radio frequency, radar, and other media. We assume that the field of computer vision is mature enough to detect and track invading UAVs. Thus we propose a visible light mode dataset called Dalian University of Technology Anti-UAV dataset, DUT Anti-UAV for short. It contains a detection dataset with a total of 10,000 images and a tracking dataset with 20 videos that include short-term and long-term sequences. All frames and images are manually annotated precisely. We use this dataset to train several existing detection algorithms and evaluate the algorithms’ performance. Several tracking methods are also tested on our tracking dataset. Furthermore, we propose a clear and simple tracking algorithm combined with detection that inherits the detector’s high precision. Extensive experiments show that the tracking performance is improved considerably after fusing detection, thus providing a new attempt at UAV tracking using our dataset. The datasets and results are publicly available at:https://github.com/wangdongdut/DUT-Anti-UAV. Jie Zhao 0014, Jingshu Zhang, Dongdong Li 0004, Dong Wang 0004 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Transformer TrackingabstractCorrelation acts as a critical role in the tracking field, especially in recent popular Siamese-based trackers. The correlation operation is a simple fusion manner to consider the similarity between the template and the search region. However, the correlation operation itself is a local linear matching process, leading to lose semantic information and fall into local optimum easily, which may be the bottleneck of designing high-accuracy tracking algorithms. Is there any better feature fusion method than correlation? To address this issue, inspired by Transformer, this work presents a novel attention-based feature fusion network, which effectively combines the template and search region features solely using attention. Specifically, the proposed method includes an ego-context augment module based on self-attention and a cross-feature augment module based on cross-attention. Finally, we present a Transformer tracking (named TransT) method based on the Siamese-like feature extraction backbone, the designed attention-based fusion mechanism, and the classification and regression head. Experiments show that our TransT achieves very promising results on six challenging datasets, especially on large-scale LaSOT, TrackingNet, and GOT-10k benchmarks. Our tracker runs at approximatively 50 fps on GPU. Code and models are available at https://github.com/chenxin-dlut/TransT. Xin Chen 0032, Bin Yan 0004, Jiawen Zhu 0003, Dong Wang 0004, Xiaoyun Yang, Huchuan Lu |
CVPR | 4 |
| 2021 | LightTrack: Finding Lightweight Neural Networks for Object Tracking via One-Shot Architecture SearchabstractObject tracking has achieved significant progress over the past few years. However, state-of-the-art trackers become increasingly heavy and expensive, which limits their deployments in resource-constrained applications. In this work, we present LightTrack, which uses neural architecture search (NAS) to design more lightweight and efficient object trackers. Comprehensive experiments show that our LightTrack is effective. It can find trackers that achieve superior performance compared to handcrafted SOTA trackers, such as SiamRPN++ [30] and Ocean [56], while using much fewer model Flops and parameters. Moreover, when deployed on resource-constrained mobile chipsets, the discovered trackers run much faster. For example, on Snapdragon 845 Adreno GPU, LightTrack runs 12× faster than Ocean, while using 13× fewer parameters and 38× fewer Flops. Such improvements might narrow the gap between academic models and industrial deployments in object tracking task. LightTrack is released at here. Bin Yan 0004, Houwen Peng, Dong Wang 0004, Jianlong Fu, Huchuan Lu |
CVPR | 4 |
| 2021 | Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box EstimationabstractVisual object tracking aims to precisely estimate the bounding box for the given target, which is a challenging problem due to factors such as deformation and occlusion. Many recent trackers adopt the multiple-stage strategy to improve bounding box estimation. These methods first coarsely locate the target and then refine the initial prediction in the following stages. However, existing approaches still suffer from limited precision, and the coupling of different stages severely restricts the method’s transferability. This work proposes a novel, flexible, and accurate refinement module called Alpha-Refine (AR), which can significantly improve the base trackers’ box estimation quality. By exploring a series of design options, we conclude that the key to successful refinement is extracting and maintaining detailed spatial information as much as possible. Following this principle, Alpha-Refine adopts a pixel-wise correlation, a corner prediction head, and an auxiliary mask head as the core components. Comprehensive experiments on TrackingNet, LaSOT, GOT-10K, and VOT2020 benchmarks with multiple base trackers show that our approach significantly improves the base tracker’s performance with little extra latency. The proposed Alpha-Refine method leads to a series of strengthened trackers, among which the ARSiamRPN (AR strengthened SiamRPNpp) and the ARDiMP50 (AR strengthened DiMP50) achieve good efficiency-precision trade-off, while the ARDiMPsuper (AR strengthened DiMPsuper) achieves very competitive performance at a realtime speed. Code and pretrained models are available at https://github.com/MasterBin-IIAU/AlphaRefine. Bin Yan 0004, Xinyu Zhang 0017, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
CVPR | 3 |
| 2021 | Learning Spatio-Temporal Transformer for Visual TrackingabstractIn this paper, we present a new tracking architecture with an encoder-decoder transformer as the key component. The encoder models the global spatio-temporal feature dependencies between target objects and search regions, while the decoder learns a query embedding to predict the spatial positions of the target objects. Our method casts object tracking as a direct bounding box prediction problem, without using any proposals or predefined anchors. With the encoder-decoder transformer, the prediction of objects just uses a simple fully-convolutional network, which estimates the corners of objects directly. The whole method is end-to-end, does not need any postprocessing steps such as cosine window and bounding box smoothing, thus largely simplifying existing tracking pipelines. The proposed tracker achieves state-of-the-art performance on multiple challenging short-term and long-term benchmarks, while running at real-time speed, being 6× faster than Siam R-CNN [54]. Code and models are open-sourced at https://github.com/researchmm/Stark. Bin Yan 0004, Houwen Peng, Jianlong Fu, Dong Wang 0004, Huchuan Lu |
ICCV | 4 |
| 2021 | Video Annotation for Visual Tracking via Selection and RefinementabstractDeep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-refinement strategy to automatically improve the preliminary annotations generated by tracking algorithms. A temporal assessment network (T-Assess Net) is proposed which is able to capture the temporal coherence of target locations and select reliable tracking results by measuring their quality. Meanwhile, a visual-geometry refinement network (VG-Refine Net) is also designed to further enhance the selected tracking results by considering both target appearance and temporal geometry constraints, allowing inaccurate tracking results to be corrected. The combination of the above two networks provides a principled approach to ensure the quality of automatic video annotation. Experiments on large scale tracking benchmarks demonstrate that our method can deliver highly accurate bounding box annotations and significantly reduce human labor by 94.0%, yielding an effective means to further boost tracking performance with augmented training data. Kenan Dai, Jie Zhao 0014, Lijun Wang 0001, Dong Wang 0004, Huchuan Lu, Xuesheng Qian, Xiaoyun Yang |
ICCV | 4 |
| 2021 | Pyramid Spatial-Temporal Aggregation for Video-based Person Re-IdentificationabstractVideo-based person re-identification aims to associate the video clips of the same person across multiple non-overlapping cameras. Spatial-temporal representations can provide richer and complementary information between frames, which are crucial to distinguish the target person when occlusion occurs. This paper proposes a novel Pyramid Spatial-Temporal Aggregation (PSTA) framework to aggregate the frame-level features progressively and fuse the hierarchical temporal features into a final video-level representation. Thus, short-term and long-term temporal information could be well exploited by different hierarchies. Furthermore, a Spatial-Temporal Aggregation Module (STAM) is proposed to enhance the aggregation capability of PSTA. It mainly consists of two novel attention blocks: Spatial Reference Attention (SRA) and Temporal Reference Attention (TRA). SRA explores the spatial correlations within a frame to determine the attention weight of each location. While TRA extends SRA with the correlations between adjacent frames, temporal consistency information can be fully explored to suppress the interference features and strengthen the discriminative ones. Extensive experiments on several challenging benchmarks demonstrate the effectiveness of the proposed PSTA, and our full model reaches 91.5% and 98.3% Rank-1 accuracy on MARS and DukeMTMC-VID benchmarks. The source code is available at https://github.com/WangYQ9/VideoReID-PSTA. Yingquan Wang, Shang Gao 0012, Xia Geng, Hu Lu, Dong Wang 0004 |
ICCV | 6 |
| 2021 | Learning Adaptive Attribute-Driven Representation for Real-Time RGB-T Tracking
Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
Int. J. Comput. Vis. | 2 |
| 2021 | Learning Regression and Verification Networks for Robust Long-term Tracking
Lijun Wang 0001, Dong Wang 0004, Jinqing Qi, Huchuan Lu |
Int. J. Comput. Vis. | 3 |
| 2021 | Deep mutual learning for visual object tracking
Haojie Zhao, Gang Yang 0002, Dong Wang 0004, Huchuan Lu |
Pattern Recognit. | 3 |
| 2021 | Jointly Modeling Motion and Appearance Cues for Robust RGB-T TrackingabstractIn this study, we propose a novel RGB-T tracking framework by jointly modeling both appearance and motion cues. First, to obtain a robust appearance model, we develop a novel late fusion method to infer the fusion weight maps of both RGB and thermal (T) modalities. The fusion weights are determined by using offline-trained global and local multimodal fusion networks, and then adopted to linearly combine the response maps of RGB and T modalities. Second, when the appearance cue is unreliable, we comprehensively take motion cues, i.e., target and camera motions, into account to make the tracker robust. We further propose a tracker switcher to switch the appearance and motion trackers flexibly. Numerous results on three recent RGB-T tracking datasets show that the proposed tracker performs significantly better than other state-of-the-art algorithms. Jie Zhao 0014, Chunjuan Bo, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
IEEE Trans. Image Process. | 4 |
| 2020 | High-Performance Long-Term Tracking With Meta-UpdaterabstractLong-term visual tracking has drawn increasing attention because it is much closer to practical applications than short-term tracking. Most top-ranked long-term trackers adopt the offline-trained Siamese architectures, thus,they cannot benefit from great progress of short-term trackers with online update. However, it is quite risky to straightforwardly introduce online-update-based trackers to solve the long-term problem, due to long-term uncertain and noisy observations. In this work, we propose a novel offline-trained Meta-Updater to address an important but unsolved problem: Is the tracker ready for updating in the current frame? The proposed meta-updater can effectively integrate geometric, discriminative, and appearance cues in a sequential manner, and then mine the sequential information with a designed cascaded LSTM module. Our meta-updater learns a binary output to guide the tracker’s update and can be easily embedded into different trackers. This work also introduces a long-term tracking framework consisting of an online local tracker, an online verifier, a SiamRPN-based re-detector, and our meta-updater. Numerous experimental results on the VOT2018LT,VOT2019LT, OxUvALT, TLP, and LaSOT benchmarks show that our tracker performs remarkably better than other competing algorithms. Our project is available on the website: https://github.com/Daikenan/LTMU. Kenan Dai, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
CVPR | 3 |
| 2020 | Cooling-Shrinking Attack: Blinding the Tracker With Imperceptible NoisesabstractAdversarial attack of CNN aims at deceiving models to misbehave by adding imperceptible perturbations to images. This feature facilitates to understand neural networks deeply and to improve the robustness of deep learning models. Although several works have focused on attacking image classifiers and object detectors, an effective and efficient method for attacking single object trackers of any target in a model-free way remains lacking. In this paper, a cooling-shrinking attack method is proposed to deceive state-of-the-art SiameseRPN-based trackers. An effective and efficient perturbation generator is trained with a carefully designed adversarial loss, which can simultaneously cool hot regions where the target exists on the heatmaps and force the predicted bounding box to shrink, making the tracked target invisible to trackers. Numerous experiments on OTB100, VOT2018, and LaSOT datasets show that our method can effectively fool the state-of-the-art SiameseRPN++ tracker by adding small perturbations to the template or the search regions. Besides, our method has good transferability and is able to deceive other top-performance trackers such as DaSiamRPN, DaSiamRPN-UpdateNet, and DiMP. The source codes are available at https://github.com/MasterBin-IIAU/CSA. Bin Yan 0004, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
CVPR | 2 |
| 2020 | Online Filtering Training Samples for Robust Visual TrackingabstractIn recent years, discriminative trackers show its great tracking performance, that is mainly due to the online updating using samples collected during tracking. The model could adapt appearance changes of objects and the background well after updating. But these trackers have a serious disadvantage that wrong samples may cause severe model degradation. Most of the training samples in the tracking phase are obtained according to the tracking result of the current frame. Wrong training samples will be collected when the tracking result is inaccurate, seriously affecting the discrimination ability of the model. Besides, partial occlusion also leads to the same problem. In this paper, we propose an optimization module named MetricNet for online filtering training samples. It applies a matching network containing the classification and distance branches, and uses multiple metric methods for different type samples. MetricNet optimizes the training sample set by recognizing wrong and redundant samples, thereby improving the tracking performance. The proposed MetricNet can be regarded as an independent optimization module and integrated into all discriminative trackers updated online. Extensive experiments on three tracking datasets show its effectiveness and generalization ability. After applying MetricNet to MDNet, the tracking result is increased by 5.3% in terms of the success plot on the LaSOT dataset. Our project is available at https://github.com/zj5559/MetricNet. Jie Zhao 0014, Kenan Dai, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
ACM Multimedia | 3 |
| 2020 | Deep-Sea Organisms Tracking Using Dehazing and Deep Learning
Huimin Lu 0001, Tomoki Uemura, Dong Wang 0004, Jihua Zhu, Zi Huang, Hyoungseop Kim |
Mob. Networks Appl. | 3 |
| 2020 | Correction to: Deep-Sea Organisms Tracking Using Dehazing and Deep Learning
Huimin Lu 0001, Tomoki Uemura, Dong Wang 0004, Jihua Zhu, Zi Huang, Hyoungseop Kim |
Mob. Networks Appl. | 3 |
| 2020 | Defocus Blur Detection via Multi-Stream Bottom-Top-Bottom NetworkabstractDefocus blur detection (DBD) is aimed to estimate the probability of each pixel being in-focus or out-of-focus. This process has been paid considerable attention due to its remarkable potential applications. Accurate differentiation of homogeneous regions and detection of low-contrast focal regions, as well as suppression of background clutter, are challenges associated with DBD. To address these issues, we propose a multi-stream bottom-top-bottom fully convolutional network (BTBNet), which is the first attempt to develop an end-to-end deep network to solve the DBD problems. First, we develop a fully convolutional BTBNet to gradually integrate nearby feature levels of bottom to top and top to bottom. Then, considering that the degree of defocus blur is sensitive to scales, we propose multi-stream BTBNets that handle input images with different scales to improve the performance of DBD. Finally, a cascaded DBD map residual learning architecture is designed to gradually restore finer structures from the small scale to the large scale. To promote further study and evaluation of the DBD models, we construct a new database of 1100 challenging images and their pixel-wise defocus blur annotations. Experimental results on the existing and our new datasets demonstrate that the proposed method achieves significantly better performance than other state-of-the-art algorithms. Wenda Zhao 0003, Fan Zhao 0006, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Visual tracking by dynamic matching-classification network switching
Peixia Li, Dong Wang 0004, Huchuan Lu |
Pattern Recognit. | 3 |
| 2020 | Non-rigid object tracking via deep multi-scale spatial-temporal discriminative saliency maps
Wei Liu 0044, Dong Wang 0004, Yinjie Lei, Hongyu Wang 0001, Huchuan Lu |
Pattern Recognit. | 3 |
| 2019 | Visual Tracking via Adaptive Spatially-Regularized Correlation FiltersabstractIn this work, we propose a novel adaptive spatially-regularized correlation filters (ASRCF) model to simultaneously optimize the filter coefficients and the spatial regularization weight. First, this adaptive spatial regularization scheme could learn an effective spatial weight for a specific object and its appearance variations, and therefore result in more reliable filter coefficients during the tracking process. Second, our ASRCF model can be effectively optimized based on the alternating direction method of multipliers, where each subproblem has the closed-from solution. Third, our tracker applies two kinds of CF models to estimate the location and scale respectively. The location CF model exploits ensembles of shallow and deep features to determine the optimal position accurately. The scale CF model works on multi-scale shallow features to estimate the optimal scale efficiently. Extensive experiments on five recent benchmarks show that our tracker performs favorably against many state-of-the-art algorithms, with real-time performance of 28fps. Kenan Dai, Dong Wang 0004, Huchuan Lu |
CVPR | 2 |
| 2019 | ROI Pooled Correlation Filters for Visual TrackingabstractThe ROI (region-of-interest) based pooling method performs pooling operations on the cropped ROI regions for various samples and has shown great success in the object detection methods. It compresses the model size while preserving the localization accuracy, thus it is useful in the visual tracking field. Though being effective, the ROI-based pooling operation is not yet considered in the correlation filter formula. In this paper, we propose a novel ROI pooled correlation filter (RPCF) algorithm for robust visual tracking. Through mathematical derivations, we show that the ROI-based pooling can be equivalently achieved by enforcing additional constraints on the learned filter weights, which makes the ROI-based pooling feasible on the virtual circular samples. Besides, we develop an efficient joint training formula for the proposed correlation filter algorithm, and derive the Fourier solvers for efficient model training. Finally, we evaluate our RPCF tracker on OTB-2013, OTB-2015 and VOT-2017 benchmark datasets. Experimental results show that our tracker performs favourably against other state-of-the-art trackers. Yuxuan Sun 0003, Dong Wang 0004, You He 0002, Huchuan Lu |
CVPR | 3 |
| 2019 | A Mutual Learning Method for Salient Object Detection With Intertwined Multi-SupervisionabstractThough deep learning techniques have made great progress in salient object detection recently, the predicted saliency maps still suffer from incomplete predictions due to the internal complexity of objects and inaccurate boundaries caused by strides in convolution and pooling operations. To alleviate these issues, we propose to train saliency detection networks by exploiting the supervision from not only salient object detection, but also foreground contour detection and edge detection. First, we leverage salient object detection and foreground contour detection tasks in an intertwined manner to generate saliency maps with uniform highlight. Second, the foreground contour and edge detection tasks guide each other simultaneously, thereby leading to preciser foreground contour prediction and reducing the local noises for edge prediction. In addition, we develop a novel mutual learning module (MLM) which serves as the building block of our method. Each MLM consists of multiple network branches trained in a mutual learning manner, which improves the performance by a large margin. Extensive experiments on seven challenging datasets demonstrate that the proposed method has delivered state-of-the-art results in both salient object detection and edge detection. Runmin Wu, Mengyang Feng, Wenlong Guan, Dong Wang 0004, Huchuan Lu, Errui Ding |
CVPR | 4 |
| 2019 | Language Person Search with Mutually Connected Classification LossabstractIn this work, we develop an effective person search algorithm with natural language descriptions. The contributions of this work mainly include two aspects. First, we design a baseline language person search framework including three basic components: a deep CNN model to extract visual features, a bi-directional LSTM to encode language descriptions and the triplet loss to conduct cross-modal feature embedding. Second, we propose a novel mutually connected classification loss to fully exploit the identity-level information, which not only introduces the identification information into both image and language descriptions but also encourages the cross-modal classification probabilities of the same identity to be more similar. The experimental results on the CUHK-PEDES dataset demonstrate that our method achieves significantly better performance than other state-of-the-art algorithms. Chunjuan Bo, Dong Wang 0004, Yunwei Qi, Huchuan Lu |
ICASSP | 3 |
| 2019 | Online Single Person Tracking for Unmanned Aerial Vehicles: Benchmark and New BaselineabstractOnline tracking a specific person from a low-altitude unmanned aerial vehicle (UAV) is a very interesting and challenging problem to be solved. However, there exists no large-scale aerial video dataset regarding this online single person tracking (OSPT) task. To promote the study of the OSPT problem in UAV, we first construct a new benchmark dataset including 100 fully annotated aerial videos with nearly 130K frames and 11 challenging factors. Second, we evaluate several state-of-the-art online trackers with real-time performance using our dataset, considering the potential applications in the UAV platform. In addition, with respect to the OSPT problem, we attempt to design a new baseline method with the combination of tracking, detection and re-identification and conduct detailed analysis of different components. This method achieves much better performance than the existing online trackers, which will serve as a new baseline for our benchmark. Zhihui Wang 0001, Dong Wang 0004, Yunwei Qi, Huchuan Lu |
ICASSP | 3 |
| 2019 | GradNet: Gradient-Guided Network for Visual Object TrackingabstractThe fully-convolutional siamese network based on template matching has shown great potentials in visual tracking. During testing, the template is fixed with the initial target feature and the performance totally relies on the general matching ability of the siamese network. However, this manner cannot capture the temporal variations of targets or background clutter. In this work, we propose a novel gradient-guided network to exploit the discriminative information in gradients and update the template in the siamese network through feed-forward and backward operations. To be specific, the algorithm can utilize the information from the gradient to update the template in the current frame. In addition, a template generalization training method is proposed to better use gradient information and avoid overfitting. To our knowledge, this work is the first attempt to exploit the information in the gradient for template update in siamese-based trackers. Extensive experiments on recent benchmarks demonstrate that our method achieves better performance than other state-of-the-art trackers. Peixia Li, Wanli Ouyang, Dong Wang 0004, Xiaoyun Yang, Huchuan Lu |
ICCV | 4 |
| 2019 | 'Skimming-Perusal' Tracking: A Framework for Real-Time and Robust Long-Term TrackingabstractCompared with traditional short-term tracking, long-term tracking poses more challenges and is much closer to realistic applications. However, few works have been done and their performance have also been limited. In this work, we present a novel robust and real-time long-term tracking framework based on the proposed skimming and perusal modules. The perusal module consists of an effective bounding box regressor to generate a series of candidate proposals and a robust target verifier to infer the optimal candidate with its confidence score. Based on this score, our tracker determines whether the tracked object being present or absent, and then chooses the tracking strategies of local search or global search respectively in the next frame. To speed up the image-wide global search, a novel skimming module is designed to efficiently choose the most possible regions from a large number of sliding windows. Numerous experimental results on the VOT-2018 long-term and OxUvA long-term benchmarks demonstrate that the proposed method achieves the best performance and runs in real-time. The source codes are available at https://github.com/iiau-tracker/SPLT. Bin Yan 0004, Haojie Zhao, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang |
ICCV | 3 |
| 2019 | Scene Text Detection with Feature Pyramid Network and Linking SegmentsabstractScene text detection is one of the most challenging problems in computer vision and has attracted great interest. Different from generic object detection, scene text detection mainly suffers from the large variance of scale, aspect ratio, and orientation in scene text. In this paper, we propose an effective and efficient model (SEG-FPN) for scene text detection, which is based on Feature Pyramid Network (FPN) and Linking Segments (SegLink). We incorporate feature pyramid mechanism with Single Shot Detector (SSD) framework to deal with different scale texts, and link locally detectable elements to detect texts of different orientations and aspect ratios. Moreover, compared with SSD, we enlarge the feature map of deep layers to better localize the large texts and recognize the small texts accurately. Experiments on ICDAR2015 and ICDAR2013 datasets demonstrate that our method can achieve comparable performance in terms of both accuracy and time. Specifically, SEG-FPN achieves an f-measure of 0.820 at 10.3 fps for 1280*768 ICDAR 2015 Incidental text images, and an f-measure of 0.879 at 19.2 fps for 512*512 ICDAR 2013 focused scene text images. Rui Zhang 0056, Yongsheng Zhou, Dong Wang 0004 |
ICDAR | 4 |
| 2019 | ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on SignboardabstractChinese scene text reading is one of the most challenging problems in computer vision and has attracted great interest. Different from English text, Chinese has more than 6000 commonly used characters and Chinese characters can be arranged in various layouts with numerous fonts. The Chinese signboards in street view are a good choice for Chinese scene text images since they have different backgrounds, fonts and layouts. We organized a competition called ICDAR2019-ReCTS, which mainly focuses on reading Chinese text on signboard. This report presents the final results of the competition. A large-scale dataset of 25,000 annotated signboard images, in which all the text lines and characters are annotated with locations and transcriptions, were released. Four tasks, namely character recognition, text line recognition, text line detection and end-to-end recognition were set up. Besides, considering the Chinese text ambiguity issue, we proposed a multi ground truth (multi-GT) evaluation method to make evaluation fairer. The competition started on March 1, 2019 and ended on April 30, 2019. 262 submissions from 46 teams are received. Most of the participants come from universities, research institutes, and tech companies in China. There are also some participants from the United States, Australia, Singapore, and Korea. 21 teams submit results for Task 1, 23 teams submit results for Task 2, 24 teams submit results for Task 3, and 13 teams submit results for Task 4. The official website for the competition is http://rrc.cvc.uab.es/?ch=12. Rui Zhang 0056, Xiang Bai, Baoguang Shi, Dimosthenis Karatzas, Shijian Lu, C. V. Jawahar, Yongsheng Zhou, Qianyi Jiang, Nan Li 0071, Dong Wang 0004, Minghui Liao |
ICDAR | 14 |
| 2019 | Lightweight Deep Neural Network for Real-Time Visual Tracking with Mutual LearningabstractIn this work, we develop a real-time tracking algorithm with a lightweight deep neural network. The contributions of this work mainly include two aspects. First, we reformulate the discriminative correlation filter (DCF) based tracker as a fully convolutional neural network and design an effective end-to-end tracking framework. Second, we build our tracker with a pruned convolutional neural network, which is trained by a mutual learning approach to further improve the location accuracy. The proposed tracking algorithm can track objects at 60 FPS. Extensive experiments on OTB2013, OTB2015 and VOT2017 Real-time demonstrate that the proposed tracker performs favorably against state-of-the-art methods. Haojie Zhao, Gang Yang 0002, Dong Wang 0004, Huchuan Lu |
ICIP | 3 |
| 2019 | A Preliminary Study on Data Augmentation of Deep Learning for Image ClassificationabstractDeep learning models have a large number of free parameters that need to be calculated by effective training of the models on a great deal of training data to improve their generalization performance. However, data obtaining and labeling is expensive in practice. Data augmentation is one of the methods to alleviate this problem. In this paper, we conduct a preliminary study on how four variables (augmentation method, augmentation rate, size of basic dataset per label, and method combination) can affect the accuracy of deep learning for image classification. The study provides some guidelines: (1) altering the geometry of the images is not always better than those just lighting and color. (2) 2-3 times augmentation rate is good enough for training. (3) the combination of two geometry methods degrade the performance, while combinations with at least one photometric method, will improve the performance, especially when one method is a photometric method and another is a geometry method. (4) the sequence of methods in combination has little effect on the performance. Benlin Hu, Dong Wang 0004, Shu Zhang 0009, Zhenyu Chen 0001 |
Internetware | 3 |
| 2019 | Multi attention module for visual tracking
Peixia Li, Dong Wang 0004, Gang Yang 0002, Huchuan Lu |
Pattern Recognit. | 4 |
| 2019 | Multi-Focus Image Fusion With a Natural Enhancement via a Joint Multi-Level Deeply Supervised Convolutional Neural NetworkabstractCommon non-focused areas are often present in multi-focus images due to the limitation of the number of focused images. This factor severely degrades the fusion quality of multi-focus images. To address this problem, we propose a novel end-to-end multi-focus image fusion with a natural enhancement method based on deep convolutional neural network (CNN). Several end-to-end CNN architectures that are specifically adapted to this task are first designed and researched. On the basis of the observation that low-level feature extraction can capture low-frequency content, whereas high-level feature extraction effectively captures high-frequency details, we further combine multi-level outputs such that the most visually distinctive features can be extracted, fused, and enhanced. In addition, the multi-level outputs are simultaneously supervised during training to boost the performance of image fusion and enhancement. Extensive experiments show that the proposed method can deliver superior fusion and enhancement performance than the state-of-the-art methods in the presence of multi-focus images with common non-focused areas, anisotropic blur, and misregistration. Wenda Zhao 0003, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Video Person Re-Identification by Temporal Residual LearningabstractIn this paper, we propose a novel feature learning framework for video person re-identification (re-ID). The proposed framework largely aims to exploit the adequate temporal information of video sequences and tackle the poor spatial alignment of moving pedestrians. More specifically, for exploiting the temporal information, we design a temporal residual learning (TRL) module to simultaneously extract the generic and specific features of consecutive frames. The TRL module is equipped with two bi-directional LSTM (BiLSTM), which are respectively responsible to describe a moving person in different aspects, providing complementary information for better feature representations. To deal with the poor spatial alignment in video re- ID datasets, we propose a spatial-temporal transformer network (ST2N) module. Transformation parameters in the ST2N module are learned by leveraging the high-level semantic information of the current frame as well as the temporal context knowledge from other frames. The proposed ST2N module with less learnable parameters allows effective person alignments under significant appearance changes. Extensive experimental results on the largescale MARS, PRID2011, ILIDS-VID and SDU-VID datasets demonstrate that the proposed method achieves consistently superior performance and outperforms most of the very recent state-of-the-art methods. Ju Dai, Dong Wang 0004, Huchuan Lu, Hongyu Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Correlation Tracking via Joint Discrimination and Reliability LearningabstractFor visual tracking, an ideal filter learned by the correlation filter (CF) method should take both discrimination and reliability information. However, existing attempts usually focus on the former one while pay less attention to reliability learning. This may make the learned filter be dominated by the unexpected salient regions on the feature map, thereby resulting in model degradation. To address this issue, we propose a novel CF-based optimization problem to jointly model the discrimination and reliability information. First, we treat the filter as the element-wise product of a base filter and a reliability term. The base filter is aimed to learn the discrimination information between the target and backgrounds, and the reliability term encourages the final filter to focus on more reliable regions. Second, we introduce a local response consistency regular term to emphasize equal contributions of different regions and avoid the tracker being dominated by unreliable regions. The proposed optimization problem can be solved using the alternating direction method and speeded up in the Fourier domain. We conduct extensive experiments on the OTB-2013, OTB-2015 and VOT-2016 datasets to evaluate the proposed tracker. Experimental results show that our tracker performs favorably against other state-of-the-art trackers. Dong Wang 0004, Huchuan Lu, Ming-Hsuan Yang 0001 |
CVPR | 2 |
| 2018 | Learning Spatial-Aware Regressions for Visual TrackingabstractIn this paper, we analyze the spatial information of deep features, and propose two complementary regressions for robust visual tracking. First, we propose a kernelized ridge regression model wherein the kernel value is defined as the weighted sum of similarity scores of all pairs of patches between two samples. We show that this model can be formulated as a neural network and thus can be efficiently solved. Second, we propose a fully convolutional neural network with spatially regularized kernels, through which the filter kernel corresponding to each output channel is forced to focus on a specific region of the target. Distance transform pooling is further exploited to determine the effectiveness of each output channel of the convolution layer. The outputs from the kernelized ridge regression model and the fully convolutional neural network are combined to obtain the ultimate response. Experimental results on two benchmark datasets validate the effectiveness of the proposed method. Dong Wang 0004, Huchuan Lu, Ming-Hsuan Yang 0001 |
CVPR | 2 |
| 2018 | Defocus Blur Detection via Multi-Stream Bottom-Top-Bottom Fully Convolutional NetworkabstractDefocus blur detection (DBD) is the separation of in-focus and out-of-focus regions in an image. This process has been paid considerable attention because of its remarkable potential applications. Accurate differentiation of homogeneous regions and detection of low-contrast focal regions, as well as suppression of background clutter, are challenges associated with DBD. To address these issues, we propose a multi-stream bottom-top-bottom fully convolutional network (BTBNet), which is the first attempt to develop an end-to-end deep network for DBD. First, we develop a fully convolutional BTBNet to integrate low-level cues and high-level semantic information. Then, considering that the degree of defocus blur is sensitive to scales, we propose multi-stream BTBNets that handle input images with different scales to improve the performance of DBD. Finally, we design a fusion and recurrent reconstruction network to recurrently refine the preceding blur detection maps. To promote further study and evaluation of the DBD models, we construct a new database of 500 challenging images and their pixel-wise defocus blur annotations. Experimental results on the existing and our new datasets demonstrate that the proposed method achieves significantly better performance than other state-of-the-art algorithms. Wenda Zhao 0003, Fan Zhao 0006, Dong Wang 0004, Huchuan Lu |
CVPR | 3 |
| 2018 | Real-Time 'Actor-Critic' Tracking
Dong Wang 0004, Peixia Li, Huchuan Lu |
ECCV (7) | 2 |
| 2018 | Structured Siamese Network for Real-Time Visual Tracking
Lijun Wang 0001, Jinqing Qi, Dong Wang 0004, Mengyang Feng, Huchuan Lu |
ECCV (9) | 4 |
| 2018 | Motor Anomaly Detection for Unmanned Aerial Vehicles Using Reinforcement LearningabstractUnmanned aerial vehicles (UAVs) are used in many fields including weather observation, farming, infrastructure inspection, and monitoring of disaster areas. However, the currently available UAVs are prone to crashing. The goal of this paper is the development of an anomaly detection system to prevent the motor of the drone from operating at abnormal temperatures. In this anomaly detection system, the temperature of the motor is recorded using DS18B20 sensors. Then, using reinforcement learning, the motor is judged to be operating abnormally by a Raspberry Pi processing unit. A specially built user interface allows the activity of the Raspberry Pi to be tracked on a Tablet for observation purposes. The proposed system provides the ability to land a drone when the motor temperature exceeds an automatically generated threshold. The experimental results confirm that the proposed system can safely control the drone using information obtained from temperature sensors attached to the motor. Huimin Lu 0001, Yujie Li 0001, Shenglin Mu, Dong Wang 0004, Hyoungseop Kim, Seiichi Serikawa |
IEEE Internet Things J. | 4 |
| 2018 | Online single target tracking in WAMI: benchmark and evaluation
Dong Wang 0004, Meng Yi, Fan Yang 0035, Erik Blasch, Carolyn Sheaff, Genshe Chen, Haibin Ling |
Multim. Tools Appl. | 1 |
| 2018 | Spectral-spatial K-Nearest Neighbor approach for hyperspectral image classification
Chunjuan Bo, Huchuan Lu, Dong Wang 0004 |
Multim. Tools Appl. | 3 |
| 2018 | Deep visual tracking: Review and experimental comparison
Peixia Li, Dong Wang 0004, Lijun Wang 0001, Huchuan Lu |
Pattern Recognit. | 2 |
| 2018 | Robust linear representation via exploiting structure prior
Dong Wang 0004, Ran He 0001, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. | 1 |
| 2018 | Tracking With Static and Dynamic Structured Correlation FiltersabstractTracking methods based on correlation filters have recently attracted attention for achieving fast tracking. However, their performance is somewhat limited in long-term tracking tasks, especially in an occlusion situation. To address this issue, we propose a novel structured correlation filter, which depends on coupled interactions between a static model and a dynamic model. Specifically, the static model exploits the star graph to capture spatial information and provides an initial estimation. The dynamic model based on Bayesian inference uses the rough location as a reference to estimate the final target state. Then, the dynamic model provides a feedback to the static regarding their updates. Finally, the dynamic model provides a scale adaptivity mechanism, which makes the proposed tracker effectively deal with not only partial occlusion but also scale variation. Both qualitative and quantitative evaluations on challenging image sequences demonstrate that the proposed method performs favorably against the state-of-the-art tracking algorithms. Dong Wang 0004, Huchuan Lu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Multisensor Image Fusion and Enhancement in Spectral Total Variation DomainabstractMost existing image fusion methods assume that at least one input image contains high-quality information at any place of an observed scene. Thus, these fusion methods will fail if every input image is degraded. To address this issue, this study proposes a novel fusion framework that integrates image fusion based on spectral total variation (TV) method and image enhancement. For spatially varying multiscale decompositions generated by the spectral TV framework, this study verifies that the decomposition components can be modeled efficiently by tailed α-stable-based random variable distribution (TRD) rather than the commonly used Gaussian distribution. Consequently, salience and match measures based on TRD are proposed to fuse each sub-band decomposition. The spatial intensity information is also adopted to fuse the remainder of the image decomposition components. A sub-band adaptive gain function family based on TV spectrum and space variation is constructed for fused multiscale decompositions to enhance fused image simultaneously. Finally, numerous experiments with various multisensor image pairs are conducted to evaluate the proposed method. Experimental results show that even if the input images are degraded, the fused image obtained by the proposed method achieves significant improvement in terms of edge details and contrast while extracting the main features of the input images, thereby achieving better performance compared with the state-of-the-art methods. Wenda Zhao 0003, Huimin Lu 0001, Dong Wang 0004 |
IEEE Trans. Multim. | 3 |
| 2017 | Learning to Detect Salient Objects with Image-Level SupervisionabstractDeep Neural Networks (DNNs) have substantially improved the state-of-the-art in salient object detection. However, training DNNs requires costly pixel-level annotations. In this paper, we leverage the observation that image-level tags provide important cues of foreground salient objects, and develop a weakly supervised learning method for saliency detection using image-level tags only. The Foreground Inference Network (FIN) is introduced for this challenging task. In the first stage of our training method, FIN is jointly trained with a fully convolutional network (FCN) for image-level tag prediction. A global smooth pooling layer is proposed, enabling FCN to assign object category tags to corresponding object regions, while FIN is capable of capturing all potential foreground regions with the predicted saliency maps. In the second stage, FIN is fine-tuned with its predicted saliency maps as ground truth. For refinement of ground truth, an iterative Conditional Random Field is developed to enforce spatial label consistency and further boost performance. Our method alleviates annotation efforts and allows the usage of existing large scale training sets with image-level tags. Our model runs at 60 FPS, outperforms unsupervised ones with a large margin, and achieves comparable or even superior performance than fully supervised counterparts. Lijun Wang 0001, Huchuan Lu, Yifan Wang 0004, Mengyang Feng, Dong Wang 0004, Xiang Ruan |
CVPR | 5 |
| 2017 | Stepwise Metric Promotion for Unsupervised Video Person Re-identificationabstractThe intensive annotation cost and the rich but unlabeled data contained in videos motivate us to propose an unsupervised video-based person re-identification (re-ID) method. We start from two assumptions: 1) different video tracklets typically contain different persons, given that the tracklets are taken at distinct places or with long intervals; 2) within each tracklet, the frames are mostly of the same person. Based on these assumptions, this paper propose a stepwise metric promotion approach to estimate the identities of training tracklets, which iterates between cross-camera tracklet association and feature learning. Specifically, We use each training tracklet as a query, and perform retrieval in the cross-camera training set. Our method is built on reciprocal nearest neighbor search and can eliminate the hard negative label matches, i.e., the cross-camera nearest neighbors of the false matches in the initial rank list. The tracklet that passes the reciprocal nearest neighbor check is considered to have the same ID with the query. Experimental results on the PRID 2011, ILIDS-VID, and MARS datasets show that the proposed method achieves very competitive re-ID accuracy compared with its supervised counterparts. Zimo Liu, Dong Wang 0004, Huchuan Lu |
ICCV | 2 |
| 2017 | Amulet: Aggregating Multi-level Convolutional Features for Salient Object DetectionabstractFully convolutional neural networks (FCNs) have shown outstanding performance in many dense labeling problems. One key pillar of these successes is mining relevant information from features in convolutional layers. However, how to better aggregate multi-level convolutional feature maps for salient object detection is underexplored. In this work, we present Amulet, a generic aggregating multi-level convolutional feature framework for salient object detection. Our framework first integrates multi-level feature maps into multiple resolutions, which simultaneously incorporate coarse semantics and fine details. Then it adaptively learns to combine these feature maps at each resolution and predict saliency maps with the combined features. Finally, the predicted results are efficiently fused to generate the final saliency map. In addition, to achieve accurate boundary inference and semantic enhancement, edge-aware feature maps in low-level layers and the predicted results of low resolution features are recursively embedded into the learning framework. By aggregating multi-level convolutional features in this efficient and flexible manner, the proposed saliency model provides accurate salient object labeling. Comprehensive experiments demonstrate that our method performs favorably against state-of-the-art approaches in terms of near all compared evaluation metrics. Dong Wang 0004, Huchuan Lu, Hongyu Wang 0001, Xiang Ruan |
ICCV | 2 |
| 2017 | Learning Uncertain Convolutional Features for Accurate Saliency DetectionabstractDeep convolutional neural networks (CNNs) have delivered superior performance in many computer vision tasks. In this paper, we propose a novel deep fully convolutional network model for accurate salient object detection. The key contribution of this work is to learn deep uncertain convolutional features (UCF), which encourage the robustness and accuracy of saliency detection. We achieve this via introducing a reformulated dropout (R-dropout) after specific convolutional layers to construct an uncertain ensemble of internal feature units. In addition, we propose an effective hybrid upsampling method to reduce the checkerboard artifacts of deconvolution operators in our decoder network. The proposed methods can also be applied to other deep convolutional networks. Compared with existing saliency detection methods, the proposed UCF model is able to incorporate uncertainties for more accurate object boundary inference. Extensive experiments demonstrate that our proposed saliency model performs favorably against state-of-the-art approaches. The uncertain feature learning mechanism as well as the upsampling method can significantly improve performance on other pixel-wise vision tasks. Dong Wang 0004, Huchuan Lu, Hongyu Wang 0001 |
ICCV | 2 |
| 2017 | Person Re-Identification via Distance Metric Learning With Latent VariablesabstractIn this paper, we propose an effective person re-identification method with latent variables, which represents a pedestrian as the mixture of a holistic model and a number of flexible models. Three types of latent variables are introduced to model uncertain factors in the re-identification problem, including vertical misalignments, horizontal misalignments and leg posture variations. The distance between two pedestrians can be determined by minimizing a given distance function with respect to latent variables, and then be used to conduct the re-identification task. In addition, we develop a latent metric learning method for learning the effective metric matrix, which can be solved via an iterative manner: once latent information is specified, the metric matrix can be obtained based on some typical metric learning methods; with the computed metric matrix, the latent variables can be determined by searching the state space exhaustively. Finally, extensive experiments are conducted on seven databases to evaluate the proposed method. The experimental results demonstrate that our method achieves better performance than other competing algorithms. Dong Wang 0004, Huchuan Lu |
IEEE Trans. Image Process. | 2 |
| 2016 | Online Object Tracking Based on Convex Hull RepresentationabstractThis paper presents a novel tracking algorithm based on the convex hull representation model with sparse representation. The tracked object is assumed to be within the object convex hull and the candidate convex hull in the meanwhile. The object convex hull consists of a principle component analysis (PCA) subspace, and the candidate convex hull is constructed by all candidate samples with the sparsity constraint. Then we propose the objective function for our convex hull representation model, and design an iterative algorithm to solve it effectively. Finally, we present a tracking framework based on the proposed convex hull model and a simple online update scheme. Both qualitative and quantitative evaluations on some challenging video clips show that our tracker achieves better performance than other state-of-theart methods. Chunjuan Bo, Dong Wang 0004 |
ICPADS | 2 |
| 2016 | Human running detection: Benchmark and baseline
Shihong Lao, Dong Wang 0004, Fu Li 0003, Haihong Zhang |
Comput. Vis. Image Underst. | 2 |
| 2016 | Multi-feature tracking via adaptive weights
Huilan Jiang, Dong Wang 0004, Huchuan Lu |
Neurocomputing | 3 |
| 2016 | Hyperspectral Image Classification via JCR and SVM Models With Decision FusionabstractIn this letter, we propose a novel hyperspectral image (HSI) classification method based on the joint collaborative representation (JCR) and support vector machine (SVM) models with decision fusion. First, motivated by the joint model, we adopt a JCR model to deal with HSI classification and develop an effective method to learn contextual basis vectors for the JCR model. Second, the mid-features are first extracted based on representation coefficients obtained by the JCR method and then used to train a multiclass SVM classifier. After that, we exploit a multiplicative fusion rule to combine the JCR and SVM models. We conduct numerous experiments to evaluate our method in comparison with other algorithms. The experimental results on three standard data sets demonstrate that our method achieves better performance than other competing ones. Chunjuan Bo, Huchuan Lu, Dong Wang 0004 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Dual Group Structured TrackingabstractThe sparse representation (SR)-based tracking framework generally considers the testing candidates and dictionary atoms individually, thus failing to model the structured information within data. In this paper, we present a robust tracking framework by exploiting the dual group structure of both candidate samples and dictionary templates, and formulate the SR at group level. The similar samples are encoded simultaneously by a few atom groups, which induces the inter-group sparsity, and also each group enjoys different internal sparsity. In this way, not only the potential commonality shared by the related candidates is taken into account but also the individual differences between samples are reflected. Then, we provide two effective optimization methods to solve our formulation by block-coordinate gradient descent and alternating direction method of multipliers, respectively, and make a comparison between them in terms of both effectiveness and efficiency. Finally, we embed the dual group structure model into the particle filter framework for visual tracking. Extensive experimental results demonstrate that our tracker achieves favorable performance against the state-of-the-art tracking methods. Fu Li 0003, Huchuan Lu, Dong Wang 0004, Yi Wu 0001, Kaihua Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Robust Visual Tracking via Least Soft-Threshold SquaresabstractIn this paper, we propose an online tracking algorithm based on a novel robust linear regression estimator. In contrast to existing methods, the proposed least soft-threshold squares (LSS) algorithm models the error term with the Gaussian-Laplacian distribution, which can be efficiently solved. For visual tracking, the Gaussian-Laplacian noise assumption enables our LSS model to handle the normal appearance change and outlier simultaneously. Based on the maximum joint likelihood of parameters, we derive an LSS distance metric to measure the difference between an observation sample and a dictionary of positive templates. Compared with the distance derived from ordinary least squares methods, the proposed metric is more effective in dealing with the outliers. In addition, we provide insights on the relationships among the LSS problem, Huber loss function, and trivial templates, which facilitate better understandings of the existing tracking methods. Finally, we develop a robust tracking algorithm based on the LSS distance metric with an update scheme and negative templates, and speed it up with a particle selection mechanism. Experimental results on numerous challenging image sequences demonstrate that the proposed tracking algorithm performs favorably than the state-of-the-art methods. Dong Wang 0004, Huchuan Lu, Ming-Hsuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Occlusion-Aware Fragment-Based Tracking With Spatial-Temporal ConsistencyabstractIn this paper, we present a robust tracking method by exploiting a fragment-based appearance model with consideration of both temporal continuity and discontinuity information. From the perspective of probability theory, the proposed tracking algorithm can be viewed as a two-stage optimization problem. In the first stage, by adopting the estimated occlusion state as a prior, the optimal state of the tracked object can be obtained by solving an optimization problem, where the objective function is designed based on the classification score, occlusion prior, and temporal continuity information. In the second stage, we propose a discriminative occlusion model, which exploits both foreground and background information to detect the possible occlusion, and also models the consistency of occlusion labels among different frames. In addition, a simple yet effective training strategy is introduced during the model training (and updating) process, with which the effects of spatial-temporal consistency are properly weighted. The proposed tracker is evaluated by using the recent benchmark data set, on which the results demonstrate that our tracker performs favorably against other state-of-the-art tracking algorithms. Dong Wang 0004, Huchuan Lu |
IEEE Trans. Image Process. | 2 |
| 2015 | Multi-view Clustering via Structured Low-rank RepresentationabstractIn this paper, we present a novel solution to multi-view clustering through a structured low-rank representation. When assuming similar samples can be linearly reconstructed by each other, the resulting representational matrix reflects the cluster structure and should ideally be block diagonal. We first impose low-rank constraint on the representational matrix to encourage better grouping effect. Then representational matrices under different views are allowed to communicate with each other and share their mutual cluster structure information. We develop an effective algorithm inspired by iterative re-weighted least squares for solving our formulation. During the optimization process, the intermediate representational matrix from one view serves as a cluster structure constraint for that from another view. Such mutual structural constraint fine-tunes the cluster structures from both views and makes them more and more agreeable. Extensive empirical study manifests the superiority and efficacy of the proposed method. Dong Wang 0004, Qiyue Yin, Ran He 0001, Liang Wang 0001, Tieniu Tan |
CIKM | 1 |
| 2015 | Kernel collaborative face recognition
Dong Wang 0004, Huchuan Lu, Ming-Hsuan Yang 0001 |
Pattern Recognit. | 1 |
| 2015 | Visual Tracking via Structure Constrained GroupingabstractThis letter introduces a novel two-pass structural grouping algorithm and casts visual tracking as foreground superpixels grouping problem. In the first step, pairwise superpixel grouping is conducted in four orientations. Grouping prototypes containing the prior information of foreground and background are generated to determine whether any pair of neighboring superpixels should be grouped together. In the second step, superpixels selected by the first step are grouped into a single region which serves as the object region. The proposed grouping method has two benefits over the state-of-the-art ones. First, pairwise grouping is independently conducted in four orientations, which exploits the local structure of the foregound/backgroud and facilitates a more robust grouping process. Second, rather than considering the similarity of two neighboring superpixels, the grouping process is performed via accounting for the prior information of the object and the background, which is more suitable for visual tracking. Many experiments on challenging video clips demonstrate that our method achieves good performance than the state-of-the-art trackers in a wide range of tracking scenarios. Lijun Wang 0001, Huchuan Lu, Dong Wang 0004 |
IEEE Signal Process. Lett. | 3 |
| 2015 | Visual Tracking via Weighted Local Cosine SimilarityabstractIn this paper, we propose a novel weighted local cosine similarity (WLCS) and apply it to visual tracking. First, we present the local cosine similarity to measure the similarities between the target template and candidates, and provide some theoretical insights on it. Second, we develop an objective function to model the discriminative ability of local components, and use a quadratic programming method to solve the objective function and to obtain the discriminative weights. Finally, we design an effective and efficient tracker based on the WLCS method and a simple update manner within the particle filter framework. Experimental results on several challenging image sequences show that the proposed tracker achieves better performance than other competing methods. Dong Wang 0004, Huchuan Lu, Chunjuan Bo |
IEEE Trans. Cybern. | 1 |
| 2015 | Fast and Robust Object Tracking via Probability Continuous Outlier ModelabstractThis paper presents a novel visual tracking method based on linear representation. First, we present a novel probability continuous outlier model (PCOM) to depict the continuous outliers within the linear representation model. In the proposed model, the element of the noisy observation sample can be either represented by a principle component analysis subspace with small Guassian noise or treated as an arbitrary value with a uniform prior, in which a simple Markov random field model is adopted to exploit the spatial consistency information among outliers (or inliners). Then, we derive the objective function of the PCOM method from the perspective of probability theory. The objective function can be solved iteratively by using the outlier-free least squares and standard max-flow/min-cut steps. Finally, for visual tracking, we develop an effective observation likelihood function based on the proposed PCOM method and background information, and design a simple update scheme. Both qualitative and quantitative evaluations demonstrate that our tracker achieves considerable performance in terms of both accuracy and speed. Dong Wang 0004, Huchuan Lu, Chunjuan Bo |
IEEE Trans. Image Process. | 1 |
| 2015 | Inverse Sparse Tracker With a Locally Weighted Distance MetricabstractSparse representation has been recently extensively studied for visual tracking and generally facilitates more accurate tracking results than classic methods. In this paper, we propose a sparsity-based tracking algorithm that is featured with two components: 1) an inverse sparse representation formulation and 2) a locally weighted distance metric. In the inverse sparse representation formulation, the target template is reconstructed with particles, which enables the tracker to compute the weights of all particles by solving only one l1 optimization problem and thereby provides a quite efficient model. This is in direct contrast to most previous sparse trackers that entail solving one optimization problem for each particle. However, we notice that this formulation with normal Euclidean distance metric is sensitive to partial noise like occlusion and illumination changes. To this end, we design a locally weighted distance metric to replace the Euclidean one. Similar ideas of using local features appear in other works, but only being supported by popular assumptions like local models could handle partial noise better than holistic models, without any solid theoretical analysis. In this paper, we attempt to explicitly explain it from a mathematical view. On that basis, we further propose a method to assign local weights by exploiting the temporal and spatial continuity. In the proposed method, appearance changes caused by partial occlusion and shape deformation are carefully considered, thereby facilitating accurate similarity measurement and model update. The experimental validation is conducted from two aspects: 1) self validation on key components and 2) comparison with other state-of-the-art algorithms. Results over 15 challenging sequences show that the proposed tracking algorithm performs favorably against the existing sparsity-based trackers and the other state-of-the-art methods. Dong Wang 0004, Huchuan Lu, Ziyang Xiao, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Robust Visual Tracking with Dual Group Structure
Fu Li 0003, Huchuan Lu, Dong Wang 0004 |
ACCV (4) | 3 |
| 2014 | Visual Tracking via Probability Continuous Outlier ModelabstractIn this paper, we present a novel online visual tracking method based on linear representation. First, we present a novel probability continuous outlier model (PCOM) to depict the continuous outliers that occur in the linear representation model. In the proposed model, the element of the noisy observation sample can be either represented by a PCA subspace with small Guassian noise or treated as an arbitrary value with a uniform prior, in which the spatial consistency prior is exploited by using a binary Markov random field model. Then, we derive the objective function of the PCOM method, the solution of which can be iteratively obtained by the outlier-free least squares and standard max-flow/min-cut steps. Finally, based on the proposed PCOM method, we design an effective observation likelihood function and a simple update scheme for visual tracking. Both qualitative and quantitative evaluations demonstrate that our tracker achieves very favorable performance in terms of both accuracy and speed. Dong Wang 0004, Huchuan Lu |
CVPR | 1 |
| 2014 | Semi-supervised subspace segmentationabstractSubspace segmentation methods usually rely on the raw explicit feature vectors in an unsupervised manner. In many applications, it is cheap to obtain some pairwise link information that tells whether two data points are in the same subspace or not. Though partially available, such link information serves as some kind of high-level semantics, which can be further used as a constraint to improve the segmentation accuracy. By constructing a link matrix and using it as a regularizer, we propose a semi-supervised subspace segmentation model where the partially observed subspace membership prior can be encoded. Specificly, under the common linear representation assumption, we enforce the representational coefficient to be consistent with the link matrix. Thus the low-level and high-level information about the data can be integrated to produce more precise segmentation results. We then develop an effective algorithm to optimize our model in an alternating minimization way. Experimental results for both motion segmentation and face clustering validate that incorporating such link information is helpful to assist and bias the unsupervised subspace segmentation methods. Dong Wang 0004, Qiyue Yin, Ran He 0001, Liang Wang 0001, Tieniu Tan |
ICIP | 1 |
| 2014 | Online Visual Tracking via Two View Sparse RepresentationabstractIn this letter, we present a novel online tracking method based on sparse representation. In contrast to existing “sparse representation”-based tracking algorithms, this work adopts the sparse representation method to construct both object and state models. The tracked object can be sparsely represented by a series of object templates, and also can be sparsely represented by candidate samples in the current frame. Furthermore, we propose a unified objective function to integrate object and state models, and cast the tracking problem as an optimization problem that can be solved in an iteration manner. Finally, we compare the proposed tracker with nine state-of-the-art tracking methods by using some challenging image sequences. Both qualitative and quantitative evaluations demonstrate that our tracker achieves favorable performance in terms of both accuracy and speed. Dong Wang 0004, Huchuan Lu, Chunjuan Bo |
IEEE Signal Process. Lett. | 1 |
| 2014 | L2-RLS-Based Object TrackingabstractIn this paper, we present a robust and fast tracking algorithm in which object tracking is achieved by solving ℓ2-regularized least square (ℓ2-RLS) problems in a Bayesian inference framework. First, the changing appearance of the tracked target is modeled with PCA basis vectors and square templates, which makes the tracker not only exploit the strength of subspace representation but also explicitly take partial occlusion into consideration. They can together represent both the intact and corrupted objects well. Second, we adopt the ℓ2-regularized least square method to solve the proposed representation model. Compared with the complex ℓ1-based algorithm, it provides a very fast performance without the loss of accuracy in handling the tracking problem. In addition, a novel likelihood function and a refined update scheme further help to improve the robustness of our tracker. Both qualitative and quantitative evaluations on several challenging image sequences demonstrate that the proposed method performs favorably against several state-of-the-art tracking algorithms. Ziyang Xiao, Huchuan Lu, Dong Wang 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Visual Tracking via Discriminative Sparse Similarity MapabstractIn this paper, we cast the tracking problem as finding the candidate that scores highest in the evaluation model based upon a matrix called discriminative sparse similarity map (DSS map). This map demonstrates the relationship between all the candidates and the templates, and it is constructed based on the solution to an innovative optimization formulation named multitask reverse sparse representation formulation, which searches multiple subsets from the whole candidate set to simultaneously reconstruct multiple templates with minimum error. A customized APG method is derived for getting the optimum solution (in matrix form) within several iterations. This formulation allows the candidates to be evaluated accurately in parallel rather than one-by-one like most sparsity-based trackers do and meanwhile considers the relationship between candidates, therefore it is more superior in terms of cost-performance ratio. The discriminative information containing in this map comes from a large template set with multiple positive target templates and hundreds of negative templates. A Laplacian term is introduced to keep the coefficients similarity level in accordance with the candidates similarities, thereby making our tracker more robust. A pooling approach is proposed to extract the discriminative information in the DSS map for easily yet effectively selecting good candidates from bad ones and finally get the optimum tracking results. Plenty experimental evaluations on challenging image sequences demonstrate that the proposed tracking algorithm performs favorably against the state-of-the-art methods. Bohan Zhuang, Huchuan Lu, Ziyang Xiao, Dong Wang 0004 |
IEEE Trans. Image Process. | 4 |
| 2013 | Least Soft-Threshold Squares TrackingabstractIn this paper, we propose a generative tracking method based on a novel robust linear regression algorithm. In contrast to existing methods, the proposed Least Soft-thresold Squares (LSS) algorithm models the error term with the Gaussian-Laplacian distribution, which can be solved efficiently. Based on maximum joint likelihood of parameters, we derive a LSS distance to measure the difference between an observation sample and the dictionary. Compared with the distance derived from ordinary least squares methods, the proposed metric is more effective in dealing with outliers. In addition, we present an update scheme to capture the appearance change of the tracked target and ensure that the model is properly updated. Experimental results on several challenging image sequences demonstrate that the proposed tracker achieves more favorable performance than the state-of-the-art methods. Dong Wang 0004, Huchuan Lu, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2013 | Fast and effective color-based object tracking by boosted color distribution
Dong Wang 0004, Huchuan Lu, Ziyang Xiao, Yen-Wei Chen 0001 |
Pattern Anal. Appl. | 1 |
| 2013 | On-line learning parts-based representation via incremental orthogonal projective non-negative matrix factorization
Dong Wang 0004, Huchuan Lu |
Signal Process. | 1 |
| 2013 | Online Object Tracking With Sparse PrototypesabstractOnline object tracking is a challenging problem as it entails learning an effective model to account for appearance change caused by intrinsic and extrinsic factors. In this paper, we propose a novel online object tracking algorithm with sparse prototypes, which exploits both classic principal component analysis (PCA) algorithms with recent sparse representation schemes for learning effective appearance models. We introduce l(1) regularization into the PCA reconstruction, and develop a novel algorithm to represent an object by sparse prototypes that account explicitly for data and noise. For tracking, objects are represented by the sparse prototypes learned online with update. In order to reduce tracking drift, we present a method that takes occlusion and motion blur into account rather than simply includes image observations for model update. Both qualitative and quantitative evaluations on challenging image sequences demonstrate that the proposed tracking algorithm performs favorably against several state-of-the-art methods. Dong Wang 0004, Huchuan Lu, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | Baseline Results for Violence Detection in Still ImagesabstractRecognizing objectionable content draws more and more attention nowadays given the rapid proliferation of images and videos on the Internet. Although there are some investigations about violence video detection and pornographic information filtering, very few existing methods touch on the problem of violence detection in still images. However, given its potential use in violence webpage filtering, online public opinion monitoring and some other aspects, recognizing violence in still images is worth being deeply investigated. To this end, we first establish a new database containing 500 violence images and 1500 non-violence images. And we use the Bag-of-Words (BoW) model which is frequently adopted in image classification domain to discriminate violence images and non-violence images. The effectiveness of four different feature representations are tested within the BoW framework. Finally the baseline results for violence image detection on our newly built database are reported. Dong Wang 0004, Zhang Zhang 0001, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
AVSS | 1 |
| 2012 | Fragment-based tracking using online multiple kernel learningabstractFragment-based tracking methods have shown its robustness in handling partial occlusion and pose change. In this paper, we propose a novel fragment-based tracking approach using on online multiple kernel learning (MKL) method. An online MKL method for object tracking is implemented by considering temporal continuity explicitly. Instead of directly using multiple features of objects, we employ MKL to make full use of multiple fragments of the object. This can automatically assign different weights to the fragments according to their discriminative power. In addition, for better robustness two kinds of independent features are computed to enrich the representation of patches. We build a classifier for each type of feature and assign them different weights according to their performance on classification. Both qualitative and quantitative evaluations on challenging image sequences demonstrate that the proposed tracking approach performs favorably against several state-of-the-art methods. Xu Jia 0012, Dong Wang 0004, Huchuan Lu |
ICIP | 2 |
| 2012 | Object tracking with L2-RLS
Ziyang Xiao, Huchuan Lu, Dong Wang 0004 |
ICPR | 3 |
| 2012 | Object Tracking via 2DPCA and ℓ1-RegularizationabstractIn this letter, we present a novel online object tracking algorithm by using 2DPCA and ℓ1-regularization. Firstly, we introduce ℓ1-regularization into the 2DPCA reconstruction, and develop an iterative algorithm to represent an object by 2DPCA bases and a sparse error matrix. Secondly, we propose a novel likelihood function that considers both the reconstruction error and the sparsity of the error matrix. This likelihood function not only handles partial occlusion effectively but also encourages the tracked object to be well-aligned. Finally, to further reduce tracking drift, we enhance the tracker updates by considering the sparsity of the error matrix. Based on our observations, a dense error matrix usually relates to partial occlusion or mis-alignment. Both qualitative and quantitative evaluations on challenging image sequences demonstrate that the proposed tracking algorithm achieves more favorable performance than several state-of-the-art methods. Dong Wang 0004, Huchuan Lu |
IEEE Signal Process. Lett. | 1 |
| 2012 | Pixel-Wise Spatial Pyramid-Based Hybrid TrackingabstractIn this paper, we propose a novel tracking algorithm that combines complementary tracking modules with a new object representation model to balance between stability and adaptivity. To reduce the update error of online tracking, we present three complementary modules (a stable module, a soft stable module, and an adaptive module) and fuse them by using a biased multiplicative criterion. The combination of those modules not only facilitates the accurate location of the tracked object but also makes our tracker adaptive to appearance change. For objection representation, we present an appearance model named pixel-wise spatial pyramid (PSP), which employs pixel feature vector to combine several pixel characteristics. During the updating process, we update the codebook by using the reserved pixel feature vectors that are selected by a distance-based scheme. Then, we generate an evolving target representation by using a hybrid feature map that consists of the reserved pixel vectors and antipart of the previous hybrid feature map. Numerous experiments on various challenging image sequences demonstrate that the proposed algorithm performs favorably against several state-of-the-art algorithms, especially for drastic appearance change. Huchuan Lu, Shipeng Lu, Dong Wang 0004, Henry Leung 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | A co-training framework for visual tracking with multiple instance learningabstractThis paper proposes a Co-training Multiple Instance Learning algorithm (CoMIL). Our framework is based on the co-training approach which labels incoming data continuously, and then uses the prediction from each classifier to enlarge the training set of the other. The discriminative classifier is implemented using online multiple instance learning (MIL), which can deal with inaccurate positive samples in the updating process and allow some flexibility while finding a decision boundary. Firstly, two classifiers are improved mutually in our CoMIL tracking system. Secondly, our update mechanism uses multiple potential positives according to the MIL which handles the update error due to the risk of extracting only one positive example. Experiments show that our CoMIL tracking algorithm performs better than several state-of-the-art tracking algorithms on challenging sequences. Huchuan Lu, Qiuhong Zhou, Dong Wang 0004, Xiang Ruan |
FG | 3 |
| 2011 | Incremental orthogonal projective non-negative matrix factorization and its applicationsabstractIn this paper, we propose an incremental orthogonal projective non-negative matrix factorization algorithm (IOPNMF), which aims to learn a parts-based subspace that reveals dynamic data streams. There exist two main contributions. Firstly, our proposed algorithm can learn parts-based representations in an online fashion. Secondly, by using projection and orthogonality constrains, our IOPNMF algorithm can guarantee to learn a linear parts-based subspace. To demonstrate the effectiveness of our method, we conduct two kinds of experiments, incremental learning parts-based components on facial database and visual tracking on several challenging video clips. The experimental results show that our IOPNMF algorithm learns parts-based representations successfully. Dong Wang 0004, Huchuan Lu |
ICIP | 1 |
| 2011 | Two dimensional principal components of natural images and its application
Dong Wang 0004, Huchuan Lu, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2010 | Object tracking by multi-cues spatial pyramid matchingabstractIn this paper, we propose a novel tracking framework, multi-cues spatial pyramid matching (MSPM). Different cues are used to generate a set of probability maps, where the value of each pixel indicates the probability that it belongs to the foreground. Then those probability maps are combined into a single probability map by a weighted linear function. There exist two main contributions. First, a generic probability maps fusion mechanism is proposed. The weights of different probability maps are updated dynamically to maintain local discriminative power, which is achieved by solving a regression problem efficiently. Second, spatial pyramid matching kernel is adopted as a likelihood function, which considers spatial information of object and is able to cope with occlusions naturally. Experiments performed on several challenging public video sequences demonstrate that our proposed framework achieves considerable performance, compared to algorithms with individual cues or equal weights combination, and other state-of-the-art ones. Dong Wang 0004, Huchuan Lu, Yen-Wei Chen 0001 |
ICIP | 1 |
| 2010 | Incremental MPCA for Color Object TrackingabstractThe task of visual tracking is to deal with dynamic image streams that change over time. For color object tracking, although a color object is a 3-order tensor in essence, little attention has been focused on this attribute. In this paper, we propose a novel Incremental Multiple Principal Component Analysis (IMPCA) method for online learning dynamic tensor streams. When newly added tensor set arrives, the mean tenor and the covariance matrices of different modes can be updated easily, and then projection matrices can be effectively calculated based on covariance matrices. Finally, we apply our IMPCA method to color object tracking using Bayes inference framework. Experiments are performed on some changeling public and our own video sequences. The experimental results demonstrate that the proposed method achieves considerable performance. Dong Wang 0004, Huchuan Lu, Yen-Wei Chen 0001 |
ICPR | 1 |