Xin Li 0034

dblp:09/1365-34 · DBLP profile ↗
← Back
65ranked-venue papers
11as first author
49since 2021 · last 2026
0000-0002-1670-1368ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 5 first-author · 30 since 2021Artificial intelligence and machine learning · 32 · 9 first-author · 23 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Semantic-aware multi-view person image generation for re-identification
Si Wu 0002, Xin Li 0034, Yong Xu 0007, Yaowei Wang 0001
Image Vis. Comput.3
2026 AlignMamba-2: Enhancing multimodal fusion and sentiment analysis with modality-aware Mamba
Yan Li 0121, Yifei Xing 0001, Xiangyuan Lan, Xin Li 0034, Dongmei Jiang
Pattern Recognit.4
2026 Adversarial flow-based generative models for visible-to-Infrared person re-Identification
Honghu Pan, Yongyong Chen, Xin Li 0034, Zhenyu He 0001
Pattern Recognit.3
2026 Seg-LLaVA: Empowering pixel-level understanding with large vision language model
Fan Yang 0089, Yousong Zhu, Yufei Zhan, Hongyin Zhao, Xin Li 0034, Yaowei Wang 0001, Ming Tang 0001, Jinqiao Wang
Pattern Recognit.5
2026 PPIFuse: Physical Priors Injected Infrared and Visible Image Fusion
abstract
Existing infrared and visible image fusion methods commonly use two structurally identical networks to extract deep features from source images, followed by a handcrafted or learnable feature fusion strategy. These methods overlook the modality-specific characteristics of the two image types, impairing the model’s ability to fully exploit their complementary information. Additionally, their fusion results often exhibit issues such as texture detail loss or unclear thermal targets. This is because the fusion rules they used are either too simple or too redundant. To address these challenges, we start from the infrared physics priors that are naturally complementary to visible images and incorporate the thermal diffusion equation and Stefan-Boltzmann Law into the image fusion architecture. Based on these two physical priors, we design a Thermal Diffusion Convolution (TDC) and a Stefan Thermal Attention (STA) to better extract infrared-specific features. Specifically, the TDC module leverages the anisotropic and isotropic characteristics of thermal diffusion adaptively to sharpen the edges of thermal targets and remove infrared noise, minimizing artifacts in the fused results. By decomposing the Stefan-Boltzmann Law, STA pays more attention on thermal features while suppressing redundant information, enabling more effective aggregation of complementary modality-specific details. To make full use of layer-wise complementary features, we propose an Interactive Injection Fusion framework(IIF) that hierarchically integrates these features, enhancing the richness of fused image content. Furthermore, an energy conservation constraint is designed to ensure the fused images adhere to physical principles. Extensive experimental results on five datasets demonstrate that our method sets a new state-of-the-art. Code is available at https://github.com/QiaoLiuHit/PPIFuse.
Qianhong Zhang, Qiao Liu 0001, Di Yuan 0002, Xin Li 0034, Yunpeng Liu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Harnessing Vision-Language Pretrained Models With Temporal-Aware Adaptation for Referring Video Object Segmentation
abstract
Referring Video Object Segmentation (RVOS) is a task that involves segmenting target objects in a video based on the given referring expressions. It is critical for video editing and analysis. The crux of RVOS is to model dense text-video relations to associate abstract linguistic concepts with dynamic visual contents at pixel-level. Most RVOS methods typically use vision and language models pretrained independently as backbones, mapping images and texts to uncoupled feature spaces. As a result, they must learn Vision-Language (VL) relation modeling from scratch. Vision-Language Pretrained (VLP) models have achieved remarkable success. Inspired by this, we propose to explore relation modeling for RVOS based on their aligned VL feature space. Nevertheless, transferring VLP models to RVOS is deceptively challenging, due to the gap between static image/region-level pretraining and dynamic pixel-level prediction. To bridge this gap, we introduce a framework named VLP-RVOS, which harnesses VLP models for RVOS through temporal-aware adaptation. We first propose temporal-aware prompt-tuning to adapt pretrained representations for pixel-level prediction and empower the vision encoder to model temporal contexts. We further customize a cube-frame attention mechanism for robust spatial-temporal reasoning. Besides, we propose to perform multi-stage VL relation modeling while and after feature extraction for comprehensive understanding. Extensive experiments demonstrate that VLP-RVOS performs favorably against state-of-the-art algorithms and generalizes well. Our codes are available at https://github.com/xwt909090/VLP-RVOS.
Zikun Zhou, Wentao Xiong, Li Zhou 0017, Xin Li 0034, Zhenyu He 0001, Yaowei Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Unsupervised Domain Adaptive Thermal Infrared Tracking
abstract
Existing deep Thermal InfraRed (TIR) trackers often use RGB datasets for training due to the lack of large-scale labeled TIR datasets. However, the performance of these methods on TIR image sequences is significantly degraded, because of the domain shift problem between the RGB and TIR datasets. To solve this problem, in this paper, we propose an unsupervised Dual-level Domain Adaptation TIR Tracking framework (DDAT), which can benefit from training on large-scale labeled RGB datasets and unlabeled TIR datasets. Specifically, to transfer the useful knowledge learned from RGB dataset to TIR tracking, we first propose an adversarial-based adaptation module on both the semantic-level and the feature-level. While the semantic-level adaptation can reduce the semantic gap between the TIR and RGB tracking tasks, the feature-level adaptation can learn domain-invariant features for more robust tracking. Second, we propose a partial domain adaptation module to alleviate the negative transfer problem because the RGB and TIR tracking domains have a non-identical class and feature spaces. Instead of aligning the entire feature space, this module adaptively selects partial similarity samples and features for alignment, thus getting more fine-grained aligned results. Third, we collect a currently largest-scale unlabeled TIR dataset to train the proposed framework. Extensive experiments on five TIR tracking benchmarks demonstrate the proposed method is effective and sets a new state-of-the-art.
Qiao Liu 0001, Xin Li 0034, Jiatian Pi, Di Yuan 0002, Yunpeng Liu 0001
IEEE Trans. Multim.3
2026 Prototype Perturbation for Relaxing Alignment Constraints in Backward-Compatible Learning
abstract
The traditional paradigm to update retrieval models requires re-computing the embeddings of the gallery data, a time-consuming and computationally intensive process known as backfilling. To circumvent backfilling, Backward-Compatible Learning (BCL) has been widely explored, which aims to train a new model compatible with the old one. Many previous works focus on effectively aligning the embeddings of the new model with those of the old one to enhance backward compatibility. Nevertheless, such strong alignment constraints would compromise the discriminative ability of the new model, particularly when different classes are closely clustered and hard to distinguish in the old feature space. To address this issue, we propose to relax the constraints by introducing perturbations to the old feature prototypes. This allows us to align the new feature space with a pseudo-old feature space defined by these perturbed prototypes, thereby preserving the discriminative ability of the new model in backward-compatible learning. We have developed two approaches for calculating the perturbations: Neighbor-Driven Prototype Perturbation (NDPP) and Optimization-Driven Prototype Perturbation (ODPP). Particularly, they take into account the feature distributions of not only the old but also the new models to obtain proper perturbations along with new model updating. Extensive experiments on the landmark and commodity datasets demonstrate that our approaches perform favorably against state-of-the-art BCL algorithms.
Zikun Zhou, Yushuai Sun, Wenjie Pei, Xin Li 0034, Yaowei Wang 0001
IEEE Trans. Multim.4
2025 AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
abstract
Cross-Modal alignment is crucial for multimodal representation fusion due to the inherent heterogeneity between modalities. While Transformer-Based methods have shown promising results in modeling inter-modal relationships, their quadratic computational complexity limits their applicability to long-sequence or large-scale data. Although recent Mamba-Based approaches achieve linear complexity, their sequential scanning mechanism poses fundamental challenges in comprehensively modeling cross-modal relationships. To address this limitation, we propose Align-Mamba, an efficient and effective method for multimodal fusion. Specifically, grounded in Optimal Transport, we introduce a local cross-modal alignment module that explicitly learns token-level correspondences between different modalities. Moreover, we propose a global cross-modal alignment loss based on Maximum Mean Discrepancy to implicitly enforce the consistency between different modal distributions. Finally, the unimodal representations after local and global alignment are passed to the Mamba backbone for further cross-modal interaction and multimodal fusion. Extensive experiments on complete and incomplete multimodal fusion tasks demonstrate the effectiveness and efficiency of the proposed method. For instance, on the CMU-MOSI dataset, AlignMamba improves classification accuracy by 0.9%, reduces GPU memory usage by 20.3%, and decreases inference time by 83.3%.
Yan Li 0121, Yifei Xing 0001, Xiangyuan Lan, Xin Li 0034, Dongmei Jiang
CVPR4
2025 Efficient Hierarchical Domain Adaptive Thermal Infrared Tracking
abstract
Constrained by the scarcity of labeled Thermal InfraRed (TIR) training data, current TIR trackers commonly rely on pre-trained RGB trackers. However, the domain discrepancy between TIR and RGB images limits effective utilization of RGB features, significantly degrades TIR tracking performance. To solve this challenge, we propose a hierarchical domain adaptation model to transfer useful pre-trained RGB features into TIR tracking more effective and efficient. Specifically, we first design a reflectance consistency network to learn style-invariant representations. Second, we present a target-aware adversarial network to align the target semantic features of the two domains. These two modules respectively narrow the distribution gap at the stylistic and semantic levels in a hierarchical manner. Third, to solve the inefficiency problem of domain adaptive training, we also propose a Bi-rank adapter side network to accelerate this process. While significantly reducing training time by 90%, our method achieves a new state-of-the-art on four TIR tracking benchmarks.
Kanlun Tan, Qiao Liu 0001, Di Yuan 0002, Xin Li 0034, Yunpeng Liu 0001
ICASSP5
2025 Synergistic Spotting and Recognition of Micro-Expression via Temporal State Transition
abstract
Micro-expressions are involuntary facial movements that cannot be consciously controlled, conveying subtle cues with substantial real-world applications. The analysis of micro-expressions generally involves two main tasks: spotting micro-expression intervals in long videos and recognizing the emotions associated with these intervals. Previous deep-learning methods have primarily relied on classification networks utilizing sliding windows. However, fixed window sizes and window-level hard classification introduce numerous constraints. Additionally, these methods have not fully exploited the potential of complementary pathways for spotting and recognition. In this paper, we present a novel temporal state transition architecture grounded in the state space model, which replaces conventional window-level classification with video-level regression. Furthermore, by leveraging the inherent connections between spotting and recognition tasks, we propose a synergistic strategy that enhances overall analysis performance. Extensive experiments demonstrate that our method achieves state-of-the-art performance. The codes are available at https://github.com/zizheng-guo/ME-TST.
Bochao Zou, Zizheng Guo 0002, Wenfeng Qin, Xin Li 0034, Kangsheng Wang, Huimin Ma 0001
ICASSP4
2025 Controllable 3D Outdoor Scene Generation via Scene Graphs
Lu Qi 0001, Xin Li 0034, Wenping Wang 0001, Chongshou Li, Ming-Hsuan Yang 0001
ICCV5
2025 Learning Spatial-Semantic Features for Robust Video Object Segmentation
abstract
Tracking and segmenting multiple similar objects with distinct or complex parts in long-term videos is particularly challenging due to the ambiguity in identifying target components and the confusion caused by occlusion, background clutter, and changes in appearance or environment over time. In this paper, we propose a robust video object segmentation framework that learns spatial-semantic features and discriminative object queries to address the above issues. Specifically, we construct a spatial-semantic block comprising a semantic embedding component and a spatial dependency modeling part for associating global semantic features and local spatial features, providing a comprehensive target representation. In addition, we develop a masked cross-attention module to generate object queries that focus on the most discriminative parts of target objects during query propagation, alleviating noise accumulation to ensure effective long-term query propagation. The experimental results show that the proposed method sets new state-of-the-art performance on multiple data sets, including the DAVIS2017 test (\textbf{87.8\%}), YoutubeVOS 2019 (\textbf{88.1\%}), MOSE val (\textbf{74.0\%}), and LVOS test (\textbf{73.0\%}), which demonstrate the effectiveness and generalization capacity of the proposed method. We will make all the source code and trained models publicly available.
Xin Li 0034, Deshui Miao, Zhenyu He 0001, Yaowei Wang 0001, Huchuan Lu, Ming-Hsuan Yang 0001
ICLR1
2025 OSDA Agent: Leveraging Large Language Models for De Novo Design of Organic Structure Directing Agents
abstract
Zeolites are crystalline porous materials that have been widely utilized in petrochemical industries as well as sustainable chemistry areas. Synthesis of zeolites often requires small molecules termed Organic Structure Directing Agents (OSDAs), which are critical in forming the porous structure. Molecule generation models can aid the design of OSDAs, but they are limited by single functionality and lack of interactivity. Meanwhile, large language models (LLMs) such as GPT-4, as general-purpose artificial intelligence systems, excel in instruction comprehension, logical reasoning, and interactive communication. However, LLMs lack in-depth chemistry knowledge and first-principle computation capabilities, resulting in uncontrollable outcomes even after fine-tuning. In this paper, we propose OSDA Agent, an interactive OSDA design framework that leverages LLMs as the brain, coupled with computational chemistry tools. The OSDA Agent consists of three main components: the Actor, responsible for generating potential OSDA structures; the Evaluator, which assesses and scores the generated OSDAs using computational chemistry tools; and the Self-reflector, which produces reflective summaries based on the Evaluator's feedback to refine the Actor's subsequent outputs. Experiments on representative zeolite frameworks show the generation-evaluation-reflection-refinement workflow can perform de novo design of OSDAs with superior generation quality than the pure LLM model, generating candidates consistent with experimentally validated OSDAs and optimizing known OSDAs.
Zhaolin Hu, Yixiao Zhou 0001, Zhongan Wang, Xin Li 0034, Weimin Yang, Hehe Fan, Yi Yang 0001
ICLR4
2025 Puzzle-MAE: A Puzzle-Inspired Mask Autoencoder for Multi-Modal Fusion
abstract
Most unsupervised methods in the video domain rely on simple encoder-decoder structures, often resulting in discrepancies between the features extracted from unmasked patches and those from the original patches. To address this issue, we propose a novel self-supervised learning framework, PuzzleMAE, which extracts features from both masked and unmasked patches and aligns them with original image representations to improve feature consistency. Inspired by the human ability to solve puzzles through holistic image recognition and the exploitation of spatial adjacency, we propose the Global-Local Attention Module, which effectively integrates global contextual information with local feature representations. Furthermore, we introduce 3D Relative Position Embedding and Structural Position Embedding to emulate human-like spatial and structural awareness of positional relationships during the puzzle-solving process. The effectiveness of our method is validated on two downstream tasks: the First Impression V2 and DFEW datasets.
Xin Li 0034, Bochao Zou, Rongquan Wang, Huimin Ma 0001
ICME1
2025 ZeroBP: Learning Position-Aware Correspondence for Zero-Shot 6D Pose Estimation in Bin-Picking
abstract
Bin-picking is a practical and challenging robotic manipulation task, where accurate 6D pose estimation plays a pivotal role. The workpieces in bin-picking are typically texture-less and randomly stacked in a bin, which poses a significant challenge to 6D pose estimation. Existing solutions are typically learning-based methods, which require object-specific training. Their efficiency of practical deployment for novel workpieces is highly limited by data collection and model retraining. Zero-shot 6D pose estimation is a potential approach to address the issue of deployment efficiency. Nevertheless, existing zero-shot 6D pose estimation methods are designed to leverage feature matching to establish point-to-point correspondences for pose estimation, which is less effective for workpieces with textureless appearances and ambiguous local regions. In this paper, we propose ZeroBP, a zero-shot pose estimation frame-work designed specifically for the bin-picking task. ZeroBP learns Position-Aware Correspondence (PAC) between the scene instance and its CAD model, leveraging both local features and global positions to resolve the mismatch issue caused by ambiguous regions with similar shapes and appearances. Extensive experiments on the ROBI dataset demonstrate that ZeroBP outperforms state-of-the-art zero-shot pose estimation methods, achieving an improvement of 9.1 % in average recall of correct poses.
Jianqiu Chen, Zikun Zhou, Xin Li 0034, Tianpeng Bao, Zhenyu He 0001
ICRA3
2025 Prior-Free Augmentation for Cloth-Changing Person Re-Identification
abstract
Cloth-changing Person Re-Identification (CCReID) aims to recognize individuals across clothing variations by learning clothing-invariant representations. However, obtaining sufficient samples of the same person in diverse outfits is often impractical. While synthesizing realistic person images provides an effective solution, existing augmentation methods require labeled data and external priors (e.g., pose skeletons, semantic maps), resulting in high costs and limited generalization. To this end, we propose a Prior-Free Augmentation method for Cloth-changing person re-identification (PFAC), which leverages text guidance to synthesize images with clothing variations while maintaining identity consistency. Our approach features: (1) a truncated diffusion model that preserves clothing-invariant structural cues from intermediate noisy images, (2) a dual-branch denoising network that decouples text-guided clothing synthesis from identity consistency via cross-modal alignment, and (3) a joint optimization strategy with identity-focused losses and image filtering to enhance realism and discriminability. Experimental results on PRCC, LTCC, and Celeb-reID datasets demonstrate that PFAC achieves state-of-the-art CCReID performance, effectively generating high-fidelity, identity-consistent images for robust augmentation without external priors.
Xin Li 0034, Si Wu 0002, Yong Xu 0007, Yaowei Wang 0001
ACM Multimedia2
2025 FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
abstract
Recent Large Vision Language Models (LVLMs) demonstrate promising capabilities in unifying visual understanding and generative modeling, enabling both accurate content understanding and flexible editing. However, current approaches treat \textbf{\textit{"what to see"}} and \textbf{\textit{"how to edit"}} separately: they either perform isolated object segmentation or utilize segmentation masks merely as conditional prompts for local edit generation tasks, often relying on multiple disjointed models. To bridge these gaps, we introduce FOCUS, a unified LVLM that integrates segmentation-aware perception and controllable object-centric generation within an end-to-end framework. FOCUS employs a dual-branch visual encoder to simultaneously capture global semantic context and fine-grained spatial details. In addition, we leverage a MoVQGAN-based visual tokenizer to produce discrete visual tokens that enhance generation quality. To enable accurate and controllable image editing, we propose a progressive multi-stage training pipeline, where segmentation masks are jointly optimized and used as spatial condition prompts to guide the diffusion decoder. This strategy aligns visual encoding, segmentation, and generation modules, effectively bridging segmentation-aware perception with fine-grained visual synthesis. Extensive experiments across three core tasks, including multimodal understanding, referring segmentation accuracy, and controllable image generation, demonstrate that FOCUS achieves strong performance by jointly optimizing visual perception and generative capabilities.
Fan Yang 0089, Yousong Zhu, Xin Li 0034, Yufei Zhan, Hongyin Zhao, Shurong Zheng, Yaowei Wang 0001, Ming Tang 0001, Jinqiao Wang
NeurIPS3
2025 D2Fusion: Dual-domain feature decoupling for infrared and visible image fusion
Yan Fan 0006, Wei Ran, Kanlun Tan, Qiao Liu 0001, Di Yuan 0002, Xin Li 0034, Yunpeng Liu 0001
Knowl. Based Syst.6
2025 Toward Long Video Understanding via Fine-Detailed Video Story Generation
abstract
Long video understanding has become a critical task in computer vision, driving advancements across numerous applications from surveillance to content retrieval. Existing video understanding methods suffer from two challenges when dealing with long video understanding: intricate long-context relationship modeling and interference from redundancy. To tackle these challenges, we introduce Fine-Detailed Video Story generation (FDVS), which interprets long videos into detailed textual representations. Specifically, to achieve fine-grained modeling of long-temporal content, we propose a Bottom-up Video Interpretation Mechanism that progressively interprets video content from clips to video. To avoid interference from redundant information in videos, we introduce a Semantic Redundancy Reduction mechanism that removes redundancy at both the visual and textual levels. Our method transforms long videos into hierarchical textual representations that contain multi-granularity information of the video. With these representations, FDVS is applicable to various tasks without any fine-tuning. We evaluate the proposed method across eight datasets spanning three tasks. The performance demonstrates the effectiveness and versatility of our method.
Zeng You, Zhiquan Wen, Yaofo Chen, Xin Li 0034, Runhao Zeng, Yaowei Wang 0001, Mingkui Tan
IEEE Trans. Circuits Syst. Video Technol.4
2025 Spatial-Temporal Saliency Guided Unbiased Contrastive Learning for Video Scene Graph Generation
abstract
Accurately detecting objects and their interrelationships for Video Scene Graph Generation (VidSGG) confronts two primary challenges. The first involves the identification of active objects interacting with humans from the numerous background objects, while the second challenge is long-tailed distribution among predicate classes. To tackle these challenges, we propose STABILE, a novel framework with a spatial-temporal saliency-guided contrastive learning scheme. For the first challenge, STABILE features an active object retriever that includes an object saliency fusion block for enhancing object embeddings with motion cues alongside an object temporal encoder to capture temporal dependencies. For the second challenge, STABILE introduces an unbiased relationship representation learning module with an Unbiased Multi-Label (UML) contrastive loss to mitigate the effect of long-tailed distribution. With the enhancements in both aspects, STABILE substantially boosts the accuracy of scene graph generation. Extensive experiments demonstrate the superiority of STABILE, setting new benchmarks in the field by offering enhanced accuracy and unbiased scene graph generation.
Weijun Zhuang, Bowen Dong 0001, Zhilin Zhu 0001, Zhijun Li 0002, Jie Liu 0001, Yaowei Wang 0001, Xiaopeng Hong, Xin Li 0034, Wangmeng Zuo
IEEE Trans. Multim.8
2025 Enhanced Crowdsourced Test Report Prioritization via Image-and-Text Semantic Understanding and Feature Integration
abstract
Crowdsourced testing has gained prominence in the field of software testing due to its ability to effectively address the challenges posed by the fragmentation problem in mobile app testing. The inherent openness of crowdsourced testing brings diversity to the testing outcome. However, it also presents challenges for app developers in inspecting a substantial quantity of test reports. To help app developers inspect the bugs in crowdsourced test reports as early as possible, crowdsourced test report prioritization has emerged as an effective technology by establishing a systematic optimal report inspecting sequence. Nevertheless, crowdsourced test reports consist of app screenshots and textual descriptions, but current prioritization approaches mostly rely on textual descriptions, and some may add vectorized image features at the image-as-a-whole level or widget level. They still lack precision in accurately characterizing the distinctive features of crowdsourced test reports. In terms of prioritization strategy, prevailing approaches adopt simple prioritization based on features combined merely using weighted coefficients, without adequately considering the semantics, which may result in biased and ineffective outcomes. In this paper, we proposeEncrePrior, an enhanced crowdsourced test report prioritization approach via image-and-text semantic understanding and feature integration.EncrePriorextracts distinctive features from crowdsourced test reports. For app screenshots,EncrePriorconsiders the structure (i.e., GUI layout) and the contents (i.e., GUI widgets), viewing the app screenshot from the macroscopic and microscopic perspectives, respectively. For textual descriptions,EncrePriorconsiders the Bug Description and Reproduction Step as the bug context. During the prioritization, we do not directly merge the features with weights to guide the prioritization. Instead, in order to comprehensively consider the semantics, we adopt a prioritize-reprioritize strategy. This practice combines different features together by considering their individual ranks. The reports are first prioritized on four features separately. Then, the ranks on four sequences are used to lexicographically reprioritize the test reports with an integration of features from app screenshots and textual descriptions. Results of an empirical study show thatEncrePrioroutperforms the representative baseline approachDeepPriorby 15.61% on average, ranging from 2.99% to 63.64% on different apps, and the novelly proposed features and prioritization strategy all contribute to the excellent performance ofEncrePrior.
Chunrong Fang, Shengcheng Yu, Quanjun Zhang, Xin Li 0034, Yulei Liu, Zhenyu Chen 0001
IEEE Trans. Software Eng.4
2024 MobileInst: Video Instance Segmentation on the Mobile
abstract
Video instance segmentation on mobile devices is an important yet very challenging edge AI problem. It mainly suffers from (1) heavy computation and memory costs for frame-by-frame pixel-level instance perception and (2) complicated heuristics for tracking objects. To address these issues, we present MobileInst, a lightweight and mobile-friendly framework for video instance segmentation on mobile devices. Firstly, MobileInst adopts a mobile vision transformer to extract multi-level semantic features and presents an efficient query-based dual-transformer instance decoder for mask kernels and a semantic-enhanced mask decoder to generate instance segmentation per frame. Secondly, MobileInst exploits simple yet effective kernel reuse and kernel association to track objects for video instance segmentation. Further, we propose temporal query passing to enhance the tracking ability for kernels. We conduct experiments on COCO and YouTube-VIS datasets to demonstrate the superiority of MobileInst and evaluate the inference latency on one single CPU core of the Snapdragon 778G Mobile Platform, without other methods of acceleration. On the COCO dataset, MobileInst achieves 31.2 mask AP and 433 ms on the mobile CPU, which reduces the latency by 50% compared to the previous SOTA. For video instance segmentation, MobileInst achieves 35.0 AP and 30.1 AP on YouTube-VIS 2019 & 2021.
Renhong Zhang, Tianheng Cheng, Shusheng Yang, Haoyi Jiang, Shuai Zhang 0050, Jiancheng Lyu, Xin Li 0034, Xiaowen Ying, Dashan Gao 0001, Wenyu Liu 0001, Xinggang Wang
AAAI7
2024 RTracker: Recoverable Tracking via PN Tree Structured Memory
abstract
Existing tracking methods mainly focus on learning better target representation or developing more robust prediction models to improve tracking performance. While tracking performance has significantly improved, the target loss issue occurs frequently due to tracking failures, complete occlusion, or out-of-view situations. However, con-siderably less attention is paid to the self-recovery issue of tracking methods, which is crucial for practical applications. To this end, we propose a recoverable tracking framework, RTracker, that uses a tree-structured memory to dynamically associate a tracker and a detector to enable self-recovery ability. Specifically, we propose a Positive-Negative Tree-structured memory to chronologically store and maintain positive and negative target samples. Upon the PN tree memory, we develop corresponding walking rules for determining the state of the target and define a set of control flows to unite the tracker and the detector in different tracking scenarios. Our core idea is to use the support samples of positive and negative target categories to establish a relative distance-based criterion for a reliable assessment of target loss. The favorable performance in comparison against the state-of-the-art methods on nu-merous challenging benchmarks demonstrates the effectiveness of the proposed algorithm. All the source code and trained models will be released at https://github.com/NorahGreen/RTracker.
Yuqing Huang, Xin Li 0034, Zikun Zhou, Yaowei Wang 0001, Zhenyu He 0001, Ming-Hsuan Yang 0001
CVPR2
2024 Spatial-Temporal Multi-level Association for Video Object Segmentation
Deshui Miao, Xin Li 0034, Zhenyu He 0001, Huchuan Lu, Ming-Hsuan Yang 0001
ECCV (67)2
2024 PLS: Unsupervised Domain Adaptation for 3d Object Detection Via Pseudo-Label Sizes
abstract
3D object detection has gained increasing attention in modern autonomous driving systems. However, the performance of the detector significantly degrades during cross-domain deployment due to domain shift. The detector is inevitably biased towards its training dataset when employed on a target dataset, particularly towards object sizes. State-of-the-art unsupervised domain adaptation approaches explicitly address the variation in object sizes by appropriately scaling the source data. However, such methods require additional target domain statistics information, which contradicts the original unsupervised assumption. In this work, we present PLS, a novel unsupervised domain adaptation method for 3D object detection to overcome the object sizes bias via Pseudo-Label Sizes, which utilizes only source domain annotations. PLS alternates between generating high-quality pseudo-label sizes through the detector and model training with the pseudo-label sizes to scale and augment the source data. This iterative process enables the detector to be trained with augmented data that resembles the target domain sizes, thereby improving the performance of detector in cross-domain scenarios. Our experimental results show the outstanding performance of our PLS in various scenarios. In addition, PLS is a plug-and-play module that can be used to directly replace existing weakly-supervised scaling methods. Experimental results show that existing excellent architectures with PLS are able to achieve better performance, and making them completely unsupervised.
Rongquan Wang, Xin Li 0034, Haizhuang Liu, Jiansheng Chen 0001, Huimin Ma 0001
ICASSP3
2024 EMo Transformer: Transformer-Based Depression Detection via Eye Movements
abstract
Depressive disorder has become a prevalent psychological illness that significantly impacts individuals’ daily lives. Traditional questionnaire assessment and clinical interviews suffer from issues such as subjectivity and a high consumption of medical resources. With the advancement of artificial intelligence, there is a growing number of depression detection methods based on statistical features. However, these methods have problems of insufficient stimulus extraction and neglecting temporal information. In order to solve these problems, we propose a transformer-based model named EMo Transformer, designed for detecting depression by effectively extracting features from stimuli and combining them with eye movements. Additionally, due to challenge in collecting data from depression patients, we design a simple and effective data augmentation method to solve this challenge. Subsequently, we design an ensemble model using the models with and without data augmentation. The experimental results of accuracy 91.95% demonstrate that our method is effective.
Xin Li 0034, Haizhuang Liu, Rongquan Wang, Bochao Zou, Huimin Ma 0001
ICME1
2024 CMT: Co-training Mean-Teacher for Unsupervised Domain Adaptation on 3D Object Detection
Junbao Zhuo, Xin Li 0034, Haizhuang Liu, Rongquan Wang, Jiansheng Chen 0001, Huimin Ma 0001
ACM Multimedia3
2024 VisEvent: Reliable Object Tracking via Collaboration of Frame and Event Flows
abstract
Different from visible cameras which record intensity images frame by frame, the biologically inspired event camera produces a stream of asynchronous and sparse events with much lower latency. In practice, visible cameras can better perceive texture details and slow motion, while event cameras can be free from motion blurs and have a larger dynamic range which enables them to work well under fast motion and low illumination (LI). Therefore, the two sensors can cooperate with each other to achieve more reliable object tracking. In this work, we propose a large-scale Visible-Event benchmark (termed VisEvent) due to the lack of a realistic and scaled dataset for this task. Our dataset consists of 820 video pairs captured under LI, high speed, and background clutter scenarios, and it is divided into a training and a testing subset, each of which contains 500 and 320 videos, respectively. Based on VisEvent, we transform the event flows into event images and construct more than 30 baseline methods by extending current single-modality trackers into dual-modality versions. More importantly, we further build a simple but effective tracking algorithm by proposing a cross-modality transformer, to achieve more effective feature fusion between visible and event data. Extensive experiments on the proposed VisEvent dataset, FE108, COESOT, and two simulated datasets (i.e., OTB-DVS and VOT-DVS), validated the effectiveness of our model. The dataset and source code have been released on: https://github.com/wangxiao5791509/VisEvent_SOT_Benchmark.
Xiao Wang 0014, Jianing Li 0001, Lin Zhu 0012, Zhe Chen 0013, Xin Li 0034, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001
IEEE Trans. Cybern.6
2024 Unified Conditional Image Generation for Visible-Infrared Person Re-Identification
abstract
This paper proposes a unified multi-modal image generation method to address two critical challenges in visible-infrared (VI) person re-identification (ReID): the insufficiency of training samples and the large cross-modality discrepancy. To be specific, we propose to generate cross-modal and middle-modal images to explicitly reduce the modality discrepancy, and generate intra-modal images to serve as training samples for datasets augmentation. To this end, we adapt the conditional diffusion model for multi-modal image generation. The condition includes a binary modality indicator and modal-irrelative pedestrian contour to control the target modality and pedestrian identity, respectively. For the intra-modality and cross-modality image generation, we modify the structure of UNet to take as input the conditions, and estimate the conditional probability density by optimizing its variational lower bound. Furthermore, we devise modal discriminators and adversarial training strategies to achieve modality alignment. The middle-modality image generation method shares the same network architecture with intra- and cross-modality generation, but has specific training objectives. We define the middle modality as the distribution equidistant from the visible modality and infrared modality. We employ the adversarial training to measure the distance from the visible or infrared modality to the middle modality, and thus minimize the difference between these two adversarial losses, serving as an equidistant constraint. Experimental results on SYSU-MM01 and RegDB demonstrate the effectiveness and generalization of the intra-modality, cross-modality, and middle-modality image generation.
Honghu Pan, Wenjie Pei, Xin Li 0034, Zhenyu He 0001
IEEE Trans. Inf. Forensics Secur.3
2024 Context-Guided Black-Box Attack for Visual Tracking
abstract
With the recent advancement of deep neural networks, visual tracking has achieved substantial progress in tracking accuracy. However, the robustness and security of tracking methods developed based on current deep models have not been thoroughly explored, a critical consideration for real-world applications. In this study, we propose a context-guided black-box attack method to investigate the robustness of recent advanced deep trackers against spatial and temporal interference. For spatial interference, the proposed algorithm generates adversarial target samples by mixing the information of the target object and the similar background regions around it in an embedded feature space of an encoder-decoder model, which evaluates the ability of trackers to handle background distractors. For temporal interference, we use the target state in the previous frame to generate the adversarial sample, which easily fools the trackers that rely too heavily on tracking prior assumptions, such as that the appearance changes and movements of a video target object are small between two consecutive frames. We assess the proposed attack method under both CNN-based and transformer-based tracking frameworks on four diverse datasets: OTB100, VOT2018, GOT-10k, and LaSOT. The experimental results demonstrate that our approach substantially deteriorates the performance of all these deep trackers across numerous datasets, even in the black-box attack mode. This reveals the weak robustness of recent deep tracking methods against background distractors and prior dependencies.
Xingsen Huang, Deshui Miao, Hongpeng Wang 0002, Yaowei Wang 0001, Xin Li 0034
IEEE Trans. Multim.5
2024 Self-Supervised Tracking via Target-Aware Data Synthesis
abstract
While deep-learning-based tracking methods have achieved substantial progress, they entail large-scale and high-quality annotated data for sufficient training. To eliminate expensive and exhaustive annotation, we study self-supervised (SS) learning for visual tracking. In this work, we develop the crop-transform-paste operation, which is able to synthesize sufficient training data by simulating various appearance variations during tracking, including appearance variations of objects and background interference. Since the target state is known in all synthesized data, existing deep trackers can be trained in routine ways using the synthesized data without human annotation. The proposed target-aware data-synthesis method adapts existing tracking approaches within a SS learning framework without algorithmic changes. Thus, the proposed SS learning mechanism can be seamlessly integrated into existing tracking frameworks to perform training. Extensive experiments show that our method: 1) achieves favorable performance against supervised (Su) learning schemes under the cases with limited annotations; 2) helps deal with various tracking challenges such as object deformation, occlusion (OCC), or background clutter (BC) due to its manipulability; 3) performs favorably against the state-of-the-art unsupervised tracking methods; and 4) boosts the performance of various state-of-the-art Su learning frameworks, including SiamRPN++, DiMP, and TransT.
Xin Li 0034, Wenjie Pei, Yaowei Wang 0001, Zhenyu He 0001, Huchuan Lu, Ming-Hsuan Yang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 LSOTB-TIR: A Large-Scale High-Diversity Thermal Infrared Single Object Tracking Benchmark
abstract
Unlike visual object tracking, thermal infrared (TIR) object tracking methods can track the target of interest in poor visibility such as rain, snow, and fog, or even in total darkness. This feature brings a wide range of application prospects for TIR object-tracking methods. However, this field lacks a unified and large-scale training and evaluation benchmark, which has severely hindered its development. To this end, we present a large-scale and high-diversity unified TIR single object tracking benchmark, called LSOTB-TIR, which consists of a tracking evaluation dataset and a general training dataset with a total of 1416 TIR sequences and more than 643 K frames. We annotate the bounding box of objects in every frame of all sequences and generate over 770 K bounding boxes in total. To the best of our knowledge, LSOTB-TIR is the largest and most diverse TIR object tracking benchmark to date. We spilt the evaluation dataset into a short-term tracking subset and a long-term tracking subset to evaluate trackers using different paradigms. What's more, to evaluate a tracker on different attributes, we also define four scenario attributes and 12 challenge attributes in the short-term tracking evaluation subset. By releasing LSOTB-TIR, we encourage the community to develop deep learning-based TIR trackers and evaluate them fairly and comprehensively. We evaluate and analyze 40 trackers on LSOTB-TIR to provide a series of baselines and give some insights and future research directions in TIR object tracking. Furthermore, we retrain several representative deep trackers on LSOTB-TIR, and their results demonstrate that the proposed training dataset significantly improves the performance of deep TIR trackers. Codes and dataset are available at https://github.com/QiaoLiuHit/LSOTB-TIR.
Qiao Liu 0001, Xin Li 0034, Di Yuan 0002, Xiaojun Chang, Zhenyu He 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Effective, Platform-Independent GUI Testing via Image Embedding and Reinforcement Learning
abstract
Software applications (apps) have been playing an increasingly important role in various aspects of society. In particular, mobile apps and web apps are the most prevalent among all applications and are widely used in various industries as well as in people’s daily lives. To help ensure mobile and web app quality, many approaches have been introduced to improve app GUI testing via automated exploration, including random testing, model-based testing, learning-based testing, and so on. Despite the extensive effort, existing approaches are still limited in reaching high code coverage, constructing high-quality models, and being generally applicable. Reinforcement learning-based approaches, as a group of representative and advanced approaches for automated GUI exploration testing, are faced with difficult challenges, including effective app state abstraction, reward function design, and so on. Moreover, they heavily depend on the specific execution platforms (i.e., Android or Web), thus leading to poor generalizability and being unable to adapt to different platforms. This work specifically tackles these challenges based on the high-level observation that apps from distinct platforms share commonalities in GUI design. Indeed, we propose PIRLTest , an effective platform-independent approach for app testing. Specifically, PIRLTest utilizes computer vision and reinforcement learning techniques in a novel, synergistic manner for automated testing. It extracts the GUI widgets from GUI pages and characterizes the corresponding GUI layouts, embedding the GUI pages as states. The app GUI state combines the macroscopic perspective (app GUI layout) and the microscopic perspective (app GUI widget) and attaches the critical semantic information from GUI images. This enables PIRLTest to be platform-independent and makes the testing approach generally applicable on different platforms. PIRLTest explores apps with the guidance of a curiosity-driven strategy, which uses a Q-network to estimate the values of specific state-action pairs to encourage more exploration in uncovered pages without platform dependency. The exploration will be assigned with rewards for all actions, which are designed considering both the app GUI states and the concrete widgets, to help the framework explore more uncovered pages. We conduct an empirical study on 20 mobile apps and 5 web apps, and the results show that PIRLTest is zero-cost when being adapted to different platforms, and can perform better than the baselines, covering 6.3–41.4% more code on mobile apps and 1.5–51.1% more code on web apps. PIRLTest is capable of detecting 128 unique bugs on mobile and web apps, including 100 bugs that cannot be detected by the baselines.
Shengcheng Yu, Chunrong Fang, Xin Li 0034, Yuchen Ling, Zhenyu Chen 0001, Zhendong Su 0001
ACM Trans. Softw. Eng. Methodol.3
2023 CiteTracker: Correlating Image and Text for Visual Tracking
abstract
Existing visual tracking methods typically take an image patch as the reference of the target to perform tracking. However, a single image patch cannot provide a complete and precise concept of the target object as images are limited in their ability to abstract and can be ambiguous, which makes it difficult to track targets with drastic variations. In this paper, we propose the CiteTracker to enhance target modeling and inference in visual tracking by connecting images and text. Specifically, we develop a text generation module to convert the target image patch into a descriptive text containing its class and attribute information, providing a comprehensive reference point for the target. In addition, a dynamic description module is designed to adapt to target variations for more effective target representation. We then associate the target description and the search image using an attention-based correlation module to generate the correlated features for target state reference. Extensive experiments on five diverse datasets are conducted to evaluate the proposed algorithm and the favorable performance against the state-of-the-art methods demonstrates the effectiveness of the proposed tracking method. The source code and trained models will be released at https://github.com/NorahGreen/CiteTracker.
Xin Li 0034, Yuqing Huang, Zhenyu He 0001, Yaowei Wang 0001, Huchuan Lu, Ming-Hsuan Yang 0001
ICCV1
2023 Micro-Expression Spotting with Face Alignment and Optical Flow
abstract
Facial expression spotting holds significant importance as it can signify emotional changes. Particularly, micro-expressions possess the potential to reveal genuine emotions, making them even more valuable in practical domains such as public safety and finance. However, spotting micro-expressions proves challenging due to their subtle movements and brief duration. This paper proposes an expression spotting method based on face alignment and optical flow. We first use a finer crop-align technique to preprocess the facial videos by aligning the face and the nose tip. Then, regions of interest (ROIs) are defined by analyzing the statistics of action units. The optical flow features are then extracted and subjected to low-pass filtering to eliminate high-frequency noise. Furthermore, candidate expression segments are identified based on the magnitude of the processed optical flows. Finally, non-maximum suppression is utilized to remove overlapping segments. The effectiveness of the proposed method is evaluated on the challenge test set, resulting in an overall F1-score of 0.19. Additional results obtained from CAS(ME)2 and SAMM Long videos provide further verification of the method's efficacy. The code is available online.
Wenfeng Qin, Bochao Zou, Xin Li 0034, Weiping Wang 0007, Huimin Ma 0001
ACM Multimedia3
2023 Siamese residual network for efficient visual tracking
Nana Fan, Qiao Liu 0001, Xin Li 0034, Zikun Zhou, Zhenyu He 0001
Inf. Sci.3
2023 Learning Dual-Level Deep Representation for Thermal Infrared Tracking
abstract
The feature models used by existing Thermal InfraRed (TIR) tracking methods are usually learned from RGB images due to the lack of a large-scale TIR image training dataset. However, these feature models are less effective in representing TIR objects and they are difficult to effectively distinguish distractors because they do not contain fine-grained discriminative information. To this end, we propose a dual-level feature model containing the TIR-specific discriminative feature and fine-grained correlation feature for robust TIR object tracking. Specifically, to distinguish inter-class TIR objects, we first design an auxiliary multi-classification network to learn the TIR-specific discriminative feature. Then, to recognize intra-class TIR objects, we propose a fine-grained aware module to learn the fine-grained correlation feature. These two kinds of features complement each other and represent TIR objects in the levels of inter-class and intra-class respectively. These two feature models are constructed using a multi-task matching framework and are jointly optimized on the TIR object tracking task. In addition, we develop a large-scale TIR image dataset to train the network for learning TIR-specific feature patterns. To the best of our knowledge, this is the largest TIR tracking training dataset with the richest object class and scenario. To verify the effectiveness of the proposed dual-level feature model, we propose an offline TIR tracker (MMNet) and an online TIR tracker (ECO-MM) based on the feature model and evaluate them on three TIR tracking benchmarks. Extensive experimental results on these benchmarks demonstrate that the proposed algorithms perform favorably against the state-of-the-art methods.
Qiao Liu 0001, Di Yuan 0002, Nana Fan, Peng Gao 0005, Xin Li 0034, Zhenyu He 0001
IEEE Trans. Multim.5
2022 Multi-Object Tracking Meets Moving UAV
abstract
Multi-object tracking in unmanned aerial vehicle (UAV) videos is an important vision task and can be applied in a wide range of applications. However, conventional multi-object trackers do not work well on UAV videos due to the challenging factors of irregular motion caused by moving camera and view change in 3D directions. In this paper, we propose a UAVMOT network specially for multi-object tracking in UAV views. The UAVMOT introduces an ID feature update module to enhance the object's feature association. To better handle the complex motions under UAV views, we develop an adaptive motion filter module. In addition, a gradient balanced focal loss is used to tackle the imbalance categories and small objects detection problem. Experimental results on the VisDrone2019 and UAVDT datasets demonstrate that the proposed UAVMOT achieves considerable improvement against the state-of-the-art tracking methods on UAV videos.
Shuai Liu 0009, Xin Li 0034, Huchuan Lu, You He 0002
CVPR2
2022 Unsupervised Learning of Accurate Siamese Tracking
abstract
Unsupervised learning has been popular in various computer vision tasks, including visual object tracking. However, prior unsupervised tracking approaches rely heavily on spatial supervision from templatesearch pairs and are still unable to track objects with strong variation over a long time span. As unlimited self-supervision signals can be obtained by tracking a video along a cycle in time, we investigate evolving a Siamese tracker by tracking videos forward-backward. We present a novel unsupervised tracking framework, in which we can learn temporal correspondence both on the classification branch and regression branch. Specifically, to propagate reliable template feature in the forward propagation process so that the tracker can be trained in the cycle, we first propose a consistency propagation transformation. We then identify an ill-posed penalty problem in conventional cycle training in backward propagation process. Thus, a differentiable region mask is proposed to select features as well as to implicitly penalize tracking errors on intermediate frames. Moreover, since noisy labels may degrade training, we propose a mask-guided loss reweighting strategy to assign dynamic weights based on the quality of pseudo labels. In extensive experiments, our tracker outperforms preceding unsupervised methods by a substantial margin, performing on par with supervised methods on large-scale datasets such as TrackingNet and LaSOT. Code is available at https://github.com/FlorinShum/ULAST.
Qiuhong Shen, Lei Qiao 0004, Jinyang Guo 0002, Peixia Li, Xin Li 0034, Bo Li 0114, Weihao Gan, Wei Wu 0021, Wanli Ouyang
CVPR5
2022 UniRLTest: universal platform-independent testing with reinforcement learning via image understanding
abstract
GUI testing has been prevailing in software testing. However, existing automated GUI testing tools mostly rely on frameworks of a specific platform. Testers have to fully understand platform features before developing platform-dependent GUI testing tools. Starting from the perspective of tester’s vision, we observe that GUIs on different platforms share commonalities of widget images and layout designs, which can be leveraged to achieve platform-independent testing. We propose UniRLTest, an automated software testing framework, to achieve platform independence testing. UniRLTest utilizes computer vision techniques to capture all the widgets in the screenshot and constructs a widget tree for each page. A set of all the executable actions in each tree will be generated accordingly. UniRLTest adopts a Deep Q-Network, a reinforcement learning (RL) method, to the exploration process and formalize the Android GUI testing problem to a Marcov Decision Process (MDP), where RL could work. We have conducted evaluation experiments on 25 applications from different platforms. The result shows that UniRLTest outperforms baselines in terms of efficiency and effectiveness.
Yulei Liu, Shengcheng Yu, Xin Li 0034, Yexiao Yun, Chunrong Fang, Zhenyu Chen 0001
ISSTA4
2022 Noise-Suppressing Deep Tracking
abstract
In visual tracking, it is challenging to distinguish the target from similar objects called noises in the background. As deep trackers use convolutional neural networks for image classification as feature extractors, the extracted features are insensitive to different instances in the same class, which is prone to make prediction models confuse the target and the similar noises in the background. To this end, we propose a noise-suppressing algorithm to learn the discriminative representation for distinguishing the target from the noises in the background. First, we learn polynomial kernels for a search patch under the semantic guidance to increase the difference between representations of the target and the noises in the background. Second, we formulate the online foreground-background functions for the target and the noises in the background to learn an adaptive kernel, which suppresses the features positive for the noises and promotes the features positive for the target. We evaluate the proposed method on seven public datasets including OTB-2013, OTB-2015, VOT-2018, LaSOT, TrackingNet, GOT10k, and NFS. The comprehensive experimental results show that the proposed algorithm performs favorably against state-of-the-art methods, while running at real-time speed.
Nana Fan, Xin Li 0034, Zikun Zhou, Qiao Liu 0001, Zhenyu He 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Target-Aware State Estimation for Visual Tracking
abstract
Trackers based on the IoU prediction network (IoU-Net) have shown superior performance, which refines a coarse bounding box to an accurate one by maximizing the IoU between the target and the coarse box. However, the traditional IoU-Net is less effective in exploiting the limited but crucial supervision information contained in the initial frame, including the discriminative information between the target and backgrounds and the structure information of the initial target. Missing such information makes the IoU-Net less robust to background distractors and diverse variations of the target appearance. To address this issue, we propose a target-aware state estimation network for visual tracking. A gradient-guided feature adjustment module is built on an online discriminative model to generate target-aware features for constructing the state estimation network; it conveys the online learned discriminative information into the offline trained state estimation network. In addition, we propose a structure-aware integration module and embed it into the state estimation network, enabling the tracker to explicitly model the structure information of the initial target. Extensive experimental results on the VOT2018, OTB2015, UAV123, NFS30, TC128, TrackingNet, LaSOT, and VOT2018-LT datasets demonstrate that the proposed approach performs favorably against state-of-the-art trackers.
Zikun Zhou, Xin Li 0034, Nana Fan, Hongpeng Wang 0002, Zhenyu He 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Object Tracking via Spatial-Temporal Memory Network
abstract
Temporal and spatial contexts, characterizing target appearance variations and target-background differences, respectively, are crucial for improving the online adaptive ability and instance-level discriminative ability of object tracking. However, most existing trackers focus on either the temporal context or the spatial context during tracking and have not exploited these contexts simultaneously and effectively. In this paper, we propose a Spatial-TEmporal Memory (STEM) network to exploit these contexts jointly for object tracking. Specifically, we develop a key-value structured memory model equipped with a key-value index-based memory reading mechanism to model the spatial and temporal contexts simultaneously. To update the memory with new target states and ensure the diversity of the memory, we introduce a similarity-aware memory update scheme. In addition, we construct an entropy-guided ensemble strategy to fuse the prediction models based on these two contexts, such that these two contexts can be exploited to estimate the target state jointly. Extensive experimental results on eight challenging datasets, including OTB2015, TC128, UAV123, VOT2018, LaSOT, TrackingNet, GOT-10k, and OxUvA, demonstrate that the proposed method performs favorably against state-of-the-art trackers.
Zikun Zhou, Xin Li 0034, Tianzhu Zhang 0001, Hongpeng Wang 0002, Zhenyu He 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 SiamCorners: Siamese Corner Networks for Visual Tracking
abstract
The current Siamese network based on region proposal network (RPN) has attracted great attention in visual tracking due to its excellent accuracy and high efficiency. However, the design of the RPN involves the selection of the number, scale, and aspect ratios of anchor boxes, which will affect the applicability and convenience of the model. Furthermore, these anchor boxes require complicated calculations, such as calculating their intersection-over-union (IoU) with ground truth bounding boxes. Due to the problems related to anchor boxes, we propose a simple yet effective anchor-free tracker (named Siamese corner networks, SiamCorners), which is end-to-end trained offline on large-scale image pairs. Specifically, we introduce a modified corner pooling layer to convert the bounding box estimate of the target into a pair of corner predictions (the bottom-right and the top-left corners). By tracking a target as a pair of corners, we avoid the need to design the anchor boxes. This will make the entire tracking algorithm more flexible and simple than anchor-based trackers. In our network design, we further introduce a layer-wise feature aggregation strategy that enables the corner pooling module to predict multiple corners for a tracking target in deep networks. We then introduce a new penalty term that is used to select an optimal tracking box in these candidate corners. Finally, SiamCorners achieves experimental results that are comparable to the state-of-art tracker while maintaining a high running speed. In particular, SiamCorners achieves a 53.7% AUC on NFS30 and a 61.4% AUC on UAV123, while still running at 42 frames per second (FPS).
Kai Yang 0018, Zhenyu He 0001, Wenjie Pei, Zikun Zhou, Xin Li 0034, Di Yuan 0002, Haijun Zhang 0002
IEEE Trans. Multim.5
2021 Saliency-Associated Object Tracking
abstract
Most existing trackers based on deep learning perform tracking in a holistic strategy, which aims to learn deep representations of the whole target for localizing the target. It is arduous for such methods to track targets with various appearance variations. To address this limitation, another type of methods adopts a part-based tracking strategy which divides the target into equal patches and tracks all these patches in parallel. The target state is inferred by summarizing the tracking results of these patches. A potential limitation of such trackers is that not all patches are equally informative for tracking. Some patches that are not discriminative may have adverse effects. In this paper, we propose to track the salient local parts of the target that are discriminative for tracking. In particular, we propose a fine-grained saliency mining module to capture the local saliencies. Further, we design a saliency-association modeling module to associate the captured saliencies together to learn effective correlation representations between the exemplar and the search image for state estimation. Extensive experiments on five diverse datasets demonstrate that the proposed method performs favorably against state-of-the-art trackers.
Zikun Zhou, Wenjie Pei, Xin Li 0034, Hongpeng Wang 0002, Feng Zheng 0001, Zhenyu He 0001
ICCV3
2021 Interactive convolutional learning for visual tracking
Nana Fan, Qiao Liu 0001, Xin Li 0034, Zikun Zhou, Zhenyu He 0001
Knowl. Based Syst.3
2021 Learning dual-margin model for visual tracking
Nana Fan, Xin Li 0034, Zikun Zhou, Qiao Liu 0001, Zhenyu He 0001
Neural Networks2
2021 Learning Deep Multi-Level Similarity for Thermal Infrared Object Tracking
abstract
Existing deep Thermal InfraRed (TIR) trackers only use semantic features to represent the TIR object, which lack the sufficient discriminative capacity for handling distractors. This becomes worse when the feature extraction network is only trained on RGB images. To address this issue, we propose a multi-level similarity model under a Siamese framework for robust TIR object tracking. Specifically, we compute different pattern similarities using the proposed multi-level similarity network. One of them focuses on the global semantic similarity and the other computes the local structural similarity of the TIR object. These two similarities complement each other and hence enhance the discriminative capacity of the network for handling distractors. In addition, we design a simple while effective relative entropy based ensemble subnetwork to integrate the semantic and structural similarities. This subnetwork can adaptive learn the weights of the semantic and structural similarities at the training stage. To further enhance the discriminative capacity of the tracker, we propose a large-scale TIR video sequence dataset for training the proposed model. To the best of our knowledge, this is the first and the largest TIR object tracking training dataset to date. The proposed TIR dataset not only benefits the training for TIR object tracking but also can be applied to numerous TIR visual tasks. Extensive experimental results on three benchmarks demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods.
Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Nana Fan, Di Yuan 0002, Hongpeng Wang 0002
IEEE Trans. Multim.2
2020 Multi-Task Driven Feature Models for Thermal Infrared Tracking
abstract
Existing deep Thermal InfraRed (TIR) trackers usually use the feature models of RGB trackers for representation. However, these feature models learned on RGB images are neither effective in representing TIR objects nor taking fine-grained TIR information into consideration. To this end, we develop a multi-task framework to learn the TIR-specific discriminative features and fine-grained correlation features for TIR tracking. Specifically, we first use an auxiliary classification network to guide the generation of TIR-specific discriminative features for distinguishing the TIR objects belonging to different classes. Second, we design a fine-grained aware module to capture more subtle information for distinguishing the TIR objects belonging to the same class. These two kinds of features complement each other and recognize TIR objects in the levels of inter-class and intra-class respectively. These two feature models are learned using a multi-task matching framework and are jointly optimized on the TIR tracking task. In addition, we develop a large-scale TIR training dataset to train the network for adapting the model to the TIR domain. Extensive experimental results on three benchmarks show that the proposed algorithm achieves a relative gain of 10% over the baseline and performs favorably against the state-of-the-art methods. Codes and the proposed TIR dataset are available at https://github.com/QiaoLiuHit/MMNet.
Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Nana Fan, Di Yuan 0002, Wei Liu 0065, Yongsheng Liang 0001
AAAI2
2020 Negative-Aware Training: Be Aware of Negative Samples
Xin Li 0034, Xiaodong Jia 0005, Xiaoyuan Jing
ECAI1
2020 LSOTB-TIR: A Large-Scale High-Diversity Thermal Infrared Object Tracking Benchmark
abstract
In this paper, we present a Large-Scale and high-diversity general Thermal InfraRed (TIR) Object Tracking Benchmark, called LSOTB-TIR, which consists of an evaluation dataset and a training dataset with a total of 1,400 TIR sequences and more than 600K frames. We annotate the bounding box of objects in every frame of all sequences and generate over 730K bounding boxes in total. To the best of our knowledge, LSOTB-TIR is the largest and most diverse TIR object tracking benchmark to date. To evaluate a tracker on different attributes, we define 4 scenario attributes and 12 challenge attributes in the evaluation dataset. By releasing LSOTB-TIR, we encourage the community to develop deep learning based TIR trackers and evaluate them fairly and comprehensively. We evaluate and analyze more than 30 trackers on LSOTB-TIR to provide a series of baselines, and the results show that deep trackers achieve promising performance. Furthermore, we re-train several representative deep trackers on LSOTB-TIR, and their results demonstrate that the proposed training dataset significantly improves the performance of deep TIR trackers. Codes and dataset are available at https://github.com/QiaoLiuHit/LSOTB-TIR.
Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Chenglong Li 0002, Zikun Zhou, Di Yuan 0002, Jing Li 0071, Kai Yang 0018, Nana Fan, Feng Zheng 0001
ACM Multimedia2
2020 Visual object tracking with adaptive structural convolutional network
Di Yuan 0002, Xin Li 0034, Zhenyu He 0001, Qiao Liu 0001, Shuwei Lu
Knowl. Based Syst.2
2020 Dual-regression model for visual tracking
Xin Li 0034, Qiao Liu 0001, Nana Fan, Zikun Zhou, Zhenyu He 0001, Xiaoyuan Jing
Neural Networks1
2020 PTB-TIR: A Thermal Infrared Pedestrian Tracking Benchmark
abstract
Thermal infrared (TIR) pedestrian tracking is one of the important components among numerous applications of computer vision, which has a major advantage: it can track pedestrians in total darkness. The ability to evaluate the TIR pedestrian tracker fairly, on a benchmark dataset, is significant for the development of this field. However, there is not a benchmark dataset. In this paper, we develop a TIR pedestrian tracking dataset for the TIR pedestrian tracker evaluation. The dataset includes 60 thermal sequences with manual annotations. Each sequence has nine attribute labels for the attribute based evaluation. In addition to the dataset, we carry out the large-scale evaluation experiments on our benchmark dataset using nine publicly available trackers. The experimental results help us understand the strengths and weaknesses of these trackers. In addition, in order to gain more insight into the TIR pedestrian tracker, we divide its functions into three components: feature extractor, motion model, and observation model. Then, we conduct three comparison experiments on our benchmark dataset to validate how each component affects the tracker's performance. The findings of these experiments provide some guidelines for future research.
Qiao Liu 0001, Zhenyu He 0001, Xin Li 0034, Yuan Zheng 0002
IEEE Trans. Multim.3
2019 Target-Aware Deep Tracking
abstract
Existing deep trackers mainly use convolutional neural networks pre-trained for the generic object recognition task for representations. Despite demonstrated successes for numerous vision tasks, the contributions of using pre-trained deep features for visual tracking are not as significant as that for object recognition. The key issue is that in visual tracking the targets of interest can be arbitrary object class with arbitrary forms. As such, pre-trained deep features are less effective in modeling these targets of arbitrary forms for distinguishing them from the background. In this paper, we propose a novel scheme to learn target-aware features, which can better recognize the targets undergoing significant appearance variations than pre-trained deep features. To this end, we develop a regression loss and a ranking loss to guide the generation of target-active and scale-sensitive features. We identify the importance of each convolutional filter according to the back-propagated gradients and select the target-aware features based on activations for representing the targets. The target-aware features are integrated with a Siamese matching network for visual tracking. Extensive experimental results show that the proposed algorithm performs favorably against the state-of-the-art methods in terms of accuracy and speed.
Xin Li 0034, Chao Ma 0004, Baoyuan Wu, Zhenyu He 0001, Ming-Hsuan Yang 0001
CVPR1
2019 Region-filtering correlation tracking
Nana Fan, Jing Li 0071, Zhenyu He 0001, Chunkai Zhang, Xin Li 0034
Knowl. Based Syst.5
2019 Hierarchical spatial-aware Siamese network for thermal infrared object tracking
Xin Li 0034, Qiao Liu 0001, Nana Fan, Zhenyu He 0001, Hongzhi Wang 0001
Knowl. Based Syst.1
2017 Wound intensity correction and segmentation with convolutional neural networks
abstract
Summary Wound area changes over multiple weeks are highly predictive of the wound healing process. A big data eHealth system would be very helpful in evaluating these changes. We usually analyze images of the wound bed for diagnosing injury. Unfortunately, accurate measurements of wound region changes from images are difficult. Many factors affect the quality of images, such as intensity inhomogeneity and color distortion. To this end, we propose a fast level set model‐based method for intensity inhomogeneity correction and a spectral properties‐based color correction method to overcome these obstacles. State‐of‐the‐art level set methods can segment objects well. However, such methods are time‐consuming and inefficient. In contrast to conventional approaches, the proposed model integrates a new signed energy force function that can detect contours at weak or blurred edges efficiently. It ensures the smoothness of the level set function and reduces the computational complexity of re‐initialization. To increase the speed of the algorithm further, we also include an additive operator‐splitting algorithm in our fast level set model. In addition, we consider using a camera, lighting, and spectral properties to recover the actual color. Numerical synthetic and real‐world images demonstrate the advantages of the proposed method over state‐of‐the‐art methods. Experimental results also show that the proposed model is at least twice as fast as methods used widely. Copyright © 2016 John Wiley & Sons, Ltd.
Huimin Lu 0001, Bin Li 0006, Junwu Zhu, Yujie Li 0001, Yun Li 0010, Xing Xu 0001, Li He 0001, Xin Li 0034, Jianru Li, Seiichi Serikawa
Concurr. Comput. Pract. Exp.8
2016 An Efficient Auction Mechanism Toward Heterogeneous Spectrum Allocation
Haiyan Qin, Xin Li 0034, Yonglong Zhang 0001, Bin Li 0006
IDEAL2
2016 Underwater image enhancement method using weighted guided trigonometric filtering and artificial light correction
Huimin Lu 0001, Yujie Li 0001, Xing Xu 0001, Jian-Ru Lin, Zhifei Liu, Xin Li 0034, Jianmin Yang, Seiichi Serikawa
J. Vis. Commun. Image Represent.6
2016 A multi-view model for visual tracking via correlation filters
Xin Li 0034, Qiao Liu 0001, Zhenyu He 0001, Hongpeng Wang 0002, Chunkai Zhang
Knowl. Based Syst.1
2016 Connected Component Model for Multi-Object Tracking
abstract
In multi-object tracking, it is critical to explore the data associations by exploiting the temporal information from a sequence of frames rather than the information from the adjacent two frames. Since straightforwardly obtaining data associations from multi-frames is an NP-hard multi-dimensional assignment (MDA) problem, most existing methods solve this MDA problem by either developing complicated approximate algorithms, or simplifying MDA as a 2D assignment problem based upon the information extracted only from adjacent frames. In this paper, we show that the relation between associations of two observations is the equivalence relation in the data association problem, based on the spatial-temporal constraint that the trajectories of different objects must be disjoint. Therefore, the MDA problem can be equivalently divided into independent subproblems by equivalence partitioning. In contrast to existing works for solving the MDA problem, we develop a connected component model (CCM) by exploiting the constraints of the data association and the equivalence relation on the constraints. Based upon CCM, we can efficiently obtain the global solution of the MDA problem for multi-object tracking by optimizing a sequence of independent data association subproblems. Experiments on challenging public data sets demonstrate that our algorithm outperforms the state-of-the-art approaches.
Zhenyu He 0001, Xin Li 0034, Xinge You, Dacheng Tao, Yuan Yan Tang
IEEE Trans. Image Process.2
2015 A robust local sparse tracker with global consistency constraint
Xinhua You, Xin Li 0034, Zhenyu He 0001, Xiaofeng Zhang 0002
Signal Process.2
2014 A novel joint tracker based on occlusion detection
Xin Li 0034, Zhenyu He 0001, Xinge You, C. L. Philip Chen
Knowl. Based Syst.1