Yinchao Ma

dblp:189/1326 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2026
0009-0003-0506-5217ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021
YearPublicationVenuePosition
2026 Contribution-aware Token Compression for Efficient Video Understanding via Reinforcement Learning
abstract
Video large language models have demonstrated remarkable capabilities in video understanding tasks. However, the redundancy of video tokens introduces significant computational overhead during inference, limiting their practical deployment. Many compression algorithms are proposed to prioritize retaining features with the highest attention scores to minimize perturbations in attention computations. However, the correlation between attention scores and their actual contribution to correct answers remains ambiguous. To address the above limitation, we propose a novel contribution-aware token compression algorithm for video understanding (CaCoVID) that explicitly optimizes the token selection policy based on the contribution of tokens to correct predictions. First, we introduce a reinforcement learning-based framework that optimizes a policy network to select video token combinations with the greatest contribution to correct predictions. This paradigm shifts the focus from passive token preservation to active discovery of optimal compressed token combinations. Secondly, we propose a combinatorial policy optimization algorithm with online combination space sampling, which dramatically reduces the exploration space for video token combinations and accelerates the convergence speed of policy optimization. Extensive experiments on diverse video understanding benchmarks demonstrate the effectiveness of CaCoVID. Codes will be released.
Yinchao Ma, Xianing Chen, Hanqing Yang 0010, Bo Zheng 0007
AAAI1
2026 Unified Thinker: A General Reasoning Core for Image Generation
abstract
Sashuai Zhou, Qiang Zhou, Jijin Hu, Hanqing Yang, Yue Cao, Junpeng Ma, Yinchao Ma, Jun Song, Tiezheng Ge, Cheng Yu, Bo Zheng, Zhou Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sashuai Zhou, Jijin Hu, Hanqing Yang 0010, Junpeng Ma, Yinchao Ma, Tiezheng Ge, Bo Zheng 0007, Zhou Zhao 0001
ACL (1)7
2026 UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
abstract
Single object tracking aims to localize target object with specific reference modalities (bounding box, natural language or both) in a sequence of specific video modalities (RGB, RGB+Depth, RGB+Thermal or RGB+Event.). Different reference modalities enable various human-machine interactions, and different video modalities are demanded in complex scenarios to enhance tracking robustness. Existing trackers are designed for single or several video modalities with single or several reference modalities, which leads to separate model designs and limits practical applications. Practically, a unified tracker is needed to handle various requirements. To the best of our knowledge, there is still no tracker that can perform tracking with these above reference modalities across these video modalities simultaneously. Thus, in this paper, we present a unified tracker, UniSOT, for different combinations of three reference modalities and four video modalities with uniform parameters. Extensive experimental results on 18 visual tracking, vision-language tracking and RGB+X tracking benchmarks demonstrate that UniSOT shows superior performance against modality-specific counterparts. Notably, UniSOT outperforms previous counterparts by over 3.0% AUC on TNL2K across all three reference modalities and outperforms Un-Track by over 2.0% main metric across all three RGB+X video modalities.
Yinchao Ma, Yuyang Tang 0001, Wenfei Yang, Tianzhu Zhang 0001, Feng Wu 0005
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 State Space Models for Natural Language Tracking: Exploring Context-Adaptive Language Cues
abstract
Natural language tracking aims to locate the target of a video based on a language description. The rich contextual information inside the language and video sequence is essential to describe the target movements and appearance variations. However, existing natural language trackers design a fixed-length memory to store historical target information, which merely uses limited context information and necessitates manually designed modules, resulting in sub-optimal localization performance and numerous computational costs. Inspired by the success of the state space model, we propose a novel Context-adaptive Mamba Tracker (CMTrack). It enjoys several merits. First, we propose a novel context-aware state space model that enables language features to serve as hidden states to interact with relevant image features adaptively. Second, CMTrack transfers the hidden states frame-by-frame to continuously incorporate contextual target information into language features, enabling context-adaptive language cues. Third, the proposed context-adaptive language cues can effectively capture the long-range behavior of the target and guide the tracker in locating the target accurately without any extra design. Finally, CMTrack provides a neat pipeline for training and tracking with linear complexity. Experimental results demonstrate that CMTrack achieves new state-of-the-art performance.
Yuyang Tang 0001, Yinchao Ma, Dengqing Yang, Tianzhu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 State Space Models for Long-Term Temporal Context in 3D Single Object Tracking
abstract
3D single object tracking (3D SOT) remains a challenging task due to the sparsity of point clouds, appearance variations caused by occlusions, and the difficulty of modeling long-term temporal context. Although recent Transformer-based approaches leverage memory mechanisms to propagate temporal information, their quadratic complexity and reliance on discrete historical snapshots limit both efficiency and temporal coherence. To address these limitations, we propose SSMTrack, a novel 3D SOT framework built upon state space models (SSMs), which efficiently models long-term temporal dependencies through a continuously evolving hidden state with linear complexity. Specifically, we introduce a serialization and bidirectional scanning (SBS) strategy to enhance intra-frame feature interactions and design a Target-Aware Encoder (TAE) to extract target cues while maintaining stable temporal representations. Furthermore, we propose a Temporal Causal Shape Learning (TCSL) mechanism that preserves critical historical information while adaptively integrating current inputs, progressively enriching target feature representations over time. Extensive experiments on three benchmark datasets demonstrate that SSMTrack achieves state-of-the-art performance with strong temporal coherence and high efficiency. The code will be released upon publication.
Yinchao Ma, Yuyang Tang 0001, Chuxin Wang, Tianzhu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 SMTrack: State-Aware Mamba for Efficient Temporal Modeling in Visual Tracking
abstract
Visual tracking aims to automatically estimate the state of a target object in a video sequence, which is challenging especially in dynamic scenarios. Thus, numerous methods are proposed to introduce temporal cues to enhance tracking robustness. However, conventional CNN and Transformer architectures exhibit inherent limitations in modeling long-range temporal dependencies in visual tracking, often necessitating either complex customized modules or substantial computational costs to integrate temporal cues. Inspired by the success of the state space model, we propose a novel temporal modeling paradigm for visual tracking, termed State-aware Mamba Tracker (SMTrack), providing a neat pipeline for training and tracking without needing customized modules or substantial computational costs to build long-range temporal dependencies. It enjoys several merits. First, we propose a novel selective state-aware space model with state-wise parameters to capture more diverse temporal cues for robust tracking. Second, SMTrack facilitates long-range temporal interactions with linear computational complexity during training. Third, SMTrack enables each frame to interact with previously tracked frames via hidden state propagation and updating, which releases computational costs of handling temporal cues during tracking. Extensive experimental results demonstrate that SMTrack achieves promising performance with low computational costs.
Yinchao Ma, Dengqing Yang, Zhangyu He, Wenfei Yang, Tianzhu Zhang 0001
IEEE Trans. Image Process.1
2025 Learning Discriminative Features for Visual Tracking via Scenario Decoupling
Yinchao Ma, Qianjin Yu, Wenfei Yang, Tianzhu Zhang 0001
Int. J. Comput. Vis.1
2025 Semantic-Aware Network for Natural Language Tracking
abstract
Natural language tracking aims to locate the position of a target specified by a natural language description. Existing methods are trained on vision-language datasets with a small number of language descriptions, which may lead to limited semantic generalization. Moreover, they extract visual and language features separately, which limits visual-semantic capabilities. To overcome these limitations, we propose a novel semantic-aware tracking framework, SATrack, which integrates a semantic-aware attention module and a cross-modal aggregation module. The proposed SATrack enjoys several merits. First, the semantic-aware attention module utilizes language semantics as a bridge to build associations between visual features, enabling stronger visual-semantic capabilities. Second, the cross-modal aggregation module transfers the semantic knowledge of CLIP into the tracking framework for semantic generalization. Extensive experimental results demonstrate that SATrack outperforms previous state-of-the-art trackers on four natural language tracking benchmarks.
Yuyang Tang 0001, Yinchao Ma, Tianzhu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Learning Adaptive Conceptual Prototypes for 3D Single Object Tracking
abstract
3D single object tracking (3D SOT) in LiDAR point clouds plays a crucial role in autonomous driving. It remains a challenging problem due to the incompleteness and the sparsity of points caused by occlusion and limited sensor capabilities. Previous methods design various modules to propagate target perceptual cues to the current frame for target localization. However, perceptual cues may contain less information for occluded or distant objects, which brings great challenges to estimating the target state accurately. To address the above limitations, we propose a novel 3D SOT framework based on the adaptive conceptual prototypes named ACPTrack, which first learns the conceptual prototype from the prior knowledge of the category structure, and then associates weak perceptual cues with the learned conceptual prototypes to improve tracking performance. The proposed ACPTrack enjoys several merits. First, we propose a universal learning method of adaptive conceptual prototype, which can quickly adapt to target-specific structure with given perceptual cues. Second, we design two modules based on the conceptual prototype for structure completion and positioning refinement, which can exploit the rich structure information of the conceptual prototype to deal with sparse and incomplete targets for robust tracking. Third, our framework is generic and compatible with various 3D trackers and brings consistent performance gains. Extensive experiments validate that our method achieves competitive performance on three large-scale datasets.
Yinchao Ma, Wenfei Yang, Tianzhu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Unifying Visual and Vision-Language Tracking via Contrastive Learning
abstract
Single object tracking aims to locate the target object in a video sequence according to the state specified by different modal references, including the initial bounding box (BBOX), natural language (NL), or both (NL+BBOX). Due to the gap between different modalities, most existing trackers are designed for single or partial of these reference settings and overspecialize on the specific modality. Differently, we present a unified tracker called UVLTrack, which can simultaneously handle all three reference settings (BBOX, NL, NL+BBOX) with the same parameters. The proposed UVLTrack enjoys several merits. First, we design a modality-unified feature extractor for joint visual and language feature learning and propose a multi-modal contrastive loss to align the visual and language features into a unified semantic space. Second, a modality-adaptive box head is proposed, which makes full use of the target reference to mine ever-changing scenario features dynamically from video contexts and distinguish the target in a contrastive way, enabling robust performance in different reference settings. Extensive experimental results demonstrate that UVLTrack achieves promising performance on seven visual tracking datasets, three vision-language tracking datasets, and three visual grounding datasets. Codes and models will be open-sourced at https://github.com/OpenSpaceAI/UVLTrack.
Yinchao Ma, Yuyang Tang 0001, Wenfei Yang, Tianzhu Zhang 0001, Mengxue Kang
AAAI1
2023 Foreground-Background Distribution Modeling Transformer for Visual Object Tracking
abstract
Visual object tracking is a fundamental research topic with a broad range of applications. Benefiting from the rapid development of Transformer, pure Transformer trackers have achieved great progress. However, the feature learning of these Transformer-based trackers is easily disturbed by complex backgrounds. To address the above limitations, we propose a novel foreground-background distribution modeling transformer for visual object tracking (F-BDMTrack), including a fore-background agent learning (FBAL) module and a distribution-aware attention (DA2) module in a unified transformer architecture. The proposed F-BDMTrack enjoys several merits. First, the proposed FBAL module can effectively mine fore-background information with designed fore-background agents. Second, the DA2module can suppress the incorrect interaction between foreground and background by modeling fore-background distribution similarities. Finally, F-BDMTrack can extract discriminative features under ever-changing tracking scenarios for more accurate target state estimation. Extensive experiments show that our F-BDMTrack outperforms previous state-of-the-art trackers on eight tracking benchmarks.
Yinchao Ma, Qianjin Yu, Tianzhu Zhang 0001
ICCV3
2023 Adaptive Part Mining for Robust Visual Tracking
abstract
Visual tracking aims to estimate object state in a video sequence, which is challenging when facing drastic appearance changes. Most existing trackers conduct tracking with divided parts to handle appearance variations. However, these trackers commonly divide target objects into regular patches by a hand-designed splitting way, which is too coarse to align object parts well. Besides, a fixed part detector is difficult to partition targets with arbitrary categories and deformations. To address the above issues, we propose a novel adaptive part mining tracker (APMT) for robust tracking via a transformer architecture, including an object representation encoder, an adaptive part mining decoder, and an object state estimation decoder. The proposed APMT enjoys several merits. First, in the object representation encoder, object representation is learned by distinguishing target object from background regions. Second, in the adaptive part mining decoder, we introduce multiple part prototypes to adaptively capture target parts through cross-attention mechanisms for arbitrary categories and deformations. Third, in the object state estimation decoder, we propose two novel strategies to effectively handle appearance variations and distractors. Extensive experimental results demonstrate that our APMT achieves promising results with high FPS. Notably, our tracker is ranked the first place in the VOT-STb2022 challenge.
Yinchao Ma, Tianzhu Zhang 0001, Feng Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 A Keypoint-based Global Association Network for Lane Detection
abstract
Lane detection is a challenging task that requires predicting complex topology shapes of lane lines and distinguishing different types of lanes simultaneously. Earlier works follow a top-down roadmap to regress predefined anchors into various shapes of lane lines, which lacks enough flexibility to fit complex shapes of lanes due to the fixed anchor shapes. Lately, some works propose to formulate lane detection as a keypoint estimation problem to describe the shapes of lane lines more flexibly and gradually group adjacent keypoints belonging to the same lane line in a point-by-point manner, which is inefficient and time-consuming during postprocessing. In this paper, we propose a Global Association Network (GANet) to formulate the lane detection problem from a new perspective, where each keypoint is directly regressed to the starting point of the lane line instead of point-by-point extension. Concretely, the association of keypoints to their belonged lane line is conducted by predicting their offsets to the corresponding starting points of lanes globally without dependence on each other, which could be done in parallel to greatly improve efficiency. In addition, we further propose a Lane-aware Feature Aggregator (LFA), which adaptively captures the local correlations between adjacent keypoints to supplement local information to the global association. Extensive experiments on two popular lane detection benchmarks show that our method outperforms previous methods with F1 score of 79.63% on CULane and 97.71% on Tusimple dataset with high FPS.
Jinsheng Wang, Yinchao Ma, Shaofei Huang 0001, Tianrui Hui, Fei Wang 0032, Tianzhu Zhang 0001
CVPR2
2021 Interpreting and Boosting Dropout from a Game-Theoretic View
Hao Zhang 0063, Yinchao Ma, Yichen Xie 0002, Quanshi Zhang
ICLR3