Mingyu Cao

dblp:165/7683 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0005-2440-9956ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 AGFT-Tracker: Adaptive Game-Based PEFT for Object Tracking with PLMs
abstract
The rise of pre-trained large models (PLMs) has sparked interest in vision tasks like object tracking. However, as PLMs scale, fully fine-tuning all parameters becomes impractical, highlighting the need for parameter-efficient fine-tuning (PEFT). While adapter tuning, which adds tunable parameters to Multi-Head Attention (MHA) or Feed-Forward Networks (FFN), is common, critical parameters like Layer Normalization (LN), vital for stability and convergence, are often overlooked. Furthermore, traditional fine-tuning strategies fail to differentiate module importance, limiting performance improvements. To solve these issues, we propose a new PEFT method for unlocking large model potential in object tracking: Adaptive Game-Based Fine-tuning Tracker (AGFT-Tracker). AGFT-Tracker combines adapter tuning with direct LN fine-tuning and adaptively allocates parameter budgets based on tracking attention losses. Important sensitive modules use higher-rank LoRA and frozen LN, while stable modules undergo lower-rank LoRA and LN adjustments. This approach improves effectiveness and efficiency, achieving state-of-the-art results on challenging benchmarks.
Mingyu Cao, Xihuai He, Xueqiong Li, Kedi Zhang, Yuhua Tang, Wanrong Huang, Huibin Tan
ICME1
2025 Adaptive Distribution-Aware Modeling for Transformer Tracking
abstract
Adapting to changes in data distribution is a major challenge in visual object tracking. In Transformer-based tracking, Layer Normalization (LN) is often applied uniformly to both template and search features, limiting feature diversity. Additionally, models tend to converge to trivial solutions, and tracking samples are sensitive to distribution shifts, affecting robustness. To address these issues, we propose the Adaptive Distribution-Aware Transformer Tracker (ADAT), incorporating three key components: the Target-Aware Module (TAM), the Region-Aware Module (RAM), and the Self-Feedback-Aware Module (SFAM). TAM normalizes template and search features separately, preserving flexibility and enhancing target learning. RAM refines target perception by distinguishing between near and far target regions. SFAM filters out noisy samples and fine-tunes normalization parameters through self-feedback. While TAM and RAM regulate feature-level distribution, SFAM adjusts at the sample level. Extensive experiments show that ADAT outperforms existing methods, achieving superior performance on challenging benchmarks.
Mingyu Cao, Huibin Tan, Xueqiong Li, Wanrong Huang, Kedi Zhang, Yuhua Tang, Shaowu Yang
ICME1
2025 Wave-wise Discriminative Tracking by Phase-Amplitude Separation, Augmentation and Mixture
abstract
Distinguishing key features in complex visual tasks is challenging. A novel approach treats image patches (tokens) as waves. By using both phase and amplitude, it captures richer semantics and specific invariances compared to pixel-based methods, and allows for feature fusion across regions for a holistic image representation. Based on this, we propose the Wave-wise Discriminative Transformer Tracker (WDT). During tracking, WDT represents features via phase-amplitude separation, enhancement, and mixture. First, we designed a Mutual Exclusive Phase-Amplitude Extractor (MEPAE) to separate phase and amplitude features with distinct semantics, representing spatial target info and background brightness respectively. Then, Wave-wise Feature Augmentation is carried out with two submodules: Phase-Amplitude Feature Augmentation and Mixture. The augmentation module disrupts the separated features in the same batch, and the mixture module recombines them to generate positive and negative waves. The original features are aggregated into the original wave. Positive waves have the same phase but different amplitudes, and negative waves have different phase components. Finally, self-supervised and tracking-supervised losses guide the global and local representation learning for original, positive, and negative waves, enhancing wave-level discrimination. Experiments on five benchmarks prove the effectiveness of our method.
Huibin Tan, Mingyu Cao, Xihuai He, Hao Li 0025, Long Lan, Mengzhu Wang
IJCAI2
2024 Sequential Fusion Based Multi-Granularity Consistency for Space-Time Transformer Tracking
abstract
Regarded as a template-matching task for a long time, visual object tracking has witnessed significant progress in space-wise exploration. However, since tracking is performed on videos with substantial time-wise information, it is important to simultaneously mine the temporal contexts which have not yet been deeply explored. Previous supervised works mostly consider template reform as the breakthrough point, but they are often limited by additional computational burdens or the quality of chosen templates. To address this issue, we propose a Space-Time Consistent Transformer Tracker (STCFormer), which uses a sequential fusion framework with multi-granularity consistency constraints to learn spatiotemporal context information. We design a sequential fusion framework that recombines template and search images based on tracking results from chronological frames, fusing updated tracking states in training. To further overcome the over-reliance on the fixed template without increasing computational complexity, we design three space-time consistent constraints: Label Consistency Loss (LCL) for label-level consistency, Attention Consistency Loss (ACL) for patch-level ROI consistency, and Semantic Consistency Loss (SCL) for feature-level semantic consistency. Specifically, in ACL and SCL, the label information is used to constrain the attention and feature consistency of the target and the background, respectively, to avoid mutual interference. Extensive experiments have shown that our STCFormer outperforms many of the best-performing trackers on several popular benchmarks.
Wenjing Yang 0002, Wanrong Huang, Xianchen Zhou, Mingyu Cao, Huibin Tan
AAAI5
2023 Enhanced Dcf Tracker Regularized by Reliable Sample Construction
abstract
Discriminative correlation filter (DCF) is a highly efficient tracking technique using the circulant shifted samples of search images to update the template, so the reliability of input samples determines template quality. In this paper, we rethink the reliability problem of input samples in advance during template updating and propose an enhanced DCF tracking method regularized by a novel sparse representation based reliable sample construction term, called enhanced sparse correlation filter (ESCF). Specifically, the reconstructed reliable samples are the sparse representation of circulant shifted samples of unfiltered input samples, in which the target will approach the center to preserve target visual cues into the template when using the cosine window. Besides, we jointly perform template learning and reliable sample construction into a unified learning paradigm to benefit from each other, which further can be carried out in the frequency domain without incurring excessive time cost by skillful decomposition. Experiments on several popular visual tracking datasets verify the efficacy of ESCF and show that ESCF performs favorably against several well-established representative counterparts.
Mingyu Cao, Mengzhu Wang, Long Lan, Wenjing Yang 0002, Huibin Tan
ICASSP2
2023 Progressive Perception Learning for Distribution Modulation in Siamese Tracking
abstract
We explore an innovative view on distribution modulation to boost Siamese trackers. Specially, we observed two cases of possible distribution inconsistency in Siamese tracking: 1) Two branches with different sizes may be in different distribution ranges after a shared backbone (including BN layers). 2) The background data may affect the total feature distribution of the search branch. To address these issues, we proposed a plug-and-play component named Progressive Perception Learning Module (P2LM) to modulate the distribution using three feature normalization blocks successively, i.e., Self-Aware Block (SAB), Target-Aware Block (TAB), and Region-Aware Block (RAB). SAB regulates the distribution of each branch independently for the first issue. TAB uses the target information to guide the distribution adjustments of the two branches. RAB divides the search image into foreground and background with a region mask and normalizes them separately to filter the background distractors for robust tracking. TAB and RAB synergistically alleviate the distribution shifts caused by environmental variance. Experiments on OTB100, UAV123, LaSOT, and GOT-10k verify the compelling effects of our module.
Xianchen Zhou, Mingyu Cao, Mengzhu Wang, Guangjie Gao, Wenjing Yang 0002, Huibin Tan
ICASSP3
2023 Decomposition, Interaction, Reconstruction Meets Global Context Learning In Visual Tracking
abstract
Tensor decomposition and reconstruction attention is a promising global context learning approach because it can remain efficient while avoiding feature compression. To exploit its potential even further in visual tracking, we redesign a 3D tensor modeling paradigm, namely tensor Decomposition, Interaction, Reconstruction attention (DIR), respectively corresponding to three function components, Tensor Decomposition Module (TDM), Tensor Interaction Module (TIM) and Context Reconstruction Module (CRM). Specifically, TDM decomposes a 3D tensor feature into rank-1 context fragments in different dimension views. The ingenuity here lies in the introduction of Circular Convolution for processing features at arbitrary scales and channel-sharing segments to enhance the interaction of the two branches in the Siamese network architecture. TIM obtains the tensor planes of each dimension by the Cross-Similarity operation of rank-1 tensors and fused cubic features, which brings more interactions between all feature dimensions. CRM reconstructs 3D context representations with the outputs of the above modules. In experiments, DIR is embedded into the tracker to verify its effectiveness.
Huibin Tan, Mingyu Cao, Mengzhu Wang, Wenjing Yang 0002
ICASSP3
2021 Fixed-time Bearing-based Distributed Network Localization
abstract
This paper studies the fixed-time distributed localization problem for directed network based on bearing measurements. The orientation of the global coordinate system is not available to nodes whose local coordinate systems do not to be aligned. First, the barycentric coordinate representation is obtained relying only on the bearing information. Then, a fixed-time distributed network localization algorithm is proposed. By applying the proposed algorithm, the localization problem is converted to the consensus tracking problem for localization errors. When nodes distribution and communication topology meet the requirements, one can prove that the position estimates can convert to the truth after a fixed time. Finally, the simulation verifies the validity of the algorithm.
Mingyu Cao, Hao Zhang 0008, Zhuping Wang, Changzhu Zhang, Chao Huang 0018
SMC1
2020 A neural network-based joint learning approach for biomedical entity and relation extraction from biomedical literature
Ling Luo 0001, Mingyu Cao, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin
J. Biomed. Informatics3
2015 Interferometric Phase Denoising by Median Patch-Based Locally Optimal Wiener Filter
abstract
This letter presents a new filtering technique for interferometric synthetic aperture radar (InSAR) phase images. Traditional local denoising algorithms all suffer from the drawback of removing texture detail information. In contrast, several nonlocal methods have attained good performance in InSAR applications. However, these methods only take radiometric similarity. The patch-based locally optimal Wiener filter (PLOW) utilizes both geometrically and radiometrically similar patch information by clustering analysis and nonlocal filtering; thus, it can better balance between detail preservation and denoising. Nevertheless, PLOW itself is not fitted for InSAR. In this letter, we modify and improve the original algorithm while considering the coherence coefficient and the characteristics of InSAR. This new method provides better filtering results, but with a higher computational cost. Moreover, we introduce the box dimension in fractal geometry as a new index. Experiments on simulated and real data demonstrate that this algorithm outperforms traditional denoising methods.
Mingyu Cao, Shiqiang Li, Robert Wang 0001, Ning Li 0002
IEEE Geosci. Remote. Sens. Lett.1