Manqi Zhao

dblp:53/1377 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Theory of computation · 2
YearPublicationVenuePosition
2026 SAM2MOT: A Novel Paradigm of Multi-Object Tracking by Segmentation
abstract
Inspired by Segment Anything 2, which generalizes segmentation from images to videos, we propose SAM2MOT—a novel segmentation-driven paradigm for multi-object tracking that breaks away from the conventional detection-association framework. In contrast to previous approaches that treat segmentation as auxiliary information, SAM2MOT places it at the heart of the tracking process, systematically tackling challenges like false positives and occlusions. Its effectiveness has been thoroughly validated on major MOT benchmarks. Furthermore, SAM2MOT integrates pre-trained detector, pre-trained segmentor with tracking logic into a zero-shot MOT system that requires no fine-tuning. This significantly reduces dependence on labeled data and paves the way for transitioning MOT research from task-specific solutions to general-purpose systems. Experiments on DanceTrack, UAVDT, and BDD100K show state-of-the-art results. Notably, SAM2MOT outperforms existing methods on DanceTrack by +2.1 HOTA and +4.5 IDF1, highlighting its effectiveness in MOT.
Manqi Zhao, Dongsheng Jiang
AAAI3
2025 Parameter-Efficient Reparameterization Tuning for Remote Sensing Image-Text Retrieval
abstract
Vision language models have been gradually adapted to various tasks of remote sensing domain with full fine-tuning paradigm,e.g., remote sensing image-text retrieval (RSITR). Superior performance enhancements have proven the powerful generalization and robustness of vision language models. However, full fine-tuning vision language models is resource-intensive and poses risks of overfitting. Moreover, existing RSITR methods usually assume that remote sensing images correspond to text captions one by one and utilize the bidirectional matching training objective, which is not aligned with evaluation benchmarks and real-world applications. To tackle the mentioned problems, we propose a novel Parameter-Efficient Reparameterization Tuning with Ranking and Matching (PERT-RaMa) framework, which effectively migrates the vision language model (i.e.CLIP) to RSITR task. To overcome the overfitting issue, we build a lightweight, plug-and-play module called Kronecker product for low-rank adaptation (KPLoRA). KPLoRA obtains higher intrinsic rank with fewer parameters. Furthermore, we design the Ranking and Matching (RaMa) training method that converts RSITR task into one-to-one matching and one-to-many ranking, which is aligned with current RSITR benchmarks and accelerates training speed through removal of unessential computations. Comprehensive experiments on three public RSITR benchmarks demonstrate that the effectiveness and efficiency of the proposed retrieval model. Our method outperforms full fine-tuning methods without CLIP by nearly 5-10%, and achieves comparable or superior retrieval capability than CLIP with full fine-tuning and parameter-efficient fine-tuning. Furthermore, our RaMa training method increases the training speed by 5x compared to current training method.
Shengyang Li, Manqi Zhao
IEEE Trans. Geosci. Remote. Sens.3
2024 Video process detection for space electrostatic suspension material experiment in China's Space Station
Manqi Zhao, Shengyang Li
Eng. Appl. Artif. Intell.3
2024 MP2Net: Mask Propagation and Motion Prediction Network for Multiobject Tracking in Satellite Videos
abstract
Mainstream multi-object tracking (MOT) algorithms employ global object detection and association methods. However, when dealing with scenarios involving crowded tiny objects in satellite videos, existing global trackers often yield numerous missed detections and unstable trajectories. To address this issue, we propose a novel joint-detection-and-tracking framework, MP2Net, which integrates local detection enhancements for tiny targets and bridges the gap between detection and association. Specifically, our approach incorporates a mask propagation network that enhances feature representation for tiny targets by matching frame-by-frame to capture local details. Additionally, we utilize an implicit and explicit motion prediction strategy that merges tracking information into detection at both feature and instance levels, thereby improving tracking robustness. Experimental results on two large-scale datasets demonstrate the effectiveness and robustness of MP2Net, achieving state-of-the-art performance on typical moving objects in satellite videos, such as 66.7% MOTA and 75.9% IDF1 on the SatVideoDT challenge dataset. The code will be available at https://github.com/DonDominic/MP2Net.
Manqi Zhao, Shengyang Li, Han Wang 0049, Yuhan Sun 0004, Yanfeng Gu
IEEE Trans. Geosci. Remote. Sens.1
2024 Transformer Tracking for Satellite Video: Matching, Propagation, and Prediction
abstract
Recently, transformer-based trackers have brought overwhelming advantages in general video. However, their performance in satellite video has been hindered by insufficient satellite-specific training and a lack of designs tailored to satellite targets and scene characteristics. To tackle these challenges, we propose a novel transformer-based tracking framework for satellite video object tracking: Transformer Matching, Propagation, and Prediction (TransMPP). TransMPP combines three stages: static matching, dynamic propagation, and prediction, to ensure accurate tracking in satellite videos. Specifically, the Matching model uses a one-stream pipeline for simultaneous feature extraction and relationship modeling across extensive search and template areas, thereby improving foreground and background discrimination capabilities. In addition, the Propagation and Prediction models enhance temporal modeling capabilities through local long-term and short-term feature propagation and global sequence prediction, respectively, boosting tracking robustness. Moreover, to ensure a fair comparison and evaluation, we also developed SatSOT-train, a large-scale training dataset for the SatSOT benchmark. After comprehensive training, TransMPP demonstrates state-of-the-art (SOTA) performance on the SatSOT dataset, achieving an area under the curve (AUC) score of 59.9% and a precision score of 71.5%, bringing improvements of 6.3% and 5.3%, respectively. The code will be available athttps://github.com/DonDominic/TransMPP.
Manqi Zhao, Shengyang Li
IEEE Trans. Geosci. Remote. Sens.1
2023 Frequency and Spatial Domain Filter Network for Visual Object Tracking
Manqi Zhao, Shenyang Li, Han Wang 0049
PRCV (6)1
2023 Siamese Graph Attention Networks for robust visual object tracking
abstract
Siamese-based trackers usually convert the object tracking task into a similarity matching problem between the target template and the search region. Since fixed or manually updated templates are not robust when tracking moving objects with dramatically changing appearance, this paper proposes an improved siamese graph attention network with adaptive template update called SiamGT. By establishing spatiotemporal and context dependencies between historical images and search regions, a frame selection mechanism is added to improve the richness of information. In addition, a graph attention network with residual connections is used in the template update mechanism which enables the propagation and aggregation of information to generate robust templates. Extensive experimental results on challenging benchmarks such as UAV123, OTB100, and VOT2019 demonstrate that the proposed SiamGT has achieved state-of-the-art performance in visual object tracking.
Shengyang Li, Weilong Guo, Manqi Zhao, Yunfei Liu 0003
Comput. Vis. Image Underst.4
2023 A Multitask Benchmark Dataset for Satellite Video: Object Detection, Tracking, and Segmentation
abstract
Video satellites can continuously image large areas and provide dynamic, real-time monitoring of hotspots and objects. The intelligent processing and analysis of satellite video have become a research hotspot in the field of remote sensing. However, the lack of high-quality satellite video datasets limits the development of relevant object detection, object tracking, and object segmentation. In this paper, we build the largest scale satellite video dataset with the most task types supported and object categories, named Satellite Video Multi-Mission Benchmark (SAT-MTB). First, multi-task annotation of aircraft, ships, cars, trains, and their corresponding 14 categories of fine-grained objects in 249 satellite videos is performed based on horizontal bounding boxes (HBB), oriented bounding boxes (OBB), masks, which cover more than 50,000 frames and 1,033,511 annotated object instances. Then, we review the tasks of object detection, object tracking, and object segmentation based on satellite videos, providing a comprehensive overview of progress in related datasets and algorithm research. Finally, we establish the first public benchmark of multi-task algorithms for satellite video object detection, object tracking, and object segmentation, evaluating and analyzing the performance of a total of 47 representative algorithms under different tasks on the constructed dataset. The proposed SAT-MTB will significantly advance research in intelligent processing and analysis of satellite video and related applications.
Shengyang Li, Manqi Zhao, Weilong Guo, Yixuan Lv, Longxuan Kou, Han Wang 0049, Yanfeng Gu
IEEE Trans. Geosci. Remote. Sens.3
2022 The First Challenge on Moving Object Detection and Tracking in Satellite Videos: Methods and Results
abstract
In this paper, we briefly summarize the first challenge on moving object detection and tracking in satellite videos (SatVideoDT). This challenge has three tracks related to satellite video analysis, including moving object detection (Track 1), single object tracking (Track 2), and multiple-object tracking (Track 3). 123, 89, and 70 participants successfully registered, while 37, 42, and 29 teams submitted their final results on the test datasets for Tracks 1-3, respectively. The top-performing methods and their results in each track are described with details. This challenge establishes a new benchmark for satellite video analysis.
Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Chenguang Dai, Weilong Guo, Xiyu Qi, Kelong Tu, Shudan Zhu, Lai Chen, Bin Lin 0013, Chaocan Xue, Jinlei Zheng, Limei Qin, Ying Li 0017, Manqi Zhao, Lu Ruan 0003, Mingpeng Cui, Guanchen Ding, Guangwei Jiang, Zhenzhong Chen 0001, Kaiyang Cao, Lingyu Kong, Shaodong Chen, Zhicheng Zhao 0001, Qin Shen, Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003
ICPR22
2022 Cooperation of Boundary Attention and Negative Matrix L1 Regularization Loss Function for Polyp Segmentation
abstract
Most colorectal cancers are caused by colorectal adenomatous polyps, early screening for colonic polyps is of great clinical importance. However, polyps have diverse appearance, such as shape and size; polyps are difficult to distinguish from mucosa. To address these problems, this paper proposes a boundary attention and negative matrix L1 regularization loss function synergistic segmentation method. Firstly, a branch of boundary attention is added to the last detail finding branch of PraNet [16]. Secondly, the inner polyp region of the boundary is complemented using negative matrix L1 regularization iterations. The synergy of the two terms can refine object areas, which will improve the segmentation accuracy. We conducted extensive experimental evaluations on four publicly available datasets, results show that our method is superior to other models. Especially on CVC-ClinicDB dataset, compared with PraNet, Dice is improved by 4.3% and IoU is improved by 5.9%.
Guoqi Liu, Manqi Zhao
ICPR2
2022 SatSOT: A Benchmark Dataset for Satellite Video Single Object Tracking
abstract
By imaging a specific area continuously, satellite video shows excellent capability in various applications such as surveillance and traffic management. Although object tracking has made significant progress in recent years, development in satellite object tracking is limited by the lack of open-source satellite datasets. It is thus essential to establish a satellite video object-tracking benchmark to fill the gap and advance the research. In this work, we present SatSOT, the first densely annotated satellite video single object-tracking benchmark dataset. SatSOT consists of 105 sequences with 27664 frames, 11 attributes, and four categories of typical moving targets in satellite videos: car, plane, ship, and train. Based on the proposed dataset and the significant challenges in satellite video object tracking, such as small targets, background interference, and severe occlusion, detailed evaluation and analysis are performed on 15 among the best and most representative tracking algorithms, which provides a basis for further research on satellite video object tracking.
Manqi Zhao, Shengyang Li, Shiyu Xuan, Longxuan Kou, Shuai Gong
IEEE Trans. Geosci. Remote. Sens.1
2011 Thresholded Basis Pursuit: LP Algorithm for Order-Wise Optimal Support Recovery for Sparse and Approximately Sparse Signals From Noisy Random Measurements
abstract
In this paper, we present a linear programming solution for sign pattern recovery of a sparse signal, x, from noisy random projections of the signal. We consider two types of noise models: input noise, where noise enters before the random projection, and output noise, where noise enters after the random projection. Sign pattern recovery involves the estimation of sign pattern of a sparse signal. Our idea is to pretend that no noise exists and solve the noiseless ℓ1problem, namely, min ||β||1s.t. y - Gβ and quantizing the resulting solution. We show that the quantized solution perfectly reconstructs the sign pattern of a sufficiently sparse signal. Specifically, we show that the sign pattern of an arbitrary k-sparse, n-dimensional signal x can be recovered with SNR - Ω(log n) and measurements scaling as m = Ώ,(log n /k) for all sparsity levels k satisfying 01problem, in that, we estimate the maximum admissible noise level before sign pattern recovery fails.
Venkatesh Saligrama, Manqi Zhao
IEEE Trans. Inf. Theory2
2010 On compressed blind de-convolution of filtered sparse processes
abstract
Suppose the signal x ∈ 葷nis realized by driving a k-sparse signal z ∈ 葷nthrough an arbitrary unknown stable discrete-linear time invariant system H, namely, x(t) = (h * z)(t), where h(·) is the impulse response of the operator H. Is x(·) compressible in the conventional sense of compressed sensing? Namely, can x(t) be reconstructed from small set of measurements obtained through suitable random projections? For the case when the unknown system H is auto-regressive (i.e. all pole) of a known order it turns out that x can indeed be reconstructed from O(k log(n)) measurements. We develop a novel LP optimization algorithm and show that both the unknown filter H and the sparse input z can be reliably estimated.
Manqi Zhao, Venkatesh Saligrama
ICASSP1
2010 Information theoretic bounds for compressed sensing
abstract
In this paper, we derive information theoretic performance bounds to sensing and reconstruction of sparse phenomena from noisy projections. We consider two settings: output noise models where the noise enters after the projection and input noise models where the noise enters before the projection. We consider two types of distortion for reconstruction: support errors and mean-squared errors. Our goal is to relate the number of measurements, m , and SNR, to signal sparsity, k, distortion level, d, and signal dimension, n . We consider support errors in a worst-case setting. We employ different variations of Fano's inequality to derive necessary conditions on the number of measurements and SNR required for exact reconstruction. To derive sufficient conditions, we develop new insights on max-likelihood analysis based on a novel superposition property. In particular, this property implies that small support errors are the dominant error events. Consequently, our ML analysis does not suffer the conservatism of the union bound and leads to a tighter analysis of max-likelihood. These results provide order-wise tight bounds. For output noise models, we show that asymptotically an SNR of ((n)) together with (k (n/k)) measurements is necessary and sufficient for exact support recovery. Furthermore, if a small fraction of support errors can be tolerated, a constant SNR turns out to be sufficient in the linear sparsity regime. In contrast for input noise models, we show that support recovery fails if the number of measurements scales as o(n(n)/SNR), implying poor compression performance for such cases. Motivated by the fact that the worst-case setup requires significantly high SNR and substantial number of measurements for input and output noise models, we consider a Bayesian setup. To derive necessary conditions, we develop novel extensions to Fano's inequality to handle continuous domains and arbitrary distortions. We then develop a new max-likelihood analysis over the set of rate distortion quantization points to characterize tradeoffs between mean-squared distortion and the number of measurements using rate-distortion theory. We show that with constant SNR the number of measurements scales linearly with the rate-distortion function of the sparse phenomena.
Shuchin Aeron, Venkatesh Saligrama, Manqi Zhao
IEEE Trans. Inf. Theory3
2009 Anomaly Detection with Score functions based on Nearest Neighbor Graphs
abstract
We propose a novel non-parametric adaptive anomaly detection algorithm for high dimensional data based on score functions derived from nearest neighbor graphs on n-point nominal data. Anomalies are declared whenever the score of a test sample falls below q, which is supposed to be the desired false alarm level. The resulting anomaly detector is shown to be asymptotically optimal in that it is uniformly most powerful for the specified false alarm level, q, for the case when the anomaly density is a mixture of the nominal and a known density. Our algorithm is computationally efficient, being linear in dimension and quadratic in data size. It does not require choosing complicated tuning parameters or function approximation classes and it can adapt to local structure such as local change in dimensionality. We demonstrate the algorithm on both artificial and real data sets in high dimensional feature spaces.
Manqi Zhao, Venkatesh Saligrama
NIPS1