Mingqiao Ye

dblp:285/9253 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Segmentation and scene understanding · 37% Video understanding and tracking · 20% Image recognition and object detection · 14%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
prompt-based segmentation
1.522025
Stable Segment Anything Model · ICLR 2025
Segment Anything in High Quality · NeurIPS 2023
Machine learning › Trustworthy machine learning
robustness
0.922025
Cascade-DETR: Delving into High-Quality Universal Object Detection · ICCV 2023
Stable Segment Anything Model · ICLR 2025
Computer vision › Video understanding and tracking › video object segmentation
unsupervised video object segmentation
0.912025
EntitySAM: Segment Everything in Video · CVPR 2025
Computer vision › Segmentation and scene understanding › 3d segmentation
3d scene segmentation
0.812024
Gaussian Grouping: Segment and Edit Anything in 3D Scenes · ECCV (29) 2024
Computer vision › 3D vision
3d scene understanding
0.812024
Gaussian Grouping: Segment and Edit Anything in 3D Scenes · ECCV (29) 2024
Machine learning › Transfer learning and domain adaptation
domain generalization
0.712023
Cascade-DETR: Delving into High-Quality Universal Object Detection · ICCV 2023
Computer vision › Segmentation and scene understanding
image segmentation
0.712023
Segment Anything in High Quality · NeurIPS 2023
Computer vision › Image recognition and object detection
object detection
0.712023
Cascade-DETR: Delving into High-Quality Universal Object Detection · ICCV 2023
Computer vision › Image recognition and object detection
object localization
0.712023
Cascade-DETR: Delving into High-Quality Universal Object Detection · ICCV 2023
Computer vision › Segmentation and scene understanding › open-world segmentation
zero-shot segmentation
0.712023
Segment Anything in High Quality · NeurIPS 2023
Computer vision › Video understanding and tracking
object tracking
0.312025
EntitySAM: Segment Everything in Video · CVPR 2025
Natural language and speech › Language models and text generation › prompting
prompt robustness
0.312025
Stable Segment Anything Model · ICLR 2025
Computer vision › 3D vision
3d scene editing
0.212024
Gaussian Grouping: Segment and Edit Anything in 3D Scenes · ECCV (29) 2024

Methods — techniques the papers use, named apart from their topics

semantic encoder · 0.9query-based entity discovery · 0.9dynamic routing · 0.9deformable sampling · 0.9automatic prompt generation · 0.9gaussian grouping · 0.8transformer · 0.7prompt tuning · 0.7iou prediction · 0.7attention mechanism · 0.7
YearPublicationVenuePosition
2025 EntitySAM: Segment Everything in Video
abstract
Automatically tracking and segmenting every video entity remains a significant challenge. Despite rapid advancements in video segmentation, even state-of-the-art models like SAM 2 struggle to consistently track all entities across a video—a task we refer to as Video Entity Segmentation. We propose EntitySAM, a framework for zero-shot video entity segmentation. EntitySAM extends SAM 2 by removing the need for explicit prompts, allowing automatic discovery and tracking of all entities, including those appearing in later frames. We incorporate query-based entity discovery and association into SAM 2, inspired by transformer-based object detectors. Specifically, we introduce an entity decoder to facilitate inter-object communication and an automatic prompt generator using learnable object queries. Additionally, we add a semantic encoder to enhance SAM 2’s semantic awareness, improving segmentation quality. Trained on image-level mask annotations without category information from the COCO dataset, EntitySAM demonstrates strong generalization on four zero-shot video segmentation tasks: Video Entity, Panoptic, Instance, and Semantic Segmentation. Results on six popular benchmarks show that EntitySAM outperforms previous unified video segmentation methods and strong baselines, setting new standards for zero-shot video segmentation. Our code and models are at github.com/ymq2017/entitysam.
Mingqiao Ye, Seoung Wug Oh, Lei Ke, Joon-Young Lee
CVPR1
2025 Stable Segment Anything Model
abstract
The Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to casual prompts, this paper presents the first comprehensive analysis on SAM’s segmentation stability across a diverse spectrum of prompt qualities, notably imprecise bounding boxes and insufficient points. Our key finding reveals that given such low-quality prompts, SAM’s mask decoder tends to activate image features that are biased towards the background or confined to specific object parts. To mitigate this issue, our key idea consists of calibrating solely SAM’s mask attention by adjusting the sampling locations and amplitudes of image features, while the original SAM model architecture and weights remain unchanged. Consequently, our deformable sampling plugin (DSP) enables SAM to adaptively shift attention to the prompted target regions in a data-driven manner. During inference, dynamic routing plugin (DRP) is proposed that toggles SAM between the deformable and regular grid sampling modes, conditioned on the input prompt quality. Thus, our solution, termed Stable-SAM, offers several advantages: 1) improved SAM’s segmentation stability across a wide range of prompt qualities, while 2) retaining SAM’s powerful promptable segmentation efficiency and generality, with 3) minimal learnable parameters (0.08 M) and fast adaptation. Extensive experiments validate the effectiveness and advantages of our approach, underscoring Stable-SAM as a more robust solution for segmenting anything. Codes are at https://github.com/fanq15/Stable-SAM.
Xin Tao 0001, Lei Ke, Mingqiao Ye, Di Zhang 0026, Pengfei Wan 0001, Yu-Wing Tai, Chi-Keung Tang
ICLR4
2024 Gaussian Grouping: Segment and Edit Anything in 3D Scenes
Mingqiao Ye, Martin Danelljan, Fisher Yu 0001, Lei Ke
ECCV (29)1
2023 Cascade-DETR: Delving into High-Quality Universal Object Detection
abstract
Object localization in general environments is a fundamental part of vision systems. While dominating on the COCO benchmark, recent Transformer-based detection methods are not competitive in diverse domains. Moreover, these methods still struggle to very accurately estimate the object bounding boxes in complex environments.We introduce Cascade-DETR for high-quality universal object detection. We jointly tackle the generalization to diverse domains and localization accuracy by proposing the Cascade Attention layer, which explicitly integrates object-centric information into the detection decoder by limiting the attention to the previous box prediction. To further enhance accuracy, we also revisit the scoring of queries. Instead of relying on classification scores, we predict the expected IoU of the query, leading to substantially more well-calibrated confidences. Lastly, we introduce a universal object detection benchmark, UDB10, that contains 10 datasets from diverse domains. While also advancing the state-of-the-art on COCO, Cascade-DETR substantially improves DETR-based detectors on all datasets in UDB10, even by over 10 mAP in some cases. The improvements under stringent quality requirements are even more pronounced. Our code and pretrained models are at https://github.com/SysCV/cascade-detr.
Mingqiao Ye, Lei Ke, Siyuan Li 0008, Yu-Wing Tai, Chi-Keung Tang, Martin Danelljan, Fisher Yu 0001
ICCV1
2023 Segment Anything in High Quality
abstract
The recent Segment Anything Model (SAM) represents a big leap in scaling up segmentation models, allowing for powerful zero-shot capabilities and flexible prompting. Despite being trained with 1.1 billion masks, SAM's mask prediction quality falls short in many cases, particularly when dealing with objects that have intricate structures. We propose HQ-SAM, equipping SAM with the ability to accurately segment any object, while maintaining SAM's original promptable design, efficiency, and zero-shot generalizability. Our careful design reuses and preserves the pre-trained model weights of SAM, while only introducing minimal additional parameters and computation. We design a learnable High-Quality Output Token, which is injected into SAM's mask decoder and is responsible for predicting the high-quality mask. Instead of only applying it on mask-decoder features, we first fuse them with early and final ViT features for improved mask details. To train our introduced learnable parameters, we compose a dataset of 44K fine-grained masks from several sources. HQ-SAM is only trained on the introduced detaset of 44k masks, which takes only 4 hours on 8 GPUs. We show the efficacy of HQ-SAM in a suite of 10 diverse segmentation datasets across different downstream tasks, where 8 out of them are evaluated in a zero-shot transfer protocol. Our code and pretrained models are at https://github.com/SysCV/SAM-HQ.
Lei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu 0001, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001
NeurIPS2
2020 Wireless D2D Network Link Scheduling based on Graph Embedding
abstract
Wireless link scheduling in D2D communication systems aims at maximizing the weighted sum rate of D2D pairs by determining which subset of D2D pairs should be activated. However, it is a non-convex combinatorial optimization problem, which is generally NP-hard and difficult to achieve the optimal solution. Inspired by the recent attempt of introducing machine learning and graph embedding to reach the general goal, we propose an efficient method to solve the weighted sum rate maximization problem. We first model the system as a one-nearest neighbor graph, in which each D2D pair is a node and the strongest interference link for each node is an edge. Then we compute the feature vectors of both weights and nodes by graph embedding and use the feature vectors as the input of a subsequent multi-layer classifier. The parameters of classifier and graph embedding are trained jointly in a supervised manner. Simulation result shows that the proposed method can obtain near-optimal performance with only hundreds of training samples and is capable to be generalized to more complicated scenarios.
Jingyun Fu, Mingqiao Ye, Mengyuan Lee, Guanding Yu
VTC Fall3