EDBT 2026 Demo / reviewers in the wild / expert
Mingqiao Ye
dblp:285/9253
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Segmentation and scene understanding · 37% Video understanding and tracking · 20% Image recognition and object detection · 14% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
prompt-based segmentation |
1.5 | 2 | 2025 | Stable Segment Anything Model · ICLR 2025 Segment Anything in High Quality · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 2 | 2025 | Cascade-DETR: Delving into High-Quality Universal Object Detection · ICCV 2023 Stable Segment Anything Model · ICLR 2025 |
Computer vision › Video understanding and tracking › video object segmentation
unsupervised video object segmentation |
0.9 | 1 | 2025 | EntitySAM: Segment Everything in Video · CVPR 2025 |
Computer vision › Segmentation and scene understanding › 3d segmentation
3d scene segmentation |
0.8 | 1 | 2024 | Gaussian Grouping: Segment and Edit Anything in 3D Scenes · ECCV (29) 2024 |
Computer vision › 3D vision
3d scene understanding |
0.8 | 1 | 2024 | Gaussian Grouping: Segment and Edit Anything in 3D Scenes · ECCV (29) 2024 |
Machine learning › Transfer learning and domain adaptation
domain generalization |
0.7 | 1 | 2023 | Cascade-DETR: Delving into High-Quality Universal Object Detection · ICCV 2023 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.7 | 1 | 2023 | Segment Anything in High Quality · NeurIPS 2023 |
Computer vision › Image recognition and object detection
object detection |
0.7 | 1 | 2023 | Cascade-DETR: Delving into High-Quality Universal Object Detection · ICCV 2023 |
Computer vision › Image recognition and object detection
object localization |
0.7 | 1 | 2023 | Cascade-DETR: Delving into High-Quality Universal Object Detection · ICCV 2023 |
Computer vision › Segmentation and scene understanding › open-world segmentation
zero-shot segmentation |
0.7 | 1 | 2023 | Segment Anything in High Quality · NeurIPS 2023 |
Computer vision › Video understanding and tracking
object tracking |
0.3 | 1 | 2025 | EntitySAM: Segment Everything in Video · CVPR 2025 |
Natural language and speech › Language models and text generation › prompting
prompt robustness |
0.3 | 1 | 2025 | Stable Segment Anything Model · ICLR 2025 |
Computer vision › 3D vision
3d scene editing |
0.2 | 1 | 2024 | Gaussian Grouping: Segment and Edit Anything in 3D Scenes · ECCV (29) 2024 |
Methods — techniques the papers use, named apart from their topics
semantic encoder · 0.9query-based entity discovery · 0.9dynamic routing · 0.9deformable sampling · 0.9automatic prompt generation · 0.9gaussian grouping · 0.8transformer · 0.7prompt tuning · 0.7iou prediction · 0.7attention mechanism · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EntitySAM: Segment Everything in VideoabstractAutomatically tracking and segmenting every video entity remains a significant challenge. Despite rapid advancements in video segmentation, even state-of-the-art models like SAM 2 struggle to consistently track all entities across a video—a task we refer to as Video Entity Segmentation. We propose EntitySAM, a framework for zero-shot video entity segmentation. EntitySAM extends SAM 2 by removing the need for explicit prompts, allowing automatic discovery and tracking of all entities, including those appearing in later frames. We incorporate query-based entity discovery and association into SAM 2, inspired by transformer-based object detectors. Specifically, we introduce an entity decoder to facilitate inter-object communication and an automatic prompt generator using learnable object queries. Additionally, we add a semantic encoder to enhance SAM 2’s semantic awareness, improving segmentation quality. Trained on image-level mask annotations without category information from the COCO dataset, EntitySAM demonstrates strong generalization on four zero-shot video segmentation tasks: Video Entity, Panoptic, Instance, and Semantic Segmentation. Results on six popular benchmarks show that EntitySAM outperforms previous unified video segmentation methods and strong baselines, setting new standards for zero-shot video segmentation. Our code and models are at github.com/ymq2017/entitysam. Mingqiao Ye, Seoung Wug Oh, Lei Ke, Joon-Young Lee |
CVPR | 1 |
| 2025 | Stable Segment Anything ModelabstractThe Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to casual prompts, this paper presents the first comprehensive analysis on SAM’s segmentation stability across a diverse spectrum of prompt qualities, notably imprecise bounding boxes and insufficient points. Our key finding reveals that given such low-quality prompts, SAM’s mask decoder tends to activate image features that are biased towards the background or confined to specific object parts. To mitigate this issue, our key idea consists of calibrating solely SAM’s mask attention by adjusting the sampling locations and amplitudes of image features, while the original SAM model architecture and weights remain unchanged. Consequently, our deformable sampling plugin (DSP) enables SAM to adaptively shift attention to the prompted target regions in a data-driven manner. During inference, dynamic routing plugin (DRP) is proposed that toggles SAM between the deformable and regular grid sampling modes, conditioned on the input prompt quality. Thus, our solution, termed Stable-SAM, offers several advantages: 1) improved SAM’s segmentation stability across a wide range of prompt qualities, while 2) retaining SAM’s powerful promptable segmentation efficiency and generality, with 3) minimal learnable parameters (0.08 M) and fast adaptation. Extensive experiments validate the effectiveness and advantages of our approach, underscoring Stable-SAM as a more robust solution for segmenting anything. Codes are at https://github.com/fanq15/Stable-SAM. Xin Tao 0001, Lei Ke, Mingqiao Ye, Di Zhang 0026, Pengfei Wan 0001, Yu-Wing Tai, Chi-Keung Tang |
ICLR | 4 |
| 2024 | Gaussian Grouping: Segment and Edit Anything in 3D Scenes
Mingqiao Ye, Martin Danelljan, Fisher Yu 0001, Lei Ke |
ECCV (29) | 1 |
| 2023 | Cascade-DETR: Delving into High-Quality Universal Object DetectionabstractObject localization in general environments is a fundamental part of vision systems. While dominating on the COCO benchmark, recent Transformer-based detection methods are not competitive in diverse domains. Moreover, these methods still struggle to very accurately estimate the object bounding boxes in complex environments.We introduce Cascade-DETR for high-quality universal object detection. We jointly tackle the generalization to diverse domains and localization accuracy by proposing the Cascade Attention layer, which explicitly integrates object-centric information into the detection decoder by limiting the attention to the previous box prediction. To further enhance accuracy, we also revisit the scoring of queries. Instead of relying on classification scores, we predict the expected IoU of the query, leading to substantially more well-calibrated confidences. Lastly, we introduce a universal object detection benchmark, UDB10, that contains 10 datasets from diverse domains. While also advancing the state-of-the-art on COCO, Cascade-DETR substantially improves DETR-based detectors on all datasets in UDB10, even by over 10 mAP in some cases. The improvements under stringent quality requirements are even more pronounced. Our code and pretrained models are at https://github.com/SysCV/cascade-detr. Mingqiao Ye, Lei Ke, Siyuan Li 0008, Yu-Wing Tai, Chi-Keung Tang, Martin Danelljan, Fisher Yu 0001 |
ICCV | 1 |
| 2023 | Segment Anything in High QualityabstractThe recent Segment Anything Model (SAM) represents a big leap in scaling up segmentation models, allowing for powerful zero-shot capabilities and flexible prompting. Despite being trained with 1.1 billion masks, SAM's mask prediction quality falls short in many cases, particularly when dealing with objects that have intricate structures. We propose HQ-SAM, equipping SAM with the ability to accurately segment any object, while maintaining SAM's original promptable design, efficiency, and zero-shot generalizability. Our careful design reuses and preserves the pre-trained model weights of SAM, while only introducing minimal additional parameters and computation. We design a learnable High-Quality Output Token, which is injected into SAM's mask decoder and is responsible for predicting the high-quality mask. Instead of only applying it on mask-decoder features, we first fuse them with early and final ViT features for improved mask details. To train our introduced learnable parameters, we compose a dataset of 44K fine-grained masks from several sources. HQ-SAM is only trained on the introduced detaset of 44k masks, which takes only 4 hours on 8 GPUs. We show the efficacy of HQ-SAM in a suite of 10 diverse segmentation datasets across different downstream tasks, where 8 out of them are evaluated in a zero-shot transfer protocol. Our code and pretrained models are at https://github.com/SysCV/SAM-HQ. Lei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu 0001, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001 |
NeurIPS | 2 |
| 2020 | Wireless D2D Network Link Scheduling based on Graph EmbeddingabstractWireless link scheduling in D2D communication systems aims at maximizing the weighted sum rate of D2D pairs by determining which subset of D2D pairs should be activated. However, it is a non-convex combinatorial optimization problem, which is generally NP-hard and difficult to achieve the optimal solution. Inspired by the recent attempt of introducing machine learning and graph embedding to reach the general goal, we propose an efficient method to solve the weighted sum rate maximization problem. We first model the system as a one-nearest neighbor graph, in which each D2D pair is a node and the strongest interference link for each node is an edge. Then we compute the feature vectors of both weights and nodes by graph embedding and use the feature vectors as the input of a subsequent multi-layer classifier. The parameters of classifier and graph embedding are trained jointly in a supervised manner. Simulation result shows that the proposed method can obtain near-optimal performance with only hundreds of training samples and is capable to be generalized to more complicated scenarios. Jingyun Fu, Mingqiao Ye, Mengyuan Lee, Guanding Yu |
VTC Fall | 3 |