Feng Li 0040

dblp:92/2954-40 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
27since 2021 · last 2026
0000-0001-7853-0448ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 7 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 16 since 2021
YearPublicationVenuePosition
2026 T-Rex2++: Toward Generic Object Perception via Text-Visual Prompt Synergy
abstract
We present T-Rex2++, a unified and highly practical framework for generic open-set object perception, encompassing both object detection and instance segmentation. Previous methods relying on text prompts effectively encapsulate the abstract concept of common objects, but struggle with rare or complex object representation due to data scarcity and descriptive limitations. Conversely, visual prompts excel in depicting novel objects through concrete visual examples, but fall short in conveying the abstract concept of objects as effectively as text prompts. Recognizing these complementary strengths, we introduce a text-visual synergy mechanism that aligns both modalities within a single feature space via contrastive learning. Crucially, T-Rex2++ advances beyond the passive perception paradigm of its predecessor by introducing a novel Universal Prompt. This learnable component models generic objectness, empowering the system to autonomously discover and localize arbitrary objects without any user-provided cues, thereby closing the loop between human-guided interaction and fully automatic perception. Furthermore, we extend the synergy verification to the pixel level by integrating a zero-shot instance segmentation module, demonstrating that our contrastive alignment generalizes robustly to fine-grained masks. Comprehensive experiments demonstrate that T-Rex2++ exhibits strong zero-shot object perception capabilities across a wide spectrum of scenarios, validating T-Rex2++ as a versatile foundation for generic object perception.
Feng Li 0040, Zhaoyang Zeng, Tianhe Ren, Shilong Liu 0004, Lei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
abstract
Visual instruction tuning has made considerable strides in enhancing the capabilities of Large Multimodal Models (LMMs). However, existing open LMMs largely focus on single-image tasks, their applications to multi-image scenarios remains less explored. Additionally, prior LMM research separately tackles different scenarios, leaving it impossible to generalize cross scenarios with new emerging capabilities. To this end, we introduce LLaVA-Interleave, which simultaneously tackles Multi-image, Multi-frame (video), Multi-view (3D), and Multi-patch (single-image) scenarios in LMMs. To enable these capabilities, we regard the interleaved data format as a general template and compile the M4-Instruct dataset with 1,177.6k samples, spanning 4 primary domains with 14 tasks and 41 datasets. We also curate the LLaVA-Interleave Bench to comprehensively evaluate the multi-image performance of LMMs. Through extensive experiments, LLaVA-Interleave achieves leading results in multi-image, video, and 3D benchmarks, while maintaining the performance of single-image tasks. Besides, our model also exhibits several emerging capabilities, e.g., transferring tasks across different settings and modalities.
Feng Li 0040, Renrui Zhang, Hao Zhang 0097, Yuanhan Zhang, Bo Li 0080, Wei Li 0119, Zejun Ma 0001, Chunyuan Li
ICLR1
2025 A Mutual Supervision Framework for Referring Expression Segmentation and Generation
Shijia Huang, Feng Li 0040, Hao Zhang 0097, Shilong Liu 0004, Lei Zhang 0001, Liwei Wang 0009
Int. J. Comput. Vis.2
2025 ED-Pose++: Enhanced Explicit Box Detection for Conventional and Interactive Multi-Object Keypoint Detection
abstract
Detecting keypoints on diverse objects is essential for fine-grained visual understanding and analysis. This paper introduces Enhanced Explicit Box Detection (ED-Pose++), an end-to-end framework that leverages cascade box regression to realize both conventional and interactive multi-object keypoint detection. Unlike traditional one-stage methods, ED-Pose++ innovatively redefines multi-object keypoint detection as a dual-phase explicit box detection, achieving a unified representation and regression optimization process. Specifically, an object detection decoder first extracts each object's position and global features, establishing a good initialization for subsequent keypoint detection. To bring in contextual information near keypoints, we also regard each keypoint as a small box to learn both positions and their related local contents. In practice, an object-to-keypoint detection decoder adopts a collaborative learning strategy between object and keypoint features, facilitating efficient information propagation between global and local perspectives. Rooted on the architecture, we further equip dual-phase box detection with an interactive mechanism that enables the model to refine its predictions based on limited user feedback. During training, we incorporate an error correction scheme to equip the model with an adept self-correction capability for use during inference. The comprehensive experiments demonstrate ED-Pose++'s superior performance in conventional multi-object keypoint detection tasks. For the first time, ED-Pose++ outperforms heatmap-based top-down approaches across various benchmarks, despite operating within a fully end-to-end architecture. The interactive variant also dramatically reduces more than 10 times the labeling effort of 2D keypoint annotation compared with manual-only annotation.
Ailing Zeng, Tianhe Ren, Shilong Liu 0004, Feng Li 0040, Ruimao Zhang, Lei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Visual in-Context Prompting
abstract
In-context prompting in large language models (LLMs) has become a prevalent approach to improve zero-shot capabilities, but this idea is less explored in the vision domain. Existing visual prompting methods focus on referring segmentation to segment the most relevant object, falling short of addressing many generic vision tasks like open-set segmentation and detection. In this paper, we introduce a universal visual in-context prompting framework for both tasks, as shown in Fig. 1. In particular, we build on top of an encoder-decoder architecture, and develop a versatile prompt encoder to support a variety of prompts like strokes, boxes, and points. We further enhance it to take an arbitrary number of reference image segments as the context. Our extensive explorations show that the proposed visual in-context prompting elicits extraordinary referring and generic segmentation capabilities to refer and detect, yielding competitive performance to close-set in-domain datasets and showing promising results on many open-set segmentation datasets. By joint training on COCO and SA-1B, DINOv achieves 57.7 PQ on COCO and 23.2 PQ on ADE20K. Code will be available at https://github.com/UX-Decoder/DINOv
Feng Li 0040, Hao Zhang 0097, Tianhe Ren, Shilong Liu 0004, Xueyan Zou, Huaizhe Xu, Hongyang Li 0003, Chunyuan Li, Lei Zhang 0001, Jianfeng Gao 0001
CVPR1
2024 T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
Feng Li 0040, Zhaoyang Zeng, Tianhe Ren, Shilong Liu 0004, Lei Zhang 0001
ECCV (33)2
2024 TAPTR: Tracking Any Point with Transformers as Detection
Hongyang Li 0003, Hao Zhang 0097, Shilong Liu 0004, Zhaoyang Zeng, Tianhe Ren, Feng Li 0040, Lei Zhang 0001
ECCV (16)6
2024 Segment and Recognize Anything at Any Granularity
Feng Li 0040, Hao Zhang 0097, Peize Sun, Xueyan Zou, Shilong Liu 0004, Chunyuan Li, Lei Zhang 0001, Jianfeng Gao 0001
ECCV (48)1
2024 LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Shilong Liu 0004, Hao Cheng 0002, Hao Zhang 0097, Feng Li 0040, Tianhe Ren, Xueyan Zou, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001, Jianfeng Gao 0001, Chunyuan Li
ECCV (47)5
2024 Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection
Shilong Liu 0004, Zhaoyang Zeng, Tianhe Ren, Feng Li 0040, Hao Zhang 0097, Chunyuan Li, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001
ECCV (47)4
2024 LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
Hao Zhang 0097, Hongyang Li 0003, Feng Li 0040, Tianhe Ren, Xueyan Zou, Shilong Liu 0004, Shijia Huang, Jianfeng Gao 0001, Leizhang, Chunyuan Li, Jainwei Yang
ECCV (43)3
2024 TAPTRv2: Attention-based Position Update Improves Tracking Any Point
abstract
In this paper, we present TAPTRv2, a Transformer-based approach built upon TAPTR for solving the Tracking Any Point (TAP) task. TAPTR borrows designs from DEtection TRansformer (DETR) and formulates each tracking point as a point query, making it possible to leverage well-studied operations in DETR-like algorithms. TAPTRv2 improves TAPTR by addressing a critical issue regarding its reliance on cost-volume, which contaminates the point query’s content feature and negatively impacts both visibility prediction and cost-volume computation. In TAPTRv2, we propose a novel attention-based position update (APU) operation and use key-aware deformable attention to realize. For each query, this operation uses key-aware attention weights to combine their corresponding deformable sampling positions to predict a new query position. This design is based on the observation that local attention is essentially the same as cost-volume, both of which are computed by dot-production between a query and its surrounding features. By introducing this new operation, TAPTRv2 not only removes the extra burden of cost-volume computation, but also leads to a substantial performance improvement. TAPTRv2 surpasses TAPTR and achieves state-of-the-art performance on many challenging datasets, demonstrating the effectiveness of our approach.
Hongyang Li 0003, Hao Zhang 0097, Shilong Liu 0004, Zhaoyang Zeng, Feng Li 0040, Tianhe Ren, Lei Zhang 0006
NeurIPS5
2024 Interfacing Foundation Models' Embeddings
abstract
Foundation models possess strong capabilities in reasoning and memorizing across modalities. To further unleash the power of foundation models, we present FIND, a generalized interface for aligning foundation models' embeddings with unified image and dataset-level understanding spanning modality and granularity. As shown in Fig.1, a lightweight transformer interface without tuning any foundation model weights is enough for segmentation, grounding, and retrieval in an interleaved manner. The proposed interface has the following favorable attributes: (1) Generalizable. It applies to various tasks spanning retrieval, segmentation, etc., under the same architecture and weights. (2) Interleavable. With the benefit of multi-task multi-modal training, the proposed interface creates an interleaved shared embedding space. (3) Extendable. The proposed interface is adaptive to new tasks, and new models. In light of the interleaved embedding space, we introduce FIND-Bench, which introduces new training and evaluation annotations to the COCO dataset for interleaved segmentation and retrieval. We are the first work aligning foundations models' embeddings for interleave understanding. Meanwhile, our approach achieves state-of-the-art performance on FIND-Bench and competitive performance on standard retrieval and segmentation settings.
Xueyan Zou, Mingyu Ding, Zhengyuan Yang, Feng Li 0040, Hao Zhang 0097, Shilong Liu 0004, Arul Aravinthan, Yong Jae Lee
NeurIPS8
2024 DN-DETR: Accelerate DETR Training by Introducing Query DeNoising
abstract
We present in this paper a novel denoising training method to speed up DETR (DEtection TRansformer) training and offer a deepened understanding of the slow convergence issue of DETR-like methods. We show that the slow convergence results from the instability of bipartite graph matching which causes inconsistent optimization goals in early training stages. To address this issue, except for the Hungarian loss, our method additionally feeds GT bounding boxes with noises into the Transformer decoder and trains the model to reconstruct the original boxes, which effectively reduces the bipartite graph matching difficulty and leads to faster convergence. Our method is universal and can be easily plugged into any DETR-like method by adding dozens of lines of code to achieve a remarkable improvement. As a result, our DN-DETR results in a remarkable improvement ( +1.9AP) under the same setting and achieves 46.0 AP and 49.5 AP trained for 12 and 50 epochs with the ResNet-50 backbone. Compared with the baseline under the same setting, DN-DETR achieves comparable performance with 50% training epochs. We also demonstrate the effectiveness of denoising training in CNN-based detectors (Faster R-CNN), segmentation models (Mask2Former, Mask DINO), and more DETR-based models (DETR, Anchor DETR, Deformable DETR).
Feng Li 0040, Hao Zhang 0097, Shilong Liu 0004, Jian Guo 0016, Lionel M. Ni, Lei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding
abstract
In this paper, we study the problem of visual grounding by considering both phrase extraction and grounding (PEG). In contrast to the previous phrase-known-at-test setting, PEG requires a model to extract phrases from text and locate objects from image simultaneously, which is a more practical setting in real applications. As phrase extraction can be regarded as a 1D text segmentation problem, we formulate PEG as a dual detection problem and propose a novel DQ-DETR model, which introduces dual queries to probe different features from image and text for object prediction and phrase mask prediction. Each pair of dual queries are designed to have shared positional parts but different content parts. Such a design effectively alleviates the difficulty of modality alignment between image and text (in contrast to a single query design) and empowers Transformer decoder to leverage phrase mask-guided attention to improve the performance. To evaluate the performance of PEG, we also propose a new metric CMAP (cross-modal average precision), analogous to the AP metric in object detection. The new metric overcomes the ambiguity of Recall@1 in many-box-to-one-phrase cases in phrase grounding. As a result, our PEG pre-trained DQ-DETR establishes new state-of-the-art results on all visual grounding benchmarks with a ResNet-101 backbone. For example, it achieves 91.04% and 83.51% in terms of recall rate on RefCOCO testA and testB with a ResNet-101 backbone.
Shilong Liu 0004, Shijia Huang, Feng Li 0040, Hao Zhang 0097, Yaoyuan Liang, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001
AAAI3
2023 Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation
abstract
In this paper we present Mask DINO, a unified object detection and segmentation framework. Mask DINO extends DINO (DETR with Improved Denoising Anchor Boxes) by adding a mask prediction branch which supports all image segmentation tasks (instance, panoptic, and semantic). It makes use of the query embeddings from DINO to dot-product a high-resolution pixel embedding map to predict a set of binary masks. Some key components in DINO are extended for segmentation through a shared architecture and training process. Mask DINO is simple, efficient, and scalable, and it can benefit from joint large-scale detection and segmentation datasets. Our experiments show that Mask DINO significantly outperforms all existing specialized segmentation methods, both on a ResNet-50 backbone and a pre-trained model with SwinL backbone. Notably, Mask DINO establishes the best results to date on instance segmentation (54.5 AP on COCO), panoptic segmentation (59.4 PQ on COCO), and semantic segmentation (60.8 mIoU on ADE20K) among models under one billion parameters. Code is available at https://github.com/IDEA-Research/MaskDINO.
Feng Li 0040, Hao Zhang 0097, Huaizhe Xu, Shilong Liu 0004, Lei Zhang 0001, Lionel M. Ni, Harry Shum
CVPR1
2023 Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR
abstract
Recent DEtection TRansformer-based (DETR) models have obtained remarkable performance. Its success cannot be achieved without the re-introduction of multi-scale feature fusion in the encoder. However, the excessively increased tokens in multi-scale features, especially for about 75% of low-level features, are quite computationally inefficient, which hinders real applications of DETR models. In this paper, we present Lite DETR, a simple yet efficient end-to-end object detection framework that can effectively reduce the GFLOPs of the detection head by 60% while keeping 99% of the original performance. Specifically, we design an efficient encoder block to update high-level features (corresponding to small-resolution feature maps) and low-level features (corresponding to large-resolution feature maps) in an interleaved way. In addition, to better fuse cross-scale features, we develop a key-aware deformable attention to predict more reliable attention weights. Comprehensive experiments validate the effectiveness and efficiency of the proposed Lite DETR, and the efficient encoder strategy can generalize well across existing DETR-based models. The code will be available in https://github.com/IDEA-Research/Lite-DETR.
Feng Li 0040, Ailing Zeng, Shilong Liu 0004, Hao Zhang 0097, Hongyang Li 0003, Lei Zhang 0001, Lionel M. Ni
CVPR1
2023 MP-Former: Mask-Piloted Transformer for Image Segmentation
abstract
We present a mask-piloted Transformer which improves masked-attention in Mask2Former for image segmentation. The improvement is based on our observation that Mask2Former suffers from inconsistent mask predictions between consecutive decoder layers, which leads to inconsistent optimization goals and low utilization of decoder queries. To address this problem, we propose a mask-piloted training approach, which additionally feeds noised ground-truth masks in masked-attention and trains the model to reconstruct the original ones. Compared with the predicted masks used in mask-attention, the ground-truth masks serve as a pilot and effectively alleviate the negative impact of inaccurate mask predictions in Mask2Former. Based on this technique, our MP-Former achieves a remarkable performance improvement on all three image segmentation tasks (instance, panoptic, and semantic), yielding +2.3AP and +1.6mIoU on the Cityscapes instance and semantic segmentation tasks with a ResNet-50 backbone. Our method also significantly speeds up the training, outperforming Mask2Former with half of the number of training epochs on ADE20K with both a ResNet-50 and a Swin-L backbones. Moreover, our method only introduces little computation during training and no extra computation during inference. Our code will be released at https://github.com/IDEA-Research/MP-Former.
Hao Zhang 0097, Feng Li 0040, Huaizhe Xu, Shijia Huang, Shilong Liu 0004, Lionel M. Ni, Lei Zhang 0001
CVPR2
2023 DFA3D: 3D Deformable Attention For 2D-to-3D Feature Lifting
abstract
In this paper, we propose a new operator, called 3D DeFormable Attention (DFA3D), for 2D-to-3D feature lifting, which transforms multi-view 2D image features into a unified 3D space for 3D object detection. Existing feature lifting approaches, such as Lift-Splat-based and 2D attention-based, either use estimated depth to get pseudo LiDAR features and then splat them to a 3D space, which is a one-pass operation without feature refinement, or ignore depth and lift features by 2D attention mechanisms, which achieve finer semantics while suffering from a depth ambiguity problem. In contrast, our DFA3D-based method first leverages the estimated depth to expand each view’s 2D feature map to 3D and then utilizes DFA3D to aggregate features from the expanded 3D feature maps. With the help of DFA3D, the depth ambiguity problem can be effectively alleviated from the root, and the lifted features can be progressively refined layer by layer, thanks to the Transformerlike architecture. In addition, we propose a mathematically equivalent implementation of DFA3D which can significantly improve its memory efficiency and computational speed. We integrate DFA3D into several methods that use 2D attention-based feature lifting with only a few modifications in code and evaluate on the nuScenes dataset. The experiment results show a consistent improvement of +1.41% mAP on average, and up to +15.1% mAP improvement when high-quality depth information is available, demonstrating the superiority, applicability, and huge potential of DFA3D. The code is available at https://github.com/IDEAResearch/3D-deformable-attention.git.
Hongyang Li 0003, Hao Zhang 0097, Zhaoyang Zeng, Shilong Liu 0004, Feng Li 0040, Tianhe Ren, Lei Zhang 0001
ICCV5
2023 Detection Transformer with Stable Matching
abstract
This paper is concerned with the matching stability problem across different decoder layers in DEtection TRansformers (DETR). We point out that the unstable matching in DETR is caused by a multi-optimization path problem, which is highlighted by the one-to-one matching design in DETR. To address this problem, we show that the most important design is to use and only use positional metrics (like IOU) to supervise classification scores of positive examples. Under the principle, we propose two simple yet effective modifications by integrating positional metrics to DETR’s classification loss and matching cost, named position-supervised loss and position-modulated cost. We verify our methods on several DETR variants. Our methods show consistent improvements over baselines. By integrating our methods with DINO, we achieve 50.4 and 51.5 AP on the COCO detection benchmark using ResNet-50 backbones under 1× (12 epochs) and 2× (24 epochs) training settings, achieving a new record under the same setting. We achieve 63.8 AP on COCO detection test-dev with a Swin-Large backbone. Our code will be made available at https://github.com/IDEA-Research/Stable-DINO.
Shilong Liu 0004, Tianhe Ren, Zhaoyang Zeng, Hao Zhang 0097, Feng Li 0040, Hongyang Li 0003, Jun Huang 0007, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001
ICCV6
2023 Neural Interactive Keypoint Detection
abstract
This work proposes an end-to-end neural interactive keypoint detection framework named Click-Pose, which can significantly reduce more than 10 times labeling costs of 2D keypoint annotation compared with manual-only annotation. Click-Pose explores how user feedback can cooperate with a neural keypoint detector to correct the predicted keypoints in an interactive way for a faster and more effective annotation process. Specifically, we design the pose error modeling strategy that inputs the ground truth pose combined with four typical pose errors into the decoder and trains the model to reconstruct the correct poses, which enhances the self-correction ability of the model. Then, we attach an interactive human-feedback loop that allows receiving users’ clicks to correct one or several predicted keypoints and iteratively utilizes the decoder to update all other keypoints with a minimum number of clicks (NoC) for efficient annotation. We validate Click-Pose in in-domain, out-of-domain scenes, and a new task of keypoint adaptation. For annotation, Click-Pose only needs 1.97 and 6.45 NoC@95 (at precision 95%) on COCO and Human-Art, reducing 31.4% and 36.3% efforts than the SOTA model (ViTPose) with manual correction, respectively. Besides, without user clicks, Click-Pose surpasses the previous end-to-end model by 1.4 AP on COCO and 3.0 AP on Human-Art.
Ailing Zeng, Feng Li 0040, Shilong Liu 0004, Ruimao Zhang, Lei Zhang 0001
ICCV3
2023 A Simple Framework for Open-Vocabulary Segmentation and Detection
abstract
We present OpenSeeD, a simple Open-vocabulary Segmentation and Detection framework that jointly learns from different segmentation and detection datasets. To bridge the gap of vocabulary and annotation granularity, we first introduce a pre-trained text encoder to encode all the visual concepts in two tasks and learn a common semantic space for them. This gives us reasonably good results compared with the counterparts trained on segmentation task only. To further reconcile them, we identify two discrepancies: i) task discrepancy – segmentation requires extracting masks for both foreground objects and background stuff, while detection merely cares about the former; ii) data discrepancy – box and mask annotations are with different spatial granularity, and thus not directly interchangeable. To address these issues, we propose a decoupled decoding to reduce the interference between foreground/background and a conditioned mask decoding to assist in generating masks for given boxes. To this end, we develop a simple encoder-decoder model encompassing all three techniques and train it jointly on COCO and Objects365. After pre-training, our model exhibits competitive or stronger zero-shot transferability for both segmentation and detection. Specifically, OpenSeeD beats the state-of-the-art method for open-vocabulary instance and panoptic segmentation across 5 datasets, and outperforms previous work for open-vocabulary detection on LVIS and ODinW under similar settings. When transferred to specific tasks, our model achieves new SoTA for panoptic segmentation on COCO and ADE20K, and instance segmentation on ADE20K and Cityscapes (The bottom row in Fig. 1 shows a comparison of the performance of OpenSeeD and previous SoTA methods). Finally, we note that OpenSeeD is the first to explore the potential of joint training on segmentation and detection, and hope it can be received as a strong baseline for developing a single model for both tasks in the open world. Code will be released at https://github.com/IDEA-Research/OpenSeeD.
Hao Zhang 0097, Feng Li 0040, Xueyan Zou, Shilong Liu 0004, Chunyuan Li, Lei Zhang 0001
ICCV2
2023 DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Hao Zhang 0097, Feng Li 0040, Shilong Liu 0004, Lei Zhang 0001, Hang Su 0006, Jun Zhu 0001, Lionel M. Ni, Harry Shum
ICLR2
2023 Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation
Ailing Zeng, Shilong Liu 0004, Feng Li 0040, Ruimao Zhang, Lei Zhang 0001
ICLR4
2023 Segment Everything Everywhere All at Once
abstract
In this work, we present SEEM, a promotable and interactive model for segmenting everything everywhere all at once in an image. In SEEM, we propose a novel and versatile decoding mechanism that enables diverse prompting for all types of segmentation tasks, aiming at a universal interface that behaves like large language models (LLMs). More specifically, SEEM is designed with four desiderata: i) Versatility. We introduce a new visual prompt to unify different spatial queries including points, boxes, scribbles, and masks, which can further generalize to a different referring image; ii) Compositionality. We learn a joint visual-semantic space between text and visual prompts, which facilitates the dynamic composition of two prompt types required for various segmentation tasks, as shown in Fig. 1; iii) Interactivity. We further incorporate learnable memory prompts into the decoder to retain segmentation history through mask-guided cross-attention from the decoder to image features; iv) Semantic awareness. We use a text encoder to encode text queries and mask labels into the same semantic space for open-vocabulary segmentation. We conduct a comprehensive empirical study to validate the effectiveness of SEEM across diverse segmentation tasks. The results demonstrate that SEEM exhibits robust generalizing to unseen user intents as it learns to compose prompts of different types in a unified representation space. Our approach achieves competitive performance on interactive segmentation, generic segmentation, referring segmentation, and video object segmentation on 9 datasets with minimum 1/100 supervision in a single set of weights.
Xueyan Zou, Hao Zhang 0097, Feng Li 0040, Jianfeng Gao 0001, Yong Jae Lee
NeurIPS4
2022 DN-DETR: Accelerate DETR Training by Introducing Query DeNoising
abstract
We present in this paper a novel denoising training method to speedup DETR (DEtection TRansformer) training and offer a deepened understanding of the slow convergence issue of DETR-like methods. We show that the slow convergence results from the instability of bipartite graph matching which causes inconsistent optimization goals in early training stages. To address this issue, except for the Hungarian loss, our method additionally feeds ground-truth bounding boxes with noises into Transformer decoder and trains the model to reconstruct the original boxes, which effectively reduces the bipartite graph matching difficulty and leads to a faster convergence. Our method is universal and can be easily plugged into any DETR-like methods by adding dozens of lines of code to achieve a remarkable improvement. As a result, our DN-DETR results in a remarkable improvement (+1.9AP) under the same setting and achieves the best result (AP 43.4 and 48.6 with 12 and 50 epochs of training respectively) among DETR-like methods with ResNet-50 backbone. Compared with the baseline under the same setting, DN-DETR achieves comparable performance with 50% training epochs. Code is available at https://github.com/FengLi-ust/DN-DETR.
Feng Li 0040, Hao Zhang 0097, Shilong Liu 0004, Jian Guo 0016, Lionel M. Ni, Lei Zhang 0001
CVPR1
2022 DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
Shilong Liu 0004, Feng Li 0040, Hao Zhang 0097, Xiao Yang 0028, Xianbiao Qi, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001
ICLR2