EDBT 2026 Demo / reviewers in the wild / expert
Huaizhe Xu
dblp:322/0982
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0003-1416-9360ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Segmentation and scene understanding · 52% Image recognition and object detection · 26% Vision and language · 12% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
panoptic segmentation |
1.3 | 2 | 2023 | MP-Former: Mask-Piloted Transformer for Image Segmentation · CVPR 2023 Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation · CVPR 2023 |
Computer vision › Vision and language
multimodal fusion |
0.8 | 1 | 2024 | Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation · ECCV (81) 2024 |
Computer vision › Image recognition and object detection › object detection › open-world object detection
open-set object detection |
0.8 | 1 | 2024 | Visual in-Context Prompting · CVPR 2024 |
Computer vision › Segmentation and scene understanding › open-world segmentation
open-set segmentation |
0.8 | 1 | 2024 | Visual in-Context Prompting · CVPR 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation · ECCV (81) 2024 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.7 | 1 | 2023 | MP-Former: Mask-Piloted Transformer for Image Segmentation · CVPR 2023 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.7 | 1 | 2023 | Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation · CVPR 2023 |
Computer vision › Image recognition and object detection › object detection › multi-task detection
joint detection and segmentation |
0.7 | 1 | 2023 | Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation · CVPR 2023 |
Computer vision › Image recognition and object detection
object detection |
0.7 | 1 | 2023 | Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation · CVPR 2023 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.7 | 1 | 2023 | Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation · CVPR 2023 |
Computer vision › Vision and language
visual prompting |
0.2 | 1 | 2024 | Visual in-Context Prompting · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
prompt encoder · 0.8joint training · 0.8encoder-decoder architecture · 0.8transformer · 0.7query embedding · 0.7masked-attention transformer · 0.7mask-piloted training · 0.7mask prediction · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Visual in-Context PromptingabstractIn-context prompting in large language models (LLMs) has become a prevalent approach to improve zero-shot capabilities, but this idea is less explored in the vision domain. Existing visual prompting methods focus on referring segmentation to segment the most relevant object, falling short of addressing many generic vision tasks like open-set segmentation and detection. In this paper, we introduce a universal visual in-context prompting framework for both tasks, as shown in Fig. 1. In particular, we build on top of an encoder-decoder architecture, and develop a versatile prompt encoder to support a variety of prompts like strokes, boxes, and points. We further enhance it to take an arbitrary number of reference image segments as the context. Our extensive explorations show that the proposed visual in-context prompting elicits extraordinary referring and generic segmentation capabilities to refer and detect, yielding competitive performance to close-set in-domain datasets and showing promising results on many open-set segmentation datasets. By joint training on COCO and SA-1B, DINOv achieves 57.7 PQ on COCO and 23.2 PQ on ADE20K. Code will be available at https://github.com/UX-Decoder/DINOv Feng Li 0040, Hao Zhang 0097, Tianhe Ren, Shilong Liu 0004, Xueyan Zou, Huaizhe Xu, Hongyang Li 0003, Chunyuan Li, Lei Zhang 0001, Jianfeng Gao 0001 |
CVPR | 7 |
| 2024 | Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
Shaozhe Hao, Bojia Zi, Huaizhe Xu, Kwan-Yee Kenneth Wong |
ECCV (81) | 4 |
| 2023 | Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and SegmentationabstractIn this paper we present Mask DINO, a unified object detection and segmentation framework. Mask DINO extends DINO (DETR with Improved Denoising Anchor Boxes) by adding a mask prediction branch which supports all image segmentation tasks (instance, panoptic, and semantic). It makes use of the query embeddings from DINO to dot-product a high-resolution pixel embedding map to predict a set of binary masks. Some key components in DINO are extended for segmentation through a shared architecture and training process. Mask DINO is simple, efficient, and scalable, and it can benefit from joint large-scale detection and segmentation datasets. Our experiments show that Mask DINO significantly outperforms all existing specialized segmentation methods, both on a ResNet-50 backbone and a pre-trained model with SwinL backbone. Notably, Mask DINO establishes the best results to date on instance segmentation (54.5 AP on COCO), panoptic segmentation (59.4 PQ on COCO), and semantic segmentation (60.8 mIoU on ADE20K) among models under one billion parameters. Code is available at https://github.com/IDEA-Research/MaskDINO. Feng Li 0040, Hao Zhang 0097, Huaizhe Xu, Shilong Liu 0004, Lei Zhang 0001, Lionel M. Ni, Harry Shum |
CVPR | 3 |
| 2023 | MP-Former: Mask-Piloted Transformer for Image SegmentationabstractWe present a mask-piloted Transformer which improves masked-attention in Mask2Former for image segmentation. The improvement is based on our observation that Mask2Former suffers from inconsistent mask predictions between consecutive decoder layers, which leads to inconsistent optimization goals and low utilization of decoder queries. To address this problem, we propose a mask-piloted training approach, which additionally feeds noised ground-truth masks in masked-attention and trains the model to reconstruct the original ones. Compared with the predicted masks used in mask-attention, the ground-truth masks serve as a pilot and effectively alleviate the negative impact of inaccurate mask predictions in Mask2Former. Based on this technique, our MP-Former achieves a remarkable performance improvement on all three image segmentation tasks (instance, panoptic, and semantic), yielding +2.3AP and +1.6mIoU on the Cityscapes instance and semantic segmentation tasks with a ResNet-50 backbone. Our method also significantly speeds up the training, outperforming Mask2Former with half of the number of training epochs on ADE20K with both a ResNet-50 and a Swin-L backbones. Moreover, our method only introduces little computation during training and no extra computation during inference. Our code will be released at https://github.com/IDEA-Research/MP-Former. Hao Zhang 0097, Feng Li 0040, Huaizhe Xu, Shijia Huang, Shilong Liu 0004, Lionel M. Ni, Lei Zhang 0001 |
CVPR | 3 |