Huaizhe Xu

dblp:322/0982 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0003-1416-9360ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Segmentation and scene understanding · 52% Image recognition and object detection · 26% Vision and language · 12%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
panoptic segmentation
1.322023
MP-Former: Mask-Piloted Transformer for Image Segmentation · CVPR 2023
Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation · CVPR 2023
Computer vision › Vision and language
multimodal fusion
0.812024
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation · ECCV (81) 2024
Computer vision › Image recognition and object detection › object detection › open-world object detection
open-set object detection
0.812024
Visual in-Context Prompting · CVPR 2024
Computer vision › Segmentation and scene understanding › open-world segmentation
open-set segmentation
0.812024
Visual in-Context Prompting · CVPR 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation · ECCV (81) 2024
Computer vision › Segmentation and scene understanding
image segmentation
0.712023
MP-Former: Mask-Piloted Transformer for Image Segmentation · CVPR 2023
Computer vision › Segmentation and scene understanding
instance segmentation
0.712023
Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation · CVPR 2023
Computer vision › Image recognition and object detection › object detection › multi-task detection
joint detection and segmentation
0.712023
Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation · CVPR 2023
Computer vision › Image recognition and object detection
object detection
0.712023
Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation · CVPR 2023
Computer vision › Segmentation and scene understanding
semantic segmentation
0.712023
Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation · CVPR 2023
Computer vision › Vision and language
visual prompting
0.212024
Visual in-Context Prompting · CVPR 2024

Methods — techniques the papers use, named apart from their topics

prompt encoder · 0.8joint training · 0.8encoder-decoder architecture · 0.8transformer · 0.7query embedding · 0.7masked-attention transformer · 0.7mask-piloted training · 0.7mask prediction · 0.7
YearPublicationVenuePosition
2024 Visual in-Context Prompting
abstract
In-context prompting in large language models (LLMs) has become a prevalent approach to improve zero-shot capabilities, but this idea is less explored in the vision domain. Existing visual prompting methods focus on referring segmentation to segment the most relevant object, falling short of addressing many generic vision tasks like open-set segmentation and detection. In this paper, we introduce a universal visual in-context prompting framework for both tasks, as shown in Fig. 1. In particular, we build on top of an encoder-decoder architecture, and develop a versatile prompt encoder to support a variety of prompts like strokes, boxes, and points. We further enhance it to take an arbitrary number of reference image segments as the context. Our extensive explorations show that the proposed visual in-context prompting elicits extraordinary referring and generic segmentation capabilities to refer and detect, yielding competitive performance to close-set in-domain datasets and showing promising results on many open-set segmentation datasets. By joint training on COCO and SA-1B, DINOv achieves 57.7 PQ on COCO and 23.2 PQ on ADE20K. Code will be available at https://github.com/UX-Decoder/DINOv
Feng Li 0040, Hao Zhang 0097, Tianhe Ren, Shilong Liu 0004, Xueyan Zou, Huaizhe Xu, Hongyang Li 0003, Chunyuan Li, Lei Zhang 0001, Jianfeng Gao 0001
CVPR7
2024 Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
Shaozhe Hao, Bojia Zi, Huaizhe Xu, Kwan-Yee Kenneth Wong
ECCV (81)4
2023 Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation
abstract
In this paper we present Mask DINO, a unified object detection and segmentation framework. Mask DINO extends DINO (DETR with Improved Denoising Anchor Boxes) by adding a mask prediction branch which supports all image segmentation tasks (instance, panoptic, and semantic). It makes use of the query embeddings from DINO to dot-product a high-resolution pixel embedding map to predict a set of binary masks. Some key components in DINO are extended for segmentation through a shared architecture and training process. Mask DINO is simple, efficient, and scalable, and it can benefit from joint large-scale detection and segmentation datasets. Our experiments show that Mask DINO significantly outperforms all existing specialized segmentation methods, both on a ResNet-50 backbone and a pre-trained model with SwinL backbone. Notably, Mask DINO establishes the best results to date on instance segmentation (54.5 AP on COCO), panoptic segmentation (59.4 PQ on COCO), and semantic segmentation (60.8 mIoU on ADE20K) among models under one billion parameters. Code is available at https://github.com/IDEA-Research/MaskDINO.
Feng Li 0040, Hao Zhang 0097, Huaizhe Xu, Shilong Liu 0004, Lei Zhang 0001, Lionel M. Ni, Harry Shum
CVPR3
2023 MP-Former: Mask-Piloted Transformer for Image Segmentation
abstract
We present a mask-piloted Transformer which improves masked-attention in Mask2Former for image segmentation. The improvement is based on our observation that Mask2Former suffers from inconsistent mask predictions between consecutive decoder layers, which leads to inconsistent optimization goals and low utilization of decoder queries. To address this problem, we propose a mask-piloted training approach, which additionally feeds noised ground-truth masks in masked-attention and trains the model to reconstruct the original ones. Compared with the predicted masks used in mask-attention, the ground-truth masks serve as a pilot and effectively alleviate the negative impact of inaccurate mask predictions in Mask2Former. Based on this technique, our MP-Former achieves a remarkable performance improvement on all three image segmentation tasks (instance, panoptic, and semantic), yielding +2.3AP and +1.6mIoU on the Cityscapes instance and semantic segmentation tasks with a ResNet-50 backbone. Our method also significantly speeds up the training, outperforming Mask2Former with half of the number of training epochs on ADE20K with both a ResNet-50 and a Swin-L backbones. Moreover, our method only introduces little computation during training and no extra computation during inference. Our code will be released at https://github.com/IDEA-Research/MP-Former.
Hao Zhang 0097, Feng Li 0040, Huaizhe Xu, Shijia Huang, Shilong Liu 0004, Lionel M. Ni, Lei Zhang 0001
CVPR3