Yu Wu 0014

dblp:22/0-14 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 38% Segmentation and scene understanding · 28% Representation and self-supervised learning · 19%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding › image segmentation › deep learning segmentation
continual segmentation
0.912025
Rethinking Query-based Transformer for Continual Image Segmentation · CVPR 2025
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.912025
Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement · CVPR 2025
Computer vision › Segmentation and scene understanding › category discovery
generalized category discovery
0.912025
Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning
part-based representation learning
0.912025
Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.912025
Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement · CVPR 2025
Computer vision › Vision and language
image captioning
0.812024
Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language Reasoning · CVPR 2024
Computer vision › Vision and language
vision-language model
0.812024
Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language Reasoning · CVPR 2024
Computer vision › Vision and language
visual question answering
0.812024
Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language Reasoning · CVPR 2024
Computer vision › Vision and language › cross-modal matching
image-text matching
0.712023
Grounded Image Text Matching with Mismatched Relation Reasoning · ICCV 2023
Computer vision › Vision and language
visual grounding
0.712023
Grounded Image Text Matching with Mismatched Relation Reasoning · ICCV 2023
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.312025
Rethinking Query-based Transformer for Continual Image Segmentation · CVPR 2025
Machine learning › Transfer learning and domain adaptation
knowledge transfer
0.312025
Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement · CVPR 2025

Methods — techniques the papers use, named apart from their topics

visual query replay · 0.9vision transformer · 0.9learnable part queries · 0.9cross-stage consistency · 0.9contrastive learning · 0.9instruction tuning · 0.8dependency parser · 0.8caption correction · 0.8transformer · 0.7bidirectional message propagation · 0.7
YearPublicationVenuePosition
2025 Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement
abstract
Generalized Category Discovery (GCD) aims to recognize unlabeled images from known and novel classes by distinguishing novel classes from known ones, while also transferring knowledge from another set of labeled images with known classes. Existing GCD methods rely on self-supervised vision transformers such as DINO for representation learning. However, focusing solely on the global representation of the DINO CLS token introduces an inherent trade-off between discriminability and generalization. In this paper, we introduce an adaptive part discovery and learning method, called APL, which generates consistent object parts and their correspondences across different similar images using a set of shared learnable part queries and DINO part priors, without requiring any additional annotations. More importantly, we propose a novel all-min contrastive loss to learn discriminative yet generalizable part representation, which adaptively highlights discriminative object parts to distinguish similar categories for enhanced discriminability while simultaneously sharing other parts to facilitate knowledge transfer for improved generalization. Our APL can easily be incorporated into different GCD frameworks by replacing their CLS token feature with our part representations, showing significant enhancements on fine-grained datasets.
Qiyuan Dai 0001, Hanzhuo Huang, Yu Wu 0014, Sibei Yang
CVPR3
2025 Rethinking Query-based Transformer for Continual Image Segmentation
abstract
Class-Incremental/Continual image segmentation (CIS) aims to train an image segmenter in stages, where the set of available categories differs at each stage. To leverage the built-in objectness of query-based transformers, which mitigates catastrophic forgetting of mask proposals, current methods often decouple mask generation from the continual learning process. This study, however, identifies two key issues with decoupled frameworks: loss of plasticity and heavy reliance on input data order. To address these, we conduct an in-depth investigation of the built-in objectness and find that highly aggregated image features provide a shortcut for queries to generate masks through simple feature alignment. Based on this, we propose SimCIS, a simple yet powerful baseline for CIS. Its core idea is to directly select image features for query assignment, ensuring "perfect alignment" to preserve objectness, while simultaneously allowing queries to select new classes to promote plasticity. To further combat catastrophic forgetting of categories, we introduce cross-stage consistency in selection and an innovative "visual query"-based replay mechanism. Experiments demonstrate that SimCIS consistently outperforms state-of-the-art methods across various segmentation tasks, settings, splits, and input data orders. All models and codes will be made publicly available at https://github.com/SooLab/SimCIS.
Cheng Shi 0001, Dingyou Wang, Jiajin Tang, Zhengxuan Wei, Yu Wu 0014, Guanbin Li, Sibei Yang
CVPR6
2024 Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language Reasoning
abstract
Generative vision-language models (VLMs) have shown impressive performance in zero-shot vision-language tasks like image captioning and visual question answering. How-ever, improving their zero-shot reasoning typically requires second-stage instruction tuning, which relies heavily on human-labeled or large language model-generated annotation, incurring high labeling costs. To tackle this challenge, we introduce Image-Conditioned Caption Correction (ICCC), a novel pre-training task designed to enhance VLMs' zero-shot performance without the need for labeled task-aware data. The ICCC task compels VLMs to rectify mismatches between visual and language concepts, thereby enhancing instruction following and text generation conditioned on visual inputs. Leveraging language structure and a lightweight dependency parser, we construct data samples of ICCC taskfrom image-text datasets with low labeling and computation costs. Experimental results on BLIP-2 and InstructBLIP demonstrate significant improvements in zero-shot image-text generation-based VL tasks through ICCC instruction tuning.
Rongjie Li, Yu Wu 0014, Xuming He 0001
CVPR2
2023 Grounded Image Text Matching with Mismatched Relation Reasoning
abstract
This paper introduces Grounded Image Text Matching with Mismatched Relation (GITM-MR), a novel visual-linguistic joint task that evaluates the relation understanding capabilities of transformer-based pre-trained models. GITM-MR requires a model to first determine if an expression describes an image, then localize referred objects or ground the mismatched parts of the text. We provide a benchmark for evaluating vision-language (VL) models on this task, with a focus on the challenging settings of limited training data and out-of-distribution sentence lengths. Our evaluation demonstrates that pre-trained VL models often lack data efficiency and length generalization ability. To address this, we propose the Relation-sensitive Correspondence Reasoning Network (RCRN), which incorporates relation-aware reasoning via bi-directional message propagation guided by language structure. Our RCRN can he interpreted as a modular program and delivers strong performance in terms of both length generalization and data efficiency. The code and data are available on https://githuh.coin/SHTUPLUS/GITM-MR.
Yu Wu 0014, Yana Wei, Haozhe Wang 0002, Yongfei Liu, Sibei Yang, Xuming He 0001
ICCV1