EDBT 2026 Demo / reviewers in the wild / expert
Yang Liu 0357
dblp:51/3710-357
· DBLP profile ↗
20ranked-venue papers
7as first author
20since 2021 · last 2026
0009-0003-8540-9154ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Modal Primitive Retrieval for Compositional Zero-Shot Learning
Chenchen Jing, Haozhe Zhang 0002, Junbo Lu, Yang Liu 0357, Hao Chen 0041, Xiaoqin Zhang 0002, Chunhua Shen |
Int. J. Comput. Vis. | 4 |
| 2025 | ROS-SAM: High-Quality Interactive Segmentation for Remote Sensing Moving ObjectabstractThe availability of large-scale remote sensing video data underscores the importance of high-quality interactive segmentation. However, challenges such as small object sizes, ambiguous features, and limited generalization make it difficult for current methods to achieve this goal. In this work, we propose ROS-SAM, a method designed to achieve high-quality interactive segmentation while preserving generalization across diverse remote sensing data. The ROS-SAM is built upon three key innovations: 1) LoRA-based fine-tuning, which enables efficient domain adaptation while maintaining SAM’s generalization ability, 2) Enhancement of network deep layers to improve the discriminability of extracted features, thereby reducing misclassifications, and 3) Integration of global context with local boundary details in the mask decoder to generate high-quality segmentation masks. Additionally, we redesign the data pipeline to ensure the model learns to better handle objects at varying scales during training while focusing on high-quality predictions during inference. Experiments on remote sensing video datasets show that the data pipeline boosts the IoU by 6%, while ROS-SAM increases the IoU by 13%. Finally, when evaluated on existing remote sensing object tracking datasets, ROS-SAM demonstrates impressive zero-shot capabilities, generating masks that closely resemble manual annotations. These results confirm ROS-SAM as a powerful tool for fine-grained segmentation in remote sensing applications. Code is available at: https://github.com/ShanZard/ROS-SAM. Yang Liu 0357, Lei Zhou 0008 |
CVPR | 2 |
| 2025 | SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator TrajectoriesabstractWhile MLLMs have demonstrated impressive image understanding capabilities, they still struggle with pixel-level comprehension, limiting their practical applications. Current evaluation tasks such as VQA and visual grounding remain too coarse to assess fine-grained pixel comprehension accurately. Although segmentation is foundational for pixel-level understanding, existing methods often require MLLMs to generate implicit tokens, decoded through external pixel decoders. This approach disrupts the MLLM’s text output space, potentially compromising language capabilities and reducing flexibility and extensibility while failing to reflect the model’s intrinsic pixel-level understanding. Thus, we introduce the Human-Like Mask Annotation Task (HLMAT), a new paradigm where MLLMs mimic human annotators using interactive segmentation tools. Modelling segmentation as a multi-step Markov Decision Process, HLMAT enables MLLMs to iteratively generate text-based click points, achieving high-quality masks without architectural changes or implicit tokens. Through this setup, we develop SegAgent, a model fine-tuned on human-like annotation trajectories, which achieves performance comparable to SoTA methods and supports additional tasks like mask refinement and annotation filtering. HLMAT provides a protocol for assessing fine-grained pixel understanding in MLLMs and introduces a vision-centric, multi-step decision-making task that facilitates the exploration of MLLMs’ visual reasoning abilities. Our adaptations of policy improvement method StaR and PRM guided tree search further enhance model robustness in complex segmentation tasks, laying a foundation for future advancements in fine-grained visual perception and multi-step decision-making for MLLMs. Code can be found at https://github.com/aim-uofa/SegAgent. Muzhi Zhu, Yuzhuo Tian, Hao Chen 0041, Chunluan Zhou, Qingpei Guo, Yang Liu 0357, Ming Yang 0007, Chunhua Shen |
CVPR | 6 |
| 2025 | Unified Open-World Segmentation with Multi-Modal Prompts
Yang Liu 0357, Yufei Yin, Chenchen Jing, Muzhi Zhu, Hao Chen 0041, Yuling Xi, Hao Wang 0052, Chunhua Shen |
ICCV | 1 |
| 2025 | Learning Visual Proxy for Compositional Zero-Shot Learning
Yang Liu 0357, Chenchen Jing, Lei Zhou 0008, Wenjun Wang 0002 |
ICCV | 3 |
| 2025 | Masked Channel Modeling for Bootstrapping Visual Pre-training
Yang Liu 0357, Muzhi Zhu, Yue Cao 0001, Tiejun Huang 0001, Chunhua Shen |
Int. J. Comput. Vis. | 1 |
| 2025 | Segment Anything in Context with Vision Foundation Models
Yang Liu 0357, Muzhi Zhu, Hao Chen 0041, Hao Wang 0052, Raviteja Vemulapalli, Chunhua Shen |
Int. J. Comput. Vis. | 1 |
| 2024 | DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative DataabstractInstance segmentation is data-hungry, and as model capacity increases, data scale becomes crucial for improving the accuracy. Most instance segmentation datasets today require costly manual annotation, limiting their data scale. Models trained on such data are prone to overfitting on the training set, especially for those rare categories. While recent works have delved into exploiting generative models to create synthetic datasets for data augmentation, these approaches do not efficiently harness the full potential of generative models. To address these issues, we introduce a more efficient strategy to construct generative datasets for data augmentation, termed DiverGen. Firstly, we provide an explanation of the role of generative data from the perspective of distribution discrepancy. We investigate the impact of different data on the distribution learned by the model. We argue that generative data can expand the data distribution that the model can learn, thus mitigating overfitting. Additionally, we find that the diversity of generative data is crucial for improving model performance and enhance it through various strategies, including category diversity, prompt diversity, and generative model diversity. With these strategies, we can scale the data to millions while maintaining the trend of model performance improvement. On the LVIS dataset, DiverGen significantly outperforms the strong model X-Paste, achieving +1.1 box AP and +1.1 mask AP across all categories, and +1.9 box AP and +2.5 mask AP for rare categories. Our codes are available at https://github.com/aim-uofa/DiverGen. Chengxiang Fan, Muzhi Zhu, Hao Chen 0041, Yang Liu 0357, Weijia Wu 0001, Huaqi Zhang, Chunhua Shen |
CVPR | 4 |
| 2024 | Matcher: Segment Anything with One Shot Using All-Purpose Feature MatchingabstractPowered by large-scale pre-training, vision foundation models exhibit significant potential in open-world image understanding. However, unlike large language models that excel at directly tackling various language tasks, vision foundation models require a task-specific model structure followed by fine-tuning on specific tasks. In this work, we present $\textbf{Matcher}$, a novel perception paradigm that utilizes off-the-shelf vision foundation models to address various perception tasks. Matcher can segment anything by using an in-context example without training. Additionally, we design three effective components within the Matcher framework to collaborate with these foundation models and unleash their full potential in diverse perception tasks. Matcher demonstrates impressive generalization performance across various segmentation tasks, all without training. For example, it achieves 52.7% mIoU on COCO-20$^i$ with one example, surpassing the state-of-the-art specialist model by 1.6%. In addition, Matcher achieves 33.0% mIoU on the proposed LVIS-92$^i$ for one-shot semantic segmentation, outperforming the state-of-the-art generalist model by 14.4%. Our visualization results further showcase the open-world generality and flexibility of Matcher when applied to images in the wild. Yang Liu 0357, Muzhi Zhu, Hengtao Li, Hao Chen 0041, Chunhua Shen |
ICLR | 1 |
| 2024 | Generative Active Learning for Long-tailed Instance SegmentationabstractRecently, large-scale language-image generative models have gained widespread attention and many works have utilized generated data from these models to further enhance the performance of perception tasks. However, not all generated data can positively impact downstream models, and these methods do not thoroughly explore how to better select and utilize generated data. On the other hand, there is still a lack of research oriented towards active learning on generated data. In this paper, we explore how to perform active learning specifically for generated data in the long-tailed instance segmentation task. Subsequently, we propose BSGAL, a new algorithm that estimates the contribution of the current batch-generated data based on gradient cache. BSGAL is meticulously designed to cater for unlimited generated data and complex downstream segmentation tasks. BSGAL outperforms the baseline approach and effectually improves the performance of long-tailed segmentation. Muzhi Zhu, Chengxiang Fan, Hao Chen 0041, Yang Liu 0357, Weian Mao, Xiaogang Xu 0002, Chunhua Shen |
ICML | 4 |
| 2024 | A Simple Image Segmentation Framework via In-Context ExamplesabstractRecently, there have been explorations of generalist segmentation models that can effectively tackle a variety of image segmentation tasks within a unified in-context learning framework. However, these methods still struggle with task ambiguity in in-context segmentation, as not all in-context examples can accurately convey the task information. In order to address this issue, we present SINE, a simple image $\textbf{S}$egmentation framework utilizing $\textbf{in}$-context $\textbf{e}$xamples. Our approach leverages a Transformer encoder-decoder structure, where the encoder provides high-quality image representations, and the decoder is designed to yield multiple task-specific output masks to eliminate task ambiguity effectively. Specifically, we introduce an In-context Interaction module to complement in-context information and produce correlations between the target image and the in-context example and a Matching Transformer that uses fixed matching and a Hungarian algorithm to eliminate differences between different tasks. In addition, we have further perfected the current evaluation system for in-context image segmentation, aiming to facilitate a holistic appraisal of these models. Experiments on various segmentation tasks show the effectiveness of the proposed method. Yang Liu 0357, Chenchen Jing, Hengtao Li, Muzhi Zhu, Hao Chen 0041, Chunhua Shen |
NeurIPS | 1 |
| 2024 | Unleashing the Potential of the Diffusion Model in Few-shot Semantic SegmentationabstractThe Diffusion Model has not only garnered noteworthy achievements in the realm of image generation
but has also demonstrated its potential as an effective pretraining method utilizing unlabeled data.
Drawing from the extensive potential unveiled by the Diffusion Model in both semantic correspondence and open vocabulary segmentation, our work initiates an investigation into employing the Latent Diffusion Model for Few-shot Semantic Segmentation.
Recently, inspired by the in-context learning ability of large language models, Few-shot Semantic Segmentation has evolved into In-context Segmentation tasks, morphing into a crucial element in assessing generalist segmentation models.
In this context, we concentrate
on Few-shot Semantic Segmentation,
establishing a solid foundation for the future development of a Diffusion-based generalist model for segmentation. Our initial focus lies in understanding how to facilitate interaction between the query image and the support image, resulting in the proposal of a KV fusion method within the self-attention framework.
Subsequently, we delve deeper into optimizing the infusion of information from the support mask and simultaneously re-evaluating how to provide reasonable supervision from the query mask.
Based on our analysis, we establish a simple and effective framework named DiffewS, maximally retaining the original Latent Diffusion Model's generative framework and effectively utilizing the pre-training prior. Experimental results demonstrate that our method significantly outperforms the previous SOTA models in multiple settings. Muzhi Zhu, Yang Liu 0357, Zekai Luo, Chenchen Jing, Hao Chen 0041, Guangkai Xu, Chunhua Shen |
NeurIPS | 2 |
| 2023 | Feature Prediction Diffusion Model for Video Anomaly DetectionabstractAnomaly detection in the video is an important research area and a challenging task in real applications. Due to the unavailability of large-scale annotated anomaly events, most existing video anomaly detection (VAD) methods focus on learning the distribution of normal samples to detect the substantially deviated samples as anomalies. To well learn the distribution of normal motion and appearance, many auxiliary networks are employed to extract foreground object or action information. These high-level semantic features effectively filter the noise from the background to decrease its influence on detection models. However, the capability of these extra semantic models heavily affects the performance of the VAD methods. Motivated by the impressive generative and anti-noise capacity of diffusion model (DM), in this work, we introduce a novel DM-based method to predict the features of video frames for anomaly detection. We aim to learn the distribution of normal samples without any extra high-level semantic feature extraction models involved. To this end, we build two denoising diffusion implicit modules to predict and refine the features. The first module concentrates on feature motion learning, while the last focuses on feature appearance learning. To the best of our knowledge, it is the first DM-based method to predict frame features for VAD. The strong capacity of DMs also enables our method to more accurately predict the normal features than non-DM-based feature prediction-based VAD methods. Extensive experiments show that the proposed approach substantially outperforms state-of-the-art competing methods. The code is available atFPDM. Yang Liu 0357, Guansong Pang, Wenjun Wang 0002 |
ICCV | 3 |
| 2023 | Generalized Zero-Shot Learning via Implicit Attribute CompositionabstractZero-shot learning (ZSL) is an important but challenging task in computer vision that aims to identify unseen classes without matching training samples. Current cutting-edge ZSL methods based on locality focus on acquiring the explicit locality of distinguishing characteristics, which could face a lack of adequate supervision at the class attribute level. This paper introduces a novel approach called IAC, which aims to learn Implicit Attribute Composition for ZSL. This method is more comprehensive compared to attribute localization that solely focuses on class-level attribute supervision. IAC utilizes subspace representations that efficiently capture the inherent structure of high-dimensional image features. Then, we learn implicit attribute composition through subspace representation learning. The superiority of the proposed IAC compared to the state-of-the-art is demonstrated through sufficient experiments conducted on three commonly used ZSL datasets, CUB, SUN, and AwA2. Lei Zhou 0008, Yang Liu 0357, Qiang Li 0060 |
SMC | 2 |
| 2023 | Information bottleneck and selective noise supervision for zero-shot learning
Lei Zhou 0008, Yang Liu 0357, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Jun Zhou 0001, Yazhou Yao, Tatsuya Harada, Edwin R. Hancock |
Mach. Learn. | 2 |
| 2023 | Attribute subspaces for zero-shot learning
Lei Zhou 0008, Yang Liu 0357, Xiao Bai 0001, Na Li 0014, Xiaohan Yu 0001, Jun Zhou 0001, Edwin R. Hancock |
Pattern Recognit. | 2 |
| 2022 | Where to Focus: Investigating Hierarchical Attention Relationship for Fine-Grained Visual Classification
Yang Liu 0357, Lei Zhou 0008, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Xiaohan Yu 0001, Jun Zhou 0001, Edwin R. Hancock |
ECCV (24) | 1 |
| 2022 | Learning Prototype via Placeholder for Zero-shot RecognitionabstractZero-shot learning (ZSL) aims to recognize unseen classes by exploiting semantic descriptions shared between seen classes and unseen classes. Current methods show that it is effective to learn visual-semantic alignment by projecting semantic embeddings into the visual space as class prototypes. However, such a projection function is only concerned with seen classes. When applied to unseen classes, the prototypes often perform suboptimally due to domain shift. In this paper, we propose to learn prototypes via placeholders, termed LPL, to eliminate the domain shift between seen and unseen classes. Specifically, we combine seen classes to hallucinate new classes which play as placeholders of the unseen classes in the visual and semantic space. Placed between seen classes, the placeholders encourage prototypes of seen classes to be highly dispersed. And more space is spared for the insertion of well-separated unseen ones. Empirically, well-separated prototypes help counteract visual-semantic misalignment caused by domain shift. Furthermore, we exploit a novel semantic-oriented fine-tuning method to guarantee the semantic reliability of placeholders. Extensive experiments on five benchmark datasets demonstrate the significant performance gain of LPL over the state-of-the-art methods. Zaiquan Yang, Yang Liu 0357, Wenjia Xu, Lei Zhou 0008 |
IJCAI | 2 |
| 2021 | Goal-Oriented Gaze Estimation for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen classes. Since semantic knowledge is built on attributes shared between different classes, which are highly local, strong prior for localization of object attribute is beneficial for visual-semantic embedding. Interestingly, when recognizing unseen images, human would also automatically gaze at regions with certain semantic clue. Therefore, we introduce a novel goal-oriented gaze estimation module (GEM) to improve the discriminative attribute localization based on the class-level attributes for ZSL. We aim to predict the actual human gaze location to get the visual attention regions for recognizing a novel object guided by attribute description. Specifically, the task-dependent attention is learned with the goal-oriented GEM, and the global image features are simultaneously optimized with the regression of local attribute features. Experiments on three ZSL benchmarks, i.e., CUB, SUN and AWA2, show the superiority or competitiveness of our proposed method against the state-of-the-art ZSL methods. The ablation analysis on real gaze data CUB-VWSW also validates the benefits and accuracy of our gaze estimation module. This work implies the promising benefits of collecting human gaze dataset and automatic gaze estimation algorithms on high-level computer vision tasks. The code is available at https://github.com/osierboy/GEM-ZSL. Yang Liu 0357, Lei Zhou 0008, Xiao Bai 0001, Yifei Huang 0002, Lin Gu 0003, Jun Zhou 0001, Tatsuya Harada |
CVPR | 1 |
| 2021 | Relation-Aware Reasoning with Graph Convolutional Network
Lei Zhou 0008, Yang Liu 0357, Xiao Bai 0001, Xiang Wang 0014, Chen Wang 0026, Liang Zhang 0044, Lin Gu 0003 |
ICIG (1) | 2 |