Yuanpei Liu

dblp:135/8947 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0008-6144-6547ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Segmentation and scene understanding · 34% Representation and self-supervised learning · 26% Efficient and distributed learning · 13%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding › category discovery
generalized category discovery
2.632025
SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery · NeurIPS 2025
DebGCD: Debiased Learning with Distribution Guidance for Generalized Category Discovery · ICLR 2025
Hyperbolic Category Discovery · CVPR 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery · NeurIPS 2025
Computer vision › Segmentation and scene understanding
hierarchical semantic learning
0.912025
SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning
0.912025
Hyperbolic Category Discovery · CVPR 2025
Machine learning › Representation and self-supervised learning › contrastive learning
soft contrastive learning
0.912025
SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.612022
Distilled Siamese Networks for Visual Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Computer vision › Video understanding and tracking
object tracking
0.612022
Distilled Siamese Networks for Visual Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Computer vision › Face, body and person analysis
person clustering
0.612022
An Efficient Person Clustering Algorithm for Open Checkout-free Groceries · ECCV (38) 2022
Computer vision › Face, body and person analysis
person re-identification
0.612022
An Efficient Person Clustering Algorithm for Open Checkout-free Groceries · ECCV (38) 2022
Computer vision › Video understanding and tracking › object tracking › deep tracking
siamese tracking
0.612022
Distilled Siamese Networks for Visual Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Efficient and distributed learning › distillation
teacher-student distillation
0.612022
Distilled Siamese Networks for Visual Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Trustworthy machine learning › debiasing
debiased learning
0.312025
DebGCD: Debiased Learning with Distribution Guidance for Generalized Category Discovery · ICLR 2025
Machine learning › Efficient and distributed learning
model compression
0.212022
Distilled Siamese Networks for Visual Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Trustworthy machine learning › open-world recognition
open-set recognition
0.212022
An Efficient Person Clustering Algorithm for Open Checkout-free Groceries · ECCV (38) 2022

Methods — techniques the papers use, named apart from their topics

self-supervised pretraining · 0.9self-distillation · 0.9hyperbolic distance · 0.9hierarchical similarity · 0.9debiased classifier · 0.9curriculum learning · 0.9cross-granularity consistency · 0.9mutual learning · 0.6knowledge distillation · 0.6clustering · 0.6
YearPublicationVenuePosition
2025 ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
abstract
The objective in this paper is to improve the per-formance of text-to-image retrieval. To this end, we introduce a new framework that can boost the performance of large-scale pre-trained vision-language models, so that they can be used for text-to-image re-ranking. The approach, Enhanced Language-Image Pre-training (ELIP), uses the text query, via a simple MLP mapping network, to predict a set of visual prompts to condition the ViT image encoding. ELIP can easily be applied to the commonly used CLIP, SigLIP and BLIP-2 networks. On the evaluation side, we set up two new out-of-distribution (OOD) benchmarks, Occluded COCO and ImageNet-R, to assess the zero-shot generalisation of the models to different domains. The results demonstrate that ELIP significantly boosts CLIP/SigLIP/SigLIP-2 text-to-image retrieval performance and outperforms BLIP-2 on several benchmarks, as well as providing an easy means to adapt to OOD datasets.
Guanqi Zhan, Yuanpei Liu, Kai Han 0001, Weidi Xie, Andrew Zisserman
CBMI2
2025 Hyperbolic Category Discovery
abstract
Generalized Category Discovery (GCD) is an intriguing open-world problem that has garnered increasing attention. Given a dataset that includes both labelled and unlabelled images, GCD aims to categorize all images in the unlabelled subset, regardless of whether they belong to known or unknown classes. In GCD, the common practice typically involves applying a spherical projection operator at the end of the self-supervised pretrained backbone, operating within Euclidean or spherical space. However, both of these spaces have been shown to be suboptimal for encoding samples that possesses hierarchical structures. In contrast, hyperbolic space exhibits exponential volume growth relative to radius, making it inherently strong at capturing the hierarchical structure of samples from both seen and unseen categories. Therefore, we propose to tackle the category discovery challenge in the hyperbolic space. We introduce HypCD, a simple Hyperbolic framework for learning hierarchy-aware representations and classifiers for generalized Category Discovery. HypCD first transforms the Euclidean embedding space of the backbone network into hyperbolic space, facilitating subsequent representation and classification learning by considering both hyperbolic distance and the angle between samples. This approach is particularly helpful for knowledge transfer from known to unknown categories in GCD. We thoroughly evaluate HypCD on public GCD benchmarks, by applying it to various baseline and state-of-the-art methods, consistently achieving significant improvements. Project page: https://visual-ai.github.io/hypcd/
Yuanpei Liu, Zhenqi He, Kai Han 0001
CVPR1
2025 DebGCD: Debiased Learning with Distribution Guidance for Generalized Category Discovery
abstract
In this paper, we tackle the problem of Generalized Category Discovery (GCD). Given a dataset containing both labelled and unlabelled images, the objective is to categorize all images in the unlabelled subset, irrespective of whether they are from known or unknown classes. In GCD, an inherent label bias exists between known and unknown classes due to the lack of ground-truth labels for the latter. State-of-the-art methods in GCD leverage parametric classifiers trained through self-distillation with soft labels, leaving the bias issue unattended. Besides, they treat all unlabelled samples uniformly, neglecting variations in certainty levels and resulting in suboptimal learning. Moreover, the explicit identification of semantic distribution shifts between known and unknown classes, a vital aspect for effective GCD, has been neglected. To address these challenges, we introduce DebGCD, a Debiased learning with distribution guidance framework for GCD. Initially, DebGCD co-trains an auxiliary debiased classifier in the same feature space as the GCD classifier, progressively enhancing the GCD features. Moreover, we introduce a semantic distribution detector in a separate feature space to implicitly boost the learning efficacy of GCD. Additionally, we employ a curriculum learning strategy based on semantic distribution certainty to steer the debiased learning at an optimized pace. Thorough evaluations on GCD benchmarks demonstrate the consistent state-of-the-art performance of our framework, highlighting its superiority. Project page: [https://visual-ai.github.io/debgcd/](https://visual-ai.github.io/debgcd/)
Yuanpei Liu, Kai Han 0001
ICLR1
2025 SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery
abstract
This paper investigates the problem of Generalized Category Discovery (GCD). Given a partially labelled dataset, GCD aims to categorize all unlabelled images, regardless of whether they belong to known or unknown classes. Existing approaches typically depend on either single-level semantics or manually designed abstract hierarchies, which limit their generalizability and scalability. To address these limitations, we introduce a SEmantic-aware hierArchical Learning framework (SEAL), guided by naturally occurring and easily accessible hierarchical structures. Within SEAL, we propose a Hierarchical Semantic-Guided Soft Contrastive Learning approach that exploits hierarchical similarity to generate informative soft negatives, addressing the limitations of conventional contrastive losses that treat all negatives equally. Furthermore, a Cross-Granularity Consistency (CGC) module is designed to align the predictions from different levels of granularity. SEAL consistently achieves state-of-the-art performance on fine-grained benchmarks, including the SSB benchmark, Oxford-Pet, and the Herbarium19 dataset, and further demonstrates generalization on coarse-grained datasets. Project page: https://visual-ai.github.io/seal/
Zhenqi He, Yuanpei Liu, Kai Han 0001
NeurIPS2
2022 An Efficient Person Clustering Algorithm for Open Checkout-free Groceries
Yu Zhang 0091, Yuanpei Liu
ECCV (38)4
2022 Distilled Siamese Networks for Visual Tracking
abstract
In recent years, Siamese network based trackers have significantly advanced the state-of-the-art in real-time tracking. Despite their success, Siamese trackers tend to suffer from high memory costs, which restrict their applicability to mobile devices with tight memory budgets. To address this issue, we propose a distilled Siamese tracking framework to learn small, fast and accurate trackers (students), which capture critical knowledge from large Siamese trackers (teachers) by a teacher-students knowledge distillation model. This model is intuitively inspired by the one teacher versus multiple students learning method typically employed in schools. In particular, our model contains a single teacher-student distillation module and a student-student knowledge sharing mechanism. The former is designed using a tracking-specific distillation strategy to transfer knowledge from a teacher to students. The latter is utilized for mutual learning between students to enable in-depth knowledge understanding. Extensive empirical evaluations on several popular Siamese trackers demonstrate the generality and effectiveness of our framework. Moreover, the results on five tracking benchmarks show that the proposed distilled trackers achieve compression rates of up to 18× and frame-rates of 265 FPS, while obtaining comparable tracking accuracy compared to base models.
Jianbing Shen, Yuanpei Liu, Xingping Dong, Xiankai Lu, Fahad Shahbaz Khan, Steven C. H. Hoi
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Image Co-segmentation with Multi-Scale Dual-Cross Correlation Network
abstract
Considering that the global correlation between images is very important for image co-segmentation, we propose a multi-scale Dual-Cross Correlation Network (DCNet) that can efficiently capture global matching information across images to obtain segmentation results. Specifically, the low-dimensional index feature is used to calculate the correlation and the high-dimensional content features are combined with the correlation matrix for final segmentation. Meanwhile, we specially design a Dual-Cross Correlation Module (DCCM) which harvests the spatial and channel correlation with the adjacent pixels of another image on the cross path to enhance the representation of correlation efficiently. By utilizing a further loop operation, each feature can capture the global dependencies from all pixels of another feature. Furthermore, we fuse multi-scale correlation and features into the decoder, which is called Multi-scale Correlation Fusing Decoder (MCFD), to refine the final segmentation results. Moreover, we introduce a new dice loss function to train the whole network by averaging the dice loss value of the foreground and background. Finally, we validate our method on three co-segmentation benchmarks and the results show that our method achieves the state-of-the-art performance.
Yushuo Li, Yuanpei Liu, Xiaopeng Gong, Xiabi Liu
IJCNN2
2020 Multiple people tracking with articulation detection and stitching strategy
Yuanpei Liu, Junbo Yin, Dajiang Yu, Sanyuan Zhao, Jianbing Shen
Neurocomputing1