Yue Zhang 0065

dblp:47/722-65 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-0179-1396ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Query-guided predicate decoupling and prototype approximation learning for scene graph generation
Shichao Kan, Yue Zhang 0065, Yi-Gang Cen, Wanru Xu, Yi Jin 0001, Yidong Li
Expert Syst. Appl.3
2025 Noise-Guided Predicate Representation Extraction and Diffusion-Enhanced Discretization for Scene Graph Generation
abstract
Scene Graph Generation (SGG) is a fundamental task in visual understanding, aimed at providing more precise local detail comprehension for downstream applications. Existing SGG methods often overlook the diversity of predicate representations and the consistency among similar predicates when dealing with long-tail distributions. As a result, the model's decision layer fails to effectively capture details from the tail end, leading to biased predictions. To address this, we propose a Noise-Guided Predicate Representation Extraction and Diffusion-Enhanced Discretization (NoDIS) method. On the one hand, expanding the predicate representation space enhances the model's ability to learn both common and rare predicates, thus reducing prediction bias caused by data scarcity. We propose a conditional diffusion model to reconstructs features and increase the diversity of representations for same category predicates. On the other hand, independent predicate representations in the decision phase increase the learning complexity of the decision layer, making accurate predictions more challenging. To address this issue, we introduce a discretization mapper that learns consistent representations among similar predicates, reducing the learning difficulty and decision ambiguity in the decision layer. To validate the effectiveness of our method, we integrate NoDIS with various SGG baseline models and conduct experiments on multiple datasets. The results consistently demonstrate superior performance.
Shichao Kan, Fanghui Zhang, Wanru Xu, Yue Zhang 0065, Yi-Gang Cen
ICML5
2025 Hierarchical Meta-prototypes Network for Few-shot Action Recognition
abstract
Existing few-shot action recognition (FSAR) studies predominantly follow a metric learning framework, where prototypes are generated directly from features extracted by an encoder, and classification is performed via distance-based matching. However, due to the limited number of available samples, significant variations exist between different video features of the same class. As a result, the same query video may yield different classification results when matched against different sets of support videos. To address this issue, we propose a novel Hierarchical Meta-Prototypes Network (HMP-Net). The key innovation of our approach lies in the introduction of a category-agnostic and feature-agnostic meta-prototype module, which guides video feature mapping into a more suitable feature space. To optimize this meta-prototype, we design an alternating meta-prototype training strategy, where the model first learns to transform features under a fixed meta-prototype, and then the meta-prototype is refined to better guide feature mapping. Additionally, to adapt image-based metric learning models to video-based FSAR tasks, we introduce a series of lightweight adaptation modules. Specifically, we integrate an adapter into the encoder to improve video frame feature extraction, design a hierarchical prototype generation mechanism to enhance overall video understanding, and incorporate a task-specific perception module to extract unique features for each task. These adaptations make our model better suited for FSAR, significantly improving performance. We evaluate HMP-Net on five challenging benchmarks, and experimental results demonstrate that our model achieves new state-of-the-art performance on HMDB51, UCF101, Kinetics, and SthSthV2-Small. Extensive empirical evaluations further highlight the effectiveness and robustness of HMP-Net.
Yi-Gang Cen, Wanru Xu, Yue Zhang 0065, Yi Jin 0001, Yidong Li, Linna Zhang
ACM Multimedia4
2025 Ora-NSC: A Novel Semi-supervised Approach for Oracle Bone Fragment Classification with Imbalanced Classes
abstract
Oracle bone inscriptions (OBI), as precious historical records of The Shang Dynasty, approximately 3,000 years ago, marked the origins of Chinese characters and the development of the Shang civilization. Classifying oracle bone fragments is essential for enhancing fragment assembly efficiency and facilitating the interpretation of divination. To address data imbalance and label scarcity in oracle bone fragment datasets, we propose a novel semi-supervised approach for oracle bone fragment classification, which is named Ora-NSC. The approach integrates Mean Teacher with FixMatch, utilizing the Exponential Moving Average (EMA) strategy to enhance the accuracy and stability of pseudo-label prediction. Additionally, an improved loss function is developed to address class imbalances in fragment segments, thereby further enhancing classification performance. Experimental results demonstrate that our method achieves superior performance on both the Oracle Fragments and CIFAR-100 datasets compared to several state-of-the-art approaches. It improves precision by 4.14%, recall by 4.03%, and F1-score by 4.46%, achieving an overall accuracy of 84.44%. This study represents the first application of semi-supervised learning for oracle bone fragment recognition, offering new possibilities for the preservation of cultural heritage.
Yunyue Hu, Weijie Gao, Yijing Si, Yue Zhang 0065
MMAsia4
2025 Pedestrian Open-Attribute Recognition via Dynamic Semantic Masking
Yue Zhang 0065, Sen Feng, Fanghui Zhang, Guoqi Liu, Yi-Gang Cen
PRCV (7)1
2023 HOI-aware Adaptive Network for Weakly-supervised Action Segmentation
abstract
In this paper, we propose an HOI-aware adaptive network named AdaAct for weakly-supervised action segmentation. Most existing methods learn a fixed network to predict the action of each frame with the neighboring frames. However, this would result in ambiguity when estimating similar actions, such as pouring juice and pouring coffee. To address this, we aim to exploit temporally global but spatially local human-object interactions (HOI) as video-level prior knowledge for action segmentation. The long-term HOI sequence provides crucial contextual information to distinguish ambiguous actions, where our network dynamically adapts to the given HOI sequence at test time. More specifically, we first design a video HOI encoder that extracts, selects, and integrates the most representative HOI throughout the video. Then, we propose a two-branch HyperNetwork to learn an adaptive temporal encoder, which automatically adjusts the parameters based on the HOI information of various videos on the fly. Extensive experiments on two widely-used datasets including Breakfast and 50Salads demonstrate the effectiveness of our method under different evaluation metrics.
Runzhong Zhang, Suchen Wang, Yueqi Duan, Yansong Tang, Yue Zhang 0065, Yap-Peng Tan
IJCAI5
2023 POAR: Towards Open Vocabulary Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition (PAR) aims to predict the attributes of a target pedestrian. Recent methods often address the PAR problem by training a multi-label classifier with predefined attribute classes, but they can hardly exhaust all possible pedestrian attributes in the real world. To tackle this problem, we propose a novel Pedestrian Open-Attribute Recognition (POAR) approach by formulating the problem as a task of image-text search. Our approach employs a Transformer-based Encoder with a Masking Strategy (TEMS) to focus on the attributes of specific pedestrian parts (e.g., head, upper body, lower body, feet, etc.), and introduces a set of attribute tokens to encode the corresponding attributes into visual embeddings. Each attribute category is described as a natural language sentence and encoded by the text encoder. Then, we compute the similarity between the visual and text embeddings to find the best attribute descriptions for the input images. To handle multiple attributes of a single pedestrian, we propose a Many-To-Many Contrastive (MTMC) loss with masked tokens. In addition, we propose a Grouped Knowledge Distillation (GKD) method to minimize the disparity between visual embeddings and unseen attribute text embeddings. We evaluate our proposed method on three PAR datasets with an open-attribute setting. The results demonstrate the effectiveness of our method as a strong baseline for the POAR task. Our code is available at https://github.com/IvyYZ/POAR.
Yue Zhang 0065, Suchen Wang, Shichao Kan, Zhenyu Weng, Yi-Gang Cen, Yap-Peng Tan
ACM Multimedia1
2023 End-to-end feature diversity person search with rank constraint of cross-class matrix
Yue Zhang 0065, Shuqin Wang 0001, Shichao Kan, Yi-Gang Cen, Linna Zhang
Neurocomputing1
2023 Local Correlation Ensemble with GCN Based on Attention Features for Cross-domain Person Re-ID
abstract
Person re-identification (Re-ID) has achieved great success in single-domain. However, it remains a challenging task to adapt a Re-ID model trained on one dataset to another one. Unsupervised domain adaption (UDA) was proposed to migrate a model from a labeled source domain to an unlabeled target domain. The main difference in the cross-domain is different background styles. Although the style transfer approach effectively reduces inter-domain gaps, it ignores the reduction of intra-class differences. Clustering-based pipelines maintain state-of-the-art performance for UDA by learning domain-independent features; however, most existing models do not sufficiently exploit the rich unlabeled samples in target domains due to unsatisfactory clustering. Thus, we propose a novel local correlation ensemble model that focuses on the diversity of intra-class information and the reliability of class centers. Specifically, a pedestrian attention module is proposed to enable the encoder to pay more attention to the person’s features to relieve interference caused by the shared background style. Furthermore, we propose a priority-distance graph convolutional network (PDGCN) module that employs a graph convolutional network network to predict the priority of a node as a class center and then calculates the distance between nodes with high priority values to screen out the class center nodes. Finally, the encoder features (local) and PDGCN features (context-aware) are combined to perform person Re-ID. The results of experiments on the large-scale public Re-ID datasets verified the effectiveness of the proposed method.
Yue Zhang 0065, Fanghui Zhang, Yi Jin 0001, Yi-Gang Cen, Viacheslav V. Voronin, Shaohua Wan 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2022 A GAN-based input-size flexibility model for single image dehazing
Shichao Kan, Yue Zhang 0065, Fanghui Zhang, Yi-Gang Cen
Signal Process. Image Commun.2
2021 Cross-domain Person Re-identification Based on the Sample Relation Guidance
Yue Zhang 0065, Fanghui Zhang, Shichao Kan, Linna Zhang, Jiaping Zong, Yi-Gang Cen
ICIG (2)1
2021 Pedestrian detection with super-resolution reconstruction for low-quality image
Yi Jin 0001, Yue Zhang 0065, Yi-Gang Cen, Yidong Li, Vladimir Mladenovic, Viacheslav V. Voronin
Pattern Recognit.2