Zihan Ye

dblp:227/4545 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DCS-Mamba: Linear-Complexity Spatiotemporal Modeling with Semantic Modulation for High-Resolution Remote Sensing
Jiachen Yuan, Wenzheng Huang, Yihan Huang, Zihan Ye, Yanqin Shi, Xianyi Yang, Ning Xin, Md Maruf Hasan
ICIC (3)4
2026 MSSTGIN: A novel adaptive graph method for long-term traffic speed forecasting considering global-local multiscale spatiotemporal correlations
Chentao Yao, Fumin Zou, Zihan Ye
Expert Syst. Appl.4
2026 A negative-anchored self-relabeling strategy for multi-label class-incremental learning
Kaile Du, Junzhou Xie, Fan Lyu, Zihan Ye, Guangcan Liu
Neurocomputing5
2026 Negative-weighted knowledge distillation regularized graph convolutional network for multi-label class-incremental learning
Kaile Du, Junzhou Xie, Fan Lyu, Zihan Ye, Yuyang Li 0005, Guangcan Liu
Pattern Recognit.5
2026 Adversarial Robustness in Zero-Shot Learning: An Empirical Study on Class and Concept-Level Vulnerabilities
abstract
Zero-shot Learning (ZSL) aims to enable image classifiers to recognize images from unseen classes that were not included during training. Unlike traditional supervised classification, ZSL typically relies on learning a mapping from visual features to predefined, human-understandable class concepts. While ZSL models promise to improve generalization and interpretability, their robustness under systematic input perturbations remain unclear. In this study, we present an empirical analysis about the robustness of existing ZSL methods at both class-level and concept-level. Specifically, we successfully disrupted their class prediction by the well-known non-target class attack (clsA). However, in the Generalized Zero-shot Learning (GZSL) setting, we observe that the success of clsA is only at the original best-calibrated point. After the attack, the optimal best-calibration point shifts, and ZSL models maintain relatively strong performance at other calibration points, indicating that clsA results in a spurious attack success in the GZSL. To address this, we propose the Class-Bias Enhanced Attack (CBEA), which completely eliminates GZSL accuracy across all calibrated points by enhancing the gap between seen and unseen class probabilities. Next, at concept-level attack, we introduce two novel attack modes: Class-Preserving Concept Attack (CPconA) and Non-Class-Preserving Concept Attack (NCPconA). Our extensive experiments evaluate three typical ZSL models across various architectures from the past three years and reveal that ZSL models are vulnerable not only to the traditional class attack but also to concept-based attacks. These attacks allow malicious actors to easily manipulate class predictions by erasing or introducing concepts. Our findings highlight a significant performance gap between existing approaches, emphasizing the need for improved adversarial robustness in current ZSL models. Our codes are available at https://github.com/FouriYe/AttackZSL_TIP26.
Zihan Ye, Shreyank N. Gowda, Yuping Yan, Ling Shao 0001
IEEE Trans. Image Process.1
2025 ZeroDiff: Solidified Visual-semantic Correlation in Zero-Shot Learning
abstract
Zero-shot Learning (ZSL) aims to enable classifiers to identify unseen classes. This is typically achieved by generating visual features for unseen classes based on learned visual-semantic correlations from seen classes. However, most current generative approaches heavily rely on having a sufficient number of samples from seen classes. Our study reveals that a scarcity of seen class samples results in a marked decrease in performance across many generative ZSL techniques. We argue, quantify, and empirically demonstrate that this decline is largely attributable to spurious visual-semantic correlations. To address this issue, we introduce ZeroDiff, an innovative generative framework for ZSL that incorporates diffusion mechanisms and contrastive representations to enhance visual-semantic correlations. ZeroDiff comprises three key components: (1) Diffusion augmentation, which naturally transforms limited data into an expanded set of noised data to mitigate generative model overfitting; (2) Supervised-contrastive (SC)-based representations that dynamically characterize each limited sample to support visual feature generation; and (3) Multiple feature discriminators employing a Wasserstein-distance-based mutual learning approach, evaluating generated features from various perspectives, including pre-defined semantics, SC-based representations, and the diffusion process. Extensive experiments on three popular ZSL benchmarks demonstrate that ZeroDiff not only achieves significant improvements over existing ZSL methods but also maintains robust performance even with scarce training data. Our codes are available at https://github.com/FouriYe/ZeroDiff_ICLR25.
Zihan Ye, Shreyank N. Gowda, Shiming Chen 0002, Xiaowei Huang 0001, Fahad Shahbaz Khan, Yaochu Jin, Kaizhu Huang, Xiao-Bo Jin
ICLR1
2025 QCSH: Quantization Controlled Semantic Hashing for Effective Similar Text Search
abstract
With the rapid growth of digital information, efficient similarity search has become a crucial challenge in large-scale information retrieval. Semantic hashing provides an effective solution by encoding high-dimensional data into compact binary representations, thereby significantly improving retrieval efficiency. While deep learning-based semantic hashing has shown promising results, existing methods struggle to preserve fine-grained features and suffer from excessive quantization errors due to fixed-threshold hard binarization. To address these issues, we propose an unsupervised novel semantic text hashing framework, Quantization Controlled Semantic Hashing (QCSH), which enhances feature representation while refines the binarization process. QCSH integrates several specialized modules to enhance hashing performance. QCSH enhances the extraction of effective features, and by jointly optimizing hash codes and quantization factors, it effectively captures their interdependence. Additionally, QCSH applies an optimized binarization strategy that preserves information fidelity and maximizes mutual information, effectively reducing quantization errors and ensuring a balanced bit distribution. Extensive experiments on three public datasets show that QCSH outperforms state-of-the-art baselines on various numbers of bits of hash codes.
Zihan Ye, Zhitian Hou, Ge Lin 0002
SMC1
2025 An improved adaptive decomposition and reconstruction-based signal denoising method for wind turbine drive systems
Haining Lu, Zihan Ye, Yanyan Nie
Adv. Eng. Informatics4
2024 Effective Message Hiding with Order-Preserving Mechanisms
Yu Gao 0042, Xuchong Qiu, Zihan Ye
BMVC3
2023 Semantic Similarity Distance: Towards better text-image consistency metric in text-to-image generation
Zhaorui Tan, Xi Yang 0008, Zihan Ye, Qiufeng Wang 0001, Yuyao Yan, Anh Nguyen 0003, Kaizhu Huang
Pattern Recognit.3
2023 Rebalanced Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to identify unseen classes with zero samples during training. Broadly speaking, present ZSL methods usually adopt class-level semantic labels and compare them with instance-level semantic predictions to infer unseen classes. However, we find that such existing models mostly produce imbalanced semantic predictions, i.e. these models could perform precisely for some semantics, but may not for others. To address the drawback, we aim to introduce an imbalanced learning framework into ZSL. However, we find that imbalanced ZSL has two unique challenges: (1) Its imbalanced predictions are highly correlated with the value of semantic labels rather than the number of samples as typically considered in the traditional imbalanced learning; (2) Different semantics follow quite different error distributions between classes. To mitigate these issues, we first formalize ZSL as an imbalanced regression problem which offers empirical evidences to interpret how semantic labels lead to imbalanced semantic predictions. We then propose a re-weighted loss termed Re-balanced Mean-Squared Error (ReMSE), which tracks the mean and variance of error distributions, thus ensuring rebalanced learning across classes. As a major contribution, we conduct a series of analyses showing that ReMSE is theoretically well established. Extensive experiments demonstrate that the proposed method effectively alleviates the imbalance in semantic prediction and outperforms many state-of-the-art ZSL methods.
Zihan Ye, Guanyu Yang 0002, Xiao-Bo Jin, Youfa Liu, Kaizhu Huang
IEEE Trans. Image Process.1
2022 Disentangling Semantic-to-Visual Confusion for Zero-Shot Learning
abstract
Using generative models to synthesize visual features from semantic distribution is one of the most popular solutions to ZSL image classification in recent years. The triplet loss (TL) is popularly used to generate realistic visual distributions from semantics by automatically searching discriminative representations. However, the traditional TL cannot search reliable unseen disentangled representations due to the unavailability of unseen classes in ZSL. To alleviate this drawback, we propose in this work a multi-modal triplet loss (MMTL) which utilizes multi-modal information to search adisentangledrepresentation space. As such, all classes can interplay which can benefit learning disentangled class representations in the searched space. Furthermore, we develop a novel model called Disentangling Class Representation Generative Adversarial Network (DCR-GAN) focusing on exploiting the disentangled representations in training, feature synthesis, and final recognition stages. Benefiting from the disentangled representations, DCR-GAN could fit a more realistic distribution over both seen and unseen features. Extensive experiments show that our proposed model can lead to superior performance to the state-of-the-arts on four benchmark datasets.
Zihan Ye, Fuyuan Hu, Fan Lyu, Kaizhu Huang
IEEE Trans. Multim.1
2021 Multi-Domain Multi-Task Rehearsal for Lifelong Learning
abstract
Rehearsal, seeking to remind the model by storing old knowledge in lifelong learning, is one of the most effective ways to mitigate catastrophic forgetting, i.e., biased forgetting of previous knowledge when moving to new tasks. However, the old tasks of the most previous rehearsal-based methods suffer from the unpredictable domain shift when training the new task. This is because these methods always ignore two significant factors. First, the Data Imbalance between the new task and old tasks that makes the domain of old tasks prone to shift. Second, the Task Isolation among all tasks will make the domain shift toward unpredictable directions; To address the unpredictable domain shift, in this paper, we propose Multi-Domain Multi-Task (MDMT) rehearsal to train the old tasks and new task parallelly and equally to break the isolation among tasks. Specifically, a two-level angular margin loss is proposed to encourage the intra-class/task compactness and inter-class/task discrepancy, which keeps the model from domain chaos. In addition, to further address domain shift of the old tasks, we propose an optional episodic distillation loss on the memory to anchor the knowledge for each old task. Experiments on benchmark datasets validate the proposed approach can effectively mitigate the unpredictable domain shift.
Fan Lyu, Wei Feng 0005, Zihan Ye, Fuyuan Hu, Song Wang 0002
AAAI4
2020 Associating Multi-Scale Receptive Fields For Fine-Grained Recognition
abstract
Extracting and fusing part features have become the key of fined-grained image recognition. Recently, Non-local (NL) module has shown excellent improvement in image recognition. However, it lacks the mechanism to model the interactions between multi-scale part features, which is vital for fine-grained recognition. In this paper, we propose a novel cross-layer non-local (CNL) module to associate multi-scale receptive fields by two operations. First, CNL computes correlations between features of a query layer and all response layers. Second, all response features are weighted according to the correlations and are added to the query features. Due to the interactions of cross-layer features, our model builds spatial dependencies among multi-level layers and learns more discriminative features. In addition, we can reduce the aggregation cost if we set low-dimensional deep layer as query layer. Experiments are conducted to show our model achieves or surpasses state-of-the-art results on three benchmark datasets of fine-grained classification. Our codes can be found at github.com/FouriYe/CNL-ICIP2020.
Zihan Ye, Fuyuan Hu, Zhenping Xia, Fan Lyu, Pengqing Liu
ICIP1
2019 SR-GAN: Semantic Rectifying Generative Adversarial Network for Zero-shot Learning
abstract
The existing Zero-Shot learning (ZSL) methods may suffer from the vague class attributes that are highly overlapped for different classes. Unlike these methods that ignore the discrimination among classes, in this paper, we propose to classify unseen image by rectifying the semantic space guided by the visual space. First, we pre-train a Semantic Rectifying Network (SRN) to rectify semantic space with a semantic loss and a rectifying loss. Then, a Semantic Rectifying Generative Adversarial Network (SR-GAN) is built to generate plausible visual feature of unseen class from both semantic feature and rectified semantic feature. To guarantee the effectiveness of rectified semantic features and synthetic visual features, a pre-reconstruction and a post reconstruction networks are proposed, which keep the consistency between visual feature and semantic feature. Experimental results demonstrate that our approach significantly outperforms the state-of-the-arts on four benchmark datasets.
Zihan Ye, Fan Lyu, Qiming Fu 0001, Jinchang Ren, Fuyuan Hu
ICME1
2019 An Object Attribute Guided Framework for Robot Learning Manipulations from Human Demonstration Videos
abstract
Learning manipulations from videos is an inspiriting way for robots to acquire new skills. In this paper, we propose a framework that can generate robotic manipulation plans by observing human demonstration videos without special marks or unnatural demonstrated behaviors. More specifically, the framework contains a video parsing module and a robot execution module. The first module recognizes the demonstrator's actions using two-stream convolution neural networks, and classifies the operated objects by adopting a Mask R-CNN. After that, two XGBoost classifiers are applied to further classify the objects into subject object and patient object respectively, according to the demonstrator's actions. In the second module, a grammar-based parser is used to summarize the videos and generate the common instructions for robot execution. Extensive experiments are conducted on a publicly available video datasets consisting of 273 videos and manifest that our approach is able to learn manipulation plans from demonstration videos with high accuracy (73.36%). Furthermore, we integrate our framework with a humanoid robot Baxter to perform the manipulation learning from demonstration videos, which effectively verifies the performance of our framework.
Dayong Liang, Xiaojing Zhou, Zihan Ye, Wenyin Liu
IROS6
2019 Social Media Popularity Prediction Based on Visual-Textual Features with XGBoost
abstract
Popularity prediction for social media is an efficient way for scientists to explore advanced predictive trend and make better strategic decisions for future. In this paper, we propose a framework that uses visual-textual features combined with XGBoost for popularity prediction. More specifically, the framework contains three procedures, including visual-textual features extraction, features fusion and XGBoost regression. In order to extract the visual-textual data, on the one hand, we first adopt one-hot encoder to encode the metadata of the posts, and then apply a word2vec model to produce word embeddings. On the other hand, we adopt a shape descriptor called Hu moment to extract the visual features from the images. What's more, we exploit user's information, e.g. users' followings and users' followers, for providing extra social features to the regression. After that, we fuse the multi-modal features and input them to the XGBoost directly for popularity prediction. Extensive experiments conducted on the SMPD2019 dataset manifest the effectiveness of our system. Furthermore, our approach achieves the 3nd place on the leader board of the Grand Challenge in ACM Multimedia 2019.
Dayong Liang, Zhanmo Zhu, Xiaojing Zhou, Zihan Ye, Xiuyun Mo
ACM Multimedia5