EDBT 2026 Demo / reviewers in the wild / expert
Yucan Zhou
dblp:146/8406
· DBLP profile ↗
40ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0002-1316-9118ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 18 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Federated Class-Incremental Learning via Spatial-Temporal Statistics AggregationabstractThe growing presence of mobile and IoT devices has led to massive decentralized and evolving data, driving the rise of Federated Learning (FL) to enable collaborative training without data sharing. However, traditional FL assumes static data distributions, which is unrealistic for dynamic real-world environments. To address this challenge, Federated Class-Incremental Learning (FCIL) has emerged as a promising framework that enables flexible adaptation to newly introduced classes over time. Existing FCIL methods typically integrate old knowledge preservation into local client training. However, these methods cannot avoid spatial-temporal client drift caused by data heterogeneity and often incur significant computational and communication overhead, limiting practical deployment. To address these challenges simultaneously, we propose a novel approach, Spatial-Temporal Statistics Aggregation (STSA), which provides a unified framework to aggregate feature statistics both spatially (across clients) and temporally (across stages). The aggregated feature statistics are unaffected by data heterogeneity and can be used to update the classifier in closed form at each stage. Additionally, we introduce STSA-E, a communication-efficient variant that enables the server to approximate global second-order feature statistics using first-order statistics uploaded from clients. Theoretical analysis shows that it achieves similar performance to STSA with much lower communication overhead. Extensive experiments on three widely used FCIL datasets, with varying degrees of data heterogeneity, show that our method outperforms state-of-the-art FCIL methods in terms of performance, flexibility, and both communication and computation efficiency. The code is available at https://github.com/Yuqin-G/STSA. Zenghao Guan, Guojun Zhu, Yucan Zhou, Wu Liu 0005, Weiping Wang 0005, Jiebo Luo 0001, Xiaoyan Gu 0001 |
WWW | 3 |
| 2026 | Decoupling and alleviating knowledge bias for federated learning under data heterogeneity
Yucan Zhou, Hong Zhao 0002, Xiaoyan Gu 0001 |
Knowl. Based Syst. | 2 |
| 2025 | Capture Global Feature Statistics for One-Shot Federated LearningabstractTraditional Federated Learning (FL) necessitates numerous rounds of communication between the server and clients, posing significant challenges including high communication costs, connection drop risks and susceptibility to privacy attacks. One-shot FL has become a compelling learning paradigm to overcome above drawbacks by enabling the training of a global server model via a single communication round. However, existing one-shot FL methods suffer from expensive computation cost on the server or clients and cannot deal with non-IID (Independent and Identically Distributed) data stably and effectively. To address these challenges, this paper proposes FedCGS, a novel Federated learning algorithm that Capture Global feature Statistics leveraging pre-trained models. With global feature statistics, we achieve training-free and heterogeneity-resistant one-shot FL. Furthermore, we expand its application to personalization scenario, where clients only need execute one extra communication round with server to download global statistics. Extensive experimental results demonstrate the effectiveness of our methods across diverse data-heterogeneity settings. Zenghao Guan, Yucan Zhou, Xiaoyan Gu 0001 |
AAAI | 2 |
| 2025 | Diversity-Enhanced Distribution Alignment for Dataset Distillation
Yucan Zhou, Xiaoyan Gu 0001, Bo Li 0063, Weiping Wang 0005 |
ICCV | 2 |
| 2025 | Take What I Need: Active Data Distillation for Federated LearningabstractFederated Learning (FL) addresses data privacy and security in distributed machine learning. However, it performs poorly in real-world applications due to non-IID client data. Fortunately, recent advancements in data distillation offer a solution to generate a compressed and privacy-preserving distilled dataset that approximates the original dataset. Nonetheless, the current data distillation for FL usually indiscriminately distills information from all the samples through multiple communication rounds, which will inevitably generate redundant information and reduce the convergence speed. In this paper, we propose a novel local active data distillation method, which enables the global model to actively distill data that are beneficial for its training from each local dataset. Subsequently, we design a selective knowledge distillation to refine the global model by aligning it with an accurate expert model. Experimental results on multiple datasets have demonstrated the effectiveness of our method compared with the state-of-the-art methods. Yucan Zhou, Xiaoyan Gu 0001, Bo Li 0063, Weiping Wang 0005 |
ICME | 2 |
| 2025 | FLR: Feature-based Label Recovery in Federated Learning with Classifier-free CommunicationabstractFederated learning has been recently threatened by gradient inversion attacks, which reconstruct clients’ data from shared gradients. Among them, label recovery is proposed to restore training label distribution with the relationship between training labels and classifier gradients. However, these label recovery attacks are inapplicable for algorithms where classifier gradients are not shared. In this paper, we propose Feature-based Label Recovery (FLR) to restore label distribution without classifier gradients. Specifically, the honest-but-curious server simulates training process on an auxiliary dataset using various label distributions to collect feature anchors of a proxy dataset. It learns the relationship between the feature anchor and training label distribution by training a meta-model, which is then utilized to predict client’s label distribution. Importantly, FLR performs well even when the auxiliary and proxy dataset are random noise. Extensive experiments show that FLR can accurately restore the label distribution. Yucan Zhou, Xiaoyan Gu 0001, Weiping Wang 0005 |
ICME | 2 |
| 2025 | Teach Structure Features to Cooperate with Node Embeddings in Link PredictionabstractStructure-enhanced models get a leading performance on the link prediction task as they utilize selected structure features and Graph Neural Network (GNN) based node embeddings simultaneously. However, we observe that when graphs get sparser, these methods perform worse than classical GNN-based methods, which has a severe impact on their practical use. We prove that when the graph gets sparser, the distance between structure features gets smaller. We induce this is the underlying reason for hindering the model from giving a reasonable prediction and leading to performance degeneration. To overcome this problem, we first claim that models need to learn the importance of node embeddings based on the distance between structure features. However, there is a lack of research to efficiently estimate the distance information, and existing models fail to assign the importance properly. Then we design a method called DIP, which satisfies the relation requirement with a weighted term of node embeddings, and use node degree to estimate the distance information to get the weight. Experimental results show that DIP can significantly improve the accuracy of structure-enhanced link prediction models and solve the performance degeneration problem effectively. The code of DIP is publicly available at https://github.com/lzwqbh/DIP. Feifei Dai, Yucan Zhou, Haihui Fan, Xiaoyan Gu 0001, Dan Meng 0002 |
IJCNN | 3 |
| 2025 | Diverse and Public Features Cooperation via Gradient Rectification for Federated Prompt LearningabstractFederated Prompt Learning (FPL) efficiently alleviates data heterogeneity and reduces communication costs by introducing pre-trained models and prompt tuning. However, local prompts tend to favor diverse features and ignore public features captured under extreme data heterogeneity, which compromises the generalization ability of the global prompt by only aggregating local prompts. To address this challenge, we present Federated Prompt Learning with Gradient Rectification (FedGR), which modifies the gradient directions of local and global prompts to enhance the generalization of the global prompt. Specifically, we first introduce a zero-shot prompt as public knowledge and constrain the gradient of local prompts to consistently deviate from the public feature space to capture diverse features adequately. Then, we compute the angular bisector of local and zero-shot prompt gradients and replace the gradient of the global prompt with the gradient of the angular bisector to capture both diverse features and public features. Finally, the server-side global prompt can enhance generalization by aggregating all client-side global prompts. Extensive experiments with various types of heterogeneities have demonstrated that our FedGR outperforms the state-of-the-art methods. Yucan Zhou, XingYou Yang, Xiaoyan Gu 0001 |
ACM Multimedia | 2 |
| 2025 | Statistics Caching Test-Time Adaptation for Vision-Language ModelsabstractTest-time adaptation (TTA) for Vision-Language Models (VLMs) aims to enhance performance on unseen test data. However, existing methods struggle to achieve robust and continuous knowledge accumulation during test time. To address this, we propose Statistics Caching test-time Adaptation (SCA), a novel cache-based approach. Unlike traditional feature-caching methods prone to forgetting, SCA continuously accumulates task-specific knowledge from all encountered test samples. By formulating the reuse of past features as a least squares problem, SCA avoids storing raw features and instead maintains compact, incrementally updated feature statistics. This design enables efficient online adaptation without the limitations of fixed-size caches, ensuring that the accumulated knowledge grows persistently over time. Furthermore, we introduce adaptive strategies that leverage the VLM's prediction uncertainty to reduce the impact of noisy pseudo-labels and dynamically balance multiple prediction sources, leading to more robust and reliable performance. Extensive experiments demonstrate that SCA achieves compelling performance while maintaining competitive computational efficiency. Zenghao Guan, Yucan Zhou, Wu Liu 0005, Xiaoyan Gu 0001 |
NeurIPS | 2 |
| 2025 | LaAeb: A comprehensive log-text analysis based approach for insider threat detection
Kexiong Fei, Yucan Zhou, Xiaoyan Gu 0001, Haihui Fan, Bo Li 0063, Weiping Wang 0005, Yong Chen 0001 |
Comput. Secur. | 3 |
| 2025 | Adaptive Diversity Induced Reweighting for long-tailed classification
Xiaohua Chen 0002, Yucan Zhou, Haihui Fan, Qinghang Su, Weiping Wang 0005 |
Neural Networks | 2 |
| 2024 | Meta-Knowledge Enhanced Data Augmentation for Federated Person Re-Identificationabstractfederated learning has been introduced into person re-identification (Re-ID) to avoid personal image leakage in traditional centralized training. To address the key issue of statistic heterogeneity in different clients, several optimization methods have been proposed to alleviate the bias of the local models. However, besides statistic heterogeneity, feature heterogeneity (e.g., various angles, different illuminations) in different clients is more challenging in federated Re-ID. In this paper, we propose a meta-knowledge enhanced data augmentation method, where the global cross semantic feature transformations are provided to each client to perform local infinite augmentation to reduce the feature difference in different clients. Specifically, to capture the cross semantic feature transformations in each client, we calculate the covariance matrix of features with the local dataset as the transferable meta-knowledge. Then, this local meta-knowledge is propagated to the server for global aggregation. Subsequently, the aggregated meta-knowledge is sent back to each client for infinite data augmentation. Moreover, since the covariance matrix indicates variations in a client, we design a variation-balanced aggregation to replace the traditional data-size-balanced aggregation. To imitate the more challenging scenario of feature heterogeneity, we focus on the federated-by-camera setting to conduct experiments, where images collected in a camera are regarded as the dataset of a client. Extensive experimental results show that our method outperforms other state-of-the-art methods. Code is available at https://github.com/songchunli1999/MEDA. Chunli Song, Xiaohua Chen 0002, Wenqiu Zhu, Yucan Zhou, Xiaoyan Gu 0001, Bo Li 0063 |
ICASSP | 4 |
| 2024 | GIE : Gradient Inversion with EmbeddingsabstractFederated Learning (FL) has emerged as a promising approach to preserve privacy for distributed machine learning, through aggregating gradients calculated from local data of multiple clients. Recent studies show that gradients can be inverted to reconstruct private data retained by local clients in a process called gradient inversion. However, the performance of existing gradient inversion attacks declines as batch size increases. In this work, we propose a simple but effective method to tackle this challenge. Our method utilizes an interesting property between weights, embeddings and their corresponding gradients in fully connected layer, which enables the rapid and accurate recovery of embeddings. And these embeddings benefit the recovery of private data as additional prior knowledge. Extensive experiments demonstrate that our method can perform high-quality recovery of private data. Zenghao Guan, Yucan Zhou, Xiaoyan Gu 0001, Bo Li 0063 |
ICME | 2 |
| 2024 | Tackling Feature Skew in Heterogeneous Federated Learning with Semantic EnhancementabstractA critical challenge in federated learning is data heterogeneity, compounded by varying local data distribution, thereby significantly impacting the performance of both local and global models. Prior works struggle to effectively address data heterogeneity but ignore the feature skew. In this paper, we propose a novel method, Federated Learning with Semantic Enhancement (FedSE) to address this challenge. Specifically, we introduce semantic regularization terms and adaptive local aggregation to mitigate the drawback of local knowledge insufficiency and enhance the semantics of representations. Then, based on the neural collapse theory, we initialize a simplex equiangular tight frame structure (ETF) of the classifier and maintain its stability during local training, to achieve alignment in the feature space across different clients. Extensive experiments demonstrate our method achieves SOTA performance on several standard benchmark datasets, effectively alleviating the feature skew. Yucan Zhou, Xiaoyan Gu 0001, Bo Li 0063 |
ICME | 2 |
| 2024 | Noise-Aware Person Re-identification via Local Uncertainty EstimationabstractPerson re-identification aims to retrieve the target person from a large-scale dataset, and its performance highly relies on a large, well-annotated training dataset. However, label noise generated by manual mislabeling is usually inevitable in real-world scenarios. An intuitive way to alleviate the effect of label noise is to identify and discard noisy data. Recently, most works recognize noisy samples with their global features, which will omit some noisy samples when they are subtly different from the clean ones. In this paper, we propose to recognize noisy data with local features based on uncertainty estimation. Specifically, for each input image, we randomly erase a part of its features to generate a local feature and propagate it to the softmax layer to obtain a local soft label. Then, we repeat the above operations to obtain multiple local soft labels. Subsequently, we estimate the uncertainty with these multiple local soft labels and utilize them to identify whether this sample is noisy. Finally, we design a noise-aware cross-entropy loss based on the estimated uncertainty to reduce the effect of noisy samples automatically. Experimental results on Market-1501 and CUHK03 datasets show the effectiveness of our proposed model, where at least 4.8% and 1.5% improvements are achieved compared with other methods under 0% and 10% noise settings on Market-1501 datasets. Code is available at https://github.com/songchunli1999/NACE. Chunli Song, Yucan Zhou, Wenqiu Zhu, Xiaoyan Gu 0001, Bo Li 0063 |
IJCNN | 2 |
| 2024 | Diversified Semantic Distribution Matching for Dataset DistillationabstractDataset distillation, also known as dataset condensation, offers a possibility for compressing a large-scale dataset into a small-scale one (i.e., distilled dataset) while achieving similar performance during model training. This method effectively tackles the challenges of training efficiency and storage cost posed by the large-scale dataset. Existing dataset distillation methods can be categorized into Optimization-Oriented (OO)-based and Distribution-Matching (DM)-based methods. Since OO-based methods require bi-level optimization to alternately optimize the model and the distilled data, they face challenges due to high computational overhead in practical applications. Thus, DM-based methods have emerged as an alternative by aligning the prototypes of the distilled data to those of the original data. Although efficient, these methods overlook the diversity of the distilled data, which will limit the performance of evaluation tasks. In this paper, we propose a novel Diversified Semantic Distribution Matching (DSDM) approach for dataset distillation. To accurately capture semantic features, we first pre-train models for dataset distillation. Subsequently, we estimate the distribution of each category by calculating its prototype and covariance matrix, where the covariance matrix indicates the direction of semantic feature transformations for each category. Then, in addition to the prototypes, the covariance matrices are also matched to obtain more diversity for the distilled data. However, since the distilled data are optimized by multiple pre-trained models, the training process will fluctuate severely. Therefore, we match the distilled data of the current pre-trained model with the historical integrated prototypes. Experimental results demonstrate that our DSDM achieves state-of-the-art results on both image and speech datasets. Code is available at https://github.com/Li-Hongcheng/DSDM. Yucan Zhou, Xiaoyan Gu 0001, Bo Li 0063, Weiping Wang 0005 |
ACM Multimedia | 2 |
| 2023 | AREA: Adaptive Reweighting via Effective Area for Long-Tailed ClassificationabstractLarge-scale data from the real-world usually follow a long-tailed distribution (i.e., a few majority classes occupy plentiful training data, while most minority classes have few samples), making the hyperplanes heavily skewed to the minority classes. Traditionally, reweighting is adopted to make the hyperplanes fairly split the feature space, where the weights are designed according to the number of samples. However, we find that the number of samples in a class can not accurately measure the size of its spanned space, especially for the majority class, where the size of its spanned space is usually larger than the samples’ number because of the high diversity. Therefore, weights designed based on the samples’ number will still compress the space of minority classes. In this paper, we reconsider reweighting from a totally new perspective of analyzing the spanned space of each class. We argue that, besides statistical numbers, relations between samples are also significant for sufficiently depicting the spanned space. Consequently, we estimate the size of the spanned space for each category, namely effective area, by detailedly analyzing its samples’ distribution. By treating samples of a class as identically distributed random variables and analyzing their correlations, a simple and non-parametric formula is derived to estimate the effective area. Then, the weight simply calculated inversely proportional to the effective area of each class is adopted to achieve fairer training. Note that our weights are more flexible as they can be adaptively adjusted along with the optimizing features during training. Experiments on four long-tailed datasets show that the proposed weights outperform the state-of-the-art reweighting methods. Moreover, our method can also achieve better results on statistically balanced CIFAR-10/100. Code is available at https://github.com/xiaohua-chen/AREA. Xiaohua Chen 0002, Yucan Zhou, Dayan Wu, Chule Yang, Bo Li 0063, Qinghua Hu, Weiping Wang 0005 |
ICCV | 2 |
| 2023 | Preserving Potential Neighbors for Low-Degree Nodes via Reweighting in Link Prediction
Yucan Zhou, Haihui Fan, Xiaoyan Gu 0001, Bo Li 0063, Dan Meng 0002 |
ICONIP (2) | 2 |
| 2023 | Decoupled Contrastive Learning for Long-Tailed Distribution
Xiaohua Chen 0002, Yucan Zhou, Lin Wang 0108, Dayan Wu, Wanqian Zhang, Bo Li 0063, Weiping Wang 0005 |
PRCV (9) | 2 |
| 2022 | Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed ClassificationabstractReal-world data often follows a long-tailed distribution, which makes the performance of existing classification algorithms degrade heavily. A key issue is that the samples in tail categories fail to depict their intra-class diversity. Humans can imagine a sample in new poses, scenes and view angles with their prior knowledge even if it is the first time to see this category. Inspired by this, we propose a novel reasoning-based implicit semantic data augmentation method to borrow transformation directions from other classes. Since the covariance matrix of each category represents the feature transformation directions, we can sample new directions from similar categories to generate definitely different instances. Specifically, the long-tailed distributed data is first adopted to train a backbone and a classifier. Then, a covariance matrix for each category is estimated, and a knowledge graph is constructed to store the relations of any two categories. Finally, tail samples are adaptively enhanced via propagating information from all the similar categories in the knowledge graph. Experimental results on CIFAR-LT-100, ImageNet-LT, and iNaturalist 2018 have demonstrated the effectiveness of our proposed method compared with the state-of-the-art methods. Xiaohua Chen 0002, Yucan Zhou, Dayan Wu, Wanqian Zhang, Yu Zhou 0015, Bo Li 0063, Weiping Wang 0005 |
AAAI | 2 |
| 2022 | Deep collaborative multi-task network: A human decision process inspired model for hierarchical image classification
Yu Zhou 0015, Xiaoni Li, Yucan Zhou, Yu Wang 0106, Qinghua Hu, Weiping Wang 0005 |
Pattern Recognit. | 3 |
| 2022 | Uncertainty-Aware and Multigranularity Consistent Constrained Model for Semi-Supervised HashingabstractRecently, deep semi-supervised hashing methods have attracted increasing attention, which can significantly improve retrieval performance by leveraging abundant unlabeled data. These methods usually generate surrogate supervision signals to learn with unlabeled data, such as neighborhood information and augmentation invariant requirements. However, an essential issue of these methods is that the supervised signals are not always reliable, which may damage the performance. In this paper, we propose a novel Uncertainty-Aware and Multi-Granularity Consistent Constrained Semi-Supervised Hashing (UMCSH) method to alleviate the negative effects of noisy supervised signals and enlarge the inter-class distance. Specifically, our UMCSH mainly consists of an Uncertainty-Aware Instance-Level Consistency (UAILC) model and a Cluster-Based Class-Level Consistency (CBCLC) model. UAILC introduces an uncertainty estimation method to select reliable supervised signals to extract discriminative features for each unlabeled data. CBCLC establishes connections between labeled data and unlabeled data by encouraging each unlabeled sample to be close to the hash center (calculated with the labeled data) according to its pseudo-label. Extensive experimental results demonstrate the superior performance of our proposed approach compared with several state-of-the-art semi-supervised hashing methods. Shuai Cheng 0002, Yucan Zhou, Wanqian Zhang, Dayan Wu, Chule Yang, Bo Li 0063, Weiping Wang 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Hierarchical Semantic Risk Minimization for Large-Scale ClassificationabstractHierarchical structures of labels usually exist in large-scale classification tasks, where labels can be organized into a tree-shaped structure. The nodes near the root stand for coarser labels, while the nodes close to leaves mean the finer labels. We label unseen samples from the root node to a leaf node, and obtain multigranularity predictions in the hierarchical classification. Sometimes, we cannot obtain a leaf decision due to uncertainty or incomplete information. In this case, we should stop at an internal node, rather than going ahead rashly. However, most existing hierarchical classification models aim at maximizing the percentage of correct predictions, and do not take the risk of misclassifications into account. Such risk is critically important in some real-world applications, and can be measured by the distance between the ground truth and the predicted classes in the class hierarchy. In this work, we utilize the semantic hierarchy to define the classification risk and design an optimization technique to reduce such risk. By defining the conservative risk and the precipitant risk as two competing risk factors, we construct the balanced conservative/precipitant semantic (BCPS) risk matrix across all nodes in the semantic hierarchy with user-defined weights to adjust the tradeoff between two kinds of risks. We then model the classification process on the semantic hierarchy as a sequential decision-making task. We design an algorithm to derive the risk-minimized predictions. There are two modules in this model: 1) multitask hierarchical learning and 2) deep reinforce multigranularity learning. The first one learns classification confidence scores of multiple levels. These scores are then fed into deep reinforced multigranularity learning for obtaining a global risk-minimized prediction with flexible granularity. Experimental results show that the proposed model outperforms state-of-the-art methods on seven large-scale classification datasets with the semantic tree. Yu Wang 0106, Zhou Wang 0001, Qinghua Hu, Yucan Zhou, Honglei Su |
IEEE Trans. Cybern. | 4 |
| 2022 | Exploring Relations in Untrimmed Videos for Self-Supervised LearningabstractExisting video self-supervised learning methods mainly rely on trimmed videos for model training. They apply their methods and verify the effectiveness on trimmed video datasets including UCF101 and Kinetics-400, among others. However, trimmed datasets are manually annotated from untrimmed videos. In this sense, these methods are not truly unsupervised. In this article, we propose a novel self-supervised method, referred to as Exploring Relations in Untrimmed Videos (ERUV), which can be straightforwardly applied to untrimmed videos (real unlabeled) to learn spatio-temporal features. ERUV first generates single-shot videos by shot change detection. After that, some designed sampling strategies are used to model relations for video clips. The strategies are saved as our self-supervision signals. Finally, the network learns representations by predicting the category of relations between the video clips. ERUV is able to compare the differences and similarities of video clips, which is also an essential procedure for video-related tasks. We validate our learned models with action recognition, video retrieval, and action similarity labeling tasks with four kinds of 3D convolutional neural networks. Experimental results show that ERUV is able to learn richer representations with untrimmed videos, and it outperforms state-of-the-art self-supervised methods with significant margins. Dezhao Luo, Yu Zhou 0015, Bo Fang 0003, Yucan Zhou, Dayan Wu, Weiping Wang 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | MMF: Multi-task Multi-structure Fusion for Hierarchical Image Classification
Xiaoni Li, Yucan Zhou, Yu Zhou 0015, Weiping Wang 0005 |
ICANN (4) | 2 |
| 2021 | Disturbance Consistent Self-Ensembling for Semi-Supervised HashingabstractRecently, deep semi-supervised hashing methods have attracted increasing attention, where the visual similarity of unlabeled data is usually adopted to guide the hash codes learning. However, samples with similar appearance may come from different categories, making their hash codes similar will lead to sub-optimal retrieval results. In this paper, we propose a novel Disturbance Consistent Self-Ensembling (DCSE) method to alleviate the drawback of visual similarity constraint. Specially, DCSE forms consensus hash codes for the same sample under different augmentations. These ensemble hash codes can capture the discriminative characteristics of a sample. Therefore, as more augmented data is involved, more ensemble hash codes in one category can become similar gradually. Then, we design a disturbance consistent loss to learn the discriminative hash codes by minimizing the distance between outputs of the hash layer and ensemble hash codes. Extensive experiments show that our proposed approach significantly outperforms state-of-the-art semi-supervised hashing methods. Shuai Cheng 0002, Yucan Zhou, Dayan Wu, Haisu Zhang, Bo Li 0063, Weiping Wang 0005 |
ICME | 2 |
| 2021 | What Matters: Attentive and Relational Feature Aggregation Network for Video-Text RetrievalabstractCross-modal video-text retrieval has been an emerging task due to the rapid growth of user-generated videos on the Internet. Most existing approaches focus on extracting visual feature for the video, while audio and caption on the screen containing rich information are ignored. Recently, the aggregations of multi-modal features in videos boost the benchmark of video-text retrieval. However, since these multi-modal features are high-dimensional and heterogeneous, their intrinsically structural relations have not been attached with enough importance and are often overlooked in previous methods. To address this issue, we propose a novel Attentive and Relational Feature Aggregation Network (ARFAN). Specifically, we introduce the self-attention mechanism to make videos adaptively assign higher weights to the representative modalities. Then, the graph convolutional layers are inserted to capture the relations among the multi-modal features to combine them. Our method achieves 15% and 12.9% relative improvements on R@1 when compared with the state-of-the-art method on MSR-VTT and MSVD datasets, respectively. Xiaoshuai Hao, Yucan Zhou, Dayan Wu, Wanqian Zhang, Bo Li 0063, Weiping Wang 0005, Dan Meng 0002 |
ICME | 2 |
| 2021 | Rescuing Deep Hashing from Dead Bits ProblemabstractDeep hashing methods have shown great retrieval accuracy and efficiency in large-scale image retrieval. How to optimize discrete hash bits is always the focus in deep hashing methods. A common strategy in these methods is to adopt an activation function, e.g. sigmoid() or tanh(), and minimize a quantization loss to approximate discrete values. However, this paradigm may make more and more hash bits stuck into the wrong saturated area of the activation functions and never escaped. We call this problem "Dead Bits Problem (DBP)". Besides, the existing quantization loss will aggravate DBP as well. In this paper, we propose a simple but effective gradient amplifier which acts before activation functions to alleviate DBP. Moreover, we devise an error-aware quantization loss to further alleviate DBP. It avoids the negative effect of quantization loss based on the similarity between two images. The proposed gradient amplifier and error-aware quantization loss are compatible with a variety of deep hashing methods. Experimental results on three datasets demonstrate the efficiency of the proposed gradient amplifier and the error-aware quantization loss. Shu Zhao 0006, Dayan Wu, Yucan Zhou, Bo Li 0063, Weiping Wang 0005 |
IJCAI | 3 |
| 2021 | Multi-Feature Graph Attention Network for Cross-Modal Video-Text RetrievalabstractCross-modal retrieval between videos and texts has attracted growing attention due to the rapid growth of user-generated videos on the web. To solve this problem, most approaches try to learn a joint embedding space to measure the cross-modal similarities, while paying little attention to the representation of each modality. Video is more complicated than the commonly used visual feature, since the audio and caption on the screen also contain rich information. Recently, the aggregations of multiple features in videos boost the benchmark of the video-text retrieval system. However, they usually handle each feature independently, which ignores the interchange of high-level semantic relations among these multiple features. Moreover, despite the inter-modal ranking constraint where semantically-similar texts and videos should stay closer, the modality-specific requirement, i.e. two similar videos/texts should have similar representations, is also significant. In this paper, we propose a novel Multi-Feature Graph ATtention Network (MFGATN) for cross-modal video-text retrieval. Specifically, we introduce a multi-feature graph attention module, which enriches the representation of each feature in videos with the interchange of high-level semantic information among them. Moreover, we elaborately design a novel Dual Constraint Ranking Loss (DCRL), which simultaneously considers the inter-modal ranking constraint and the intra-modal structure constraint to preserve both the cross-modal semantic similarity and the modality-specific consistency in the embedding space. Experiments on two datasets, i.e. MSR-VTT and MSVD, demonstrate that our method achieves significant performance gain compared with the state-of-the-arts. Xiaoshuai Hao, Yucan Zhou, Dayan Wu, Wanqian Zhang, Bo Li 0063, Weiping Wang 0005 |
ICMR | 2 |
| 2021 | Robust hierarchical feature selection driven by data and knowledge
Xinxin Liu 0011, Yucan Zhou, Hong Zhao 0002 |
Inf. Sci. | 2 |
| 2020 | SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text RecognitionabstractScene text recognition is a hot research topic in computer vision. Recently, many recognition methods based on the encoder-decoder framework have been proposed, and they can handle scene texts of perspective distortion and curve shape. Nevertheless, they still face lots of challenges like image blur, uneven illumination, and incomplete characters. We argue that most encoder-decoder methods are based on local visual features without explicit global semantic information. In this work, we propose a semantics enhanced encoder-decoder framework to robustly recognize low-quality scene texts. The semantic information is used both in the encoder module for supervision and in the decoder module for initializing. In particular, the state-of-the-art ASTER method is integrated into the proposed framework as an exemplar. Extensive experiments demonstrate that the proposed framework is more robust for low-quality text images, and achieves state-of-the-art results on several benchmark datasets. The source code will be available. Yu Zhou 0015, Dongbao Yang, Yucan Zhou, Weiping Wang 0005 |
CVPR | 4 |
| 2018 | Deep super-class learning for long-tail distributed image classification
Yucan Zhou, Qinghua Hu, Yu Wang 0106 |
Pattern Recognit. | 1 |
| 2018 | Kernel-Based Semantic Hashing for Gait RetrievalabstractIt is very important to retrieve a specific person in locating and tracking the missing people as well as the suspects quickly. However, the well-studied face-based and appearance-based individual retrieval methods are ineffective in the surveillance scenarios because of the far photograph distances, the low camera resolutions, the long time intervals, and the complex lighting conditions. To avoid the disadvantages of face-based and appearance-based methods, we propose to retrieve individuals from the surveillance videos with the gait biometric, which has been proved to be beneficial to remote person recognition and robust to lighting variations. What's more, the gait biometric can be collected without conscious cooperation, making the data collection much easier. But it varies greatly with the view angles, the clothing style, and the carrying conditions. Therefore, the videos of the target person from a similar view angle with the same clothing style and carrying conditions should rank higher than the others. To achieve this purpose and improve the efficiency, this paper proposes a kernel-based semantic hashing (KSH) model, which is learnt by optimizing a semantic triplet ranking loss. Specifically, in the training phase, a semantic similarity score, which depends on the view angles, the clothing style, and the carrying conditions, is calculated for each training pair. Then, a weighted triplet loss considering these semantic scores is designed, which encourages videos with a higher score to stay closer to the gallery in the binary Hamming space. To evaluate the performance of the proposed method, we compare it with several methods on the CASIA Gait Database B and the OU-ISIR Gait Database. The experimental results demonstrate that the KSH is effective and efficient. Yucan Zhou, Yongzhen Huang, Qinghua Hu, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Large-Scale Multimodality Attribute Reduction With Multi-Kernel Fuzzy Rough SetsabstractIn complex pattern recognition tasks, objects are typically characterized by means of multimodality attributes, including categorical, numerical, text, image, audio, and even videos. In these cases, data are usually high dimensional, structurally complex, and granular. Those attributes exhibit some redundancy and irrelevant information. The evaluation, selection, and combination of multimodality attributes pose great challenges to traditional classification algorithms. Multikernel learning handles multimodality attributes by using different kernels to extract information coming from different attributes. However, it cannot consider the aspects fuzziness in fuzzy classification. Fuzzy rough sets emerge as a powerful vehicle to handle fuzzy and uncertain attribute reduction. In this paper, we design a framework of multimodality attribute reduction based on multikernel fuzzy rough sets. First, a combination of kernels based on set theory is defined to extract fuzzy similarity for fuzzy classification with multimodality attributes. Then, a model of multikernel fuzzy rough sets is constructed. Finally, we design an efficient attribute reduction algorithm for large scale multimodality fuzzy classification based on the proposed model. Experimental results demonstrate the effectiveness of the proposed model and the corresponding algorithm. Qinghua Hu, Lingjun Zhang, Yucan Zhou, Witold Pedrycz |
IEEE Trans. Fuzzy Syst. | 3 |
| 2017 | Local Bayes Risk Minimization Based Stopping Strategy for Hierarchical ClassificationabstractIn large-scale data classification tasks, it is becoming more and more challenging in finding a true class from a huge amount of candidate categories. Fortunately, a hierarchical structure usually exists in these massive categories. The task of utilizing this structure for effective classification is called hierarchical classification. It usually follows a top-down fashion which predicts a sample from the root node with a coarse-grained category to a leaf node with a fine-grained category. However, misclassification is inevitable if the information is insufficient or large uncertainty exists in the prediction process. In this scenario, we can design a stopping strategy to stop the sample at an internal node with a coarser category, instead of predicting a wrong leaf node. Several studies address the problem by improving performance in terms of hierarchical accuracy and informative prediction. However, all of these researches ignore an important issue: when predicting a sample at the current node, the error is inclined to occur if large uncertainty exists in the next lower level children nodes. In this paper, we integrate this uncertainty into a risk problem: when predicting a sample at a decision node, it will take precipitance risk in predicting the sample to a children node in the next lower level on one hand, and take conservative risk in stopping at the current node on the other. We address the risk problem by designing a Local Bayes Risk Minimization (LBRM) framework, which divides the prediction process into recursively deciding to stop or to go down at each decision node by balancing these two risks in a top-down fashion. Rather than setting a global loss function in the traditional Bayes risk framework, we replace it with different uncertainty in the two risks for each decision node. The uncertainty on the precipitance risk and the conservative risk are measured by information entropy on children nodes and information gain from the current node to children nodes, respectively. We propose a Weighted Tree Induced Error (WTIE) to obtain the predictions of minimum risk with different emphasis on the two risks. Experimental results on various datasets show the effectiveness of the proposed LBRM algorithm. Yu Wang 0106, Qinghua Hu, Yucan Zhou, Hong Zhao 0002, Jiye Liang |
ICDM | 3 |
| 2017 | A Pixel-to-Pixel Convolutional Neural Network for Single Image Dehazing
Chengkai Zhu, Yucan Zhou, Zongxia Xie |
ICONIP (3) | 2 |
| 2015 | Heterogeneous Features Integration via Semi-supervised Multi-modal Deep Networks
Qinghua Hu, Yucan Zhou |
ICONIP (4) | 3 |
| 2015 | Improved multi-kernel SVM for multi-modal and imbalanced dialogue act classificationabstractDialogue act recognition is recognized as an important step for computers to understand human dialogues as it is closely related to the human intention. There are two main challenges in dialogue act recognition. Firstly, multimodal features should be taken into consideration, which include lexical, syntactic, prosodic cues, even facial appearance and gesture. Secondly, samples distribution in the dialogue act corpus is highly imbalanced. Thus traditional classification algorithms produce poor performance when they are applied on these imbalanced multi-modal tasks. In this paper, the multi-kernel SVM model is investigated to deal with these problems. Multi-kernel SVM is an effective technique for leaning from multi-modal data, but it is sensitive to imbalance. So an improved multi-kernel SVM model is proposed. To show the effectiveness of the proposed model, we test it on some open classification tasks and a Chinese dialogue act recognition task. Significant improvements are observed from the experimental results. Yucan Zhou, Xiaowei Cui, Qinghua Hu, Yuan Jia |
IJCNN | 1 |
| 2015 | Combining heterogeneous deep neural networks with conditional random fields for Chinese dialogue act recognition
Yucan Zhou, Qinghua Hu, Jie Liu 0007, Yuan Jia |
Neurocomputing | 1 |
| 2014 | Learning conditional random field with hierarchical representations for dialogue act recognition
Yucan Zhou, Qinghua Hu, Jie Liu 0007, Yuan Jia |
INTERSPEECH | 1 |