VLDB 2026 Research / reviewers in the wild / expert
Yi Zhang 0113
dblp:64/6544-113
· DBLP profile ↗
14ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0003-0843-7170ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Noisy label effects for out-of-distribution detection in single-positive multi-label settings
Yi Zhang 0113, Xiaohong Zhang 0002, Dan Yang 0001, Sheng Huang 0001 |
Pattern Anal. Appl. | 2 |
| 2025 | Learning complementary visual information for few-shot food recognition by Regional Erasure and Reactivation
Yi Zhang 0113, Luwen Huangfu, Lili Balazs, Sheng Huang 0001 |
Expert Syst. Appl. | 1 |
| 2025 | Rethinking the sample relations for few-shot classification
Guowei Yin, Sheng Huang 0001, Luwen Huangfu, Yi Zhang 0113, Xiaohong Zhang 0002 |
Image Vis. Comput. | 4 |
| 2025 | Adaptive learning of instance representatives in dual spaces for medical image classification
Sheng Huang 0001, Yi Zhang 0113, Xiaoxian Zhang, Chen Liu 0026, Xiahong Zhang |
Neural Comput. Appl. | 4 |
| 2025 | Learning the Difference of Few-Shot Food Data Using Multivariate Knowledge-Guided Variational AutoencoderabstractRecent advancements in food image recognition have underscored its importance in dietary monitoring, which promotes a healthy lifestyle and aids in the prevention of diseases such as diabetes and obesity. While mainstream food recognition methods excel in scenarios with large-scale annotated datasets, they falter in few-shot regimes where data is limited. This paper addresses this challenge by introducing a variational generative method, the Multivariate Knowledge-guided Variational AutoEncoder (MK-VAE), for few-shot food recognition. MK-VAE leverages handcrafted features and semantic embeddings as multivariate prior knowledge to strengthen feature learning and feature generation in different phases. Specifically, we design a lightweight and flexible feature distillation module that distills handcrafted features to enhance the feature learning network for capturing the salient visual information in few-shot samples. During the feature generation phase, we utilize a variational autoencoder to learn the difference distribution of food data and explicitly boost the latent representation with category-level semantic embeddings to pull homogeneous features closer together while pushing inhomogeneous features apart. Experimental results demonstrate that our proposed MK-VAE significantly outperforms state-of-the-art few-shot food recognition methods in both 5-way 1-shot and 5-way 5-shot settings on three widely-used benchmark datasets: Food-101, VIREO Food-172, and UECFood-256. Yi Zhang 0113, Sheng Huang 0001, Mingjian Hong, Dan Yang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Learning Feature Exploration and Selection With Handcrafted Features for Few-Shot LearningabstractInterest in few-shot learning (FSL) has grown recently, but the value of feature learning, which bridges the gap between base and novel classes, remains largely understudied. The limited availability of labeled samples for each class poses a major challenge. To tackle this, we propose a simple yet effective approach called deep discriminative handcrafted feature regression (DDHFR) to explore intrinsic information and select improved discriminative features in few-shot data by mining knowledge from classical handcrafted features. To explore intrinsic information, we design several deep handcrafted feature regression (DHFR) modules and plugged them separately into different layers of the backbone to use feature engineering knowledge for feature learning optimization at different granularities. To achieve discriminative feature selection, we incorporate an auxiliary classifier (AC) into each DHFR module to enhance the acquisition of discriminative information. Furthermore, we employed self-distillation to boost ability of ACs ot be classified. Experimental results in three backbones on three datasets show that DDHFR can generally improve the performance of existing FSL methods. On average, it improves the recognition accuracy by 1.16% in two common few-shot settings. Yi Zhang 0113, Sheng Huang 0001, Luwen Huangfu, Daniel Dajun Zeng |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyabstractMultiple instance learning (MIL) is the most widely used framework in computational pathology, encompassing sub-typing, diagnosis, prognosis, and more. However, the ex-isting MIL paradigm typically requires an offline instance feature extractor, such as a pre-trained ResNet or a foun-dation model. This approach lacks the capability for feature fine-tuning within the specific downstream tasks, limiting its adaptability and performance. To address this issue, we propose a Re-embedded Regional Transformer (R2T) for re-embedding the instance features online, which captures fine-grained local features and establishes connections across different regions. Unlike existing works that focus on pre-training powerful feature extractor or designing sophisticated instance aggregator, R2T is tailored to re-embed instance features online. It serves as a portable module that can seamlessly integrate into mainstream MIL models. Extensive experimental results on common computational pathology tasks validate that: 1) feature re-embedding improves the performance of MIL models based on ResNet-50 features to the level of foundation model features, and further enhances the performance of foundation model features; 2) the R2T can introduce more signifi-cant performance improvements to various MIL models; 3) R2T-MIL, as an R2T-enhanced AB-MIL, outperforms other latest methods by a large margin. The code is available at: https://github.com/DearCaat/RRT-MIL. Fengtao Zhou, Sheng Huang 0001, Yi Zhang 0113, Bo Liu 0005 |
CVPR | 5 |
| 2024 | Semi-Identical Twins Variational AutoEncoder for Few-Shot LearningabstractData augmentation is a popular way for few-shot learning (FSL). It generates more samples as supplements and then transforms the FSL task into a common supervised learning problem for a solution. However, most data-augmentation-based FSL approaches only consider the prior visual knowledge for feature generation, thereby leading to low diversity and poor quality of generated data. In this study, we attempt to address this issue by incorporating both prior visual and prior semantic knowledge to condition the feature generation process. Inspired by some genetic characteristics of semi-identical twins, a novel multimodal generative FSL approach was developed named semi-identical twins variational autoencoder (STVAE) to better exploit the complementarity of these modality information by considering the multimodal conditional feature generation process as a process that semi-identical twins are born and collaborate to simulate their father. STVAE conducts feature synthesis by pairing two conditional variational autoencoders (CVAEs) with the same seed but different modality conditions. Subsequently, the generated features of two CVAEs are considered as semi-identical twins and adaptively combined to yield the final feature, which is considered as their fake father. STVAE requires that the final feature can be converted back into its paired conditions while ensuring these conditions remain consistent with the original in both representation and function. Moreover, STVAE is able to work in the partial modality-absence case due to the adaptive linear feature combination strategy. STVAE essentially provides a novel idea to exploit the complementarity of different modality prior information inspired by genetics in FSL. Extensive experimental results demonstrate that our work achieves promising performances in comparison to the recent state-of-the-art approaches, as well as validate its effectiveness on FSL under various modality settings. Yi Zhang 0113, Sheng Huang 0001, Xi Peng 0005, Dan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Multiple Instance Learning Framework with Masked Hard Instance Mining for Whole Slide Image ClassificationabstractThe whole slide image (WSI) classification is often formulated as a multiple instance learning (MIL) problem. Since the positive tissue is only a small fraction of the gigapixel WSI, existing MIL methods intuitively focus on identifying salient instances via attention mechanisms. However, this leads to a bias towards easy-to-classify instances while neglecting hard-to-classify instances. Some literature has revealed that hard examples are beneficial for modeling a discriminative boundary accurately. By applying such an idea at the instance level, we elaborate a novel MIL framework with masked hard instance mining (MHIM-MIL), which uses a Siamese structure (Teacher-Student) with a consistency constraint to explore the potential hard instances. With several instance masking strategies based on attention scores, MHIM-MIL employs a momentum teacher to implicitly mine hard instances for training the student model, which can be any attention-based MIL model. This counter-intuitive strategy essentially enables the student to learn a better discriminating boundary. Moreover, the student is used to update the teacher with an exponential moving average (EMA), which in turn identifies new hard instances for subsequent training iterations and stabilizes the optimization. Experimental results on the CAMELYON-16 and TCGA Lung Cancer datasets demonstrate that MHIM-MIL outperforms other latest methods in terms of performance and training cost. The code is available at: https://github.com/DearCaat/MHIM-MIL. Sheng Huang 0001, Xiaoxian Zhang, Fengtao Zhou, Yi Zhang 0113, Bo Liu 0005 |
ICCV | 5 |
| 2023 | ASDFL: An adaptive super-pixel discriminative feature-selective learning for vehicle matchingabstractAbstract There are a large number of cameras in modern transportation system that capture numerous vehicle images continuously. Therefore, automatic analysis of these vehicle images is helpful for traffic flow management, criminal investigations and vehicle inspections. Vehicle matching, which aims to determine whether two input images depict an identical vehicle, is one of the core tasks in vehicle analysis. Recent relevant studies have focused on local feature extraction instead of global extraction, since local details can provide crucial cues to distinguish between cars. However, these methods do not select local features; that is, they do not assign weights to local features. Therefore, in this research, we systematically study the vehicle matching task, and present a novel annotation‐free local‐based deep learning method called Adaptive super‐pixel discriminative feature‐selective learning (ASDFL) to address this issue. In ASDFL, vehicle images are segmented into clusters of super‐pixels of similar size by considering the location and colour similarities of pixels without using any component‐level annotation. These super‐pixels are deemed to be the virtual components of vehicles. Moreover, a convolutional neural network is used to extract the deep features of these virtual components. Thereafter, an instance‐specific mask generation module driven by the extracted global features is enhanced to produce a mask to select the most distinctive virtual components of each vehicle image pair in the feature space. Finally, the vehicle matching task is accomplished by classifying the selected virtual component features of each imaged vehicle pair. Extensive experiments on two popular vehicle identification benchmarks demonstrate that our method is 1.57% and 0.8% more accurate than the previous baselines in a vehicle matching task on the VeRi and VehicleID datasets, respectively, which demonstrates the effectiveness of our method. Rong Qin 0001, Huanhuan Lv, Yi Zhang 0113, Luwen Huangfu, Sheng Huang 0001 |
Expert Syst. J. Knowl. Eng. | 3 |
| 2023 | Anchor-based discriminative dual distribution calibration for transductive zero-shot learning
Yi Zhang 0113, Sheng Huang 0001, Xiaohong Zhang 0002, Dan Yang 0001 |
Image Vis. Comput. | 1 |
| 2022 | Dual Space Multiple Instance Representative Learning for Medical Image Classification
Xiaoxian Zhang, Sheng Huang 0001, Yi Zhang 0113, Xiaohong Zhang 0002, Mingchen Gao, Chen Liu 0026 |
BMVC | 3 |
| 2022 | Adversarial Bidirectional Feature Generation for Generalized Zero-Shot Learning Under Unreliable Semantics
Guowei Yin, Yi Zhang 0113, Sheng Huang 0001 |
PRCV (2) | 3 |
| 2021 | Generally Boosting Few-Shot Learning with HandCrafted FeaturesabstractExisting Few-Shot Learning (FSL) methods predominantly focus on developing different types of sophisticated models to extract the transferable prior knowledge for recognizing novel classes, while they almost pay less attention to the feature learning part in FSL which often simply leverage some well-known CNN as the feature learner. However, feature is the core medium for encoding such transferable knowledge. Feature learning is easy to be trapped in the over-fitting particularly in the scarcity of the training data, and thereby degenerates the performances of FSL. The handcrafted features, such as Histogram of Oriented Gradient (HOG) and Local Binary Pattern (LBP), have no requirement on the amount of training data, and used to perform quite well in many small-scale data scenarios, since their extractions involve no learning process, and are mainly based on the empirically observed and summarized prior feature engineering knowledge. In this paper, we intend to develop a general and simple approach for generally boosting FSL via exploiting such prior knowledge in the feature learning phase. To this end, we introduce two novel handcrafted feature regression modules, namely HOG and LBP regression, to the feature learning parts of deep learning-based FSL models. These two modules are separately plugged into the different convolutional layers of backbone based on the characteristics of the corresponding handcrafted features to guide the backbone optimization from different feature granularity, and also ensure that the learned feature can encode the handcrafted feature knowledge which improves the generalization ability of feature and alleviate the over-fitting of the models. Three recent state-of-the-art FSL approaches are leveraged for examining the effectiveness of our method. Extensive experiments on miniImageNet, CIFAR-FS and FC100 datasets show that the performances of all these FSL approaches are well boosted via applying our method on all three datasets. Our codes and models have been released. Yi Zhang 0113, Sheng Huang 0001, Fengtao Zhou |
ACM Multimedia | 1 |