EDBT 2026 Demo / reviewers in the wild / expert
Zhen Zhao 0001
dblp:29/1773-1
· DBLP profile ↗
29ranked-venue papers
7as first author
27since 2021 · last 2025
0000-0002-0796-4078ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 19 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy LabelsabstractMedical Visual Question Answering (Med-VQA) systems benefit the interpretation of medical images containing critical clinical information. However, the challenge of noisy labels and limited high-quality datasets remains underexplored. To address this, we establish the first benchmark for noisy labels in Med-VQA by simulating human mislabeling with semantically designed noise types. More importantly, we introduce the DiN framework, which leverages a diffusion model to handle noisy labels in Med-VQA. Unlike the dominant classification-based VQA approaches that directly predict answers, our Answer Diffuser (AD) module employs a coarse-to-fine process, refining answer candidates with a diffusion model for improved accuracy. The Answer Condition Generator (ACG) further enhances this process by generating task-specific conditional information via integrating answer embeddings with fused image-question features. To address label noise, our Noisy Label Refinement(NLR) module introduces a robust loss function and dynamic answer adjustment to further boost the performance of the AD module. Our DiN framework consistently outperforms existing methods across multiple benchmarks with varying noise levels1. Erjian Guo, Zhen Zhao 0001, Zicheng Wang 0012, Tong Chen 0011, Yunyi Liu, Luping Zhou |
CVPR | 2 |
| 2025 | TimeKAN: KAN-based Frequency Decomposition Learning Architecture for Long-term Time Series ForecastingabstractReal-world time series often have multiple frequency components that are intertwined with each other, making accurate time series forecasting challenging. Decomposing the mixed frequency components into multiple single frequency components is a natural choice. However, the information density of patterns varies across different frequencies, and employing a uniform modeling approach for different frequency components can lead to inaccurate characterization. To address this challenges, inspired by the flexibility of the recent Kolmogorov-Arnold Network (KAN), we propose a KAN-based Frequency Decomposition Learning architecture (TimeKAN) to address the complex forecasting challenges caused by multiple frequency mixtures. Specifically, TimeKAN mainly consists of three components: Cascaded Frequency Decomposition (CFD) blocks, Multi-order KAN Representation Learning (M-KAN) blocks and Frequency Mixing blocks. CFD blocks adopt a bottom-up cascading approach to obtain series representations for each frequency band. Benefiting from the high flexibility of KAN, we design a novel M-KAN block to learn and represent specific temporal patterns within each frequency band. Finally, Frequency Mixing blocks is used to recombine the frequency bands into the original format. Extensive experimental results across multiple real-world time series datasets demonstrate that TimeKAN achieves state-of-the-art performance as an extremely lightweight architecture. Code is available at https://github.com/huangst21/TimeKAN. Songtao Huang, Zhen Zhao 0001, Can Li 0014, Lei Bai 0001 |
ICLR | 2 |
| 2025 | Mutual Prompt Leaning for Vision Language Models
Sifan Long 0001, Zhen Zhao 0001, Junkun Yuan, Zichang Tan, Jiangjiang Liu 0006, Jingyuan Feng, Sheng-Sheng Wang 0001, Jingdong Wang 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | UniMatch V2: Pushing the Limit of Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation (SSS) aims at learning rich visual knowledge from cheap unlabeled images to enhance semantic segmentation capability. Among recent works, UniMatch (Yang et al. 2023) improves its precedents tremendously by amplifying the practice of weak-to-strong consistency regularization. Subsequent works typically follow similar pipelines and propose various delicate designs. Despite the achieved progress, strangely, even in this flourishing era of numerous powerful vision models, almost all SSS works are still sticking to 1) using outdated ResNet encoders with small-scale ImageNet-1 K pre-training, and 2) evaluation on simple Pascal and Cityscapes datasets. In this work, we argue that, it is necessary to switch the baseline of SSS from ResNet-based encoders to more capable ViT-based encoders (e.g., DINOv2) that are pre-trained on massive data. A simple update on the encoder (even using 2× fewer parameters) can bring more significant improvement than careful method designs. Built on this competitive baseline, we present our upgraded and simplified UniMatch V2, inheriting the core spirit of weak-to-strong consistency from V1, but requiring less training cost and providing consistently better results. Additionally, witnessing the gradually saturated performance on Pascal and Cityscapes, we appeal that we should focus on more challenging benchmarks with complex taxonomy, such as ADE20K and COCO datasets. Lihe Yang, Zhen Zhao 0001, Hengshuang Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Imbalanced Medical Image Segmentation With Pixel-Dependent Noisy LabelsabstractAccurate medical image segmentation is often hindered by noisy labels in training data, due to the challenges of annotating medical images. Prior research works addressing noisy labels tend to make class-dependent assumptions, overlooking the pixel-dependent nature of most noisy labels. Furthermore, existing methods typically apply fixed thresholds to filter out noisy labels, risking the removal of minority classes and consequently degrading segmentation performance. To bridge these gaps, our proposed framework, Collaborative Learning with Curriculum Selection (CLCS), addresses pixel-dependent noisy labels with class imbalance. CLCS advances the existing works by i) treating noisy labels as pixel-dependent and addressing them through a collaborative learning framework, and ii) employing a curriculum dynamic thresholding approach adapting to model learning progress to select clean data samples to mitigate the class imbalance issue, and iii) applying a noise balance loss to noisy data samples to improve data utilization instead of discarding them outright. Specifically, our CLCS contains two modules: Curriculum Noisy Label Sample Selection (CNS) and Noise Balance Loss (NBL). In the CNS module, we designed a two-branch network with discrepancy loss for collaborative learning so that different feature representations of the same instance could be extracted from distinct views and used to vote the class probabilities of pixels. Besides, a curriculum dynamic threshold is adopted to select clean-label samples through probability voting. In the NBL module, instead of directly dropping the suspiciously noisy labels, we further adopt a robust loss to leverage such instances to boost the performance. We verify our CLCS on two benchmarks with different types of segmentation noise. Our method can obtain new state-of-the-art performance in different settings, yielding more than 3% Dice and mIoU improvements. Our code is available at https://github.com/Erjian96/CLCS.git. Erjian Guo, Zicheng Wang 0012, Zhen Zhao 0001, Luping Zhou |
IEEE Trans. Medical Imaging | 3 |
| 2025 | SOEDiff: Efficient Distillation for Small Object EditingabstractIn this article, we delve into a new task known as Small Object Editing (SOE), which focuses on text-based image inpainting within a constrained, small-sized area. Despite the remarkable success have been achieved by current image inpainting approaches, their application to the SOE task generally results in failure cases such as Object Missing, Text-Image Mismatch, and Distortion . These failures stem from the limited use of small-sized objects in training datasets and the down-sampling operations employed by U-Net models, which hinders accurate generation. To overcome these challenges, we introduce a novel training-based approach, SOEDiff, aimed at enhancing the capability of baseline models like StableDiffusion in editing small-sized objects while minimizing training costs. Specifically, our method involves two key components: SO-LoRA , which efficiently fine-tunes low-rank matrices, and Cross-scale score distillation , which leverages high-resolution predictions from the pre-trained teacher diffusion model. Our method presents significant improvements on the test dataset collected from MSCOCO and OpenImage, validating the effectiveness of our proposed method in SOE. In particular, when comparing SOEDiff with SD-I model on the OpenImage-small-val dataset, we observe a 0.99 improvement in CLIP-Score and a reduction of 2.87 in FID. Yiming Wu 0005, Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Ronghua Liang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Roll with the Punches: Expansion and Shrinkage of Soft Label Selection for Semi-supervised Fine-Grained LearningabstractWhile semi-supervised learning (SSL) has yielded promising results, the more realistic SSL scenario remains to be explored, in which the unlabeled data exhibits extremely high recognition difficulty, e.g., fine-grained visual classification in the context of SSL (SS-FGVC). The increased recognition difficulty on fine-grained unlabeled data spells disaster for pseudo-labeling accuracy, resulting in poor performance of the SSL model. To tackle this challenge, we propose Soft Label Selection with Confidence-Aware Clustering based on Class Transition Tracking (SoC) by reconstructing the pseudo-label selection process by jointly optimizing Expansion Objective and Shrinkage Objective, which is based on a soft label manner. Respectively, the former objective encourages soft labels to absorb more candidate classes to ensure the attendance of ground-truth class, while the latter encourages soft labels to reject more noisy classes, which is theoretically proved to be equivalent to entropy minimization. In comparisons with various state-of-the-art methods, our approach demonstrates its superior performance in SS-FGVC. Checkpoints and source code are available at https://github.com/NJUyued/SoC4SS-FGVC. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
AAAI | 2 |
| 2024 | Progressive Classifier and Feature Extractor Adaptation for Unsupervised Domain Adaptation on Point Clouds
Zicheng Wang 0012, Zhen Zhao 0001, Yiming Wu 0005, Luping Zhou, Dong Xu 0001 |
ECCV (28) | 2 |
| 2024 | Alternate Diverse Teaching for Semi-supervised Medical Image Segmentation
Zhen Zhao 0001, Zicheng Wang 0012, Longyue Wang, Dian Yu 0001, Yixuan Yuan, Luping Zhou |
ECCV (5) | 1 |
| 2024 | Towards Small Object Editing: A Benchmark Dataset and A Training-Free ApproachabstractA plethora of text-guided image editing methods has recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models especially Stable Diffusion. Despite the success of diffusion models in producing high-quality images, their application to small object generation has been limited due to difficulties in aligning cross-modal attention maps between text and these objects. Our approach offers a training-free method that significantly mitigates this alignment issue with local and global attention guidance, enhancing the model's ability to accurately render small objects in accordance with textual descriptions. We detail the methodology in our approach, emphasizing its divergence from traditional generation techniques and highlighting its advantages. What's more important is that we also provide SOEBench (Small Object Editing), a standardized benchmark for quantitatively evaluating text-based small object generation collected from MSCOCO[22] and OpenImage[18]. Preliminary results demonstrate the effectiveness of our method, showing marked improvements in the fidelity and accuracy of small object generation compared to existing models. This advancement not only contributes to the field of AI and computer vision but also opens up new possibilities for applications in various industries where precise image generation is critical.We will release our dataset on our project page: https://soebench.github.io/ Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Yiming Wu 0005, Wei Ji 0008, Haoran Liang 0001, Ronghua Liang |
ACM Multimedia | 2 |
| 2024 | GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation
Zhanyu Wang, Longyue Wang, Zhen Zhao 0001, Minghao Wu, Chenyang Lyu, Deng Cai 0002, Luping Zhou, Shuming Shi 0001, Zhaopeng Tu |
ACM Multimedia | 3 |
| 2024 | Depth Anything V2abstractThis work presents Depth Anything V2. Without pursuing fancy techniques, we aim to reveal crucial findings to pave the way towards building a powerful monocular depth estimation model. Notably, compared with V1, this version produces much finer and more robust depth predictions through three key practices: 1) replacing all labeled real images with synthetic images, 2) scaling up the capacity of our teacher model, and 3) teaching student models via the bridge of large-scale pseudo-labeled real images. Compared with the latest models built on Stable Diffusion, our models are significantly more efficient (more than 10x faster) and more accurate. We offer models of different scales (ranging from 25M to 1.3B params) to support extensive scenarios. Benefiting from their strong generalization capability, we fine-tune them with metric depth labels to obtain our metric depth models. In addition to our models, considering the limited diversity and frequent noise in current test sets, we construct a versatile evaluation benchmark with sparse depth annotations to facilitate future research. Models are available at https://github.com/DepthAnything/Depth-Anything-V2. Lihe Yang, Bingyi Kang, Zhen Zhao 0001, Xiaogang Xu 0002, Jiashi Feng, Hengshuang Zhao |
NeurIPS | 4 |
| 2024 | Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generationabstractdemands substantial human effort and incurs high training Yunxin Li, Baotian Hu, Longyue Wang, Jiashun Zhu, Jinyi Xu, Zhen Zhao 0001, Min Zhang 0005 |
SIGGRAPH Asia | 7 |
| 2024 | MutexMatch: Semi-Supervised Learning With Mutex-Based Consistency RegularizationabstractThe core issue in semi-supervised learning (SSL) lies in how to effectively leverage unlabeled data, whereas most existing methods tend to put a great emphasis on the utilization of high-confidence samples yet seldom fully explore the usage of low-confidence samples. In this article, we aim to utilize low-confidence samples in a novel way with our proposed mutex-based consistency regularization, namely MutexMatch. Specifically, the high-confidence samples are required to exactly predict "what it is" by the conventional true-positive classifier (TPC), while low-confidence samples are employed to achieve a simpler goal-to predict with ease "what it is not" by the true-negative classifier (TNC). In this sense, we not only mitigate the pseudo-labeling errors but also make full use of the low-confidence unlabeled data by the consistency of dissimilarity degree. MutexMatch achieves superior performance on multiple benchmark datasets, i.e., Canadian Institute for Advanced Research (CIFAR)-10, CIFAR-100, street view house numbers (SVHN), self-taught learning 10 (STL-10), and mini-ImageNet. More importantly, our method further shows superiority when the amount of labeled data is scarce, e.g., 92.23% accuracy with only 20 labeled data on CIFAR-10. Code has been released at https://github.com/NJUyued/MutexMatch4SSL. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Lei Wang 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision TransformersabstractVision transformers have achieved significant improvements on various vision tasks but their quadratic interactions between tokens significantly reduce computational efficiency. Many pruning methods have been proposed to remove redundant tokens for efficient vision transformers recently. However, existing studies mainly focus on the token importance to preserve local attentive tokens but completely ignore the global token diversity. In this paper, we emphasize the cruciality of diverse global semantics and propose an efficient token decoupling and merging method that can jointly consider the token importance and diversity for token pruning. According to the class token attention, we decouple the attentive and inattentive tokens. In addition to preserving the most discriminative local tokens, we merge similar inattentive tokens and match homogeneous attentive tokens to maximize the token diversity. Despite its simplicity, our method obtains a promising trade-off between model complexity and classification accuracy. On DeiT-S, our method reduces the FLOPs by 35% with only a 0.2% accuracy drop. Notably, benefiting from maintaining the token diversity, our method can even improve the accuracy of DeiT-T by 0.1% after reducing its FLOPs by 40%. Sifan Long 0001, Zhen Zhao 0001, Jimin Pi, Sheng-Sheng Wang 0001, Jingdong Wang 0001 |
CVPR | 2 |
| 2023 | Conflict-Based Cross-View Consistency for Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation (SSS) has recently gained increasing research interest as it can reduce the requirement for large-scale fully-annotated training data. The current methods often suffer from the confirmation bias from the pseudo-labelling process, which can be alleviated by the co-training framework. The current co-training-based SSS methods rely on hand-crafted perturbations to prevent the different sub-nets from collapsing into each other, but these artificial perturbations cannot lead to the optimal solution. In this work, we propose a new conflict-based cross-view consistency (CCVC) method based on a two-branch co-training framework which aims at enforcing the two sub-nets to learn informative features from irrelevant views. In particular, we first propose a new cross-view consistency (CVC) strategy that encourages the two sub-nets to learn distinct features from the same input by introducing a feature discrepancy loss, while these distinct features are expected to generate consistent prediction scores of the input. The CVC strategy helps to prevent the two sub-nets from stepping into the collapse. In addition, we further propose a conflict-based pseudo-labelling (CPL) method to guarantee the model will learn more useful information from conflicting predictions, which will lead to a stable training process. We validate our new CCVC approach on the SSS benchmark datasets where our method achieves new state-of-the-art performance. Our code is available at https://github.com/xiaoyao3302/CCVC. Zicheng Wang 0012, Zhen Zhao 0001, Xiaoxia Xing, Dong Xu 0001, Luping Zhou |
CVPR | 2 |
| 2023 | Instance-Specific and Model-Adaptive Supervision for Semi-Supervised Semantic SegmentationabstractRecently, semi-supervised semantic segmentation has achieved promising performance with a small fraction of labeled data. However, most existing studies treat all unlabeled data equally and barely consider the differences and training difficulties among unlabeled instances. Differentiating unlabeled instances can promote instance-specific supervision to adapt to the model's evolution dynamically. In this paper, we emphasize the cruciality of instance differences and propose an instance-specific and model-adaptive supervision for semi-supervised semantic segmentation, named iMAS. Relying on the model's performance, iMAS employs a class-weighted symmetric intersection-over-union to evaluate quantitative hardness of each unlabeled instance and supervises the training on unlabeled data in a model-adaptive manner. Specifically, iMAS learns from unlabeled instances progressively by weighing their corresponding consistency losses based on the evaluated hardness. Besides, iMAS dynamically adjusts the augmentation for each instance such that the distortion degree of augmented instances is adapted to the model's generalization capability across the training course. Not integrating additional losses and training procedures, iMAS can obtain remarkable performance gains against current state-of-the-art approaches on segmentation benchmarks under different semi-supervised partition protocols11Code and logs: https://github.com/zhenzhao/iMAS. Zhen Zhao 0001, Sifan Long 0001, Jimin Pi, Jingdong Wang 0001, Luping Zhou |
CVPR | 1 |
| 2023 | Augmentation Matters: A Simple-Yet-Effective Approach to Semi-Supervised Semantic SegmentationabstractRecent studies on semi-supervised semantic segmentation (SSS) have seen fast progress. Despite their promising performance, current state-of-the-art methods tend to increasingly complex designs at the cost of introducing more network components and additional training procedures. Differently, in this work, we follow a standard teacher-student framework and propose AugSeg, a simple and clean approach that focuses mainly on data perturbations to boost the SSS performance. We argue that various data augmentations should be adjusted to better adapt to the semi-supervised scenarios instead of directly applying these techniques from supervised learning. Specifically, we adopt a simplified intensity-based augmentation that selects a random number of data transformations with uniformly sampling distortion strengths from a continuous space. Based on the estimated confidence of the model on different unlabeled samples, we also randomly inject labelled information to augment the unlabeled samples in an adaptive manner. Without bells and whistles, our simple AugSeg can readily achieve new state-of-the-art performance on SSS benchmarks under different partition protocols11Code and logs: https://github.com/zhenzhao/AugSeg.. Zhen Zhao 0001, Lihe Yang, Sifan Long 0001, Jimin Pi, Luping Zhou, Jingdong Wang 0001 |
CVPR | 1 |
| 2023 | Towards Semi-supervised Learning with Non-random Missing LabelsabstractSemi-supervised learning (SSL) tackles the label missing problem by enabling the effective usage of unlabeled data. While existing SSL methods focus on the traditional setting, a practical and challenging scenario called label Missing Not At Random (MNAR) is usually ignored. In MNAR, the labeled and unlabeled data fall into different class distributions resulting in biased label imputation, which deteriorates the performance of SSL models. In this work, class transition tracking based Pseudo-Rectifying Guidance (PRG) is devised for MNAR. We explore the class-level guidance information obtained by the Markov random walk, which is modeled on a dynamically created graph built over the class tracking matrix. PRG unifies the historical information of class distribution and class transitions caused by the pseudo-rectifying procedure to maintain the model’s unbiased enthusiasm towards assigning pseudo-labels to all classes, so as the quality of pseudo-labels on both popular classes and rare classes in MNAR could be improved. Finally, we show the superior performance of PRG across a variety of MNAR scenarios, outperforming the latest SSL approaches combining bias removal solutions by a large margin. Code and model weights are available at https://github.com/NJUyued/PRG4SSL-MNAR. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
ICCV | 2 |
| 2023 | Enhancing Sample Utilization through Sample Adaptive Augmentation in Semi-Supervised LearningabstractIn semi-supervised learning, unlabeled samples can be utilized through augmentation and consistency regularization. However, we observed certain samples, even undergoing strong augmentation, are still correctly classified with high confidence, resulting in a loss close to zero. It indicates that these samples have been already learned well and do not provide any additional optimization benefits to the model. We refer to these samples as "naive samples". Unfortunately, existing SSL models overlook the characteristics of naive samples, and they just apply the same learning strategy to all samples. To further optimize the SSL model, we emphasize the importance of giving attention to naive samples and augmenting them in a more diverse manner. Sample adaptive augmentation (SAA) is proposed for this stated purpose and consists of two modules: 1) sample selection module; 2) sample augmentation module. Specifically, the sample selection module picks out naive samples based on historical training information at each epoch, then the naive samples will be augmented in a more diverse manner in the sample augmentation module. Thanks to the extreme ease of implementation of the above modules, SAA is advantageous for being simple and lightweight. We add SAA on top of FixMatch and FlexMatch respectively, and experiments demonstrate SAA can significantly improve the models. For example, SAA helped improve the accuracy of FixMatch from 92.50% to 94.76% and that of FlexMatch from 95.01% to 95.31% on CIFAR-10 with 40 labels. The code is available at https://github.com/GuanGui-nju/SAA. Guan Gui 0002, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
ICCV | 2 |
| 2023 | Task-Oriented Multi-Modal Mutual Learning for Vision-Language ModelsabstractPrompt learning has become one of the most efficient paradigms for adapting large pre-trained vision-language models to downstream tasks. Current state-of-the-art methods, like CoOp and ProDA, tend to adopt soft prompts to learn an appropriate prompt for each specific task. Recent CoCoOp further boosts the base-to-new generalization performance via an image-conditional prompt. However, it directly fuses identical image semantics to prompts of different labels and significantly weakens the discrimination among different classes as shown in our experiments. Motivated by this observation, we first propose a class-aware text prompt (CTP) to enrich generated prompts with label-related image information. Unlike CoCoOp, CTP can effectively involve image semantics and avoid introducing extra ambiguities into different prompts. On the other hand, instead of reserving the complete image representations, we propose text-guided feature tuning (TFT) to make the image branch attend to class-related representation. A contrastive loss is employed to align such augmented text and image representations on downstream tasks. In this way, the image-to-text CTP and text-to-image TFT can be mutually promoted to enhance the adaptation of VLMs for downstream tasks. Extensive experiments demonstrate that our method outperforms the existing methods by a significant margin. Especially, compared to CoCoOp, we achieve an average improvement of 4.03% on new classes and 3.19% on harmonic-mean over eleven classification benchmarks. Sifan Long 0001, Zhen Zhao 0001, Junkun Yuan, Zichang Tan, Jiangjiang Liu 0006, Luping Zhou, Sheng-Sheng Wang 0001, Jingdong Wang 0001 |
ICCV | 2 |
| 2023 | Shrinking Class Space for Enhanced Certainty in Semi-Supervised LearningabstractSemi-supervised learning is attracting blooming attention, due to its success in combining unlabeled data. To mitigate potentially incorrect pseudo labels, recent frameworks mostly set a fixed confidence threshold to discard uncertain samples. This practice ensures high-quality pseudo labels, but incurs a relatively low utilization of the whole unlabeled set. In this work, our key insight is that these uncertain samples can be turned into certain ones, as long as the confusion classes for the top-1 class are detected and removed. Invoked by this, we propose a novel method dubbed ShrinkMatch to learn uncertain samples. For each uncertain sample, it adaptively seeks a shrunk class space, which merely contains the original top-1 class, as well as remaining less likely classes. Since the confusion ones are removed in this space, the re-calculated top-1 confidence can satisfy the pre-defined threshold. We then impose a consistency regularization between a pair of strongly and weakly augmented samples in the shrunk space to strive for discriminative representations. Furthermore, considering the varied reliability among uncertain samples and the gradually improved model during training, we correspondingly design two reweighting principles for our uncertain loss. Our method exhibits impressive performance on widely adopted benchmarks. Lihe Yang, Zhen Zhao 0001, Lei Qi 0001, Yu Qiao 0001, Yinghuan Shi, Hengshuang Zhao |
ICCV | 2 |
| 2023 | Entropy-based Optimization on Individual and Global Predictions for Semi-Supervised LearningabstractPseudo-labelling-based semi-supervised learning (SSL) has demonstrated remarkable success in enhancing model performance by effectively leveraging a large amount of unlabeled data. However, existing studies focus mainly on rectifying individual predictions (i.e., pseudo-labels) on each unlabeled instance but ignore the overall prediction statistics from a global perspective. Such neglect may lead to model collapse and performance degradation in SSL, especially in label-scarce scenarios. In this paper, we emphasize the cruciality of global prediction constraints and propose a new SSL method that employs Entropy-based optimization on both Individual and Global predictions of unlabeled instances, dubbed EntInG. Specifically, we propose two criteria for leveraging unlabeled data in SSL: individual prediction entropy minimization (IPEM) and global distribution entropy maximization (GDEM). On the one hand, we show that current dominant SSL methods can be viewed as an implicit form of IPEM improved by recent augmentation techniques. On the other hand, we construct a new distribution loss to encourage GDEM, which greatly benefits producing better pseudo-labels for unlabeled data. Theoretical analysis also demonstrates that our proposed criteria can be derived by enforcing mutual information maximization on unlabeled instances. Despite its simplicity, our proposed method can achieve significant accuracy gains on popular SSL classification benchmarks. Zhen Zhao 0001, Meng Zhao 0003, Ye Liu 0013, Luping Zhou |
ACM Multimedia | 1 |
| 2022 | LaSSL: Label-Guided Self-Training for Semi-supervised LearningabstractThe key to semi-supervised learning (SSL) is to explore adequate information to leverage the unlabeled data. Current dominant approaches aim to generate pseudo-labels on weakly augmented instances and train models on their corresponding strongly augmented variants with high-confidence results. However, such methods are limited in excluding samples with low-confidence pseudo-labels and under-utilization of the label information. In this paper, we emphasize the cruciality of the label information and propose a Label-guided Self-training approach to Semi-supervised Learning (LaSSL), which improves pseudo-label generations from two mutually boosted strategies. First, with the ground-truth labels and iteratively-polished pseudo-labels, we explore instance relations among all samples and then minimize a class-aware contrastive loss to learn discriminative feature representations that make same-class samples gathered and different-class samples scattered. Second, on top of improved feature representations, we propagate the label information to the unlabeled samples across the potential data manifold at the feature-embedding level, which can further improve the labelling of samples with reference to their neighbours. These two strategies are seamlessly integrated and mutually promoted across the whole training process. We evaluate LaSSL on several classification benchmarks under partially labeled settings and demonstrate its superiority over the state-of-the-art approaches. Zhen Zhao 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001 |
AAAI | 1 |
| 2022 | DC-SSL: Addressing Mismatched Class Distribution in Semi-supervised LearningabstractConsistency-based Semi-supervised learning (SSL) has achieved promising performance recently. However, the success largely depends on the assumption that the labeled and unlabeled data share an identical class distribution, which is hard to meet in real practice. The distribution mismatch between the labeled and unlabeled sets can cause severe bias in the pseudo-labels of SSL, resulting in significant performance degradation. To bridge this gap, we put forward a new SSL learning framework, named Distribution Consistency SSL (DC-SSL), which rectifies the pseudolabels from a distribution perspective. The basic idea is to directly estimate a reference class distribution (RCD), which is regarded as a surrogate of the ground truth class distribution about the unlabeled data, and then improve the pseudo-labels by encouraging the predicted class distribution (PCD) of the unlabeled data to approach RCD gradually. To this end, this paper revisits the Exponentially Moving Average (EMA) model and utilizes it to estimate RCD in an iteratively improved manner, which is achieved with a momentum-update scheme throughout the training procedure. On top of this, two strategies are proposed for RCD to rectify the pseudo-label prediction, respectively. They correspond to an efficient training-free scheme and a training-based alternative that generates more accurate and reliable predictions. DC-SSL is evaluated on multiple SSL benchmarks and demonstrates remarkable performance improvement over competitive methods under matched- and mismatched-distribution scenarios. Zhen Zhao 0001, Luping Zhou, Yue Duan, Lei Wang 0001, Lei Qi 0001, Yinghuan Shi |
CVPR | 1 |
| 2022 | AoTI Minimization for Multi-Type Data Sampling in Industrial Wireless Sensor NetworksabstractFor practical industrial wireless sensor networks (IWSNs), the system freshness of a specific task is usually related to multiple and multitype sensing data. However, most existing research on freshness metrics, such as Age of Information (AoI) or Age of Processing (AoP), only considers a single-package setting with a single type of data. To fill this gap, we propose the Age of Task-oriented Information (AoTI) for measuring the freshness of industrial tasks in IWSNs. It measures the time elapsed of the latest analyzed results before arriving at the receiver since the generation of any type of sampling data belonging to one certain task. Furthermore, we aim to minimize the long-term AoTI for IWSNs applications by jointly optimizing access modes and sampling frequencies for all sensors. By first formulating the problem as a Mixed Integer Nonlinear Program-ming problem, we then transform it to a constrained Markov Decision Process (CMDP) and relax it as an un-constrained MDP using Lagrangian method. Finally, we develop a Learning-based Access mode selection and Sampling frequency Control (LASC) algorithm and verify its superiority through simulations. Chen Ying, Zhen Zhao 0001, Changyan Yi, You Shi, Ran Wang 0004 |
EUC | 2 |
| 2022 | Improving Barely Supervised Learning by Discriminating Unlabeled Samples with Super-ClassabstractIn semi-supervised learning (SSL), a common practice is to learn consistent information from unlabeled data and discriminative information from labeled data to ensure both the immutability and the separability of the classification model. Existing SSL methods suffer from failures in barely-supervised learning (BSL), where only one or two labels per class are available, as the insufficient labels cause the discriminative information being difficult or even infeasible to learn. To bridge this gap, we investigate a simple yet effective way to leverage unlabeled samples for discriminative learning, and propose a novel discriminative information learning module to benefit model training. Specifically, we formulate the learning objective of discriminative information at the super-class level and dynamically assign different classes into different super-classes based on model performance improvement. On top of this on-the-fly process, we further propose a distribution-based loss to learn discriminative information by utilizing the similarity relationship between samples and super-classes. It encourages the unlabeled samples to stay closer to the distribution of their corresponding super-class than those of others. Such a constraint is softer than the direct assignment of pseudo labels, while the latter could be very noisy in BSL. We compare our method with state-of-the-art SSL and BSL methods through extensive experiments on standard SSL benchmarks. Our method can achieve superior results, \eg, an average accuracy of 76.76\% on CIFAR-10 with merely 1 label per class. Guan Gui 0002, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
NeurIPS | 2 |
| 2017 | Energy Efficient Packet Transmission Strategies for Wireless Body Area Networks with Rechargeable SensorsabstractIn this paper, we investigate energy efficient packet transmission strategies for wireless body area networks (WBANs) with rechargeable sensors. For practical implementations, we propose a multi-threshold based transmission strategy by taking into account the channel state, battery state and number of buffered packets in the system. A discrete Markov arrival process (DMAP) is introduced to jointly model channel correlations and energy allocations. After that, with given thresholds and corresponding energy allocations, a level dependent Quasi-Birth- and-Death Markov chain is constructed to evaluate the system performance. According to the derived performance metrics, we formulate an optimization problem to find optimal thresholds for energy efficiency maximization with reasonable performance provisioning. Extensive simulations are conducted to verify our proposed queueing analytical model and demonstrate perfor- mance gains of our proposed strategy. Zhen Zhao 0001, Shiwei Huang, Jun Cai 0001 |
VTC Fall | 1 |
| 2016 | Priority-aware pricing-based capacity sharing scheme for beyond-wireless body area networks
Changyan Yi, Zhen Zhao 0001, Jun Cai 0001, Ricardo Lobato de Faria, Gong Zhang 0010 |
Comput. Networks | 2 |