EDBT 2026 Demo / reviewers in the wild / expert
Zhun Zhong
dblp:32/6525
· DBLP profile ↗
116ranked-venue papers
15as first author
89since 2021 · last 2026
0000-0002-8202-0544ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 75 · 11 first-author · 54 since 2021Artificial intelligence and machine learning · 72 · 11 first-author · 64 since 2021Security and privacy · 4 · 4 since 2021Computer networks · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Open-World Deepfake Attribution via Confidence-Aware Asymmetric LearningabstractThe proliferation of synthetic facial imagery has intensified the need for robust Open-World DeepFake Attribution (OW-DFA), which aims to attribute both known and unknown forgeries using labeled data for known types and unlabeled data containing a mixture of known and novel types. However, existing OW-DFA methods face two critical limitations: 1) A confidence skew that leads to unreliable pseudo-labels for novel forgeries, resulting in biased training. 2) An unrealistic assumption that the number of unknown forgery types is known a priori. To address these challenges, we propose a Confidence-aware Asymmetric Learning (CAL) framework, which adaptively balances model confidence across known and novel forgery types. CAL mainly consists of two components: Confidence-aware Consistency Regularization (CCR) and Asymmetric Confidence Reinforcement (ACR). CCR mitigates pseudo-label bias by dynamically scaling sample losses based on normalized confidence, gradually shifting the training focus from high- to low-confidence samples. ACR complements this by separately calibrating confidence for known and novel classes through selective learning on high-confidence samples, guided by their confidence gap. Together, CCR and ACR form a mutually reinforcing loop that significantly improves the model's OW-DFA performance. Moreover, we introduce a Dynamic Prototype Pruning (DPP) strategy that automatically estimates the number of novel forgery types in a coarse-to-fine manner, removing the need for unrealistic prior assumptions and enhancing the scalability of our methods to real-world OW-DFA scenarios. Extensive experiments on the standard and OW-DFA benchmark and a newly extended benchmark incorporating advanced manipulations demonstrate that CAL consistently outperforms previous methods, achieving new state-of-the-art performance on both known and novel forgery attribution. Haiyang Zheng, Nan Pu, Wenjing Li 0005, Nicu Sebe, Zhun Zhong |
AAAI | 6 |
| 2026 | Parameter-efficient multimodal adaptation for adverse condition depth estimation
Guanglei Yang, Yongqiang Zhang 0007, Zhun Zhong, Wangmeng Zuo |
Expert Syst. Appl. | 4 |
| 2026 | Memory Consistency Guided Divide-and-Conquer Learning for Generalized Category Discovery
Yuanpeng Tu, Zhun Zhong, Hengshuang Zhao |
Int. J. Comput. Vis. | 2 |
| 2026 | Plug-and-play Class-aware Knowledge Injection for Prompt Learning with Visual-Language Model
Junhui Yin, Nan Pu, Lingfeng Yang, Zhun Zhong |
Int. J. Comput. Vis. | 7 |
| 2026 | Towards Stable Source-Free Domain Adaptive Semantic Segmentation
Dong Zhao 0007, Qi Zang, Nan Pu, Jinlong Li 0003, Shuang Wang 0001, Nicu Sebe, Zhun Zhong |
Int. J. Comput. Vis. | 7 |
| 2026 | Identity-Compensated Style Distillation for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared Person Re-Identification (VI-ReID) that matches pedestrian images across visible and infrared modalities suffers from substantial modality discrepancies and intra-class variations. While existing methods typically address the modality gap via style alignment, they often lose identity-relevant semantics and overlook fine-grained inter-class nuances, such as body part contours and structural cues around the head, shoulders, or feet. To tackle these challenges, we propose an Identity-Compensated Style Distillation (ICSD) network that enforces cross-modality style consistency and enhances the discriminative power of modality-invariant features. Specifically, ICSD comprises two core components: (1) a Style Knowledge Distillation (SKD) module, which integrates Style Discrepancy Reduction (SDR) and Identity Knowledge Compensation (IKC) to align modality styles while preserving identity-relevant semantics; (2) an Identity Discrimination Amplification (IDA) module, which captures and enhances subtle inter-class differences by refining identity-specific cues, thereby facilitating more accurate discrimination between different pedestrians. Extensive experiments on three public benchmarks-SYSU-MM01, RegDB, and LLCM-demonstrate that ICSD consistently outperforms state-of-the-art methods, validating the effectiveness and complementarity of its components. Yongguo Ling, Zihao Hu, Nan Pu, Zhun Zhong, Xudong Jiang 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Boosting Semi-Supervised Learning With Entropy-Guided Adaptive Reward MaximizationabstractExisting semi-supervised learning (SSL) methods rely predominantly on pseudo-labeling and consistency regularization to leverage unlabeled data, demonstrating significant performance improvements. However, we pinpoint that these methods suffer from a confidence-for-weighting issue, overvaluing high-confidence pseudo-labels while undervaluing low-confidence yet informative samples that are critical for robust generalization. In this paper, we introduce EntropyMatch, an entropy-driven SSL framework that redefines sample importance through prediction entropy rather than confidence alone. EntropyMatch employs a bidirectional weighting strategy: upward exploitation exploits reliable hard samples to refine decision boundaries while downward exploration cautiously explores uncertain ones to reduce noise. Additionally, EntropyMatch features an adaptive training mechanism that aligns with model maturity, shifting focus from safe exploration to strategic exploitation as training progresses. Experiments on eight benchmarks across various SSL tasks-spanning image classification, facial expression recognition, and human action recognition-validate EntropyMatch's robustness and effectiveness. It consistently achieves state-of-the-art results, notably matching state-of-the-art LION's performance on RAF-DB with just half the labeled data, demonstrating superior data efficiency and generalization. Anyang Tong, Zenglin Shi, Zhun Zhong, Meng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Prior-Constrained Association Learning for Fine-Grained Generalized Category DiscoveryabstractThis paper addresses generalized category discovery (GCD), the task of clustering unlabeled data from potentially known or unknown categories with the help of labeled instances from each known category. Compared to traditional semi-supervised learning, GCD is more challenging because unlabeled data could be from novel categories not appearing in labeled data. Current state-of-the-art methods typically learn a parametric classifier assisted by self-distillation. While being effective, these methods do not make use of cross-instance similarity to discover class-specific semantics which are essential for representation learning and category discovery. In this paper, we revisit the association-based paradigm and propose a Prior-constrained Association Learning method to capture and learn the semantic relations within data. In particular, the labeled data from known categories provides a unique prior for the association of unlabeled data. Unlike previous methods that only adopts the prior as a pre or post-clustering refinement, we fully incorporate the prior into the association process, and let it constrain the association towards a reliable grouping outcome. The estimated semantic groups are utilized through non-parametric prototypical contrast to enhance the representation learning. A further combination of both parametric and non-parametric classification complements each other and leads to a model that outperforms existing methods by a significant margin. On multiple GCD benchmarks, we perform extensive experiments and validate the effectiveness of our proposed method. Menglin Wang 0001, Zhun Zhong, Xiaojin Gong |
AAAI | 2 |
| 2025 | ChangeDiff: A Multi-Temporal Change Detection Data Generator with Flexible Text Prompts via Diffusion ModelabstractData-driven deep learning models have enabled tremendous progress in change detection (CD) with the support of pixel-level annotations. However, collecting diverse data and manually annotating them is costly, laborious, and knowledge-intensive. Existing generative methods for CD data synthesis show competitive potential in addressing this issue but still face the following limitations: 1) difficulty in flexibly controlling change events, 2) dependence on additional data to train the data generators, 3) focus on specific change detection tasks. To this end, this paper focuses on the semantic CD (SCD) task and develops a multi-temporal SCD data generator ChangeDiff by exploring powerful diffusion models. ChangeDiff innovatively generates change data in two steps: first, it uses text prompts and a text-to-layout (T2L) model to create continuous layouts, and then it employs layout-to-image (L2I) to convert these layouts into images. Specifically, we propose multi-class distribution-guided text prompts (MCDG-TP), allowing for layouts to be generated flexibly through controllable classes and their corresponding ratios. Subsequently, to generalize the T2L model to the proposed MCDG-TP, a class distribution refinement loss is further designed as training supervision. Our generated data shows significant progress in temporal continuity, spatial diversity, and quality realism, empowering change detectors with accuracy and transferability. Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Wenjun Yi, Zhun Zhong |
AAAI | 6 |
| 2025 | Multi-Scale Global-Instance Prompt Tuning for Continual Test-Time Adaptation in Medical Image SegmentationabstractDistribution shift is a common challenge in medical images obtained from different clinical centers, significantly hindering the deployment of pre-trained semantic segmentation models in real-world applications across multiple domains. Continual Test-Time Adaptation (CTTA) has emerged as a promising approach to address cross-domain distribution shifts during continually evolving target domains. Most existing CTTA methods rely on incrementally updating model parameters, which inevitably suffer from error accumulation and catastrophic forgetting, especially in long-term adaptation. Recent prompt-tuning-based works have shown potential to mitigate the two issues above by updating only visual prompts. While these approaches have demonstrated promising performance, several limitations remain: 1) lacking multi-scale prompt diversity, 2) inadequate incorporation of instance-specific knowledge, and 3) risk of privacy leakage. To overcome these limitations, we propose Multi-scale Global-Instance Prompt Tuning (MGIPT), to enhance scale diversity of prompts as well as capture both globaland instance-level knowledge for robust CTTA. Specifically, MGIPT consists of an Adaptive-scale Instance Prompt (AIP) and a Multi-scale Global-level Prompt (MGP). AIP dynamically learns lightweight and instance-specific prompts to mitigate error accumulation with adaptive optimal-scale selection mechanism. MGP captures domain-level knowledge across different scales to ensure robust adaptation with anti-forgetting capabilities. These complementary components are combined through a weighted ensemble approach, enabling effective dual-level adaptation that integrates both global and local information. Extensive experiments on medical image segmentation benchmarks (five optic disc/cup datasets and four polyp datasets) demonstrate that our MGIPT outperforms state-of-the-art methods, achieving robust adaptation across continually changing target domains. Notably, our MGIPT exhibits particularly strong performance in longterm CTTA scenarios, showing great anti-forgetting ability. Lingrui Li, Yanfeng Zhou, Nan Pu, Xin Chen 0003, Zhun Zhong |
BIBM | 5 |
| 2025 | Feature Spectrum Learning for Remote Sensing Change DetectionabstractChange detection (CD) holds significant implications for Earth observation, in which pseudo-changes between bitemporal images induced by imaging environmental factors are key challenges. Existing methods mainly regard pseudo-changes as a kind of style shift and alleviate it by transforming bitemporal images into the same style using generative adversarial networks (GANs). Nevertheless, their efforts are limited by the complexity of optimizing GANs and the absence of guidance from physical properties. This paper finds that the spectrum transformation (ST) has the potential to mitigate pseudo-changes by aligning in the frequency domain carrying the style. However, the benefit of ST is largely constrained by two drawbacks: 1) limited transformation space and 2) inefficient parameter search. To address these limitations, we propose a Feature Spectrum learning (FeaSpect) that adaptively eliminate pseudo-changes in the latent space. For the drawback 1), FeaSpect directs the transformation towards stylealigned discriminative features via feature spectrum transformation (FST). For the drawback 2), FeaSpect allows FST to be trainable, efficiently discovering optimal parameters via extraction box with adaptive attention and extraction box with learnable strides. Extensive experiments on challenging datasets demonstrate that our method remarkably outperforms existing methods and achieves a commendable trade-off between accuracy and efficiency. Importantly, our method can be easily injected into other frameworks, achieving consistent improvements. Qi Zang, Dong Zhao 0007, Shuang Wang 0001, Dou Quan, Zhun Zhong |
CVPR | 5 |
| 2025 | ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and GroundingabstractWe present ASAP, a new framework for detecting and grounding multi-modal media manipulation (DGM4). Upon thorough examination, we observe that accurate fine-grained cross-modal semantic alignment between the image and text is vital for accurately manipulation detection and grounding. While existing DGM4methods pay rare attention to the cross-modal alignment, hampering the accuracy of manipulation detecting to step further. To remedy this issue, this work targets to advance the semantic alignment learning to promote this task. Particularly, we utilize the off-the-shelf large models to construct paired image-text pairs, especially for the manipulated instances. Subsequently, a cross-modal alignment learning is performed to enhance the semantic alignment. Besides the explicit auxiliary clues, we further design a Manipulation-Guided Cross Attention (MGCA) to provide implicit guidance for augmenting the manipulation perceiving. With the grounding truth available during training, MGCA encourages the model to concentrate more on manipulated components while downplaying normal ones, enhancing the model’s ability to capture manipulations. Extensive experiments are conducted on the DGM4dataset, the results demonstrate that our model can surpass the comparison method with a clear margin. Code will be released at https://github.com/CriliasMiller/ASAP. Yaxiong Wang, Lechao Cheng, Zhun Zhong, Dan Guo 0001, Meng Wang 0001 |
CVPR | 4 |
| 2025 | FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized SegmentationabstractVision Foundation Models (VFMs) excel in generalization due to large-scale pretraining, but fine-tuning them for Domain Generalized Semantic Segmentation (DGSS) while maintaining this ability remains a challenge. Existing approaches either selectively fine-tune parameters or freeze the VFMs and update only the adapters, both of which may underutilize the VFMs’ full potential in DGSS tasks. We observe that domain-sensitive parameters in VFMs, arising from task and distribution differences, can hinder generalization. To address this, we propose FisherTune, a robust fine-tuning method guided by the Domain-Related Fisher Information Matrix (DR-FIM). DR-FIM measures parameter sensitivity across tasks and domains, enabling selective updates that preserve generalization and enhance DGSS adaptability. To stabilize DR-FIM estimation, FisherTune incorporates variational inference, treating parameters as Gaussian-Distributed variables and leveraging pre-trained priors. Extensive experiments show that Fisher-Tune achieves superior cross-domain segmentation while maintaining generalization, outperforming both selective-parameter and adapter-based methods. Dong Zhao 0007, Jinlong Li 0003, Shuang Wang 0001, Qi Zang, Nicu Sebe, Zhun Zhong |
CVPR | 7 |
| 2025 | Generate, Refine, and Encode: Leveraging Synthesized Novel Samples for On-the-Fly Fine-Grained Category DiscoveryabstractIn this paper, we investigate a practical yet challenging task: On-the-fly Category Discovery (OCD). This task focuses on the online identification of newly arriving stream data that may belong to both known and unknown categories, utilizing the category knowledge from only labeled data. Existing OCD methods are devoted to fully mining transferable knowledge from only labeled data. However, the transferability learned by these methods is limited because the knowledge contained in known categories is often insufficient, especially when few annotated data/categories are available in fine-grained recognition. To mitigate this limitation, we propose a diffusion-based OCD framework, dubbed DiffGRE, which integrates Generation, Refinement, and Encoding in a multi-stage fashion. Specifically, we first design an attribute-composition generation method based on cross-image interpolation in the diffusion latent space to synthesize novel samples. Then, we propose a diversity-driven refinement approach to select the synthesized images that differ from known categories for subsequent OCD model training. Finally, we leverage a semi-supervised leader encoding to inject additional category knowledge contained in synthesized data into the OCD models, which can benefit the discovery of both known and unknown categories during the on-the-fly inference process. Extensive experiments demonstrate the superiority of our DiffGRE over previous methods on six fine-grained datasets. Nan Pu, Haiyang Zheng, Wenjing Li 0005, Nicu Sebe, Zhun Zhong |
ICCV | 6 |
| 2025 | Pseudo-SD: Pseudo Controlled Stable Diffusion for Semi-Supervised and Cross-Domain Semantic Segmentation
Dong Zhao 0007, Qi Zang, Shuang Wang 0001, Nicu Sebe, Zhun Zhong |
ICCV | 5 |
| 2025 | Noisy Test-Time Adaptation in Vision-Language ModelsabstractTest-time adaptation (TTA) aims to address distribution shifts between source and target data by relying solely on target data during testing. In open-world scenarios, models often encounter noisy samples, i.e., samples outside the in-distribution (ID) label space. Leveraging the zero-shot capability of pre-trained vision-language models (VLMs), this paper introduces Zero-Shot Noisy TTA (ZS-NTTA), focusing on adapting the model to target data with noisy samples during test-time in a zero-shot manner. In the preliminary study, we reveal that existing TTA methods suffer from a severe performance decline under ZS-NTTA, often lagging behind even the frozen model. We conduct comprehensive experiments to analyze this phenomenon, revealing that the negative impact of unfiltered noisy data outweighs the benefits of clean data during model updating. In addition, as these methods adopt the adapting classifier to implement ID classification and noise detection sub-tasks, the ability of the model in both sub-tasks is largely hampered. Based on this analysis, we propose a novel framework that decouples the classifier and detector, focusing on developing an individual detector while keeping the classifier (including the backbone) frozen. Technically, we introduce the Adaptive Noise Detector (AdaND), which utilizes the frozen model's outputs as pseudo-labels to train a noise detector for detecting noisy samples effectively. To address clean data streams, we further inject Gaussian noise during adaptation, preventing the detector from misclassifying clean samples as noisy. Beyond the ZS-NTTA, AdaND can also improve the zero-shot out-of-distribution (ZS-OOD) detection ability of VLMs. Extensive experiments show that our method outperforms in both ZS-NTTA and ZS-OOD detection. On ImageNet, AdaND achieves a notable improvement of $8.32\%$ in harmonic mean accuracy ($\text{Acc}_\text{H}$) for ZS-NTTA and $9.40\%$ in FPR95 for ZS-OOD detection, compared to state-of-the-art methods. Importantly, AdaND is computationally efficient and comparable to the model-frozen method. The code is publicly available at: https://github.com/tmlr-group/ZS-NTTA. Chentao Cao, Zhun Zhong, Zhanke Zhou, Tongliang Liu, Yang Liu 0018, Kun Zhang 0001, Bo Han 0003 |
ICLR | 2 |
| 2025 | Knowledge Swapping via Learning and UnlearningabstractWe introduce Knowledge Swapping, a novel task designed to selectively regulate knowledge of a pretrained model by enabling the forgetting of user-specified information, retaining essential knowledge, and acquiring new knowledge simultaneously. By delving into the analysis of knock-on feature hierarchy, we find that incremental learning typically progresses from low-level representations to higher-level semantics, whereas forgetting tends to occur in the opposite direction—starting from high-level semantics and moving down to low-level features. Building upon this, we propose to benchmark the knowledge swapping task with the strategy of Learning Before Forgetting. Comprehensive experiments on various tasks like image classification, object detection, and semantic segmentation validate the effectiveness of the proposed strategy. The source code is available at https://github.com/xingmingyu123456/KnowledgeSwapping. Mingyu Xing, Lechao Cheng, Shengeng Tang, Yaxiong Wang, Zhun Zhong, Meng Wang 0001 |
ICML | 5 |
| 2025 | Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training ApproachabstractMicro-Action Recognition (MAR) aims to classify subtle human actions in video. However, annotating MAR datasets is particularly challenging due to the subtlety of actions. To this end, we introduce the setting of Semi-Supervised MAR (SSMAR), where only a part of samples are labeled. We first evaluate traditional Semi-Supervised Learning (SSL) methods to SSMAR and find that these methods tend to overfit on inaccurate pseudo-labels, leading to error accumulation and degraded performance. This issue primarily arises from the common practice of directly using the predictions of classifier as pseudo-labels to train the model. To solve this issue, we propose a novel framework, called Asynchronous Pseudo Labeling and Training (APLT), which explicitly separates the pseudo-labeling process from model training. Specifically, we introduce a semi-supervised clustering method during the offline pseudo-labeling phase to generate more accurate pseudo-labels. Moreover, a self-adaptive thresholding strategy is proposed to dynamically filter noisy labels of different classes. We then build a memory-based prototype classifier based on the filtered pseudo-labels, which is fixed and used to guide the subsequent model training phase. By alternating the two pseudo-labeling and model training phases in an asynchronous manner, the model can not only be learned with more accurate pseudo-labels but also avoid the overfitting issue. Experiments on three MAR datasets show that our APLT largely outperforms state-of-the-art SSL methods. For instance, APLT improves accuracy by 14.5% over FixMatch on the MA-12 dataset when using only 50% labeled data. Code is available at https://github.com/zy-hfut/APLT Yan Zhang 0053, Lechao Cheng, Yaxiong Wang, Zhun Zhong, Meng Wang 0001 |
IJCAI | 4 |
| 2025 | Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal ManipulationsabstractThe detection and grounding of manipulated content in multimodal data has emerged as a critical challenge in media forensics. While existing benchmarks demonstrate technical progress, they suffer from misalignment artifacts that poorly reflect real-world manipulation patterns: practical attacks typically maintain semantic consistency across modalities, whereas current datasets artificially disrupt cross-modal alignment, creating easily detectable anomalies. To bridge this gap, we pioneer the detection of semantically-coordinated manipulations where visual edits are systematically paired with semantically consistent textual descriptions. Our approach begins with constructing the first Semantic-Aligned Multimodal Manipulation (SAMM) dataset, generated through a two-stage pipeline: 1) applying state-of-the-art image manipulations, followed by 2) generation of contextually-plausible textual narratives that reinforce the visual deception. Building on this foundation, we propose a Retrieval-Augmented Manipulation Detection and Grounding (RamDG) framework. RamDG commences by harnessing external knowledge repositories to retrieve contextual evidence, which serves as the auxiliary texts and encoded together with the inputs through our image forgery grounding and deep manipulation detection modules to trace all manipulations. Extensive experiments demonstrate our framework significantly outperforms existing methods, achieving 2.06% higher detection accuracy on SAMM compared to state-of-the-art approaches. The dataset and code are publicly available at https://github.com/shen8424/SAMM-RamDG-CAP. Jinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu, Zhun Zhong |
ACM Multimedia | 5 |
| 2025 | Beyond General Alignment: Fine-Grained Entity-Centric Image-Text Matching with Multimodal Attentive ExpertsabstractRecent progress in aligning images with texts has achieved remarkable results, however, existing models tend to serve general queries and often fall short when dealing with detailed query requirements. In this paper, we work towards Entity-centric Image-Text Matching (EITM), a finer-grained image-text matching task that aligns texts and images centered around specific entities. The main challenge in EITM lies in bridging the substantial semantic gap between entity-related information in texts and images, which is more pronounced than in general image-text matching problems. To address this challenge, we adopt CLIP as our foundational model and devise a Multimodal Attentive Experts (MMAE)-based contrastive learning to adapt CLIP into an expert for EITM problem. Particularly, the core of our multimodal attentive experts learning is to generate explanation texts by Large Language Models (LLMs) as bridging clues. In specific, we first employ off-the-shelf LLMs to generate explanatory text. This text, along with the original image and text, is then fed into our Multimodal Attentive Experts module to narrow the semantic gap within a unified semantic space. Upon the enriched feature representations generated by MMAE, we have further developed an effective Gated Integrative Image-text Matching (GI-ITM) strategy. GI-ITM utilizes an adaptive gating mechanism to combine features from MMAE, followed by applying image-text matching constraints to enhance the alignment precision. Our method has been extensively evaluated on three social media news benchmarks: N24News, VisualNews, and GoodNews. The experimental results demonstrate that our approach significantly outperforms competing methods. Our code is available at: https://github.com/wangyxxjtu/ETE. Yaxiong Wang, Lianwei Wu, Lechao Cheng, Zhun Zhong, Yujiao Wu, Meng Wang 0001 |
SIGIR | 4 |
| 2025 | Guest Editorial: Special Issue on Open-World Visual Recognition
Zhun Zhong, Hong Liu 0009, Yin Cui, Shin'ichi Satoh 0001, Nicu Sebe, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Modeling the Label Distributions for Weakly-Supervised Semantic SegmentationabstractWeakly-Supervised Semantic Segmentation (WSSS) aims to train segmentation models by weak labels, which is receiving significant attention due to its low annotation cost. Existing approaches focus on generating pseudo labels for supervision while largely ignoring to leverage the inherent semantic correlation among different pseudo labels. We observe that pseudo-labeled pixels that are close to each other in the feature space are more likely to share the same class, and those closer to the distribution centers tend to have higher confidence. Motivated by this, we propose to model the underlying label distributions and employ cross-label constraints to generate more accurate pseudo labels. In this paper, we develop a unified WSSS framework named Adaptive Gaussian Mixtures Model, which leverages a GMM to model the label distributions. Specifically, we calculate the feature distribution centers of pseudo-labeled pixels and build the GMM by measuring the distance between the centers and each pseudo-labeled pixel. Then, we introduce an Online Expectation-Maximization (OEM) algorithm and a novel maximization loss to optimize the GMM adaptively, aiming to learn more discriminative decision boundaries between different class-wise Gaussian mixtures. Based on the label distributions, we leverage the GMM to generate high-quality pseudo labels for more reliable supervision. Our framework is capable of solving different forms of weak labels: image-level labels, points, scribbles, blocks, and bounding-boxes. Extensive experiments on PASCAL, COCO, Cityscapes, and ADE20 K datasets demonstrate that our framework can effectively provide more reliable supervision and outperform the state-of-the-art methods under all settings. Linshan Wu, Zhun Zhong, Jiayi Ma 0001, Yunchao Wei, Hao Chen 0011, Leyuan Fang, Shutao Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | SeCoV2: Semantic Connectivity-Driven Pseudo-Labeling for Robust Cross-Domain Semantic SegmentationabstractPseudo-labeling is a dominant strategy for cross-domain semantic segmentation (CDSS), yet its effectiveness is limited by fragmented and noisy pixel-level predictions under severe domain shifts. To address this, we propose a semantic connectivity-driven pseudo-labeling framework, SeCo, which constructs and refines pseudo-labels at the connectivity level by aggregating high-confidence pixels into coherent semantic regions. The framework includes two key components: Pixel Semantic Aggregation (PSA), which leverages a dual prompting strategy to preserve category-specific granularity, and Semantic Connectivity Correction with Loss Distribution (SCC-LD), which filters noisy regions based on early-loss statistics. Building upon this foundation, we further present SeCoV2, which introduces SCC-Unc, a novel uncertainty-aware correction module that constructs a connectivity graph and enforces relational consistency for robust refinement in ambiguous regions. SeCoV2 also broadens the applicability of SeCo by extending evaluation to more challenging scenarios, including open-set and multimodal adaptation, semi-supervised domain generalization, and by validating compatibility with different interactive foundation segmentation models such as SAM Kirillov et al. 2023, SEEM Zou et al. 2023, and Fast-SAM Zhao et al. 2023. Extensive experiments across six CDSS tasks demonstrate that SeCoV2 achieves consistent improvements over previous methods, with an average performance gain of up to +4.6%, establishing new state-of-the-art results. These findings highlight the effectiveness and generalization ability for robust adaptation in diverse real-world environments. Dong Zhao 0007, Qi Zang, Nan Pu, Shuang Wang 0001, Nicu Sebe, Zhun Zhong |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Cross-modality average precision optimization for visible thermal person re-identification
Yongguo Ling, Zhiming Luo, Dazhen Lin, Shaozi Li, Min Jiang 0005, Nicu Sebe, Zhun Zhong |
Pattern Recognit. | 7 |
| 2025 | Joint Style and Layout Synthesizing: Toward Generalizable Remote Sensing Semantic SegmentationabstractThis paper studies the domain generalized remote sensing semantic segmentation (RSSS), aiming to generalize a model trained only on the source domain to unseen domains. Existing methods in computer vision treat style information as domain characteristics to achieve domain-agnostic learning. Nevertheless, their generalizability to RSSS remains constrained, due to the incomplete consideration of domain characteristics. We argue that remote sensing scenes have layout differences beyond just style. Considering this, we devise a joint style and layout synthesizing framework, enabling the model to jointly learn out-of-domain samples synthesized from these two perspectives. For style, we estimate the variant intensities of per-class representations affected by domain shift and randomly sample within this modeled scope to reasonably expand the boundaries of style-carrying feature statistics. For layout, we explore potential scenes with diverse layouts in the source domain and propose granularity-fixed and granularity-learnable masks to perturb layouts, forcing the model to learn characteristics of objects rather than variable positions. The mask is designed to learn more context-robust representations by discovering difficult-to-recognize perturbation directions. Subsequently, we impose gradient angle constraints between the samples synthesized using the two ways to correct conflicting optimization directions. Extensive experiments demonstrate the superior generalization ability of our method over existing methods. Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Zhun Zhong, Biao Hou, Licheng Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | M3-ReID: Unifying Multi-View, Granularity, and Modality for Video-Based Visible-Infrared Person Re-IdentificationabstractVideo-based visible-infrared person re-identification (VVI-ReID) task focuses on cross-modality retrieval of pedestrian videos, which are captured in visible and infrared modalities by non-overlapping cameras across diverse scenes, and holds significant value for security surveillance scenarios. The challenges of this task mainly stem from three issues: the difficulty of capturing comprehensive spatio-temporal cues, intra-class variations within video sequences, and inter-modality discrepancies between visible and infrared data. Existing methods mainly try to address the modality gap or focus on one of the other aspects, but rarely do they jointly consider these key factors. Motivated by these core challenges, we propose the M3-ReID (Multi-View & Granularity & Modality) method, a unified framework that simultaneously enhances spatio-temporal feature extraction, intra-class discrimination, and cross-modality consistency. Specifically, to capture diverse spatio-temporal patterns, we design a Multi-View Learning module that leverages different spatial and temporal-spatial perspectives to adaptively emphasize diverse key regions and motion cues. To enhance intra-class modeling of each identity, we introduce a Multi-Granularity Representation strategy that optimizes features across both fine-grained frame level and coarse-grained video level by minimizing mutual information among redundant frames while enhancing identity representations. Furthermore, to bridge the visible-infrared gap, we propose a Multi-Modality Alignment mechanism that explicitly aligns metric learning and cross-modality matching goals, transforming features into a unified embedding space with modality consistency and class discrimination. Extensive experiments on benchmark VVI-ReID datasets demonstrate the superiority of our proposed M3-ReID framework against existing methods. Tengfei Liang, Yi Jin 0001, Zhun Zhong, Xin Chen 0003, Xianjia Meng, Tao Wang 0011, Yidong Li |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Person Re-Identification With Arbitrary Modalities: A Multi-Modal Dataset and a Unified FrameworkabstractThis paper proposes a unified visual person re-identification (re-id) framework capable of handling various re-id tasks, including modal-fusion re-id, cross-modal re-id, and single-modal re-id, to accommodate diverse modal scenarios. We begin by constructing a Multi-modal Person Re-identification (MPR) dataset comprising RGB, infrared (IR), and depth modalities. Then, the unified re-id framework is established by integrating an Adaptive Modality Aggregation Module (AMAM) and Multi-modal Auto-aligned Learning (MAL). The former autonomously aggregates distinct modalities by thoroughly exploring their relationships. It not only benefits modal-fusion re-id by promoting the modal-fusion representations, but also enhances cross-modal re-id by performing modal consistency learning on the modal-fusion features to narrow modal gaps. The latter automatically aligns multiple modalities through contrastive learning constraints to lessen modal gaps for multiple cross-modal re-id tasks. So, these two modules respectively balance the tasks of distinct types and various tasks of the same type, which are beneficial to realize more re-id tasks with diverse modal scenarios. Moreover, we evaluate state-of-the-art (SOTA) multi-modal methods in terms of plentiful testing settings constructed on MPR dataset. The experiments demonstrate that the proposed unified method that only needs to be trained once outperforms existing methods that require multiple training processes with specific modalities. Besides, it can cope with more scenarios. Extensive ablation studies investigate the effects of the proposed modules on all re-id tasks. Our datasets and code will be publicly available soon: https://github.com/hfutwujingjing/A-Multi-Modal-Dataset-and-A-Unified-Framework. Jingjing Wu 0001, Zhun Zhong, Yanrong Guo, Shejiao Hu, Richang Hong |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | A Self-Adaptive Feature Extraction Method for Aerial-View Geo-LocalizationabstractCross-view geo-localization aims to match the same geographic location from different view images, e.g., drone-view images and geo-referenced satellite-view images. Due to UAV cameras' different shooting angles and heights, the scale of the same captured target building in the drone-view images varies greatly. Meanwhile, there is a difference in size and floor area for different geographic locations in the real world, such as towers and stadiums, which also leads to scale variants of geographic targets in the images. However, existing methods mainly focus on extracting the fine-grained information of the geographic targets or the contextual information of the surrounding area, which overlook the robust feature for scale changes and the importance of feature alignment. In this study, we argue that the key underpinning of this task is to train a network to mine a discriminative representation against scale variants. To this end, we design an effective and novel end-to-end network called Self-Adaptive Feature Extraction Network (Safe-Net) to extract powerful scale-invariant features in a self-adaptive manner. Safe-Net includes a global representation-guided feature alignment module and a saliency-guided feature partition module. The former applies an affine transformation guided by the global feature for adaptive feature alignment. Without extra region annotations, the latter computes saliency distribution for different regions of the image and adopts the saliency information to guide a self-adaptive feature partition on the feature map to learn a visual representation against scale variants. Experiments on two prevailing large-scale aerial-view geo-localization benchmarks, i.e., University-1652 and SUES-200, show that the proposed method achieves state-of-the-art results. In addition, our proposed Safe-Net has a significant scale adaptive capability and can extract robust feature representations for those query images with small target buildings. The source code of this study is available at: https://github.com/AggMan96/Safe-Net. Jinliang Lin, Zhiming Luo, Dazhen Lin, Shaozi Li, Zhun Zhong |
IEEE Trans. Image Process. | 5 |
| 2024 | Diversity-Authenticity Co-constrained Stylization for Federated Domain Generalization in Person Re-identificationabstractThis paper tackles the problem of federated domain generalization in person re-identification (FedDG re-ID), aiming to learn a model generalizable to unseen domains with decentralized source domains. Previous methods mainly focus on preventing local overfitting. However, the direction of diversifying local data through stylization for model training is largely overlooked. This direction is popular in domain generalization but will encounter two issues under federated scenario: (1) Most stylization methods require the centralization of multiple domains to generate novel styles but this is not applicable under decentralized constraint. (2) The authenticity of generated data cannot be ensured especially given limited local data, which may impair the model optimization. To solve these two problems, we propose the Diversity-Authenticity Co-constrained Stylization (DACS), which can generate diverse and authentic data for learning robust local model. Specifically, we deploy a style transformation model on each domain to generate novel data with two constraints: (1) A diversity constraint is designed to increase data diversity, which enlarges the Wasserstein distance between the original and transformed data; (2) An authenticity constraint is proposed to ensure data authenticity, which enforces the transformed data to be easily/hardly recognized by the local-side global/local model. Extensive experiments demonstrate the effectiveness of the proposed DACS and show that DACS achieves state-of-the-art performance for FedDG re-ID. Fengxiang Yang, Zhun Zhong, Zhiming Luo, Yifan He 0002, Shaozi Li, Nicu Sebe |
AAAI | 2 |
| 2024 | Frequency Decoupling for Motion Magnification Via Multi-Level Isomorphic ArchitectureabstractVideo Motion Magnification (VMM) aims to reveal subtle and imperceptible motion information of objects in the macroscopic world. Prior methods directly model the motion field from the Eulerian perspective by Representation Learning that separates shape and texture or Multi-domain Learning from phase fluctuations. Inspired by the frequency spectrum, we observe that the low-frequency components with stable energy always possess spatial structure and less noise, making them suitable for modeling the subtle motion field. To this end, we present FD4MM, a new paradigm of Frequency Decoupling for Motion Magnification with a Multi-level Isomorphic Architecture to capture multi-level high-frequency details and a stable low-frequency structure (motion field) in video space. Since high-frequency details and subtle motions are susceptible to information degra-dation due to their inherent subtlety and unavoidable ex-ternal interference from noise, we carefully design Sparse High/Low-pass Filters to enhance the integrity of details and motion structures, and a Sparse Frequency Mixer to promote seamless recoupling. Besides, we innovatively design a contrastive regularization for this task to strengthen the model's ability to discriminate irrelevant features, re-ducing undesired motion magnification. Extensive experiments on both Real-world and Synthetic Datasets show that our FD4MM outperforms SOTA methods. Meanwhile, FD4MM reduces FLOPs by 1.63x and boosts inference speed by 1.68x than the latest method. Our code is available at https://github.com/Jiafei127/FD4MM. Fei Wang 0067, Dan Guo 0001, Kun Li 0008, Zhun Zhong, Meng Wang 0001 |
CVPR | 4 |
| 2024 | Active Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) is a pragmatic and challenging open-world task, which endeavors to cluster unlabeled samples from both novel and old classes, leveraging some labeled data of old classes. Given that knowledge learned from old classes is not fully transferable to new classes, and that novel categories are fully unlabeled, GCD inherently faces intractable problems, including imbalanced classification performance and inconsistent confidence between old and new classes, especially in the low-labeling regime. Hence, some annotations of new classes are deemed necessary. However, labeling new classes is extremely costly. To address this issue, we take the spirit of active learning and propose a new setting called Active Generalized Category Discovery (AGCD). The goal is to improve the performance of GCD by actively selecting a limited amount of valuable samples for labeling from the oracle. To solve this problem, we devise an adaptive sampling strategy, which jointly considers novelty, informativeness and diversity to adaptively select novel samples with proper uncertainty. However, owing to the varied orderings of label indices caused by the clustering of novel classes, the queried labels are not directly applicable to subsequent training. To overcome this issue, we further propose a stable label mapping algorithm that transforms ground truth labels to the label space of the classifier, thereby ensuring consistent training across different active selection stages. Our method achieves state-of-the-art performance on both generic and fine-grained datasets. Our code is available at https://github.com/mashijie1028/ActiveGCD Shijie Ma, Fei Zhu 0004, Zhun Zhong, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 3 |
| 2024 | Federated Generalized Category DiscoveryabstractGeneralized category discovery (GCD) aims at grouping unlabeled samples from known and unknown classes, given labeled data of known classes. To meet the recent decen-tralization trend in the community, we introduce a practical yet challenging task, Federated GCD (Fed-GCD), where the training data are distributed among local clients and cannot be shared among clients. Fed-GCD aims to train a generic GCD model by client collaboration under the privacy-protected constraint. The Fed-GCD leads to two challenges: 1) representation degradation caused by training each client model with fewer data than centralized GCD learning, and 2) highly heterogeneous label spaces across different clients. To this end, we propose a novel Asso-ciated Gaussian Contrastive Learning (AGCL) framework based on learnable GMMs, which consists of a Client Se-mantics Association (CSA) and a global-local GMM Contrastive Learning (GCL). On the server, CSA aggregates the heterogeneous categories of local-client GMMs to generate a global GMM containing more comprehensive category knowledge. On each client, GCL builds class-level contrastive learning with both local and global GMMs. The local GCL learns robust representation with limited local data. The global GCL encourages the model to produce more discriminative representation with the comprehensive category relationships that may not exist in local data. We build a benchmark based on six visual datasets to facilitate the study of Fed-GCD. Extensive experiments show that our AGCL outperforms multiple baselines on all datasets. Code is available at https://github.com/TPCD/FedGCD. Nan Pu, Wenjing Li 0005, Xingyuan Ji, Yalan Qin, Nicu Sebe, Zhun Zhong |
CVPR | 6 |
| 2024 | Stable Neighbor Denoising for Source-free Domain Adaptive SegmentationabstractWe study source-free unsupervised domain adaptation (SFUDA) for semantic segmentation, which aims to adapt a source-trained model to the target domain without accessing the source data. Many works have been proposed to address this challenging problem, among which uncertainty-based self-training is a predominant approach. However, without comprehensive denoising mechanisms, they still largely fall into biased estimates when dealing with different domains and confirmation bias. In this paper, we observe that pseudo-label noise is mainly contained in unstable samples in which the predictions of most pixels undergo significant variations during self-training. Inspired by this, we propose a novel mechanism to denoise unstable samples with stable ones. Specifically, we introduce the Stable Neighbor Denoising (SND) approach, which effectively discovers highly correlated stable and unstable samples by nearest neighbor retrieval and guides the reliable optimization of unstable samples by bi-level learning. Moreover, we compensate for the stable set by object-level object paste, which can further eliminate the bias caused by less learned classes. Our SND enjoys two advantages. First, SND does not require a specific segmentor structure, endowing its universality. Second, SND simultaneously addresses the issues of class, domain, and confirmation biases during adaptation, ensuring its effectiveness. Extensive experiments show that SND consistently outperforms state-of-the-art methods in various SFUDA semantic segmentation settings. In addition, SND can be easily integrated with other approaches, obtaining further improvements. The source code is available at https://github.com/DZhaoXd/SND. Dong Zhao 0007, Shuang Wang 0001, Qi Zang, Licheng Jiao, Nicu Sebe, Zhun Zhong |
CVPR | 6 |
| 2024 | ReMamber: Referring Image Segmentation with Mamba Twister
Yuhuan Yang, Chaofan Ma, Jiangchao Yao, Zhun Zhong, Ya Zhang 0002, Yanfeng Wang 0001 |
ECCV (10) | 4 |
| 2024 | Learning to Distinguish Samples for Generalized Category Discovery
Fengxiang Yang, Nan Pu, Wenjing Li 0005, Zhiming Luo, Shaozi Li, Nicu Sebe, Zhun Zhong |
ECCV (65) | 7 |
| 2024 | Textual Knowledge Matters: Cross-Modality Co-teaching for Generalized Visual Class Discovery
Haiyang Zheng, Nan Pu, Wenjing Li 0005, Nicu Sebe, Zhun Zhong |
ECCV (52) | 5 |
| 2024 | Democratizing Fine-grained Visual Recognition with Large Language ModelsabstractIdentifying subordinate-level categories from images is a longstanding task in computer vision and is referred to as fine-grained visual recognition (FGVR). It has tremendous significance in real-world applications since an average layperson does not excel at differentiating species of birds or mushrooms due to subtle differences among the species. A major bottleneck in developing FGVR systems is caused by the need of high-quality paired expert annotations. To circumvent the need of expert knowledge we propose Fine-grained Semantic Category Reasoning (FineR) that internally leverages the world knowledge of large language models (LLMs) as a proxy in order to reason about fine-grained category names. In detail, to bridge the modality gap between images and LLM, we extract part-level visual attributes from images as text and feed that information to a LLM. Based on the visual attributes and its internal world knowledge the LLM reasons about the subordinate-level category names. Our training-free FineR outperforms several state-of-the-art FGVR and language and vision assistant models and shows promise in working in the wild and in new domains where gathering expert annotation is arduous. Subhankar Roy, Wenjing Li 0005, Zhun Zhong, Nicu Sebe, Elisa Ricci 0001 |
ICLR | 4 |
| 2024 | Envisioning Outlier Exposure by Large Language Models for Out-of-Distribution DetectionabstractDetecting out-of-distribution (OOD) samples is essential when deploying machine learning models in open-world scenarios. Zero-shot OOD detection, requiring no training on in-distribution (ID) data, has been possible with the advent of vision-language models like CLIP. Existing methods build a text-based classifier with only closed-set labels. However, this largely restricts the inherent capability of CLIP to recognize samples from large and open label space. In this paper, we propose to tackle this constraint by leveraging the expert knowledge and reasoning capability of large language models (LLM) to Envision potential Outlier Exposure, termed EOE, without access to any actual OOD data. Owing to better adaptation to open-world scenarios, EOE can be generalized to different tasks, including far, near, and fine-grained OOD detection. Technically, we design (1) LLM prompts based on visual similarity to generate potential outlier class labels specialized for OOD detection, as well as (2) a new score function based on potential outlier penalty to distinguish hard OOD samples effectively. Empirically, EOE achieves state-of-the-art performance across different OOD tasks and can be effectively scaled to the ImageNet-1K dataset. The code is publicly available at: https://github.com/tmlr-group/EOE. Chentao Cao, Zhun Zhong, Zhanke Zhou, Yang Liu 0018, Tongliang Liu, Bo Han 0003 |
ICML | 2 |
| 2024 | Large-Scale Pre-trained Models are Surprisingly Strong in Incremental Novel Class Discovery
Subhankar Roy, Zhun Zhong, Nicu Sebe, Elisa Ricci 0001 |
ICPR (16) | 3 |
| 2024 | Mitigating robust overfitting via self-residual-calibration regularization (Abstract Reprint)
Hong Liu 0009, Zhun Zhong, Nicu Sebe, Shin'ichi Satoh 0001 |
IJCAI | 2 |
| 2024 | MORE'24 Multimedia Object Re-ID: Advancements, Challenges, and OpportunitiesabstractObject re-identification (or object re-id) has gained significant attention in recent years, fueled by the increasing demand for advanced video analysis and safety systems. In object re-id, a query can be of different modalities, such as an image, a video, or natural language, containing or describing the object of interest. This workshop aims to bring together researchers, practitioners, and enthusiasts interested in object re-id to delve into the latest advancements, challenges, and opportunities in this dynamic field. The workshop covers a spectrum of topics related to object re-id, including but not limited to deep metric learning, multi-view data generation, video-based object re-id, cross-domain object re-id and real-world applications. The workshop provides a platform for researchers to showcase their work, exchange ideas, and foster potential collaborations. Additionally, it serves as a valuable opportunity for practitioners to stay abreast of the latest developments in object re-id technology. Zhedong Zheng, Yaxiong Wang, Xuelin Qian, Zhun Zhong, Zheng Wang 0007, Liang Zheng 0001 |
ICMR | 4 |
| 2024 | Generalized Source-Free Domain-adaptive Segmentation via Reliable Knowledge Propagation
Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Dou Quan, Jinlong Li 0003, Nicu Sebe, Zhun Zhong |
ACM Multimedia | 8 |
| 2024 | Cross-Modality Perturbation Synergy Attack for Person Re-identificationabstractIn recent years, there has been significant research focusing on addressing security concerns in single-modal person re-identification (ReID) systems that are based on RGB images. However, the safety of cross-modality scenarios, which are more commonly encountered in practical applications involving images captured by infrared cameras, has not received adequate attention. The main challenge in cross-modality ReID lies in effectively dealing with visual differences between different modalities. For instance, infrared images are typically grayscale, unlike visible images that contain color information. Existing attack methods have primarily focused on the characteristics of the visible image modality, overlooking the features of other modalities and the variations in data distribution among different modalities. This oversight can potentially undermine the effectiveness of these methods in image retrieval across diverse modalities. This study represents the first exploration into the security of cross-modality ReID models and proposes a universal perturbation attack specifically designed for cross-modality ReID. This attack optimizes perturbations by leveraging gradients from diverse modality data, thereby disrupting the discriminator and reinforcing the differences between modalities. We conducted experiments on three widely used cross-modality datasets, namely RegDB, SYSU, and LLCM. The results not only demonstrate the effectiveness of our method but also provide insights for future improvements in the robustness of cross-modality ReID systems. Yunpeng Gong, Zhun Zhong, Yansong Qu, Zhiming Luo, Rongrong Ji, Min Jiang 0005 |
NeurIPS | 2 |
| 2024 | Happy: A Debiased Learning Framework for Continual Generalized Category DiscoveryabstractConstantly discovering novel concepts is crucial in evolving environments. This paper explores the underexplored task of Continual Generalized Category Discovery (C-GCD), which aims to incrementally discover new classes from *unlabeled* data while maintaining the ability to recognize previously learned classes. Although several settings are proposed to study the C-GCD task, they have limitations that do not reflect real-world scenarios. We thus study a more practical C-GCD setting, which includes more new classes to be discovered over a longer period, without storing samples of past classes. In C-GCD, the model is initially trained on labeled data of known classes, followed by multiple incremental stages where the model is fed with unlabeled data containing both old and new classes. The core challenge involves two conflicting objectives: discover new classes and prevent forgetting old ones. We delve into the conflicts and identify that models are susceptible to *prediction bias* and *hardness bias*. To address these issues, we introduce a debiased learning framework, namely **Happy**, characterized by **H**ardness-**a**ware **p**rototype sampling and soft entro**py** regularization. For the *prediction bias*, we first introduce clustering-guided initialization to provide robust features. In addition, we propose soft entropy regularization to assign appropriate probabilities to new classes, which can significantly enhance the clustering performance of new classes. For the *harness bias*, we present the hardness-aware prototype sampling, which can effectively reduce the forgetting issue for previously seen classes, especially for difficult classes. Experimental results demonstrate our method proficiently manages the conflicts of C-GCD and achieves remarkable performance across various datasets, e.g., 7.5% overall gains on ImageNet-100. Our code is publicly available at https://github.com/mashijie1028/Happy-CGCD. Shijie Ma, Fei Zhu 0004, Zhun Zhong, Wenzhuo Liu, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
NeurIPS | 3 |
| 2024 | Connectivity-Driven Pseudo-Labeling Makes Stronger Cross-Domain SegmentersabstractPresently, pseudo-labeling stands as a prevailing approach in cross-domain semantic segmentation, enhancing model efficacy by training with pixels assigned with reliable pseudo-labels. However, we identify two key limitations within this paradigm: (1) under relatively severe domain shifts, most selected reliable pixels appear speckled and remain noisy. (2) when dealing with wild data, some pixels belonging to the open-set class may exhibit high confidence and also appear speckled. These two points make it difficult for the pixel-level selection mechanism to identify and correct these speckled close- and open-set noises. As a result, error accumulation is continuously introduced into subsequent self-training, leading to inefficiencies in pseudo-labeling. To address these limitations, we propose a novel method called Semantic Connectivity-driven Pseudo-labeling (SeCo). SeCo formulates pseudo-labels at the connectivity level, which makes it easier to locate and correct closed and open set noise. Specifically, SeCo comprises two key components: Pixel Semantic Aggregation (PSA) and Semantic Connectivity Correction (SCC). Initially, PSA categorizes semantics into ``stuff'' and ``things'' categories and aggregates speckled pseudo-labels into semantic connectivity through efficient interaction with the Segment Anything Model (SAM). This enables us not only to obtain accurate boundaries but also simplifies noise localization. Subsequently, SCC introduces a simple connectivity classification task, which enables us to locate and correct connectivity noise with the guidance of loss distribution. Extensive experiments demonstrate that SeCo can be flexibly applied to various cross-domain semantic segmentation tasks, \textit{i.e.} domain generalization and domain adaptation, even including source-free, and black-box domain adaptation, significantly improving the performance of existing state-of-the-art methods. The code is provided in the appendix and will be open-source. Dong Zhao 0007, Qi Zang, Shuang Wang 0001, Nicu Sebe, Zhun Zhong |
NeurIPS | 5 |
| 2024 | Prototypical Hash Encoding for On-the-Fly Fine-Grained Category DiscoveryabstractIn this paper, we study a practical yet challenging task, On-the-fly Category Discovery (OCD), aiming to online discover the newly-coming stream data that belong to both known and unknown classes, by leveraging only known category knowledge contained in labeled data. Previous OCD methods employ the hash-based technique to represent old/new categories by hash codes for instance-wise inference. However, directly mapping features into low-dimensional hash space not only inevitably damages the ability to distinguish classes and but also causes ``high sensitivity'' issue, especially for fine-grained classes, leading to inferior performance. To address these drawbacks, we propose a novel Prototypical Hash Encoding (PHE) framework consisting of Category-aware Prototype Generation (CPG) and Discriminative Category Encoding (DCE) to mitigate the sensitivity of hash code while preserving rich discriminative information contained in high-dimension feature space, in a two-stage projection fashion. CPG enables the model to fully capture the intra-category diversity by representing each category with multiple prototypes. DCE boosts the discrimination ability of hash code with the guidance of the generated category prototypes and the constraint of minimum separation distance. By jointly optimizing CPG and DCE, we demonstrate that these two components are mutually beneficial towards an effective OCD. Extensive experiments show the significant superiority of our PHE over previous methods, e.g. obtaining an improvement of +5.3% in ALL ACC averaged on all datasets. Moreover, due to the nature of the interpretable prototypes, we visually analyze the underlying mechanism of how PHE helps group certain samples into either known or unknown categories. Code is available at https://github.com/HaiyangZheng/PHE. Haiyang Zheng, Nan Pu, Wenjing Li 0005, Nicu Sebe, Zhun Zhong |
NeurIPS | 5 |
| 2024 | Style-Hallucinated Dual Consistency Learning: A Unified Framework for Visual Domain Generalization
Zhun Zhong, Na Zhao 0004, Nicu Sebe, Gim Hee Lee |
Int. J. Comput. Vis. | 2 |
| 2024 | Bridge Gap in Pixel and Feature Level for Cross-Modality Person Re-IdentificationabstractVisible thermal person re-identification (VT-ReID) plays a vital role in intelligent surveillance systems, particularly in weak lighting environments. VT-ReID faces substantial challenges, including the cross-modality gap and intra-class variations. Existing methods address these challenges through either pixel-level image translation techniques or feature-level metric learning techniques. However, the former approaches require additional computational costs and often generate noisy images, making model training challenging. The latter methods focus on constraining the relations between individual instances or class centers, while often ignoring joint consideration of the relationship between the two aspects. In addition, these works do not fully investigate the mutual benefits at both pixel-level and feature-level. To address these limitations, we propose a unified Dual-level Smooth Gap (DSG) learning framework that simultaneously smooths the cross-modality gap at the pixel and feature levels. Specifically, on the one hand, we develop a parameter-free Class-aware Modality Mix (CMM) to smooth the cross-modality gap at the pixel level. CMM can capture and explore internal information between the two modalities by mixing images from different modalities belonging to the same class. On the other hand, we devise an efficient Center-guided Metric Learning (CML) to reduce the inter-modality discrepancy and intra-class variations at the feature level. CML enhances model discrimination and generalization by enforcing constraints on both class centers and instances. Experiments on two benchmark datasets demonstrate the mutual benefits of our proposed and show the superior performance of our method over state-of-the-art methods. Yongguo Ling, Zhun Zhong, Zhiming Luo, Shaozi Li, Nicu Sebe |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Boosting Adversarial Transferability via Logits Mixup With Dominant Decomposed FeatureabstractRecent research has shown that adversarial samples are highly transferable and can be used to attack other unknown black-box Deep Neural Networks (DNNs). To improve the transferability of adversarial samples, several feature-based adversarial attack methods have been proposed to disrupt neuron activation in the middle layers. However, current state-of-the-art feature-based attack methods typically require additional computation costs for estimating the importance of neurons. To address this challenge, we propose a Singular Value Decomposition (SVD)-based feature-level attack method. Our approach is inspired by the discovery that eigenvectors associated with the larger singular values decomposed from the middle layer features exhibit superior generalization and attention properties. Specifically, we conduct the attack by retaining the dominant decomposed feature that corresponds to the largest singular value (i.e., Rank-1 decomposed feature) for computing the output logits before the final softmax. These logits are later integrated with the original logits to optimize adversarial examples. Our extensive experimental results verify the effectiveness of our proposed method, which can be easily integrated into various baselines to significantly enhance the transferability of adversarial samples for disturbing normally trained CNNs and advanced defense strategies. The source code is available at Link. Juanjuan Weng, Zhiming Luo, Shaozi Li, Dazhen Lin, Zhun Zhong |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Depth Matters: Spatial Proximity-Based Gaze Cone Generation for Gaze Following in WildabstractGaze following aims to predict where a person is looking in a scene. Existing methods tend to prioritize traditional 2D RGB visual cues or require burdensome prior knowledge and extra expensive datasets annotated in 3D coordinate systems to train specialized modules to enhance scene modeling. In this work, we introduce a novel framework deployed on a simple ResNet backbone, which exclusively uses image and depth maps to mimic human visual preferences and realize 3D-like depth perception. We first leverage depth maps to formulate spatial-based proximity information regarding the objects with the target person. This process sharpens the focus of the gaze cone on the specific region of interest pertaining to the target while diminishing the impact of surrounding distractions. To capture the diverse dependence of scene context on the saliency gaze cone, we then introduce a learnable grid-level regularized attention that anticipates coarse-grained regions of interest, thereby refining the mapping of the saliency feature to pixel-level heatmaps. This allows our model to better account for individual differences when predicting others’ gaze locations. Finally, we employ the KL-divergence loss to super the grid-level regularized attention, which combines the gaze direction, heatmap regression, and in/out classification losses, providing comprehensive supervision for model optimization. Experimental results on two publicly available datasets demonstrate the comparable performance of our model with less help of modal information. Quantitative visualization results further validate the interpretability of our method. The source code will be available at https://github.com/VUT-HFUT/DepthMatters . Kun Li 0008, Zhun Zhong, Wei Jia 0001, Bin Hu 0001, Xun Yang 0001, Meng Wang 0001, Dan Guo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Multi-View Domain Adaptive Object Detection on Camera NetworksabstractIn this paper, we study a new domain adaptation setting on camera networks, namely Multi-View Domain Adaptive Object Detection (MVDA-OD), in which labeled source data is unavailable in the target adaptation process and target data is captured from multiple overlapping cameras. In such a challenging context, existing methods including adversarial training and self-training fall short due to multi-domain data shift and the lack of source data. To tackle this problem, we propose a novel training framework consisting of two stages. First, we pre-train the backbone using self-supervised learning, in which a multi-view association is developed to construct an effective pretext task. Second, we fine-tune the detection head using robust self-training, where a tracking-based single-view augmentation is introduced to achieve weak-hard consistency learning. By doing so, an object detection model can take advantage of informative samples generated by multi-view association and single-view augmentation to learn discriminative backbones as well as robust detection classifiers. Experiments on two real-world multi-camera datasets demonstrate significant advantages of our approach over the state-of-the-art domain adaptive object detection methods. Yan Lu 0006, Zhun Zhong, Yuanchao Shu |
AAAI | 2 |
| 2023 | Cross-Modality Earth Mover's Distance for Visible Thermal Person Re-identificationabstractVisible thermal person re-identification (VT-ReID) suffers from inter-modality discrepancy and intra-identity variations. Distribution alignment is a popular solution for VT-ReID, however, it is usually restricted to the influence of the intra-identity variations. In this paper, we propose the Cross-Modality Earth Mover's Distance (CM-EMD) that can alleviate the impact of the intra-identity variations during modality alignment. CM-EMD selects an optimal transport strategy and assigns high weights to pairs that have a smaller intra-identity variation. In this manner, the model will focus on reducing the inter-modality discrepancy while paying less attention to intra-identity variations, leading to a more effective modality alignment. Moreover, we introduce two techniques to improve the advantage of CM-EMD. First, Cross-Modality Discrimination Learning (CM-DL) is designed to overcome the discrimination degradation problem caused by modality alignment. By reducing the ratio between intra-identity and inter-identity variances, CM-DL leads the model to learn more discriminative representations. Second, we construct the Multi-Granularity Structure (MGS), enabling us to align modalities from both coarse- and fine-grained levels with the proposed CM-EMD. Extensive experiments show the benefits of the proposed CM-EMD and its auxiliary techniques (CM-DL and MGS). Our method achieves state-of-the-art performance on two VT-ReID benchmarks. Yongguo Ling, Zhun Zhong, Zhiming Luo, Fengxiang Yang, Donglin Cao, Yaojin Lin, Shaozi Li, Nicu Sebe |
AAAI | 2 |
| 2023 | Exploring Non-target Knowledge for Improving Ensemble Universal Adversarial AttacksabstractThe ensemble attack with average weights can be leveraged for increasing the transferability of universal adversarial perturbation (UAP) by training with multiple Convolutional Neural Networks (CNNs). However, after analyzing the Pearson Correlation Coefficients (PCCs) between the ensemble logits and individual logits of the crafted UAP trained by the ensemble attack, we find that one CNN plays a dominant role during the optimization. Consequently, this average weighted strategy will weaken the contributions of other CNNs and thus limit the transferability for other black-box CNNs. To deal with this bias issue, the primary attempt is to leverage the Kullback–Leibler (KL) divergence loss to encourage the joint contribution from different CNNs, which is still insufficient. After decoupling the KL loss into a target-class part and a non-target-class part, the main issue lies in that the non-target knowledge will be significantly suppressed due to the increasing logit of the target class. In this study, we simply adopt a KL loss that only considers the non-target classes for addressing the dominant bias issue. Besides, to further boost the transferability, we incorporate the min-max learning framework to self-adjust the ensemble weights for each CNN. Experiments results validate that considering the non-target KL loss can achieve superior transferability than the original KL loss by a large margin, and the min-max training can provide a mutual benefit in adversarial ensemble attacks. The source code is available at: https://github.com/WJJLL/ND-MM. Juanjuan Weng, Zhiming Luo, Zhun Zhong, Dazhen Lin, Shaozi Li |
AAAI | 3 |
| 2023 | Dynamic Conceptional Contrastive Learning for Generalized Category DiscoveryabstractGeneralized category discovery (GCD) is a recently proposed open-world problem, which aims to automatically cluster partially labeled data. The main challenge is that the unlabeled data contain instances that are not only from known categories of the labeled data but also from novel categories. This leads traditional novel category discovery (NCD) methods to be incapacitated for GCD, due to their assumption of unlabeled data are only from novel categories. One effective way for GCD is applying self-supervised learning to learn discriminate representation for unlabeled data. However, this manner largely ignores underlying relationships between instances of the same concepts (e.g., class, super-class, and sub-class), which results in inferior representation learning. In this paper, we propose a Dynamic Conceptional Contrastive Learning (DCCL)framework, which can effectively improve clustering accuracy by alternately estimating underlying visual conceptions and learning conceptional representation. In addition, we design a dynamic conception generation and update mechanism, which is able to ensure consistent conception learning and thus further facilitate the optimization of DCCL. Extensive experiments show that DCCL achieves new state-of-the-art performances on six generic and fine-grained visual recognition datasets, especially on fine-grained ones. For example, our method significantly surpasses the best competitor by 16.2% on the new classes for the CUB-200 dataset. Code is available at https://github.com/TPCD/DCCL Nan Pu, Zhun Zhong, Nicu Sebe |
CVPR | 2 |
| 2023 | Dynamically Instance-Guided Adaptation: A Backward-free Approach for Test-Time Domain Adaptive Semantic SegmentationabstractIn this paper, we study the application of Test-time domain adaptation in semantic segmentation (TTDA-Seg) where both efficiency and effectiveness are crucial. Existing methods either have low efficiency (e.g., backward optimization) or ignore semantic adaptation (e.g., distribution alignment). Besides, they would suffer from the accumulated errors caused by unstable optimization and abnormal distributions. To solve these problems, we propose a novel backward-free approach for TTDA-Seg, called Dynamically Instance-Guided Adaptation (DIGA). Our principle is utilizing each instance to dynamically guide its own adaptation in a non-parametric way, which avoids the error accumulation issue and expensive optimizing cost. Specifically, DIGA is composed of a distribution adaptation module (DAM) and a semantic adaptation module (SAM), enabling us to jointly adapt the model in two indispensable aspects. DAM mixes the instance and source BN statistics to encourage the model to capture robust representation. SAM combines the historical prototypes with instance-level prototypes to adjust semantic predictions, which can be associated with the parametric classifier to mutually benefit the final results. Extensive experiments evaluated on five target domains demonstrate the effectiveness and efficiency of the proposed method. Our DIGA establishes new state-of-the-art performance in TTDA-Seg. Source code is available at: https://github.com/Waybaba/DIGA. Zhun Zhong, Weijie Wang 0002, Charles Ling 0001, Boyu Wang 0004, Nicu Sebe |
CVPR | 2 |
| 2023 | Sparsely Annotated Semantic Segmentation with Adaptive Gaussian MixturesabstractSparsely annotated semantic segmentation (SASS) aims to learn a segmentation model by images with sparse labels (i.e., points or scribbles). Existing methods mainly focus on introducing low-level affinity or generating pseudo labels to strengthen supervision, while largely ignoring the inherent relation between labeled and unlabeled pixels. In this paper, we observe that pixels that are close to each other in the feature space are more likely to share the same class. Inspired by this, we propose a novel SASS framework, which is equipped with an Adaptive Gaussian Mixture Model (AGMM). Our AGMM can effectively endow reliable supervision for unlabeled pixels based on the distributions of labeled and unlabeled pixels. Specifically, we first build Gaussian mixtures using labeled pixels and their relatively similar unlabeled pixels, where the labeled pixels act as centroids, for modeling the feature distribution of each class. Then, we leverage the reliable information from labeled pixels and adaptively generated GMM predictions to supervise the training of unlabeled pixels, achieving online, dynamic, and robust selfsupervision. In addition, by capturing category-wise Gaussian mixtures, AGMM encourages the model to learn discriminative class decision boundaries in an end-to-end contrastive learning manner. Experimental results conducted on the PASCAL VOC 2012 and Cityscapes datasets demonstrate that our AGMM can establish new state-of-the-art SASS performance. Code is available at https://github.com/Luffy03/AGMM-SASS Linshan Wu, Zhun Zhong, Leyuan Fang, Xingxin He, Jiayi Ma 0001, Hao Chen 0011 |
CVPR | 2 |
| 2023 | Multi-Domain Lifelong Visual Question Answering via Self-Critical DistillationabstractVisual Question Answering (VQA) has achieved significant success over the last few years, while most studies focus on training a VQA model on a stationary domain (e.g., a given dataset). In real-world application scenarios, however, these methods are often inefficient because VQA systems are always supposed to extend their knowledge and meet the ever-changing demands of users. In this paper, we introduce a new and challenging multi-domain lifelong VQA task, dubbed MDL-VQA, which encourages the VQA model to continuously learn across multiple domains while mitigating the forgetting on previously-learned domains. Furthermore, we propose a novel replay-free Self-Critical Distillation (SCD) framework tailor-made for MDL-VQA, which alleviates forgetting issue via transferring previous-domain knowledge from teacher to student models. First, we propose to introspect the teacher's understanding over original and counterfactual samples, thereby creating informative instance-relevant and domain-relevant knowledge for logits-based distillation. Second, on the side of feature-based distillation, we propose to introspect the reasoning behavior of student model to establish the harmful domain-specific knowledge acquired in current domain, and further leverage the metric learning strategy to encourage student to learn useful knowledge in new domain. Extensive experiments demonstrate that SCD framework outperforms state-of-the-art competitors with different training orders. Mingrui Lao, Nan Pu, Yu Liu 0012, Zhun Zhong, Erwin M. Bakker, Nicu Sebe, Michael S. Lew |
ACM Multimedia | 4 |
| 2023 | FedVQA: Personalized Federated Visual Question Answering over Heterogeneous ScenesabstractThis paper presents a new setting for visual question answering (VQA) called personalized federated VQA (FedVQA) that addresses the growing need for decentralization and data privacy protection. FedVQA is both practical and challenging, requiring clients to learn well-personalized models on scene-specific datasets with severe feature/label distribution skews. These models then collaborate to optimize a generic global model on a central server, which is desired to generalize well on both seen and unseen scenes without sharing raw data with the server and other clients. The primary challenge of FedVQA is that, client models tend to forget the global knowledge initialized from central server during the personalized training, which impairs their personalized capacity due to the potential overfitting issue on local data. This further leads to divergence issues when aggregating distinct personalized knowledge at the central server, resulting in an inferior generalization ability on unseen scenes. To address the challenge, we propose a novel federated pairwise preference preserving (FedP3) framework to improve personalized learning via preserving generic knowledge under FedVQA constraints. Specifically, we first design a differentiable pairwise preference (DPP) to improve knowledge preserving by formulating a flexible yet effective global knowledge. Then, we introduce a forgotten-knowledge filter (FKF) to encourage the client models to selectively consolidate easily-forgotten knowledge. By aggregating the DPP and the FKF, FedP3 coordinates the generic and the personalized knowledge to enhance the personalized ability of clients and generalizability of the server. Extensive experiments show that FedP3 consistently surpasses the state-of-the-art in FedVQA task. Mingrui Lao, Nan Pu, Zhun Zhong, Nicu Sebe, Michael S. Lew |
ACM Multimedia | 3 |
| 2023 | Discover and Align Taxonomic Context Priors for Open-world Semi-Supervised LearningabstractOpen-world Semi-Supervised Learning (OSSL) is a realistic and challenging task, aiming to classify unlabeled samples from both seen and novel classes using partially labeled samples from the seen classes.
Previous works typically explore the relationship of samples as priors on the pre-defined single-granularity labels to help novel class recognition. In fact, classes follow a taxonomy and samples can be classified at multiple levels of granularity, which contains more underlying relationships for supervision. We thus argue that learning with single-granularity labels results in sub-optimal representation learning and inaccurate pseudo labels, especially with unknown classes. In this paper, we take the initiative to explore and propose a uniformed framework, called Taxonomic context prIors Discovering and Aligning (TIDA), which exploits the relationship of samples under various granularity. It allows us to discover multi-granularity semantic concepts as taxonomic context priors (i.e., sub-class, target-class, and super-class), and then collaboratively leverage them to enhance representation learning and improve the quality of pseudo labels.
Specifically, TIDA comprises two components: i) A taxonomic context discovery module that constructs a set of hierarchical prototypes in the latent space to discover the underlying taxonomic context priors; ii) A taxonomic context-based prediction alignment module that enforces consistency across hierarchical predictions to build the reliable relationship between classes among various granularity and provide additions supervision. We demonstrate that these two components are mutually beneficial for an effective OSSL framework, which is theoretically explained from the perspective of the EM algorithm. Extensive experiments on seven commonly used datasets show that TIDA can significantly improve the performance and achieve a new state of the art. The source codes are publicly available at https://github.com/rain305f/TIDA. Yu Wang 0027, Zhun Zhong, Pengchong Qiao, Xuxin Cheng, Xiawu Zheng, Chang Liu 0030, Nicu Sebe, Rongrong Ji, Jie Chen 0001 |
NeurIPS | 2 |
| 2023 | Mitigating robust overfitting via self-residual-calibration regularization
Hong Liu 0009, Zhun Zhong, Nicu Sebe, Shin'ichi Satoh 0001 |
Artif. Intell. | 2 |
| 2023 | Self-training transformer for source-free domain adaptation
Guanglei Yang, Zhun Zhong, Mingli Ding, Nicu Sebe, Elisa Ricci 0001 |
Appl. Intell. | 2 |
| 2023 | A Hard Knowledge Regularization Method with Probability Difference in Thorax Disease Images
Qingji Guan, Qinrun Chen, Zhun Zhong |
Knowl. Based Syst. | 3 |
| 2023 | A Memorizing and Generalizing Framework for Lifelong Person Re-IdentificationabstractIn this paper, we introduce a challenging yet practical setting for person re-identification (ReID) task, named lifelong person re-identification (LReID), which aims to continuously train a ReID model across multiple domains and the trained model is required to generalize well on both seen and unseen domains. It is therefore critical to learn a ReID model that can learn a generalized representation without forgetting knowledge of seen domains. In this paper, we propose a new MEmorizing and GEneralizing framework (MEGE) for LReID, which can jointly prevent the model from forgetting and improve its generalization ability. Specifically, our MEGE is composed of two novel modules, i.e., Adaptive Knowledge Accumulation (AKA) and differentiable Ranking Consistency Distillation (RCD). Taking inspiration from the cognitive processes in the human brain, we endow AKA with two special capacities, knowledge representation and knowledge operation by graph convolution networks. AKA can effectively mitigate catastrophic forgetting on seen domains while improving the generalization ability to unseen domains. By considering the ranking factor that is specifically important in ReID, RCD is designed to distill the ranking knowledge in a differentiable manner, which can further prevent the catastrophic forgetting. To supporting the study of LReID, we build a new and large-scale benchmark with two practical evaluation protocols that consider the metrics of non-forgetting and generalization. Experiments demonstrate that 1) our MEGE framework can effectively improve the performance on seen and unseen domains under the domain-incremental learning constraint, and that 2) the proposed MEGE outperforms state-of-the-art competitors by large margins. Nan Pu, Zhun Zhong, Nicu Sebe, Michael S. Lew |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Querying Labeled for Unlabeled: Cross-Image Semantic Consistency Guided Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation aims to learn a semantic segmentation model via limited labeled images and adequate unlabeled images. The key to this task is generating reliable pseudo labels for unlabeled images. Existing methods mainly focus on producing reliable pseudo labels based on the confidence scores of unlabeled images while largely ignoring the use of labeled images with accurate annotations. In this paper, we propose a Cross-Image Semantic Consistency guided Rectifying (CISC-R) approach for semi-supervised semantic segmentation, which explicitly leverages the labeled images to rectify the generated pseudo labels. Our CISC-R is inspired by the fact that images belonging to the same class have a high pixel-level correspondence. Specifically, given an unlabeled image and its initial pseudo labels, we first query a guiding labeled image that shares the same semantic information with the unlabeled image. Then, we estimate the pixel-level similarity between the unlabeled image and the queried labeled image to form a CISC map, which guides us to achieve a reliable pixel-level rectification for the pseudo labels. Extensive experiments on the PASCAL VOC 2012, Cityscapes, and COCO datasets demonstrate that the proposed CISC-R can significantly improve the quality of the pseudo labels and outperform the state-of-the-art methods. Code is available at https://github.com/Luffy03/CISC-R. Linshan Wu, Leyuan Fang, Xingxin He, Jiayi Ma 0001, Zhun Zhong |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Towards Robust Person Re-Identification by Defending Against Universal AttackersabstractRecent studies show that deep person re-identification (re-ID) models are vulnerable to adversarial examples, so it is critical to improving the robustness of re-ID models against attacks. To achieve this goal, we explore the strengths and weaknesses of existing re-ID models, i.e., designing learning-based attacks and training robust models by defending against the learned attacks. The contributions of this paper are three-fold: First, we build a holistic attack-defense framework to study the relationship between the attack and defense for person re-ID. Second, we introduce a combinatorial adversarial attack that is adaptive to unseen domains and unseen model types. It consists of distortions in pixel and color space (i.e., mimicking camera shifts). Third, we propose a novel virtual-guided meta-learning algorithm for our attack-defense system. We leverage a virtual dataset to conduct experiments under our meta-learning framework, which can explore the cross-domain constraints for enhancing the generalization of the attack and the robustness of the re-ID model. Comprehensive experiments on three large-scale re-ID benchmarks demonstrate that: 1) Our combinatorial attack is effective and highly universal in cross-model and cross-dataset scenarios; 2) Our meta-learning algorithm can be readily applied to different attack and defense approaches, which can reach consistent improvement; 3) The defense model trained on the learning-to-learn framework is robust to recent SOTA attacks that are not even used during training. Fengxiang Yang, Juanjuan Weng, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Zhiming Luo, Donglin Cao, Shaozi Li, Shin'ichi Satoh 0001, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Logit Margin Matters: Improving Transferable Targeted Adversarial Attack by Logit CalibrationabstractPrevious works have extensively studied the transferability of adversarial samples in untargeted black-box scenarios. However, it still remains challenging to craft targeted adversarial examples with higher transferability than non-targeted ones. Recent studies reveal that the traditional Cross-Entropy (CE) loss function is insufficient to learn transferable targeted adversarial examples due to the issue of vanishing gradient. In this work, we provide a comprehensive investigation of the CE loss function and find that the logit margin between the targeted and untargeted classes will quickly obtain saturation in CE, which largely limits the transferability. Therefore, in this paper, we devote to the goal of continually increasing the logit margin along the optimization to deal with the saturation issue and propose two simple and effective logit calibration methods, which are achieved by downscaling the logits with a temperature factor and an adaptive margin, respectively. Both of them can effectively encourage optimization to produce a larger logit margin and lead to higher transferability. Besides, we show that minimizing the cosine distance between the adversarial examples and the classifier weights of the target class can further improve the transferability, which is benefited from downscaling logits via L2-normalization. Experiments conducted on the ImageNet dataset validate the effectiveness of the proposed methods, which outperform the state-of-the-art methods in black-box targeted attacks. The source code is available at Link. Juanjuan Weng, Zhiming Luo, Shaozi Li, Nicu Sebe, Zhun Zhong |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Structure-Guided Cross-Attention Network for Cross-Domain OCT Fluid SegmentationabstractAccurate retinal fluid segmentation on Optical Coherence Tomography (OCT) images plays an important role in diagnosing and treating various eye diseases. The art deep models have shown promising performance on OCT image segmentation given pixel-wise annotated training data. However, the learned model will achieve poor performance on OCT images that are obtained from different devices (domains) due to the domain shift issue. This problem largely limits the real-world application of OCT image segmentation since the types of devices usually are different in each hospital. In this paper, we study the task of cross-domain OCT fluid segmentation, where we are given a labeled dataset of the source device (domain) and an unlabeled dataset of the target device (domain). The goal is to learn a model that can perform well on the target domain. To solve this problem, in this paper, we propose a novel Structure-guided Cross-Attention Network (SCAN), which leverages the retinal layer structure to facilitate domain alignment. Our SCAN is inspired by the fact that the retinal layer structure is robust to domains and can reflect regions that are important to fluid segmentation. In light of this, we build our SCAN in a multi-task manner by jointly learning the retinal structure prediction and fluid segmentation. To exploit the mutual benefit between layer structure and fluid segmentation, we further introduce a cross-attention module to measure the correlation between the layer-specific feature and the fluid-specific feature encouraging the model to concentrate on highly relative regions during domain alignment. Moreover, an adaptation difficulty map is evaluated based on the retinal structure predictions from different domains, which enforces the model focus on hard regions during structure-aware adversarial learning. Extensive experiments on the three domains of the RETOUCH dataset demonstrate the effectiveness of the proposed method and show that our approach produces state-of-the-art performance on cross-domain OCT fluid segmentation. Xingxin He, Zhun Zhong, Leyuan Fang, Nicu Sebe |
IEEE Trans. Image Process. | 2 |
| 2023 | Efficient Layer Compression Without PruningabstractNetwork pruning is one of the chief means for improving the computational efficiency of Deep Neural Networks (DNNs). Pruning-based methods generally discard network kernels, channels, or layers, which however inevitably will disrupt original well-learned network correlation and thus lead to performance degeneration. In this work, we propose an Efficient Layer Compression (ELC) approach to efficiently compress serial layers by decoupling and merging rather than pruning. Specifically, we first propose a novel decoupling module to decouple the layers, enabling us readily merge serial layers that include both nonlinear and convolutional layers. Then, the decoupled network is losslessly merged based on the equivalent conversion of the parameters. In this way, our ELC can effectively reduce the depth of the network without destroying the correlation of the convolutional layers. To our best knowledge, we are the first to exploit the mergeability of serial convolutional layers for lossless network layer compression. Experimental results conducted on two datasets demonstrate that our method retains superior performance with a FLOPs reduction of 74.1% for VGG-16 and 54.6% for ResNet-56, respectively. In addition, our ELC improves the inference speed by 2× on Jetson AGX Xavier edge device. Jie Wu 0035, Dingshun Zhu, Leyuan Fang, Yue Deng 0001, Zhun Zhong |
IEEE Trans. Image Process. | 5 |
| 2023 | Win-Win by Competition: Auxiliary-Free Cloth-Changing Person Re-IdentificationabstractRecent person Re-IDentification (ReID) systems have been challenged by changes in personnel clothing, leading to the study of Cloth-Changing person ReID (CC-ReID). Commonly used techniques involve incorporating auxiliary information (e.g., body masks, gait, skeleton, and keypoints) to accurately identify the target pedestrian. However, the effectiveness of these methods heavily relies on the quality of auxiliary information and comes at the cost of additional computational resources, ultimately increasing system complexity. This paper focuses on achieving CC-ReID by effectively leveraging the information concealed within the image. To this end, we introduce an Auxiliary-free Competitive IDentification (ACID) model. It achieves a win-win situation by enriching the identity (ID)-preserving information conveyed by the appearance and structure features while maintaining holistic efficiency. In detail, we build a hierarchical competitive strategy that progressively accumulates meticulous ID cues with discriminating feature extraction at the global, channel, and pixel levels during model inference. After mining the hierarchical discriminative clues for appearance and structure features, these enhanced ID-relevant features are crosswise integrated to reconstruct images for reducing intra-class variations. Finally, by combing with self- and cross-ID penalties, the ACID is trained under a generative adversarial learning framework to effectively minimize the distribution discrepancy between the generated data and real-world data. Experimental results on four public cloth-changing datasets (i.e., PRCC-ReID, VC-Cloth, LTCC-ReID, and Celeb-ReID) demonstrate the proposed ACID can achieve superior performance over state-of-the-art methods. The code is available soon at: https://github.com/BoomShakaY/Win-CCReID. Zhengwei Yang 0001, Xian Zhong, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Shin'ichi Satoh 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | 100-Driver: A Large-Scale, Diverse Dataset for Distracted Driver ClassificationabstractDistracted driver classification (DDC) plays an important role in ensuring driving safety. Although many datasets are introduced to support the study of DDC, most of them are small in data size and are short of diversity in environmental variations. This largely limits the development of DDC since many practical problems such as the cross-modality setting cannot be fully studied. In this paper, we introduce 100-Driver, a large-scale, diverse posture-based distracted diver dataset, with more than 470K images taken by 4 cameras observing 100 drivers over 79 hours from 5 vehicles. 100-Driver involves different types of variations that closely meet real-world applications, including changes in the vehicle, person, camera view, lighting, and modality. We provide a detailed analysis of 100-Driver and present 4 settings for investigating practical problems of DDC, including the traditional setting without domain shift and 3 challenging settings (i.e., cross-modality, cross-view, and cross-vehicle) with domain shifts. We conduct comprehensive experiments on these 4 settings with state-the-of-art techniques and show several insights to the future study of DDC. Our 100-Driver will be publicly available offering new opportunities to advance the development of DDC. The 100-driver dataset, source code, and evaluation protocols are available athttps://100-driver.github.io. Jing Wang 0092, Wenjing Li 0005, Jun Zhang 0034, ZhongCheng Wu, Zhun Zhong, Nicu Sebe |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Novel Class Discovery in Semantic SegmentationabstractWe introduce a new setting of Novel Class Discovery in Semantic Segmentation (NCDSS), which aims at segmenting unlabeled images containing new classes given prior knowledge from a labeled set of disjoint classes. In contrast to existing approaches that look at novel class dis-covery in image classification, we focus on the more chal-lenging semantic segmentation. In NCDSS, we need to dis-tinguish the objects and background, and to handle the existence of multiple classes within an image, which in-creases the difficulty in using the unlabeled data. To tackle this new setting, we leverage the labeled base data and a saliency model to coarsely cluster novel classes for model training in our basic framework. Additionally, we propose the Entropy-based Uncertainty Modeling and Self-training (EUMS) framework to overcome noisy pseudo-labels, fur-ther improving the model performance on the novel classes. Our EUMS utilizes an entropy ranking technique and a dy-namic reassignment to distill clean labels, thereby making full use of the noisy data via self-supervised learning. We build the NCDSS benchmark on the PASCAL-5idataset and COCO-20i dataset. Extensive experiments demonstrate the feasibility of the basic framework (achieving an average mIoU of 49.81% on PASCAL-5i) and the effectiveness of EUMS framework (outperforming the basic framework by 9.28% mIoU on PASCAL-5i). Zhun Zhong, Nicu Sebe, Gim Hee Lee |
CVPR | 2 |
| 2022 | Class-Incremental Novel Class Discovery
Subhankar Roy, Zhun Zhong, Nicu Sebe, Elisa Ricci 0001 |
ECCV (33) | 3 |
| 2022 | 3D-Aware Semantic-Guided Generative Model for Human Synthesis
Jichao Zhang, Enver Sangineto, Hao Tang 0005, Aliaksandr Siarohin, Zhun Zhong, Nicu Sebe, Wei Wang 0108 |
ECCV (15) | 5 |
| 2022 | Style-Hallucinated Dual Consistency Learning for Domain Generalized Semantic Segmentation
Zhun Zhong, Na Zhao 0004, Nicu Sebe, Gim Hee Lee |
ECCV (28) | 2 |
| 2022 | Attentive Decoupling Network for Cloth-Changing Re-IdentificationabstractRecently, Cloth-Changing person Re-IDentification (CC-ReID) plays a vital role in the public security system and social livelihood, and suffers the problem of considerable intra-class variation. This paper demonstrates that coarse-grained appearance and body shape features are helpful for CC-ReID. We propose an Attentive DeCoupling (ADC) Network for CC-ReID without auxiliary information. The proposed network is built on two core designs. First, a joint identification structure is proposed to retain ID-relevant information at appearance and shape levels. Second, Competitive Attention (CA) is adopted, where the model progressively updates attention to accumulate sound cues for discriminating identity (ID). The proposed decoupling process is continuously improved through constant self-defeating competition of the network. Experimental results on the public cloth-changing dataset show the proposed method's effectiveness and generalizability. Zhengwei Yang 0001, Xian Zhong, Hong Liu 0009, Zhun Zhong, Zheng Wang 0007 |
ICME | 4 |
| 2022 | Adversarial Style Augmentation for Domain Generalized Urban-Scene SegmentationabstractIn this paper, we consider the problem of domain generalization in semantic segmentation, which aims to learn a robust model using only labeled synthetic (source) data. The model is expected to perform well on unseen real (target) domains. Our study finds that the image style variation can largely influence the model's performance and the style features can be well represented by the channel-wise mean and standard deviation of images. Inspired by this, we propose a novel adversarial style augmentation (AdvStyle) approach, which can dynamically generate hard stylized images during training and thus can effectively prevent the model from overfitting on the source domain. Specifically, AdvStyle regards the style feature as a learnable parameter and updates it by adversarial training. The learned adversarial style feature is used to construct an adversarial image for robust model training. AdvStyle is easy to implement and can be readily applied to different models. Experiments on two synthetic-to-real semantic segmentation benchmarks demonstrate that AdvStyle can significantly improve the model performance on unseen real domains and show that we can achieve the state of the art. Moreover, AdvStyle can be employed to domain generalized image classification and produces a clear improvement on the considered datasets. Zhun Zhong, Gim Hee Lee, Nicu Sebe |
NeurIPS | 1 |
| 2022 | Source-Free Open Compound Domain Adaptation in Semantic SegmentationabstractIn this work, we introduce a new concept, named source-free open compound domain adaptation (SF-OCDA), and study it in semantic segmentation. SF-OCDA is more challenging than the traditional domain adaptation but it is more practical. It jointly considers (1) the issues of data privacy and data storage and (2) the scenario of multiple target domains and unseen open domains. In SF-OCDA, only the source pre-trained model and the target data are available to learn the target model. The model is evaluated on the samples from the target and unseen open domains. To solve this problem, we present an effective framework by separating the training process into two stages: (1) pre-training a generalized source model and (2) adapting a target model with self-supervised learning. In our framework, we propose the Cross-Patch Style Swap (CPSS) to diversify samples with various patch styles in the feature-level, which can benefit the training of both stages. First, CPSS can significantly improve the generalization ability of the source model, providing more accurate pseudo-labels for the latter stage. Second, CPSS can reduce the influence of noisy pseudo-labels and also avoid the model overfitting to the target domain during self-supervised learning, consistently boosting the performance on the target and open domains. Experiments demonstrate that our method produces state-of-the-art results on the C-Driving dataset. Furthermore, our model also achieves the leading performance on CityScapes for domain generalization. Zhun Zhong, Zhiming Luo, Gim Hee Lee, Nicu Sebe |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Joint Representation Learning and Keypoint Detection for Cross-View Geo-LocalizationabstractIn this paper, we study the cross-view geo-localization problem to match images from different viewpoints. The key motivation underpinning this task is to learn a discriminative viewpoint-invariant visual representation. Inspired by the human visual system for mining local patterns, we propose a new framework called RK-Net to jointly learn the discriminative Representation and detect salient Keypoints with a single Network. Specifically, we introduce a Unit Subtraction Attention Module (USAM) that can automatically discover representative keypoints from feature maps and draw attention to the salient regions. USAM contains very few learning parameters but yields significant performance improvement and can be easily plugged into different networks. We demonstrate through extensive experiments that (1) by incorporating USAM, RK-Net facilitates end-to-end joint learning without the prerequisite of extra annotations. Representation learning and keypoint detection are two highly-related tasks. Representation learning aids keypoint detection. Keypoint detection, in turn, enriches the model capability against large appearance changes caused by viewpoint variants. (2) USAM is easy to implement and can be integrated with existing methods, further improving the state-of-the-art performance. We achieve competitive geo-localization accuracy on three challenging datasets, i. e., University-1652, CVUSA and CVACT. Our code is available at https://github.com/AggMan96/RK-Net. Jinliang Lin, Zhedong Zheng, Zhun Zhong, Zhiming Luo, Shaozi Li, Yi Yang 0001, Nicu Sebe |
IEEE Trans. Image Process. | 3 |
| 2021 | Learning to Attack Real-World Models for Person Re-identification via Virtual-Guided Meta-LearningabstractRecent advances in person re-identification (re-ID) have led to impressive retrieval accuracy. However, existing re-ID models are challenged by the adversarial examples crafted by adding quasi-imperceptible perturbations. Moreover, re-ID systems face the domain shift issue that training and testing domains are not consistent. In this study, we argue that learning powerful attackers with high universality that works well on unseen domains is an important step in promoting the robustness of re-ID systems. Therefore, we introduce a novel universal attack algorithm called ``MetaAttack'' for person re-ID. MetaAttack can mislead re-ID models on unseen domains by a universal adversarial perturbation. Specifically, to capture common patterns across different domains, we propose a meta-learning scheme to seek the universal perturbation via the gradient interaction between meta-train and meta-test formed by two datasets. We also take advantage of a virtual dataset (PersonX), instead of real ones, to conduct meta-test. This scheme not only enables us to learn with more comprehensive variation factors but also mitigates the negative effects caused by biased factors of real datasets. Experiments on three large-scale re-ID datasets demonstrate the effectiveness of our method in attacking re-ID models on unseen domains. Our final visualization results reveal some new properties of existing re-ID systems, which can guide us in designing a more robust re-ID model. Code and supplemental material are available at \url{https://github.com/FlyingRoastDuck/MetaAttack_AAAI21}. Fengxiang Yang, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Zhiming Luo, Shaozi Li, Nicu Sebe, Shin'ichi Satoh 0001 |
AAAI | 2 |
| 2021 | Curriculum Graph Co-Teaching for Multi-Target Domain AdaptationabstractIn this paper we address multi-target domain adaptation (MTDA), where given one labeled source dataset and multiple unlabeled target datasets that differ in data distributions, the task is to learn a robust predictor for all the target domains. We identify two key aspects that can help to alleviate multiple domain-shifts in the MTDA: feature aggregation and curriculum learning. To this end, we propose Curriculum Graph Co-Teaching (CGCT) that uses a dual classifier head, with one of them being a graph convolutional network (GCN) which aggregates features from similar samples across the domains. To prevent the classifiers from over-fitting on its own noisy pseudo-labels we develop a co-teaching strategy with the dual classifier head that is assisted by curriculum learning to obtain more reliable pseudo-labels. Furthermore, when the domain labels are available, we propose Domain-aware Curriculum Learning (DCL), a sequential adaptation strategy that first adapts on the easier target domains, followed by the harder ones. We experimentally demonstrate the effectiveness of our proposed frameworks on several benchmarks and advance the state-of-the-art in the MTDA by large margins (e.g. +5.6% on the DomainNet). Subhankar Roy, Evgeny Krivosheev, Zhun Zhong, Nicu Sebe, Elisa Ricci 0001 |
CVPR | 3 |
| 2021 | Joint Noise-Tolerant Learning and Meta Camera Shift Adaptation for Unsupervised Person Re-IdentificationabstractThis paper considers the problem of unsupervised person re-identification (re-ID), which aims to learn discriminative models with unlabeled data. One popular method is to obtain pseudo-label by clustering and use them to optimize the model. Although this kind of approach has shown promising accuracy, it is hampered by 1) noisy labels produced by clustering and 2) feature variations caused by camera shift. The former will lead to incorrect optimization and thus hinders the model accuracy. The latter will result in assigning the intra-class samples of different cameras to different pseudo-label, making the model sensitive to camera variations. In this paper, we propose a unified framework to solve both problems. Concretely, we propose a Dynamic and Symmetric Cross-Entropy loss (DSCE) to deal with noisy samples and a camera-aware meta-learning algorithm (MetaCam) to adapt camera shift. DSCE can alleviate the negative effects of noisy samples and accommodate the change of clusters after each clustering step. MetaCam simulates cross-camera constraint by splitting the training data into meta-train and meta-test based on camera IDs. With the interacted gradient from meta-train and meta-test, the model is enforced to learn camera-invariant features. Extensive experiments on three re-ID benchmarks show the effectiveness and the complementary of the proposed DSCE and MetaCam. Our method outperforms the state-of-the-art methods on both fully unsupervised re-ID and unsupervised domain adaptive re-ID. Fengxiang Yang, Zhun Zhong, Zhiming Luo, Yuanzheng Cai, Yaojin Lin, Shaozi Li, Nicu Sebe |
CVPR | 2 |
| 2021 | Learning to Generalize Unseen Domains via Memory-based Multi-Source Meta-Learning for Person Re-IdentificationabstractRecent advances in person re-identification (ReID) obtain impressive accuracy in the supervised and unsupervised learning settings. However, most of the existing methods need to train a new model for a new domain by accessing data. Due to public privacy, the new domain data are not always accessible, leading to a limited applicability of these methods. In this paper, we study the problem of multi-source domain generalization in ReID, which aims to learn a model that can perform well on unseen domains with only several labeled source domains. To address this problem, we propose the Memory-based Multi-Source Meta-Learning (M3L) framework to train a generalizable model for unseen domains. Specifically, a meta-learning strategy is introduced to simulate the train-test process of domain generalization for learning more generalizable models. To overcome the unstable meta-optimization caused by the parametric classifier, we propose a memory-based identification loss that is non-parametric and harmonizes with meta-learning. We also present a meta batch normalization layer (MetaBN) to diversify meta-test features, further establishing the advantage of meta-learning. Experiments demonstrate that our M3L can effectively enhance the generalization ability of the model for unseen domains and can outperform the state-of-the-art methods on four large-scale ReID datasets. Zhun Zhong, Fengxiang Yang, Zhiming Luo, Yaojin Lin, Shaozi Li, Nicu Sebe |
CVPR | 2 |
| 2021 | Neighborhood Contrastive Learning for Novel Class DiscoveryabstractIn this paper, we address Novel Class Discovery (NCD), the task of unveiling new classes in a set of unlabeled samples given a labeled dataset with known classes. We exploit the peculiarities of NCD to build a new framework, named Neighborhood Contrastive Learning (NCL), to learn discriminative representations that are important to clustering performance. Our contribution is twofold. First, we find that a feature extractor trained on the labeled set generates representations in which a generic query sample and its neighbors are likely to share the same class. We exploit this observation to retrieve and aggregate pseudo-positive pairs with contrastive learning, thus encouraging the model to learn more discriminative representations. Second, we notice that most of the instances are easily discriminated by the network, contributing less to the contrastive loss. To overcome this issue, we propose to generate hard negatives by mixing labeled and unlabeled samples in the feature space. We experimentally demonstrate that these two ingredients significantly contribute to clustering performance and lead our model to outperform state-of-the-art methods by a large margin (e.g., clustering accuracy +13% on CIFAR-100 and +8% on ImageNet). Zhun Zhong, Enrico Fini, Subhankar Roy, Zhiming Luo, Elisa Ricci 0001, Nicu Sebe |
CVPR | 1 |
| 2021 | OpenMix: Reviving Known Knowledge for Discovering Novel Visual Categories in an Open WorldabstractIn this paper, we tackle the problem of discovering new classes in unlabeled visual data given labeled data from disjoint classes. Existing methods typically first pre-train a model with labeled data, and then identify new classes in unlabeled data via unsupervised clustering. However, the labeled data that provide essential knowledge are often underexplored in the second step. The challenge is that the labeled and unlabeled examples are from non-overlapping classes, which makes it difficult to build a learning relationship between them. In this work, we introduce Open-Mix to mix the unlabeled examples from an open set and the labeled examples from known classes, where their non-overlapping labels and pseudo-labels are simultaneously mixed into a joint label distribution. OpenMix dynamically compounds examples in two ways. First, we produce mixed training images by incorporating labeled examples with unlabeled examples. With the benefit of unique prior knowledge in novel class discovery, the generated pseudo-labels will be more credible than the original unlabeled predictions. As a result, OpenMix helps preventing the model from overfitting on unlabeled samples that may be assigned with wrong pseudo-labels. Second, the first way encourages the unlabeled examples with high class-probabilities to have considerable accuracy. We introduce these examples as reliable anchors and further integrate them with un-labeled samples. This enables us to generate more combinations in unlabeled examples and exploit finer object relations among the new classes. Experiments on three classification datasets demonstrate the effectiveness of the proposed OpenMix, which is superior to state-of-the-art methods in novel class discovery. Zhun Zhong, Linchao Zhu, Zhiming Luo, Shaozi Li, Yi Yang 0001, Nicu Sebe |
CVPR | 1 |
| 2021 | A Unified Objective for Novel Class DiscoveryabstractIn this paper, we study the problem of Novel Class Discovery (NCD). NCD aims at inferring novel object categories in an unlabeled set by leveraging from prior knowledge of a labeled set containing different, but related classes. Existing approaches tackle this problem by considering multiple objective functions, usually involving specialized loss terms for the labeled and the unlabeled samples respectively, and often requiring auxiliary regularization terms. In this paper we depart from this traditional scheme and introduce a UNified Objective function (UNO) for discovering novel classes, with the explicit purpose of favoring synergy between supervised and unsupervised learning. Using a multi-view self-labeling strategy, we generate pseudo-labels that can be treated homogeneously with ground truth labels. This leads to a single classification objective operating on both known and unknown classes. De-spite its simplicity, UNO outperforms the state of the art by a significant margin on several benchmarks (≈ +10% on CIFAR-100 and +8% on ImageNet). The project page is available at : https://ncd-uno.github.io. Enrico Fini, Enver Sangineto, Stéphane Lathuilière, Zhun Zhong, Moin Nabi, Elisa Ricci 0001 |
ICCV | 4 |
| 2021 | High-Order-Interaction for weakly supervised Fine-Grained Visual Categorization
Nanyu Li, Zhiming Luo, Zhun Zhong, Shaozi Li |
Neurocomputing | 4 |
| 2021 | Divide-and-Merge the embedding space for cross-modality person search
Chengji Wang, Zhiming Luo, Zhun Zhong, Shaozi Li |
Neurocomputing | 3 |
| 2021 | SAFD: single shot anchor free face detector
Chengji Wang, Zhiming Luo, Zhun Zhong, Shaozi Li |
Multim. Tools Appl. | 3 |
| 2021 | Learning to Adapt Invariance in Memory for Person Re-IdentificationabstractThis work considers the problem of unsupervised domain adaptation in person re-identification (re-ID), which aims to transfer knowledge from the source domain to the target domain. Existing methods are primary to reduce the inter-domain shift between the domains, which however usually overlook the relations among target samples. This paper investigates into the intra-domain variations of the target domain and proposes a novel adaptation framework w.r.t three types of underlying invariance, i.e., Exemplar-Invariance, Camera-Invariance, and Neighborhood-Invariance. Specifically, an exemplar memory is introduced to store features of samples, which can effectively and efficiently enforce the invariance constraints over the global dataset. We further present the Graph-based Positive Prediction (GPP) method to explore reliable neighbors for the target domain, which is built upon the memory and is trained on the source samples. Experiments demonstrate that 1) the three invariance properties are complementary and indispensable for effective domain adaptation, 2) the memory plays a key role in implementing invariance learning and improves the performance with limited extra computation cost, 3) GPP can facilitate the invariance learning and thus significantly improves the results, and 4) our approach produces new state-of-the-art adaptation accuracy on three re-ID large-scale benchmarks. Zhun Zhong, Liang Zheng 0001, Zhiming Luo, Shaozi Li, Yi Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Asymmetric Co-Teaching for Unsupervised Cross-Domain Person Re-IdentificationabstractPerson re-identification (re-ID), is a challenging task due to the high variance within identity samples and imaging conditions. Although recent advances in deep learning have achieved remarkable accuracy in settled scenes, i.e., source domain, few works can generalize well on the unseen target domain. One popular solution is assigning unlabeled target images with pseudo labels by clustering, and then retraining the model. However, clustering methods tend to introduce noisy labels and discard low confidence samples as outliers, which may hinder the retraining process and thus limit the generalization ability. In this study, we argue that by explicitly adding a sample filtering procedure after the clustering, the mined examples can be much more efficiently used. To this end, we design an asymmetric co-teaching framework, which resists noisy labels by cooperating two models to select data with possibly clean labels for each other. Meanwhile, one of the models receives samples as pure as possible, while the other takes in samples as diverse as possible. This procedure encourages that the selected training samples can be both clean and miscellaneous, and that the two models can promote each other iteratively. Extensive experiments show that the proposed framework can consistently benefit most clustering based methods, and boost the state-of-the-art adaptation accuracy. Our code is available at https://github.com/FlyingRoastDuck/ACT_AAAI20. Fengxiang Yang, Ke Li 0015, Zhun Zhong, Zhiming Luo, Xing Sun 0001, Hao Cheng 0012, Feiyue Huang, Rongrong Ji, Shaozi Li |
AAAI | 3 |
| 2020 | Random Erasing Data AugmentationabstractIn this paper, we introduce Random Erasing, a new data augmentation method for training the convolutional neural network (CNN). In training, Random Erasing randomly selects a rectangle region in an image and erases its pixels with random values. In this process, training images with various levels of occlusion are generated, which reduces the risk of over-fitting and makes the model robust to occlusion. Random Erasing is parameter learning free, easy to implement, and can be integrated with most of the CNN-based recognition models. Albeit simple, Random Erasing is complementary to commonly used data augmentation techniques such as random cropping and flipping, and yields consistent improvement over strong baselines in image classification, object detection and person re-identification. Code is available at: https://github.com/zhunzhong07/Random-Erasing. Zhun Zhong, Liang Zheng 0001, Guoliang Kang, Shaozi Li, Yi Yang 0001 |
AAAI | 1 |
| 2020 | Class-Aware Modality Mix and Center-Guided Metric Learning for Visible-Thermal Person Re-IdentificationabstractVisible thermal person re-identification (VT-REID) is an important and challenging task in that 1) weak lighting environments are inevitably encountered in real-world settings and 2) the inter-modality discrepancy is serious. Most existing methods either aim at reducing the cross-modality gap in pixel- and feature-level or optimizing cross-modality network by metric learning techniques. However, few works have jointly considered these two aspects and studied their mutual benefits. In this paper, we design a novel framework to jointly bridge the modality gap in pixel- and feature-level without additional parameters, as well as reduce the inter- and intra-modalities variations by a center-guided metric learning constraint. Specifically, we introduce the Class-aware Modality Mix (CMM) to generate internal information of the two modalities for reducing the modality gap in pixel-level. In addition, we exploit the KL-divergence to further align modality distributions on feature-level. On the other hand, we propose an efficient Center-guided Metric Learning (CML) method for decreasing the discrepancy within the inter- and intra-modalities, by enforcing constraints on class centers and instances. Extensive experiments on two datasets show the mutual advantage of the proposed components and demonstrate the superiority of our method over the state of the art. Yongguo Ling, Zhun Zhong, Zhiming Luo, Paolo Rota, Shaozi Li, Nicu Sebe |
ACM Multimedia | 2 |
| 2020 | Thorax disease classification with attention guided convolutional neural network
Qingji Guan, Zhun Zhong, Zhedong Zheng, Liang Zheng 0001, Yi Yang 0001 |
Pattern Recognit. Lett. | 3 |
| 2020 | Leveraging Virtual and Real Person for Unsupervised Person Re-IdentificationabstractPerson re-identification (re-ID) is a challenging instance retrieval problem, especially when identity annotations are not available for training. Although modern deep re-ID approaches have achieved great improvement, it is still difficult to optimize the deep re-ID model and learn discriminative person representation without annotations in training data. To address this challenge, this study considers the problem of unsupervised person re-ID and introduces a novel approach to solve this problem by leveraging virtual and real data. Our approach includes two components: virtual person generation and training of the deep re-ID model. For virtual person generation, we learn a person generation model and a camera style transfer model using unlabeled real data to generate virtual persons with different poses and camera styles. The virtual data is formed as labeled training data, enabling subsequent training deep re-ID model in supervision. For training of the deep re-ID model, we divide it into three steps: 1) pre-training a coarse re-ID model by using virtual data; 2) collaborative filtering based positive pair mining from the real data; and 3) fine-tuning of the coarse re-ID model by leveraging the mined positive pairs and virtual data. The final re-ID model is achieved by iterating between step 2 and step 3 until convergence. Extensive experiments demonstrate the effectiveness of our method. Experimental results on two large-scale datasets, Market-1501 and DukeMTMC-reID, show the advantages of our method over state-of-the-art approaches in unsupervised person re-ID. Our code is now available online1. Fengxiang Yang, Zhun Zhong, Zhiming Luo, Sheng Lian, Shaozi Li |
IEEE Trans. Multim. | 2 |
| 2019 | Invariance Matters: Exemplar Memory for Domain Adaptive Person Re-IdentificationabstractThis paper considers the domain adaptive person re-identification (re-ID) problem: learning a re-ID model from a labeled source domain and an unlabeled target domain. Conventional methods are mainly to reduce feature distribution gap between the source and target domains. However, these studies largely neglect the intra-domain variations in the target domain, which contain critical factors influencing the testing performance on the target domain. In this work, we comprehensively investigate into the intra-domain variations of the target domain and propose to generalize the re-ID model w.r.t three types of the underlying invariance, i.e., exemplar-invariance, camera-invariance and neighborhood-invariance. To achieve this goal, an exemplar memory is introduced to store features of the target domain and accommodate the three invariance properties. The memory allows us to enforce the invariance constraints over global training batch without significantly increasing computation cost. Experiment demonstrates that the three invariance properties and the proposed memory are indispensable towards an effective domain adaptation system. Results on three re-ID domains show that our domain adaptation accuracy outperforms the state of the art by a large margin. Code is available at: https://github.com/zhunzhong07/ECN. Zhun Zhong, Liang Zheng 0001, Zhiming Luo, Shaozi Li, Yi Yang 0001 |
CVPR | 1 |
| 2019 | CamStyle: A Novel Data Augmentation Method for Person Re-IdentificationabstractPerson re-identification (re-ID) is a cross-camera retrieval task that suffers from image style variations caused by different cameras. The art implicitly addresses this problem by learning a camera-invariant descriptor subspace. In this paper, we explicitly consider this challenge by introducing camera style (CamStyle). CamStyle can serve as a data augmentation approach that reduces the risk of deep network overfitting and that smooths the CamStyle disparities. Specifically, with a style transfer model, labeled training images can be style transferred to each camera, and along with the original training samples, form the augmented training set. This method, while increasing data diversity against overfitting, also incurs a considerable level of noise. In the effort to alleviate the impact of noise, the label smooth regularization (LSR) is adopted. The vanilla version of our method (without LSR) performs reasonably well on few camera systems in which overfitting often occurs. With LSR, we demonstrate consistent improvement in all systems regardless of the extent of overfitting. We also report competitive accuracy compared with the state of the art on Market-1501 and DukeMTMC-re-ID. Importantly, CamStyle can be employed to the challenging problems of one view learning and unsupervised domain adaptation (UDA) in person re-identification (re-ID), both of which have critical research and application significance. The former only has labeled data in one camera view and the latter only has labeled data in the source domain. Experimental results show that CamStyle significantly improves the performance of the baseline in the two problems. Specially, for UDA, CamStyle achieves state-of-the-art accuracy based on a baseline deep re-ID model on Market-1501 and DukeMTMC-reID. Our code is available at: https://github.com/zhunzhong07/CamStyle . Zhun Zhong, Liang Zheng 0001, Zhedong Zheng, Shaozi Li, Yi Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Camera Style Adaptation for Person Re-IdentificationabstractBeing a cross-camera retrieval task, person re-identification suffers from image style variations caused by different cameras. The art implicitly addresses this problem by learning a camera-invariant descriptor subspace. In this paper, we explicitly consider this challenge by introducing camera style (CamStyle) adaptation. CamStyle can serve as a data augmentation approach that smooths the camera style disparities. Specifically, with CycleGAN, labeled training images can be style-transferred to each camera, and, along with the original training samples, form the augmented training set. This method, while increasing data diversity against over-fitting, also incurs a considerable level of noise. In the effort to alleviate the impact of noise, the label smooth regularization (LSR) is adopted. The vanilla version of our method (without LSR) performs reasonably well on few-camera systems in which over-fitting often occurs. With LSR, we demonstrate consistent improvement in all systems regardless of the extent of over-fitting. We also report competitive accuracy compared with the state of the art. Code is available at: https://github.com/zhunzhong07/CamStyle. Zhun Zhong, Liang Zheng 0001, Zhedong Zheng, Shaozi Li, Yi Yang 0001 |
CVPR | 1 |
| 2018 | Generalizing a Person Retrieval Model Hetero- and Homogeneously
Zhun Zhong, Liang Zheng 0001, Shaozi Li, Yi Yang 0001 |
ECCV (13) | 1 |
| 2018 | Attention guided U-Net for accurate iris segmentation
Sheng Lian, Zhiming Luo, Zhun Zhong, Songzhi Su, Shaozi Li |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Re-ranking Person Re-identification with k-Reciprocal EncodingabstractWhen considering person re-identification (re-ID) as a retrieval process, re-ranking is a critical step to improve its accuracy. Yet in the re-ID community, limited effort has been devoted to re-ranking, especially those fully automatic, unsupervised solutions. In this paper, we propose a k-reciprocal encoding method to re-rank the re-ID results. Our hypothesis is that if a gallery image is similar to the probe in the k-reciprocal nearest neighbors, it is more likely to be a true match. Specifically, given an image, a k-reciprocal feature is calculated by encoding its k-reciprocal nearest neighbors into a single vector, which is used for re-ranking under the Jaccard distance. The final distance is computed as the combination of the original distance and the Jaccard distance. Our re-ranking method does not require any human interaction or any labeled data, so it is applicable to large-scale datasets. Experiments on the large-scale Market-1501, CUHK03, MARS, and PRW datasets confirm the effectiveness of our method. Zhun Zhong, Liang Zheng 0001, Donglin Cao, Shaozi Li |
CVPR | 1 |
| 2017 | Feature++: Cross dimension feature fusion for road detectionabstractRoad detection is a key component of Advanced Driving Assistance Systems, which provides valid space and candidate regions of objects for vehicles. Mainstream road detection methods have focused on extracting discriminative features. In this paper, we propose a robust feature fusion framework, called “Feature++”, which is combined with superpixel feature and 3D feature extracted from stereo images. Then a neural network classifier is been trained to decide whether a superpixel is road region or not. Finally, the classified results are further refined by conditional random field. Experiments conducted on the KITTI ROAD benchmark show that the proposed “Feature++” method outperforms most manually designed features, and are comparable with state-of-the-art methods that based on deep learning architecture. Wenli He, Guo-Rong Cai, Zhun Zhong, Songzhi Su |
ICASSP | 3 |
| 2017 | Class-specific object proposals re-ranking for object detection in automatic driving
Zhun Zhong, Mingyi Lei, Donglin Cao, Jianping Fan 0001, Shaozi Li |
Neurocomputing | 1 |
| 2017 | Detecting ground control points via convolutional neural network for stereo matching
Zhun Zhong, Songzhi Su, Donglin Cao, Shaozi Li, Zhihan Lyu |
Multim. Tools Appl. | 1 |
| 2004 | Spectrum agile radio: radio resource measurements for opportunistic spectrum usageabstractRadio spectrum allocation is undergoing radical rethinking. Regulators, government agencies, industry, and the research community have recently established many initiatives for new spectrum policies and seek approaches to more efficiently manage the radio spectrum. In this paper, we examine new approaches, namely, spectrum agile radios, for opportunistic spectrum usage. Spectrum agile radios use parts of the radio spectrum that were originally licensed to other radio services. A spectrum agile radio device seeks opportunities, i.e. unused radio resources. Devices communicate using the identified opportunities, without interfering with the operation of licensed radio devices. The identification of spectrum opportunities is coordinated by policies, which are defined by, and under the control of, the radio regulator. Our approach is motivated by the publications of the next generation communications, XG, research project of the USA-based Defense Advanced Research Projects Agency, DARPA. We focus on IEEE 802.11k for radio resource measurements as an approach to facilitate the development of spectrum agile radios. Stefan Mangold, Zhun Zhong, Kiran S. Challapali, Chun-Ting Chou |
GLOBECOM | 2 |
| 2004 | IEEE 802.11e/802.11k wireless LAN: spectrum awareness for distributed resource sharingabstractAbstract Coordinating priorities in wireless medium access is difficult when radio networks operate with contention‐based medium access. Contention‐based medium access protocols such as listen‐before‐talk are widely employed today, and for example used in the popular IEEE 802.11 protocol. Contention‐based protocols are used for wireless communication in unreliable radio environments such as the unlicensed frequency bands with their typically irregular and unpredictable interferences. However, to support time‐bounded traffic with a certain quality of service (QoS) support is extremely difficult, because it requires the knowledge of how aggressive other radio stations, which also contend for radio resources, access the medium. In this contribution, we discuss a new measurement in the IEEE 802.11k draft standard, together with the IEEE 802.11e draft standard for coordinating priorities. By combining the two extensions of IEEE 802.11, we develop an algorithm that allows radio stations to estimate the achievable throughput per radio station (the saturation throughput) in the presence of other radio stations. Our algorithm further allows predicting the saturation throughput per radio station in the presence of other non‐802.11 radio networks, because it only relies on the information about how the medium is used by other stations, i.e. for what duration other stations have to sense the medium as idle before initiating transmissions. The algorithm does not require knowledge about the contention‐parameters (like, e.g. minimum contention window sizes) used by other radio stations, and only relies on medium sensing information. For this reason, we refer to spectrum awareness in this work. We modify an existing model that was originally developed for calculating the saturation throughput in IEEE 802.11, to calculate the saturation throughput for IEEE 802.11e with one single priority. We then describe a new measurement, which is part of the IEEE 802.11k draft standard. The measurement provides information about medium access probabilities of other radio stations per contention window slots. These probabilities provide the information about how aggressive the medium is utilized by other stations. The probabilities are used in our model for approximating the saturation throughput per station and priority in the presence of other radio stations. As a result, with the help of the new model, a radio station is able to estimate its own expected saturation throughput. The comparison of the model with stochastic simulation stations indicates that our model approximates the saturation throughput per station and priority sufficiently in many scenarios, and hence allows to predict expected saturation throughputs per radio station. Copyright © 2004 John Wiley & Sons, Ltd. Stefan Mangold, Zhun Zhong, Guido R. Hiertz, Bernhard Walke |
Wirel. Commun. Mob. Comput. | 2 |
| 2003 | IEEE 802.11 downlink traffic shaping scheme for multi-user service enhancementabstractIn IEEE 802.11 wireless LAN (WLAN), due to the combination of a first-in-first-out (FIFO) transmission queue and the retransmission policy in the serving access point (AP), a single station that is moving away from the AP, e.g., thus handing off to a neighbouring AP, can have a significant negative impact on the services other users receive in the same basic service set (BSS). In this paper, we present a traffic shaping algorithm to control this undesirable effect. Experimental results show that our algorithm manages to maintain stable throughput levels in networking applications that otherwise would have experienced high throughput degradation. This algorithm constitutes an innovative implementation of a traffic shaper that can be easily adopted by the current state-of-the-art IEEE 802.11 AP products to enhance the performance of WLANs, since the algorithm can be implemented in the device driver and no specific modification of hardware is needed. Marc Portoles-Comeras, Zhun Zhong, Sunghyun Choi 0001 |
PIMRC | 2 |
| 2003 | IEEE 802.11 link-layer forwarding for smooth handoffabstractIn this paper, we present a link-layer packet forwarding scheme to reduce packet losses during a handoff in IEEE 802.11 WLAN. Through a novel scheme utilizing buffer and image queues in the device driver, the scheme is able to recover most packets that would otherwise be lost during the handoff, including those held in the network interface card. Our experimental results from a test-bed show that the proposed scheme can eliminate or significantly reduce the packet losses in UDP traffic, and help preserve throughput for TCP traffic when the retransmission time out (RTO) is large. To achieve the maximum benefit, we propose to apply the forwarding scheme adaptively depending on the traffic type. Marc Portoles-Comeras, Zhun Zhong, Sunghyun Choi 0001, Chun-Ting Chou |
PIMRC | 2 |
| 2002 | Natural interaction synthesizing in virtual teleconferencingabstractThe virtual teleconferencing system provides a collaborative virtual environment for applications such as collaborative seminar. The challenge of the natural interaction is for the participants to be aware of the virtual environment. A video avatar is a computer-synthesized three-dimension image created using live video. We study the factors contributing to the participant's natural interaction and introduce a video avatar based participant model to implement the natural, accurate and convenient interaction among participants. Our virtual teleconferencing prototype provides a seamless virtual collaboration work environment including eye contacts and gaze awareness. Our experiments on a three parties collaborative seminar show that natural interaction synthesizing can significantly improve the communication and that natural multi-model human computer interaction is necessary in collaboration. Lifeng Sun, Yuzhuo Zhong, Zhun Zhong |
ICIP (2) | 3 |
| 2002 | Complexity regulation for real-time video encodingabstractPast research in video compression has focused on improving the rate-distortion performance without active consideration of the encoding complexity. However, realtime software or processor based encoding has tight complexity constraints. In this paper we study real-time software-based video encoding as a problem of dynamic tradeoffs among rate, distortion, and complexity. We present a compression system where the compressed video quality is pre-set and fixed, but the compression complexity is constrained and regulated to ensure realtime performance, by changing the encoding configurations with varying compression efficiency. We propose a dynamic encoding complexity control scheme utilizing an input frame buffer. Our experimental results show that the proposed encoding control algorithm effectively avoids frame drops caused by complexity peaks for real-time encoding. Zhun Zhong, Yingwei Chen |
ICIP (1) | 1 |
| 2002 | Greedy heuristic placement algorithms in distributed cooperative proxy system
Changjie Guo, Zhe Xiang, Zhun Zhong, Yuzhuo Zhong |
VCIP | 3 |
| 2002 | Resource-constrained complexity-scalable video decoding via adaptive B-residual computation
Sharon S. Peng, Zhun Zhong |
VCIP | 2 |
| 2002 | MobileCache: An efficient proxy cache mechanism for cell-based wireless multimedia streaming
Zhe Xiang, Zhun Zhong, Yuzhuo Zhong |
VCIP | 2 |
| 2002 | Signal adaptive processing in MPEG-2 decoders with embedded resizing for interlaced video
Zhun Zhong, Yingwei Chen, Tse-Hua Lan |
VCIP | 1 |
| 2002 | Regulated complexity scalable MPEG-2 video decoding for media processorsabstractVideo processing on programmable platforms offers significant advantages over dedicated hardware such as shorter development cycles, flexibility, and upgradability, among others. Although the emergence of powerful media processors is making video processing on programmable platforms closer to reality, practical media-processor-based systems benefit from complexity-scalable video processing algorithms that can trade off video quality against computation while ensuring real-time performance. We develop complexity-scalable MPEG-2 video decoding algorithms with both SNR quality degradation and decoding resolution reduction. We improve the decoder's complexity-quality performance by employing signal-adaptive processing in decoding blocks and picture-type-dependent processing with unequal computation resource allocation to exploit each picture type's varying impact on the overall video quality. Realizing that software video decoding operates in a dynamic environment where both the actual computation (depending on video data) and computation allocation (depending on other functions running simultaneously) vary, we develop dynamic complexity regulation techniques to ensure that the decoding complexity stays within allocation. Simulation results on complexity-scalable MPEG-2 video decoding and dynamic complexity regulation on media processors are presented. Yingwei Chen, Zhun Zhong, Tse-Hua Lan, Sharon S. Peng, Kees van Zon |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | Picture-wise resource allocation in MPEG-2 decodingabstractPast efforts in lowering the complexity of MPEG-2 decoding through internal scaling or partial-quality decoding have led to video quality degradation. Improvement in degradation of quality can be achieved at the expense of more computation resources. We introduce the notion of picture-type-dependent (PTD) processing where more computational resources are allocated to pictures that contribute more to the overall video quality. Specifically, we advocate decoding and processing more critical pictures in MPEG-2 at higher quality while keeping the resource allocation to non-critical pictures low. Our simulation results show PTD is a very effective way to keep the average resource consumption low while maintaining satisfactory video quality. Yingwei Chen, Zhun Zhong |
ICIP (1) | 2 |
| 2001 | MPEG2 decoding complexity regulation for a media processorabstractCurrent implementations of real-time video decoders employ hard-decoding without run-time computation regulation. The result is fluctuation in the video decoding time that requires over-specified dedicated hardware or general-purpose processors to guarantee real-time performance. We propose a computation regulation system for MPEG-2 video decoding on media processors that obviates over-engineering while maintaining real-time performance and video quality. Our computation regulation scheme consists of dynamic complexity prediction and control. The complexity prediction algorithm takes implementation issues specific to media processors into account and yields accurate prediction results that correlate with actual measurements to 97%. The predicted complexity is then used as a control signal to adjust the decoding algorithm so that peaks in computation load are suppressed. Simulation results show that our system can lower the overall CPU cycle budget by a large factor (25%) with very little degradation in video quality. Tse-Hua Lan, Yingwei Chen, Zhun Zhong |
MMSP | 3 |