Zilei Wang

dblp:49/1878 · DBLP profile ↗
← Back
114ranked-venue papers
8as first author
74since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 83 · 2 first-author · 58 since 2021Graphics, computer vision, multimedia, augmented reality and games · 82 · 6 first-author · 49 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Rethinking Open-world Prompt Tuning: A Systematic Framework for Evaluation and Optimization
abstract
Prompt Tuning (PT) is a widely used strategy for adapting pre-trained Vision-Language Models (VLMs) to various downstream tasks. Conventional PT methods evaluate performance separately on known (base) and unknown (new) classes. However, in real-world scenarios, models often encounter inputs without prior knowledge of their class domain. This challenge has motivated the development of Open-world Prompt Tuning (OPT), which requires models to first determine whether a sample belongs to base or new classes and then classify it accordingly. In this work, we carefully review existing OPT methods and identify three key limitations: (L1) incomplete evaluation metrics, (L2) time-consuming and memory-intensive OOD detection methods, and (L3) insufficiently comprehensive optimization strategies. To address these issues, we first tackle L1 by proposing two novel metrics to explicitly evaluate adaptability and generalization under the OPT setting, forming a more comprehensive evaluation framework. For L2, we propose a training-free OOD detection method called Entropy-weighted Rank-normalized Fusion (ERF), which first applies rank normalization to both the maximum and the sum of base-class probabilities, followed by an entropy-weighted fusion of the normalized values. For L3, we propose a plug-and-play Gated Dual-Merging (GDM) strategy to strengthen the classifier’s capability. GDM performs selective merging at the weight level based on an adaptive criterion and combines fine-tuned and LLM-boosted logits at the output level. Extensive experiments on three PT baselines across 11 datasets demonstrate the effectiveness of our proposed ERF and GDM.
Zilei Wang
AAAI2
2026 PosPrune: Visual Token Pruning with Positional Bias Correction for Efficient Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) enhance performance on vision-language tasks by integrating visual features from pre-trained vision encoders into large language models (LLMs). However, the large number of visual tokens introduces significant computational overhead. Existing token pruning methods either perform global selection via [CLS]-based attention in the vision encode or prune within LLM decoding layers. These approaches face two key challenges: (1) [CLS]-based attention primarily focuses on visually salient regions across the entire image, often overlooking semantically important tokens essential for reasoning; and (2) strong positional bias in the shallow decoder layers causes the model to favor later-positioned tokens, while neglecting earlier ones that may carry critical reasoning cues. To address these issues, we propose PosPrune, a training-free, two-stage visual token pruning framework. At the vision encoder, we introduce an Asymmetric Region-aware Pruning (ARP) strategy that retains more tokens in semantically rich regions while discarding more tokens from semantically less informative regions, thus preserving spatial diversity and task-relevant details. In the LLM decoding stage, we find that the positional bias in shallow layers is primarily driven by model architecture rather than task semantics. Based on this insight, we propose a novel Positional Bias Correction (PBC) mechanism to mitigate this bias. To further reduce redundancy, we apply Maximal Marginal Relevance (MMR) to select tokens that best balance textual relevance and diversity. Extensive experiments on various LVLMs and benchmarks demonstrate the general effectiveness of our approach. Notably, when applied to LLaVA-1.5-7B, PosPrune achieves a reduction of 85% in FLOPs while preserving 98.5% of the original performance.
Zilei Wang
AAAI5
2026 Training-free Boosting for Few-shot Segmentation via Generalizing Semantic Mining
abstract
Few-shot Semantic Segmentation (FSS) aims to segment the novel target objects with the guidance of minimal annotated reference examples. The affinity-based method has great advantages in the FSS inference stage for both specialist model and foundation model. However, current affinity calculation merely relies on only support-query matching, without considering the query-specific semantic or the semantic correlation among inter-support samples, which limits the representation ability of affinity map. In this paper, we propose the Generalizing Semantic Mining (GSM) that focuses on exploiting generalizing semantic to improve the affinity calculation. Concretely, we first organize the affinity-based inference into three main steps to reveal the crucial role of affinity map. To address the low-data problem, Target Semantic Reusing module considers the query sample as a proxy reference and assigns it with proxy mask identifying its most generalizing semantic regions. Then, to generate the high-fidelity proxy mask, Query-specific Semantic Modeling module pinpoints the most generalizing regions through prior semantic analysis. Finally, Representative Re-weighting module explicitly modulates affinity calculation via generalization-aware weighting. Experiments on FSS benchmarks demonstrate that our GSM can serve as a plug-and-play free lunch for both specialist models and foundation models.
Kangyu Xiao, Zilei Wang, Yixin Zhang 0007, Junjie Li 0002
AAAI2
2026 From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning
abstract
Jiajun Zhang, Zeyu Cui, Jiaxi Yang, Lei Zhang, Yuheng Jing, Zeyao Ma, Tianyi Bai, Zilei Wang, Qiang Liu, Liang Wang, Binyuan Hui, Junyang Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiajun Zhang 0012, Zeyu Cui, Jiaxi Yang 0004, Lei Zhang 0201, Yuheng Jing, Zeyao Ma, Tianyi Bai, Zilei Wang, Qiang Liu 0006, Liang Wang 0001, Binyuan Hui, Junyang Lin
ACL (1)8
2026 RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task Evaluation
abstract
Jiajun Zhang, Yuying Li, Zhixun Li, Xingyu Guo, Jingzhuo Wu, Leqi Zheng, Yiran Yang, Jianke Zhang, Qingbin Li, Shannan Yan, Changguo Jia, Junfei Wu, Zilei Wang, Qiang Liu, Liang Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhixun Li, Xingyu Guo, Jingzhuo Wu, Leqi Zheng, Jianke Zhang, Qingbin Li, Shannan Yan, Changguo Jia, Junfei Wu, Zilei Wang
ACL (1)13
2026 MFP-DETR: A Multi-scale Frequency-Aware and Prototype-Guided Transformer Detection Model for Power Inspection Defects
Feng Zhao 0004, Zilei Wang
ICIC (10)6
2026 Layer-wise contrastive network for unsupervised graph representation learning
abstract
Unsupervised graph representation learning has emerged as a cornerstone for extracting meaningful insights from complex relational data. While contrastive learning paradigms have achieved notable success, they traditionally rely on stochastic data augmentations to generate multiple input views—a process that often incurs significant computational overhead and sensitivity to augmentation quality. To transcend these limitations, we propose the Layer-wise Contrastive Network (LCN), a novel and efficient paradigm that redefines the construction of contrastive views. Unlike conventional methods that rely on extrinsic data perturbations, LCN exploits the intrinsic architectural hierarchy of Graph Convolutional Networks. By treating distinct neural layers as different views of the same graph instance, we introduce a contrastive objective that enforces consistency between shallow and deep representations. This mechanism not only eliminates the need for expensive augmentation operations but also distills and preserves fundamental node characteristics from the original graph throughout deeper layers. Furthermore, LCN serves as a flexible plug-and-play framework, exemplified by its extension into Wide LCN, which integrates with traditional augmentation-based methods. Extensive evaluations across both transductive and inductive benchmarks demonstrate that our method achieves superior representational robustness and computational efficiency, offering a scalable and principled perspective for future graph-based contrastive learning. Our code is available at https://github.com/XiangluZhu/LCN.git . • We introduce a novel contrastive loss to learn graph representations by contrasting shallow and deep features. • Our method is flexible and can be readily combined with existing graph contrastive learning techniques that utilize data augmentation. • We demonstrate outstanding performance with our method across four benchmark datasets for node classification.
Xianglu Zhu, Zhang Zhang 0001, Zilei Wang, Liang Wang 0001, Tieniu Tan
Neurocomputing3
2026 Improving Zero-Shot Generalization for CLIP With Prompt Ensemble Self-Distillation
abstract
Prompt tuning has emerged as an effective alternative for adapting pre-trained Vision-Language Models (VLMs) to various downstream tasks. In our experiments utilizing prompt tuning methods, we observed that modifying the prompt initialization led to inconsistencies in the model’s predictions, particularly with pronounced variability on specific datasets. Motivated by this observation, we examine the predictive performance of two ensemble methods: prompt fusion and logits fusion. Experimental results indicate that logits fusion results in considerable performance improvements, while prompt fusion does not yield any enhancements. However, a significant downside of logits fusion is the enormous rise in inference time. To investigate a practical approach for integrating knowledge derived from multiple prompts without incurring additional inference costs, we propose a straightforward Prompt Ensemble self-Distillation (PED) framework that considerably improves the generalization capacity of prompt tuning. Specifically, we initialize multiple groups of prompts, and during the training process, we integrate the prediction outputs from each group to facilitate the learning of the fused prompts. The proposed self-distillation approach offers dual benefits: enhancing the performance of both the fused prompts and the fused logits. We utilize fused prompts for prediction during the inference process, thereby achieving performance that is comparable to that of fused logits without incurring additional inference time. We evaluate the effectiveness of our methodology across four distinct tasks. Our PED consistently demonstrates superior performance in all assessments when contrasted with numerous state-of-the-art methods. Moreover, our method can be seamlessly integrated into existing prompt learning approaches and consistently improves their performance. Our code is publicly available at https://github.com/vim-wei/PED.
Zilei Wang, Yixin Zhang 0007
IEEE Trans. Circuits Syst. Video Technol.2
2025 Target Semantics Clustering via Text Representations for Robust Universal Domain Adaptation
abstract
Universal Domain Adaptation (UniDA) focuses on transferring source domain knowledge to the target domain under both domain shift and unknown category shift. Its main challenge lies in identifying common class samples and aligning them. Current methods typically obtain target domain semantics centers from an unconstrained continuous image representation space. Due to domain shift and the unknown number of clusters, these centers often result in complex and less robust alignment algorithm. In this paper, based on vision-language models, we search for semantic centers in a semantically meaningful and discrete text representation space. The constrained space ensures almost no domain bias and appropriate semantic granularity for these centers, enabling a simple and robust adaptation algorithm. Specifically, we propose TArget Semantics Clustering (TASC) via Text Representations, which leverages information maximization as a unified objective and involves two stages. First, with the frozen encoders, a greedy search-based framework is used to search for an optimal set of text embeddings to represent target semantics. Second, with the search results fixed, encoders are refined based on gradient descent, simultaneously achieving robust domain alignment and private class clustering. Additionally, we propose Universal Maximum Similarity (UniMS), a scoring function tailored for detecting open-set samples in UniDA. Experimentally, we evaluate the universality of UniDA algorithms under four category shift scenarios. Extensive experiments on four benchmarks demonstrate the effectiveness and robustness of our method, which has achieved state-of-the-art performance.
Weinan He 0004, Zilei Wang
AAAI2
2025 Exploring Vacant Classes in Label-Skewed Federated Learning
abstract
Label skews, characterized by disparities in local label distribution across clients, pose a significant challenge in federated learning. As minority classes suffer from worse accuracy due to overfitting on local imbalanced data, prior methods often incorporate class-balanced learning techniques during local training. Although these methods improve the mean accuracy across all classes, we observe that vacant classes—referring to categories absent from a client's data distribution—remain poorly recognized. Besides, there is still a gap in the accuracy of local models on minority classes compared to the global model. This paper introduces FedVLS, a novel approach to label-skewed federated learning that integrates both vacant-class distillation and logit suppression simultaneously. Specifically, vacant-class distillation leverages knowledge distillation during local training on each client to retain essential information related to vacant classes from the global model. Moreover, logit suppression directly penalizes network logits for non-label classes, effectively addressing misclassifications in minority classes that may be biased toward majority classes. Extensive experiments validate the efficacy of FedVLS, demonstrating superior performance compared to previous state-of-the-art (SOTA) methods across diverse datasets with varying degrees of label skews.
Kuangpu Guo, Yuhe Ding, Jian Liang 0001, Zilei Wang, Ran He 0001, Tieniu Tan
AAAI4
2025 Protecting Model Adaptation from Trojans in the Unlabeled Data
abstract
Model adaptation tackles the distribution shift problem with a pre-trained model instead of raw data, which has become a popular paradigm due to its great privacy protection. Existing methods always assume adapting to a clean target domain, overlooking the security risks of unlabeled samples. This paper for the first time explores the potential trojan attacks on model adaptation launched by well-designed poisoning target data. Concretely, we provide two trigger patterns with two poisoning strategies for different prior knowledge owned by attackers. These attacks achieve a high success rate while maintaining the normal performance on clean samples in the test stage. To defend against such backdoor injection, we propose a plug-and-play method named DiffAdapt, which can be seamlessly integrated with existing adaptation algorithms. Experiments across commonly used benchmarks and adaptation methods demonstrate the effectiveness of DiffAdapt. We hope this work will shed light on the safety of transfer learning with unlabeled data.
Lijun Sheng, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan
AAAI4
2025 GCD: Advancing Vision-Language Models for Incremental Object Detection via Global Alignment and Correspondence Distillation
abstract
Incremental object detection (IOD) is a challenging task that requires detection models to continuously learn from newly arriving data. This work focuses on incremental learning for vision-language detectors (VLDs), an under explored domain. Existing research typically adopts a local alignment paradigm to avoid label conflicts, where different tasks are learned separately without interaction. However, we reveal that this practice fails to effectively preserve the semantic structure. Specifically, aligned relationships between objects and texts would collapse when handling novel categories, ultimately leading to catastrophic forgetting. Though knowledge distillation (KD) is a common approach for tackling this, traditional KD performs poorly when directly applied to VLDs, as for different phases, a natural knowledge gap exists in both encoding and decoding processes. To address above issues, we propose a novel method called Global alignment and Correspondence Distillation (GCD). Differently, we first integrate knowledge across phases within the same embedding space to construct global semantic structure. We then enable effective knowledge distillation in VLDs through a semantic correspondence mechanism, ensuring consistent proposal generation and decoding. On the top of that, we distill teacher model’s informative predictions and topological relationships to maintain stable local semantic structure. Extensive experiments on COCO 2017 demonstrate that our method significantly outperforms existing approaches, achieving new state-of-the-art in various IOD scenarios.
Zilei Wang
AAAI2
2025 Divide-Then-Align: Honest Alignment based on the Knowledge Boundary of RAG
abstract
Xin Sun, Jianan Xie, Zhongqi Chen, Qiang Liu, Shu Wu, Yuehe Chen, Bowen Song, Zilei Wang, Weiqiang Wang, Liang Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jianan Xie, Zhongqi Chen, Qiang Liu 0006, Yuehe Chen, Zilei Wang, Liang Wang 0001
ACL (1)8
2025 Towards Continual Universal Segmentation
abstract
Despite the significant progress in continual image segmentation, existing arts still strive to balance between stability and plasticity. Additionally, they are specialist to specific tasks and models, which hinders the extension to more general situations. In this work, we present CUE, a novel Continual Universal sEgmentation pipeline that not only inherently tackles the stability-plasticity dilemma, but unifies any segmentation across tasks and models as well. Our key insight: any segmentation task can be reformulated as an understanding-then-refinement paradigm, which is inspired by humans’ visual perception system to first perform high-level semantic understanding, then focus on low-level vision cues. We claim three desiderata for this design: Continuity by inherently avoiding the stability-plasticity dilemma via exploiting the natural differences between high-level and low-level knowledge. Generality by unifying and simplifying the landscape towards various segmentation tasks. Efficiency as an interesting by-product by significantly reducing the research effort. Our resulting model, built upon this pipeline by complementary expert models, shows significant improvements over previous state-of-the-arts across various segmentation tasks and datasets. We believe that our work is a significant step towards making continual segmentation more universal and practicable.
Zilei Wang
CVPR2
2025 R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
abstract
Vision-language models (VLMs), such as CLIP, have gained significant popularity as foundation models, with numerous fine-tuning methods developed to enhance performance on downstream tasks. However, due to their inherent vulnerability and the common practice of selecting from a limited set of open-source models, VLMs suffer from a higher risk of adversarial attacks than traditional vision models. Existing defense techniques typically rely on adversarial fine-tuning during training, which requires labeled data and lacks of flexibility for downstream tasks. To address these limitations, we propose robust test-time prompt tuning (R-TPT), which mitigates the impact of adversarial attacks during the inference stage. We first reformulate the classic marginal entropy objective by eliminating the term that introduces conflicts under adversarial conditions, retaining only the pointwise entropy minimization. Furthermore, we introduce a plug-and-play reliability-based weighted ensembling strategy, which aggregates useful information from reliable augmented views to strengthen the defense. R-TPT enhances defense against adversarial attacks without requiring labeled training data while offering high flexibility for inference tasks. Extensive experiments on widely used benchmarks with various attacks demonstrate the effectiveness of R-TPT. The code is available in https://github.com/TomSheng21/R-TPT.
Lijun Sheng, Jian Liang 0001, Zilei Wang, Ran He 0001
CVPR3
2025 Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
abstract
Current multimodal large language models (MLLMs) struggle with fine-grained or precise understanding of visuals although they give comprehensive perception and reasoning in a spectrum of vision applications. Recent studies either develop tool-using or unify specific visual tasks into the autoregressive framework, often at the expense of overall multimodal performance. To address this issue and enhance MLLMs with visual tasks in a scalable fashion, we propose Task Preference Optimization (TPO), a novel method that utilizes differentiable task preferences derived from typical fine-grained visual tasks. TPO introduces learnable task tokens that establish connections between multiple task-specific heads and the MLLM. By leveraging rich visual labels during training, TPO significantly enhances the MLLM’s multimodal capabilities and task-specific performance. Through multi-task co-training within TPO, we observe synergistic benefits that elevate individual task performance beyond what is achievable through single-task training methodologies. Our instantiation of this approach with VideoChat and LLaVA demonstrates an overall 14.6% improvement in multimodal performance compared to baseline models. Additionally, MLLM-TPO demonstrates robust zero-shot capabilities across various tasks, performing comparably to state-of-the-art supervised models.
Ziang Yan, Yinan He, Chenting Wang, Kunchang Li 0002, Xinhao Li 0004, Xiangyu Zeng 0004, Zilei Wang, Yali Wang 0001, Yu Qiao 0001, Limin Wang 0002, Yi Wang 0074
CVPR8
2025 Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference
abstract
Multimodal large language models (MLLMs) improve performance on vision-language tasks by integrating visual features from pre-trained vision encoders into large language models (LLMs). However, how MLLMs process and utilize visual information remains unclear. In this paper, a shift in the dominant flow of visual information is uncovered: (1) in shallow layers, strong interactions are observed between image tokens and instruction tokens, where most visual information is injected into instruction tokens to form cross-modal semantic representations; (2) in deeper layers, image tokens primarily interact with each other, aggregating the remaining visual information to optimize semantic representations within visual modality. Based on these insights, we propose Hierarchical Modality-Aware Pruning (HiMAP), a plug-and-play inference acceleration method that dynamically prunes image tokens at specific layers, reducing computational costs by approximately 65% without sacrificing performance. Our findings offer a new understanding of visual information processing in MLLMs and provide a state-of-the-art solution for efficient inference. The code is available at https://github.com/ustc-hyin/HiMAP.
Guangzong Si, Zilei Wang
CVPR3
2025 ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large Language Models
abstract
Contrastive decoding strategies are widely used to mitigate object hallucinations in multimodal large language models (MLLMs). By reducing over-reliance on language priors, these strategies ensure that generated content remains closely grounded in visual inputs, producing contextually accurate outputs. Since contrastive decoding requires no additional training or external tools, it offers both computational efficiency and versatility, making it highly attractive. However, these methods present two main limitations: (1) bluntly suppressing language priors can compromise coherence and accuracy of generated content, and (2) processing contrastive inputs adds computational load, significantly slowing inference speed. To address these challenges, we propose Visual Amplification Fusion (VAF), a plug-and-play technique that enhances attention to visual signals within the model’s middle layers, where modality fusion predominantly occurs. This approach enables more effective capture of visual features, reducing the model’s bias toward language modality. Experimental results demonstrate that VAF significantly reduces hallucinations across various MLLMs without affecting inference speed, while maintaining coherence and accuracy in generated outputs. The code is available at https://github.com/ustc-hyin/ClearSight.
Guangzong Si, Zilei Wang
CVPR3
2025 Cooperative Pseudo Labeling for Unsupervised Federated Classification
Kuangpu Guo, Lijun Sheng, Yongcan Yu, Jian Liang 0001, Zilei Wang, Ran He 0001
ICCV5
2025 Progressive Distribution Bridging: Unsupervised Adaptation for Large-Scale Pre-Trained Models via Adaptive Auxiliary Data
Weinan He 0004, Zilei Wang
ICCV3
2025 LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
abstract
Low-rank adaptation, also known as LoRA, has emerged as a prominent method for parameter-efficient fine-tuning of foundation models. Despite its computational efficiency, LoRA still yields inferior performance compared to full fine-tuning. In this paper, we first uncover a fundamental connection between the optimization processes of LoRA and full fine-tuning: using LoRA for optimization is mathematically equivalent to full fine-tuning using a low-rank gradient for parameter updates. And this low-rank gradient can be expressed in terms of the gradients of the two low-rank matrices in LoRA. Leveraging this insight, we introduce LoRA-Pro, a method that enhances LoRA's performance by strategically adjusting the gradients of these low-rank matrices. This adjustment allows the low-rank gradient to more accurately approximate the full fine-tuning gradient, thereby narrowing the performance gap between LoRA and full fine-tuning. Furthermore, we theoretically derive the optimal solutions for adjusting the gradients of the low-rank matrices, applying them during fine-tuning in LoRA-Pro. We conduct extensive experiments across natural language understanding, dialogue generation, mathematical reasoning, code generation, and image classification tasks, demonstrating that LoRA-Pro substantially improves LoRA's performance, effectively narrowing the gap with full fine-tuning. Our code is publicly available at https://github.com/mrflogs/LoRA-Pro.
Zhengbo Wang, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan
ICLR4
2025 The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models
abstract
Test-time adaptation (TTA) methods have gained significant attention for enhancing the performance of vision-language models (VLMs) such as CLIP during inference, without requiring additional labeled data. However, current TTA researches generally suffer from major limitations such as duplication of baseline results, limited evaluation metrics, inconsistent experimental settings, and insufficient analysis. These problems hinder fair comparisons between TTA methods and make it difficult to assess their practical strengths and weaknesses. To address these challenges, we introduce TTA-VLM, a comprehensive benchmark for evaluating TTA methods on VLMs. Our benchmark implements 8 episodic TTA and 7 online TTA methods within a unified and reproducible framework, and evaluates them across 15 widely used datasets. Unlike prior studies focused solely on CLIP, we extend the evaluation to SigLIP—a model trained with a Sigmoid loss—and include training-time tuning methods such as CoOp, MaPLe, and TeCoA to assess generality. Beyond classification accuracy, TTA-VLM incorporates various evaluation metrics, including robustness, calibration, out-of-distribution detection, and stability, enabling a more holistic assessment of TTA methods. Through extensive experiments, we find that 1) existing TTA methods produce limited gains compared to the previous pioneering work; 2) current TTA methods exhibit poor collaboration with training-time fine-tuning methods; 3) accuracy gains frequently come at the cost of reduced model trustworthiness. We release TTA-VLM to provide fair comparison and comprehensive evaluation of TTA methods for VLMs, and we hope it encourages the community to develop more reliable and generalizable TTA strategies. The code is available in https://github.com/TomSheng21/tta-vlm.
Lijun Sheng, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan
NeurIPS4
2025 The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
abstract
Contrastive decoding strategies are widely used to reduce object hallucinations in multimodal large language models (MLLMs). These methods work by constructing contrastive samples to induce hallucinations and then suppressing them in the output distribution. However, this paper demonstrates that such approaches fail to effectively mitigate the hallucination problem. The performance improvements observed on POPE Benchmark are largely driven by two misleading factors: (1) crude, unidirectional adjustments to the model’s output distribution and (2) the adaptive plausibility constraint, which reduces the sampling strategy to greedy search. To further illustrate these issues, we introduce a series of spurious improvement methods and evaluate their performance against contrastive decoding techniques. Experimental results reveal that the observed performance gains in contrastive decoding are entirely unrelated to its intended goal of mitigating hallucinations. Our findings challenge common assumptions about the effectiveness of contrastive decoding strategies and pave the way for developing genuinely effective solutions to hallucinations in MLLMs.
Guangzong Si, Zilei Wang
NeurIPS3
2025 Category-instance distillation based on visual-language models for rehearsal-free class incremental learning
abstract
Abstract Recently, visual‐language models (VLMs) have displayed potent capabilities in the field of computer vision. Their emerging trend as the backbone of visual tasks necessitates studying class incremental learning (CIL) issues within the VLM architecture. However, the pre‐training data for many VLMs is proprietary, and during the incremental phase, old task data may also raise privacy issues. Moreover, replay‐based methods can introduce new problems like class imbalance, the selection of data for replay and a trade‐off between replay cost and performance. Therefore, the authors choose the more challenging rehearsal‐free settings. In this paper, the authors study class‐incremental tasks based on the large pre‐trained vision‐language models like CLIP model. Initially, at the category level, the authors combine traditional optimisation and distillation techniques, utilising both pre‐trained models and models trained in previous incremental stages to jointly guide the training of the new model. This paradigm effectively balances the stability and plasticity of the new model, mitigating the issue of catastrophic forgetting. Moreover, utilising the VLM infrastructure, the authors redefine the relationship between instances. This allows us to glean fine‐grained instance relational information from the a priori knowledge provided during pre‐training. The authors supplement this approach with an entropy‐balancing method that allows the model to adaptively distribute optimisation weights across training samples. The authors’ experimental results validate that their method, within the framework of VLMs, outperforms traditional CIL methods.
Weilong Jin, Zilei Wang
IET Comput. Vis.2
2025 Context Sensitive Network for weakly-supervised fine-grained temporal action localization
Cerui Dong, Qinying Liu, Zilei Wang, Yixin Zhang 0007, Feng Zhao 0004
Neural Networks3
2025 Multilevel semantic and adaptive actionness learning for weakly supervised temporal action localization
Zilei Wang, Cerui Dong
Neural Networks2
2025 Semantic-Aware Late-Stage Supervised Contrastive Learning for Fine-Grained Action Recognition
abstract
Fine-grained action recognition typically faces challenges with lower inter-class variances and higher intra-class variances. Supervised contrastive learning is inherently suitable for this task, as it can decrease intra-class feature distances while increasing inter-class ones. However, directly applying it into fine-grained action recognition encounters two main problems. The first problem stems from the heavy training cost associated with supervised contrastive learning, which requires numerous training epochs, each involving double augmentation views per instance. To address this issue, we propose the late-stage supervised contrastive learning (late-SC) strategy, which effectively reduces the number of training epochs needed for the contrastive learning process. The second problem is that supervised contrastive loss does not explicitly consider the semantic distances between fine-grained actions when adjusting representation distances. This results in less reasonable and efficient adjustments to the representation space. To overcome this limitation, we introduce the semantic-aware temperature adaptation (STA) mechanism, enhancing the suitability of the supervised contrastive loss for fine-grained action recognition. We conduct experiments on several benchmark datasets for fine-grained action recognition, including Epic-Kitchens-55/100, SomethingSomething-V1, and Diving48-V2. The results demonstrate that our proposed method (referred to as LSC-STA) consistently enhances performance across various base feature extractors, without introducing additional inference overhead and incurring only a marginal increase in training expenses.
Yijun Pan, Yueyi Zhang 0001, Zilei Wang, Xiaoyan Sun 0001, Feng Wu 0005
IEEE Trans. Circuits Syst. Video Technol.4
2024 A Dynamic Learning Method towards Realistic Compositional Zero-Shot Learning
abstract
To tackle the challenge of recognizing images of unseen attribute-object compositions, Compositional Zero-Shot Learning (CZSL) methods have been previously addressed. However, test images in realistic scenarios may also incorporate other forms of unknown factors, such as novel semantic concepts or novel image styles. As previous CZSL works have overlooked this critical issue, in this research, we first propose the Realistic Compositional Zero-Shot Learning (RCZSL) task which considers the various types of unknown factors in an unified experimental setting. To achieve this, we firstly conduct re-labelling on MIT-States and use the pre-trained generative models to obtain images of various domains. Then the entire dataset is split into a training set and a test set, with the latter containing images of unseen concepts, unseen compositions, unseen domains as well as their combinations. Following this, we show that the visual-semantic relationship changes on unseen images, leading us to construct two dynamic modulators to adapt the visual features and composition prototypes in accordance with the input image. We believe that such a dynamic learning method could effectively alleviate the domain shift problem caused by various types of unknown factors. We conduct extensive experiments on benchmark datasets for both the conventional CZSL setting and the proposed RCZSL setting. The effectiveness of our method has been proven by empirical results, which significantly outperformed both our baseline method and state-of-the-art approaches.
Zilei Wang
AAAI2
2024 Evolving to the Future: Unseen Event Adaptive Fake News Detection on Social Media
Jiajun Zhang 0012, Zhixun Li, Qiang Liu 0006, Zilei Wang, Liang Wang 0001
CIKM5
2024 Infer from What You Have Seen Before: Temporally-dependent Classifier for Semi-supervised Video Segmentation
abstract
Due to high expense of human labor, one major challenge for semantic segmentation in real-world scenarios is the lack of sufficient pixel-level labels, which is more serious when processing video data. To exploit unlabeled data for model training, semi-supervised learning methods attempt to construct pseudo labels or various auxiliary constraints as supervision signals. However, most of them just process video data as a set of independent images in a per-frame manner. The rich temporal relationships are ignored, which can serve as valuable clues for representation learning. Besides, this per-frame recognition paradigm is quite different from that of humans. Actually, benefited from the internal temporal relevance of video data, human would wisely use the distinguished semantic concepts in historical frames to aid the recognition of the current frame. Motivated by this observation, we propose a novel temporally-dependent classifier (TDC) to mimic the human-like recognition procedure. Comparing to the conventional classifier, TDC can guide the model to learn a group of temporally-consistent semantic concepts across frames, which essentially provides an implicit and effective constraint. We conduct extensive experiments on Cityscapes and Cam Vid, and the results demonstrate the superiority of our proposed method to previous state-of-the-art methods. The code is available at https://github.com/jfzhuang/TDC.
Jiafan Zhuang, Zilei Wang, Zhun Fan
CVPR2
2024 Efficient Active Domain Adaptation for Semantic Segmentation by Selecting Information-Rich Superpixels
Zilei Wang, Bohai Tu
ECCV (34)2
2024 Semantic-Guided Robustness Tuning for Few-Shot Transfer Across Extreme Domain Shift
Kangyu Xiao, Zilei Wang, Junjie Li 0002
ECCV (49)2
2024 A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation
abstract
Contrastive Language-Image Pretraining (CLIP) has gained popularity for its remarkable zero-shot capacity. Recent research has focused on developing efficient fine-tuning methods, such as prompt learning and adapter, to enhance CLIP's performance in downstream tasks. However, these methods still require additional training time and computational resources, which is undesirable for devices with limited resources. In this paper, we revisit a classical algorithm, Gaussian Discriminant Analysis (GDA), and apply it to the downstream classification of CLIP. Typically, GDA assumes that features of each class follow Gaussian distributions with identical covariance. By leveraging Bayes' formula, the classifier can be expressed in terms of the class means and covariance, which can be estimated from the data without the need for training. To integrate knowledge from both visual and textual modalities, we ensemble it with the original zero-shot classifier within CLIP. Extensive results on 17 datasets validate that our method surpasses or achieves comparable results with state-of-the-art methods on few-shot classification, imbalanced learning, and out-of-distribution generalization. In addition, we extend our method to base-to-new generalization and unsupervised learning, once again demonstrating its superiority over competing approaches. Our code is publicly available at https://github.com/mrflogs/ICLR24.
Zhengbo Wang, Jian Liang 0001, Lijun Sheng, Ran He 0001, Zilei Wang, Tieniu Tan
ICLR5
2024 Delve into Source and Target Collaboration in Semi-supervised Domain Adaptation for Semantic Segmentation
abstract
Semi-supervised domain adaptation (SSDA) for semantic segmentation aims to train a model that performs well on target domain by learning from both fully-labeled source domain and partially-labeled target domain data. The key to this task is how to collaborate the labeled data from both domains, as well as the unlabeled data, to benefit the model training complementarily. In this paper, we innovatively achieve this goal from both data combination and model mergence perspectives. To this end, we propose a co-training framework based on siamese networks, where two networks are encouraged to learn from each other by cross-supervision with pseudo labels of unlabeled data. Meanwhile, for the labeled data, we enforce two networks to separately learn the knowledge dominated by source domain and target domain. Specifically, we propose domain-specific initialization and differentiated cross-domain combination of labeled data. Moreover, we propose a target-preferred alignment method to encourage the source-biased network to optimize towards target domain, as the target-biased network is more in line with the task than the source-biased network. We conduct extensive experiments on two challenging benchmarks, and the results demonstrate the effectiveness of our method, which outperforms previous state-of-the-art methods with considerable performance improvement. Our code is available at https://github.com/EdenHazardan/DSTC-SSDA.
Zilei Wang
ICME2
2024 Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
abstract
With the emergence of pretrained vision-language models (VLMs), considerable efforts have been devoted to fine-tuning them for downstream tasks. Despite the progress made in designing efficient fine-tuning methods, such methods require access to the model’s parameters, which can be challenging as model owners often opt to provide their models as a black box to safeguard model ownership. This paper proposes a Collaborative Fine-Tuning (CraFT) approach for fine-tuning black-box VLMs to downstream tasks, where one only has access to the input prompts and the output predictions of the model. CraFT comprises two modules, a prompt generation module for learning text prompts and a prediction refinement module for enhancing output predictions in residual style. Additionally, we introduce an auxiliary prediction-consistent loss to promote consistent optimization across these modules. These modules are optimized by a novel collaborative training algorithm. Extensive experiments on few-shot classification over 15 datasets demonstrate the superiority of CraFT. The results show that CraFT achieves a decent gain of about 12% with 16-shot datasets and only 8,000 queries. Moreover, CraFT trains faster and uses only about 1/80 of the memory footprint for deployment, while sacrificing only 1.62% compared to the white-box method. Our code is publicly available at https://github.com/mrflogs/CraFT.
Zhengbo Wang, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan
ICML4
2024 Probabilistic Contrastive Learning for Domain Adaptation
Junjie Li 0002, Yixin Zhang 0007, Zilei Wang, Saihui Hou, Keyu Tu, Man Zhang 0005
IJCAI3
2024 Learning Energy-Based Models for 3D Human Pose Estimation
abstract
Recently, 3D human pose estimation has attracted more attention due to its promising applications. In general, existing methods usually directly predict a target 3D pose for a given input using a Deep Neural Network (DNN), and train the DNN by minimizing the mean squared error (MSE) loss. Despite the impressive performance of these methods, they create a fixed-variance Gaussian model of the conditional target density (the distribution for the target 3D pose given the input) from a probabilistic perspective, which significantly restricts the expressive capabilities of the learned conditional target density. Thus, this hinders the complete utilization of the predictive potential embedded within the DNN. We tackle this problem by delving into the latest developments in conditional energy-based models (EBMs) for probabilistic regression. In this work, we design a simple yet effective network to learn an energy function from 2D and 3D joints pairs. Then a gradient-based refinement procedure is adopted to minimize the energy function to find the corresponding target 3D pose. In this way, we can apply the energy-based model to refine the initial 3D joints estimated by the state-of-the-art 3D human pose estimator. Extensive experiments are conducted on two popular benchmarks on human pose estimation and the results demonstrate the superiority of our method over existing state-of-the-art approaches.
Xianglu Zhu, Zhang Zhang 0001, Wei Wang 0115, Zilei Wang, Liang Wang 0001
IJCNN4
2024 DIVE: Subgraph Disagreement for Graph Out-of-Distribution Generalization
abstract
This paper addresses the challenge of out-of-distribution (OOD) generalization in graph machine learning, a field rapidly advancing yet grappling with the discrepancy between source and target data distributions. Traditional graph learning algorithms, based on the assumption of uniform distribution between training and test data, falter in real-world scenarios where this assumption fails, resulting in suboptimal performance. A principal factor contributing to this suboptimal performance is the inherent simplicity bias of neural networks trained through Stochastic Gradient Descent (SGD), which prefer simpler features over more complex yet equally or more predictive ones. This bias leads to a reliance on spurious correlations, adversely affecting OOD performance in various tasks such as image recognition, natural language understanding, and graph classification. Current methodologies, including subgraph-mixup and information bottleneck approaches, have achieved partial success but struggle to overcome simplicity bias, often reinforcing spurious correlations. To tackle this, our study introduces a new learning paradigm for graph OOD issue. We propose DIVE, training a collection of models to focus on all label-predictive subgraphs by encouraging the models to foster divergence on the subgraph mask, which circumvents the limitation of a model solely focusing on the subgraph corresponding to simple structural patterns. Specifically, we employs a regularizer to punish overlap in extracted subgraphs across models, thereby encouraging different models to concentrate on distinct structural patterns. Model selection for robust OOD performance is achieved through validation accuracy. Tested across four datasets from GOOD benchmark and one dataset from DrugOOD benchmark, our approach demonstrates significant improvement over existing methods, effectively addressing the simplicity bias and enhancing generalization in graph machine learning.
Liang Wang 0056, Qiang Liu 0006, Zilei Wang, Liang Wang 0001
KDD5
2024 Restoration towards decomposition: A simple approach for domain generalization
Zilei Wang
Inf. Sci.2
2024 Weakly supervised temporal action localization with actionness-guided false positive suppression
Zilei Wang, Qinying Liu
Neural Networks2
2024 RISTRA: Recursive Image Super-Resolution Transformer With Relativistic Assessment
abstract
Many recent image restoration methods use Transformer as the backbone network and redesign the Transformer blocks. Differently, we explore the parameter-sharing mechanism over Transformer blocks and propose a dynamic recursive process to address the image super-resolution task efficiently. We firstly present a Recursive Image Super-resolution Transformer (RIST). By sharing the weights across different blocks, a plain forward process through the whole Transformer network can be folded into recursive iterations through a Transformer block. Such a parameter-sharing based recursive process can not only reduce the model size greatly, but also enable restoring images progressively. Features in the recursive process are modeled as a sequence and propagated with a temporal attention network. Besides, by analyzing the prediction variation across different iterations in RIST, we design a dynamic recursive process that can allocate adaptive computation costs to different samples. Specifically, a quality assessment network estimates the restoration quality and terminates the recursive process dynamically. We propose a relativistic learning strategy to simplify the objective from absolute image quality assessment to relativistic quality comparison. The proposed Recursive Image Super-resolution Transformer with Relativistic Assessment (RISTRA) reduces the model size greatly with the parameter-sharing mechanism, and achieves an instance-wise dynamic restoration process as well. Extensive experiments on several image super-resolution benchmarks show the superiority of our approach over state-of-the-art counterparts
Xiaoqiang Zhou, Huaibo Huang, Zilei Wang, Ran He 0001
IEEE Trans. Multim.3
2024 Learning Lightweight Dynamic Kernels With Attention Inside via Local-Global Context Fusion
abstract
Traditional convolutional neural networks (CNNs) share their kernels among all positions of the input, which may constrain the representation ability in feature extraction. Dynamic convolution proposes to generate different kernels for different inputs to improve the model capacity. However, the total parameters of the dynamic network can be significantly huge. In this article, we propose a lightweight dynamic convolution method to strengthen traditional CNNs with an affordable increase of total parameters and multiply-adds. Instead of generating the whole kernels directly or combining several static kernels, we choose to "look inside," learning the attention within convolutional kernels. An extra network is used to adjust the weights of kernels for every feature aggregation operation. By combining local and global contexts, the proposed approach can capture the variance among different samples, the variance in different positions of the feature maps, and the variance in different positions inside sliding windows. With a minor increase in the number of model parameters, remarkable improvements in image classification on CIFAR and ImageNet with multiple backbones have been obtained. Experiments on object detection also verify the effectiveness of the proposed method.
Yonglin Tian, Xiao Wang 0002, Jiangong Wang, Kunfeng Wang, Weiping Ding 0001, Zilei Wang, Fei-Yue Wang 0001
IEEE Trans. Neural Networks Learn. Syst.7
2023 Exploit Domain-Robust Optical Flow in Domain Adaptive Video Semantic Segmentation
abstract
Domain adaptive semantic segmentation aims to exploit the pixel-level annotated samples on source domain to assist the segmentation of unlabeled samples on target domain. For such a task, the key is to construct reliable supervision signals on target domain. However, existing methods can only provide unreliable supervision signals constructed by segmentation model (SegNet) that are generally domain-sensitive. In this work, we try to find a domain-robust clue to construct more reliable supervision signals. Particularly, we experimentally observe the domain-robustness of optical flow in video tasks as it mainly represents the motion characteristics of scenes. However, optical flow cannot be directly used as supervision signals of semantic segmentation since both of them essentially represent different information. To tackle this issue, we first propose a novel Segmentation-to-Flow Module (SFM) that converts semantic segmentation maps to optical flows, named the segmentation-based flow (SF), and then propose a Segmentation-based Flow Consistency (SFC) method to impose consistency between SF and optical flow, which can implicitly supervise the training of segmentation model. The extensive experiments on two challenging benchmarks demonstrate the effectiveness of our method, and it outperforms previous state-of-the-art methods with considerable performance improvement. Our code is available at https://github.com/EdenHazardan/SFC.
Zilei Wang, Jiafan Zhuang, Yixin Zhang 0007, Junjie Li 0002
AAAI2
2023 Leveraging Sub-class Discimination for Compositional Zero-Shot Learning
abstract
Compositional Zero-Shot Learning (CZSL) aims at identifying unseen compositions composed of previously seen attributes and objects during the test phase. In real images, the visual appearances of attributes and objects (primitive concepts) generally interact with each other. Namely, the visual appearances of an attribute may change when composed with different objects, and vice versa. But previous works overlook this important property. In this paper, we introduce a simple yet effective approach with leveraging sub-class discrimination. Specifically, we define the primitive concepts in different compositions as sub-classes, and then maintain the sub-class discrimination to address the above challenge. More specifically, inspired by the observation that the composed recognition models could account for the differences across sub-classes, we first propose to impose the embedding alignment between the composed and disentangled recognition to incorporate sub-class discrimination at the feature level. Then we develop the prototype modulator networks to adjust the class prototypes w.r.t. the composition information, which can enhance sub-class discrimination at the classifier level. We conduct extensive experiments on the challenging benchmark datasets, and the considerable performance improvement over state-of-the-art approaches is achieved, which indicates the effectiveness of our method. Our code is available at https://github.com/hxm97/SCD-CZSL.
Zilei Wang
AAAI2
2023 Actionness Inconsistency-Guided Contrastive Learning for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (WTAL) aims to detect action instances given only video-level labels. To address the challenge, recent methods commonly employ a two-branch framework, consisting of a class-aware branch and a class-agnostic branch. In principle, the two branches are supposed to produce the same actionness activation. However, we observe that there are actually many inconsistent activation regions. These inconsistent regions usually contain some challenging segments whose semantic information (action or background) is ambiguous. In this work, we propose a novel Actionness Inconsistency-guided Contrastive Learning (AICL) method which utilizes the consistent segments to boost the representation learning of the inconsistent segments. Specifically, we first define the consistent and inconsistent segments by comparing the predictions of two branches and then construct positive and negative pairs between consistent segments and inconsistent segments for contrastive learning. In addition, to avoid the trivial case where there is no consistent sample, we introduce an action consistency constraint to control the difference between the two branches. We conduct extensive experiments on THUMOS14, ActivityNet v1.2, and ActivityNet v1.3 datasets, and the results show the effectiveness of AICL with state-of-the-art performance. Our code is available at https://github.com/lizhilin-ustc/AAAI2023-AICL.
Zilei Wang, Qinying Liu
AAAI2
2023 SimpleNet: A Simple Network for Image Anomaly Detection and Localization
abstract
We propose a simple and application-friendly network (called SimpleNet) for detecting and localizing anoma-lies. SimpleNet consists of four components: (1) a pre-trained Feature Extractor that generates local features, (2) a shallow Feature Adapter that transfers local features to-wards target domain, (3) a simple Anomaly Feature Gener-ator that counterfeits anomaly features by adding Gaussian noise to normal features, and (4) a binary Anomaly Discriminator that distinguishes anomaly features from normal features. During inference, the Anomaly Feature Generator would be discarded. Our approach is based on three in-tuitions. First, transforming pre-trained features to target-oriented features helps avoid domain bias. Second, gen-erating synthetic anomalies in feature space is more effective, as defects may not have much commonality in the image space. Third, a simple discriminator is much efficient and practical. In spite of simplicity, SimpleNet outper-forms previous methods quantitatively and qualitatively. On The MVTec AD benchmark, SimpleNet achieves an anomaly detection AUROC of 99.6%, reducing the error by 55.5% compared to the next best performing model. Further-more, SimpleNet is faster than existing methods, with a high frame rate of 77 FPS on a 3080ti GPU. Additionally, SimpleNet demonstrates significant improvements in per-formance on the One-Class Novelty Detection task. Code: https://github.com/DonaldRR/SimpleNet.
Zhikang Liu, Yuansheng Xu, Zilei Wang
CVPR4
2023 Boundary-enhanced Co-training for Weakly Supervised Semantic Segmentation
abstract
The existing weakly supervised semantic segmentation (WSSS) methods pay much attention to generating accurate and complete class activation maps (CAMs) as pseudo-labels, while ignoring the importance of training the segmentation networks. In this work, we observe that there is an inconsistency between the quality of the pseudo-labels in CAMs and the performance of the final segmentation model, and the mislabeled pixels mainly lie on the boundary areas. Inspired by these findings, we argue that the focus of WSSS should be shifted to robust learning given the noisy pseudo-labels, and further propose a boundary-enhanced co-training (BECO) method for training the segmentation model. To be specific, we first propose to use a co-training paradigm with two interactive networks to improve the learning of uncertain pixels. Then we propose a boundary-enhanced strategy to boost the prediction of difficult boundary areas, which utilizes reliable predictions to construct artificial boundaries. Benefiting from the design of co-training and boundary enhancement, our method can achieve promising segmentation performance for different CAMs. Extensive experiments on PASCAL VOC 2012 and MS COCO 2014 validate the superiority of our BECO over other state-of-the-art methods.11The code and models are available at https://github.com/ShenghaiRong/BECO.
Shenghai Rong, Bohai Tu, Zilei Wang, Junjie Li 0002
CVPR3
2023 Class Relationship Embedded Learning for Source-Free Unsupervised Domain Adaptation
abstract
This work focuses on a practical knowledge transfer task defined as Source-Free Unsupervised Domain Adaptation (SFUDA), where only a well-trained source model and unlabeled target data are available. To fully utilize source knowledge, we propose to transfer the class relationship, which is domain-invariant but still under-explored in previous works. To this end, we first regard the classifier weights of the source model as class prototypes to compute class relationship, and then propose a novel probability-based similarity between target-domain samples by embedding the source-domain class relationship, resulting in Class Relationship embedded Similarity (CRS). Here the inter-class term is particularly considered in order to more accurately represent the similarity between two samples, in which the source prior of class relationship is utilized by weighting. Finally, we propose to embed CRS into contrastive learning in a unified form. Here both class-aware and instance discrimination contrastive losses are employed, which are complementary to each other. We combine the proposed method with existing representative methods to evaluate its efficacy in multiple SFUDA settings. Extensive experimental results reveal that our method can achieve state-of-the-art performance due to the transfer of domain-invariant class relationship.11Code is available at https://github.com/zhyx12/CRCo
Zilei Wang, Weinan He 0004
CVPR2
2023 Notice of Removal: Exploiting Semantic Attributes for Transductive Zero-Shot Learning
abstract
Removed.
Zhengbo Wang, Jian Liang 0001, Zilei Wang, Tieniu Tan
ICASSP3
2023 Preparing the Future for Continual Semantic Segmentation
abstract
In this study, we focus on Continual Semantic Segmentation (CSS) and present a novel approach to tackle the issue of existing methods struggling to learn new classes. The primary challenge of CSS is to learn new knowledge while retaining old knowledge, which is commonly known as the rigidity-plasticity dilemma. Existing approaches strive to address this by carefully balancing the learning of new and old classes during training on new data. Differently, this work aims to avoid this dilemma fundamentally rather than handling the difficulties involved in it. Specifically, we reveal that this dilemma mainly arises from the greater fluctuation of knowledge for new classes because they have never been learned before the current step. Additionally, the data available in incremental steps are usually inadequate, which can impede the model’s ability to learn discriminative features for both new and old classes. To address these challenges, we introduce a novel concept of pre-learning for future knowledge. Our approach entails optimizing the feature space and output space for unlabeled data, which thus enables the model to acquire knowledge for future classes. With this approach, updating the model for new classes becomes as smooth as for old classes, effectively avoiding the rigidity-plasticity dilemma. We conducted extensive experiments and the results demonstrate a significant improvement in the learning of new classes compared to previous state-of-the-art methods.
Zilei Wang
ICCV2
2023 Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based Approach
abstract
Weakly-supervised temporal action localization aims to localize action instances in videos with only video-level action labels. Existing methods mainly embrace a localization-by-classification pipeline that optimizes the snippet-level prediction with a video classification loss. However, this formulation suffers from the discrepancy between classification and detection, resulting in inaccurate separation of foreground and background (F&B) snippets. To alleviate this problem, we propose to explore the underlying structure among the snippets by resorting to unsupervised snippet clustering, rather than heavily relying on the video classification loss. Specifically, we propose a novel clustering-based F&B separation algorithm. It comprises two core components: a snippet clustering component that groups the snippets into multiple latent clusters and a cluster classification component that further classifies the cluster as foreground or background. As there are no ground-truth labels to train these two components, we introduce a unified self-labeling mechanism based on optimal transport to produce high-quality pseudo-labels that match several plausible prior distributions. This ensures that the cluster assignments of the snippets can be accurately associated with their F&B labels, thereby boosting the F&B separation. We evaluate our method on three benchmarks: THUMOS14, ActivityNet v1.2 and v1.3. Our method achieves promising performance on all three benchmarks while being significantly more lightweight than previous methods. Code is available at https://github.com/Qinying-Liu/CASE
Qinying Liu, Zilei Wang, Shenghai Rong, Junjie Li 0002, Yixin Zhang 0007
ICCV2
2023 Towards Effective Instance Discrimination Contrastive Loss for Unsupervised Domain Adaptation
abstract
Domain adaptation (DA) aims to transfer knowledge from a label-rich source domain to a related but label-scarce target domain. Recently, increasing research has focused on exploring data structure of the target domain. In light of the recent success of Instance Discrimination Contrastive (IDCo) loss in self-supervised learning, we try directly applying it to domain adaptation tasks. However, the improvement is very limited, which motivates us to rethink its underlying limitations for domain adaptation tasks. An intuitive limitation is that a pair of samples belonging to the same class could be treated as negatives. Here we argue that using low-confidence samples to construct positive and negative pairs can alleviate this issue and is more suitable for IDCo loss. Another limitation is that IDCo loss cannot capture enough semantic information. We address this by introducing domain-invariant and accurate semantic information from classifier weights and input data. Specifically, we propose a class relationship enhanced features. It uses probability weighted class prototpyes as the input features of IDCo loss, which can implicitly transfer the domain-invariant class relationship. We further propose a target-dominated cross-domain mixup that can incorporate accurate semantic information from the source domain. We evaluate the proposed method in unsupervised DA and other DA settings, and extensive experimental results reveal that our method can make IDCo loss more effective and achieve state-of-the-art performance.1
Yixin Zhang 0007, Zilei Wang, Junjie Li 0002, Jiafan Zhuang
ICCV2
2023 Unleashing the Potential of Adjacent Snippets for Weakly-supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (WTAL) intends to detect action instances with only weak supervision, e.g., video-level labels. The current de facto pipeline locates action instances by thresholding and grouping continuous high-score regions on temporal class activation sequences. In this route, the capacity of the model to recognize the relationships between adjacent snippets is of vital importance which determines the quality of the action boundaries. However, it is error-prone since the variations between adjacent snippets are typically subtle, and unfortunately this is overlooked in the literature. To tackle the issue, we propose a novel WTAL approach named Convex Combination Consistency between Neighbors (C3BN). C3BN consists of two key ingredients: a micro data augmentation strategy that increases the diversity in-between adjacent snippets by convex combination of adjacent snippets, and a macro-micro consistency regularization that enforces the model to be invariant to the transformations w.r.t. video semantics, snippet predictions, and snippet representations. Consequently, fine-grained patterns in-between adjacent snippets are enforced to be explored, thereby resulting in a more robust action boundary localization. Experimental results demonstrate the effectiveness of C3BN on top of various baselines for WTAL with video-level and point-level supervision. Code is at: https://github.com/canbaoburen/C3BN.
Qinying Liu, Zilei Wang, Ruoxi Chen
ICME2
2023 Exploiting Low-confidence Pseudo-labels for Source-free Object Detection
abstract
Source-free object detection (SFOD) aims to adapt a source-trained detector to an unlabeled target domain without access to the labeled source data. Current SFOD methods utilize a threshold-based pseudo-label approach in the adaptation phase, which is typically limited to high-confidence pseudo-labels and results in a loss of information. To address this issue, we propose a new approach to take full advantage of pseudo-labels by introducing high and low confidence thresholds. Specifically, the pseudo-labels with confidence scores above the high threshold are used conventionally, while those between the low and high thresholds are exploited using the Low-confidence Pseudo-labels Utilization (LPU) module. The LPU module consists of Proposal Soft Training (PST) and Local Spatial Contrastive Learning (LSCL). PST generates soft labels of proposals for soft training, which can mitigate the label mismatch problem. LSCL exploits the local spatial relationship of proposals to improve the model's ability to differentiate between spatially adjacent proposals, thereby optimizing representational features further. Combining the two components overcomes the challenges faced by traditional methods in utilizing low-confidence pseudo-labels. Extensive experiments on five cross-domain object detection benchmarks demonstrate that our proposed method outperforms the previous SFOD methods, achieving state-of-the-art performance.
Zilei Wang, Yixin Zhang 0007
ACM Multimedia2
2023 Sparse Sharing Relation Network for Panoptic Driving Perception
abstract
Efficient and accurate perception system is critical for autonomous driving, including traffic object detection, drivable area segmentation, and lane detection. Most previous works do not consider the spatial and semantic cues in traffic scenes. In this paper, we propose a novel multi-task learning network to exploit these priors. Specifically, to model the co-occurrence and spatial relationships of traffic objects, we propose to use a Graph Convolutional Network (GCN) block operating on the patches of feature maps. It enables adaptive discovery and incorporation of semantic and spatial relationships in the feature space. Furthermore, we propose a sub-feature sharing method to mitigate negative transfer in multi-task learning. On the basis of a fully shared base network, we split the feature space of different tasks along the channel dimension, resulting in the shared and private features for each task. It allows the network parameters to be selectively updated by different tasks during training. Experimental results on the challenging BDD100K dataset demonstrate that our proposed approach gets consistent improvement with fewer parameters, and achieves new state-of-the-art performance in terms of accuracy and speed.
Fan Jiang 0009, Zilei Wang
ACM Multimedia2
2023 Semi-supervised Domain Adaptation via Joint Contrastive Learning with Sensitivity
abstract
Semi-supervised Domain Adaptation (SSDA) aims to learn a well-performed model using fully labeled source samples and scarcely labeled target samples, along with unlabeled target samples. Due to the dominant presence of labeled samples from the source domain in the training data, both the feature extractor and classifier can display bias towards the source domain. This can result in sub-optimal feature extraction for the challenging target samples that have notable differences from the source domain. Moreover, the source-favored classifier can hinder the classification performance of the target domain. To this end, we propose a novel Joint Contrastive Learning with Sensitivity (JCLS) in this paper, which consists of sensitivity-aware feature contrastive learning (SFCL) and class-wise probabilistic contrastive learning (CPCL). Different from the traditional contrastive learning, SFCL pays more attention to the sensitive samples during optimizing the feature extractor, and consequently the feature discrimination of unlabeled samples can be enhanced. CPCL performs class-wise contrastive learning in the probabilistic space to enforce the cross-domain classifier to match the real distribution of source and target samples. By combining these two components, our JCLS is able to extract domain-invariant and compact features and obtain a well-performed classifier. We conduct the experiments on the DomainNet and Office-Home benchmarks, and the results show that our approach achieves state-of-the-art performance.
Keyu Tu, Zilei Wang, Junjie Li 0002, Yixin Zhang 0007
ACM Multimedia2
2023 Few-shot learning with unsupervised part discovery and part-aligned similarity
Zhang Zhang 0001, Wei Wang 0025, Liang Wang 0001, Zilei Wang, Tieniu Tan
Pattern Recognit.5
2023 Improve Temporal Action Proposals using Hierarchical Context
Qinying Liu, Zilei Wang, Shenghai Rong
Pattern Recognit.2
2023 Learning complementary semantic information for zero-shot recognition
Zilei Wang, Junjie Li 0002
Signal Process. Image Commun.2
2022 Semi-Supervised Video Semantic Segmentation with Inter-Frame Feature Reconstruction
abstract
One major challenge for semantic segmentation in realworld scenarios is only limited pixel-level labels available due to high expense of human labor though a vast volume of video data is provided. Existing semi-supervised methods attempt to exploit unlabeled data in model training, but they just regard video as a set of independent images. To better explore semi-supervised segmentation problem with video data, we formulate a semi-supervised video semantic segmentation task in this paper. For this task, we observe that the overfitting is surprisingly severe between labeled and unlabeled frames within a training video although they are very similar in style and contents. This is called inner-video overfitting, and it would actually lead to inferior performance. To tackle this issue, we propose a novel interframe feature reconstruction (IFR) technique to leverage the ground-truth labels to supervise the model training on unlabeled frames. IFR is essentially to utilize the internal relevance of different frames within a video. During training, IFR would enforce the feature distributions between labeled and unlabeled frames to be narrowed. Consequently, the inner-video overfitting issue can be effectively alleviated. We conduct extensive experiments on Cityscapes and CamVid, and the results demonstrate the superiority of our proposed method to previous state-of-the-art methods. The code is available at https://github.com/jfzhuang/IFR.
Jiafan Zhuang, Zilei Wang
CVPR2
2022 Cross-Domain Cross-Set Few-Shot Learning via Learning Compact and Aligned Representations
Zhang Zhang 0001, Wei Wang 0115, Liang Wang 0001, Zilei Wang, Tieniu Tan
ECCV (20)5
2022 Continual Semantic Segmentation via Structure Preserving and Projected Feature Alignment
Zilei Wang, Yixin Zhang 0007
ECCV (29)2
2022 Collaborating Domain-Shared and Target-Specific Feature Clustering for Cross-domain 3D Action Recognition
Qinying Liu, Zilei Wang
ECCV (4)2
2022 Disentangled Federated Learning for Tackling Attributes Skew via Invariant Aggregation and Diversity Transferring
abstract
Attributes skew hinders the current federated learning (FL) frameworks from consistent optimization directions among the clients, which inevitably leads to performance reduction and unstable convergence. The core problems lie in that: 1) Domain-specific attributes, which are non-causal and only locally valid, are indeliberately mixed into global aggregation. 2) The one-stage optimizations of entangled attributes cannot simultaneously satisfy two conflicting objectives, i.e., generalization and personalization. To cope with these, we proposed disentangled federated learning (DFL) to disentangle the domain-specific and cross-invariant attributes into two complementary branches, which are trained by the proposed alternating local-global optimization independently. Importantly, convergence analysis proves that the FL system can be stably converged even if incomplete client models participate in the global aggregation, which greatly expands the application scope of FL. Extensive experiments verify that DFL facilitates FL with higher performance, better interpretability, and faster convergence rate, compared with SOTA FL methods on both manually synthesized and realistic attributes skew datasets.
Zhengquan Luo, Yunlong Wang 0003, Zilei Wang, Zhenan Sun, Tieniu Tan
ICML3
2022 Exploring High-quality Target Domain Information for Unsupervised Domain Adaptive Semantic Segmentation
abstract
In unsupervised domain adaptive (UDA) semantic segmentation, the distillation based methods are currently dominant in performance. However, the distillation technique requires complicate multi-stage process and many training tricks. In this paper, we propose a simple yet effective method that can achieve competitive performance to the advanced distillation methods. Our core idea is to fully explore the target-domain information from the views of boundaries and features. First, we propose a novel mix-up strategy to generate high-quality target-domain boundaries with ground-truth labels. Different from the source-domain boundaries in previous works, we select the high-confidence target-domain areas and then paste them to the source-domain images. Such a strategy can generate the object boundaries in target domain (edge of target-domain object areas) with the correct labels. Consequently, the boundary information of target domain can be effectively captured by learning on the mixed-up samples. Second, we design a multi-level contrastive loss to improve the representation of target-domain data, including pixel-level and prototype-level contrastive learning. By combining two proposed methods, more discriminative features can be extracted and hard object boundaries can be better addressed for the target domain. The experimental results on two commonly adopted benchmarks (i.e., GTA5 -> Cityscapes and SYNTHIA -> Cityscapes) show that our method achieves competitive performance to complicated distillation methods. Notably, for the SYNTHIA-> Cityscapes scenario, our method achieves the state-of-the-art performance with 57.8% mIoU and 64.6% mIoU on 16 classes and 13 classes. Code is available at https://github.com/ljjcoder/EHTDI.
Junjie Li 0002, Zilei Wang
ACM Multimedia2
2022 Context-Aware Dynamic Feature Extraction for 3D Object Detection in Point Clouds
abstract
Varying density of point clouds increases the difficulty of 3D detection. In this paper, we present a context-aware dynamic network (CADNet) to capture the variance of density by considering both point context and semantic context. Point-level contexts are generated from original point clouds to enlarge the effective receptive filed. They are extracted around the voxelized pillars based on our extended voxelization method and processed with the context encoder in parallel with the pillar features. With a large perception range, we are able to capture the variance of features for potential objects and generate attentive spatial guidance to help adjust the strengths for different regions. In the region proposal network, considering the limited representation ability of traditional convolution where same kernels are shared among different samples and positions, we propose a decomposable dynamic convolutional layer to adapt to the variance of input features by learning from the local semantic context. It adaptively generates the position-dependent coefficients for multiple fixed kernels and combines them to convolve with local features. Based on our dynamic convolution, we design a dual-path convolution block to further improve the representation ability. We conduct experiments on KITTI dataset and the proposed CADNet has achieved superior performance of 3D detection outperforming SECOND and PointPillars by a large margin at the speed of 30 FPS.
Yonglin Tian, Lichao Huang, Hui Yu 0001, Xiangbin Wu, Xuesong Li 0004, Kunfeng Wang, Zilei Wang, Fei-Yue Wang 0001
IEEE Trans. Intell. Transp. Syst.7
2021 Learning Intact Features by Erasing-Inpainting for Few-shot Classification
abstract
Few-shot classification aims to categorize the samples from unseen classes with only few labeled samples. To address such a challenge, many methods exploit a base set consisting of massive labeled samples to learn an instance embedding function, i.e., image feature extractor, and it is expected to possess good transferability among different tasks. Such characteristics of few-shot learning are essentially different from that of traditional image classification only pursuing to get discriminative image representations. In this paper, we propose to learn intact features by erasing-inpainting for few-shot classification. Specifically, we argue that extracting intact features of target objects is more transferable, and then propose a novel cross-set erasing-inpainting (CSEI) method. CSEI processes the images in the support set using erasing and inpainting, and then uses them to augment the query set of the same task. Consequently, the feature embedding produced by our proposed method can contain more complete information of target objects. In addition, we propose task-specific feature modulation to make the features adaptive to the current task. The extensive experiments on two widely used benchmarks well demonstrates the effectiveness of our proposed method, which can consistently get considerable performance gains for different baseline methods.
Junjie Li 0002, Zilei Wang
AAAI2
2021 Efficient License Plate Recognition via Holistic Position Attention
abstract
License plate recognition (LPR) is a fundamental component of various intelligent transportation systems, and is always expected to be accurate and efficient enough in real-world applications. Nowadays, recognition of single character has been sophisticated benefiting from the power of deep learning, and extracting position information for forming a character sequence becomes the main bottleneck of LPR. To tackle this issue, we propose a novel holistic position attention (HPA) in this paper that consists of position network and shared classifier. Specifically, the position network explicitly encodes the character position into the maps of HPA, and then the shared classifier performs the character recognition in a unified and parallel way. Here the extracted features are modulated by the attention maps before feeding into the classifier to yield the final recognition results. Note that our proposed method is end-to-end trainable, character recognition can be concurrently performed, and no post-processing is needed. Thus our LPR system can achieve good effectiveness and efficiency simultaneously. The experimental results on four public datasets, including AOLP, Media Lab, CCPD, and CLPD, well demonstrate the superiority of our method to previous state-of-the-art methods in both accuracy and speed.
Yesheng Zhang, Zilei Wang, Jiafan Zhuang
AAAI2
2021 RPN Prototype Alignment for Domain Adaptive Object Detector
abstract
Recent years have witnessed great progress of object detection. However, due to the domain shift problem, applying the knowledge of an object detector learned from one specific domain to another one often suffers severe performance degradation. Most existing methods adopt feature alignment either on the backbone network or instance classifier to increase the transferability of object detector. Differently, we propose to perform feature alignment in the RPN stage such that the foreground and background RPN proposals in target domain can be effectively distinguished. Specifically, we first construct one set of learnable RPN prototpyes, and then enforce the RPN features to align with the prototypes for both source and target domains. It essentially cooperates the learning of RPN prototypes and features to align the source and target RPN features. Particularly, we propose a simple yet effective method suitable for RPN feature alignment to generate high-quality pseudo label of proposals in target domain, i.e., using the filtered detection results with IoU. Furthermore, we adopt Grad CAM to find the discriminative region within a foreground proposal and use it to increase the discriminability of RPN features for alignment. We conduct extensive experiments on multiple cross-domain detection scenarios, and the results show the effectiveness of our proposed method against previous state-of-the-art methods.
Zilei Wang, Yushi Mao
CVPR2
2021 Few-Shot Learning with Part Discovery and Augmentation from Unlabeled Images
abstract
Few-shot learning is a challenging task since only few instances are given for recognizing an unseen class. One way to alleviate this problem is to acquire a strong inductive bias via meta-learning on similar tasks. In this paper, we show that such inductive bias can be learned from a flat collection of unlabeled images, and instantiated as transferable representations among seen and unseen classes. Specifically, we propose a novel part-based self-supervised representation learning scheme to learn transferable representations by maximizing the similarity of an image to its discriminative part. To mitigate the overfitting in few-shot classification caused by data scarcity, we further propose a part augmentation strategy by retrieving extra images from a base dataset. We conduct systematic studies on miniImageNet and tieredImageNet benchmarks. Remarkably, our method yields impressive results, outperforming the previous best unsupervised methods by 7.74% and 9.24% under 5-way 1-shot and 5-way 5-shot settings, which are comparable with state-of-the-art supervised methods.
Chenyang Si, Wei Wang 0115, Liang Wang 0001, Zilei Wang, Tieniu Tan
IJCAI5
2021 Multi-level Discriminator and Wavelet Loss for Image Inpainting with Large Missing Area
Junjie Li 0002, Zilei Wang
PRCV (3)2
2021 Separated smooth sampling for fine-grained image classification
Shenghai Rong, Zilei Wang, Jie Wang 0005
Neurocomputing2
2021 Video Semantic Segmentation With Distortion-Aware Feature Correction
abstract
Video semantic segmentation aims to generate an accurate semantic map for each frame in a video. For such a task, conducting per-frame image segmentation is generally unacceptable in practice due to high computation cost. To address this issue, many works perform the flow-based feature propagation to reuse the features of previous frames, which essentially exploits the content continuity of consecutive frames. However, the estimated optical flow would inevitably suffer inaccuracy and then make the propagated features distorted. In this article, we propose a distortion-aware feature correction method with the goal of improving video segmentation performance at a low price. Our core idea is to correct the features on distorted regions using the current frame while reserving the propagated features for other regions. In this way, a lightweight network is enough for achieving promising segmentation results. In particular, we propose to predict the distorted regions by utilizing the consistency of distortion patterns in images and features, such that the high-cost feature extraction from current frames can be avoided. We conduct extensive experiments on Cityscapes, CamVid, and UAVid, and the results show that our proposed method significantly outperforms previous methods and achieves the state-of-the-art performance on both segmentation accuracy and speed. Code and pretrained models are available at https://github.com/jfzhuang/DAVSS.
Jiafan Zhuang, Zilei Wang, Bingke Wang
IEEE Trans. Circuits Syst. Video Technol.2
2021 Meta-USR: A Unified Super-Resolution Network for Multiple Degradation Parameters
abstract
Recent research on single image super-resolution (SISR) has achieved great success due to the development of deep convolutional neural networks. However, most existing SISR methods merely focus on super-resolution of a single fixed integer scale factor. This simplified assumption does not meet the complex conditions for real-world images which often suffer from various blur kernels or various levels of noise. More importantly, previous methods lack the ability to cope with arbitrary degradation parameters (scale factors, blur kernels, and noise levels) with a single model. A few methods can handle multiple degradation factors, e.g., noninteger scale factors, blurring, and noise, simultaneously within a single SISR model. In this work, we propose a simple yet powerful method termed meta-USR which is the first unified super-resolution network for arbitrary degradation parameters with meta-learning. In Meta-USR, a meta-restoration module (MRM) is proposed to enhance the traditional upscale module with the capability to adaptively predict the weights of the convolution filters for various combinations of degradation parameters. Thus, the MRM can not only upscale the feature maps with arbitrary scale factors but also restore the SR image with different blur kernels and noise levels. Moreover, the lightweight MRM can be placed at the end of the network, which makes it very efficient for iteratively/repeatedly searching the various degradation factors. We evaluate the proposed method through extensive experiments on several widely used benchmark data sets on SISR. The qualitative and quantitative experimental results show the superiority of our Meta-USR.
Xuecai Hu, Zhang Zhang 0001, Caifeng Shan, Zilei Wang, Liang Wang 0001, Tieniu Tan
IEEE Trans. Neural Networks Learn. Syst.4
2020 Progressive Boundary Refinement Network for Temporal Action Detection
abstract
Temporal action detection is a challenging task due to vagueness of action boundaries. To tackle this issue, we propose an end-to-end progressive boundary refinement network (PBRNet) in this paper. PBRNet belongs to the family of one-stage detectors and is equipped with three cascaded detection modules for localizing action boundary more and more precisely. Specifically, PBRNet mainly consists of coarse pyramidal detection, refined pyramidal detection, and fine-grained detection. The first two modules build two feature pyramids to perform the anchor-based detection, and the third one explores the frame-level features to refine the boundaries of each action instance. In the fined-grained detection module, three frame-level classification branches are proposed to augment the frame-level features and update the confidence scores of action instances. Evidently, PBRNet integrates the anchor-based and frame-level methods. We experimentally evaluate the proposed PBRNet and comprehensively investigate the effect of the main components. The results show PBRNet achieves the state-of-the-art detection performances on two popular benchmarks: THUMOS'14 and ActivityNet, and meanwhile possesses a high inference speed.
Qinying Liu, Zilei Wang
AAAI2
2020 Joint Adversarial Learning for Domain Adaptation in Semantic Segmentation
abstract
Unsupervised domain adaptation in semantic segmentation is to exploit the pixel-level annotated samples in the source domain to aid the segmentation of unlabeled samples in the target domain. For such a task, the key point is to learn domain-invariant representations and adversarial learning is usually used, in which the discriminator is to distinguish which domain the input comes from, and the segmentation model targets to deceive the domain discriminator. In this work, we first propose a novel joint adversarial learning (JAL) to boost the domain discriminator in output space by introducing the information of domain discriminator from low-level features. Consequently, the training of the high-level decoder would be enhanced. Then we propose a weight transfer module (WTM) to alleviate the inherent bias of the trained decoder towards source domain. Specifically, WTM changes the original decoder into a new decoder, which is learned only under the supervision of adversarial loss and thus mainly focuses on reducing domain divergence. The extensive experiments on two widely used benchmarks show that our method can bring considerable performance improvement over different baseline methods, which well demonstrates the effectiveness of our method in the output space adaptation.
Zilei Wang
AAAI2
2020 Polynomial Regression Network for Variable-Number Lane Detection
Bingke Wang, Zilei Wang, Yixin Zhang 0007
ECCV (18)2
2020 Image Inpainting with Contrastive Relation Network
abstract
Image inpainting faces the challenging issue of the requirements on structure reasonableness and texture coherence. In this paper, we propose a two-stage inpainting framework to address this issue. The basic idea is to address the two requirements in two separate stages. Completed segmentation of the corrupted image is firstly predicted through segmentation reconstruction network, while fine-grained image details are restored in the second stage through an image generator. The two stages are connected in series as the image details are generated under the guidance of completed segmentation map that predicted in the first stage. Specifically, in the second stage, we propose a novel graph-based relation network to model the relationship existed in corrupted image. In relation network, both intra-relationship for pixels in the same semantic region and inter-relationship between different semantic parts are considered, improving the consistency and compatibility of image textures. Besides, contrastive loss is designed to facilitate the relation network training. Such a framework not only simplifies the inpainting problem directly, but also exploits the relationship in corrupted image explicitly. Extensive experiments on various public datasets quantitatively and qualitatively demonstrate the superiority of our approach compared with the state-of-the-art.
Xiaoqiang Zhou, Junjie Li 0002, Zilei Wang, Ran He 0001, Tieniu Tan
ICPR3
2020 Adaptive and azimuth-aware fusion network of multimodal local features for 3D object detection
Yonglin Tian, Kunfeng Wang, Yuang Wang, Zilei Wang, Fei-Yue Wang 0001
Neurocomputing5
2019 Weighted Channel Dropout for Regularization of Deep Convolutional Neural Network
abstract
In this work, we propose a novel method named Weighted Channel Dropout (WCD) for the regularization of deep Convolutional Neural Network (CNN). Different from Dropout which randomly selects the neurons to set to zero in the fully-connected layers, WCD operates on the channels in the stack of convolutional layers. Specifically, WCD consists of two steps, i.e., Rating Channels and Selecting Channels, and three modules, i.e., Global Average Pooling, Weighted Random Selection and Random Number Generator. It filters the channels according to their activation status and can be plugged into any two consecutive layers, which unifies the original Dropout and Channel-Wise Dropout. WCD is totally parameter-free and deployed only in training phase with very slight computation cost. The network in test phase remains unchanged and thus the inference cost is not added at all. Besides, when combining with the existing networks, it requires no re-pretraining on ImageNet and thus is well-suited for the application on small datasets. Finally, WCD with VGGNet-16, ResNet-101, Inception-V3 are experimentally evaluated on multiple datasets. The extensive results demonstrate that WCD can bring consistent improvements over the baselines.
Saihui Hou, Zilei Wang
AAAI2
2019 Learning a Unified Classifier Incrementally via Rebalancing
abstract
Conventionally, deep neural networks are trained offline, relying on a large dataset prepared in advance. This paradigm is often challenged in real-world applications, e.g. online services that involve continuous streams of incoming data. Recently, incremental learning receives increasing attention, and is considered as a promising solution to the practical challenges mentioned above. However, it has been observed that incremental learning is subject to a fundamental difficulty -- catastrophic forgetting, namely adapting a model to new data often results in severe performance degradation on previous tasks or classes. Our study reveals that the imbalance between previous and new data is a crucial cause to this problem. In this work, we develop a new framework for incrementally learning a unified classifier, e.g. a classifier that treats both old and new classes uniformly. Specifically, we incorporate three components, cosine normalization, less-forget constraint, and inter-class separation, to mitigate the adverse effects of the imbalance. Experiments show that the proposed method can effectively rebalance the training process, thus obtaining superior performance compared to the existing methods. On CIFAR-100 and ImageNet, our method can reduce the classification errors by more than 6% and 13% respectively, under the incremental setting of 10 phases.
Saihui Hou, Chen Change Loy, Zilei Wang, Dahua Lin
CVPR4
2019 Meta-SR: A Magnification-Arbitrary Network for Super-Resolution
abstract
Recent research on super-resolution has achieved great success due to the development of deep convolutional neural networks (DCNNs). However, super-resolution of arbitrary scale factor has been ignored for a long time. Most previous researchers regard super-resolution of differentscale factors as independent tasks. They train a specific model for each scale factor which is inefficient in computing, and prior work only take the super-resolution of several integer scale factors into consideration. In this work,we propose a novel method called Meta-SR to firstly solve super-resolution of arbitrary scale factor (including non-integer scale factors) with a single model. In our Meta-SR,the Meta-Upscale Module is proposed to replace the traditional upscale module. For arbitrary scale factor, the Meta-Upscale Module dynamically predicts the weights of the up-scale filters by taking the scale factor as input and use these weights to generate the HR image of arbitrary size. For any low-resolution image, our Meta-SR can continuously zoomin it with arbitrary scale factor by only using a single model.We evaluated the proposed method through extensive experiments on widely used benchmark datasets on single image super-resolution. The experimental results show the superiority of our Meta-Upscale.
Xuecai Hu, Haoyuan Mu, Xiangyu Zhang 0005, Zilei Wang, Tieniu Tan, Jian Sun 0001
CVPR4
2019 Densely Supervised Hierarchical Policy-Value Network for Image Paragraph Generation
abstract
Image paragraph generation aims to describe an image with a paragraph in natural language. Compared to image captioning with a single sentence, paragraph generation provides more expressive and fine-grained description for storytelling. Existing approaches mainly optimize paragraph generator towards minimizing word-wise cross entropy loss, which neglects linguistic hierarchy of paragraph and results in ``sparse" supervision for generator learning. In this paper, we propose a novel Densely Supervised Hierarchical Policy-Value (DHPV) network for effective paragraph generation. We design new hierarchical supervisions consisting of hierarchical rewards and values at both sentence and word levels. The joint exploration of hierarchical rewards and values provides dense supervision cues for learning effective paragraph generator. We propose a new hierarchical policy-value architecture which exploits compositionality at token-to-token and sentence-to-sentence levels simultaneously and can preserve the semantic and syntactic constituent integrity. Extensive experiments on the Stanford image-paragraph benchmark have demonstrated the effectiveness of the proposed DHPV approach with performance improvements over multiple state-of-the-art methods.
Siying Wu, Zhengjun Zha, Zilei Wang, Houqiang Li, Feng Wu 0001
IJCAI3
2019 Feedback Convolutional Neural Network for Visual Localization and Segmentation
abstract
Feedback is a fundamental mechanism existing in the human visual system, but has not been explored deeply in designing computer vision algorithms. In this paper, we claim that feedback plays a critical role in understanding convolutional neural networks (CNNs), e.g., how a neuron in CNNs describes an object's pattern, and how a collection of neurons form comprehensive perception to an object. To model the feedback in CNNs, we propose a novel model named Feedback CNN and develop two new processing algorithms, i.e., neural pathway pruning and pattern recovering. We mathematically prove that the proposed method can reach local optimum. Note that Feedback CNN belongs to weakly supervised methods and can be trained only using category-level labels. But it possesses a powerful capability to accurately localize and segment category-specific objects. We conduct extensive visualization analysis, and the results reveal the close relationship between neurons and object parts in Feedback CNN. Finally, we evaluate the proposed Feedback CNN over the tasks of weakly supervised object localization and segmentation, and the experimental results on ImageNet and Pascal VOC show that our method remarkably outperforms the state-of-the-art ones.
Chunshui Cao, Yongzhen Huang, Yi Yang 0007, Liang Wang 0001, Zilei Wang, Tieniu Tan
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 Compressed-Domain Highway Vehicle Counting by Spatial and Temporal Regression
abstract
Counting on-road vehicles in the highway is fundamental for intelligent transportation management. This paper presents the first highway vehicle counting method in compressed domain, aiming at achieving comparable estimation performance with the pixel-domain methods. Counting in compressed domain is rather challenging due to limited information about vehicles and large variance in vehicle numbers. To address this problem, we develop new low-level features to mitigate the challenge from insufficient information in compressed videos. The new proposed features can be easily extracted from the coding-related metadata. Then, we propose a hierarchical classification-based regression (HCR) model to estimate the number of vehicles from the compressed-domain low-level features for individual frame. HCR hierarchically divides the traffic scenes into different cases according to the density of vehicles such that the large variance of traffic scenes can be effectively captured. Beside the spatial regression in each frame, we propose a locally temporal regression model to further refine the counting results, which exploits the continuous variation characteristics of the traffic flow. We extensively evaluate the proposed method on real highway surveillance videos. The experimental results consistently show that the proposed method is very competitive compared with the pixel-domain methods, which can reach similar performance with much lower computational cost.
Zilei Wang, Xu Liu 0008, Jiashi Feng, Jian Yang 0014, Hongsheng Xi
IEEE Trans. Circuits Syst. Video Technol.1
2019 Software-Defined Multimedia Streaming System Aided By Variable-Length Interval In-Network Caching
abstract
Explosive growth in video traffic volumes incurs a high percentage of redundancy in today's Internet, following the 80–20 rule. Fortunately, the advanced in-network cache is considered as an effective scheme for eliminating the repetitive traffic by caching the popular content in network nodes. Besides, the emerging software-defined networking (SDN) enables centralized control and management, as well as the collaboration between network devices and upper applications. Moreover, the Network Functions Virtualization is also developed to support for customized network functions, including caching and streaming. This inspires us to design an SDN-assisted multimedia streaming Video-on-Demand system, integrating in-network cache, to improve the quality of service. The designed architecture is capable of reducing the redundant traffic via the reusable duplications. In particular, it can achieve greater performance gains by deploying specific scheduling policy. We further propose a variable-length interval cache strategy for RTP streaming, which can realize the self-adaptive adjustment of the size of cached video segments based on their access patterns. Our goal is to efficiently utilize the limited storage resources and increase the cache hit ratio. We present the theoretical analysis to demonstrate the attainable performance of the proposed algorithm; furthermore, the integrated system design is implemented as a prototype to show its feasibility and applicability. Ultimately, emulation experiments are conducted to evaluate the achievable performance improvement more comprehensively.
Jian Yang 0014, Zhen Yao 0003, Xiaobin Tan, Zilei Wang, Quan Zheng 0002
IEEE Trans. Multim.5
2019 Dense 3D-Convolutional Neural Network for Person Re-Identification in Videos
abstract
Person re-identification aims at identifying a certain pedestrian across non-overlapping multi-camera networks in different time and places. Existing person re-identification approaches mainly focus on matching pedestrians on images; however, little attention has been paid to re-identify pedestrians in videos. Compared to images, video clips contain motion patterns of pedestrians, which is crucial to person re-identification. Moreover, consecutive video frames present pedestrian appearance with different body poses and from different viewpoints, providing valuable information toward addressing the challenge of pose variation, occlusion, and viewpoint change, and so on. In this article, we propose a Dense 3D-Convolutional Network (D3DNet) to jointly learn spatio-temporal and appearance representation for person re-identification in videos. The D3DNet consists of multiple three-dimensional (3D) dense blocks and transition layers. The 3D dense blocks enlarge the receptive fields of visual neurons in both spatial and temporal dimensions, leading to discriminative appearance representation as well as short-term and long-term motion patterns of pedestrians without the requirement of an additional motion estimation module. Moreover, we formulate a loss function consisting of an identification loss and a center loss to minimize intra-class variance and maximize inter-class variance simultaneously, toward addressing the challenge of large intra-class variance and small inter-class variance. Extensive experiments on two real-world video datasets of person identification, i.e., MARS and iLIDS-VID, have shown the effectiveness of the proposed approach.
Jiawei Liu 0001, Zhengjun Zha, Xuejin Chen, Zilei Wang, Yongdong Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2018 Lateral Inhibition-Inspired Convolutional Neural Network for Visual Attention and Saliency Detection
abstract
Lateral inhibition in top-down feedback is widely existing in visual neurobiology, but such an important mechanism has not be well explored yet in computer vision. In our recent research, we find that modeling lateral inhibition in convolutional neural network (LICNN) is very useful for visual attention and saliency detection. In this paper, we propose to formulate lateral inhibition inspired by the related studies from neurobiology, and embed it into the top-down gradient computation of a general CNN for classification, i.e. only category-level information is used. After this operation (only conducted once), the network has the ability to generate accurate category-specific attention maps. Further, we apply LICNN for weakly-supervised salient object detection.Extensive experimental studies on a set of databases, e.g., ECSSD, HKU-IS, PASCAL-S and DUT-OMRON, demonstrate the great advantage of LICNN which achieves the state-of-the-art performance. It is especially impressive that LICNN with only category-level supervised information even outperforms some recent methods with segmentation-level supervised learning.
Chunshui Cao, Yongzhen Huang, Zilei Wang, Liang Wang 0001, Ninglong Xu, Tieniu Tan
AAAI3
2018 SMC: Single-Stage Multi-location Convolutional Network for Temporal Action Detection
Zhikang Liu, Zilei Wang
ACCV (2)2
2018 CCNet: Cluster-Coordinated Net for Learning Multi-agent Communication Protocols with Reinforcement Learning
abstract
Multi-agent system is crucial for many practical applications. Recent years have witnessed numerous research on multi-agent task with reinforcement learning (RL) algorithms. Traditional reinforcement learning algorithms often fail to learn the cooperation between different agents, which is vital for multi-agent problems. A promising solution is to establish a communication protocol among agents. However, existing approaches often suffer from generalization challenges especially in tasks with partial observation and dynamic variation of agent amount. In this paper, we develop a Cluster-Coordinated Network (CCNet) to address the “Learning-to-communicate” problem in multi-agent system by utilizing the combination of a trainable Vector of Locally Aggregated Descriptor (VLAD) algorithm and reinforcement learning. Embedding with a VLAD based end-to-end trainable communication information processing module (called VLAD Processing Core), CCNet can learn efficient communication protocols even from scratch under partially observable environments and possesses robustness to the dynamic changes of agent number as well. Moreover, with the help of communication, CCNet is with less non-stationarity when training the network by common RL algorithms. We evaluated the proposed CCNet on two multi-agent partially observable tasks, \emph{i.e.}, Traffic Junction and Combat Task. The experimental results have demonstrated that CCNet is effective and improves the performance by a large margin over the state-of-the-art methods.
Zhengjun Zha, Zilei Wang, Liansheng Zhuang, Houqiang Li
ACML3
2018 Lifelong Learning via Progressive Distillation and Retrospection
Saihui Hou, Chen Change Loy, Zilei Wang, Dahua Lin
ECCV (3)4
2018 End-to-End View Synthesis for Light Field Imaging with Pseudo 4DCNN
Yunlong Wang 0003, Fei Liu 0031, Zilei Wang, Guangqi Hou, Zhenan Sun, Tieniu Tan
ECCV (2)3
2018 Towards Human-Level License Plate Recognition
Jiafan Zhuang, Saihui Hou, Zilei Wang, Zhengjun Zha
ECCV (3)3
2018 Object detection via deeply exploiting depth information
Saihui Hou, Zilei Wang, Feng Wu 0001
Neurocomputing2
2018 Dynamic Resource Allocation and Layer Selection for Scalable Video Streaming in Femtocell Networks: A Twin-Time-Scale Approach
abstract
Scalable video streaming over femtocell networks relying on two-tier spectrum-sharing is designed for coping with time-varying channel conditions, stringent video QoS requirements as well as with strong cross-tier interference between the over-sailing macro- and the femtocells. Dynamic video layer selection and resource allocation are invoked to enable the adaptation of the scalable video streaming service to the dynamics of both channel quality and interference price fluctuations. We formulate the design as a constrained stochastic optimization problem, which strikes a compelling compromise between the perceivable quality of experience and the monetary implications of the interference. Since the time scale of resource allocation is more short term than that of the video layer selection, we decompose the original long-term utility optimization problem into a pair of readily tractable subproblems with the aid of two different time-scales by invoking the powerful technique of Lyapunov drift and optimization. By exploiting the specific structure of these subproblems, low-complexity algorithms are derived for dynamic video layer selection and resource allocation, which rely on the near-instantaneously available information rather than on any prior statistical knowledge. Finally, we derive the analytical bounds of the theoretically achievable performance. Experimental results are presented for characterizing the performance attained.
Jian Yang 0014, Peng Si, Zilei Wang, Xiaofeng Jiang, Lajos Hanzo
IEEE Trans. Commun.3
2017 VegFru: A Domain-Specific Dataset for Fine-Grained Visual Categorization
abstract
In this paper, we propose a novel domain-specific dataset named VegFru for fine-grained visual categorization (FGVC). While the existing datasets for FGVC are mainly focused on animal breeds or man-made objects with limited labelled data, VegFru is a larger dataset consisting of vegetables and fruits which are closely associated with the daily life of everyone. Aiming at domestic cooking and food management, VegFru categorizes vegetables and fruits according to their eating characteristics, and each image contains at least one edible part of vegetables or fruits with the same cooking usage. Particularly, all the images are labelled hierarchically. The current version covers vegetables and fruits of 25 upper-level categories and 292 subordinate classes. And it contains more than 160,000 images in total and at least 200 images for each subordinate class. Accompanying the dataset, we also propose an effective framework called HybridNet to exploit the label hierarchy for FGVC. Specifically, multiple granularity features are first extracted by dealing with the hierarchical labels separately. And then they are fused through explicit operation, e.g., Compact Bilinear Pooling, to form a unified representation for the ultimate recognition. The experimental results on the novel VegFru, the public FGVC-Aircraft and CUB-200-2011 indicate that HybridNet achieves one of the top performance on these datasets. The dataset and code are available at https://github.com/ustc-vim/vegfru.
Saihui Hou, Yushan Feng, Zilei Wang
ICCV3
2017 DualNet: Learn Complementary Features for Image Recognition
abstract
In this work we propose a novel framework named Dual-Net aiming at learning more accurate representation for image recognition. Here two parallel neural networks are coordinated to learn complementary features and thus a wider network is constructed. Specifically, we logically divide an end-to-end deep convolutional neural network into two functional parts, i.e., feature extractor and image classifier. The extractors of two subnetworks are placed side by side, which exactly form the feature extractor of DualNet. Then the two-stream features are aggregated to the final classifier for overall classification, while two auxiliary classifiers are appended behind the feature extractor of each subnetwork to make the separately learned features discriminative alone. The complementary constraint is imposed by weighting the three classifiers, which is indeed the key of DualNet. The corresponding training strategy is also proposed, consisting of iterative training and joint fine tuning, to make the two subnetworks cooperate well with each other. Finally, DualNet based on the well-known CaffeNet, VGGNet, NIN and ResNet are thoroughly investigated and experimentally evaluated on multiple datasets including CIFAR-100, Stanford Dogs and UEC FOOD-100. The results demonstrate that DualNet can really help learn more accurate image representation, and thus result in higher accuracy for recognition. In particular, the performance on CIFAR-100 is state-of-the-art compared to the recent works.
Saihui Hou, Xu Liu 0008, Zilei Wang
ICCV3
2017 Improving human action recognitionby temporal attention
abstract
Recently, deep learning methods have been extensively applied for action recognition in videos. Most existing deep networks equally treat every video frame and directly assign a video label to all the frames sampled from it. However, discriminative action may occurs sparsely in a few key frames in a video, and other frames are less relevant or even irrelevant to the action class. Equally treating all the frames will hurt performance. To address this issue, we propose a temporal attention model which learns to recognize human actions in videos while focusing selectively on the informative frames. Our model does not need explicit annotations regarding such informative frames during training and testing. Specifically, we adopt Recurrent Neural Network (RNN) with Long Short-Term Memory (LSTM) unit and attaches higher importance to the frames which are discriminative for the task at hand. Our method consistently improves on no-attention methods, with both RGB and optical flow based deep ConvNets. We achieve state-of-the-art performance on two challenging datasets of UCF101 and HMDB51.
Zhikang Liu, Zilei Wang
ICIP3
2017 Action recognition with low observational latency via part movement model
Zhikang Liu, Zilei Wang
Multim. Tools Appl.2
2017 Salient object detection via saliency bias and diffusion
Dao Xiang, Zilei Wang
Multim. Tools Appl.2
2017 Background-Driven Salient Object Detection
abstract
The background information is a significant prior for salient object detection, especially when images contain cluttered background and diverse object parts. In this paper, we propose a background-driven salient object detection (BD-SOD) method to more comprehensively exploit the background prior, aiming at generating more accurate and robust salient maps. To be specific, we first exploit the background prior to conduct the saliency estimation, i.e., computing the regional saliency values. In this stage, the background prior is utilized in threefold: restricting the reference regions to only the background regions, weighting the contribution of reference regions, and leveraging the importance of different features. Benefiting from such an explicit utilization, the proposed model can greatly mitigate the negative interference of the cluttered background and diverse object parts. We then embed the background prior into the optimization graph for saliency refinement. Specifically, two virtual supernodes (representing the background and foreground, respectively) are introduced with extra connections, and the nonlocal feature connections between similar regions are also set up. These connections enhance the power of optimization graph to alleviate the perturbations from diverse parts, and thus help to achieve the uniformity of saliency values. Finally, we provide systematical studies to investigate the effectiveness of the proposed BD-SOD in exploiting the valuable background prior. Experimental results on multiple public benchmark datasets, including MSRA-1000, THUS-10000, PASCAL-S, and ECSSD, clearly show that BD-SOD consistently outperforms the well-established baselines and achieves state-of-the-art performance.
Zilei Wang, Dao Xiang, Saihui Hou, Feng Wu 0001
IEEE Trans. Multim.1
2016 Stacked Overcomplete Independent Component Analysis for Action Recognition
Zhikang Liu, Zilei Wang
ACCV (2)3
2016 Highway Vehicle Counting in Compressed Domain
abstract
This paper presents a highway vehicle counting method in compressed domain, aiming at achieving acceptable estimation performance approaching the pixel-domain methods. Such a task essentially is challenging because the available information (e.g. motion vector) to describe vehicles in videos is quite limited and inaccurate, and the vehicle count in realistic traffic scenes always varies greatly. To tackle this issue, we first develop a batch of low-level features, which can be extracted from the encoding metadata of videos, to mitigate the informational insufficiency of compressed videos. Then we propose a Hierarchical Classification based Regression (HCR) model to estimate the vehicle count from features. HCR hierarchically divides the traffic scenes into different cases according to vehicle density, such that the broad-variation characteristics of traffic scenes can be better approximated. Finally, we evaluated the proposed method on the real highway surveillance videos. The results show that our method is very competitive to the pixel-domain methods, which can reach similar performance along with its lower complexity.
Xu Liu 0008, Zilei Wang, Jiashi Feng, Hongsheng Xi
CVPR2
2016 A simple and robust super resolution method for light field images
abstract
Light field cameras generate low-resolution images due to the tradeoff between spatial and angular resolution. Traditional light field super-resolution (LFSR) methods depend on prior knowledge of depth information. This paper presents a projection-based LFSR solution without prior information based on redefinition of the mapping function between disparity and shearing shift. Moreover, simplified variational regularization is imposed in global optimization formulation to the rendered high-resolution images. Both a synthetic dataset and a real-world dataset of light field images captured by a self-developed light field camera are used to demonstrate the state-of-the-art performance of the proposed method.
Yunlong Wang 0003, Guangqi Hou, Zhenan Sun, Zilei Wang, Tieniu Tan
ICIP4
2015 Look and Think Twice: Capturing Top-Down Visual Attention with Feedback Convolutional Neural Networks
abstract
While feedforward deep convolutional neural networks (CNNs) have been a great success in computer vision, it is important to note that the human visual cortex generally contains more feedback than feedforward connections. In this paper, we will briefly introduce the background of feedbacks in the human visual cortex, which motivates us to develop a computational feedback mechanism in deep neural networks. In addition to the feedforward inference in traditional neural networks, a feedback loop is introduced to infer the activation status of hidden layer neurons according to the "goal" of the network, e.g., high-level semantic labels. We analogize this mechanism as "Look and Think Twice." The feedback networks help better visualize and understand how deep neural networks work, and capture visual attention on expected objects, even in images with cluttered background and multiple objects. Experiments on ImageNet dataset demonstrate its effectiveness in solving tasks such as image classification and object localization.
Chunshui Cao, Xianming Liu 0005, Yi Yang 0007, Yinan Yu, Jiang Wang 0001, Zilei Wang, Yongzhen Huang, Liang Wang 0001, Chang Huang, Wei Xu 0017, Deva Ramanan, Thomas S. Huang
ICCV6
2015 Collaborative Linear Coding for Robust Image Classification
Zilei Wang, Jiashi Feng, Shuicheng Yan
Int. J. Comput. Vis.1
2014 Autogrouped Sparse Representation for Visual Analysis
abstract
In image classification, recognition or retrieval systems, image contents are commonly described by global features. However, the global features generally contain noise from the background, occlusion, or irrelevant objects in the images. Thus, only part of the global feature elements is informative for describing the objects of interest and useful for the image analysis tasks. In this paper, we propose algorithms to automatically discover the subgroups of highly correlated feature elements within predefined global features. To this end, we first propose a novel mixture sparse regression (MSR) method, which groups the elements of a single vector according to the membership conveyed by their sparse regression coefficients. Based on MSR, we proceed to develop the autogrouped sparse representation (ASR), which groups correlated feature elements together through fusing their individual sparse representations over multiple samples. We apply ASR/MSR in two practical visual analysis tasks: 1) multilabel image classification and 2) motion segmentation. Comprehensive experimental evaluations show that our proposed methods are able to achieve superior performance compared with the state-of-the-art classification on these two tasks.
Jiashi Feng, Xiao-Tong Yuan, Zilei Wang, Huan Xu 0001, Shuicheng Yan
IEEE Trans. Image Process.3
2013 Multi-class learning from class proportions
Zilei Wang, Jiashi Feng
Neurocomputing1
2013 Linear Distance Coding for Image Classification
abstract
The feature coding-pooling framework is shown to perform well in image classification tasks, because it can generate discriminative and robust image representations. The unavoidable information loss incurred by feature quantization in the coding process and the undesired dependence of pooling on the image spatial layout, however, may severely limit the classification. In this paper, we propose a linear distance coding (LDC) method to capture the discriminative information lost in traditional coding methods while simultaneously alleviating the dependence of pooling on the image spatial layout. The core of the LDC lies in transforming local features of an image into more discriminative distance vectors, where the robust image-to-class distance is employed. These distance vectors are further encoded into sparse codes to capture the salient features of the image. The LDC is theoretically and experimentally shown to be complementary to the traditional coding methods, and thus their combination can achieve higher classification accuracy. We demonstrate the effectiveness of LDC on six data sets, two of each of three types (specific object, scene, and general object), i.e., Flower 102 and PFID 61, Scene 15 and Indoor 67, Caltech 101 and Caltech 256. The results show that our method generally outperforms the traditional coding methods, and achieves or is comparable to the state-of-the-art performance on these data sets.
Zilei Wang, Jiashi Feng, Shuicheng Yan, Hongsheng Xi
IEEE Trans. Image Process.1
2013 Image Classification via Object-Aware Holistic Superpixel Selection
abstract
In this paper, we propose an object-aware holistic superpixel selection (HPS) method to automatically select the discriminative superpixels of an image for image classification purpose. Through only considering the selected superpixels, the interference of cluttered background on the object can be alleviated effectively and thus the classification performance is significantly enhanced. In particular, for an image, HPS first selects the discriminative superpixels for the characteristics of certain class, which can together match the object template of this class well. In addition, these superpixels compose a class-specific matching region. Through performing such superpixel selection for several most probable classes, respectively, HPS generates multiple class-specific matching regions for a single image. Then, HPS merges these matching regions into an integral object region through exploiting their pixel-level intersection information. Finally, such object region instead of the original image is used for image classification. An appealing advantage of HPS is the ability to alleviate the interference of cluttered background yet not require the object to be segmented out accurately. We evaluate the proposed HPS on four challenging image classification benchmark datasets: Oxford-IIIT PET 37, Caltech-UCSD Birds 200, Caltech 101, and PASCAL VOC 2011. The experimental results consistently show that the proposed HPS can remarkably improve the classification performance.
Zilei Wang, Jiashi Feng, Shuicheng Yan, Hongsheng Xi
IEEE Trans. Image Process.1
2012 Auto-Grouped Sparse Representation for Visual Analysis
Jiashi Feng, Xiao-Tong Yuan, Zilei Wang, Huan Xu 0001, Shuicheng Yan
ECCV (1)3
2012 Purposive Hidden-Object-Game: Embedding Human Computation in Popular Game
abstract
Having sufficient training images with fully annotated object locations is undoubtedly critical for modern learning-based image annotation, retrieval, and object detection methods. Typically, collecting such annotations for large-scale datasets is notoriously tedious because the process involves amount of manual cropping and hand labeling operations. In this work, following the principle of games with a purpose (GWAP), we design a so-called purposive hidden-object-game (P-HOG), which imperceptibly embeds localizing objects into enjoyable playing game process and thus attracts many people to make voluntary contribution to annotating images. In particular, besides preserving the interestingness as popular HOG games, P-HOG is able to automatically generate satisfactory game images (i.e., “hide” certain items into target images) by integrating several semantic and visual processing techniques. P-HOG is also built in an effective mechanism to prevent the players from cheating. The mechanism inherits the merit of Recaptcha and identifies potential cheating behavior based on the annotation accuracy of some known items. Moreover, P-HOG will filter noisy annotations effectively based on a weighted majority method and improve the accuracy of the raw annotations from the players. Most importantly, players only play P-HOG for entertainment purpose and they are unaware of the background data collection procedure. The collected data are used towards constructing a large database, which may benefit general learning-based algorithms for multimedia tasks. To the best of our knowledge, this is the first work dedicated to such a specific and important task under the GWAP framework. We conduct a pilot study of the game prototype and the comprehensive experiments show that the P-HOG appeals to general players, and is effective for collecting massive object locations with satisfactory accuracy, which further boosts the algorithmic performances for both tag refinement and image annotation tasks.
Jiashi Feng, Yuzhao Ni, Jian Dong 0011, Zilei Wang, Shuicheng Yan
IEEE Trans. Multim.4
2010 A relaxing bandwidth smoothing schedule for transmitting prerecorded VBR video in periodic network
Zilei Wang, Hongsheng Xi
Multim. Syst.1
2009 Generalized PCRTT Offline Bandwidth Smoothing Based on SVM and Systematic Video Segmentation
abstract
As a trade-off technique, bandwidth smoothing can reduce the client buffer requirements and simultaneously keep transmission scheme as smooth as possible. In this paper, bandwidth smoothing is formulated into a binary classification problem of the underflow and overflow points. We propose a novel method to solve that problem based on support vector machine (SVM). Our method is proven to be able to achieve the minimum buffer requirements of constant rate transmission and transport. Furthermore, it directly computes the transmission rate without exhaustively searching buffer size and startup delay. Besides this method, this paper provides a systematic video segmentation algorithm, which can intelligently partition the playback curve into some unequal segments to naturally track the trends of playback curve. The smoothing results with the playback curve ofy=xndemonstrate that this video systematic segmentation requires smaller than half of the buffer of the equal segmentation algorithm. Finally, we construct a generalized piecewise constant rate transmission and transport algorithm with SVM and the systematic video segmentation method. The experiments of some real MPEG4 and H.264 video data confirmed the efficiency of our proposed algorithm.
Zilei Wang, Hongsheng Xi
IEEE Trans. Multim.1