EDBT 2026 Demo / reviewers in the wild / expert
Zhi Chen 0010
dblp:05/1539-10
· DBLP profile ↗
27ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0002-9385-144XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TR-DQ: Time-Rotation Diffusion QuantizationabstractDiffusion models have been widely adopted in image and video generation. However, their complex network architecture leads to high inference overhead for its generation process. Existing diffusion quantization methods primarily focus on the quantization of the model structure while ignoring the impact of time-steps variation during sampling. At the same time, most current approaches fail to account for significant activations that cannot be eliminated, resulting in substantial performance degradation after quantization. To address these issues, we propose Time-Rotation Diffusion Quantization (TR-DQ), a novel quantization method incorporating time-step and rotation-based optimization. TR-DQ first divides the sampling process based on time-steps and applies a rotation matrix to smooth activations and weights dynamically. For different time-steps, a dedicated hyperparameter is introduced for adaptive timing modeling, which enables dynamic quantization across different time steps. Additionally, we also explore the compression potential of Classifier-Free Guidance (CFG-wise) to establish a foundation for subsequent work. TR-DQ achieves state-of-the-art (SOTA) performance on image generation and video generation tasks and a 1.38-1.89× speedup and 1.97-2.58× memory reduction in inference compared to existing quantization methods. Yihua Shao, Deyang Lin, Minxi Yan, Siyu Chen 0021, Fanhu Zeng, Minwen Liao, Ao Ma 0005, Ziyang Yan, Haozhe Wang 0002, Yan Wang 0068, Zhi Chen 0010, Xiaofeng Cao 0002, Haotong Qin, Hao Tang 0005, Jingcai Guo |
AAAI | 11 |
| 2026 | Cluster-aware prompt ensemble learning for few-shot vision-language model adaptationabstract• Introduces CAPEL: ensembles multiple prompts in logit space for VLM adaptation. • Preserves multimodal class structure and supports fully batched inference. • Uses cluster-preserving entropy to stabilize diverse prompt subclassifiers. • Learns prompt-level attention and enables simple post-hoc pruning of heads. • Demonstrates generalization across backbones and tasks, including robustness and segmentation. Vision-language models (VLMs) such as CLIP achieve zero-shot transfer across various tasks by pre-training on numerous image-text pairs. These models often benefit from using an ensemble of context prompts to represent a class. Despite being effective, conventional prompt ensembling that averages textual features of context prompts often yields suboptimal results. This is because feature averaging shifts the class centroids away from the true class distribution. To address this issue, we propose the Cluster-Aware Prompt Ensemble Learning (CAPEL) framework, which preserves the cluster nature of context prompts. CAPEL classifies images into one of several class clusters, each represented by a distinct prompt. Instead of ensembling prompts in the feature space, we perform ensembling in the classification logits space, aligning better with the visual feature distribution. To further optimize prompt fine-tuning while maintaining cluster-specific discriminative power, we introduce a cluster-preserving regularization term. This ensures that prompts remain distinct and specialized for different clusters, preventing collapse into a uniform direction. Additionally, we integrate an adaptive prompt weighting technique to dynamically adjust the attention weights for flawed or ambiguous prompts, ensuring robust performance across diverse datasets and tasks. Zhi Chen 0010, Xin Yu 0002, Zi Huang |
Pattern Recognit. | 1 |
| 2026 | FastEdit: fast text-guided single-image editing via semantic-aware diffusion fine-tuningabstract• Semantic-aware diffusion: edits guided by image-text discrepancy for precise, controllable changes. • Fast and lightweight: 50-iteration fine-tune with LoRA ( 0.37% params), 17 s per image. • Better alignment without sacrificing fidelity: higher CLIP on TEdBench with competitive LPIPS. • Broadly applicable: runs on SD v1.4 and ports to SDXL (via IP-Adapter) with consistent gains. Text-guided single-image editing has emerged as a promising solution to precisely alter an input image based on the target texts, such as making a standing dog appear seated or a bird spreading its wings. While effective, conventional approaches require a two-step process, including fine-tuning the target text embedding for over 1K iterations and the generative model for another 1.5K iterations. Although it ensures that the resulting image closely aligns with both the input image and the target text, this process often requires 7 minutes per image, posing a challenge for practical application due to its time-intensive nature. To address this bottleneck, we introduce FastEdit, a fast text-guided single-image editing method with semantic-aware diffusion fine-tuning, accelerating the editing process to just 17 seconds. FastEdit streamlines the generative model’s fine-tuning phase, reducing it from 1.5K to 50 iterations. Specifically, we perform diffusion fine-tuning on certain time steps determined by the semantic discrepancy between the input image and target text. We conduct extensive experiments to validate the editing performance of our approach and show promising editing capabilities, including content addition, style transfer, background replacement, and posture manipulation, etc. Our code and more edited images are available at https://fastedit-sd.github.io . Zhi Chen 0010, Zecheng Zhao, Yadan Luo, Zi Huang |
Pattern Recognit. | 1 |
| 2025 | Dynamic Target Distribution Estimation for Source-Free Open-Set Domain AdaptationabstractUnsupervised domain adaptation (UDA) has emerged as a promising technique for transferring knowledge from a labeled domain to an unlabeled domain. However, existing UDA methods are severely constrained by data privacy and semantic inconsistencies. To alleviate these limitations, this work challenges the Source-Free Open-Set Domain Adaptation (SF-OSDA), where the pre-trained source model is directly leveraged on the open target domain for adaptation. For this purpose, we introduce the novel Dynamic Target Distribution Estimation (DTDE) method, which effectively performs known classification and unknown separation through self-supervised learning with prototypes. To construct known prototypes, a self-adaptive sampling strategy is employed to consider the category disparity. For unknown prototypes, we utilize a self-splitting and excluding principle to bypass the unknown semantics problem. Specifically, self-splitting is to evaluate the overall clustering distribution of the target domain. By excluding clusters resembling known prototypes, the remaining cluster centroids can serve as unknown prototypes. The superiority of our approach is validated across multiple benchmarks. Remarkably, DTDE outperforms the best competitor by 7.6% on the VisDA dataset. Zhiqi Yu, Zhichao Liao, Jingjing Li 0001, Zhi Chen 0010, Lei Zhu 0002 |
AAAI | 4 |
| 2025 | SVIP: Semantically Contextualized Visual Patches for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize unseen classes without labeled training examples by leveraging class-level semantic descriptors such as attributes. A fundamental challenge in ZSL is semantic misalignment, where semantic-unrelated information involved in visual features introduce ambiguity to visual-semantic interaction. Unlike existing methods that suppress semantic-unrelated information post hoc either in the feature space or the model space, we propose addressing this issue at the input stage, preventing semantic-unrelated patches from propagating through the network. To this end, we introduce Semantically contextualized VIsual Patches (SVIP) for ZSL, a transformer-based framework designed to enhance visual-semantic alignment. Specifically, we propose a self-supervised patch selection mechanism that preemptively learns to identify semantic-unrelated patches in the input space. This is trained with the supervision from aggregated attention scores across all transformer layers, which estimate each patch's semantic score. As removing semantic-unrelated patches from the input sequence may disrupt object structure, we replace them with learnable patch embeddings. With initialization from word embeddings, we can ensure they remain semantically meaningful throughout feature extraction. Extensive experiments on ZSL benchmarks demonstrate that SVIP achieves state-of-the-art performance results while providing more interpretable and semantically rich feature representations. Code is available at https://github.com/uqzhichen/SVIP. Zhi Chen 0010, Zecheng Zhao, Jingcai Guo, Jingjing Li 0001, Zi Huang |
ICCV | 1 |
| 2025 | On the Discrimination and Consistency for Exemplar-Free Class Incremental LearningabstractExemplar-free class incremental learning (EF-CIL) is a nontrivial task that requires continuously enriching model capability with new classes while maintaining previously learned knowledge without storing and replaying any old class exemplars. An emerging theory-guided framework for CIL trains task-specific models for a shared network, shifting the pressure of forgetting to task-id prediction. In EF-CIL, task-id prediction is more challenging due to the lack of inter-task interaction (e.g., replays of exemplars). To address this issue, we conduct a theoretical analysis of the importance and feasibility of preserving a discriminative and consistent feature space, upon which we propose a novel method termed DCNet. Concretely, it progressively maps class representations into a hyperspherical space, in which different classes are orthogonally distributed to achieve ample inter-class separation. Meanwhile, it also introduces compensatory training to adaptively adjust supervision intensity, thereby aligning the degree of intra-class aggregation. Extensive experiments and theoretical analysis verified the superiority of DCNet. Code is available at https://github.com/Tianqi-Wang1/DCNet. Jingcai Guo, Depeng Li 0001, Zhi Chen 0010 |
IJCAI | 4 |
| 2025 | PatAug: Augmentation of Augmentation for Test-Time AdaptationabstractThe rich pretrained knowledge in vision-language models (VLMs) endows them with the ability to discriminate common objects given only category names, but may be challenged by out-of-distribution unlabeled samples. To address this limitation, test-time adaptation (TTA) dynamically adjusts VLMs to target distributions during inference. Current TTA frameworks rely heavily on unsupervised data augmentations to enhance sample informativeness, but remain vulnerable to naive augmented views. This work introduces Patch Augmentation (PatAug), a pixel-level perturbation framework that optimizes the benefits of informative augmentations and mitigates negative transformation impacts. Implemented as trainable pixels, PatAug are prepared given only category names before inference, introducing few additional overheads. The patches encode class-related semantic information. They assist VLMs in emphasizing on the compatible visual information in the images, restoring perturbed image details, while retaining unrecognized information. Such merits inspire the design of an augmentation of augmentation framework, where PatAug is applied to standard augmentation views for reliable TTA inference results. To better fit the target distributions, we adjust patches with a cross-modal similarity alignment loss and learnable patching weights. Experiments on natural and specialized domain shifts confirm the effectiveness of PatAug. Zhekai Du, Lei Zhu 0002, Zhi Chen 0010, Jingjing Li 0001 |
ACM Multimedia | 5 |
| 2025 | Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic VideosabstractText-to-video (T2V) synthesis has advanced rapidly, yet current evaluation metrics primarily capture visual quality and temporal consistency, offering limited insight into how synthetic videos perform in downstream tasks such as text-to-video retrieval (TVR). In this work, we introduce SynTVA, a new dataset and benchmark designed to evaluate the utility of synthetic videos for building retrieval models. Based on 800 diverse user queries derived from MSRVTT training split, we generate synthetic videos using state-of-the-art T2V models and annotate each video-text pair along four key semantic alignment dimensions: Object & Scene, Action, Attribute, and Prompt Fidelity. Our evaluation framework correlates general video quality assessment (VQA) metrics with these alignment scores, and examines their predictive power for downstream TVR performance. To explore pathways of scaling up, we further develop an Auto-Evaluator to estimate alignment quality from existing metrics. Beyond benchmarking, our results show that SynTVA is a valuable asset for dataset augmentation, enabling the selection of high-utility synthetic samples that measurably improve TVR outcomes. Project page and dataset can be found at https://jasoncodemaker.github.io/SynTVA/. Zecheng Zhao, Selena Song, Tong Chen 0005, Zhi Chen 0010, Shazia Sadiq, Yadan Luo |
ACM Multimedia | 4 |
| 2025 | Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware RoutingabstractText-to-Video Retrieval (TVR) aims to retrieve relevant videos based on textual queries.However, as video content evolves continuously, adapting TVR systems to new data remains a critical yet underexplored challenge.In this paper, we introduce the first benchmark for Continual Text-to-Video Retrieval (CTVR) to address the limitations of existing approaches.Current Pre-Trained Model (PTM)based TVR methods struggle with maintaining model plasticity when adapting to new tasks, while existing Continual Learning (CL) methods suffer from catastrophic forgetting, leading to semantic misalignment between historical queries and stored video features.To address these two challenges, we propose FrameFu-sionMoE, a novel CTVR framework that comprises two key components: (1) the Frame Fusion Adapter (FFA), which captures temporal video dynamics while preserving model plasticity, and (2) the Task-Aware Mixture-of-Experts (TAME), which ensures consistent semantic alignment between queries across tasks and the stored video features.Thus, FrameFusionMoE enables effective adaptation to new video content while preserving historical textvideo relevance to mitigate catastrophic forgetting.We comprehensively evaluate FrameFusionMoE on two benchmark datasets under various task settings.Results demonstrate that FrameFusionMoE outperforms existing CL and TVR methods, achieving superior retrieval performance with minimal degradation on earlier tasks when handling continuous video streams.Our code is available at: https://github.com/JasonCodeMaker/CTVR. Zecheng Zhao, Zhi Chen 0010, Zi Huang, Shazia Sadiq, Tong Chen 0005 |
SIGIR | 2 |
| 2024 | Benchmarking In-the-Wild Multimodal Disease Recognition and A Versatile BaselineabstractExisting plant disease classification models have achieved remarkable performance in recognizing in-laboratory diseased images. However, their performance often significantly degrades in classifying in-the-wild images. Furthermore, we observed that in-the-wild plant images may exhibit similar appearances across various diseases (i.e., small inter-class discrepancy) while the same diseases may look quite different (i.e., large intra-class variance). Motivated by this observation, we propose an in-the-wild multimodal plant disease recognition dataset that contains the largest number of disease classes but also text-based descriptions for each disease. Particularly, the newly provided text descriptions are introduced to provide rich information in textual modality and facilitate in-the-wild disease classification with small inter-class discrepancy and large intra-class variance issues. Therefore, our proposed dataset can be regarded as an ideal testbed for evaluating disease recognition methods in the real world. In addition, we further present a strong yet versatile baseline that models text descriptions and visual data through multiple prototypes for a given class. By fusing the contributions of multimodal prototypes in classification, our baseline can effectively address the small inter-class discrepancy and large intra-class variance issues. Remarkably, our baseline model can not only classify diseases but also recognize diseases in few-shot or training-free scenarios. Extensive benchmarking results demonstrate that our proposed in-the-wild multimodal dataset sets many new challenges to the plant disease recognition task and there is a large space to improve for future works. Tianqi Wei 0002, Zhi Chen 0010, Zi Huang, Xin Yu 0002 |
ACM Multimedia | 2 |
| 2024 | Snap and Diagnose: An Advanced Multimodal Retrieval System for Identifying Plant Diseases in the Wild
Tianqi Wei 0002, Zhi Chen 0010, Xin Yu 0002 |
MMAsia | 2 |
| 2024 | DiPEx: Dispersing Prompt Expansion for Class-Agnostic Object DetectionabstractClass-agnostic object detection (OD) can be a cornerstone or a bottleneck for many downstream vision tasks. Despite considerable advancements in bottom-up and multi-object discovery methods that leverage basic visual cues to identify salient objects, consistently achieving a high recall rate remains difficult due to the diversity of object types and their contextual complexity. In this work, we investigate using vision-language models (VLMs) to enhance object detection via a self-supervised prompt learning strategy. Our initial findings indicate that manually crafted text queries often result in undetected objects, primarily because detection confidence diminishes when the query words exhibit semantic overlap. To address this, we propose a Dispersing Prompt Expansion (DiPEx) approach. DiPEx progressively learns to expand a set of distinct, non-overlapping hyperspherical prompts to enhance recall rates, thereby improving performance in downstream tasks such as out-of-distribution OD. Specifically, DiPEx initiates the process by self-training generic parent prompts and selecting the one with the highest semantic uncertainty for further expansion. The resulting child prompts are expected to inherit semantics from their parent prompts while capturing more fine-grained semantics. We apply dispersion losses to ensure high inter-class discrepancy among child prompts while preserving semantic consistency between parent-child prompt pairs. To prevent excessive growth of the prompt sets, we utilize the maximum angular coverage (MAC) of the semantic space as a criterion for early termination. We demonstrate the effectiveness of DiPEx through extensive class-agnostic OD and OOD-OD experiments on MS-COCO and LVIS, surpassing other prompting methods by up to 20.1% in AR and achieving a 21.3% AP improvement over SAM. Jia Syuen Lim, Zhuoxiao Chen, Zhi Chen 0010, Mahsa Baktash, Xin Yu 0002, Zi Huang, Yadan Luo |
NeurIPS | 3 |
| 2024 | Towards Cost-Efficient Federated Multi-agent RL with Learnable Aggregation
Yi Zhang 0105, Sen Wang 0001, Zhi Chen 0010, Xuwei Xu, Stanislav Funiak, Jiajun Liu 0004 |
PAKDD (2) | 3 |
| 2023 | Zero-Shot Learning by Harnessing Adversarial SamplesabstractZero-Shot Learning (ZSL) aims to recognize unseen classes by generalizing the knowledge, i.e., visual and semantic relationships, obtained from seen classes, where image augmentation techniques are commonly applied to improve the generalization ability of a model. However, this approach can also cause adverse effects on ZSL since the conventional augmentation techniques that solely depend on single-label supervision is not able to maintain semantic information and result in the semantic distortion issue consequently. In other words, image argumentation may falsify the semantic (e.g., attribute) information of an image. To take the advantage of image augmentations while mitigating the semantic distortion issue, we propose a novel ZSL approach by Harnessing Adversarial Samples (HAS). HAS advances ZSL through adversarial training which takes into account three crucial aspects: (1) robust generation by enforcing augmentations to be similar to negative classes, while maintaining correct labels, (2) reliable generation by introducing a latent space constraint to avert significant deviations from the original data manifold, and (3) diverse generation by incorporating attribute-based perturbation by adjusting images according to each semantic attribute's localization. Through comprehensive experiments on three prominent zero-shot benchmark datasets, we demonstrate the effectiveness of our adversarial samples approach in both ZSL and Generalized Zero-Shot Learning (GZSL) scenarios. Our source code is available at https://github.com/uqzhichen/HASZSL. Zhi Chen 0010, Peng-Fei Zhang 0001, Jingjing Li 0001, Sen Wang 0001, Zi Huang |
ACM Multimedia | 1 |
| 2023 | Cal-SFDA: Source-Free Domain-adaptive Semantic Segmentation with Differentiable Expected Calibration ErrorabstractThe prevalence of domain adaptive semantic segmentation has prompted concerns regarding source domain data leakage, where private information from the source domain could inadvertently be exposed in the target domain. To circumvent the requirement for source data, source-free domain adaptation has emerged as a viable solution that leverages self-training methods to pseudo-label high-confidence regions and adapt the model to the target data. However, the confidence scores obtained are often highly biased due to overconfidence and class-imbalance issues, which render both model selection and optimization problematic. In this paper, we propose a novel calibration-guided source-free domain adaptive semantic segmentation (Cal-SFDA) framework. The core idea is to estimate the expected calibration error (ECE) from the segmentation predictions, serving as a strong indicator of the model's generalization capability to the unlabeled target domain. The estimated ECE scores, in turn, assist the model training and fair selection in both source training and target adaptation stages. During model pre-training on the source domain, we ensure the differentiability of the ECE objective by leveraging the LogSumExp trick and using ECE scores to select the best source checkpoints for adaptation. To enable ECE estimation on the target domain without requiring labels, we train a value net for ECE estimation and apply statistic warm-up on its BatchNorm layers for stability. The estimated ECE scores assist in determining the reliability of prediction and enable class-balanced pseudo-labeling by positively guiding the adaptation progress and inhibiting potential error accumulation. Extensive experiments on two widely-used synthetic-to-real transfer tasks show that the proposed approach surpasses previous state-of-the-art by up to 5.25% of mIoU with fair model selection criteria. Yadan Luo, Zhi Chen 0010, Sen Wang 0001, Zi Huang |
ACM Multimedia | 3 |
| 2023 | GSMFlow: Generation Shifts Mitigating Flow for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) aims to recognize images not only for seen classes but also for unseen ones by transferring semantic-visual relationships from the seen to the unseen classes. It is an intuitive solution to take the advantage of generative models to hallucinate realistic unseen samples based on the knowledge learned from the seen classes. However, due to the generation shifts, the synthesized samples by most existing methods may drift from the real distribution of the unseen data. To address this issue, we propose a novel flow-based generative framework that consists of multiple conditional affine coupling layers for learning unseen data generation. Specifically, we investigate and address three essential problems that trigger the generation shifts,i.e.,semantic inconsistency,variance collapse, andstructure disorder. First, to improve the reflection of the semantic information in the generated samples, we proactively embed the semantic information into the transformation in each conditional affine coupling layer. Second, to promote the intrinsic feature variance of the unseen classes, we introduce a boundary sample mining strategy with entropy maximization to discover ambiguous visual variants of semantic prototypes and hereby calibrate the decision boundary of the classifiers. Third, a relative positioning strategy is proposed to revise the attribute embeddings, guiding which to fully preserve the inter-class geometric structure and further avoid structure disorder in the semantic space. Extensive experimental results on four GZSL benchmark datasets demonstrate that GSMFlow achieves the state-of-the-art performance on GZSL. Zhi Chen 0010, Yadan Luo, Sen Wang 0001, Jingjing Li 0001, Zi Huang |
IEEE Trans. Multim. | 1 |
| 2022 | Distinguishing Unseen from Seen for Generalized Zero-shot LearningabstractGeneralized zero-shot learning (GZSL) aims to recognize samples whose categories may not have been seen at training. Recognizing unseen classes as seen ones or vice versa often leads to poor performance in GZSL. Therefore, distinguishing seen and unseen domains is naturally an effective yet challenging solution for GZSL. In this paper, we present a novel method which leverages both visual and semantic modalities to distinguish seen and unseen categories. Specifically, our method deploys two variational autoencoders to generate latent representations for visual and semantic modalities in a shared latent space, in which we align latent representations of both modalities by Wasserstein distance and reconstruct two modalities with the representations of each other. In order to learn a clearer boundary between seen and unseen classes, we propose a two-stage training strategy which takes advantage of seen and unseen semantic descriptions and searches a threshold to separate seen and unseen visual samples. At last, a seen expert and an unseen expert are used for final classification. Extensive experiments on five widely used benchmarks verify that the proposed method can significantly improve the results of GZSL. For instance, our method correctly recognizes more than 99% samples when separating domains and improves the final classification accuracy from 72.6% to 82.9% on AWA1. Hongzu Su, Jingjing Li 0001, Zhi Chen 0010, Lei Zhu 0002, Ke Lu 0001 |
CVPR | 3 |
| 2022 | FBG-Based Variable-Length Estimation for Shape Sensing of Extensible Soft Robotic ManipulatorsabstractIn this paper, we propose a novel variable-length estimation approach for shape sensing of extensible soft robots utilizing fiber Bragg gratings (FBGs). Shape reconstruction from FBG sensors has been increasingly developed for soft robots, while the narrow stretching range of FBG fiber makes it difficult to acquire accurate sensing results for extensible robots. Towards this limitation, we newly introduce an FBG-based length sensor by leveraging a rigid curved channel, through which FBGs are allowed to slide within the robot following its body extension/compression, hence we can search and match the FBGs with specific constant curvature in the fiber to determine the effective length. From the fusion with the above measurements, a model-free filtering technique is accordingly presented for simultaneous calibration of a variable-length model and temporally continuous length estimation of the robot, enabling its accurate shape sensing using solely FBGs. The performances of the proposed method have been experimentally evaluated on an extensible soft robot equipped with an FBG fiber in both free and unstructured environments. The results concerning dynamic accuracy and robustness of length estimation and shape sensing demonstrate the effectiveness of our approach. Yiang Lu, Wei Chen 0068, Zhi Chen 0010, Jianshu Zhou, Yun-Hui Liu 0001 |
IROS | 3 |
| 2022 | Pixel Exclusion: Uncertainty-aware Boundary Discovery for Active Cross-Domain Semantic SegmentationabstractUnsupervised Domain Adaptation (UDA) has been shown to alleviate the heavy annotations for semantic segmentation. Recently, numerous self-training approaches are proposed to address the challenging cross-domain semantic segmentation problem. However, there still exists two open issues: (1) The generated pseudo-labels are inevitably noisy without external supervision. (2) These is a performance gap between UDA models and the fully-supervised model. In this paper, we propose to investigate Active Learning (AL) that selects a small portion of unlabeled pixels (or images) to be annotated, which leads to an impressive performance gain. Specifically, we propose a novel Uncertainty-aware Boundary Discovery (UBD) strategy that selects the uncertain pixels in the boundary areas that contains rich contextual information. Technically, we firstly select the pixels with top entropy values, and then re-select the pixels that are exclusive to their neighbors. We leverage the Kullback-Leibler divergence between one pixel's softmax prediction and its neighbors' to measure its "exclusivity". Extensive experiments show that our approach outperforms previous methods with both pixel-level and image-level label acquisition protocols. Fuming You, Jingjing Li 0001, Zhi Chen 0010, Lei Zhu 0002 |
ACM Multimedia | 3 |
| 2021 | Semantics Disentangling for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) aims to classify samples under the assumption that some classes are not observable during training. To bridge the gap between the seen and unseen classes, most GZSL methods attempt to associate the visual features of seen classes with attributes or to generate unseen samples directly. Nevertheless, the visual features used in the prior approaches do not necessarily encode semantically related information that the shared attributes refer to, which degrades the model generalization to unseen classes. To address this issue, in this paper, we propose a novel semantics disentangling framework for the generalized zero-shot learning task (SDGZSL), where the visual features of unseen classes are firstly estimated by a conditional VAE and then factorized into semantic-consistent and semantic-unrelated latent vectors. In particular, a total correlation penalty is applied to guarantee the independence between the two factorized representations, and the semantic consistency of which is measured by the derived relation network. Extensive experiments conducted on four GZSL benchmark datasets have evidenced that the semantic-consistent features disentangled by the proposed SDGZSL are more generalizable in tasks of canonical and generalized zero-shot learning. Our source code is available at https://github.com/uqzhichen/SDGZSL. Zhi Chen 0010, Yadan Luo, Ruihong Qiu, Sen Wang 0001, Zi Huang, Jingjing Li 0001, Zheng Zhang 0006 |
ICCV | 1 |
| 2021 | Local Graph Convolutional Networks for Cross-Modal HashingabstractCross-modal hashing aims to map the data of different modalities into a common binary space to accelerate the retrieval speed. Recently, deep cross-modal hashing methods have shown promising performance by applying deep neural networks to facilitate feature learning. However, the known supervised deep methods mainly rely on the labeled information of datasets, which is insufficient to characterize the latent structures that exist among different modalities. To mitigate this problem, in this paper, we propose to use Graph Convolutional Networks (GCNs) to exploit the local structure information of datasets for cross-modal hash learning. Specifically, a local graph is constructed according to the neighborhood relationships between samples in deep feature spaces and fed into GCNs to generate graph embeddings. Then, a within-modality loss is designed to measure the inner products between deep features and graph embeddings so that hashing networks and GCNs can be jointly optimized. By taking advantage of GCNs to assist model's training, the performance of hashing networks can be improved. Extensive experiments on benchmarks verify the effectiveness of the proposed method. Yudong Chen 0002, Sen Wang 0001, Jianglin Lu, Zhi Chen 0010, Zheng Zhang 0006, Zi Huang |
ACM Multimedia | 4 |
| 2021 | Mitigating Generation Shifts for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) is the task of leveraging semantic information to recognize seen and unseen samples, where unseen classes are not observable during training. It is natural to derive generative models and hallucinate training samples for unseen classes based on the knowledge learned from the seen samples. However, most of these models suffer from the generation shifts, where the synthesized samples may drift from the real distribution of unseen data. In this paper, we propose a novel generative flow framework that consists of multiple conditional affine coupling layers for learning unseen data generation. In particular, we identify three potential problems that trigger the generation shifts, i.e., semantic inconsistency, variance collapse, and structure disorder and address them respectively. First, to reinforce the correlations between the generated samples and their corresponding attributes, we explicitly embed the semantic information into the transformations in each coupling layer. Second, to recover the intrinsic variance of the real unseen features, we introduce a visual perturbation strategy to diversify the generated data and hereby help adjust the decision boundary of the classifiers. Third, a relative positioning strategy is proposed to revise the attribute embeddings, guiding them to fully preserve the inter-class geometric structure and further avoid structure disorder in the semantic space. Experimental results demonstrate that GSMFlow achieves the state-of-the-art performance on GZSL. Zhi Chen 0010, Yadan Luo, Sen Wang 0001, Ruihong Qiu, Jingjing Li 0001, Zi Huang |
ACM Multimedia | 1 |
| 2021 | CausalRec: Causal Inference for Visual Debiasing in Visually-Aware RecommendationabstractVisually-aware recommendation on E-commerce platforms aims to leverage visual information of items to predict a user's preference for these items in addition to the historical user-item interaction records. It is commonly observed that user's attention to visual features does not always reflect the real preference. Although a user may click and view an item in light of a visual satisfaction of their expectations, a real purchase does not always occur due to the unsatisfaction of other essential features (e.g., brand, material, price). We refer to the reason for such a visually related interaction deviating from the real preference as a visual bias. Existing visually-aware models make use of the visual features as a separate collaborative signal similarly to other features to directly predict the user's preference without considering a potential bias, which gives rise to a visually biased recommendation. In this paper, we derive a causal graph to identify and analyze the visual bias of these existing methods. In this causal graph, the visual feature of an item acts as a mediator, which could introduce a spurious relationship between the user and the item. To eliminate this spurious relationship that misleads the prediction of the user's real preference, an intervention and a counterfactual inference are developed over the mediator. Particularly, the Total Indirect Effect is applied for a debiased prediction during the testing phase of the model. This causal inference framework is model agnostic such that it can be integrated into the existing methods. Furthermore, we propose a debiased visually-aware recommender system, denoted as CausalRec to effectively retain the supportive significance of the visual information and remove the visual bias. Extensive experiments are conducted on eight benchmark datasets, which shows the state-of-the-art performance of CausalRec and the efficacy of debiasing. Ruihong Qiu, Sen Wang 0001, Zhi Chen 0010, Hongzhi Yin, Zi Huang |
ACM Multimedia | 3 |
| 2021 | Domain Adaptive Semantic Segmentation without Source DataabstractDomain adaptive semantic segmentation is recognized as a promising technique to alleviate the domain shift between the labeled source domain and the unlabeled target domain in many real-world applications, such as automatic pilot. However, large amounts of source domain data often introduce significant costs in storage and training, and sometimes the source data is inaccessible due to privacy policies. To address these problems, we investigate domain adaptive semantic segmentation without source data, which assumes that the model is pre-trained on the source domain, and then adapting to the target domain without accessing source data anymore. Since there is no supervision from the source domain data, many self-training methods tend to fall into the winner-takes-all dilemma, where the majority classes totally dominate the segmentation networks and the networks fail to classify the minority classes. Consequently, we propose an effective framework for this challenging problem with two components: positive learning and negative learning. In positive learning, we select the class-balanced pseudo-labeled pixels with intra-class threshold, while in negative learning, for each pixel, we investigate which category the pixel does not belong to with the proposed heuristic complementary label selection. Notably, our framework can be easily implemented and incorporated with other methods to further enhance the performance. Extensive experiments on two widely-used synthetic-to-real benchmarks demonstrate our claims and the effectiveness of our framework, which outperforms the baseline with a large margin. Code is available at https://github.com/fumyou13/LDBE. Fuming You, Jingjing Li 0001, Lei Zhu 0002, Zhi Chen 0010, Zi Huang |
ACM Multimedia | 4 |
| 2020 | Rethinking Generative Zero-Shot Learning: An Ensemble Learning Perspective for Recognising Visual PatchesabstractZero-shot learning (ZSL) is commonly used to address the very pervasive problem of predicting unseen classes in fine-grained image classification and other tasks. One family of solutions is to learn synthesised unseen visual samples produced by generative models from auxiliary semantic information, such as natural language descriptions. However, for most of these models, performance suffers from noise in the form of irrelevant image backgrounds. Further, most methods do not allocate a calculated weight to each semantic patch. Yet, in the real world, the discriminative power of features can be quantified and directly leveraged to improve accuracy and reduce computational complexity. To address these issues, we propose a novel framework called multi-patch generative adversarial nets (MPGAN) that synthesises local patch features and labels unseen classes with a novel weighted voting strategy. The process begins by generating discriminative visual features from noisy text descriptions for a set of predefined local patches using multiple specialist generative models. The features synthesised from each patch for unseen classes are then used to construct an ensemble of diverse supervised classifiers, each corresponding to one local patch. A voting strategy averages the probability distributions output from the classifiers and, given that some patches are more discriminative than others, a discrimination-based attention mechanism helps to weight each patch accordingly. Extensive experiments show that MPGAN has significantly greater accuracy than state-of-the-art methods. Zhi Chen 0010, Sen Wang 0001, Jingjing Li 0001, Zi Huang |
ACM Multimedia | 1 |
| 2020 | Relationship graph learning network for visual relationship detectionabstractVisual relationship detection aims to predict the relationships between detected object pairs. It is well believed that the correlations between image components (i.e., objects and relationships between objects) are significant considerations when predicting objects' relationships. However, most current visual relationship detection methods only exploited the correlations among objects, and the correlations among objects' relationships remained underexplored. This paper proposes a relationship graph learning network (RGLN) to explore the correlations among objects' relationships for visual relationship detection. Specifically, RGLN obtains image objects using an object detector, and then, every pair of objects constitutes a relationship proposal. All relationship proposals construct a relationship graph, in which the proposals are treated as nodes. Accordingly, RGLN designs bi-stream graph attention subnetworks to detect relationship proposals, in which one graph attention subnetwork analyzes correlations among relationships based on visual and spatial information, and the other analyzes correlations based on semantic and spatial information. Besides, RGLN exploits a relationship selection subnetwork to ignore redundant information of object pairs with no relationships. We conduct extensive experiments on two public datasets: the VRD and the VG datasets. The experimental results compared with the state-of-the-art demonstrate the competitiveness of RGLN. Jun Yu 0002, Yibing Zhan, Zhi Chen 0010 |
MMAsia | 4 |
| 2020 | CANZSL: Cycle-Consistent Adversarial Networks for Zero-Shot Learning from Natural LanguageabstractExisting methods using generative adversarial approaches for Zero-Shot Learning (ZSL) aim to generate realistic visual features from class semantics by a single generative alignment, which is highly under-constrained. As a result, the previous methods cannot guarantee that the generated visual features can truthfully reflect the corresponding semantics. To address this issue, we propose a novel method named Cycle-consistent Adversarial Networks for Zero-Shot Learning (CANZSL). It encourages a visual feature generator to synthesize realistic visual features from semantics, and then inversely translate back the synthesized visual features to the corresponding semantic space by a semantic feature generator. Furthermore, in this paper a more challenging and practical ZSL problem is considered where the original semantics are from natural language with irrelevant words instead of clean semantics, which are widely used in previous work. Specifically, a multi-modal consistent bidirectional generative adversarial model is trained to handle unseen instances by suppressing noise in the natural language. A forward one-to-many mapping from the class level descriptions to the visual features is coupled with an inverse many-to-one mapping from the visual space to the semantic space. Thus, a multi-modal cycle-consistency loss between the synthesized semantic representations and the ground truth can be learned and leveraged to enforce the generated semantic features to approximate to the real distribution in semantic space. Extensive experiments are conducted to demonstrate that our method consistently outperforms state-of-the-art approaches on natural language-based zero-shot learning tasks. Zhi Chen 0010, Jingjing Li 0001, Yadan Luo, Zi Huang, Yangyang Yangyang |
WACV | 1 |