Zhaozhi Xie

dblp:231/2328 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0001-6698-184XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Learning Auxiliary Representations With Inconsistency-Guided Detail Regularization for Mask-Guided Matting
abstract
Mask-guided matting networks have achieved significant improvements and have shown great potential in practical applications in recent years. However, simply learning matting representation from synthetic and lack-of-real-world-diversity matting data, these approaches tend to overfit low-level details in wrong regions, lack generalization to objects with complex structures and real-world scenes such as shadows, as well as suffer from interference of background lines or textures. To address these challenges, in this paper, we propose a novel auxiliary learning framework for mask-guided matting models, incorporating three auxiliary tasks: semantic segmentation, edge detection, and background line detection besides matting, to learn different and effective auxiliary representations from different types of data and annotations. Our framework and model introduce the following key aspects: 1) to learn real-world adaptive semantic representation for objects with diverse and complex structures under real-world scenes, we introduce extra semantic segmentation and edge detection tasks on more diverse real-world data with segmentation annotations; 2) to avoid overfitting on low-level details, we propose a module to utilize the inconsistency between learned segmentation and matting representations to regularize detail refinement; 3) we propose a novel background line detection task into our auxiliary learning framework, to suppress interference of background lines or textures. In addition, we propose a high-quality matting benchmark, Plant-Mat, to evaluate matting methods on complex structures. Extensively quantitative and qualitative results show that our approach outperforms state-of-the-art mask-guided methods.
Zhaozhi Xie, Longjie Qi, Jingyong Cai, Hiroyuki Uchiyama, Yue Ding 0001, Hongtao Lu 0001
IEEE Trans. Multim.2
2024 PA-SAM: Prompt Adapter SAM for High-Quality Image Segmentation
abstract
The Segment Anything Model (SAM) has exhibited outstanding performance in various image segmentation tasks. Despite being trained with over a billion masks, SAM faces challenges in mask prediction quality in numerous scenarios, especially in real-world contexts. In this paper, we introduce a novel prompt-driven adapter into SAM, namely Prompt Adapter Segment Anything Model (PA-SAM), aiming to enhance the segmentation mask quality of the original SAM. By exclusively training the prompt adapter, PA-SAM extracts detailed information from images and optimizes the mask decoder feature at both sparse and dense prompt levels, improving the segmentation performance of SAM to produce high-quality masks. Experimental results demonstrate that our PA-SAM outperforms other SAM-based methods in high-quality, zero-shot, and open-set segmentation. We’re making the source code and models available at https://github.com/xzz2/pa-sam.
Zhaozhi Xie, Bochen Guan, Muyang Yi, Yue Ding 0001, Hongtao Lu 0001, Lei Zhang 0006
ICME1
2024 BARTENDER: A simple baseline model for task-level heterogeneous federated learning
abstract
This study presents the Task-level Heterogeneous Federated Learning (TH-FL), a novel paradigm that fuses the principles of Federated Learning (FL) and Multi-Task Learning (MTL). In the TH-FL scenario, each client can learn an indefinite number of tasks, which may vary in type and originate from distinct domains. We introduce a unique baseline model, BARTENDER, that integrates a Conditional Prompt (CP) module. This module encodes task-specific and domain-specific information, enabling the model to generate tailored outputs based on the encoding inputs. This innovative strategy not only minimizes the communication costs associated with FL but also enhances model generalization across a variety of task types. Through extensive experiments, we establish that the BARTENDER model surpasses traditional multi-decoder architecture models across diverse scenarios. We also explore the influence of the parameter decoupling strategy on model training and outline the assumptions necessary for achieving a $O\left( {1/\sqrt T } \right)$ convergence speed in the TH-FL scenario.
Yuwen Yang, Suizhi Huang, Shalayiding Sirejiding, Chang Liu 0078, Muyang Yi, Zhaozhi Xie, Yue Ding 0001, Hongtao Lu 0001
ICME7
2024 Superpixel Guided Network for Weakly Supervised Semantic Segmentation
abstract
Image-level weakly supervised semantic segmentation faces challenges in accurately capturing boundaries and representing intricate details due to the absence of pixel-level supervision. Constrained by the enormous number of pixels, pixel-level propagation has difficulty in capturing the long-range dependency, particularly in small, isolated regions. To this end, we introduce a novel approach of self-supervised segmentation integrated with superpixel, and develop a network called superpixel guided network (SPGNet) to simultaneously perform superpixel generation and segmentation mask prediction. Significantly, our framework facilitates mutual supervised learning between the segmentation branch and the superpixel branch. The superpixel guides the predicted mask for improved boundary location, while the latter provides supervision on superpixel through superpixel center generation (SCG) and union boundary extraction (UBE). Furthermore, we propose superpixel context fusion (SCF) to generate compact pseudo masks and capture long-range dependency. Experimental results demonstrate that the proposed SPGNet achieves outstanding performance on the PASCAL VOC 2012 segmentation benchmark
Zhaozhi Xie, Yuwen Yang, Hongtao Lu 0001
IEEE Signal Process. Lett.1
2023 Trimap-guided feature mining and fusion network for natural image matting
Dongdong Yu, Zhaozhi Xie, Yaoyi Li, Zehuan Yuan, Hongtao Lu 0001
Comput. Vis. Image Underst.3
2022 Exploring Category Consistency for Weakly Supervised Semantic Segmentation
abstract
Self-supervised framework has been widely used in weakly supervised semantic segmentation. Generating a reliable and detailed pseudo mask label is the main challenge for improving the quality of predicted mask. In this paper, we propose Category Consistency Mask Refinement (CCMR) to explore the category consistency cued with the input image, and inject such information to mask refinement, guaranteeing the completeness of the refined mask. Moreover, we exploit Selective Weighted Pooling (SWP) to restrict the backward propagation of background, limiting the update of the background. Experimental results demonstrate that our methods can boost the performance on the PASCAL VOC 2012 segmentation benchmark, outperforming the state-of-the-art weakly supervised semantic segmentation methods.
Zhaozhi Xie, Hongtao Lu 0001
ICASSP1
2022 Merged U-Net for Bone Tumors X-Ray Images Segmentation
abstract
Bone tumors X-ray image segmentation is crucial for many medical image processing applications such as X-ray image enhancement, lesion diagnosis, etc. In this paper, we propose a new topological structure of u-net, named merged u-net, for bone tumors X-ray image segmentation. We attach a top-down merged branch to the vanilla encoder-decoder network, enhancing the hierarchical feature aggregation. In the merged path, a multi-feature aggregation block, named merged gate, is proposed for better localization of lesion area. According to the attention matrix, the proposed merged gate can generate a reconstructed feature that contains both low-level information and high-level information. Then we incorporate the reconstructed feature with the small scare feature cued with a local feature fusion method, achieving multi-scare feature aggregation. Additionally, we propose a new X-ray image dataset for bone tumors segmentation, which consists of 88 benign images and 217 malignant images, provided with corresponding segmentation masks. Experimental results demonstrate that the proposed merged u-net outperforms other u-net based medical segmentation methods on the proposed X-ray image dataset.
Zhaozhi Xie, Keyang Zhao, Sheng-Hui Wu, Jiong Mei, Hongtao Lu 0001
ICIP1
2021 Text Detection by Jointly Learning Character and Word Regions
Deyang Wu, Xingfei Hu, Zhaozhi Xie, Usman Ali 0009, Hongtao Lu 0001
ICDAR (1)3
2018 Copy-move detection of digital audio based on multi-feature decision
Zhaozhi Xie, Wei Lu 0001, Xianjin Liu, Yingjie Xue, Yuileong Yeung
J. Inf. Secur. Appl.1