VLDB 2026 Research / reviewers in the wild / expert
Zixiao Wang 0002
dblp:141/1943-2
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0002-0009-5033ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Leveraging Text Localization for Scene Text Removal via Text-Aware Masked Image Modeling
Zixiao Wang 0002, Hongtao Xie 0001, Yuxin Wang 0002, Yadong Qu, Fengjun Guo, Pengwei Liu |
ECCV (66) | 1 |
| 2024 | Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition
Zuan Gao, Yuxin Wang 0002, Yadong Qu, Boqiang Zhang, Zixiao Wang 0002, Hongtao Xie 0001 |
IJCAI | 5 |
| 2024 | Focus on the Whole Character: Discriminative Character Modeling for Scene Text Recognition
Bangbang Zhou, Yadong Qu, Zixiao Wang 0002, Boqiang Zhang, Hongtao Xie 0001 |
IJCAI | 3 |
| 2024 | DCFP: Distribution Calibrated Filter Pruning for Lightweight and Accurate Long-Tail Semantic SegmentationabstractRecently, semantic segmentation has made promising progress, but the high cost of processing still limits its application. With focusing on removing the parameters of the networks, filter pruning using the importance criterion is a straightforward and effective technique to obtain the lightweight sub-network. However, we argue that the long-tail distribution in segmentation datasets poses two significant problems which are ignored in existing pruning algorithms: 1) The importance criterion is dominated by head classes which contain numerous positive samples, where the knowledge of tail classes is easily degenerated. 2) The degenerated knowledge of tail classes is hard to recover as their samples are also insufficient during fine-tuning. To address these issues, we propose a Distribution Calibrated Filter Pruning (DCFP) framework for segmentation. Firstly, a gradient-based Equalization Importance Criterion (EIC) is designed to generate a class-balanced pruning procedure. It avoids the bias on head classes by discarding the imbalanced positive gradients. Secondly, we introduce a Geometric-Semantic Re-balanced Loss (GSRL) to emphasize the learning on tail classes during fine-tuning. The GSRL consists of two cooperative components to calibrate the imbalanced optimization on geometric and semantic domains dynamically. Compared with previous methods, DCFP explores a novel distribution-aware pruning framework to obtain lightweight architectures with accurate results. Extensive experiments proved that DCFP achieves impressive performance on four popular segmentation benchmarks. Zixiao Wang 0002, Hongtao Xie 0001, Yuxin Wang 0002, Guoqing Jin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Symmetrical Linguistic Feature Distillation with CLIP for Scene Text RecognitionabstractIn this paper, we explore the potential of the Contrastive Language-Image Pretraining (CLIP) model in scene text recognition (STR), and establish a novel Symmetrical Linguistic Feature Distillation framework (named CLIP-OCR) to leverage both visual and linguistic knowledge in CLIP. Different from previous CLIP-based methods mainly considering feature generalization on visual encoding, we propose a symmetrical distillation strategy (SDS) that further captures the linguistic knowledge in the CLIP text encoder. By cascading the CLIP image encoder with the reversed CLIP text encoder, a symmetrical structure is built with an image-to-text feature flow that covers not only visual but also linguistic information for distillation. Benefiting from the natural alignment in CLIP, such guidance flow provides a progressive optimization objective from vision to language, which can supervise the STR feature forwarding process layer-by-layer. Besides, a new Linguistic Consistency Loss (LCL) is proposed to enhance the linguistic capability by considering second-order statistics during the optimization. Overall, CLIP-OCR is the first to design a smooth transition between image and text for the STR task. Extensive experiments demonstrate the effectiveness of CLIP-OCR with 93.8% average accuracy on six popular STR benchmarks. Code will be available at https://github.com/wzx99/CLIPOCR. Zixiao Wang 0002, Hongtao Xie 0001, Yuxin Wang 0002, Boqiang Zhang, Yongdong Zhang 0001 |
ACM Multimedia | 1 |
| 2023 | What is the Real Need for Scene Text Removal? Exploring the Background Integrity and Erasure Exhaustivity PropertiesabstractAs a crucial application in privacy protection, scene text removal (STR) has received amounts of attention in recent years. However, existing approaches coarsely erasing texts from images ignore two important properties: the background texture integrity (BI) and the text erasure exhaustivity (EE). These two properties directly determine the erasure performance, and how to maintain them in a single network is the core problem for STR task. In this paper, we attribute the lack of BI and EE properties to the implicit erasure guidance and imbalanced multi-stage erasure respectively. To improve these two properties, we propose a new ProgrEssively Region-based scene Text eraser (PERT). There are three key contributions in our study. First, a novel explicit erasure guidance is proposed to enhance the BI property. Different from implicit erasure guidance modifying all the pixels in the entire image, our explicit one accurately performs stroke-level modification with only bounding-box level annotations. Second, a new balanced multi-stage erasure is constructed to improve the EE property. By balancing the learning difficulty and network structure among progressive stages, each stage takes an equal step towards the text-erased image to ensure the erasure exhaustivity. Third, we propose two new evaluation metrics called BI-metric and EE-metric, which make up the shortcomings of current evaluation tools in analyzing BI and EE properties. Compared with previous methods, PERT outperforms them by a large margin in both BI-metric ( ↑ 6.13 %) and EE-metric ( ↑ 1.9 %), obtaining SOTA results with high speed (71 FPS) and at least 25% lower parameter complexity. Code will be available at https://github.com/wangyuxin87/PERT. Yuxin Wang 0002, Hongtao Xie 0001, Zixiao Wang 0002, Yadong Qu, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 3 |