Baole Wei

dblp:260/6320 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 GlyphSR: A Simple Glyph-Aware Framework for Scene Text Image Super-Resolution
abstract
The goal of scene text image super-resolution (STISR) is to enhance the clarity of text within line images, thereby improving readability and enabling more accurate text recognition. However, existing STISR methods often rely heavily on Text Prior (TP) derived from trained recognizers, which can be unreliable and may lead to incorrect glyph restoration. Text images contain two crucial types of information: semantic content from word meanings and structural details from glyphs. When semantic information is unreliable, accurate perception of glyph structures becomes essential. This paper introduces GlyphSR, a novel STISR framework that addresses three key challenges: precise extraction, effective learning, and optimal utilization of glyph structural information. GlyphSR incorporates the Glyph Extraction Module (GEM), a training-free approach leveraging the Segment Anything Model (SAM) to accurately extract character-level glyphs. The Glyph Perception Module (GPM) models and learns glyph structures through segmentation and classification tasks, while the Glyph Fusion Module (GFM) integrates glyph information to enhance overall STISR model performance. Extensive experiments on the TextZoom dataset demonstrate that GlyphSR achieves a new state-of-the-art performance.
Baole Wei, Liangcai Gao, Zhi Tang 0001
AAAI1
2025 Vote & Mix: Plug-and-Play Token Reduction for Efficient Vision Transformer
abstract
Despite the remarkable success of Vision Transformers (ViTs) in various visual tasks, they are often hindered by substantial computational cost. In this work, we introduce Vote&Mix (VoMix), a plug-and-play and parameter-free token reduction method, which can be readily applied to off-the-shelf ViT models without any training. VoMix tackles the computational redundancy of ViTs by identifying tokens with high homogeneity through a layer-wise token similarity voting mechanism. Subsequently, the selected tokens are mixed into the retained set, thereby preserving visual information. Experiments demonstrate VoMix significantly improves the speed-accuracy tradeoff of ViTs on both images and videos. Without any training, VoMix achieves a 2× increase in throughput of existing ViT-H on ImageNet-1K and a 2.4× increase in throughput of existing ViT-L on Kinetics-400 video dataset, with a mere 0.3% drop in top-1 accuracy.
Shuai Peng, Di Fu, Baole Wei, Liangcai Gao, Zhi Tang 0001
ICME3
2025 Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression Recognition
abstract
Handwritten Mathematical Expression Recognition (HMER) remains a persistent challenge in Optical Character Recognition (OCR) due to the inherent freedom of symbol layouts and variability in handwriting styles. Prior methods have faced performance bottlenecks by proposing isolated architectural modifications, making them difficult to integrate coherently into a unified framework. Meanwhile, recent advances in pretrained vision-language models (VLMs) have demonstrated strong cross-task generalization, offering a promising foundation for developing unified solutions. In this paper, we introduce Uni-MuMER, which fully fine-tunes a VLM for the HMER task without modifying its architecture, effectively injecting domain-specific knowledge into a generalist framework. Our method integrates three data-driven tasks: Tree-Aware Chain-of-Thought (Tree-CoT) for structured spatial reasoning, Error-Driven Learning (EDL) for reducing confusion among visually similar characters, and Symbol Counting (SC) for improving recognition consistency in long expressions. Experiments on the CROHME and HME100K datasets show that Uni-MuMER achieves super state-of-the-art performance, outperforming the best lightweight specialized model SSAN by 16.31\% and the top-performing VLM Gemini2.5-flash by 24.42\% under zero-shot setting. Our datasets, models, and code are open-sourced at: https://github.com/BFlameSwift/Uni-MuMER
Shuai Peng, Baole Wei, Liangcai Gao
NeurIPS5
2024 Maskstr: Guide Scene Text Recognition Models with Masking
abstract
Text recognition in information loss scenarios like blurriness, occlusion, and perspective distortion is challenging in real-world applications. To enhance robustness, some studies use extra unlabeled data for encoder pretraining. Others focus on improving decoder context reasoning. However, pretraining methods require abundant unlabeled data and high computing resources, while decoder-based approaches risk over-correction. In this paper, we propose MaskSTR, a dual-branch training framework for STR models, using patch masking to simulate information loss. MaskSTR guides visual representation learning, improving robustness to information loss conditions without extra data or training stages. Furthermore, we introduce Block Masking, a novel and straightforward mask generation method, for further performance enhancement. Experiments demonstrate MaskSTR’s effectiveness across CTC, attention, and Transformer decoding methods, achieving significant performance gains and setting new state-of-the-art results.
Baole Wei, Minghang He, Liangcai Gao, Duoyou Zhou, Xiang Bai, Zhi Tang 0001
ICASSP1
2024 Recognition-Guided Diffusion Model for Scene Text Image Super-Resolution
abstract
Scene Text Image Super-Resolution (STISR) aims to enhance the resolution and legibility of text within low-resolution (LR) images, consequently elevating recognition accuracy in Scene Text Recognition (STR). Previous methods predominantly employ discriminative Convolutional Neural Networks (CNNs) augmented with diverse forms of text guidance to address this issue. Nevertheless, they remain deficient when confronted with severely blurred images, due to their insufficient generation capability when little structural or semantic information can be extracted from original images. Therefore, we introduce RGDiffSR, a Recognition-Guided Diffusion model for scene text image Super-Resolution, which exhibits great generative diversity and fidelity even in challenging scenarios. Moreover, we propose a Recognition-Guided Denoising Network, to guide the diffusion model generating LR-consistent results through succinct semantic guidance. Experiments on the TextZoom dataset demonstrate the superiority of RGDiffSR over prior state-of-the-art methods in both text recognition accuracy and image fidelity.
Liangcai Gao, Zhi Tang 0001, Baole Wei
ICASSP4
2024 SegHist: A General Segmentation-Based Framework for Chinese Historical Document Text Line Detection
Xingjian Hu, Baole Wei, Liangcai Gao
ICDAR (3)2
2022 Molecular Formula Image Segmentation with Shape Constraint Loss and Data Augmentation
abstract
The increasing demand for molecular formula image data leads to formidable pressure for researchers. Most existing image segmentation approaches can not be directly utilized for molecules, and how to improve the coverage fineness and generate a large amount of labeled training data is worthy of further exploration. To this end, we establish a deep learning based molecular formula image segmentation model (DL-MFS). Specifically, we design a shape constraint loss (SCL) function to refine the detection frame position and propose a rule-based molecular formula image data augmentation method for solving the bottleneck problem that the lack of training data. Experimental results demonstrate the effectiveness of the proposed segmentation model.
Ruiqi Jia, Baole Wei, Guanren Qiao, Xiaoqing Lyu, Zhi Tang 0001
BIBM3
2022 MIR: A Benchmark for Molecular Image Retrival with a Cross-modal Pretraining Framework
abstract
Molecular image retrieval is one of the crucial steps in automatic mining and utilization of biochemistry-related literatures, which is also a relatively open and challenging task in cross fields of biochemistry and artificial intelligence. The challenges come from two aspects: 1) there is a lack of open datasets and evaluation criteria for molecular image retrieval. 2) Common retrieval methods always ignore that molecular image retrieval has cross-modal information of both images and SMILES texts. To address the first challenge, we firstly construct a new molecular image retrieval benchmark, named MIR, including 130770 molecular images, labeled structural similarity, and reasonable evaluation metrics. Faced with the second challenge, we propose an effective cross-modal pre-training framework for molecular image retrieval following CLIP. Experimental results reflect the effectiveness of our proposed benchmark MIR and cross-modal pre-training framework.
Baole Wei, Ruiqi Jia, Shihan Fu, Xiaoqing Lyu, Liangcai Gao, Zhi Tang 0001
BIBM1
2021 Adaptive Smooth L1 Loss: A Better Way to Regress Scene Texts with Extreme Aspect Ratios
abstract
In recent years, scene text detection has experienced rapid development. Regression-based methods are currently a mainstream method for scene text detection, and the effect of bounding box regression is a major factor limiting their detection performance. The regression of bounding boxes is greatly affected by the aspect ratio of texts since the text in natural scenes varies greatly in height and width. However, the existing methods ignore the difference between the height and width of the text in the bounding box regression, which leads to an imperfect regression effect and thus suppresses the performance of the scene text detection. In this paper, we propose an Adaptive Smooth L1 Loss function (abbreviated as ASLL) for bounding box regression, which can adaptively determine the weight of each regression variable according to the current state of the model during the training process, so as to guide the bounding box to regress in a more critical direction. The experimental results demonstrate that ASLL achieves promising performance on scene text detection. Specially, an F-measure of 84.56% is achieved on CTW-1500 dataset, surpassing the state-of-the-art detectors, and the detection results on TotalText and ICDAR2015 datasets are competitive to those of state-of-the-art methods.
Chao Liu 0020, Min Yu 0001, Baole Wei, Boquan Li 0002, Gang Li 0009, Weiqing Huang
ISCC4
2021 An end-to-end text spotter with text relation networks
abstract
Abstract Reading text in images automatically has become an attractive research topic in computer vision. Specifically, end-to-end spotting of scene text has attracted significant research attention, and relatively ideal accuracy has been achieved on several datasets. However, most of the existing works overlooked the semantic connection between the scene text instances, and had limitations in situations such as occlusion, blurring, and unseen characters, which result in some semantic information lost in the text regions. The relevance between texts generally lies in the scene images. From the perspective of cognitive psychology, humans often combine the nearby easy-to-recognize texts to infer the unidentifiable text. In this paper, we propose a novel graph-based method for intermediate semantic features enhancement, called Text Relation Networks. Specifically, we model the co-occurrence relationship of scene texts as a graph. The nodes in the graph represent the text instances in a scene image, and the corresponding semantic features are defined as representations of the nodes. The relative positions between text instances are measured as the weights of edges in the established graph. Then, a convolution operation is performed on the graph to aggregate semantic information and enhance the intermediate features corresponding to text instances. We evaluate the proposed method through comprehensive experiments on several mainstream benchmarks, and get highly competitive results. For example, on the , our method surpasses the previous top works by 2.1% on the word spotting task.
Baole Wei, Min Yu 0001, Gang Li 0009, Boquan Li 0002, Chao Liu 0020, Weiqing Huang
Cybersecur.2
2021 FakeFilter: A cross-distribution Deepfake detection system with domain adaptation
abstract
Abuse of face swap techniques poses serious threats to the integrity and authenticity of digital visual media. More alarmingly, fake images or videos created by deep learning technologies, also known as Deepfakes, are more realistic, high-quality, and reveal few tampering traces, which attracts great attention in digital multimedia forensics research. To address those threats imposed by Deepfakes, previous work attempted to classify real and fake faces by discriminative visual features, which is subjected to various objective conditions such as the angle or posture of a face. Differently, some research devises deep neural networks to discriminate Deepfakes at the microscopic-level semantics of images, which achieves promising results. Nevertheless, such methods show limited success as encountering unseen Deepfakes created with different methods from the training sets. Therefore, we propose a novel Deepfake detection system, named FakeFilter, in which we formulate the challenge of unseen Deepfake detection into a problem of cross-distribution data classification, and address the issue with a strategy of domain adaptation. By mapping different distributions of Deepfakes into similar features in a certain space, the detection system achieves comparable performance on both seen and unseen Deepfakes. Further evaluation and comparison results indicate that the challenge has been successfully addressed by FakeFilter.
Boquan Li 0002, Baole Wei, Gang Li 0009, Chao Liu 0020, Weiqing Huang, Meimei Li, Min Yu 0001
J. Comput. Secur.3