Mengbiao Zhao

dblp:280/0810 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2024 Sequential Transformer for End-to-End Video Text Detection
abstract
In existing methods of video text detection, the detection and tracking branches are usually independent of each other, and although they jointly optimize the backbone network, the tracking-by-detection paradigm still needs to be used during the inference stage. To address this issue, we propose a novel video text detection framework based on sequential transformer, which decodes detection and tracking tasks in parallel, without explicitly setting up a tracking branch. To achieve this, we first introduce the concept of instance query, which learns long-term context information in the video sequence. Then, based on the instance query, the transformer decoder is used to predict the entire box and mask sequence of the text instance in one pass. As a result, the tracking task is realized naturally. In addition, the proposed method can be applied to the scene text detection task seamlessly, without modifying any modules. To the best of our knowledge, this is the first framework to unify the tasks of scene text detection and video text detection. Our model achieves state-of-the-art performance on four video text datasets (YVT, RT-1K, BOVText, and BiRViT-1K), and competitive results on three scene text datasets (CTW1500, MSRA-TD500, and Total-Text). The code is available at https://github.com/zjb-1/SeqVideoText.
Mengbiao Zhao, Cheng-Lin Liu 0001
WACV2
2024 Video Text Detection With Robust Feature Representation
abstract
Existing video text detection methods mostly track texts with appearance feature only, thus are easily influenced by the change of perspective and illumination. In this paper, we propose an end-to-end video text detector that tracks texts based on robust feature representation fusing multiple descriptors. First, we introduce a character center segmentation branch to extract semantic feature, which encodes the category and position information of characters. And for extracting the topology feature of each text instance, we propose a relative position awareness branch to encode the relative position information among texts. Then, an adaptive feature fusion network is proposed to dynamically fuse multiple descriptors to generate a robust feature representation for more robust tracking. In addition, to promote the research and evaluation in this field, we also construct a large Bilingual Road scene Video Text dataset, named BiRViT-1K, which contains 1000 videos of Chinese and English texts. Experimental results show the proposed semantic and topology features are beneficial to the text detection and tracking performance, and the proposed method achieves state-of-the-art performance on four public video text benchmarks ICDAR 2015 Video, YVT, RT-1K and BOVText, and two Chinese scene text benchmarks CASIA10K and MSRA-TD500.
Wei Feng 0016, Mengbiao Zhao, Xu-Yao Zhang, Cheng-Lin Liu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Texts as points: Scene text detection with point supervision
Mengbiao Zhao, Wei Feng 0016, Cheng-Lin Liu 0001
Pattern Recognit. Lett.1
2022 A Large-Scale Database for Chemical Structure Recognition and Preliminary Evaluation
abstract
Chemical structure recognition (CSR), transforming chemical structure images into formulas in character strings (such as SMILES), is a challenging problem due to the complex 2D structures and relationships. For this research, there is not a database of sufficient scale and diversity for model design and fair evaluation. In this paper, we present a large-scale chemical structure database named CASIA-CSDB, containing 480,668 samples (images corresponding to SMILES strings). To construct the database, we select chemical structures from the ChEMBL, a well-known bioactive molecules database, and use the RDKit tool to generate images according to the chemical format SMILES strings. The selected structures represent the major types of chemical compounds covering eight weight partitions. We also select a subset of 97,309 samples of the database to form the Mini-CASIA-CSDB database. To provide a benchmark, we evaluate three state-of-the-art image-to-markup recognition methods on the database. The results demonstrate the challenge of the database. The database with its annotation is available at http://www.nlpr.ia.ac.cn/databases/CASIA-CSDB/index.html.
Longfei Ding, Mengbiao Zhao, Shuiling Zeng, Cheng-Lin Liu 0001
ICPR2
2022 Mixed-Supervised Scene Text Detection With Expectation-Maximization Algorithm
abstract
Scene text detection is an important and challenging task in computer vision. For detecting arbitrarily-shaped texts, most existing methods require heavy data labeling efforts to produce polygon-level text region labels for supervised training. In order to reduce the cost in data labeling, we study mixed-supervised arbitrarily-shaped text detection by combining various weak supervision forms (e.g., image-level tags, coarse, loose and tight bounding boxes), which are far easier to annotate. Whereas the existing weakly-supervised learning methods (such as multiple instance learning) do not promote full object coverage, to approximate the performance of fully-supervised detection, we propose an Expectation-Maximization (EM) based mixed-supervised learning framework to train scene text detector using only a small amount of polygon-level annotated data combined with a large amount of weakly annotated data. The polygon-level labels are treated as latent variables and recovered from the weak labels by the EM algorithm. A new contour-based scene text detector is also proposed to facilitate the use of weak labels in our mixed-supervised learning framework. Extensive experiments on six scene text benchmarks show that (1) using only 10% strongly annotated data and 90% weakly annotated data, our method yields comparable performance to that of fully supervised methods, (2) with 100% strongly annotated data, our method achieves state-of-the-art performance on five scene text benchmarks (CTW1500, Total-Text, ICDAR-ArT, MSRA-TD500, and C-SVT), and competitive results on the ICDAR2015 Dataset. We will make our weakly annotated datasets publicly available.
Mengbiao Zhao, Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001
IEEE Trans. Image Process.1
2021 End-to-End Detection and Recognition of Arithmetic Expressions
Jiangpeng Wan, Mengbiao Zhao, Xu-Yao Zhang, Linlin Huang 0001
PRCV (1)2
2020 Mutually Guided Dual-Task Network for Scene Text Detection
abstract
Scene text detection has been studied extensively. Existing methods detect either words or text lines and use either word-level or line-level annotated data for training. In this paper, we propose a dual-task network that can perform word-level and line-level text detection simultaneously and use training data of both levels of annotation to boost the performance. The dual-task network has two detection heads for word-level and line-level text detection, respectively. Then we propose a mutual guidance scheme for the joint training of the two tasks with two modules: line filtering module utilizes the output feature map of the text line detector to filter out the non-text regions for the word detector, and word enhancing module provides prior positions of words for the text line detector depending on the output feature map of the word detector. Experimental results of word-level and line-level text detection demonstrate the effectiveness of the proposed dual-task network and mutual guidance scheme, and the results of our method are competitive with state-of-the-art methods.
Mengbiao Zhao, Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001
ICPR1