EDBT 2026 Demo / reviewers in the wild / expert
Minghui Liao
dblp:190/7335
· DBLP profile ↗
31ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0002-2583-4314ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 9 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
Xudong Xie, Minghui Liao, Wei Chen 0088, Xiang Bai |
Int. J. Comput. Vis. | 6 |
| 2025 | Cross-Lingual Text-Rich Visual Comprehension: An Information Theory PerspectiveabstractRecent Large Vision-Language Models (LVLMs) have shown promising reasoning capabilities on text-rich images from charts, tables, and documents. However, the abundant text within such images may increase the model's sensitivity to language. This raises the need to evaluate LVLM performance on cross-lingual text-rich visual inputs, where the language in the image differs from the language of the instructions. To address this, we introduce XT-VQA (Cross-Lingual Text-Rich Visual Question Answering), a benchmark designed to assess how LVLMs handle language inconsistency between image text and questions. XT-VQA integrates five existing text-rich VQA datasets and a newly collected dataset, XPaperQA, covering diverse scenarios that require faithful recognition and comprehension of visual information despite language inconsistency. Our evaluation of prominent LVLMs on XT-VQA reveals a significant drop in performance for cross-lingual scenarios, even for models with multilingual capabilities. A mutual information analysis suggests that this performance gap stems from cross-lingual questions failing to adequately activate relevant visual information. To mitigate this issue, we propose MVCL-MI (Maximization of Vision-Language Cross-Lingual Mutual Information), where a visual-text cross-lingual alignment is built by maximizing mutual information between the model's outputs and visual information. This is achieved by distilling knowledge from monolingual to cross-lingual settings through KL divergence minimization, where monolingual output logits serve as a teacher. Experimental results on the XT-VQA demonstrate that MVCL-MI effectively reduces the visual-text cross-lingual performance disparity while preserving the inherent capabilities of LVLMs, shedding new light on the potential practice for improving LVLMs. Xinmiao Yu, Minghui Liao, Ya-Qi Yu, Xiachong Feng, Weihong Zhong, Ruihan Chen 0001, Mengkang Hu, Jihao Wu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
AAAI | 4 |
| 2025 | UI-Hawk: Unleashing the Screen Stream Understanding for Mobile GUI AgentsabstractGraphical User Interface (GUI) agents are expected to precisely operate on the screens of digital devices. Existing GUI agents merely depend on current visual observations and plain-text action history, ignoring the significance of history screens. To mitigate this issue, we propose UI-Hawk, a multi-modal GUI agent specially designed to process screen streams encountered during GUI navigation. UI-Hawk incorporates a history-aware visual encoder to handle the screen sequences. To acquire a better understanding of screen streams, we select four fundamental tasks—UI grounding, UI referring, screen question answering, and screen summarization. We further propose a curriculum learning strategy to subsequently guide the model from fundamental tasks to advanced screen-stream comprehension.Along with the efforts above, we have also created a benchmark FunUI to quantitatively evaluate the fundamental screen understanding ability of MLLMs. Extensive experiments on FunUI and GUI navigation benchmarks consistently validate that screen stream understanding is essential for GUI tasks.Our code and data are now available at https://github.com/IMNearth/UIHawk. Jiwen Zhang, Ya-Qi Yu, Minghui Liao, WenTao Li, Jihao Wu, Zhongyu Wei |
EMNLP | 3 |
| 2025 | Brain Wiring Knowledge Graph Reasoning: A Region Embedding Approach for Logical Neuronal Relation Inference
Zhengyun Zhou, Guojia Wan, Wenbin Hu 0001, Minghui Liao, Junchao Qiu, Bo Du 0001 |
MICCAI (12) | 5 |
| 2025 | Building connectome analysis tools with representation learning on neuronal skeleton and circuit topology
Minghui Liao, Guojia Wan, Wenbin Hu 0001, Bo Du 0001 |
Neural Networks | 1 |
| 2025 | Partial Scene Text RetrievalabstractThe task of partial scene text retrieval involves localizing and searching for text instances that are the same or similar to a given query text from an image gallery. However, existing methods can only handle text-line instances, leaving the problem of searching for partial patches within these text-line instances unsolved due to a lack of patch annotations in the training data. To address this issue, we propose a network that can simultaneously retrieve both text-line instances and their partial patches. Our method embeds the two types of data (query text and scene text instances) into a shared feature space and measures their cross-modal similarities. To handle partial patches, our proposed approach adopts a Multiple Instance Learning (MIL) approach to learn their similarities with query text, without requiring extra annotations. However, constructing bags, which is a standard step of conventional MIL approaches, can introduce numerous noisy samples for training, and lower inference speed. To address this issue, we propose a Ranking MIL (RankMIL) approach to adaptively filter those noisy samples. Additionally, we present a Dynamic Partial Match Algorithm (DPMA) that can directly search for the target partial patch from a text-line instance during the inference stage, without requiring bags. This greatly improves the search efficiency and the performance of retrieving partial patches. We evaluate the proposed method on both English and Chinese datasets in two tasks: retrieving text-line instances and partial patches. For English text retrieval, our method outperforms state-of-the-art approaches by 8.04% mAP and 12.71% mAP on average, respectively, among three datasets for the two tasks. For Chinese text retrieval, our approach surpasses state-of-the-art approaches by 24.45% mAP and 38.06% mAP on average, respectively, among three datasets for the two tasks. The source code and dataset are available at https://github.com/lanfeng4659/PSTR. Hao Wang 0207, Minghui Liao, Zhouyi Xie, Wenyu Liu 0001, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Joint Learning Neuronal Skeleton and Brain Circuit Topology with Permutation Invariant Encoders for Neuron ClassificationabstractDetermining the types of neurons within a nervous system plays a significant role in the analysis of brain connectomics and the investigation of neurological diseases. However, the efficiency of utilizing anatomical, physiological, or molecular characteristics of neurons is relatively low and costly. With the advancements in electron microscopy imaging and analysis techniques for brain tissue, we are able to obtain whole-brain connectome consisting neuronal high-resolution morphology and connectivity information. However, few models are built based on such data for automated neuron classification. In this paper, we propose NeuNet, a framework that combines morphological information of neurons obtained from skeleton and topological information between neurons obtained from neural circuit. Specifically, NeuNet consists of three components, namely Skeleton Encoder, Connectome Encoder, and Readout Layer. Skeleton Encoder integrates the local information of neurons in a bottom-up manner, with a one-dimensional convolution in neural skeleton's point data; Connectome Encoder uses a graph neural network to capture the topological information of neural circuit; finally, Readout Layer fuses the above two information and outputs classification results. We reprocess and release two new datasets for neuron classification task from volume electron microscopy(VEM) images of human brain cortex and Drosophila brain. Experiments on these two datasets demonstrated the effectiveness of our model with accuracies of 0.9169 and 0.9363, respectively. Code and data are available at: https://github.com/WHUminghui/NeuNet. Minghui Liao, Guojia Wan, Bo Du 0001 |
AAAI | 1 |
| 2024 | Self-Distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective ApproachabstractText recognition methods are gaining rapid development. Some advanced techniques, e.g., powerful modules, language models, and un- and semi-supervised learning schemes, consecutively push the performance on public benchmarks forward. However, the problem of how to better optimize a text recognition model from the perspective of loss functions is largely overlooked. CTC-based methods, widely used in practice due to their good balance between performance and inference speed, still grapple with accuracy degradation. This is because CTC loss emphasizes the optimization of the entire sequence target while neglecting to learn individual characters. We propose a self-distillation scheme for CTC-based model to address this issue. It incorporates a framewise regularization term in CTC loss to emphasize individual supervision, and leverages the maximizing-a-posteriori of latent alignment to solve the inconsistency problem that arises in distillation between CTC-based models. We refer to the regularized CTC loss as Distillation Connectionist Temporal Classification (DCTC) loss. DCTC loss is module-free, requiring no extra parameters, longer inference lag, or additional training data or phases. Extensive experiments on public benchmarks demonstrate that DCTC can boost text recognition model accuracy by up to 2.6%, without any of these drawbacks. Ziyin Zhang, Ning Lu 0003, Minghui Liao, Yongshuai Huang, Cheng Li 0040, Wei Peng 0011 |
AAAI | 3 |
| 2024 | Self-supervised Contrastive Graph Views for Learning Neuron-Level Circuit Network
Junchi Li, Guojia Wan, Minghui Liao, Bo Du 0001 |
MICCAI (11) | 3 |
| 2024 | Class-Aware Mask-guided feature refinement for scene text recognition
Minghui Liao, Yingying Zhu 0005, Xiang Bai |
Pattern Recognit. | 3 |
| 2024 | Sequential visual and semantic consistency for semi-supervised text recognition
Minghui Liao, Yingying Zhu 0005, Xiang Bai |
Pattern Recognit. Lett. | 3 |
| 2023 | Real-Time Scene Text Detection With Differentiable Binarization and Adaptive Scale FusionabstractRecently, segmentation-based scene text detection methods have drawn extensive attention in the scene text detection field, because of their superiority in detecting the text instances of arbitrary shapes and extreme aspect ratios, profiting from the pixel-level descriptions. However, the vast majority of the existing segmentation-based approaches are limited to their complex post-processing algorithms and the scale robustness of their segmentation models, where the post-processing algorithms are not only isolated to the model optimization but also time-consuming and the scale robustness is usually strengthened by fusing multi-scale feature maps directly. In this paper, we propose a Differentiable Binarization (DB) module that integrates the binarization process, one of the most important steps in the post-processing procedure, into a segmentation network. Optimized along with the proposed DB module, the segmentation network can produce more accurate results, which enhances the accuracy of text detection with a simple pipeline. Furthermore, an efficient Adaptive Scale Fusion (ASF) module is proposed to improve the scale robustness by fusing features of different scales adaptively. By incorporating the proposed DB and ASF with the segmentation network, our proposed scene text detector consistently achieves state-of-the-art results, in terms of both detection accuracy and speed, on five standard benchmarks. Minghui Liao, Zhisheng Zou, Zhaoyi Wan, Cong Yao, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Reading and Writing: Discriminative and Generative Modeling for Self-Supervised Text RecognitionabstractExisting text recognition methods usually need large-scale training data. Most of them rely on synthetic training data due to the lack of annotated real images. However, there is a domain gap between the synthetic data and real data, which limits the performance of the text recognition models. Recent self-supervised text recognition methods attempted to utilize unlabeled real images by introducing contrastive learning, which mainly learns the discrimination of the text images. Inspired by the observation that humans learn to recognize the texts through both reading and writing, we propose to learn discrimination and generation by integrating contrastive learning and masked image modeling in our self-supervised method. The contrastive learning branch is adopted to learn the discrimination of text images, which imitates the reading behavior of humans. Meanwhile, masked image modeling is firstly introduced for text recognition to learn the context generation of the text images, which is similar to the writing behavior. The experimental results show that our method outperforms previous self-supervised text recognition methods by 10.2%-20.2% on irregular scene text recognition datasets. Moreover, our proposed text recognizer exceeds previous state-of-the-art text recognition methods by averagely 5.3% on 11benchmarks, with similar model size. We also demonstrate that our pre-trained model can be easily applied to other text-related tasks with obvious performance gain. Minghui Liao, Pu Lu, Jing Wang 0221, Shenggao Zhu, Hualin Luo, Qi Tian 0001, Xiang Bai |
ACM Multimedia | 2 |
| 2022 | Comprehensive benchmark datasets for Amharic scene text detection and recognition
Wondimu Dikubab, Dingkang Liang, Minghui Liao, Xiang Bai |
Sci. China Inf. Sci. | 3 |
| 2021 | MOST: A Multi-Oriented Scene Text Detector With Localization RefinementabstractOver the past few years, the field of scene text detection has progressed rapidly that modern text detectors are able to hunt text in various challenging scenarios. However, they might still fall short when handling text instances of extreme aspect ratios and varying scales. To tackle such difficulties, we propose in this paper a new algorithm for scene text detection, which puts forward a set of strategies to significantly improve the quality of text localization. Specifically, a Text Feature Alignment Module (TFAM) is proposed to dynamically adjust the receptive fields of features based on initial raw detections; a Position-Aware Non-Maximum Suppression (PA-NMS) module is devised to selectively concentrate on reliable raw detections and exclude unreliable ones; besides, we propose an Instance-wise IoU loss for balanced training to deal with text instances of different scales. An extensive ablation study demonstrates the effectiveness and superiority of the proposed strategies. The resulting text detection system, which integrates the proposed strategies with a leading scene text detector EAST, achieves state-of-the-art or competitive performance on various standard benchmarks for text detection while keeping a fast running speed. Minghang He, Minghui Liao, Zhibo Yang 0003, Humen Zhong, Jun Tang 0008, Wenqing Cheng, Cong Yao, Yongpan Wang, Xiang Bai |
CVPR | 2 |
| 2021 | Scene Text Detection with Scribble Line
Yang Qiu 0002, Minghui Liao, Rui Zhang 0056, Xiaolin Wei, Xiang Bai |
ICDAR (4) | 3 |
| 2021 | Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary ShapesabstractUnifying text detection and text recognition in an end-to-end training fashion has become a new trend for reading text in the wild, as these two tasks are highly relevant and complementary. In this paper, we investigate the problem of scene text spotting, which aims at simultaneous text detection and recognition in natural images. An end-to-end trainable neural network named as Mask TextSpotter is presented. Different from the previous text spotters that follow the pipeline consisting of a proposal generation network and a sequence-to-sequence recognition network, Mask TextSpotter enjoys a simple and smooth end-to-end learning procedure, in which both detection and recognition can be achieved directly from two-dimensional space via semantic segmentation. Further, a spatial attention module is proposed to enhance the performance and universality. Benefiting from the proposed two-dimensional representation on both detection and recognition, it easily handles text instances of irregular shapes, for instance, curved text. We evaluate it on four English datasets and one multi-language dataset, achieving consistently superior performance over state-of-the-art methods in both detection and end-to-end text recognition tasks. Moreover, we further investigate the recognition module of our method separately, which significantly outperforms state-of-the-art methods on both regular and irregular text datasets for scene text recognition. Minghui Liao, Pengyuan Lv, Minghang He, Cong Yao, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Real-Time Scene Text Detection with Differentiable BinarizationabstractRecently, segmentation-based methods are quite popular in scene text detection, as the segmentation results can more accurately describe scene text of various shapes such as curve text. However, the post-processing of binarization is essential for segmentation-based detection, which converts probability maps produced by a segmentation method into bounding boxes/regions of text. In this paper, we propose a module named Differentiable Binarization (DB), which can perform the binarization process in a segmentation network. Optimized along with a DB module, a segmentation network can adaptively set the thresholds for binarization, which not only simplifies the post-processing but also enhances the performance of text detection. Based on a simple segmentation network, we validate the performance improvements of DB on five benchmark datasets, which consistently achieves state-of-the-art results, in terms of both detection accuracy and speed. In particular, with a light-weight backbone, the performance improvements by DB are significant so that we can look for an ideal tradeoff between detection accuracy and efficiency. Specifically, with a backbone of ResNet-18, our detector achieves an F-measure of 82.8, running at 62 FPS, on the MSRA-TD500 dataset. Code is available at: https://github.com/MhLiao/DB. Minghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen 0006, Xiang Bai |
AAAI | 1 |
| 2020 | Mask TextSpotter v3: Segmentation Proposal Network for Robust Scene Text Spotting
Minghui Liao, Guan Pang, Jing Huang 0020, Tal Hassner, Xiang Bai |
ECCV (11) | 1 |
| 2020 | SynthText3D: synthesizing scene text images from 3D virtual worlds
Minghui Liao, Boyu Song, Shangbang Long, Minghang He, Cong Yao, Xiang Bai |
Sci. China Inf. Sci. | 1 |
| 2020 | Scanning Imaging Restoration of Moving or Dynamically Deforming ObjectsabstractThe raster scanning imaging mode is widely used in scanning electron microscopes (SEMs), transmission electron microscopes (TEM), and atomic force microscopes (AFM), and can achieve subatomic resolution. However, only a point on the shallow surface of an object can be imaged at one time using the raster scanning imaging mode, whereas the entire surface of the object can be imaged in the image plane once and instantaneously using the optical imaging mode, which is a parallel imaging mode. Therefore, the image distortion and blur for the scanning imaging mode are different from the optical imaging. In this paper, we propose a theory to describe the mechanism of the scanning imaging process and restore the degraded image (distorted and blurred image) obtained using an SEM. The theory consists of a scanning equation, motion or deformation equations, and an assumption called the intensity-invariant hypothesis. Numerical simulations of the scanning imaging process and restoration of the degraded images are performed using the scanning imaging formulas, spatial non-uniform point spread function, and inverse restoration algorithms, including algebraic, interpolation, and their hybrid methods to verify the feasibility of our theory. In situ experiments on uniform linear motion, uniaxial tensile, and fatigue were also conducted to demonstrate the validity and efficiency of the proposed scanning imaging theory and restoration methods. We anticipate that this imaging and restoration theory will enable the scanning imaging mode to be used in in situ dynamic imaging and for mechanical property measurement of materials. Hongfu Xie, Jiecun Liang, Minghui Liao, Xide Li |
IEEE Trans. Image Process. | 4 |
| 2019 | Scene Text Recognition from Two-Dimensional PerspectiveabstractInspired by speech recognition, recent state-of-the-art algorithms mostly consider scene text recognition as a sequence prediction problem. Though achieving excellent performance, these methods usually neglect an important fact that text in images are actually distributed in two-dimensional space. It is a nature quite different from that of speech, which is essentially a one-dimensional signal. In principle, directly compressing features of text into a one-dimensional form may lose useful information and introduce extra noise. In this paper, we approach scene text recognition from a two-dimensional perspective. A simple yet effective model, called Character Attention Fully Convolutional Network (CA-FCN), is devised for recognizing the text of arbitrary shapes. Scene text recognition is realized with a semantic segmentation network, where an attention mechanism for characters is adopted. Combined with a word formation module, CA-FCN can simultaneously recognize the script and predict the position of each character. Experiments demonstrate that the proposed algorithm outperforms previous methods on both regular and irregular text datasets. Moreover, it is proven to be more robust to imprecise localizations in the text detection phase, which are very common in practice. Minghui Liao, Zhaoyi Wan, Fengming Xie, Jiajun Liang, Pengyuan Lv, Cong Yao, Xiang Bai |
AAAI | 1 |
| 2019 | Symmetry-Constrained Rectification Network for Scene Text RecognitionabstractReading text in the wild is a very challenging task due to the diversity of text instances and the complexity of natural scenes. Recently, the community has paid increasing attention to the problem of recognizing text instances with irregular shapes. One intuitive and effective way to handle this problem is to rectify irregular text to a canonical form before recognition. However, these methods might struggle when dealing with highly curved or distorted text instances. To tackle this issue, we propose in this paper a Symmetry-constrained Rectification Network (ScRN) based on local attributes of text instances, such as center line, scale and orientation. Such constraints with an accurate description of text shape enable ScRN to generate better rectification results than existing methods and thus lead to higher recognition accuracy. Our method achieves state-of-the-art performance on text with both regular and irregular shapes. Specifically, the system outperforms existing algorithms by a large margin on datasets that contain quite a proportion of irregular text instances, e.g., ICDAR 2015, SVT-Perspective and CUTE80. Yushuo Guan, Minghui Liao, Kaigui Bian, Song Bai 0001, Cong Yao, Xiang Bai |
ICCV | 3 |
| 2019 | ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on SignboardabstractChinese scene text reading is one of the most challenging problems in computer vision and has attracted great interest. Different from English text, Chinese has more than 6000 commonly used characters and Chinese characters can be arranged in various layouts with numerous fonts. The Chinese signboards in street view are a good choice for Chinese scene text images since they have different backgrounds, fonts and layouts. We organized a competition called ICDAR2019-ReCTS, which mainly focuses on reading Chinese text on signboard. This report presents the final results of the competition. A large-scale dataset of 25,000 annotated signboard images, in which all the text lines and characters are annotated with locations and transcriptions, were released. Four tasks, namely character recognition, text line recognition, text line detection and end-to-end recognition were set up. Besides, considering the Chinese text ambiguity issue, we proposed a multi ground truth (multi-GT) evaluation method to make evaluation fairer. The competition started on March 1, 2019 and ended on April 30, 2019. 262 submissions from 46 teams are received. Most of the participants come from universities, research institutes, and tech companies in China. There are also some participants from the United States, Australia, Singapore, and Korea. 21 teams submit results for Task 1, 23 teams submit results for Task 2, 24 teams submit results for Task 3, and 13 teams submit results for Task 4. The official website for the competition is http://rrc.cvc.uab.es/?ch=12. Rui Zhang 0056, Xiang Bai, Baoguang Shi, Dimosthenis Karatzas, Shijian Lu, C. V. Jawahar, Yongsheng Zhou, Qianyi Jiang, Nan Li 0071, Dong Wang 0004, Minghui Liao |
ICDAR | 15 |
| 2018 | Rotation-Sensitive Regression for Oriented Scene Text DetectionabstractText in natural images is of arbitrary orientations, requiring detection in terms of oriented bounding boxes. Normally, a multi-oriented text detector often involves two key tasks: 1) text presence detection, which is a classification problem disregarding text orientation; 2) oriented bounding box regression, which concerns about text orientation. Previous methods rely on shared features for both tasks, resulting in degraded performance due to the incompatibility of the two tasks. To address this issue, we propose to perform classification and regression on features of different characteristics, extracted by two network branches of different designs. Concretely, the regression branch extracts rotation-sensitive features by actively rotating the convolutional filters, while the classification branch extracts rotation-invariant features by pooling the rotation-sensitive features. The proposed method named Rotation-sensitive Regression Detector (RRD) achieves state-of-the-art performance on several oriented scene text benchmark datasets, including ICDAR 2015, MSRA-TD500, RCTW-17, and COCO-Text. Furthermore, RRD achieves a significant improvement on a ship collection dataset, demonstrating its generality on oriented object detection. Minghui Liao, Zhen Zhu 0006, Baoguang Shi, Gui-Song Xia, Xiang Bai |
CVPR | 1 |
| 2018 | Feature Fusion for Scene Text DetectionabstractA significant challenge in scene text detection is the large variation in text sizes. In particular, small text are usually hard to detect. This paper presents an accurate oriented text detector based on Faster R-CNN. We observe that Faster R-CNN is suitable for general object detection but inadequate for scene text detection due to the large variation in text size. We apply feature fusion both in RPN and Fast R-CNN to alleviate this problem and furthermore, enhance model's ability to detect relatively small text. Our text detector achieves comparable results to those state of the art methods on ICDAR 2015 and MSRA-TD500, showing its advantage and applicability. Zhen Zhu 0006, Minghui Liao, Baoguang Shi, Xiang Bai |
DAS | 2 |
| 2018 | Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
Pengyuan Lv, Minghui Liao, Cong Yao, Xiang Bai |
ECCV (14) | 2 |
| 2018 | TextBoxes++: A Single-Shot Oriented Scene Text DetectorabstractScene text detection is an important step of scene text recognition system and also a challenging problem. Different from general object detection, the main challenges of scene text detection lie on arbitrary orientations, small sizes, and significantly variant aspect ratios of text in natural images. In this paper, we present an end-to-end trainable fast scene text detector, named TextBoxes++, which detects arbitrary-oriented scene text with both high accuracy and efficiency in a single network forward pass. No post-processing other than an efficient non-maximum suppression is involved. We have evaluated the proposed TextBoxes++ on four public datasets. In all experiments, TextBoxes++ outperforms competing methods in terms of text localization accuracy and runtime. More specifically, TextBoxes++ achieves an f-measure of 0.817 at 11.6fps for 1024 × 1024 ICDAR 2015 Incidental text images, and an f-measure of 0.5591 at 19.8fps for 768 × 768 COCO-Text images. Furthermore, combined with a text recognizer, TextBoxes++ significantly outperforms the stateof-the-art approaches for word spotting and end-to-end text recognition tasks on popular benchmarks. Minghui Liao, Baoguang Shi, Xiang Bai |
IEEE Trans. Image Process. | 1 |
| 2018 | Cascaded Segmentation-Detection Networks for Text-Based Traffic Sign DetectionabstractIn this paper, we propose a novel text-based traffic sign detection framework with two deep learning components. More precisely, we apply a fully convolutional network to segment candidate traffic sign areas providing candidate regions of interest (RoI), followed by a fast neural network to detect texts on the extracted RoI. The proposed method makes full use of the characteristics of traffic signs to improve the efficiency and accuracy of text detection. On one hand, the proposed two-stage detection method reduces the search area of text detection and removes texts outside traffic signs. On the other hand, it solves the problem of multi-scales for the text detection part to a large extent. Extensive experimental results show that the proposed method achieves the state-of-the-art results on the publicly available traffic sign data set: Traffic Guide Panel data set. In addition, we collect a data set of text-based traffic signs including Chinese and English traffic signs. Our method also performs well on this data set, which demonstrates that the proposed method is general in detecting traffic signs of different languages. Yingying Zhu 0005, Minghui Liao, Wenyu Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | TextBoxes: A Fast Text Detector with a Single Deep Neural NetworkabstractThis paper presents an end-to-end trainable fast scene text detector, named TextBoxes, which detects scene text with both high accuracy and efficiency in a single network forward pass, involving no post-process except for a standard non-maximum suppression. TextBoxes outperforms competing methods in terms of text localization accuracy and is much faster, taking only 0.09s per image in a fast implementation. Furthermore, combined with a text recognizer, TextBoxes significantly outperforms state-of-the-art approaches on word spotting and end-to-end text recognition tasks. Minghui Liao, Baoguang Shi, Xiang Bai, Xinggang Wang, Wenyu Liu 0001 |
AAAI | 1 |
| 2017 | ICDAR2017 Competition on Reading Chinese Text in the Wild (RCTW-17)abstractChinese is the most widely used language in the world. Algorithms that read Chinese text in natural images facilitate applications of various kinds. Despite the large potential value, datasets and competitions in the past primarily focus on English, which bares very different characteristics than Chinese. This report introduces RCTW, a new competition that focuses on Chinese text reading. The competition features a large-scale dataset with over 12,000 annotated images. Two tasks, namely text localization and end-to-end recognition, are set up. The competition took place from January 20 to May 31, 2017. 23 valid submissions were received from 19 teams. This report includes dataset description, task definitions, evaluation protocols, and results summaries and analysis. Through this competition, we call for more future research on the Chinese text reading problem. Baoguang Shi, Cong Yao, Minghui Liao, Pei Xu 0006, Linyan Cui, Serge J. Belongie, Shijian Lu, Xiang Bai |
ICDAR | 3 |