VLDB 2026 Research / reviewers in the wild / expert
Zhibo Yang 0003
dblp:127/9460-3
· DBLP profile ↗
27ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0003-2343-7750ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 4 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual DocumentsabstractYifan Ji, Zhipeng Xu, Zhenghao Liu, Zulong Chen, Qian Zhang, Zhibo Yang, Junyang Lin, Yu Gu, Ge Yu, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yifan Ji, Zhenghao Liu 0001, Zulong Chen, Zhibo Yang 0003, Junyang Lin, Yu Gu 0002, Ge Yu 0001, Maosong Sun 0001 |
ACL (1) | 6 |
| 2026 | OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language ModelsabstractVisually-situated text parsing (VsTP) has recently seen notable advancements, driven by the growing demand for automated document understanding and the emergence of large language models capable of processing document-based questions. While various methods have been proposed to tackle the complexities of VsTP, existing solutions often rely on task-specific architectures and objectives for individual tasks. This leads to modal isolation and complex workflows due to the diversified targets and heterogeneous schemas. In this paper, we introduce OmniParser V2, a universal model that unifies VsTP typical tasks, including text spotting, key information extraction, table recognition, and layout analysis, into a unified framework. Central to our approach is the proposed Structured-Points-of-Thought (SPOT) prompting schemas, which improves model performance across diverse scenarios by leveraging a unified encoder-decoder architecture, objective, and input&output representation. SPOT eliminates the need for task-specific architectures and loss functions, significantly simplifying the processing pipeline. Our extensive evaluations across four tasks on eight different datasets show that OmniParser V2 achieves state-of-the-art or competitive results in VsTP. Additionally, we explore the integration of SPOT within a multimodal large language model structure, further enhancing visual text parsing capabilities on four tasks, thereby confirming the generality of SPOT prompting technique. Wenwen Yu, Zhibo Yang 0003, Jianqiang Wan, Sibo Song, Jun Tang 0008, Wenqing Cheng, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | SceneVTG++: Controllable Multilingual Visual Text Generation in the WildabstractGenerating visual text in natural scene images is a challenging task with many unsolved problems. Different from generating text on artificially designed images (such as posters, covers, and cartoons), existing methods for natural scene visual text generation still have significant deficiencies: methods based on rendering engines rely on manually crafted rules, which struggle to adapt to diverse backgrounds and leave obvious artificial traces, while their text layouts may be placed in unreasonable areas (e.g., sky or ground) and text content is semantically disconnected from the scene; diffusion model-based methods, on the other hand, face difficulties in generating small characters, depend on manually designed prompts to ensure reasonable layout and content, fail to generate text at precise locations, and cannot effectively control text attributes (e.g., font and color). In this paper, we propose a two-stage method named SceneVTG++ to address these issues. SceneVTG++ comprises two core components: a Text Layout and Content Generator (TLCG) and a Controllable Local Text Diffusion (CLTD). The former leverages the world knowledge and visual reasoning capabilities of multimodal large language models to identify reasonable text areas and recommend scene-relevant text content based on natural scene background images; the latter generates controllable multilingual text using a diffusion model, ensuring alignment with the outputs of TLCG. Through extensive experiments, we verified the effectiveness of both TLCG and CLTD, and demonstrated that SceneVTG++ achieves state-of-the-art performance in natural scene visual text generation. Additionally, the images generated by SceneVTG++ exhibit superior utility for training natural scene optical character recognition (OCR) tasks, including text detection and text recognition. Codes and datasets will be made publicly available. Jiawei Liu 0006, Feiyu Gao, Zhibo Yang 0003, Peng Wang 0028, Junyang Lin, Xinggang Wang, Wenyu Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in LiteracyabstractLarge Multimodal Models (LMMs) have demonstrated impressive performance in recognizing document images with natural language instructions. However, it remains unclear to what extent capabilities in literacy with rich structure and fine-grained visual challenges. The current landscape lacks a comprehensive benchmark to effectively measure the literate capabilities of LMMs. Existing benchmarks are often limited by narrow scenarios and specified tasks. To this end, we introduce CC-OCR, a comprehensive benchmark that possesses a diverse range of scenarios, tasks, and challenges. CC-OCR comprises four OCR-centric tracks: multi-scene text reading, multilingual text reading, document parsing, and key information extraction. It includes 39 subsets with 7,058 full annotated images, of which 41% are sourced from real applications, and released for the first time. We evaluate nine prominent LMMs and reveal both the strengths and weaknesses of these models, particularly in text grounding, multi-orientation, and hallucination of repetition. CC-OCR aims to comprehensively evaluate the capabilities of LMMs on OCR-centered tasks, facilitating continued progress in this crucial area. Zhibo Yang 0003, Jun Tang 0008, Zhaohai Li, Jianqiang Wan, Humen Zhong, Xuejing Liu, Peng Wang 0028, Shuai Bai, Junyang Lin |
ICCV | 1 |
| 2025 | DocThinker: Explainable Multimodal Large Language Models with Rule-Based Reinforcement Learning for Document UnderstandingabstractMultimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in document understanding. However, their reasoning processes remain largely black-box, making it difficult to ensure reliability and trustworthiness, especially in high-stakes domains such as legal, financial, and medical document analysis. Existing methods use fixed Chain-of-Thought (CoT) reasoning with supervised fine-tuning (SFT) but suffer from catastrophic forgetting, poor adaptability, and limited generalization across domain tasks. In this paper, we propose DocThinker, a rule-based Reinforcement Learning (RL) framework for dynamic inference-time reasoning. Instead of relying on static CoT templates, DocThinker autonomously refines reasoning strategies via policy learning, generating explainable intermediate results, including structured reasoning processes, rephrased questions, regions of interest (RoI) supporting the answer, and the final answer. By integrating multi-objective rule-based rewards and KL-constrained optimization, our method mitigates catastrophic forgetting and enhances both adaptability and transparency. Extensive experiments on multiple benchmarks demonstrate that DocThinker significantly improves generalization while producing more explainable and human-understandable reasoning steps. Our findings highlight RL as a powerful alternative for enhancing explainability and adaptability in MLLM-based document understanding. Code will be available at https://github.com/wenwenyu/DocThinker. Wenwen Yu, Zhibo Yang 0003, Xiang Bai |
ICCV | 2 |
| 2025 | HIP: Hierarchical Point Modeling and Pre-training for Visual Information Extraction
Rujiao Long, Zhibo Yang 0003, Wenqing Cheng |
ICDAR (1) | 3 |
| 2025 | LORE++: Logical location regression network for table structure recognition with pre-training
Rujiao Long, Hangdi Xing, Zhibo Yang 0003, Qi Zheng 0002, Fei Huang 0002, Cong Yao |
Pattern Recognit. | 3 |
| 2025 | Generative compositor for few-shot visual information extractionabstractVisual Information Extraction (VIE), aiming at extracting structured information from visually rich document images , plays a pivotal role in document processing. Considering various layouts, semantic scopes, and languages, VIE encompasses an extensive range of types, potentially numbering in the thousands. However, many of these types suffer from a lack of training data , which poses significant challenges. In this paper, we propose a novel generative model , named Generative Compositor, to address the challenge of few-shot VIE. The Generative Compositor is a hybrid pointer-generator network that emulates the operations of a compositor by retrieving words from the source text and assembling them based on the provided prompts. Furthermore, three pre-training strategies are employed to enhance the model’s perception of spatial context information. Besides, a prompt-aware resampler is specially designed to enable efficient matching by leveraging the entity-semantic prior contained in prompts. The introduction of the prompt-based retrieval mechanism and the pre-training strategies enable the model to acquire more effective spatial and semantic clues with limited training samples . Experiments demonstrate that the proposed method achieves highly competitive results in the full-sample training, while notably outperforms the baseline in the 1-shot, 5-shot, and 10-shot settings. Zhibo Yang 0003, Wei Hua 0005, Sibo Song, Cong Yao, Yingying Zhu 0005, Wenqing Cheng, Xiang Bai |
Pattern Recognit. | 1 |
| 2025 | HierCode: A lightweight hierarchical codebook for zero-shot Chinese text recognition
Yuyi Zhang 0002, Dezhi Peng, Peirong Zhang 0001, Zhenhua Yang, Zhibo Yang 0003, Cong Yao |
Pattern Recognit. | 6 |
| 2024 | OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table RecognitionabstractRecently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Generative Large Language Models (LLMs) capable of processing document-based questions. Various methods have been proposed to address the challenging problem of VsTP. However, due to the diversified targets and heterogeneous schemas, previous works usually design task-specific architectures and objectives for individual tasks, which in- advertently leads to modal isolation and complex workflow. In this paper, we propose a unified paradigm for parsing visually-situated text across diverse scenarios. Specifically, we devise a universal model, called OmniParser, which can simultaneously handle three typical visually-situated text parsing tasks: text spotting, key information extraction, and table recognition. In OmniParser, all tasks share the unified encoder-decoder architecture, the unified objective: point- conditioned text generation, and the unified input&output representation: prompt & structured sequences. Extensive experiments demonstrate that the proposed OmniParser achieves state-of-the-art (SOTA) or highly competitive performances on 7 datasets for the three visually-situated text parsing tasks, despite its unified, concise design. The code is available at AdvancedLiterateMachinery. Jianqiang Wan, Sibo Song, Wenwen Yu, Wenqing Cheng, Fei Huang 0002, Xiang Bai, Cong Yao, Zhibo Yang 0003 |
CVPR | 9 |
| 2024 | Platypus: A Generalized Specialist Model for Reading Text in Various Forms
Peng Wang 0028, Zhaohai Li, Jun Tang 0008, Humen Zhong, Fei Huang 0002, Zhibo Yang 0003, Cong Yao |
ECCV (35) | 6 |
| 2024 | Visual Text Generation in the Wild
Jiawei Liu 0006, Feiyu Gao, Wenyu Liu 0001, Xinggang Wang, Peng Wang 0028, Fei Huang 0002, Cong Yao, Zhibo Yang 0003 |
ECCV (53) | 9 |
| 2024 | VL-Reader: Vision and Language Reconstructor is an Effective Scene Text RecognizerabstractText recognition is an inherent integration of vision and language, encompassing the visual texture in stroke patterns and the semantic context among the character sequences. Towards advanced text recognition, there are three key challenges: (1) an encoder capable of representing the visual and semantic distributions; (2) a decoder that ensures the alignment between vision and semantics; and (3) consistency in the framework during pre-training, if it exists, and fine-tuning. Inspired by masked autoencoding, a successful pre-training strategy in both vision and language, we propose an innovative scene text recognition approach, named VL-Reader. The novelty of the VL-Reader lies in the pervasive interplay between vision and language throughout the entire process. Concretely, we first introduce a Masked Visual-Linguistic Reconstruction (MVLR) objective, which aims at simultaneously modeling visual and linguistic information. Then, we design a Masked Visual-Linguistic Decoder (MVLD) to further leverage masked vision-language context and achieve bi-modal feature interaction. The architecture of VL-Reader maintains consistency from pre-training to fine-tuning. In the pre-training stage, VL-Reader reconstructs both masked visual and text tokens, while in the fine-tuning stage, the network degrades to reconstruct all characters from an image without any masked regions. VL-reader achieves an average accuracy of 97.1% on six typical datasets, surpassing the SOTA by 1.1%. The improvement was even more significant on challenging datasets. The results demonstrate that vision and language reconstructor can serve as an effective scene text recognizer. Humen Zhong, Zhibo Yang 0003, Zhaohai Li, Peng Wang 0028, Jun Tang 0008, Wenqing Cheng, Cong Yao |
ACM Multimedia | 2 |
| 2023 | Modeling Entities as Semantic Points for Visual Information Extraction in the WildabstractRecently, Visual Information Extraction (VIE) has been becoming increasingly important in both the academia and industry, due to the wide range of real-world applications. Previously, numerous works have been proposed to tackle this problem. However, the benchmarks used to assess these methods are relatively plain, i.e., scenarios with real-world complexity are not fully represented in these benchmarks. As the first contribution of this work, we curate and release a new dataset for VIE, in which the document images are much more challenging in that they are taken from real applications, and difficulties such as blur, partial occlusion, and printing shift are quite common. All these factors may lead to failures in information extraction. Therefore, as the second contribution, we explore an alternative approach to precisely and robustly extract key information from document images under such tough conditions. Specifically, in contrast to previous methods, which usually either incorporate visual information into a multi-modal architecture or train text spotting and information extraction in an end-to-end fashion, we explicitly model entities as semantic points, i.e., center points of entities are enriched with semantic information describing the attributes and relationships of different entities, which could largely benefit entity labeling and linking. Extensive experiments on standard benchmarks in this field as well as the proposed dataset demonstrate that the proposed method can achieve significantly enhanced performance on entity labeling and linking, compared with previous state-of-the-art models. Dataset is available at https://www.modelscope.cn/datasets/damo/SIBR/summary. Zhibo Yang 0003, Rujiao Long, Sibo Song, Humen Zhong, Wenqing Cheng, Xiang Bai, Cong Yao |
CVPR | 1 |
| 2023 | ICDAR 2023 Competition on Born Digital Video Text Question Answering
Zhibo Yang 0003, Xiaoge Song, Sibo Song, Tong Lu 0002, Xiang Bai, Cheng-Lin Liu 0001, Fei Huang 0002, Cong Yao |
ICDAR (2) | 1 |
| 2022 | Revisiting Document Image Dewarping by Grid RegularizationabstractThis paper addresses the problem of document image dewarping, which aims at eliminating the geometric distortion in document images for document digitization. Instead of designing a better neural network to approximate the optical flow fields between the inputs and outputs, we pursue the best readability by taking the text lines and the document boundaries into account from a constrained optimization perspective. Specifically, our proposed method first learns the boundary points and the pixels in the text lines and then follows the most simple observation that the boundaries and text lines in both horizontal and vertical directions should be kept after dewarping to introduce a novel grid regularization scheme. To obtain the final forward mapping for dewarping, we solve an optimization problem with our proposed grid regularization. The experiments comprehensively demonstrate that our proposed approach outperforms the prior arts by large margins in terms of readability (with the metrics of Character Errors Rate and the Edit Distance) while maintaining the best image quality on the publicly-available DocUNet benchmark. Xiangwei Jiang, Rujiao Long, Nan Xue 0001, Zhibo Yang 0003, Cong Yao, Gui-Song Xia |
CVPR | 4 |
| 2022 | Vision-Language Pre-Training for Boosting Scene Text DetectorsabstractRecently, vision-language joint representation learning has proven to be highly effective in various scenarios. In this paper, we specifically adapt vision-language joint learning for scene text detection, a task that intrinsically involves cross-modal interaction between the two modalities: vision and language, since text is the written form of language. Concretely, we propose to learn contextualized, joint representations through vision-language pretraining, for the sake of enhancing the performance of scene text detectors. Towards this end, we devise a pre-training architecture with an image encoder, a text encoder and a cross-modal encoder, as well as three pretext tasks: image-text contrastive learning (ITC), masked language modeling (MLM) and word-in-image prediction (WIP). The pretrained model is able to produce more informative representations with richer semantics, which could readily benefit existing scene text detectors (such as EAST and PSENet) in the down-stream text detection task. Extensive experiments on standard benchmarks demonstrate that the proposed paradigm can significantly improve the performance of various representative text detectors, outperforming previous pre-training approaches. The code and pre-trained models will be publicly released. Sibo Song, Jianqiang Wan, Zhibo Yang 0003, Jun Tang 0008, Wenqing Cheng, Xiang Bai, Cong Yao |
CVPR | 3 |
| 2022 | PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped TextabstractScene text detection and recognition have been well explored in the past few years. Despite the progress, efficient and accurate end-to-end spotting of arbitrarily-shaped text remains challenging. In this work, we propose an end-to-end text spotting framework, termed PAN++, which can efficiently detect and recognize text of arbitrary shapes in natural scenes. PAN++ is based on the kernel representation that reformulates a text line as a text kernel (central region) surrounded by peripheral pixels. By systematically comparing with existing scene text representations, we show that our kernel representation can not only describe arbitrarily-shaped text but also well distinguish adjacent text. Moreover, as a pixel-based representation, the kernel representation can be predicted by a single fully convolutional network, which is very friendly to real-time applications. Taking the advantages of the kernel representation, we design a series of components as follows: 1) a computationally efficient feature enhancement network composed of stacked Feature Pyramid Enhancement Modules (FPEMs); 2) a lightweight detection head cooperating with Pixel Aggregation (PA); and 3) an efficient attention-based recognition head with Masked RoI. Benefiting from the kernel representation and the tailored components, our method achieves high inference speed while maintaining competitive accuracy. Extensive experiments show the superiority of our method. For example, the proposed PAN++ achieves an end-to-end text spotting F-measure of 64.9 at 29.2 FPS on the Total-Text dataset, which significantly outperforms the previous best method. Code will be available at: git.io/PAN. Wenhai Wang, Enze Xie, Xiang Li 0041, Xuebo Liu 0001, Ding Liang, Zhibo Yang 0003, Tong Lu 0002, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | MOST: A Multi-Oriented Scene Text Detector With Localization RefinementabstractOver the past few years, the field of scene text detection has progressed rapidly that modern text detectors are able to hunt text in various challenging scenarios. However, they might still fall short when handling text instances of extreme aspect ratios and varying scales. To tackle such difficulties, we propose in this paper a new algorithm for scene text detection, which puts forward a set of strategies to significantly improve the quality of text localization. Specifically, a Text Feature Alignment Module (TFAM) is proposed to dynamically adjust the receptive fields of features based on initial raw detections; a Position-Aware Non-Maximum Suppression (PA-NMS) module is devised to selectively concentrate on reliable raw detections and exclude unreliable ones; besides, we propose an Instance-wise IoU loss for balanced training to deal with text instances of different scales. An extensive ablation study demonstrates the effectiveness and superiority of the proposed strategies. The resulting text detection system, which integrates the proposed strategies with a leading scene text detector EAST, achieves state-of-the-art or competitive performance on various standard benchmarks for text detection while keeping a fast running speed. Minghang He, Minghui Liao, Zhibo Yang 0003, Humen Zhong, Jun Tang 0008, Wenqing Cheng, Cong Yao, Yongpan Wang, Xiang Bai |
CVPR | 3 |
| 2021 | Parsing Table Structures in the WildabstractThis paper tackles the problem of table structure parsing (TSP) from images in the wild. In contrast to existing studies that mainly focus on parsing well-aligned tabular images with simple layouts from scanned PDF documents, we aim to establish a practical table structure parsing system for real-world scenarios where tabular input images are taken or scanned with severe deformation, bending or occlusions. For designing such a system, we propose an approach named Cycle-CenterNet on the top of CenterNet with a novel cycle-pairing module to simultaneously detect and group tabular cells into structured tables. In the cycle-pairing module, a new pairing loss function is proposed for the network training. Alongside with our Cycle-CenterNet, we also present a large-scale dataset, named Wired Table in the Wild (WTW), which includes well-annotated structure parsing of multiple style tables in several scenes like photo, scanning files, web pages, etc.. In experiments, we demonstrate that our Cycle-CenterNet consistently achieves the best accuracy of table structure parsing on the new WTW dataset by 24.6% absolute improvement evaluated by the TEDS metric. A more comprehensive experimental analysis also validates the advantages of our proposed methods for the TSP task. Rujiao Long, Nan Xue 0001, Feiyu Gao, Zhibo Yang 0003, Yongpan Wang, Gui-Song Xia |
ICCV | 5 |
| 2021 | Arbitrarily-Oriented Text Detection in Low Light Natural Scene ImagesabstractText detection in low light natural scene images is challenging due to poor image quality and low contrast. Unlike most existing methods that focus on well-lit (normally daylight) images, the proposed method considers much darker natural scene images. For this task, our method first integrates spatial and frequency domain features through fusion to enhance fine details in the image. Next, we use Maximally Stable Extremal Regions (MSER) for detecting text candidates from the enhanced images. We then introduce Cloud of Line Distribution (COLD) features, which capture the distribution of pixels of text candidates in the polar domain. The extracted features are sent to a Convolution Neural Network (CNN) to correct the bounding boxes for arbitrarily oriented text lines by removing false positives. Experiments are conducted on a dataset of low light images to evaluate the proposed enhancement step. The results show our approach is more effective compared to existing methods in terms of standard quality measures, namely, BRISQE, NIQE and PIQE. In addition, experimental results on a variety of standard benchmark datasets, namely, ICDAR 2013, ICDAR 2015, SVT, Total-Text, ICDAR 2017-MLT and CTW1500, show that the proposed approach not only produces better results for low light images, at the same time it is also competitive for daylight images. Minglong Xue, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti, Zhibo Yang 0003 |
IEEE Trans. Multim. | 8 |
| 2020 | AE TextSpotter: Learning Visual and Linguistic Representation for Ambiguous Text Spotting
Wenhai Wang, Xuebo Liu 0001, Xiaozhong Ji, Enze Xie, Ding Liang, Zhibo Yang 0003, Tong Lu 0002, Chunhua Shen, Ping Luo 0002 |
ECCV (14) | 6 |
| 2019 | SegLink++: Detecting Dense and Arbitrary-shaped Scene Text by Instance-aware Component Grouping
Jun Tang 0008, Zhibo Yang 0003, Yongpan Wang, Qi Zheng 0002, Yongchao Xu, Xiang Bai |
Pattern Recognit. | 2 |
| 2019 | TextField: Learning a Deep Direction Field for Irregular Scene Text DetectionabstractScene text detection is an important step in the scene text reading system. The main challenges lie in significantly varied sizes and aspect ratios, arbitrary orientations, and shapes. Driven by the recent progress in deep learning, impressive performances have been achieved for multi-oriented text detection. Yet, the performance drops dramatically in detecting the curved texts due to the limited text representation (e.g., horizontal bounding boxes, rotated rectangles, or quadrilaterals). It is of great interest to detect the curved texts, which are actually very common in natural scenes. In this paper, we present a novel text detector named TextField for detecting irregular scene texts. Specifically, we learn a direction field pointing away from the nearest text boundary to each text point. This direction field is represented by an image of 2D vectors and learned via a fully convolutional neural network. It encodes both binary text mask and direction information used to separate adjacent text instances, which is challenging for the classical segmentation-based approaches. Based on the learned direction field, we apply a simple yet effective morphological-based post-processing to achieve the final detection. The experimental results show that the proposed TextField outperforms the state-of-the-art methods by a large margin (28% and 8%) on two curved text datasets: Total-Text and SCUT-CTW1500, respectively; TextField also achieves very competitive performance on multi-oriented datasets: ICDAR 2015 and MSRA-TD500. Furthermore, TextField is robust in generalizing unseen datasets. Yongchao Xu, Yukang Wang, Wei Zhou 0068, Yongpan Wang, Zhibo Yang 0003, Xiang Bai |
IEEE Trans. Image Process. | 5 |
| 2018 | ICPR2018 Contest on Robust Reading for Multi-Type Web ImagesabstractElectronic commerce has infiltrated every aspect of our daily lives, which offers great convenience for shopping, advertising, etc. Text in the web images is responsible to convey essential information for consumers. Algorithms that read text in these web images can facilitate applications of various types, such as goods surveillance, products classification, and intelligent retrieval or recommendation. Despite of various existing text reading tasks, this contest introduces a novel large-scale dataset named MTWI that contains 20,000 images, which is the first dataset that is mainly constructed by Chinese and English web text. Three tasks (web text recognition, web text detection, and end-to-end web text detection and recognition) were set up for encouraging more research on the web text reading problem. The contest was held from February 2, 2018 to May 26, 2018 with 289 valid submissions from 4,282 registered teams. Throughout this report, we describe the details of this new dataset, the purposes and definitions of the tasks, the evaluation protocols, and the summaries of the results. Mengchao He, Zhibo Yang 0003, Sheng Zhang 0024, Canjie Luo, Feiyu Gao, Qi Zheng 0002, Yongpan Wang, Xin Zhang 0013 |
ICPR | 3 |
| 2014 | Effective Palm Tracking with Integrated Tracker and Offline Detector
Zhibo Yang 0003, Bo Yuan 0003 |
ICIC (2) | 1 |
| 2013 | Vision Based Multi-pedestrian Tracking Using Adaptive Detection and Clustering
Zhibo Yang 0003, Bo Yuan 0003 |
IDEAL | 1 |