Mengchao He

dblp:231/1790 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
8since 2021 · last 2024
0009-0004-3759-4727ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2024 SideNet: Learning representations from interactive side information for zero-shot Chinese character recognition
Dezhi Peng, Mengchao He
Pattern Recognit.4
2023 Building A Mobile Text Recognizer via Truncated SVD-based Knowledge Distillation-Guided NAS
Weifeng Lin, Canyu Xie, Dezhi Peng, Cong Yao, Mengchao He
BMVC8
2023 Conditional Text Image Generation with Diffusion Models
abstract
Current text recognition systems, including those for handwritten scripts and scene text, have relied heavily on image synthesis and augmentation, since it is difficult to realize real-world complexity and diversity through collecting and annotating enough real text images. In this paper, we explore the problem of text image generation, by taking advantage of the powerful abilities of Diffusion Models in generating photo-realistic and diverse image samples with given conditions, and propose a method called Conditional Text Image Generation with Diffusion Models (CTIG-DM for short). To conform to the characteristics of text images, we devise three conditions: image condition, text condition, and style condition, which can be used to control the attributes, contents, and styles of the samples in the image generation process. Specifically, four text image generation modes, namely: (1) synthesis mode, (2) augmentation mode, (3) recovery mode, and (4) imitation mode, can be derived by combining and configuring these three conditions. Extensive experiments on both handwritten and scene text demonstrate that the proposed CTIG-DM is able to produce image samples that simulate real-world complexity and diversity, and thus can boost the performance of existing text recognizers. Besides, CTIG-DM shows its appealing potential in domain adaptation and generating images containing Out-Of-Vocabulary (OOV) words.
Zhaohai Li, Mengchao He, Cong Yao
CVPR4
2023 Read Ten Lines at One Glance: Line-Aware Semi-Autoregressive Transformer for Multi-Line Handwritten Mathematical Expression Recognition
abstract
Handwritten Mathematical Expression Recognition (HMER) plays a critical role in various applications, such as digitized education and scientific research. Although existing methods have achieved promising performance on publicly available datasets, they still struggle to recognize multi-line mathematical expressions (MEs), suffering from complex structures and slow inference speed. To address these issues, we propose a Line-Aware Semi-autoregressive Transformer (LAST) that treats multi-line mathematical expression sequences as two-dimensional dual-end structures. The proposed LAST utilizes a line-wise dual-end decoding strategy to decode multi-line mathematical expressions in parallel and perform dual-end decoding within each line. Specifically, we introduce a line-aware positional encoding module and a line-partitioned dual-end mask to endow LAST with line order awareness and directionality. Additionally, we adopt a shared-task optimization strategy to train LAST in both autoregressive and semi-autoregressive tasks. To evaluate the effectiveness of our approach in real-world scenarios, we have built a new Multi-line Mathematical Expression dataset (M2E), which, to the best of our knowledge, is the first of its kind and boasts with the largest character category, the largest samples of characters, and the longest average sequence length, compared to existing ME datasets. Experimental results on both the M2E dataset and publicly available datasets demonstrate the effectiveness of our proposed method. Notably, our semi-autoregressive decoding approach achieves significantly faster decoding speeds while still achieving state-of-the-art performance compared to the existing methods.
Wentao Yang 0003, Zhe Li 0046, Dezhi Peng, Mengchao He, Cong Yao
ACM Multimedia5
2022 TextSRNet: Scene Text Super-Resolution Based on Contour Prior and Atrous Convolution
abstract
Low-resolution (LR) scene text images are often encountered in real application scenarios. It is difficult to recognize these scene texts due to the loss of their character details. Super-resolution (SR) is an intuitive alternative to solve this problem. Unlike generic SR, which mainly focuses on image quality, text SR focuses on the accuracy of the downstream recognition tasks. Although several scene text SR methods have been recently proposed, they ignore retaining character contours that are meaningful for character readability and poorly perform on long texts. Therefore, we proposes a new scene text SR network called TextSRNet. In order to get fine character details, we adopt the segmentation maps of scene text images as the prior knowledge of character contours and embed it into the proposed TextSRNet. In addition, we incorporate Sobel loss to enhance the shape boundaries of characters. To improve the robustness of the model for long texts, we adopt a text atrous spatial pyramid pooling module. The proposed TextSRNet is trained and evaluated on TextZoom, a real-world scene text SR dataset. Extensive experiments demonstrate that our TextSRNet outperforms existing methods in terms of both recognition accuracy and image quality.
Jizhao Ma, Jiaxin Zhang 0003, Yang Xue 0001, Mengchao He
ICPR6
2022 AGTGAN: Unpaired Image Translation for Photographic Ancient Character Generation
abstract
The study of ancient writings has great value for archaeology and philology. Essential forms of material are photographic characters, but manual photographic character recognition is extremely time-consuming and expertise-dependent. Automatic classification is therefore greatly desired. However, the current performance is limited due to the lack of annotated data. Data generation is an inexpensive but useful solution to data scarcity. Nevertheless, the diverse glyph shapes and complex background textures of photographic ancient characters make the generation task difficult, leading to unsatisfactory results of existing methods. To this end, we propose an unsupervised generative adversarial network called AGTGAN in this paper. By explicitly modeling global and local glyph shape styles, followed by a stroke-aware texture transfer and an associate adversarial learning mechanism, our method can generate characters with diverse glyphs and realistic textures. We evaluate our method on photographic ancient character datasets, e.g., OBC306 and CSDD. Our method outperforms other state-of-the-art methods in terms of various metrics and performs much better in terms of the diversity and authenticity of generated samples. With our generated images, experiments on the largest photographic oracle bone character dataset show that our method can achieve a significant increase in classification accuracy, up to 16.34%. The source code is available at https://github.com/Hellomystery/AGTGAN.
Hongxiang Huang, Daihui Yang, Gang Dai 0002, Zhen Han 0003, Yuyi Wang 0001, Kin-Man Lam 0001, Fan Yang 0082, Shuangping Huang, Yongge Liu, Mengchao He
ACM Multimedia10
2021 Implicit Feature Alignment: Learn To Convert Text Recognizer to Text Spotter
abstract
Text recognition is a popular research subject with many associated challenges. Despite the considerable progress made in recent years, the text recognition task itself is still constrained to solve the problem of reading cropped line text images and serves as a subtask of optical character recognition (OCR) systems. As a result, the final text recognition result is limited by the performance of the text detector. In this paper, we propose a simple, elegant and effective paradigm called Implicit Feature Alignment (IFA), which can be easily integrated into current text recognizers, resulting in a novel inference mechanism called IFA- inference. This enables an ordinary text recognizer to process multi-line text such that text detection can be completely freed. Specifically, we integrate IFA into the two most prevailing text recognition streams (attention-based and CTC-based) and propose attention-guided dense prediction (ADP) and Extended CTC (ExCTC). Furthermore, the Wasserstein-based Hollow Aggregation Cross-Entropy (WH-ACE) is proposed to suppress negative predictions to assist in training ADP and ExCTC. We experimentally demonstrate that IFA achieves state-of-the-art performance on end-to-end document recognition tasks while maintaining the fastest speed, and ADP and ExCTC complement each other on the perspective of different application scenarios. Code will be available at https://github.com/Wang-Tianwei/Implicit-feature-alignment.
Dezhi Peng, Zhe Li 0046, Mengchao He, Yongpan Wang, Canjie Luo
CVPR6
2021 Context-Aware Selective Label Smoothing for Calibrating Sequence Recognition Model
abstract
Despite the success of deep neural network (DNN) on sequential data (i.e., scene text and speech) recognition, it suffers from the over-confidence problem mainly due to overfitting in training with the cross-entropy loss, which may make the decision-making less reliable. Confidence calibration has been recently proposed as one effective solution to this problem. Nevertheless, the majority of existing confidence calibration methods aims at non-sequential data, which is limited if directly applied to sequential data since the intrinsic contextual dependency in sequences or the class-specific statistical prior is seldom exploited. To the end, we propose a Context-Aware Selective Label Smoothing (CASLS) method for calibrating sequential data. The proposed CASLS fully leverages the contextual dependency in sequences to construct confusion matrices of contextual prediction statistics over different classes. Class-specific error rates are then used to adjust the weights of smoothing strength in order to achieve adaptive calibration. Experimental results on sequence recognition tasks, including scene text recognition and speech recognition, demonstrate that our method can achieve the state-of-the-art performance.
Shuangping Huang, Yu Luo 0007, Zhenzhou Zhuang, Jin-Gang Yu, Mengchao He, Yongpan Wang
ACM Multimedia5
2020 All You Need Is Boundary: Toward Arbitrary-Shaped Text Spotting
abstract
Recently, end-to-end text spotting that aims to detect and recognize text from cluttered images simultaneously has received particularly growing interest in computer vision. Different from the existing approaches that formulate text detection as bounding box extraction or instance segmentation, we localize a set of points on the boundary of each text instance. With the representation of such boundary points, we establish a simple yet effective scheme for end-to-end text spotting, which can read the text of arbitrary shapes. Experiments on three challenging datasets, including ICDAR2015, TotalText and COCO-Text demonstrate that the proposed method consistently surpasses the state-of-the-art in both scene text detection and end-to-end text recognition tasks.
Hao Wang 0207, Pu Lu, Hui Zhang 0085, Xiang Bai, Yongchao Xu, Mengchao He, Yongpan Wang, Wenyu Liu 0001
AAAI7
2020 RD-GAN: Few/Zero-Shot Chinese Character Style Transfer via Radical Decomposition and Rendering
Yaoxiong Huang, Mengchao He, Yongpan Wang
ECCV (6)2
2018 ICPR2018 Contest on Robust Reading for Multi-Type Web Images
abstract
Electronic commerce has infiltrated every aspect of our daily lives, which offers great convenience for shopping, advertising, etc. Text in the web images is responsible to convey essential information for consumers. Algorithms that read text in these web images can facilitate applications of various types, such as goods surveillance, products classification, and intelligent retrieval or recommendation. Despite of various existing text reading tasks, this contest introduces a novel large-scale dataset named MTWI that contains 20,000 images, which is the first dataset that is mainly constructed by Chinese and English web text. Three tasks (web text recognition, web text detection, and end-to-end web text detection and recognition) were set up for encouraging more research on the web text reading problem. The contest was held from February 2, 2018 to May 26, 2018 with 289 valid submissions from 4,282 registered teams. Throughout this report, we describe the details of this new dataset, the purposes and definitions of the tasks, the evaluation protocols, and the summaries of the results.
Mengchao He, Zhibo Yang 0003, Sheng Zhang 0024, Canjie Luo, Feiyu Gao, Qi Zheng 0002, Yongpan Wang, Xin Zhang 0013
ICPR1