Minghang He

dblp:245/0147 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Image recognition and object detection · 58% Segmentation and scene understanding · 42%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › scene text detection
multi-oriented scene text detection
0.512021
MOST: A Multi-Oriented Scene Text Detector With Localization Refinement · CVPR 2021
Computer vision › Image recognition and object detection
scene text detection
0.512021
MOST: A Multi-Oriented Scene Text Detector With Localization Refinement · CVPR 2021
Computer vision › Image recognition and object detection
scene text spotting
0.512021
Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Computer vision › Segmentation and scene understanding › image segmentation › document image segmentation
character segmentation
0.412020
TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020
Computer vision › Image recognition and object detection
scene text recognition
0.412020
TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020
Computer vision › Segmentation and scene understanding
semantic segmentation
0.412020
TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020
Visual content generation and editing
image generation
0.412020
SynthText3D: synthesizing scene text images from 3D virtual worlds · Sci. China Inf. Sci. 2020
Visual content generation and editing › visual text generation
scene text synthesis
0.412020
SynthText3D: synthesizing scene text images from 3D virtual worlds · Sci. China Inf. Sci. 2020
Visual content generation and editing
synthetic data generation
0.112020
SynthText3D: synthesizing scene text images from 3D virtual worlds · Sci. China Inf. Sci. 2020

Methods — techniques the papers use, named apart from their topics

spatial attention · 0.5non-maximum suppression · 0.5iou loss · 0.5feature alignment · 0.5end-to-end training · 0.5synthetic data generation · 0.4recurrent neural network · 0.4parallel prediction · 0.4attention mechanism · 0.43d rendering · 0.4
YearPublicationVenuePosition
2024 Maskstr: Guide Scene Text Recognition Models with Masking
abstract
Text recognition in information loss scenarios like blurriness, occlusion, and perspective distortion is challenging in real-world applications. To enhance robustness, some studies use extra unlabeled data for encoder pretraining. Others focus on improving decoder context reasoning. However, pretraining methods require abundant unlabeled data and high computing resources, while decoder-based approaches risk over-correction. In this paper, we propose MaskSTR, a dual-branch training framework for STR models, using patch masking to simulate information loss. MaskSTR guides visual representation learning, improving robustness to information loss conditions without extra data or training stages. Furthermore, we introduce Block Masking, a novel and straightforward mask generation method, for further performance enhancement. Experiments demonstrate MaskSTR’s effectiveness across CTC, attention, and Transformer decoding methods, achieving significant performance gains and setting new state-of-the-art results.
Baole Wei, Minghang He, Liangcai Gao, Duoyou Zhou, Xiang Bai, Zhi Tang 0001
ICASSP2
2021 MOST: A Multi-Oriented Scene Text Detector With Localization Refinement
abstract
Over the past few years, the field of scene text detection has progressed rapidly that modern text detectors are able to hunt text in various challenging scenarios. However, they might still fall short when handling text instances of extreme aspect ratios and varying scales. To tackle such difficulties, we propose in this paper a new algorithm for scene text detection, which puts forward a set of strategies to significantly improve the quality of text localization. Specifically, a Text Feature Alignment Module (TFAM) is proposed to dynamically adjust the receptive fields of features based on initial raw detections; a Position-Aware Non-Maximum Suppression (PA-NMS) module is devised to selectively concentrate on reliable raw detections and exclude unreliable ones; besides, we propose an Instance-wise IoU loss for balanced training to deal with text instances of different scales. An extensive ablation study demonstrates the effectiveness and superiority of the proposed strategies. The resulting text detection system, which integrates the proposed strategies with a leading scene text detector EAST, achieves state-of-the-art or competitive performance on various standard benchmarks for text detection while keeping a fast running speed.
Minghang He, Minghui Liao, Zhibo Yang 0003, Humen Zhong, Jun Tang 0008, Wenqing Cheng, Cong Yao, Yongpan Wang, Xiang Bai
CVPR1
2021 Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
abstract
Unifying text detection and text recognition in an end-to-end training fashion has become a new trend for reading text in the wild, as these two tasks are highly relevant and complementary. In this paper, we investigate the problem of scene text spotting, which aims at simultaneous text detection and recognition in natural images. An end-to-end trainable neural network named as Mask TextSpotter is presented. Different from the previous text spotters that follow the pipeline consisting of a proposal generation network and a sequence-to-sequence recognition network, Mask TextSpotter enjoys a simple and smooth end-to-end learning procedure, in which both detection and recognition can be achieved directly from two-dimensional space via semantic segmentation. Further, a spatial attention module is proposed to enhance the performance and universality. Benefiting from the proposed two-dimensional representation on both detection and recognition, it easily handles text instances of irregular shapes, for instance, curved text. We evaluate it on four English datasets and one multi-language dataset, achieving consistently superior performance over state-of-the-art methods in both detection and end-to-end text recognition tasks. Moreover, we further investigate the recognition module of our method separately, which significantly outperforms state-of-the-art methods on both regular and irregular text datasets for scene text recognition.
Minghui Liao, Pengyuan Lv, Minghang He, Cong Yao, Xiang Bai
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 TextScanner: Reading Characters in Order for Robust Scene Text Recognition
abstract
Driven by deep learning and a large volume of data, scene text recognition has evolved rapidly in recent years. Formerly, RNN-attention-based methods have dominated this field, but suffer from the problem of attention drift in certain situations. Lately, semantic segmentation based algorithms have proven effective at recognizing text of different forms (horizontal, oriented and curved). However, these methods may produce spurious characters or miss genuine characters, as they rely heavily on a thresholding procedure operated on segmentation maps. To tackle these challenges, we propose in this paper an alternative approach, called TextScanner, for scene text recognition. TextScanner bears three characteristics: (1) Basically, it belongs to the semantic segmentation family, as it generates pixel-wise, multi-channel segmentation maps for character class, position and order; (2) Meanwhile, akin to RNN-attention-based methods, it also adopts RNN for context modeling; (3) Moreover, it performs paralleled prediction for character position and class, and ensures that characters are transcripted in the correct order. The experiments on standard benchmark datasets demonstrate that TextScanner outperforms the state-of-the-art methods. Moreover, TextScanner shows its superiority in recognizing more difficult text such as Chinese transcripts and aligning with target characters.
Zhaoyi Wan, Minghang He, Xiang Bai, Cong Yao
AAAI2
2020 SynthText3D: synthesizing scene text images from 3D virtual worlds
Minghui Liao, Boyu Song, Shangbang Long, Minghang He, Cong Yao, Xiang Bai
Sci. China Inf. Sci.4