EDBT 2026 Demo / reviewers in the wild / expert
Minghang He
dblp:245/0147
· DBLP profile ↗
5ranked-venue papers
1as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Image recognition and object detection · 58% Segmentation and scene understanding · 42% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection › scene text detection
multi-oriented scene text detection |
0.5 | 1 | 2021 | MOST: A Multi-Oriented Scene Text Detector With Localization Refinement · CVPR 2021 |
Computer vision › Image recognition and object detection
scene text detection |
0.5 | 1 | 2021 | MOST: A Multi-Oriented Scene Text Detector With Localization Refinement · CVPR 2021 |
Computer vision › Image recognition and object detection
scene text spotting |
0.5 | 1 | 2021 | Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes · IEEE Trans. Pattern Anal. Mach. Intell. 2021 |
Computer vision › Segmentation and scene understanding › image segmentation › document image segmentation
character segmentation |
0.4 | 1 | 2020 | TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020 |
Computer vision › Image recognition and object detection
scene text recognition |
0.4 | 1 | 2020 | TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.4 | 1 | 2020 | TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020 |
Visual content generation and editing
image generation |
0.4 | 1 | 2020 | SynthText3D: synthesizing scene text images from 3D virtual worlds · Sci. China Inf. Sci. 2020 |
Visual content generation and editing › visual text generation
scene text synthesis |
0.4 | 1 | 2020 | SynthText3D: synthesizing scene text images from 3D virtual worlds · Sci. China Inf. Sci. 2020 |
Visual content generation and editing
synthetic data generation |
0.1 | 1 | 2020 | SynthText3D: synthesizing scene text images from 3D virtual worlds · Sci. China Inf. Sci. 2020 |
Methods — techniques the papers use, named apart from their topics
spatial attention · 0.5non-maximum suppression · 0.5iou loss · 0.5feature alignment · 0.5end-to-end training · 0.5synthetic data generation · 0.4recurrent neural network · 0.4parallel prediction · 0.4attention mechanism · 0.43d rendering · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Maskstr: Guide Scene Text Recognition Models with MaskingabstractText recognition in information loss scenarios like blurriness, occlusion, and perspective distortion is challenging in real-world applications. To enhance robustness, some studies use extra unlabeled data for encoder pretraining. Others focus on improving decoder context reasoning. However, pretraining methods require abundant unlabeled data and high computing resources, while decoder-based approaches risk over-correction. In this paper, we propose MaskSTR, a dual-branch training framework for STR models, using patch masking to simulate information loss. MaskSTR guides visual representation learning, improving robustness to information loss conditions without extra data or training stages. Furthermore, we introduce Block Masking, a novel and straightforward mask generation method, for further performance enhancement. Experiments demonstrate MaskSTR’s effectiveness across CTC, attention, and Transformer decoding methods, achieving significant performance gains and setting new state-of-the-art results. Baole Wei, Minghang He, Liangcai Gao, Duoyou Zhou, Xiang Bai, Zhi Tang 0001 |
ICASSP | 2 |
| 2021 | MOST: A Multi-Oriented Scene Text Detector With Localization RefinementabstractOver the past few years, the field of scene text detection has progressed rapidly that modern text detectors are able to hunt text in various challenging scenarios. However, they might still fall short when handling text instances of extreme aspect ratios and varying scales. To tackle such difficulties, we propose in this paper a new algorithm for scene text detection, which puts forward a set of strategies to significantly improve the quality of text localization. Specifically, a Text Feature Alignment Module (TFAM) is proposed to dynamically adjust the receptive fields of features based on initial raw detections; a Position-Aware Non-Maximum Suppression (PA-NMS) module is devised to selectively concentrate on reliable raw detections and exclude unreliable ones; besides, we propose an Instance-wise IoU loss for balanced training to deal with text instances of different scales. An extensive ablation study demonstrates the effectiveness and superiority of the proposed strategies. The resulting text detection system, which integrates the proposed strategies with a leading scene text detector EAST, achieves state-of-the-art or competitive performance on various standard benchmarks for text detection while keeping a fast running speed. Minghang He, Minghui Liao, Zhibo Yang 0003, Humen Zhong, Jun Tang 0008, Wenqing Cheng, Cong Yao, Yongpan Wang, Xiang Bai |
CVPR | 1 |
| 2021 | Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary ShapesabstractUnifying text detection and text recognition in an end-to-end training fashion has become a new trend for reading text in the wild, as these two tasks are highly relevant and complementary. In this paper, we investigate the problem of scene text spotting, which aims at simultaneous text detection and recognition in natural images. An end-to-end trainable neural network named as Mask TextSpotter is presented. Different from the previous text spotters that follow the pipeline consisting of a proposal generation network and a sequence-to-sequence recognition network, Mask TextSpotter enjoys a simple and smooth end-to-end learning procedure, in which both detection and recognition can be achieved directly from two-dimensional space via semantic segmentation. Further, a spatial attention module is proposed to enhance the performance and universality. Benefiting from the proposed two-dimensional representation on both detection and recognition, it easily handles text instances of irregular shapes, for instance, curved text. We evaluate it on four English datasets and one multi-language dataset, achieving consistently superior performance over state-of-the-art methods in both detection and end-to-end text recognition tasks. Moreover, we further investigate the recognition module of our method separately, which significantly outperforms state-of-the-art methods on both regular and irregular text datasets for scene text recognition. Minghui Liao, Pengyuan Lv, Minghang He, Cong Yao, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | TextScanner: Reading Characters in Order for Robust Scene Text RecognitionabstractDriven by deep learning and a large volume of data, scene text recognition has evolved rapidly in recent years. Formerly, RNN-attention-based methods have dominated this field, but suffer from the problem of attention drift in certain situations. Lately, semantic segmentation based algorithms have proven effective at recognizing text of different forms (horizontal, oriented and curved). However, these methods may produce spurious characters or miss genuine characters, as they rely heavily on a thresholding procedure operated on segmentation maps. To tackle these challenges, we propose in this paper an alternative approach, called TextScanner, for scene text recognition. TextScanner bears three characteristics: (1) Basically, it belongs to the semantic segmentation family, as it generates pixel-wise, multi-channel segmentation maps for character class, position and order; (2) Meanwhile, akin to RNN-attention-based methods, it also adopts RNN for context modeling; (3) Moreover, it performs paralleled prediction for character position and class, and ensures that characters are transcripted in the correct order. The experiments on standard benchmark datasets demonstrate that TextScanner outperforms the state-of-the-art methods. Moreover, TextScanner shows its superiority in recognizing more difficult text such as Chinese transcripts and aligning with target characters. Zhaoyi Wan, Minghang He, Xiang Bai, Cong Yao |
AAAI | 2 |
| 2020 | SynthText3D: synthesizing scene text images from 3D virtual worlds
Minghui Liao, Boyu Song, Shangbang Long, Minghang He, Cong Yao, Xiang Bai |
Sci. China Inf. Sci. | 4 |