VLDB 2026 Research / reviewers in the wild / expert
Maoyuan Ye
dblp:324/2398
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-4180-1096ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Image recognition and object detection · 41% Video understanding and tracking · 22% Segmentation and scene understanding · 17% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
scene text spotting |
1.4 | 2 | 2024 | GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching · NeurIPS 2024 DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text Spotting · CVPR 2023 |
Natural language and speech › Information extraction and text analysis
text segmentation |
0.9 | 1 | 2025 | Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › Video understanding and tracking › object tracking
text tracking |
0.8 | 1 | 2024 | GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching · NeurIPS 2024 |
Computer vision › Video understanding and tracking
video text spotting |
0.8 | 1 | 2024 | GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching · NeurIPS 2024 |
Computer vision › Image recognition and object detection › scene text spotting
end-to-end text spotting |
0.7 | 1 | 2023 | DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text Spotting · CVPR 2023 |
Computer vision › Image recognition and object detection
scene text detection |
0.7 | 1 | 2023 | DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in Transformer · AAAI 2023 |
Machine learning › Deep learning architectures and training
foundation model |
0.3 | 1 | 2025 | Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › Segmentation and scene understanding › prompt-based segmentation
segment anything model |
0.3 | 1 | 2025 | Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › Vision and language
multimodal understanding |
0.2 | 1 | 2024 | GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
promptable segmentation · 0.9parameter-efficient fine-tuning · 0.9hierarchical mask decoder · 0.9transformer · 0.8query-based image text spotter · 0.8long-short term matching · 0.8positional query · 0.7factorized self-attention · 0.7explicit point queries · 0.7detection transformer · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hi-SAM: Marrying Segment Anything Model for Hierarchical Text SegmentationabstractThe Segment Anything Model (SAM), a profound vision foundation model pretrained on a large-scale dataset, breaks the boundaries of general segmentation and sparks various downstream applications. This paper introduces Hi-SAM, a unified model leveraging SAM for hierarchical text segmentation. Hi-SAM excels in segmentation across four hierarchies, including pixel-level text, word, text-line, and paragraph, while realizing layout analysis as well. Specifically, we first turn SAM into a high-quality pixel-level text segmentation (TS) model through a parameter-efficient fine-tuning approach. We use this TS model to iteratively generate the pixel-level text labels in a semi-automatical manner, unifying labels across the four text hierarchies in the HierText dataset. Subsequently, with these complete labels, we launch the end-to-end trainable Hi-SAM based on the TS architecture with a customized hierarchical mask decoder. During inference, Hi-SAM offers both automatic mask generation (AMG) mode and promptable segmentation (PS) mode. In the AMG mode, Hi-SAM segments pixel-level text foreground masks initially, then samples foreground points for hierarchical text mask generation and achieves layout analysis in passing. As for the PS mode, Hi-SAM provides word, text-line, and paragraph masks with a single point click. Experimental results show the state-of-the-art performance of our TS model: 84.86% fgIOU on Total-Text and 88.96% fgIOU on TextSeg for pixel-level text segmentation. Moreover, compared to the previous specialist for joint hierarchical detection and layout analysis on HierText, Hi-SAM achieves significant improvements: 4.73% PQ and 5.39% F1 on the text-line level, 5.49% PQ and 7.39% F1 on the paragraph level layout analysis, requiring fewer training epochs. Maoyuan Ye, Jing Zhang 0037, Juhua Liu, Cong Liu 0006, Bo Du 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term MatchingabstractBeyond the text detection and recognition tasks in image text spotting, video text spotting presents an augmented challenge with the inclusion of tracking. While advanced end-to-end trainable methods have shown commendable performance, the pursuit of multi-task optimization may pose the risk of producing sub-optimal outcomes for individual tasks. In this paper, we identify a main bottleneck in the state-of-the-art video text spotter: the limited recognition capability. In response to this issue, we propose to efficiently turn an off-the-shelf query-based image text spotter into a specialist on video and present a simple baseline termed GoMatching, which focuses the training efforts on tracking while maintaining strong recognition performance. To adapt the image text spotter to video datasets, we add a rescoring head to rescore each detected instance's confidence via efficient tuning, leading to a better tracking candidate pool.
Additionally, we design a long-short term matching module, termed LST-Matcher, to enhance the spotter's tracking capability by integrating both long- and short-term matching results via Transformer. Based on the above simple designs, GoMatching delivers new records on ICDAR15-video, DSText, BOVText, and our proposed novel test set with arbitrary-shaped text termed ArTVideo, which demonstates GoMatching's capability to accommodate general, dense, small, arbitrary-shaped, Chinese and English text scenarios while saving considerable training budgets. The code will be released. Haibin He 0001, Maoyuan Ye, Jing Zhang 0037, Juhua Liu, Bo Du 0001, Dacheng Tao |
NeurIPS | 2 |
| 2023 | DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in TransformerabstractRecently, Transformer-based methods, which predict polygon points or Bezier curve control points for localizing texts, are popular in scene text detection. However, these methods built upon detection transformer framework might achieve sub-optimal training efficiency and performance due to coarse positional query modeling. In addition, the point label form exploited in previous works implies the reading order of humans, which impedes the detection robustness from our observation. To address these challenges, this paper proposes a concise Dynamic Point Text DEtection TRansformer network, termed DPText-DETR. In detail, DPText-DETR directly leverages explicit point coordinates to generate position queries and dynamically updates them in a progressive way. Moreover, to improve the spatial inductive bias of non-local self-attention in Transformer, we present an Enhanced Factorized Self-Attention module which provides point queries within each instance with circular shape guidance. Furthermore, we design a simple yet effective positional label form to tackle the side effect of the previous form. To further evaluate the impact of different label forms on the detection robustness in real-world scenario, we establish an Inverse-Text test set containing 500 manually labeled images. Extensive experiments prove the high training efficiency, robustness, and state-of-the-art performance of our method on popular benchmarks. The code and the Inverse-Text test set are available at https://github.com/ymy-k/DPText-DETR. Maoyuan Ye, Jing Zhang 0037, Shanshan Zhao 0001, Juhua Liu, Bo Du 0001, Dacheng Tao |
AAAI | 1 |
| 2023 | DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text SpottingabstractEnd-to-end text spotting aims to integrate scene text detection and recognition into a unified framework. Dealing with the relationship between the two sub-tasks plays a pivotal role in designing effective spotters. Although Transformer-based methods eliminate the heuristic postprocessing, they still suffer from the synergy issue between the sub-tasks and low training efficiency. In this paper, we present DeepSolo, a simple DETR-like baseline that lets a single Decoder with Explicit Points Solo for text detection and recognition simultaneously. Technically, for each text instance, we represent the character sequence as ordered points and model them with learnable explicit point queries. After passing a single decoder, the point queries have encoded requisite text semantics and locations, thus can be further decoded to the center line, boundary, script, and confidence of text via very simple prediction heads in parallel. Besides, we also introduce a text-matching criterion to deliver more accurate supervisory signals, thus enabling more efficient training. Quantitative experiments on public benchmarks demonstrate that DeepSolo outperforms previous state-of-the-art methods and achieves better training efficiency. In addition, DeepSolo is also compatible with line annotations, which require much less annotation cost than polygons. The code is available at https://github.com/ViTAE-Transformer/DeepSolo. Maoyuan Ye, Jing Zhang 0037, Shanshan Zhao 0001, Juhua Liu, Tongliang Liu, Bo Du 0001, Dacheng Tao |
CVPR | 1 |