Shengtian Jiang

dblp:433/8167 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0006-0217-3193ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Image recognition and object detection · 67% Language models and text generation · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › scene text detection
arbitrary-shaped text detection
1.012026
TextRSR: Enhanced Arbitrary-Shaped Scene Text Representation via Robust Subspace Recovery · IEEE Trans. Multim. 2026
Computer vision › Image recognition and object detection
scene text detection
1.012026
TextRSR: Enhanced Arbitrary-Shaped Scene Text Representation via Robust Subspace Recovery · IEEE Trans. Multim. 2026
Natural language and speech › Language models and text generation
text representation
1.012026
TextRSR: Enhanced Arbitrary-Shaped Scene Text Representation via Robust Subspace Recovery · IEEE Trans. Multim. 2026

Methods — techniques the papers use, named apart from their topics

robust subspace recovery · 1.0dynamic sparse assignment · 1.0
YearPublicationVenuePosition
2026 TextRSR: Enhanced Arbitrary-Shaped Scene Text Representation via Robust Subspace Recovery
abstract
In recent years, scene text detection research has increasingly focused on arbitrary-shaped texts, where text representation is a fundamental problem. However, most existing methods still struggle to separate adjacent or overlapping texts due to ambiguous spatial positions of points or segmentation masks. Besides, the time efficiency of the entire pipeline is often neglected, resulting in sub-optimal inference speed. To tackle these problems, we first propose a novel text representation method based on robust subspace recovery, which robustly represents complex text shapes by combining orthogonal basis vectors learned from labeled text contours. These basis vectors capture basis contour patterns with distinct information, enabling clearer boundaries even in densely populated text scenarios. Moreover, we propose a dynamic sparse assignment scheme for positive samples that adaptively adjusts their weights during training, which not only accelerates inference speed by eliminating redundant predictions but also enhances feature learning by providing sufficient supervision signals. Building on these innovations, we present TextRSR, an accurate and efficient scene text detection network. Extensive experiments on challenging benchmarks demonstrate the superior accuracy and efficiency of TextRSR compared to state-of-the-art methods. Particularly, TextRSR achieves an F-measure of 88.5% at 37.8 frames per second (FPS) for CTW1500 dataset and an F-measure of 89.1% at 23.1 FPS for Total-Text dataset.
Zhiwen Shao, Shengtian Jiang, Hancheng Zhu, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung
IEEE Trans. Multim.2