EDBT 2026 Demo / reviewers in the wild / expert
Shenggao Zhu
dblp:122/2694
· DBLP profile ↗
17ranked-venue papers
1as first author
12since 2021 · last 2023
0000-0002-3254-0058ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate ModelingabstractTable structure recognition aims to extract the logical and physical structure of unstructured table images into a machine-readable format. The latest end-to-end image-to-text approaches simultaneously predict the two structures by two decoders, where the prediction of the physical structure (the bounding boxes of the cells) is based on the representation of the logical structure. However, the previous methods struggle with imprecise bounding boxes as the logical representation lacks local visual information. To address this issue, we propose an end-to-end sequential modeling framework for table structure recognition called VAST. It contains a novel coordinate sequence decoder triggered by the representation of the non-empty cell from the logical structure decoder. In the coordinate sequence decoder, we model the bounding box coordinates as a language sequence, where the left, top, right and bottom coordinates are decoded sequentially to leverage the inter-coordinate dependency. Furthermore, we propose an auxiliary visual-alignment loss to enforce the logical representation of the non-empty cells to contain more local visual details, which helps produce better cell bounding boxes. Extensive experiments demonstrate that our proposed method can achieve state-of-the-art results in both logical and physical structure recognition. The ablation study also validates that the proposed coordinate sequence decoder and the visual-alignment loss are the keys to the success of our method. Yongshuai Huang, Ning Lu 0003, Dapeng Chen, Zecheng Xie, Shenggao Zhu, Liangcai Gao, Wei Peng 0011 |
CVPR | 6 |
| 2023 | Recognition of Handwritten Chinese Text by Segmentation: A Segment-Annotation-Free ApproachabstractOnline and offline handwritten Chinese text recognition (HTCR) has been studied for decades. Early methods adopted oversegmentation-based strategies but suffered from low speed, insufficient accuracy, and high cost of character segmentation annotations. Recently, segmentation-free methods based on connectionist temporal classification (CTC) and attention mechanism, have dominated the field of HCTR. However, people actually read text character by character, especially for ideograms such as Chinese. This raises the question: are segmentation-free strategies really the best solution to HCTR? To explore this issue, we propose a new segmentation-based method for recognizing handwritten Chinese text that is implemented using a simple yet efficient fully convolutional network. A novel weakly supervised learning method is proposed to enable the network to be trained using only transcript annotations; thus, the expensive character segmentation annotations required by previous segmentation-based methods can be avoided. Owing to the lack of context modeling in fully convolutional networks, we propose a contextual regularization method to integrate contextual information into the network during the training stage, which can further improve the recognition performance. Extensive experiments conducted on four widely used benchmarks, namely CASIA-HWDB, CASIA-OLHWDB, ICDAR2013, and SCUT-HCCDoc, show that our method significantly surpasses existing methods on both online and offline HCTR, and exhibits a considerably higher inference speed than CTC/attention-based approaches. Dezhi Peng, Weihong Ma, Canyu Xie, Hesuo Zhang, Shenggao Zhu |
IEEE Trans. Multim. | 6 |
| 2022 | SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionabstractEnd-to-end scene text spotting has attracted great attention in recent years due to the success of excavating the intrinsic synergy of the scene text detection and recognition. However, recent state-of-the-art methods usually incorporate detection and recognition simply by sharing the backbone, which does not directly take advantage of the feature interaction between the two tasks. In this paper, we propose a new end-to-end scene text spotting framework termed SwinTextSpotter. Using a transformer encoder with dynamic head as the detector, we unify the two tasks with a novel Recognition Conversion mechanism to explicitly guide text localization through recognition loss. The straightforward design results in a concise framework that requires neither additional rectification module nor character-level annotation for the arbitrarily-shaped text. Qualitative and quantitative experiments on multi-oriented datasets RoIC13 and ICDAR 2015, arbitrarily-shaped datasets Total-Text and CTW1500, and multi-lingual datasets ReCTS (Chinese) and VinText (Viet-namese) demonstrate SwinTextSpotter significantly outperforms existing methods. Code is available at https://github.com/mxin262/SwinTextSpotter. Mingxin Huang, Zhenghao Peng, Chongyu Liu, Dahua Lin, Shenggao Zhu, Nicholas Jing Yuan, Kai Ding 0009 |
CVPR | 6 |
| 2022 | Look Closer to Supervise Better: One-Shot Font Generation via Component-Based DiscriminatorabstractAutomatic font generation remains a challenging research issue due to the large amounts of characters with complicated structures. Typically, only a few samples can serve as the style/content reference (termed few-shot learning), which further increases the difficulty to preserve local style patterns or detailed glyph structures. We investigate the drawbacks of previous studies and find that a coarsegrained discriminator is insufficient for supervising a font generator. To this end, we propose a novel Component-Aware Module (CAM), which supervises the generator to decouple content and style at a more fine-grained level, i.e., the component level. Different from previous studies struggling to increase the complexity of generators, we aim to perform more effective supervision for a relatively simple generator to achieve its full potential, which is a brand new perspective for font generation. The whole framework achieves remarkable results by coupling component-level supervision with adversarial learning, hence we call it Component-Guided GAN, shortly CG-GAN. Extensive experiments show that our approach outperforms state-of-the-art one-shot font generation methods. Furthermore, it can be applied to handwritten word synthesis and scene text image editing, suggesting the generalization of our approach. Yuxin Kong, Canjie Luo, Weihong Ma, Qiyuan Zhu, Shenggao Zhu, Nicholas Jing Yuan |
CVPR | 5 |
| 2022 | Detecting Tampered Scene Text in the Wild
Yuxin Wang 0002, Hongtao Xie 0001, Mengting Xing, Jing Wang 0221, Shenggao Zhu, Yongdong Zhang 0001 |
ECCV (28) | 5 |
| 2022 | SPTS: Single-Point Text SpottingabstractExisting scene text spotting (i.e., end-to-end text detection and recognition) methods rely on costly bounding box annotations (e.g., text-line, word-level, or character-level bounding boxes). For the first time, we demonstrate that training scene text spotting models can be achieved with an extremely low-cost annotation of a single-point for each instance. We propose an end-to-end scene text spotting method that tackles scene text spotting as a sequence prediction task. Given an image as input, we formulate the desired detection and recognition results as a sequence of discrete tokens and use an auto-regressive Transformer to predict the sequence. The proposed method is simple yet effective, which can achieve state-of-the-art results on widely used benchmarks. Most significantly, we show that the performance is not very sensitive to the positions of the point annotation, meaning that it can be much easier to be annotated or even be automatically generated than the bounding box that requires precise positions. We believe that such a pioneer attempt indicates a significant opportunity for scene text spotting applications of a much larger scale than previously possible. The code is available at https://github.com/shannanyinxiang/SPTS. Dezhi Peng, Xinyu Wang 0010, Jiaxin Zhang 0003, Mingxin Huang, Songxuan Lai, Jing Li 0036, Shenggao Zhu, Dahua Lin, Chunhua Shen, Xiang Bai |
ACM Multimedia | 8 |
| 2022 | Reading and Writing: Discriminative and Generative Modeling for Self-Supervised Text RecognitionabstractExisting text recognition methods usually need large-scale training data. Most of them rely on synthetic training data due to the lack of annotated real images. However, there is a domain gap between the synthetic data and real data, which limits the performance of the text recognition models. Recent self-supervised text recognition methods attempted to utilize unlabeled real images by introducing contrastive learning, which mainly learns the discrimination of the text images. Inspired by the observation that humans learn to recognize the texts through both reading and writing, we propose to learn discrimination and generation by integrating contrastive learning and masked image modeling in our self-supervised method. The contrastive learning branch is adopted to learn the discrimination of text images, which imitates the reading behavior of humans. Meanwhile, masked image modeling is firstly introduced for text recognition to learn the context generation of the text images, which is similar to the writing behavior. The experimental results show that our method outperforms previous self-supervised text recognition methods by 10.2%-20.2% on irregular scene text recognition datasets. Moreover, our proposed text recognizer exceeds previous state-of-the-art text recognition methods by averagely 5.3% on 11benchmarks, with similar model size. We also demonstrate that our pre-trained model can be easily applied to other text-related tasks with obvious performance gain. Minghui Liao, Pu Lu, Jing Wang 0221, Shenggao Zhu, Hualin Luo, Qi Tian 0001, Xiang Bai |
ACM Multimedia | 5 |
| 2022 | Boundary TextSpotter: Toward Arbitrary-Shaped Scene Text SpottingabstractReading arbitrary-shaped text in an end-to-end fashion has received particularly growing interested in computer vision. In this paper, we study the problem of scene text spotting, which aims to detect and recognize text from cluttered images simultaneously and propose an end-to-end trainable neural network named Boundary TextSpotter. Different from existing methods that describe the shape of text instance with bounding box or shape mask, Boundary TextSpotter formulates it as a set of boundary points. Besides, the representation of such boundary points provides the order of reading text. Benefiting from the representation on both detection and recognition, Boundary TextSpotter can easily deal with the text of arbitrary shapes. Further, to efficiently detect the boundary points of the text, a single-stage text detector is proposed, which can almost perform at a real-time speed. Experiments on three challenging datasets, including ICDAR2015, Total-Text and CTW1500 demonstrate that the proposed method achieves state-of-the-art or competitive results, meanwhile significantly improving the inference speed. Pu Lu, Hao Wang 0207, Shenggao Zhu, Jing Wang 0221, Xiang Bai, Wenyu Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | PETR: Rethinking the Capability of Transformer-Based Language Model in Scene Text RecognitionabstractThe exploration of linguistic information promotes the development of scene text recognition task. Benefiting from the significance in parallel reasoning and global relationship capture, transformer-based language model (TLM) has achieved dominant performance recently. As a decoupled structure from the recognition process, we argue that TLM's capability is limited by the input low-quality visual prediction. To be specific: 1) The visual prediction with low character-wise accuracy increases the correction burden of TLM. 2) The inconsistent word length between visual prediction and original image provides a wrong language modeling guidance in TLM. In this paper, we propose a Progressive scEne Text Recognizer (PETR) to improve the capability of transformer-based language model by handling above two problems. Firstly, a Destruction Learning Module (DLM) is proposed to consider the linguistic information in the visual context. DLM introduces the recognition of destructed images with disordered patches in the training stage. Through guiding the vision model to restore patch orders and make word-level prediction on the destructed images, visual prediction with high character-wise accuracy is obtained by exploring inner relationship between the local visual patches. Secondly, a new Language Rectification Module (LRM) is proposed to optimize the word length for language guidance rectification. Through progressively implementing LRM in different language modeling steps, a novel progressive rectification network is constructed to handle some extremely challenging cases (e.g. distortion, occlusion, etc.). By utilizing DLM and LRM, PETR enhances the capability of transformer-based language model from a more general aspect, that is, focusing on the reduction of correction burden and rectification of language modeling guidance. Compared with parallel transformer-based methods, PETR obtains 1.0% and 0.8% improvement on regular and irregular datasets respectively while introducing only 1.7M additional parameters. The extensive experiments on both English and Chinese benchmarks demonstrate that PETR achieves the state-of-the-art results. Yuxin Wang 0002, Hongtao Xie 0001, Shancheng Fang, Mengting Xing, Jing Wang 0221, Shenggao Zhu, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2021 | Scene Text Retrieval via Joint Text Detection and Similarity LearningabstractScene text retrieval aims to localize and search all text instances from an image gallery, which are the same or similar with a given query text. Such a task is usually realized by matching a query text to the recognized words, outputted by an end-to-end scene text spotter. In this paper, we address this problem by directly learning a cross-modal similarity between a query text and each text instance from natural images. Specifically, we establish an end-to-end trainable network, jointly optimizing the procedures of scene text detection and cross-modal similarity learning. In this way, scene text retrieval can be simply performed by ranking the detected text instances with the learned similarity. Experiments on three benchmark datasets demonstrate our method consistently outperforms the state-of-the-art scene text spotting/retrieval approaches. In particular, the proposed framework of joint detection and similarity learning achieves significantly better performance than separated methods. Code is available at: https://github.com/lanfeng4659/STR-TDSL. Hao Wang 0207, Xiang Bai, Shenggao Zhu, Jing Wang 0221, Wenyu Liu 0001 |
CVPR | 4 |
| 2021 | From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkabstractIn this paper, we abandon the dominant complex language model and rethink the linguistic learning process in the scene text recognition. Different from previous methods considering the visual and linguistic information in two separate structures, we propose a Visual Language Modeling Network (VisionLAN), which views the visual and linguistic information as a union by directly enduing the vision model with language capability. Specially, we introduce the text recognition of character-wise occluded feature maps in the training stage. Such operation guides the vision model to use not only the visual texture of characters, but also the linguistic information in visual context for recognition when the visual cues are confused (e.g. occlusion, noise, etc.). As the linguistic information is acquired along with visual features without the need of extra language model, Vision-LAN significantly improves the speed by 39% and adaptively considers the linguistic information to enhance the visual features for accurate recognition. Furthermore, an Occlusion Scene Text (OST) dataset is proposed to evaluate the performance on the case of missing character-wise visual cues. The state of-the-art results on several benchmarks prove our effectiveness. Code and dataset are available at https://github.com/wangyuxin87/VisionLAN. Yuxin Wang 0002, Hongtao Xie 0001, Shancheng Fang, Jing Wang 0221, Shenggao Zhu, Yongdong Zhang 0001 |
ICCV | 5 |
| 2021 | Video Text Tracking With a Spatio-Temporal Complementary ModelabstractText tracking is to track multiple texts in a video, and construct a trajectory for each text. Existing methods tackle this task by utilizing the tracking-by-detection framework, i.e., detecting the text instances in each frame and associating the corresponding text instances in consecutive frames. We argue that the tracking accuracy of this paradigm is severely limited in more complex scenarios, e.g., owing to motion blur, etc., the missed detection of text instances causes the break of the text trajectory. In addition, different text instances with similar appearance are easily confused, leading to the incorrect association of the text instances. To this end, a novel spatio-temporal complementary text tracking model is proposed in this paper. We leverage a Siamese Complementary Module to fully exploit the continuity characteristic of the text instances in the temporal dimension, which effectively alleviates the missed detection of the text instances, and hence ensures the completeness of each text trajectory. We further integrate the semantic cues and the visual cues of the text instance into a unified representation via a text similarity learning network, which supplies a high discriminative power in the presence of text instances with similar appearance, and thus avoids the mis-association between them. Our method achieves state-of-the-art performance on several public benchmarks. The source code is available at https://github.com/lsabrinax/VideoTextSCM. Yuzhe Gao, Jiajian Zhang, Yu Zhou 0016, Jing Wang 0221, Shenggao Zhu, Xiang Bai |
IEEE Trans. Image Process. | 7 |
| 2018 | MANA: Designing and Validating a User-Centered Mobility Analysis SystemabstractIn this paper, we demonstrate a new IMU-based wearable system (dubbed MANA or Mobility ANAlytics) for measuring gait in a clinical setting. The design process and choices that were made to ensure that the technology was invisible and accessible are described. We collect a rich and diverse dataset of walking data from sixty participants, including forty people with Parkinson's Disease (PD). The system is then validated in a clinical setting with this dataset. We present novel and innovative algorithms to measure common gait parameters. The system is able to estimate these gait parameters with high accuracy, with a mean absolute error of 4.0 cm for stride length and 2.6 cm for step length, outperforming all state-of-the-art methods that included data from people with PD. Boyd Anderson, Shenggao Zhu, Hugh Anderson, Chao Xu Tay, Vincent Y. F. Tan, Ye Wang 0007 |
ASSETS | 2 |
| 2018 | A unified scheme of text localization and structured data extraction for joint OCR and data miningabstractBoth text detection and structured data extraction are imperative in an optical character recognition (OCR) processing pipeline. Text detection, especially for indistinct, diverse, multi-language text regions, is one of the most challenging tasks in computer vision and has attracted increasing attention recently. Moreover, although there are some studies in data mining related to structured data extraction, it has not received its deserved attention as one of important steps in OCR. The previous methods for structural data extraction, including layout template-based, rule-based, and natural language processing (NLP)-based methods, usually leads to either inaccurate results or complex modules. In this paper, we integrate text detection and structured data extraction into a unified deep learning-based Image Text Extraction (ITE) scheme. Our ITE is an end-to-end trainable model and able to handle multi-scale and multi-lingual text in a single process. Experiments on large-scale real-world passport and medical receipt datasets have demonstrated the superiority of the proposed method in terms of both effectiveness and efficiency. Yibin Ye, Shenggao Zhu, Jing Wang 0221, Qi Du, Yezhang Yang, Dandan Tu, Lanjun Wang, Jiebo Luo 0001 |
IEEE BigData | 2 |
| 2016 | A Computer Vision-Based System for Stride Length Estimation using a Mobile Phone CameraabstractConditions such as Parkinson's disease (PD), a chronic neurodegenerative disorder which severely affects the motor system, will be an increasingly common problem for our growing and aging population. Gait analysis is widely used as a noninvasive method for PD diagnosis and assessment. However, current clinical systems for gait analysis usually require highly specialized cameras and lab settings, which are expensive and not scalable. This paper presents a computer vision-based gait analysis system using a camera on a common mobile phone. A simple PVC mat was designed with markers printed on it, on which a subject can walk whilst being recorded by a mobile phone camera. A set of video analysis methods were developed to segment the walking video, detect the mat and feet locations, and calculate gait parameters such as stride length. Experiments showed that stride length measurement has a mean absolute error of 0.62 cm, which is comparable with the "gold standard" walking mat system GAITRite. We also tested our system on Parkinson's disease patients in a real clinical environment. Our system is affordable, portable, and scalable, indicating a potential clinical gait measurement tool for use in both hospitals and the homes of patients. Boyd Anderson, Shenggao Zhu, Ye Wang 0007 |
ASSETS | 3 |
| 2014 | Validating an iOS-based Rhythmic Auditory Cueing Evaluation (iRACE) for Parkinson's DiseaseabstractMovement disorders such as Parkinson's disease (PD) will affect a rapidly growing segment of the population as society continues to age. Rhythmic Auditory Cueing (RAC) is a well-supported evidence-based intervention for the treatment of gait impairments in PD. RAC interventions have not been widely adopted, however, due to limitations in access to personnel, technological, and financial resources. To help "scale up" RAC for wider distribution, we have developed an iOS-based Rhythmic Auditory Cueing Evaluation (iRACE) mobile application to deliver RAC and assess motor performance in PD patients. The touchscreen of the mobile device is used to assess motor timing during index finger tapping, and the device's built-in tri-axial accelerometer and gyroscope to assess step time and step length during walking. Novel machine learning-based gait analysis algorithms have been developed for iRACE, including heel strike detection, step length quantification, and left-versus-right foot identification. The concurrent validity of iRACE was assessed using a clinic-standard instrumented walking mat and a pair of force-sensing resistor sensors. Results from 10 PD patients reveal that iRACE has low error rates (<±1.0%) across a set of four clinically relevant outcome measures, indicating a potentially useful clinical tool. Shenggao Zhu, Robert J. Ellis, Gottfried Schlaug, Yee Sien Ng, Ye Wang 0007 |
ACM Multimedia | 1 |
| 2013 | Hearing versus Seeing Identical Twins
Li Zhang 0005, Shenggao Zhu, Terence Sim, Wee Kheng Leow, Hossein Nejati, Dong Guo 0001 |
CAIP (1) | 2 |