Lei Sun 0003

dblp:02/2264-3 · DBLP profile ↗
← Back
34ranked-venue papers
4as first author
13since 2021 · last 2024
0000-0002-4974-9122ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 15 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2024 UniVIE: A Unified Label Space Approach to Visual Information Extraction from Form-Like Documents
Jiawei Wang 0026, Weihong Lin, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICDAR (6)5
2024 Dynamic Relation Transformer for Contextual Text Block Detection
Jiawei Wang 0026, Shunchi Zhang, Chixiang Ma, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICDAR (1)6
2024 Mathematical formula detection in document images: A new dataset and a new approach
Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
Pattern Recognit.3
2024 Detect-order-construct: A tree construction based approach for hierarchical document structure analysis
Jiawei Wang 0026, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
Pattern Recognit.4
2023 A Question-Answering Approach to Key Value Pair Extraction from Form-Like Document Images
abstract
In this paper, we present a new question-answering (QA) based key-value pair extraction approach, called KVPFormer, to robustly extracting key-value relationships between entities from form-like document images. Specifically, KVPFormer first identifies key entities from all entities in an image with a Transformer encoder, then takes these key entities as questions and feeds them into a Transformer decoder to predict their corresponding answers (i.e., value entities) in parallel. To achieve higher answer prediction accuracy, we propose a coarse-to-fine answer prediction approach further, which first extracts multiple answer candidates for each identified question in the coarse stage and then selects the most likely one among these candidates in the fine stage. In this way, the learning difficulty of answer prediction can be effectively reduced so that the prediction accuracy can be improved. Moreover, we introduce a spatial compatibility attention bias into the self-attention/cross-attention mechanism for KVPFormer to better model the spatial interactions between entities. With these new techniques, our proposed KVPFormer achieves state-of-the-art results on FUNSD and XFUND datasets, outperforming the previous best-performing method by 7.2% and 13.2% in F1 score, respectively.
Zhuoyuan Wu, Zhuoyao Zhong, Weihong Lin, Lei Sun 0003, Qiang Huo
AAAI5
2023 DETRs with Hybrid Matching
abstract
One-to-one set matching is a key design for DETR to establish its end-to-end capability, so that object detection does not require a hand-crafted NMS (non-maximum suppression) to remove duplicate detections. This end-to-end signature is important for the versatility of DETR, and it has been generalized to broader vision tasks. However, we note that there are few queries assigned as positive samples and the one-to-one set matching significantly reduces the training efficacy of positive samples. We propose a simple yet effective method based on a hybrid matching scheme that combines the original one-to-one matching branch with an auxiliary one-to-many matching branch during training. Our hybrid strategy has been shown to significantly improve accuracy. In inference, only the original one-to-one match branch is used, thus maintaining the end-to-end merit and the same inference efficiency of DETR. The method is namedℋ-DETR, and it shows that a wide range of representative DETR methods can be consistently improved across a wide range of visual tasks, including Deformable-DETR, PETRv2, PETR, and TransTrack, among others. Code is available at: https://github.com/HDETR.
Ding Jia, Yuhui Yuan, Haodi He, Xiaopei Wu, Haojun Yu, Weihong Lin, Lei Sun 0003, Chao Zhang 0001, Han Hu 0001
CVPR7
2023 DQ-DETR: Dynamic Queries Enhanced Detection Transformer for Arbitrary Shape Text Detection
Chixiang Ma, Lei Sun 0003, Jiawei Wang 0026, Qiang Huo
ICDAR (2)2
2023 A Hybrid Approach to Document Layout Analysis for Heterogeneous Document Images
Zhuoyao Zhong, Jiawei Wang 0026, Haiqing Sun, Erhan Zhang, Lei Sun 0003, Qiang Huo
ICDAR (5)6
2023 Robust Table Detection and Structure Recognition from Heterogeneous Document Images
abstract
We introduce a new table detection and structure recognition approach named RobusTabNet to detect the boundaries of tables and reconstruct the cellular structure of each table from heterogeneous document images. For table detection, we propose to use CornerNet as a new region proposal network to generate higher quality table proposals for Faster R-CNN, which has significantly improved the localization accuracy of Faster R-CNN for table detection. Consequently, our table detection approach achieves state-of-the-art performance on three public table detection benchmarks, namely cTDaR TrackA, PubLayNet and IIIT-AR-13K, by only using a lightweight ResNet-18 backbone network. Furthermore, we propose a new split-and-merge based table structure recognition approach, in which a novel spatial CNN based separation line prediction module is proposed to split each detected table into a grid of cells, and a Grid CNN based cell merging module is applied to recover the spanning cells. As the spatial CNN module can effectively propagate contextual information across the whole table image, our table structure recognizer can robustly recognize tables with large blank spaces and geometrically distorted (even curved) tables. Thanks to these two techniques, our table structure recognition approach achieves state-of-the-art performance on three public benchmarks, including SciTSR, PubTabNet and cTDaR TrackB2-Modern. Moreover, we have further demonstrated the advantages of our approach in recognizing tables with complex structures, large blank spaces, as well as geometrically distorted or even curved shapes on a more challenging in-house dataset.
Chixiang Ma, Weihong Lin, Lei Sun 0003, Qiang Huo
Pattern Recognit.3
2023 Robust table structure recognition with dynamic queries enhanced detection transformer
Jiawei Wang 0026, Weihong Lin, Chixiang Ma, Lei Sun 0003, Qiang Huo
Pattern Recognit.6
2022 TSRFormer: Table Structure Recognition with Transformers
abstract
We present a new table structure recognition (TSR) approach, called TSRFormer, to robustly recognizing the structures of complex tables with geometrical distortions from various table images. Unlike previous methods, we formulate table separation line prediction as a line regression problem instead of an image segmentation problem and propose a new two-stage DETR based separator prediction approach, dubbed Sep arator RE gression TR ansformer (SepRETR), to predict separation lines from table images directly. To make the two-stage DETR framework work efficiently and effectively for the separation line prediction task, we propose two improvements: 1) A prior-enhanced matching strategy to solve the slow convergence issue of DETR; 2) A new cross attention module to sample features from a high-resolution convolutional feature map directly so that high localization accuracy is achieved with low computational cost. After separation line prediction, a simple relation network based cell merging module is used to recover spanning cells. With these new techniques, our TSRFormer achieves state-of-the-art performance on several benchmark datasets, including SciTSR, PubTabNet and WTW. Furthermore, we have validated the robustness of our approach to tables with complex structures, borderless cells, large blank spaces, empty or spanning cells as well as distorted or even curved shapes on a more challenging real-world in-house dataset.
Weihong Lin, Chixiang Ma, Jiawei Wang 0026, Lei Sun 0003, Qiang Huo
ACM Multimedia6
2021 ViBERTgrid: A Jointly Trained Multi-modal 2D Document Representation for Key Information Extraction from Documents
Weihong Lin, Qifang Gao, Lei Sun 0003, Zhuoyao Zhong, Qin Ren 0003, Qiang Huo
ICDAR (1)3
2021 ReLaText: Exploiting visual relationships for arbitrary-shaped scene text detection with graph convolutional networks
Chixiang Ma, Lei Sun 0003, Zhuoyao Zhong, Qiang Huo
Pattern Recognit.2
2019 A Relation Network Based Approach to Curved Text Detection
abstract
In this paper, a new relation network based approach to curved text detection is proposed by formulating it as a visual relationship detection problem. The key idea is to decompose curved text detection into two subproblems, namely detection of text primitives and prediction of link relationship for each nearby text primitive pair. Specifically, an anchor-free region proposal network based text detector is first used to detect text primitives of different scales from different feature maps of a feature pyramid network, from which a manageable number of text primitive pairs are selected. Then, a relation network is used to predict whether each text primitive pair belongs to a same text instance. Finally, isolated text primitives are grouped into curved text instances based on link relationships of text primitive pairs. Because pairwise link prediction has used features extracted from the bounding boxes of each text primitive and their union, the relation network can effectively leverage wider context information to improve link prediction accuracy. Furthermore, since the link relationships of relatively distant text primitives can be predicted robustly, our relation network based text detector is capable of detecting text instances with large inter-character spaces. Consequently, our proposed approach achieves superior performance on not only two public curved text detection datasets, namely Total-Text and SCUT-CTW1500, but also a multi-oriented text detection dataset, namely MSRA-TD500.
Chixiang Ma, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICDAR3
2019 A Teacher-Student Learning Based Born-Again Training Approach to Improving Scene Text Detection Accuracy
abstract
With the recent success of convolutional neural network (CNN) based text detection approaches, designing better CNN-based text detection frameworks has become a major research focus to improve text detection accuracy. In this paper, instead of following this direction, we propose to use a born-again training strategy, which is based on teacher-student learning (TSL), to improve the accuracy of the state-of-the-art CNN-based text detectors. More specifically, given a well-trained CNN-based text detector, we take it as a teacher model and train from scratch a new student model with the same topology under the supervision of both the teacher model and ground-truth labels. Furthermore, we propose a new proposal-free multi-level feature mimicking approach to making multi-level convolutional feature maps be effectively mimicked in a unified manner. Experiments demonstrate that the student models trained by the proposed approach can achieve substantially better results than their teacher models and have better generalization abilities.
Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICDAR2
2019 Mask R-CNN With Pyramid Attention Network for Scene Text Detection
abstract
In this paper, we present a new Mask R-CNN based text detection approach which can robustly detect multi-oriented and curved text from natural scene images in a unified manner. To enhance the feature representation ability of Mask R-CNN for text detection tasks, we propose to use the Pyramid Attention Network (PAN) as a new backbone network of Mask R-CNN. Experiments demonstrate that PAN can suppress false alarms caused by text-like backgrounds more effectively. Our proposed approach has achieved superior performance on both multi-oriented (ICDAR-2015, ICDAR-2017 MLT) and curved (SCUT-CTW1500) text detection benchmark tasks by only using single-scale and single-model testing.
Zhida Huang, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
WACV3
2019 Flow-guided feature propagation with occlusion aware detail enhancement for hand segmentation in egocentric videos
Lei Sun 0003, Qiang Huo
Comput. Vis. Image Underst.2
2019 An anchor-free region proposal network for Faster R-CNN-based text detection approaches
Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
Int. J. Document Anal. Recognit.2
2019 Improved localization accuracy by LocNet for Faster R-CNN based text detection in natural scene images
Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
Pattern Recognit.2
2018 A CNN-Based Approach to Detecting Text from Images of Whiteboards and Handwritten Notes
abstract
Detecting handwritten text from images of whiteboards and handwritten notes is an important yet under-researched topic. In this paper, we propose a convolutional neural network (CNN) based approach to address this problem. First, to detect text instances of different scales, a feature pyramid network is adopted as a backbone network to extract three feature maps of different scales from a given input image, where a scale-specific detection module is attached to each feature map. Then, for a pixel on each feature map, a detection module is used to predict whether there exists a text instance at its corresponding location in the input image. For positive prediction, the bounding box of the detected text segment and the links between the concerned pixel and its 8 neighbors on the feature map are predicted simultaneously. Based on the linkage information, text segments extracted from each feature map are grouped into text-lines respectively and wrongly grouped text-lines are separated by a graph-based text-line segmentation method. Finally, detection results from three different feature maps are aggregated by a skewed non-maximum suppression algorithm. Our proposed approach has achieved superior results on a testing set consisting of 285 natural scene images of whiteboards and handwritten notes.
Wei Jia 0003, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICFHR3
2018 DFF-DEN: Deep Feature Flow with Detail Enhancement Network for Hand Segmentation in Depth Video
abstract
Recently, researches are emerging to extend CNN-based segmentation approaches from still image to video. Directly applying per-frame based image segmentation networks on video is not efficient. To address this issue, a promising direction is to explore the video continuity. A state-of-the-art approach called Deep Feature Flow (DFF) runs the segmentation network only on sparse key frames and propagates feature maps to other frames via cross-frame motion. However, such approach does not work well for hand segmentation in a video as it is not robust to hand posture change. In this paper, we propose to incorporate a light-weight detail enhancement network (DEN) into the DFF framework to achieve robustness against cross-frame motion and hand posture change. Experimental results on a public depth video dataset, FingerPaint, demonstrate that our approach achieves higher segmentation accuracy than the DFF-based approach with similar speedups against the per-frame based video hand segmentation approach.
Lei Sun 0003, Qiang Huo
ICIP2
2017 A Compact CNN-DBLSTM Based Character Model for Online Handwritten Chinese Text Recognition
abstract
Recently, character model based on integrated convolutional neural network (CNN) and deep bidirectional long short-term memory (DBLSTM) has been demonstrated to be effective for online handwritten Chinese text recognition (HCTR). However, the reported CNN-DBLSTM topologies are too complex to be practically useful. In this paper, we propose a compact CNN-DBLSTM which has small footprint and low computation cost yet be able to accommodate multiple receptive fields for CNN-based feature extraction. By using the training set of a popular benchmark database, namely CASIA-OLHWDB, we trained a compact CNN-DBLSTM by a connectionist temporal classification (CTC) criterion with a multi-step training strategy. Combined this character model with a character trigram language model, our online HCTR system with a WFSTbased decoder has achieved state-of-the-art performance on both CASIA and ICDAR-2013 Chinese handwriting recognition competition test sets.
Kai Chen 0001, Haisong Ding, Lei Sun 0003, Sen Liang, Qiang Huo
ICDAR5
2017 An Open Vocabulary OCR System with Hybrid Word-Subword Language Models
abstract
The accuracy of a typical state-of-the-art optical character recognition (OCR) system benefits greatly from using a language model (LM). However, a conventional LM has a limited vocabulary, resulting in out-of-vocabulary (OOV) words that cannot be recognized by the OCR system. In this paper, we present an open vocabulary OCR system based on a hybrid LM. The vocabulary of the hybrid LM consists of both words and subwords. OOV words can be generated by combinations of subwords. A refined hybrid LM training scheme is applied by interpolating a standard hybrid LM, a word-based LM and a subword-based LM. An efficient word combination method is performed by modeling optional space symbols in a decoding network. The overall system deals with OOV words in a general, data-driven and language-independent way. We conduct experiments on an English handwriting OCR task. Evaluations on three testing sets demonstrate that the OCR system with the proposed method achieves a word error rate of 33.4% on an OOV-only testing set, yet without degrading the recognition accuracies on the other two testing sets mainly consisting of in-vocabulary words.
Wenping Hu, Kai Chen 0001, Lei Sun 0003, Sen Liang, Xiongjian Mo, Qiang Huo
ICDAR4
2017 A Compact CNN-DBLSTM Based Character Model for Offline Handwriting Recognition with Tucker Decomposition
abstract
Recently, character model based on integrated convolutional neural network (CNN) and deep bidirectional long short-term memory (DBLSTM) has achieved excellent performance for offline handwriting recognition (HWR). To deploy CNN-DBLSTM model in products, it is necessary to reduce the footprint and runtime latency as much as possible. In this paper, we study two methods to compress the CNN part: (1) Use Tucker decomposition to decompose pre-trained weights with low-rank approximation, followed by fine-tuning; (2) Use grouped convolution to construct sparse connections in channel domain. Experiments have been conducted on a large-scale offline English HWR task to compare the effectiveness of the above two techniques. Our results show that using Tucker decomposition alone offers a good solution to building a compact CNN-DBLSTM model which can reduce significantly both the footprint and latency yet without degrading recognition accuracy.
Haisong Ding, Kai Chen 0001, Lei Sun 0003, Sen Liang, Qiang Huo
ICDAR5
2017 Sequence Discriminative Training for Offline Handwriting Recognition by an Interpolated CTC and Lattice-Free MMI Objective Function
abstract
We study two sequence discriminative training criteria, i.e., Lattice-Free Maximum Mutual Information (LFMMI) and Connectionist Temporal Classification (CTC), for end-to-end training of Deep Bidirectional Long Short-Term Memory (DBLSTM) based character models of two offline English handwriting recognition systems with an input feature vector sequence extracted by Principal Component Analysis (PCA) and Convolutional Neural Network (CNN), respectively. We observe that refining CTC-trained PCA-DBLSTM model with an interpolated CTC and LFMMI objective function ("CTC+LFMMI") for several additional iterations achieves a relative Word Error Rate (WER) reduction of 24.6% and 13.9% on the public IAM test set and an in-house E2E test set, respectively. For a much better CTC-trained CNN-DBLSTM system, the proposed "CTC+LFMMI" method achieves a relative WER reduction of 19.6% and 8.3% on the above two test sets, respectively.
Wenping Hu, Kai Chen 0001, Haisong Ding, Lei Sun 0003, Sen Liang, Xiongjian Mo, Qiang Huo
ICDAR5
2017 A Robust Approach to Detecting Text from Images of Whiteboards and Handwritten Notes
abstract
Detecting text from the images of whiteboards and handwritten notes is an important yet under-researched topic. In this paper, we present a robust approach to solving this challenging problem as follows. First, given a color image, colorenhanced Contrasting Extremal Regions (CERs) are extracted from its grayscale image as candidate text connected components (CCs). Second, four shallow neural networks are used to preprune efficiently most of unambiguous non-text CCs. Third, a Fast R-CNN based approach is proposed to filter out remaining nontext CCs by leveraging contextual information and to estimate the corresponding text-line orientation in the position of each remaining text CC. Fourth, each pair of the remaining text CCs within a certain distance and orientation constraint are connected to construct a directed graph. Finally, based on the estimated textline orientations, candidate text-lines are generated easily by pruning greedily redundant edges in the graph to make each vertex have at most one direct successor and one direct predecessor, respectively. Our proposed approach has achieved promising results on an in-house testing set consisting of 285 camera-captured images of whiteboards and handwritten notes.
Wei Jia 0003, Lei Sun 0003, Zhuoyao Zhong, Xiongjian Mo, Guoen Ma, Qiang Huo
ICDAR2
2017 Improved Localization Accuracy by LocNet for Faster R-CNN Based Text Detection
abstract
Although Faster R-CNN based approaches have achieved promising results for text detection, their localization accuracy is not satisfactory in certain cases. In this paper, we propose to use a LocNet to improve the localization accuracy of a Faster R-CNN based text detector. Given a proposal generated by region proposal network (RPN), instead of predicting directly the bounding box coordinates of the concerned text instance, the proposal is enlarged to create a search region so that conditional probabilities to each row and column of this search region can be assigned, which are then used to infer accurately the concerned bounding box. Experiments demonstrate that the proposed approach boosts the localization accuracy for Faster R-CNN based text detection significantly. Consequently, our new text detector has achieved superior performance on ICDAR-2011, ICDAR-2013 and MULTILIGUL text detection benchmark tasks.
Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICDAR2
2017 Fingertip detection based on protuberant saliency from depth image
abstract
We propose a new approach for detecting a protuberant region from a depth image by leveraging a notion of protuberant saliency which describes how much the protuberant region stands out from its surroundings in the depth image. An intuitive and simple method is designed to calculate protuberant saliency from depth image, which can be used effectively together with the nearness information to detect the tip of a protuberant object such as a fingertip. We evaluate and compare our method with several state-of-the-art saliency methods for fingertip detection. Experimental results demonstrate that our method outperforms the comparing methods in terms of detection accuracy, being more robust against the rotation and isometric deformation of a fingertip region, the scale and depth noise issues, and the low resolution of a depth image.
Yuseok Ban, Lei Sun 0003, Qiang Huo
ICIP3
2016 Precise hand segmentation from a single depth image
abstract
We propose a new approach to segmenting a hand accurately from a single depth image. Given a depth image, we extract first a rough hand region of interest (RoI) including a hand and a part of an arm. Then, the RoI is partitioned into triangles by using a constrained Delaunay triangulation (CDT) approach from which hand segmentation proposals are generated. Each segmentation proposal is evaluated by a shallow convolutional neural network (CNN) which is trained as a regression function to predict a confidence score for each proposal. Finally, the segmentation proposal with the highest confidence score is selected as our hand segmentation result. To evaluate the effectiveness of our approach, we use a set of real data containing more than 370,000 frames of hand depth images collected from 40 subjects with large variations in pose, orientation and sensing distance. Compared with segmentation results achieved by a random decision forest (RDF) based approach, our approach achieves much higher accuracy.
Lei Sun 0003, Qiang Huo
ICPR2
2015 A robust approach for text detection from natural scene images
Lei Sun 0003, Qiang Huo, Wei Jia 0003, Kai Chen 0001
Pattern Recognit.1
2014 Robust Text Detection in Natural Scene Images by Generalized Color-Enhanced Contrasting Extremal Region and Neural Networks
abstract
This paper presents a robust text detection approach based on generalized color-enhanced contrasting extremal region (CER) and neural networks. Given a color natural scene image, six component-trees are built from its gray scale image, hue and saturation channel images in a perception-based illumination invariant color space, and their inverted images, respectively. From each component-tree, generalized color-enhanced CERs are extracted as character candidates. By using a "divide-and-conquer" strategy, each candidate image patch is labeled reliably by rules as one of five types, namely, Long, Thin, Fill, Square-large and Square-small, and classified as text or non-text by a corresponding neural network, which is trained by an ambiguity-free learning strategy. After pruning non-text components, repeating components in each component-tree are pruned by using color and area information to obtain a component graph, from which candidate text-lines are formed and verified by another set of neural networks. Finally, results from six component-trees are combined, and a post-processing step is used to recover lost characters and split text lines into words as appropriate. Our proposed method achieves 85.72% recall, 87.03% precision, and 86.37% F-score on ICDAR-2013 "Reading Text in Scene Images" test set.
Lei Sun 0003, Qiang Huo, Wei Jia 0003, Kai Chen 0001
ICPR1
2013 An Improved Component Tree Based Approach to User-Intention Guided Text Extraction from Natural Scene Images
abstract
We have proposed previously a component-tree based approach to user-intention guided text extraction from natural scene images. In this paper, in addition to improving the performance of text extraction algorithm for "swipe" gesture, the algorithm has also been extended to support a new mode of using "tap" gesture to indicate the intended text. Given a grayscale image, two component-trees are built and pre-pruned first by using a so-called contrasting extremal region (CER) criterion and simple rules of geometric features. The remaining nodes are enhanced by using color information in a perceptual color space. Then, a pre-trained neural network is used to classify a selected set of enhanced nodes as single-character or non-text objects. The remaining nodes are grouped into candidate text lines, where possible outliers are pruned in individual lines. Finally, the text line "swiped" or "tapped" by a user is selected as the target line and the intended text is extracted accordingly. The proposed algorithm has been evaluated on ICDAR-2003 benchmark dataset and a superior performance is achieved against the previous methods.
Lei Sun 0003, Qiang Huo
ICDAR1
2012 A component-tree based method for user-intention guided text extraction
Lei Sun 0003, Qiang Huo
ICPR1
2011 Snap and Translate Using Windows Phone
abstract
We have developed a prototype of a mobile app called "Snap and Translate" on "Windows Phone 7". A person who is reading an English menu/sign and wants a Chinese translation of an English word or phrase or paragraph can use a Windows Phone to snap an image of the text, tap the word or swipe the phrase or circle the paragraph with a finger, and get a Chinese translation displayed on the screen of the phone. This is enabled by seamless integration of three Microsoft technologies: intelligent text extraction, OCR, and machine translation based on a client-plus-cloud architecture. The current prototype also supports Chinese OCR plus Chinese-to-English translation. In this paper, we highlight the UI design of the system and the corresponding user-intention guided text extraction approach to achieving a compelling user experience.
Jun Du 0002, Qiang Huo, Lei Sun 0003
ICDAR3