VLDB 2026 Research / reviewers in the wild / expert
Zhanzhan Cheng
dblp:163/6485
· DBLP profile ↗
39ranked-venue papers
6as first author
26since 2021 · last 2025
0000-0002-5732-1513ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 3 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 16 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Real-Time Respiration Monitoring via Motion Artifact Suppression and Quality-Guided Peak DetectionabstractReal-time respiration monitoring faces several challenges including network latency in remote settings, limited computational resources, and increased motion artifacts. Although many existing non-contact respiration algorithms are designed for offline processing and thus overlook these limitations, real-time applications demand greater robustness, efficiency, and adaptability to dynamic conditions. In this study, a lightweight framework called Quality-Guided Respiration Monitoring (QGRM) is proposed. This framework integrates a two-stage motion artifact suppression module and a quality-guided peak detection (QGPD) module. The former enhances signal stability through FIR filtering and amplitude limiting, while the latter improves the estimation of the respiration rate by filtering false peaks based on amplitude and zero-crossing constraints. The experimental results obtained with both the public OVRM dataset and a self-constructed simulated dataset demonstrate that QGRM achieves superior accuracy and robustness compared to state-of-the-art methods. The dataset and code are available athttps://github.com/zxx5058/QGRM. Chenrui Niu, Zhanzhan Cheng, Nengfeng Qian, Changyin Wu, Guangtao Zhai, Menghan Hu |
IEEE Signal Process. Lett. | 3 |
| 2025 | Optimizing Video-Based Respiration Monitoring: Motion Artifact Reduction and Adaptive ROI SelectionabstractIn non-contact respiratory monitoring, reducing motion artifact and selecting the appropriate Region of Interest (ROI) pose significant challenges. Most motion artifact removal methods rely on signal periodicity assumptions, while respiratory signals usually are non-periodic in real-world scenarios. Existing automated ROI selection approaches are mostly primarily impacted by the texture of clothing, absence of chest landmarks, and obstruction of face. To improve the quality of respiratory signals, in this study, we propose a framework for automatic respiratory ROI selection based on video, namely, Optimizing Video-based Respiration Monitoring (OVRM), which consists of peak-trough adaptive motion artifact removal and characteristic-driven adaptive ROI selection. This motion artifact removal strategy removes motion artifacts by using a dynamic ratio-based judgment mechanism, and reconstructs signals using sinusoidal interpolation. The adaptive ROI method scores signals based on periodicity, similarity, smoothness, and energy, selecting the highest-scoring blocks as the ROIs to match respiratory signals efficiently. Experimental results, validated across four datasets, demonstrate that OVRM effectively reduces signal noise caused by subject movement and outperforms state-of-the-art non-contact respiratory monitoring algorithms. The dataset and code are publicly available at:https://github.com/zxx5058/OVRM. Xudong Tan, Mei Zhou, Menghan Hu, Zhanzhan Cheng, Nengfeng Qian, Changyin Wu, Guangtao Zhai, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 6 |
| 2024 | Reading order detection in visually-rich documents with multi-modal layout-aware relation prediction
Liang Qiao 0001, Zhanzhan Cheng, Yunlu Xu, Xi Li 0001 |
Pattern Recognit. | 3 |
| 2023 | Few-Shot Class-Incremental Learning via Class-Aware Bilateral DistillationabstractFew-Shot Class-Incremental Learning (FSCIL) aims to continually learn novel classes based on only few training samples, which poses a more challenging task than the well-studied Class-Incremental Learning (CIL) due to data scarcity. While knowledge distillation, a prevailing technique in CIL, can alleviate the catastrophic forgetting of older classes by regularizing outputs between current and previous model, it fails to consider the overfitting risk of novel classes in FSCIL. To adapt the powerful distillation technique for FSCIL, we propose a novel distillation structure, by taking the unique challenge of overfitting into account. Concretely, we draw knowledge from two complementary teachers. One is the model trained on abundant data from base classes that carries rich general knowledge, which can be leveraged for easing the overfitting of current novel classes. The other is the updated model from last incremental session that contains the adapted knowledge of previous novel classes, which is used for alleviating their forgetting. To combine the guidances, an adaptive strategy conditioned on the class-wise semantic similarities is introduced. Besides, for better preserving base class knowledge when accommodating novel concepts, we adopt a two-branch network with an attention-based aggregation module to dynamically merge predictions from two complementary branches. Extensive experiments on 3 popular FSCIL datasets: mini-ImageNet, CIFAR100 and CUB200 validate the effectiveness of our method by surpassing existing works by a significant margin. Code is available at https://github.com/LinglanZhao/BiDistFSCIL. Linglan Zhao, Jing Lu 0004, Yunlu Xu, Zhanzhan Cheng, Dashan Guo, Xiangzhong Fang |
CVPR | 4 |
| 2023 | HyperMatch: Noise-Tolerant Semi-Supervised Learning via Relaxed Contrastive ConstraintabstractRecent developments of the application of Contrastive Learning in Semi-Supervised Learning (SSL) have demonstrated significant advancements, as a result of its exceptional ability to learn class-aware cluster representations and the full exploitation of massive unlabeled data. However, mismatched instance pairs caused by inaccurate pseudo labels would assign an unlabeled instance to the incorrect class in feature space, hence exacerbating SSL's renowned confirmation bias. To address this issue, we introduced a novel SSL approach, HyperMatch, which is a plugin to several SSL designs enabling noise-tolerant utilization of unlabeled data. In particular, confidence predictions are combined with semantic similarities to generate a more objective class distribution, followed by a Gaussian Mixture Model to divide pseudo labels into a ‘confident’ and a ‘less confident’ subset. Then, we introduce Relaxed Contrastive Loss by assigning the ‘less-confident’ samples to a hyper-class, i.e. the union of top-K nearest classes, which effectively regularizes the interference of incorrect pseudo labels and even increases the probability of pulling a ‘less confident’ sample close to its true class. Experiments and in-depth studies demonstrate that HyperMatch delivers remarkable state-of-the-art performance, outperforming Fix-Match on CIFAR100 with 400 and 2500 labeled samples by 11.86% and 4.88%, respectively. Beitong Zhou, Jing Lu 0004, Kerui Liu, Yunlu Xu, Zhanzhan Cheng |
CVPR | 5 |
| 2023 | Rethinking Self-Supervision for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) focuses on progressively absorbing new concepts given only limited training data. For tackling this challenge, several recent FS-CIL works resort to pre-training models with Self-Supervised Learning (SSL) to obtain features that can generalize well to new classes. However, to avoid overfitting and catastrophic forgetting, previous works only leverage SSL in the base session and keep all or most parameters fixed in incremental sessions, resulting in inadequate adaptation to novel classes. Thus, in this paper, we explore the setting where more parameters can be updated for adapting to novel concepts, and discover that the model pre-trained with SSL leads to degraded performance even compared to that without SSL. It can be attributed to the severer forgetting of base class knowledge. To address this issue, we propose an imprinting-based distillation module for effectively regularizing the adaption process, and a mathematically provable routing strategy for further improved results. The effectiveness of our approach is verified on 3 popular FSCIL benchmarks by significantly outperforming previous methods. Linglan Zhao, Jing Lu 0004, Zhanzhan Cheng, Xiangzhong Fang |
ICME | 3 |
| 2023 | Divide Rows and Conquer Cells: Towards Structure Recognition for Large TablesabstractRecent advanced Table Structure Recognition (TSR) models adopt image-to-text solutions to parse table structure. These methods can be formulated as image caption problem, i.e., input a single-table image and output table structure description in a specific text format, e.g., HTML. With the impressive success of Transformer in text generation tasks, these methods use Transformer architecture to predict HTML table text in an autoregressive manner. However, tables always emerge with a large variety of shapes and sizes. Autoregressive models usually suffer from the error accumulation problem as the length of predicted text increases, which results in unsatisfactory performance for large tables. In this paper, we propose a novel image-to-text based TSR method that relieves error accumulation problems and improves performance noticeably. At the core of our method is a cascaded two-step decoder architecture with the former decoder predicting HTML table row tags non-autoregressively and the latter predicting HTML table cell tags of each row in a semi-autoregressive manner. Compared with existing methods that predict HTML text autoregressively, the superiority of our row-to-cell progressive table parsing is twofold: (1) it generates an HTML tag sequence with a vertical-and-horizontal two-step `scanning', which better fits the inherent 2D structure of image data, (2) it performs substantially better for large tables (long sequence prediction) since it alleviates error accumulation problem specific to autoregressive models. Extensive experiments demonstrate that our method achieves competitive performance on three public benchmarks. Huawen Shen, Yu Zhou 0015, Zhanzhan Cheng |
IJCAI | 7 |
| 2022 | PMAL: Open Set Recognition via Robust Prototype MiningabstractOpen Set Recognition (OSR) has been an emerging topic. Besides recognizing predefined classes, the system needs to reject the unknowns. Prototype learning is a potential manner to handle the problem, as its ability to improve intra-class compactness of representations is much needed in discrimination between the known and the unknowns. In this work, we propose a novel Prototype Mining And Learning (PMAL) framework. It has a prototype mining mechanism before the phase of optimizing embedding space, explicitly considering two crucial properties, namely high-quality and diversity of the prototype set. Concretely, a set of high-quality candidates are firstly extracted from training samples based on data uncertainty learning, avoiding the interference from unexpected noise. Considering the multifarious appearance of objects even in a single category, a diversity-based strategy for prototype set filtering is proposed. Accordingly, the embedding space can be better optimized to discriminate therein the predefined classes and between known and unknowns. Extensive experiments verify the two good characteristics (i.e., high-quality and diversity) embraced in prototype mining, and show the remarkable performance of the proposed framework compared to state-of-the-arts. Jing Lu 0004, Yunlu Xu, Zhanzhan Cheng |
AAAI | 4 |
| 2022 | Flooding-X: Improving BERT's Resistance to Adversarial Attacks via Loss-Restricted Fine-TuningabstractQin Liu, Rui Zheng, Bao Rong, Jingyi Liu, ZhiHua Liu, Zhanzhan Cheng, Liang Qiao, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Qin Liu 0010, Bao Rong, Zhanzhan Cheng, Liang Qiao 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 6 |
| 2022 | MINER: Improving Out-of-Vocabulary Named Entity Recognition from an Information Theoretic PerspectiveabstractXiao Wang, Shihan Dou, Limao Xiong, Yicheng Zou, Qi Zhang, Tao Gui, Liang Qiao, Zhanzhan Cheng, Xuanjing Huang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Xiao Wang 0001, Shihan Dou, Limao Xiong, Yicheng Zou, Qi Zhang 0001, Tao Gui, Liang Qiao 0001, Zhanzhan Cheng, Xuanjing Huang 0001 |
ACL (1) | 8 |
| 2022 | Read Extensively, Focus Smartly: A Cross-document Semantic Enhancement Method for Visual Documents NERabstractThe introduction of multimodal information and pretraining technique significantly improves entity recognition from visually-rich documents. However, most of the existing methods pay unnecessary attention to irrelevant regions of the current document while ignoring the potentially valuable information in related documents. To deal with this problem, this work proposes a cross-document semantic enhancement method, which consists of two modules: 1) To prevent distractions from irrelevant regions in the current document, we design a learnable attention mask mechanism, which is used to adaptively filter redundant information in the current document. 2) To further enrich the entity-related context, we propose a cross-document information awareness technique, which enables the model to collect more evidence across documents to assist in prediction. The experimental results on two documents understanding benchmarks covering eight languages demonstrate that our method outperforms the SOTA methods. Jun Zhao 0019, Wenyu Zhan, Tao Gui, Qi Zhang 0001, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu |
COLING | 7 |
| 2022 | Dynamic Low-Resolution Distillation for Cost-Efficient End-to-End Text Spotting
Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xi Li 0001 |
ECCV (28) | 3 |
| 2022 | Distilling Object Detectors with Global Knowledge
Sanli Tang, Zhanzhan Cheng, Jing Lu 0004, Yunlu Xu |
ECCV (9) | 3 |
| 2022 | Active Model Adaptation Under Unknown ShiftabstractSuccessful machine learning typically relies on fixed data distribution. However, due to unforeseen situations in the open world, distribution shift often occurs in applications. For instance, in the image recognition task, an unpredictable distributional shift may occur due to changes in background or lighting. Furthermore, to alleviate the harm of distribution shift, the resource budget is not infinite and often constrained. To cope with such a novel problem Resource Constrained Adaptation under Unknown Shift, in this paper we study active model adaptation both theoretically and empirically. First, we present a generalization analysis of active model adaptation for distribution shift. In theory, we show that active model adaptation could improve the generalization error from O(1/N) to O(1/N), with only a few queried samples. Second, based on the theoretical analysis, we present a systemic solution Auto, consisting of three sub-steps, that is, distribution tracking, sample selection and model adaptation. Specifically, we design a shifted distribution detection module to locate the distributional shifted samples. To fit the labeling budget, we employ a core-set algorithm to enhance the informativeness of the selected samples. Finally, we update the model through the newly queried labeled data. We conduct empirical studies of nine existing active strategies on diverse real world data sets and the results show that Auto could remarkably outperform all the baselines. Jie-Jing Shao, Yunlu Xu, Zhanzhan Cheng, Yufeng Li 0008 |
KDD | 3 |
| 2022 | End-to-End Compound Table Understanding with Multi-Modal ModelingabstractTable is a widely used data form in webpages, spreadsheets, or PDFs to organize and present structural data. Although studies on table structure recognition have been successfully used to convert image-based tables into digital structural formats, solving many real problems still relies on further understanding of the table, such as cell relationship extraction. The current datasets related to table understanding are all based on the digit format. To boost research development, we release a new benchmark named ComFinTab with rich annotations that support both table recognition and understanding tasks. Unlike previous datasets containing the basic tables, ComFinTab contains a large ratio of compound tables, which is much more challenging and requires methods using multiple information sources. Based on the dataset, we also propose a uniform, concise task form with the evaluation metric to better evaluate the model's performance on the table understanding task in compound tables. Finally, a framework named CTUNet is proposed to integrate the compromised visual, semantic, and position features with a graph attention network, which can solve the table recognition task and the challenging table understanding task as a whole. Experimental results compared with some previous advanced table understanding methods demonstrate the effectiveness of our proposed model. Code and dataset are available at \urlhttps://github.com/hikopensource/DAVAR-Lab-OCR. Zaisheng Li, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xi Li 0001 |
ACM Multimedia | 5 |
| 2022 | DavarOCR: A Toolbox for OCR and Multi-Modal Document UnderstandingabstractThis paper presents DavarOCR, an open-source toolbox for OCR and document understanding tasks. DavarOCR currently implements 19 advanced algorithms, covering 9 different task forms. DavarOCR provides detailed usage instructions and the trained models for each algorithm. Compared with the previous open-source OCR toolbox, DavarOCR has relatively more complete support for the sub-tasks of the cutting-edge technology of document understanding. In order to promote the development and application of OCR technology in academia and industry, we pay more attention to the use of modules that different sub-domains of technology can share. DavarOCR is publicly released at https://github.com/hikopensource/Davar-Lab-OCR. Liang Qiao 0001, Zaisheng Li, Baorui Zou, Dashan Guo, Yingda Xu, Yunlu Xu, Zhanzhan Cheng |
ACM Multimedia | 11 |
| 2021 | SPIN: Structure-Preserving Inner Offset Network for Scene Text RecognitionabstractArbitrary text appearance poses a great challenge in scene text recognition tasks. Existing works mostly handle with the problem in consideration of the shape distortion, including perspective distortions, line curvature or other style variations. Rectification (i.e., spatial transformers) as the preprocessing stage is one popular approach and extensively studied. However, chromatic difficulties in complex scenes have not been paid much attention on. In this work, we introduce a new learnable geometric-unrelated rectification, Structure-Preserving Inner Offset Network (SPIN), which allows the color manipulation of source data within the network. This differentiable module can be inserted before any recognition architecture to ease the downstream tasks, giving neural networks the ability to actively transform input intensity rather than only the spatial rectification. It can also serve as a complementary module to known spatial transformations and work in both independent and collaborative ways with them. Extensive experiments show the proposed transformation outperforms existing rectification networks and has comparable performance among the state-of-the-arts. Chengwei Zhang 0003, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Fei Wu 0001, Futai Zou |
AAAI | 3 |
| 2021 | MANGO: A Mask Attention Guided One-Stage Scene Text SpotterabstractRecently end-to-end scene text spotting has become a popular research topic due to its advantages of global optimization and high maintainability in real applications. Most methods attempt to develop various region of interest (RoI) operations to concatenate the detection part and the sequence recognition part into a two-stage text spotting framework. However, in such framework, the recognition part is highly sensitive to the detected results (e.g., the compactness of text contours). To address this problem, in this paper, we propose a novel Mask AttentioN Guided One-stage text spotting framework named MANGO, in which character sequences can be directly recognized without RoI operation. Concretely, a position-aware mask attention module is developed to generate attention weights on each text instance and its characters. It allows different text instances in an image to be allocated on different feature map channels which are further grouped as a batch of instance features. Finally, a lightweight sequence decoder is applied to generate the character sequences. It is worth noting that MANGO inherently adapts to arbitrary-shaped text spotting and can be trained end-to-end with only coarse position information (e.g., rectangular bounding box) and text annotations. Experimental results show that the proposed method achieves competitive and even new state-of-the-art performance on both regular and irregular text spotting benchmarks, i.e., ICDAR 2013, ICDAR 2015, Total-Text, and SCUT-CTW1500. Liang Qiao 0001, Zhanzhan Cheng, Yunlu Xu, Shiliang Pu, Fei Wu 0001 |
AAAI | 3 |
| 2021 | A Strong Baseline for Semi-Supervised Incremental Few-Shot Learning
Linglan Zhao, Dashan Guo, Yunlu Xu, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xiangzhong Fang |
BMVC | 5 |
| 2021 | ICDAR 2021 Competition on Scene Video Text Spotting
Zhanzhan Cheng, Jing Lu 0004, Baorui Zou, Shuigeng Zhou, Fei Wu 0001 |
ICDAR (4) | 1 |
| 2021 | Reciprocal Feature Learning via Explicit and Implicit Tasks in Scene Text Recognition
Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Wenqi Ren, Fei Wu 0001, Wenming Tan |
ICDAR (1) | 3 |
| 2021 | LGPMA: Complicated Table Structure Recognition with Local and Global Pyramid Mask Alignment
Liang Qiao 0001, Zaisheng Li, Zhanzhan Cheng, Peng Zhang 0075, Shiliang Pu, Wenqi Ren, Wenming Tan, Fei Wu 0001 |
ICDAR (1) | 3 |
| 2021 | VSR: A Unified Framework for Document Layout Analysis Combining Vision, Semantics and Relations
Peng Zhang 0075, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Fei Wu 0001 |
ICDAR (1) | 4 |
| 2021 | Towards Robust Model Reuse in the Presence of Latent DomainsabstractModel reuse tries to adapt well pre-trained models to a new target task, without access of raw data. It attracts much attention since it reduces the learning resources. Previous model reuse studies typically operate in a single-domain scenario, i.e., the target samples arise from one single domain. However, in practice the target samples often arise from multiple latent or unknown domains, e.g., the images for cars may arise from latent domains such as photo, line drawing, cartoon, etc. The methods based on single-domain may no longer be feasible for multiple latent domains and may sometimes even lead to performance degeneration. To address the above issue, in this paper we propose the MRL (Model Reuse for multiple Latent domains) method. Both domain characteristics and pre-trained models are considered for the exploration of instances in the target task. Theoretically, the overall considerations are packed in a bi-level optimization framework with a reliable generalization. Moreover, through an ensemble of multiple models, the model robustness is improved with a theoretical guarantee. Empirical results on diverse real-world data sets clearly validate the effectiveness of proposed algorithms. Jie-Jing Shao, Zhanzhan Cheng, Yufeng Li 0008, Shiliang Pu |
IJCAI | 2 |
| 2021 | STEP: Out-of-Distribution Detection in the Presence of Limited In-Distribution Labeled DataabstractExisting semi-supervised learning (SSL) studies typically assume that unlabeled and test data are drawn from the same distribution as labeled data. However, in many real-world applications, it is desirable to have SSL algorithms that not only classify the samples drawn from the same distribution of labeled data but also detect out-of-distribution (OOD) samples drawn from an unknown distribution. In this paper, we study a setting called semi-supervised OOD detection. Two main challenges compared with previous OOD detection settings are i) the lack of labeled data and in-distribution data; ii) OOD samples could be unseen during training. Efforts on this direction remain limited. In this paper, we present an approach STEP significantly improving OOD detection performance by introducing a new technique: Structure-Keep Unzipping. It learns a new representation space in which OOD samples could be separated well. An efficient optimization algorithm is derived to solve the objective. Comprehensive experiments across various OOD detection benchmarks clearly show that our STEP approach outperforms other methods by a large margin and achieves remarkable detection performance on several benchmarks. Zhi Zhou 0007, Lan-Zhe Guo, Zhanzhan Cheng, Yufeng Li 0008, Shiliang Pu |
NeurIPS | 3 |
| 2021 | FREE: A Fast and Robust End-to-End Video Text SpotterabstractCurrently, video text spotting tasks usually fall into the four-staged pipeline: detecting text regions in individual images, recognizing localized text regions frame-wisely, tracking text streams and post-processing to generate final results. However, they may suffer from the huge computational cost as well as sub-optimal results due to the interferences of low-quality text and the none-trainable pipeline strategy. In this article, we propose a fast and robust end-to-end video text spotting framework named FREE by only recognizing the localized text stream one-time instead of frame-wise recognition. Specifically, FREE first employs a well-designed spatial-temporal detector that learns text locations among video frames. Then a novel text recommender is developed to select the highest-quality text from text streams for recognizing. Here, the recommender is implemented by assembling text tracking, quality scoring and recognition into a trainable module. It not only avoids the interferences from the low-quality text but also dramatically speeds up the video text spotting. FREE unites the detector and recommender into a whole framework, and helps achieve global optimization. Besides, we collect a large scale video text dataset for promoting the video text spotting community, containing 100 videos from 21 real-life scenarios. Extensive experiments on public benchmarks show our method greatly speeds up the text spotting process, and also achieves the remarkable state-of-the-art. Zhanzhan Cheng, Jing Lu 0004, Baorui Zou, Liang Qiao 0001, Yunlu Xu, Shiliang Pu, Fei Wu 0001, Shuigeng Zhou |
IEEE Trans. Image Process. | 1 |
| 2020 | Text Perceptron: Towards End-to-End Arbitrary-Shaped Text SpottingabstractMany approaches have recently been proposed to detect irregular scene text and achieved promising results. However, their localization results may not well satisfy the following text recognition part mainly because of two reasons: 1) recognizing arbitrary shaped text is still a challenging task, and 2) prevalent non-trainable pipeline strategies between text detection and text recognition will lead to suboptimal performances. To handle this incompatibility problem, in this paper we propose an end-to-end trainable text spotting approach named Text Perceptron. Concretely, Text Perceptron first employs an efficient segmentation-based text detector that learns the latent text reading order and boundary information. Then a novel Shape Transform Module (abbr. STM) is designed to transform the detected feature regions into regular morphologies without extra parameters. It unites text detection and the following recognition part into a whole framework, and helps the whole network achieve global optimization. Experiments show that our method achieves competitive performance on two standard text benchmarks, i.e., ICDAR 2013 and ICDAR 2015, and also obviously outperforms existing methods on irregular text benchmarks SCUT-CTW1500 and Total-Text. Liang Qiao 0001, Sanli Tang, Zhanzhan Cheng, Yunlu Xu, Shiliang Pu, Fei Wu 0001 |
AAAI | 3 |
| 2020 | Text Recognition in Real Scenarios with a Few Labeled SamplesabstractScene text recognition (STR) is still a hot research topic in computer vision field due to its various applications. Existing works mainly focus on learning a general model with a huge number of synthetic text images to recognize unconstrained scene texts, and have achieved substantial progress. However, these methods are not quite applicable in many real-world scenarios where 1) high recognition accuracy is required, while 2) labeled samples are lacked. To tackle this challenging problem, this paper proposes a few-shot adversarial sequence domain adaptation (FASDA) approach to build sequence adaptation between the synthetic source domain (with many synthetic labeled samples) and a specific target domain (with only some or a few real labeled samples). This is done by simultaneously learning each character's feature representation with an attention mechanism and establishing the corresponding character-level latent subspace with adversarial learning. Our approach can maximize the character-level confusion between the source domain and the target domain, thus achieves the sequence-level adaptation with even a small number of labeled samples in the target domain. Extensive experiments on various datasets show that our method significantly outperforms the finetuning scheme, and obtains comparable performance to the state-of-the-art STR methods. Jinghuang Lin, Zhanzhan Cheng, Fan Bai 0001, Shiliang Pu, Shuigeng Zhou |
ICPR | 2 |
| 2020 | Recognizing Multiple Text Sequences from an Image by Pure End-to-End Learning
Zhenlong Xu, Shuigeng Zhou, Fan Bai 0001, Zhanzhan Cheng, Shiliang Pu |
ICPR | 4 |
| 2020 | TRIE: End-to-End Text Reading and Information Extraction for Document UnderstandingabstractSince real-world ubiquitous documents (e.g., invoices, tickets, resumes and leaflets) contain rich information, automatic document image understanding has become a hot topic. Most existing works decouple the problem into two separate tasks, (1) text reading for detecting and recognizing texts in images and (2) information extraction for analyzing and extracting key elements from previously extracted plain text.However, they mainly focus on improving information extraction task, while neglecting the fact that text reading and information extraction are mutually correlated. In this paper, we propose a unified end-to-end text reading and information extraction network, where the two tasks can reinforce each other. Specifically, the multimodal visual and textual features of text reading are fused for information extraction and in turn, the semantics in information extraction contribute to the optimization of text reading. On three real-world datasets with diverse document images (from fixed layout to variable layout, from structured text to semi-structured text), our proposed method significantly outperforms the state-of-the-art methods in both efficiency and accuracy. Peng Zhang 0075, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu 0004, Liang Qiao 0001, Fei Wu 0001 |
ACM Multimedia | 3 |
| 2019 | Segregated Temporal Assembly Recurrent Networks for Weakly Supervised Multiple Action DetectionabstractThis paper proposes a segregated temporal assembly recurrent (STAR) network for weakly-supervised multiple action detection. The model learns from untrimmed videos with only supervision of video-level labels and makes prediction of intervals of multiple actions. Specifically, we first assemble video clips according to class labels by an attention mechanism that learns class-variable attention weights and thus helps the noise relieving from background or other actions. Secondly, we build temporal relationship between actions by feeding the assembled features into an enhanced recurrent neural network. Finally, we transform the output of recurrent neural network into the corresponding action distribution. In order to generate more precise temporal proposals, we design a score term called segregated temporal gradient-weighted class activation mapping (ST-GradCAM) fused with attention weights. Experiments on THUMOS’14 and ActivityNet1.3 datasets show that our approach outperforms the state-of-theart weakly-supervised method, and performs at par with the fully-supervised counterparts. Yunlu Xu, Chengwei Zhang 0003, Zhanzhan Cheng, Jianwen Xie, Shiliang Pu, Fei Wu 0001 |
AAAI | 3 |
| 2019 | You Only Recognize Once: Towards Fast Video Text SpottingabstractVideo text spotting is still an important research topic due to its various real-applications. Previous approaches usually fall into the four-staged pipeline: text detection in individual images, frame-wisely recognizing localized text regions, tracking text streams and generating final results with complicated post-processing skills, which might suffer from the huge computational cost as well as the interferences of low-quality text. In this paper, we propose a fast and robust video text spotting framework by only recognizing the localized text one-time instead of frame-wisely recognition. Specifically, we first obtain text regions in videos with a well-designed spatial-temporal detector. Then we concentrate on developing a novel text recommender for selecting the highest-quality text from text streams and only recognizing the selected ones. Here, the recommender assembles text tracking, quality scoring and recognition into an end-to-end trainable module, which not only avoids the interferences from low-quality text but also dramatically speeds up the video text spotting process. In addition, we collect a larger scale video text dataset (LSVTD) for promoting the video text spotting community, which contains 100 text videos from 22 different real-life scenarios. Extensive experiments on two public benchmarks show that our method greatly speeds up the recognition process averagely by 71 times compared with the frame-wise manner, and also achieves the remarkable state-of-the-art. Zhanzhan Cheng, Jing Lu 0004, Shiliang Pu, Fei Wu 0001, Shuigeng Zhou |
ACM Multimedia | 1 |
| 2019 | Adversarial Seeded Sequence Growing for Weakly-Supervised Temporal Action LocalizationabstractTemporal action localization is an important yet challenging research topic due to its various applications. Since the frame-level or segment-level annotations of untrimmed videos require amounts of labor expenditure, studies on the weakly-supervised action detection have been springing up. However, most of existing frameworks rely on Class Activation Sequence (CAS) to localize actions by minimizing the video-level classification loss, which exploits the most discriminative parts of actions but ignores the minor regions. In this paper, we propose a novel weakly-supervised framework by adversarial learning of two modules for eliminating such demerits. Specifically, the first module is designed as a well-designed Seeded Sequence Growing (SSG) Network for progressively extending seed regions (namely the highly reliable regions initialized by a CAS-based framework) to their expected boundaries. The second module is a specific classifier for mining trivial or incomplete action regions, which is trained on the shared features after erasing the seeded regions activated by SSG. In this way, a whole network composed of these two modules can be trained in an adversarial manner. The goal of the adversary is to mine features that are difficult for the action classifier. That is, erasion from SSG will force the classifier to discover minor or even new action regions on the input feature sequence, and the classifier will drive the seeds to grow, alternately. At last, we could obtain the action locations and categories from the well-trained SSG and the classifier. Extensive experiments on two public benchmarks THUMOS'14 and ActivityNet1.3 demonstrate the impressive performance of our proposed method compared with the state-of-the-arts. Chengwei Zhang 0003, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Fei Wu 0001, Futai Zou |
ACM Multimedia | 3 |
| 2019 | REAPS: Towards Better Recognition of Fine-Grained Images by Region Attending and Part Sequencing
Peng Zhang 0075, Zhanzhan Cheng, Shuigeng Zhou |
PRCV (1) | 3 |
| 2018 | Edit Probability for Scene Text RecognitionabstractWe consider the scene text recognition problem under the attention-based encoder-decoder framework, which is the state of the art. The existing methods usually employ a frame-wise maximal likelihood loss to optimize the models. When we train the model, the misalignment between the ground truth strings and the attention's output sequences of probability distribution, which is caused by missing or superfluous characters, will confuse and mislead the training process, and consequently make the training costly and degrade the recognition accuracy. To handle this problem, we propose a novel method called edit probability (EP) for scene text recognition. EP tries to effectively estimate the probability of generating a string from the output sequence of probability distribution conditioned on the input image, while considering the possible occurrences of missing/superfluous characters. The advantage lies in that the training process can focus on the missing, superfluous and unrecognized characters, and thus the impact of the misalignment problem can be alleviated or even overcome. We conduct extensive experiments on standard benchmarks, including the IIIT-5K, Street View Text and ICDAR datasets. Experimental results show that the EP can substantially boost scene text recognition performance. Fan Bai 0001, Zhanzhan Cheng, Shiliang Pu, Shuigeng Zhou |
CVPR | 2 |
| 2018 | AON: Towards Arbitrarily-Oriented Text RecognitionabstractRecognizing text from natural images is a hot research topic in computer vision due to its various applications. Despite the enduring research of several decades on optical character recognition (OCR), recognizing texts from natural images is still a challenging task. This is because scene texts are often in irregular (e.g. curved, arbitrarily-oriented or seriously distorted) arrangements, which have not yet been well addressed in the literature. Existing methods on text recognition mainly work with regular (horizontal and frontal) texts and cannot be trivially generalized to handle irregular texts. In this paper, we develop the arbitrary orientation network (AON) to directly capture the deep features of irregular texts, which are combined into an attention-based decoder to generate character sequence. The whole network can be trained end-to-end by using only images and word-level annotations. Extensive experiments on various benchmarks, including the CUTE80, SVT-Perspective, IIIT5k, SVT and ICDAR datasets, show that the proposed AON-based method achieves the-state-of-the-art performance in irregular datasets, and is comparable to major existing methods in regular datasets. Zhanzhan Cheng, Yangliu Xu, Fan Bai 0001, Shiliang Pu, Shuigeng Zhou |
CVPR | 1 |
| 2018 | Effectively Identifying Compound-Protein Interactions by Learning from Positive and Unlabeled ExamplesabstractPrediction of compound-protein interactions (CPIs) is to find new compound-protein pairs where a protein is targeted by at least a compound, which is a crucial step in new drug design. Currently, a number of machine learning based methods have been developed to predict new CPIs in the literature. However, as there is not yet any publicly available set of validated negative CPIs, most existing machine learning based approaches use the unknown interactions (not validated CPIs) selected randomly as the negative examples to train classifiers for predicting new CPIs. Obviously, this is not quite reasonable and unavoidably impacts the CPI prediction performance. In this paper, we simply take the unknown CPIs as unlabeled examples, and propose a new method called PUCPI (the abbreviation of PU learning for Compound-Protein Interaction identification) that employs biased-SVM (Support Vector Machine) to predict CPIs using only positive and unlabeled examples. PU learning is a class of learning methods that leans from positive and unlabeled (PU) samples. To the best of our knowledge, this is the first work that identifies CPIs using only positive and unlabeled examples. We first collect known CPIs as positive examples and then randomly select compound-protein pairs not in the positive set as unlabeled examples. For each CPI/compound-protein pair, we extract protein domains as protein features and compound substructures as chemical features, then take the tensor product of the corresponding compound features and protein features as the feature vector of the CPI/compound-protein pair. After that, biased-SVM is employed to train classifiers on different datasets of CPIs and compound-protein pairs. Experiments over various datasets show that our method outperforms six typical classifiers, including random forest, L1- and L2-regularized logistic regression, naive Bayes, SVM and k-nearest neighbor (kNN), and three types of existing CPI prediction models. More information can be found at http://admis.fudan.edu.cn/projects/pucpi.html. Zhanzhan Cheng, Shuigeng Zhou, Yang Wang 0100, Hui Liu 0026, Jihong Guan, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | Focusing Attention: Towards Accurate Text Recognition in Natural ImagesabstractScene text recognition has been a hot research topic in computer vision due to its various applications. The state of the art is the attention-based encoder-decoder framework that learns the mapping between input images and output sequences in a purely data-driven way. However, we observe that existing attention-based methods perform poorly on complicated and/or low-quality images. One major reason is that existing methods cannot get accurate alignments between feature areas and targets for such images. We call this phenomenon “attention drift”. To tackle this problem, in this paper we propose the FAN (the abbreviation of Focusing Attention Network) method that employs a focusing attention mechanism to automatically draw back the drifted attention. FAN consists of two major components: an attention network (AN) that is responsible for recognizing character targets as in the existing methods, and a focusing network (FN) that is responsible for adjusting attention by evaluating whether AN pays attention properly on the target areas in the images. Furthermore, different from the existing methods, we adopt a ResNet-based network to enrich deep representations of scene text images. Extensive experiments on various benchmarks, including the IIIT5k, SVT and ICDAR datasets, show that the FAN method substantially outperforms the existing methods. Zhanzhan Cheng, Fan Bai 0001, Yunlu Xu, Shiliang Pu, Shuigeng Zhou |
ICCV | 1 |
| 2015 | Exploiting topic modeling to boost metagenomic reads binningabstractBACKGROUND: With the rapid development of high-throughput technologies, researchers can sequence the whole metagenome of a microbial community sampled directly from the environment. The assignment of these metagenomic reads into different species or taxonomical classes is a vital step for metagenomic analysis, which is referred to as binning of metagenomic data. RESULTS: In this paper, we propose a new method TM-MCluster for binning metagenomic reads. First, we represent each metagenomic read as a set of "k-mers" with their frequencies occurring in the read. Then, we employ a probabilistic topic model -- the Latent Dirichlet Allocation (LDA) model to the reads, which generates a number of hidden "topics" such that each read can be represented by a distribution vector of the generated topics. Finally, as in the MCluster method, we apply SKWIC -- a variant of the classical K-means algorithm with automatic feature weighting mechanism to cluster these reads represented by topic distributions. CONCLUSIONS: Experiments show that the new method TM-MCluster outperforms major existing methods, including AbundanceBin, MetaCluster 3.0/5.0 and MCluster. This result indicates that the exploitation of topic modeling can effectively improve the binning performance of metagenomic reads. Ruichang Zhang, Zhanzhan Cheng, Jihong Guan, Shuigeng Zhou |
BMC Bioinform. | 2 |