VLDB 2026 Research / reviewers in the wild / expert
Cheng-Lin Liu 0001
dblp:24/3006-1 · also Chenglin Liu 0001
· DBLP profile ↗
306ranked-venue papers
7as first author
117since 2021 · last 2026
0000-0002-6743-4175ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 235 · 7 first-author · 91 since 2021Graphics, computer vision, multimedia, augmented reality and games · 105 · 45 since 2021Databases, data management, data science and information retrieval · 74 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Too Long, Do Re-weighting for Efficient LLM Reasoning CompressionabstractZhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji, Peijie Wang, Haotian Xu, Xing W, Haizhen Huang, Weiwei Deng, Yeyun Gong, Zhijiang Guo, Xiao Liu, Fei Yin, Cheng-Lin Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhongzhi Li, Lei Ji 0001, Xing W, Haizhen Huang, Yeyun Gong, Zhijiang Guo, Xiao Liu 0029, Cheng-Lin Liu 0001 |
ACL (1) | 14 |
| 2026 | MVR: Diffusion-Based Multi-View Reasoning for Scene Text Detection
Debayan Das Gupta, Palaiahnakote Shivakumara, Palash Ghosal, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICDAR (2) | 5 |
| 2026 | From System 1 to System 2: A Survey of Reasoning Large Language ModelsabstractAchieving human-level intelligence requires refining the transition from the fast, intuitive System 1 to the slower, more deliberate System 2 reasoning. While System 1 excels in quick, heuristic decisions, System 2 relies on logical reasoning for more accurate judgments and reduced biases. Foundational Large Language Models (LLMs) excel at fast decision-making but lack the depth for complex reasoning, as they have not yet fully embraced the step-by-step analysis characteristic of true System 2 thinking. Recently, reasoning LLMs like OpenAI's o1/o3 and DeepSeek's R1 have demonstrated expert-level performance in fields such as mathematics and coding, closely mimicking the deliberate reasoning of System 2 and showcasing human-like cognitive abilities. This survey begins with a brief overview of the progress in foundational LLMs and the early development of System 2 technologies, exploring how their combination has paved the way for reasoning LLMs. Next, we discuss how to construct reasoning LLMs, trace the evolution of various reasoning models, and examine the core methods that enable advanced reasoning behind them. Additionally, we provide an overview of reasoning benchmarks, offering an in-depth comparison of the performance of representative reasoning LLMs. Finally, we explore promising directions for advancing reasoning LLMs and maintain a real-time GitHub Repository to track the latest developments. We hope this survey will serve as a valuable resource to inspire innovation and drive progress in this rapidly evolving field. Duzhen Zhang, Zhongzhi Li, Jiaxin Zhang 0024, Zengyan Liu, Junhao Zheng, Xiuyi Chen, Jiahua Dong 0001, Zhijiang Guo, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 15 |
| 2026 | Bayesian classifier calibration based on synthesized samples for zero-shot Chinese character recognition
Xiang Ao 0002, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2026 | FGPR: A large-scale dataset and benchmark for fine-grained product retrieval
Ruisong Zhang, Zhongzhi Li, Chuang Wang 0007, Cheng-Lin Liu 0001 |
Pattern Recognit. | 5 |
| 2026 | Det-Agent: Open-Vocabulary Object Localization and Detection With Reinforcement Learning AgentabstractObject detection, which aims to locate and recognize objects in images, is evolving toward reduced reliance on manual annotations and enhanced adaptability to open-world scenarios. This shift has led to open-vocabulary object detection (OVD), which enables zero-shot detection of objects from novel categories beyond the base categories. In this work, we identify three key challenges in detecting unseen class instances: 1) locating the instances of new classes; 2) distinguishing new class instances from the background; 3) recognizing new class instances. We propose a detection framework that leverages vision-language pre-trained (VLPT) models, such as CLIP, as the backbone to jointly address these three challenges. Specifically, we treat localization as a box-deformation decision process, where the agent interacts with the image to learn a universal deformation strategy, enhancing generalization for unseen class objects. We further reformulate the foreground-background classification as an objectness ranking task to improve objectness evaluation, utilizing a specially designed AP loss. Additionally, a feature magnitude minimization constraint is introduced for the adapter during fine-tuning, boosting recognition performance for both base and novel classes. Experiments on COCO and LVIS datasets demonstrate that our method outperforms previous approaches in open-vocabulary object detection. Ruisong Zhang, Xin-Jian Wu, Chuang Wang 0007, Cheng-Lin Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and LocatingabstractChao Deng, Jiale Yuan, Pi Bu, Peijie Wang, Zhong-Zhi Li, Jian Xu, Xiao-Hui Li, Yuan Gao, Jun Song, Bo Zheng, Cheng-Lin Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiale Yuan, Pi Bu, Zhongzhi Li, Jian Xu 0027, Bo Zheng 0007, Cheng-Lin Liu 0001 |
ACL (1) | 11 |
| 2025 | DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed LearningabstractDocument image segmentation is crucial for document analysis and recognition but remains challenging due to the diversity of document formats and segmentation tasks. Existing methods often address these tasks separately, resulting in limited generalization and resource wastage. This paper introduces DocSAM, a transformer-based unified framework designed for various document image segmentation tasks, such as document layout analysis, multi-granularity text segmentation, and table structure recognition, by modelling these tasks as a combination of instance and semantic segmentation. Specifically, DocSAM employs Sentence-BERT to map category names from each dataset into semantic queries that match the dimensionality of instance queries. These two sets of queries interact through an attention mechanism and are cross-attended with image features to predict instance and semantic segmentation masks. Instance categories are predicted by computing the dot product between instance and semantic queries, followed by softmax normalization of scores. Consequently, DocSAM can be jointly trained on heterogeneous datasets, enhancing robustness and generalization while reducing computational and storage resources. Comprehensive evaluations show that DocSAM surpasses existing methods in accuracy, efficiency, and adaptability, highlighting its potential for advancing document image understanding and segmentation across various applications. Codes are available at https://github.com/xhli-git/DocSAM. Xiao-Hui Li 0012, Cheng-Lin Liu 0001 |
CVPR | 3 |
| 2025 | MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual ContextsabstractMultimodal Large Language Models (MLLMs) have shown promising capabilities in mathematical reasoning within visual contexts across various datasets. However, most existing multimodal math benchmarks are limited to single-visual contexts, which diverges from the multi-visual scenarios commonly encountered in real-world mathematical applications. To address this gap, we introduce MV-MATH: a meticulously curated dataset of 2,009 high-quality mathematical problems. Each problem integrates multiple images interleaved with text, derived from authentic K-12 scenarios, and enriched with detailed annotations. MV-MATH includes multiple-choice, free-form, and multi-step questions, covering 11 subject areas across 3 difficulty levels, and serves as a comprehensive and rigorous benchmark for assessing MLLMs’ mathematical reasoning in multi-visual contexts. Through extensive experimentation, we observe that MLLMs encounter substantial challenges in multi-visual math tasks, with a considerable performance gap relative to human capabilities on MV-MATH. Furthermore, we analyze the performance and error patterns of various models, providing insights into MLLMs’ mathematical reasoning capabilities within multi-visual settings. The data and code: https://eternal8080.github.io/MV-MATH.github.io/. Zhongzhi Li, Dekang Ran, Cheng-Lin Liu 0001 |
CVPR | 5 |
| 2025 | HiE-VL: A Large Vision-Language Model with Hierarchical Adapter for Handwritten Mathematical Expression RecognitionabstractLarge Vision-Language Models (LVLMs) have shown impressive capabilities across various domains, but existing LVLMs have limited performance in dense perception and structured learning problems, such as Handwritten Mathematical Expression Recognition (HMER). The primary challenges stem from the complexity of formula images comprising multiple symbols and complicated inter-symbol relationships. This poses difficulties to LVLMs with locality insensitive visual encoders and structure-agnostic vision-language projectors. To overcome these challenges, we propose HiE-VL, the first LVLM for HMER containing: (1) a primitive-aware high-resolution visual encoder, (2) a hierarchical adapter, (3) a math-context enhanced large language model (LLM). Specifically, the adopted visual encoder allows locating and recognizing symbols in complex formula images. The hierarchical adapter functions as a vision-language projector to progressively capture primitive and structure information for facilitating expression decoding. The whole model is optimized in a two-stage training pipeline. In experiments on two benchmark datasets of HMER, our model achieves significantly higher performance than existing LVLMs like GPT-4V and state-of-the-art HMER models. Our codes are available at https://github.com/guohy17/HiE-VL. Hong-Yu Guo, Jian Xu 0015, Cheng-Lin Liu 0001 |
ICASSP | 4 |
| 2025 | Federated Continual Instruction Tuning
Haiyang Guo, Fanhu Zeng, Fei Zhu 0004, Wenzhuo Liu, Dahan Wang, Jian Xu 0015, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICCV | 8 |
| 2025 | A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition
Ritabrata Chakraborty, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICDAR (2) | 4 |
| 2025 | Balti-Tamko: A Spoken Words Dataset Development for Automatic Recognition of the Endangered Balti Language
Sardar Shan Ali Naqvi, Zeeshan Abbas, Cheng-Lin Liu 0001 |
ICONIP (5) | 5 |
| 2025 | SolidGeo: Measuring Multimodal Spatial Math Reasoning in Solid GeometryabstractGeometry is a fundamental branch of mathematics and plays a crucial role in evaluating the reasoning capabilities of multimodal large language models (MLLMs). However, existing multimodal mathematics benchmarks mainly focus on plane geometry and largely ignore solid geometry, which requires spatial reasoning and is more challenging than plane geometry. To address this critical gap, we introduce SolidGeo, the first large-scale benchmark specifically designed to evaluate the performance of MLLMs on mathematical reasoning tasks in solid geometry. SolidGeo consists of 3,113 real-world K–12 and competition-level problems, each paired with visual context and annotated with difficulty levels and fine-grained solid geometry categories. Our benchmark covers a wide range of 3D reasoning subjects such as projection, unfolding, spatial measurement, and spatial vector, offering a rigorous testbed for assessing solid geometry. Through extensive experiments, we observe that MLLMs encounter substantial challenges in solid geometry math tasks, with a considerable performance gap relative to human capabilities on SolidGeo. Moreover, we analyze the performance, inference effiency and error patterns of various models, offering insights into the solid geometric mathematical reasoning capabilities of MLLMs. We hope SolidGeo serves as a catalyst for advancing MLLMs toward deeper geometric reasoning and spatial intelligence. The dataset is released at https://huggingface.co/datasets/HarryYancy/SolidGeo/ Zhongzhi Li, Dekang Ran, Zhilong Ji, Jinfeng Bai, Cheng-Lin Liu 0001 |
NeurIPS | 9 |
| 2025 | Breaking the Limits of Reliable Prediction via Generated Data
Zhen Cheng 0003, Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | QDNet: Query-Denoising Network for Visual Traffic Knowledge Graph GenerationabstractTraffic scene perception underpins essential tasks like map construction and route planning in modern intelligent transportation systems, thus receiving extensive attention. However, existing methods tend to concentrate solely on specific elements, lacking a comprehensive understanding of various traffic scenes. This paper addresses the Visual Traffic Knowledge Graph Generation (VTKGG) task, aiming to extract and represent traffic information from various elements in the traffic scene image as a knowledge graph. To achieve this, we propose Query-Denoising Network (QDNet) to integrate multiple subtasks through different types of queries in an end-to-end manner. These queries facilitate information communication between different modules, streamlining the generation of visual traffic knowledge graphs by eliminating cumbersome intermediate steps. Considering the challenges in optimizing such a cascaded multi-task model, we incorporate the query-denoising method into the training process of QDNet. By introducing the noised query, enhancing the internal noise of the model, and forcing the model to recover the ground truth, our approach achieves accurate results. This strategy improves the robustness and performance of our model. We conduct extensive ablation and comparative experiments to demonstrate the superiority and effectiveness of our framework and strategy, and experiments on a similar task Panoptic Scene Graph Generation also demonstrate its superiority. Xiao-Hui Li 0012, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | ProtoGCD: Unified and Unbiased Prototype Learning for Generalized Category DiscoveryabstractGeneralized category discovery (GCD) is a pragmatic but underexplored problem, which requires models to automatically cluster and discover novel categories by leveraging the labeled samples from old classes. The challenge is that unlabeled data contain both old and new classes. Early works leveraging pseudo-labeling with parametric classifiers handle old and new classes separately, which brings about imbalanced accuracy between them. Recent methods employing contrastive learning neglect potential positives and are decoupled from the clustering objective, leading to biased representations and sub-optimal results. To address these issues, we introduce a unified and unbiased prototype learning framework, namely ProtoGCD, wherein old and new classes are modeled with joint prototypes and unified learning objectives, enabling unified modeling between old and new classes. Specifically, we propose a dual-level adaptive pseudo-labeling mechanism to mitigate confirmation bias, together with two regularization terms to collectively help learn more suitable representations for GCD. Moreover, for practical considerations, we devise a criterion to estimate the number of new classes. Furthermore, we extend ProtoGCD to detect unseen outliers, achieving task-level unification. Comprehensive experiments show that ProtoGCD achieves state-of-the-art performance on both generic and fine-grained datasets. Shijie Ma, Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | PASS++: A Dual Bias Reduction Framework for Non-Exemplar Class-Incremental LearningabstractClass-incremental learning (CIL) aims to continually recognize new classes while preserving the discriminability of previously learned ones. Most existing CIL methods are exemplar-based, relying on the storage and replay of a subset of old data during training. Without access to such data, these methods typically suffer from catastrophic forgetting. In this paper, we identify two fundamental causes of forgetting in CIL: representation bias and classifier bias. To address these challenges, we propose a simple yet effective dual-bias reduction framework, which leverages self-supervised transformation (SST) in the input space and prototype augmentation (protoAug) in the feature space. On one hand, SST mitigates representation bias by encouraging the model to learn generic, diverse representations that generalize across tasks. On the other hand, protoAug tackles classifier bias by explicitly or implicitly augmenting the prototypes of old classes in the feature space, thereby imposing stronger constraints to preserve decision boundaries. We further enhance the framework with hardness-aware prototype augmentation and multi-view ensemble strategies, yielding significant performance gains. The proposed framework can be easily integrated with pre-trained models. Without storing any samples of old classes, our method performs comparably to state-of-the-art exemplar-based approaches that rely on extensive data storage. We hope to draw the attention of researchers back to non-exemplar CIL by rethinking the necessity of storing old samples. Fei Zhu 0004, Xu-Yao Zhang, Zhen Cheng 0003, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Vision-language pre-training for graph-based handwritten mathematical expression recognition
Hong-Yu Guo, Chuang Wang 0007, Xiao-Hui Li 0012, Cheng-Lin Liu 0001 |
Pattern Recognit. | 5 |
| 2025 | Class incremental learning with self-supervised pre-training and prototype learning
Wenzhuo Liu, Xin-Jian Wu, Fei Zhu 0004, Ming-Ming Yu, Chuang Wang 0007, Cheng-Lin Liu 0001 |
Pattern Recognit. | 6 |
| 2025 | A novel domain independent scene text localizerabstract• The proposed domain independent model for scene text localization is new. • Exploring partial convolution with Yolov5-transformer for feature extraction. • Integrating the swin transformer with the novel channel attention modules. • The result of the proposed method is superior to the existing methods. Text localization across multiple domains is crucial for applications like autonomous driving and tracking marathon runners. This work introduces DIPCYT, a novel model that utilizes Domain Independent Partial Convolution and a Yolov5-based Transformer for text localization in scene images from various domains, including natural scenes, underwater, and drone images. Each domain presents unique challenges: underwater images suffer from poor quality and degradation, drone images suffer from tiny text and loss of shapes, and scene images suffer from arbitrarily oriented, shaped text. Additionally, license plates in drone images may not provide rich semantic information compared to other text types due to loss of contextual information between characters. To tackle these challenges, DIPCYT employs new partial convolution layers within Yolov5 and integrates Transformer detection heads with a novel Fourier Positional Convolutional Block Attention Module (FPCBAM). This approach leverages common text properties across domains, such as contextual (global) and spatial (local) relationships. Experimental results demonstrate that DIPCYT outperforms existing methods, achieving F-scores of 0.90, 0.90, 0.77, 0.85, 0.85, and 0.88 on Total-Text, ICDAR 2015, ICDAR 2019 MLT, CTW1500, Drone, and Underwater datasets, respectively. Ayush Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2025 | Towards reliable domain generalization: Insights from the PF2HC benchmark and dynamic evaluations
Xiang Ao 0002, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2025 | Split-net: Dual transformer encoder with splitting scene text image for script identification
Ayush Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. Lett. | 4 |
| 2025 | Cross-Modal Causal Representation Learning for Radiology Report GenerationabstractRadiology Report Generation (RRG) is essential for computer-aided diagnosis and medication guidance, which can relieve the heavy burden of radiologists by automatically generating the corresponding radiology reports according to the given radiology image. However, generating accurate lesion descriptions remains challenging due to spurious correlations from visual-linguistic biases and inherent limitations of radiological imaging, such as low resolution and noise interference. To address these issues, we propose a two-stage framework named Cross-Modal Causal Representation Learning (CMCRL), consisting of the Radiological Cross-modal Alignment and Reconstruction Enhanced (RadCARE) pre-training and the Visual-Linguistic Causal Intervention (VLCI) fine-tuning. In the pre-training stage, RadCARE introduces a degradation-aware masked image restoration strategy tailored for radiological images, which reconstructs high-resolution patches from low-resolution inputs to mitigate noise and detail loss. Combined with a multiway architecture and four adaptive training strategies (e.g., text postfix generation with degraded images and text prefixes), RadCARE establishes robust cross-modal correlations even with incomplete data. In the VLCI phase, we deploy causal front-door intervention through two modules: the Visual Deconfounding Module (VDM) disentangles local-global features without fine-grained annotations, while the Linguistic Deconfounding Module (LDM) eliminates context bias without external terminology databases. Experiments on IU-Xray and MIMIC-CXR show that our CMCRL pipeline significantly outperforms state-of-the-art methods, with ablation studies confirming the necessity of both stages. Code and models are available at https://github.com/WissingChen/CMCRL. Yang Liu 0084, Ce Wang 0001, Jiarui Zhu, Guanbin Li, Cheng-Lin Liu 0001, Liang Lin 0004 |
IEEE Trans. Image Process. | 6 |
| 2025 | Average of Pruning: Improving Performance and Stability of Out-of-Distribution DetectionabstractDetecting out-of-distribution (OOD) inputs has been a critical issue for neural networks in the open world. However, the unstable behavior of OOD detection along the optimization trajectory during training has not been explored clearly. In this article, we first find the performance of OOD detection suffers from overfitting and instability during training: 1) the performance could decrease when the training error is near zero and 2) the performance would vary sharply in the final stage of training. Based on our findings, we propose an average of pruning (AoP), consisting of model averaging (MA) and pruning, to mitigate the unstable behaviors. Specifically, MA can help achieve a stable performance by smoothing the landscape, and pruning is theoretically and empirically verified to eliminate overfitting by avoiding redundant features. Comprehensive experiments on various datasets and architectures are conducted to verify the effectiveness of our method. Zhen Cheng 0003, Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Branch-Tuning: Balancing Stability and Plasticity for Continual Self-Supervised LearningabstractThe self-supervised learning (SSL) has emerged as an effective paradigm for deriving general representations from vast amounts of unlabeled data. However, as real-world applications continually integrate new content, the high computational and resource demands of SSL necessitate continual learning (CL) rather than complete retraining. This poses a challenge in balancing between stability and plasticity when adapting to new information. In this article, we employ centered kernel alignment (CKA) for quantitatively analyzing model stability and plasticity, revealing the critical roles of batch normalization (BN) layers for stability and convolutional layers for plasticity. Motivated by this, we propose branch-tuning (BT), an efficient and straightforward method that achieves a balance between stability and plasticity in continual SSL. BT consists of branch expansion and compression and can be easily applied to various SSL methods without the need of modifying the original methods, retaining old data or models. We validate our method through experiments on various benchmark datasets, demonstrating its effectiveness and practical value in real-world scenarios. We hope our work offers new insights for future continual SSL research. The code will be made publicly available. Wenzhuo Liu, Fei Zhu 0004, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Unified Entropy Optimization for Open-Set Test-Time AdaptationabstractTest-time adaptation (TTA) aims at adapting a model pretrained on the labeled source domain to the unlabeled target domain. Existing methods usually focus on improving TTA performance under covariate shifts, while neglecting semantic shifts. In this paper, we delve into a realistic open-set TTA setting where the target domain may contain samples from unknown classes. Many state-of-the-art closed-set TTA methods perform poorly when applied to open-set scenarios, which can be attributed to the inaccurate estimation of data distribution and model confidence. To address these issues, we propose a simple but effective framework called unified entropy optimization (UniEnt), which is capable of simultaneously adapting to covariate-shifted in-distribution (csID) data and detecting covariate-shifted out-of-distribution (csOOD) data. Specifically, UniEnt first mines pseudo-csID and pseudo-csOOD samples from test data, followed by entropy min-imization on the pseudo-csID data and entropy maximization on the pseudo-csOOD data. Furthermore, we introduce UniEnt+ to alleviate the noise caused by hard data partition leveraging sample-level confidence. Extensive experiments on CIFAR benchmarks and Tiny-ImageNet-C show the superiority of our framework. The code is available at https://github.com/gaozhengqing/UniEnt. Zhengqing Gao, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 3 |
| 2024 | Active Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) is a pragmatic and challenging open-world task, which endeavors to cluster unlabeled samples from both novel and old classes, leveraging some labeled data of old classes. Given that knowledge learned from old classes is not fully transferable to new classes, and that novel categories are fully unlabeled, GCD inherently faces intractable problems, including imbalanced classification performance and inconsistent confidence between old and new classes, especially in the low-labeling regime. Hence, some annotations of new classes are deemed necessary. However, labeling new classes is extremely costly. To address this issue, we take the spirit of active learning and propose a new setting called Active Generalized Category Discovery (AGCD). The goal is to improve the performance of GCD by actively selecting a limited amount of valuable samples for labeling from the oracle. To solve this problem, we devise an adaptive sampling strategy, which jointly considers novelty, informativeness and diversity to adaptively select novel samples with proper uncertainty. However, owing to the varied orderings of label indices caused by the clustering of novel classes, the queried labels are not directly applicable to subsequent training. To overcome this issue, we further propose a stable label mapping algorithm that transforms ground truth labels to the label space of the classifier, thereby ensuring consistent training across different active selection stages. Our method achieves state-of-the-art performance on both generic and fine-grained datasets. Our code is available at https://github.com/mashijie1028/ActiveGCD Shijie Ma, Fei Zhu 0004, Zhun Zhong, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 5 |
| 2024 | RCL: Reliable Continual Learning for Unified Failure DetectionabstractDeep neural networks are known to be overconfident for what they don't know in the wild, which is undesirable for decision-making in high-stakes applications. Despite quan-tities of existing works, most of them focus on detecting out-of-distribution (OOD) samples from unseen classes, while ignoring large parts of relevant failure sources like mis-classified samples from known classes. In particular, recent studies reveal that prevalent OOD detection methods are actually harmful for misclassification detection (MisD), indicating that there seems to be a tradeoff between those two tasks. In this paper, we study the critical yet under-explored problem of unified failure detection, which aims to detect both misclassified and OOD examples. Concretely, we identify the failure of simply integrating learning objectives of misclassification and OOD detection, and show the potential of sequence learning. Inspired by this, we propose a reliable continual learning paradigm, whose spirit is to equip the model with MisD ability first, and then improve the OOD detection ability without degrading the al-ready adequate MisD performance. Extensive experiments demonstrate that our method achieves strong unified failure detection performance. The code is available at https://github.com/Impression2805/RCL. Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang, Cheng-Lin Liu 0001, Zhaoxiang Zhang 0001 |
CVPR | 4 |
| 2024 | WPS-SAM: Towards Weakly-Supervised Part Segmentation with Foundation Models
Xin-Jian Wu, Ruisong Zhang, Shijie Ma, Cheng-Lin Liu 0001 |
ECCV (44) | 5 |
| 2024 | Prototype Calibration with Synthesized Samples for Zero-Shot Chinese Character RecognitionabstractZero-shot Chinese character recognition aims to recognize unseen characters that have never appeared in training. Recently, many methods learn a cross-modal alignment between character samples and auxiliary semantic data like glyph templates in training, and directly employ it to recognize unseen characters by retrieving the class with most similar semantics. However, these approaches suffer from the domain shift problem, which means that the learned alignment shows a deviation on unseen characters. To alleviate this problem, we generate unseen character samples to calibrate the shifted prototypes in the feature space. Specifically, we train a cross-modal prototype classifier and a generator conditioned on glyph templates, then use the generator to synthesize unseen character samples to calibrate the prototypes of the classifier. The calibration process does not require any extra training. Experiments on a handwritten dataset and a nature scene dataset show the superiority of our method and the effectiveness of prototype calibration. Xiang Ao 0002, Xiao-Hui Li 0012, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICASSP | 4 |
| 2024 | GraphMLLM: A Graph-Based Multi-level Layout Language-Independent Model for Document Understanding
He-Sen Dai, Xiao-Hui Li 0012, Shuqi Mei, Cheng-Lin Liu 0001 |
ICDAR (1) | 6 |
| 2024 | A New Unsupervised Approach for Text Localization in Shaky and Non-shaky Scene Video
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Cheng-Lin Liu 0001 |
ICDAR (5) | 5 |
| 2024 | Deep Metric Learning with Cross-Writer Attention for Offline Signature Verification
Lu-Rong Ling, Heng Zhang 0028, Cheng-Lin Liu 0001 |
ICDAR (2) | 4 |
| 2024 | Context-Aware Confidence Estimation for Rejection in Handwritten Chinese Text Recognition
Yi Chen 0027, Cheng-Lin Liu 0001 |
ICDAR (1) | 4 |
| 2024 | MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any ResolutionabstractAlthough Vision Transformers (ViTs) have recently advanced computer vision tasks significantly, an important real-world problem was overlooked: adapting to variable input resolutions. Typically, images are resized to a fixed resolution, such as 224x224, for efficiency during training and inference. However, uniform input size conflicts with real-world scenarios where images naturally vary in resolution. Modifying the preset resolution of a model may severely degrade the performance. In this work, we propose to enhance the model adaptability to resolution variation by optimizing the patch embedding. The proposed method, called Multi-Scale Patch Embedding (MSPE), substitutes the standard patch embedding with multiple variable-sized patch kernels and selects the best parameters for different resolutions, eliminating the need to resize the original image. Our method does not require high-cost training or modifications to other parts, making it easy to apply to most ViT models. Experiments in image classification, segmentation, and detection tasks demonstrate the effectiveness of MSPE, yielding superior performance on low-resolution inputs and performing comparably on high-resolution inputs with existing methods. Wenzhuo Liu, Fei Zhu 0004, Shijie Ma, Cheng-Lin Liu 0001 |
NeurIPS | 4 |
| 2024 | Happy: A Debiased Learning Framework for Continual Generalized Category DiscoveryabstractConstantly discovering novel concepts is crucial in evolving environments. This paper explores the underexplored task of Continual Generalized Category Discovery (C-GCD), which aims to incrementally discover new classes from *unlabeled* data while maintaining the ability to recognize previously learned classes. Although several settings are proposed to study the C-GCD task, they have limitations that do not reflect real-world scenarios. We thus study a more practical C-GCD setting, which includes more new classes to be discovered over a longer period, without storing samples of past classes. In C-GCD, the model is initially trained on labeled data of known classes, followed by multiple incremental stages where the model is fed with unlabeled data containing both old and new classes. The core challenge involves two conflicting objectives: discover new classes and prevent forgetting old ones. We delve into the conflicts and identify that models are susceptible to *prediction bias* and *hardness bias*. To address these issues, we introduce a debiased learning framework, namely **Happy**, characterized by **H**ardness-**a**ware **p**rototype sampling and soft entro**py** regularization. For the *prediction bias*, we first introduce clustering-guided initialization to provide robust features. In addition, we propose soft entropy regularization to assign appropriate probabilities to new classes, which can significantly enhance the clustering performance of new classes. For the *harness bias*, we present the hardness-aware prototype sampling, which can effectively reduce the forgetting issue for previously seen classes, especially for difficult classes. Experimental results demonstrate our method proficiently manages the conflicts of C-GCD and achieves remarkable performance across various datasets, e.g., 7.5% overall gains on ImageNet-100. Our code is publicly available at https://github.com/mashijie1028/Happy-CGCD. Shijie Ma, Fei Zhu 0004, Zhun Zhong, Wenzhuo Liu, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
NeurIPS | 6 |
| 2024 | Adapting Vision-Language Models to Open Classes via Test-Time Prompt Tuning
Zhengqing Gao, Xiang Ao 0002, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
PRCV (5) | 4 |
| 2024 | Sequential Transformer for End-to-End Video Text DetectionabstractIn existing methods of video text detection, the detection and tracking branches are usually independent of each other, and although they jointly optimize the backbone network, the tracking-by-detection paradigm still needs to be used during the inference stage. To address this issue, we propose a novel video text detection framework based on sequential transformer, which decodes detection and tracking tasks in parallel, without explicitly setting up a tracking branch. To achieve this, we first introduce the concept of instance query, which learns long-term context information in the video sequence. Then, based on the instance query, the transformer decoder is used to predict the entire box and mask sequence of the text instance in one pass. As a result, the tracking task is realized naturally. In addition, the proposed method can be applied to the scene text detection task seamlessly, without modifying any modules. To the best of our knowledge, this is the first framework to unify the tasks of scene text detection and video text detection. Our model achieves state-of-the-art performance on four video text datasets (YVT, RT-1K, BOVText, and BiRViT-1K), and competitive results on three scene text datasets (CTW1500, MSRA-TD500, and Total-Text). The code is available at https://github.com/zjb-1/SeqVideoText. Mengbiao Zhao, Cheng-Lin Liu 0001 |
WACV | 4 |
| 2024 | SignParser: An End-to-End Framework for Traffic Sign Understanding
Wei Feng 0016, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Ensemble Quadratic Assignment Network for Graph Matching
Haoru Tan, Chuang Wang 0007, Sitong Wu, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 6 |
| 2024 | Revisiting Confidence Estimation: Towards Reliable Failure PredictionabstractReliable confidence estimation is a challenging yet fundamental requirement in many risk-sensitive applications. However, modern deep neural networks are often overconfident for their incorrect predictions, i.e., misclassified samples from known classes, and out-of-distribution (OOD) samples from unknown classes. In recent years, many confidence calibration and OOD detection methods have been developed. In this paper, we find a general, widely existing but actually-neglected phenomenon that most confidence estimation methods are harmful for detecting misclassification errors. We investigate this problem and reveal that popular calibration and OOD detection methods often lead to worse confidence separation between correctly classified and misclassified examples, making it difficult to decide whether to trust a prediction or not. Finally, we propose to enlarge the confidence gap by finding flat minima, which yields state-of-the-art failure prediction performance under various settings including balanced, long-tailed, and covariate-shift classification scenarios. Our study not only provides a strong baseline for reliable confidence estimation but also acts as a bridge between understanding calibration, OOD detection, and failure prediction. Fei Zhu 0004, Xu-Yao Zhang, Zhen Cheng 0003, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | An end-to-end model for multi-view scene text recognition
Ayan Banerjee 0002, Palaiahnakote Shivakumara, Saumik Bhattacharya, Umapada Pal 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 5 |
| 2024 | Transformer-based stroke relation encoding for online handwriting and sketches
Jing-Yu Liu, Yan-Ming Zhang 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2024 | Large-scale continual learning for ancient Chinese character recognition
Xu-Yao Zhang, Zhaoxiang Zhang 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2024 | An approach for handwritten Chinese text recognition unifying character segmentation and recognition
Mingming Yu, Heng Zhang 0028, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2024 | Polynomial kernel learning for interpolation kernel machines with application to graph classification
Cheng-Lin Liu 0001, Xiaoyi Jiang 0001 |
Pattern Recognit. Lett. | 2 |
| 2024 | Video Text Detection With Robust Feature RepresentationabstractExisting video text detection methods mostly track texts with appearance feature only, thus are easily influenced by the change of perspective and illumination. In this paper, we propose an end-to-end video text detector that tracks texts based on robust feature representation fusing multiple descriptors. First, we introduce a character center segmentation branch to extract semantic feature, which encodes the category and position information of characters. And for extracting the topology feature of each text instance, we propose a relative position awareness branch to encode the relative position information among texts. Then, an adaptive feature fusion network is proposed to dynamically fuse multiple descriptors to generate a robust feature representation for more robust tracking. In addition, to promote the research and evaluation in this field, we also construct a large Bilingual Road scene Video Text dataset, named BiRViT-1K, which contains 1000 videos of Chinese and English texts. Experimental results show the proposed semantic and topology features are beneficial to the text detection and tracking performance, and the proposed method achieves state-of-the-art performance on four public video text benchmarks ICDAR 2015 Video, YVT, RT-1K and BOVText, and two Chinese scene text benchmarks CASIA10K and MSRA-TD500. Wei Feng 0016, Mengbiao Zhao, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Social Relation Reasoning Based on Triangular ConstraintsabstractSocial networks are essentially in a graph structure where persons act as nodes and the edges connecting nodes denote social relations. The prediction of social relations, therefore, relies on the context in graphs to model the higher-order constraints among relations, which has not been exploited sufficiently by previous works, however. In this paper, we formulate the paradigm of the higher-order constraints in social relations into triangular relational closed-loop structures, i.e., triangular constraints, and further introduce the triangular reasoning graph attention network (TRGAT). Our TRGAT employs the attention mechanism to aggregate features with triangular constraints in the graph, thereby exploiting the higher-order context to reason social relations iteratively. Besides, to acquire better feature representations of persons, we introduce node contrastive learning into relation reasoning. Experimental results show that our method outperforms existing approaches significantly, with higher accuracy and better consistency in generating social relation graphs. Wei Feng 0016, Shuqi Mei, Cheng-Lin Liu 0001 |
AAAI | 7 |
| 2023 | Interpolation Kernel Machines: Reducing Multiclass to Binary
Cheng-Lin Liu 0001, Xiaoyi Jiang 0001 |
CAIP (1) | 2 |
| 2023 | OpenMix: Exploring Outlier Samples for Misclassification DetectionabstractReliable confidence estimation for deep neural classifiers is a challenging yet fundamental requirement in high-stakes applications. Unfortunately, modern deep neural networks are often overconfident for their erroneous predictions. In this work, we exploit the easily available outlier samples, i.e., unlabeled samples coming from non-target classes, for helping detect misclassification errors. Particularly, we find that the well-known Outlier Exposure, which is powerful in detecting out-of-distribution (OOD) samples from unknown classes, does not provide any gain in identifying misclassification errors. Based on these observations, we propose a novel method called OpenMix, which incorporates open-world knowledge by learning to reject uncertain pseudo-samples generated via outlier transformation. OpenMix significantly improves confidence reliability under various scenarios, establishing a strong and unified framework for detecting both misclassified samples from known classes and OOD samples from unknown classes. The code is publicly available at https://github.com/Impression2805/OpenMix. Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 4 |
| 2023 | Streaming Stroke Classification of Online HandwritingabstractStroke classification for online handwriting aims at providing each stroke with a semantic label so as to fulfill handwriting segmentation. This task has attracted considerable attention due to its significance in online handwriting analysis. Existing methods are designed for the static situation, where stroke classification is conducted on the completion of handwriting. With the popularity of pad devices and electronic whiteboards, streaming stroke classification becomes increasingly important for instant handwriting processing and feedback. However, streaming classification is much more challenging due to the lack of contextual information and is underexplored in the past. In this paper, we propose Multiple Stroke State Transformer (MSST), a novel framework to enable simultaneous real-time classification and modifiability of previous predictions. Particularly, we set multiple states with duration for each stroke and then divide all states into chunks to perform message passing by Transformer. Experiments on handwritten documents and diagrams demonstrate the superiority of our method. Jing-Yu Liu, Yan-Ming Zhang 0001, Cheng-Lin Liu 0001 |
ICASSP | 4 |
| 2023 | Visual Traffic Knowledge Graph Generation from Scene ImagesabstractAlthough previous works on traffic scene understanding have achieved great success, most of them stop at a low-level perception stage, such as road segmentation and lane detection, and few concern high-level understanding. In this paper, we present Visual Traffic Knowledge Graph Generation (VTKGG), a new task for in-depth traffic scene understanding that tries to extract multiple kinds of information and integrate them into a knowledge graph. To achieve this goal, we first introduce a large dataset named CASIA-Tencent Road Scene dataset (RS10K) with comprehensive annotations to support related research. Secondly, we propose a novel traffic scene parsing architecture containing a Hierarchical Graph ATtention network (HGAT) to analyze the heterogeneous elements and their complicated relations in traffic scene images. By hierarchizing the heterogeneous graph and equipping it with cross-level links, our approach exploits the correlation among various elements completely and acquires accurate relations. The experimental results show that our method can effectively generate visual traffic knowledge graphs and achieve state-of-the-art performance. The dataset RS10K is available at http://www.nlpr.ia.ac.cn/pal/RS10K.html. Xiao-Hui Li 0012, Shuqi Mei, Cheng-Lin Liu 0001 |
ICCV | 7 |
| 2023 | ICDAR 2023 Competition on Recognition of Multi-line Handwritten Mathematical Expressions
Chenyang Gao, Shiyu Yao, Jinfeng Bai, Xiang Bai, Cheng-Lin Liu 0001 |
ICDAR (2) | 7 |
| 2023 | ViSA: Visual and Semantic Alignment for Robust Scene Text Recognition
Zhenru Pan, Zhilong Ji, Xiao Liu 0040, Jinfeng Bai, Cheng-Lin Liu 0001 |
ICDAR (2) | 5 |
| 2023 | ICDAR 2023 Competition on Born Digital Video Text Question Answering
Zhibo Yang 0003, Xiaoge Song, Sibo Song, Tong Lu 0002, Xiang Bai, Cheng-Lin Liu 0001, Fei Huang 0002, Cong Yao |
ICDAR (2) | 6 |
| 2023 | ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao 0001, Wei Hua 0005, Bohan Li 0010, Mingrui Chen 0001, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lv, Yuechen Yu, Wanxiang Che, Errui Ding, Cheng-Lin Liu 0001, Jiebo Luo 0001, Shuicheng Yan, Min Zhang 0005, Dimosthenis Karatzas, Xing Sun 0001, Jingdong Wang 0001, Xiang Bai |
ICDAR (2) | 20 |
| 2023 | Diff-Writer: A Diffusion Model-Based Stylized Online Handwritten Chinese Character Generator
Minsi Ren, Yan-Ming Zhang 0001, Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
ICONIP (10) | 5 |
| 2023 | A Multi-Modal Neural Geometric Solver with Textual Clauses Parsed from DiagramabstractGeometry problem solving (GPS) is a high-level mathematical reasoning requiring the capacities of multi-modal fusion and geometric knowledge application. Recently, neural solvers have shown great potential in GPS but still be short in diagram presentation and modal fusion. In this work, we convert diagrams into basic textual clauses to describe diagram features effectively, and propose a new neural solver called PGPSNet to fuse multi-modal information efficiently. Combining structural and semantic pre-training, data augmentation and self-limited decoding, PGPSNet is endowed with rich knowledge of geometry theorems and geometric representation, and therefore promotes geometric understanding and reasoning. In addition, to facilitate the research of GPS, we build a new large-scale and fine-annotated GPS dataset named PGPS9K, labeled with both fine-grained diagram annotation and interpretable solution program. Experiments on PGPS9K and an existing dataset Geometry3K validate the superiority of our method over the state-of-the-art neural solvers. Our code, dataset and appendix material are available at \url{https://github.com/mingliangzhang2018/PGPS}. Mingliang Zhang 0005, Cheng-Lin Liu 0001 |
IJCAI | 3 |
| 2023 | Training with scaled logits to alleviate class-level over-fitting in few-shot learning
Rui-Qi Wang, Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Neurocomputing | 4 |
| 2023 | A New Lightweight Script Independent Scene Text Style Transfer NetworkabstractScene text style transfer without a language barrier is an open challenge for the video and scene text recognition community because this plays a vital role in poster, web design, augmenting character images, and editing characters to improve scene text recognition performance and usability. This work presents a new model, called Script Independent Scene Text Style Transfer Network (SISTSTNet), for extracting scene characters and transferring text style simultaneously. The SISTSTNet performs mapping in language-independent feature space for transferring style. It is designed based on a Style Parameter Network and Target Encoder Network through lightweight MobileNetv3 convolutional and residual blocks to capture the style and shape to generate target characters. Similarly, a generative model is explored through the Visual Geometry Group (VGG) network for character replacement. The SISTSTNet is flexible and works on different languages and arbitrary examples in a neat and unified fashion. The experimental results on images in various languages, namely, English, Chinese, Hindi, Russian, Japanese, Arabic, Greek, and Bengali and cross-language validation demonstrate the effectiveness of the proposed method. The performance of the method is superior compared to the state-of-the-art methods in terms of quality measures, language independence, shape-preserving, and efficiency. The code and dataset will be released to the public to support reproducibility. Palaiahnakote Shivakumara, Ayush Roy, Lokesh Nandanwar, Umapada Pal 0001, Yue Lu 0001, Cheng-Lin Liu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2023 | Imitating the oracle: Towards calibrated model for class incremental learning
Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Neural Networks | 4 |
| 2023 | Learning by Seeing More ClassesabstractTraditional pattern recognition models usually assume a fixed and identical number of classes during both training and inference stages. In this paper, we study an interesting but ignored question: can increasing the number of classes during training improve the generalization and reliability performance? For a k-class problem, instead of training with only these k classes, we propose to learn with k+m classes, where the additional m classes can be either real classes from other datasets or synthesized from known classes. Specifically, we propose two strategies for constructing new classes from known classes. By making the model see more classes during training, we can obtain several advantages. First, the added m classes serve as a regularization which is helpful to improve the generalization accuracy on the original k classes. Second, this will alleviate the overconfident phenomenon and produce more reliable confidence estimation for different tasks like misclassification detection, confidence calibration, and out-of-distribution detection. Lastly, the additional classes can also improve the learned feature representation, which is beneficial for new classes generalization in few-shot learning and class-incremental learning. Compared with the widely proved concept of data augmentation (dataAug), our method is driven from another dimension of augmentation based on additional classes (classAug). Comprehensive experiments demonstrated the superiority of our classAug under various open-environment metrics on benchmark datasets. Fei Zhu 0004, Xu-Yao Zhang, Rui-Qi Wang, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | A Survey on Learning to RejectabstractLearning to reject is a special kind of self-awareness (the ability to know what you do not know), which is an essential factor for humans to become smarter. Although machine intelligence has become very accurate nowadays, it lacks such kind of self-awareness and usually acts as omniscient, resulting in overconfident errors. This article presents a comprehensive overview of this topic from three perspectives: confidence, calibration, and discrimination. Confidence is an important measurement for the reliability of model predictions. Rejection can be realized by setting thresholds on confidence. However, most models, especially modern deep neural networks, are usually overconfident. Therefore, calibration is a process to ensure confidence matching the actual likelihood of correctness, including two approaches: post-calibration and self-calibration. Calibration reflects the global characteristic of confidence, and the local distinguishing property of confidence is also important. In light of this, discrimination focuses on the performance of accepting positive samples while rejecting negative samples. As a binary classification problem, the challenge of discrimination comes from the missing and nonrepresentativeness of the negative data. Three discrimination tasks are comprehensively analyzed and discussed: failure rejection, unknown rejection, and fake rejection. By rejecting failures, the risk could be controlled especially for mission-critical applications. By rejecting unknowns, the awareness of the knowledge blind zone would be enhanced. By rejecting fakes, security and privacy could be protected. We provide a general taxonomy, organization, and discussion of the methods for solving these problems, which are studied separately in the literature. The connections between different approaches and future directions that are worth further investigation are also presented. With a discriminative and calibrated confidence, learning to reject will let the decision-making process be more practical, reliable, and secure. Xu-Yao Zhang, Guosen Xie, Xiuli Li, Tao Mei 0001, Cheng-Lin Liu 0001 |
Proc. IEEE | 5 |
| 2023 | Adversarial training with distribution normalization and margin balance
Zhen Cheng 0003, Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2023 | Dynamics-aware loss for learning with label noise
Xiu-Chuan Li, Xiaobo Xia, Fei Zhu 0004, Tongliang Liu, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 6 |
| 2023 | Towards open-set text recognition via label-to-prototype learning
Chang Liu 0083, Haibo Qin, Xiaobin Zhu 0001, Cheng-Lin Liu 0001, Xu-Cheng Yin |
Pattern Recognit. | 5 |
| 2023 | DyGAT: Dynamic stroke classification of online handwritten documents and sketches
Yu-Ting Yang, Yan-Ming Zhang 0001, Xiao-Long Yun, Cheng-Lin Liu 0001 |
Pattern Recognit. | 5 |
| 2023 | Towards prior gap and representation gap for long-tailed recognition
Mingliang Zhang 0005, Xu-Yao Zhang, Chuang Wang 0007, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2023 | Deep representation learning for domain generalization with information bottleneck principle
Xu-Yao Zhang, Chuang Wang 0007, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2023 | VQAPT: A New visual question answering model for personality traits in social media images
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001, Yue Lu 0001 |
Pattern Recognit. Lett. | 4 |
| 2023 | Texts as points: Scene text detection with point supervision
Mengbiao Zhao, Wei Feng 0016, Cheng-Lin Liu 0001 |
Pattern Recognit. Lett. | 4 |
| 2023 | A New Language-Independent Deep CNN for Scene Text Detection and Style Transfer in Social Media ImagesabstractDue to the adverse effect of quality caused by different social media and arbitrary languages in natural scenes, detecting text from social media images and transferring its style is challenging. This paper presents a novel end-to-end model for text detection and text style transfer in social media images. The key notion of the proposed work is to find dominant information, such as fine details in the degraded images (social media images), and then restore the structure of character information. Therefore, we first introduce a novel idea of extracting gradients from the frequency domain of the input image to reduce the adverse effect of different social media, which outputs text candidate points. The text candidates are further connected into components and used for text detection via a UNet++ like network with an EfficientNet backbone (EffiUNet++). Then, to deal with the style transfer issue, we devise a generative model, which comprises a target encoder and style parameter networks (TESP-Net) to generate the target characters by leveraging the recognition results from the first stage. Specifically, a series of residual mapping and a position attention module are devised to improve the shape and structure of generated characters. The whole model is trained end-to-end so as to optimize the performance. Experiments on our social media dataset, benchmark datasets of natural scene text detection and text style transfer show that the proposed model outperforms the existing text detection and style transfer methods in multilingual and cross-language scenario. Palaiahnakote Shivakumara, Ayan Banerjee 0002, Umapada Pal 0001, Lokesh Nandanwar, Tong Lu 0002, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 6 |
| 2023 | Cycle-Consistent Weakly Supervised Visual Grounding With Individual and Contextual RepresentationsabstractVisual grounding, aiming to align image regions with textual queries, is a fundamental task for cross-modal learning. We study the weakly supervised visual grounding, where only image-text pairs at a coarse-grained level are available. Due to the lack of fine-grained correspondence information, existing approaches often encounter matching ambiguity. To overcome this challenge, we introduce the cycle consistency constraint into region-phrase pairs, which strengthens correlated pairs and weakens unrelated pairs. This cycle pairing makes use of the bidirectional association between image regions and text phrases to alleviate matching ambiguity. Furthermore, we propose a parallel grounding framework, where backbone networks and subsequent relation modules extract individual and contextual representations to calculate context-free and context-aware similarities between regions and phrases separately. Those two representations characterize visual/linguistic individual concepts and inter-relationships, respectively, and then complement each other to achieve cross-modal alignment. The whole framework is trained by minimizing an image-text contrastive loss and a cycle consistency loss. During inference, the above two similarities are fused to give the final region-phrase matching score. Experiments on five popular datasets about visual grounding demonstrate a noticeable improvement in our method. The source code is available at https://github.com/Evergrow/WSVG. Ruisong Zhang, Chuang Wang 0007, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Cross-Lingual Text Image Recognition via Multi-Hierarchy Cross-Modal MimicabstractOptical character recognition and machine translation are usually studied and applied separately. In this paper, we consider a new problem named cross-lingual text image recognition (CLTIR) that integrates these two tasks together. The core of this problem is to recognize source language texts shown in images and transcribe them to the target language in an end-to-end manner. Traditional cascaded systems perform text image recognition and text translation sequentially. This can lead to error accumulation and parameter redundancy problems. To overcome these problems, we propose a multihierarchy cross-modal mimic (MHCMM) framework for end-to-end CLTIR, which can be trained with a massive bilingual text corpus and a small number of bilingual annotated text images. In this framework, a plug-in machine translation model is used as a teacher to guide the CLTIR model for learning representations compatible with image and text modes. Via adversarial learning and attention mechanisms, the proposed mimic method can integrate both global and local information in the semantic space. Experiments on a newly collected dataset demonstrate the superiority of the proposed framework. Our method outperforms other pipelines while containing fewer parameters. Additionally, the MHCMM framework can utilize a large-scale bilingual corpus to further improve the performance efficiently. The visualization of attention scores indicates that the proposed model can read text images in a fashion similar to the machine translation model reading text tokens. Zhuo Chen 0051, Qing Yang 0002, Cheng-Lin Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | A Two-Level Rectification Attention Network for Scene Text RecognitionabstractScene text recognition is a challenging task in the computer vision field due to the diversity of text styles and the complexity of the image backgrounds. In recent decades, numerous text rectification and recognition methods have been proposed to solve these problems. However, most of these methods rectify texts at the geometry level or pixel level. The former is limited by geometric constraints, and the latter is prone to blurring the text. In this paper, we propose a two-level rectification attention network (TRAN) to rectify and recognize texts. This network consists of two parts: a two-level rectification network (TORN) and an attention-based recognition network (ABRN). Specifically, the TORN first rectifies texts at the geometry level and then performs a pixel-level adjustment, which not only eliminates the geometric constraints but also renders clear texts. The ABRN’s role is to recognize text in the rectified images. To improve the feature extraction ability of our model, we design a new channel-wise and kernel-wise attention unit, which enables the network to handle significant variations of character size and channel interdependencies. Furthermore, we propose a skip training strategy to make our model converge smoothly. We conduct experiments on various benchmarks, including regular and irregular datasets. The experimental results show that our method achieves a state-of-the-art performance. Lintai Wu, Yong Xu 0001, Junhui Hou, C. L. Philip Chen, Cheng-Lin Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Rethinking Confidence Calibration for Failure Prediction
Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ECCV (25) | 4 |
| 2022 | EAU-Net: A New Edge-Attention Based U-Net for Nationality Identification
Aritro Pal Choudhury, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICFHR | 4 |
| 2022 | An Efficient Prototype-Based Model for Handwritten Text Recognition with Multi-loss Fusion
Mingming Yu, Heng Zhang 0028, Cheng-Lin Liu 0001 |
ICFHR | 4 |
| 2022 | A Large-Scale Database for Chemical Structure Recognition and Preliminary EvaluationabstractChemical structure recognition (CSR), transforming chemical structure images into formulas in character strings (such as SMILES), is a challenging problem due to the complex 2D structures and relationships. For this research, there is not a database of sufficient scale and diversity for model design and fair evaluation. In this paper, we present a large-scale chemical structure database named CASIA-CSDB, containing 480,668 samples (images corresponding to SMILES strings). To construct the database, we select chemical structures from the ChEMBL, a well-known bioactive molecules database, and use the RDKit tool to generate images according to the chemical format SMILES strings. The selected structures represent the major types of chemical compounds covering eight weight partitions. We also select a subset of 97,309 samples of the database to form the Mini-CASIA-CSDB database. To provide a benchmark, we evaluate three state-of-the-art image-to-markup recognition methods on the database. The results demonstrate the challenge of the database. The database with its annotation is available at http://www.nlpr.ia.ac.cn/databases/CASIA-CSDB/index.html. Longfei Ding, Mengbiao Zhao, Shuiling Zeng, Cheng-Lin Liu 0001 |
ICPR | 5 |
| 2022 | Primitive Contrastive Learning for Handwritten Mathematical Expression RecognitionabstractContrastive learning has gained significant attention recently as it can learn a representation from a large amount of unlabeled training data to improve downstream tasks. While the existing approaches mainly focus on standard tasks of image classification and object detection, they are not easily applied to structured prediction problems. In this paper, we propose an unsupervised pre-trained model, called PrimCLR, for handwritten mathematical expression recognition. For a formula recognition model of encoder-decoder architecture, a pre-trained representation is obtained by PrimCLR, where the contrastive loss is computed from pairs of patches so as to better discriminate primitives. The pre-trained representation is transferred to downstream formula recognition with supervised fine-tuning. Experiments show that pre-training by PrimCLR can significantly improve the formula recognition performance, and PrimCLR shows superiority to conventional contrastive learning methods. Our model achieves state-of-the-art performance on standard datasets CROHME 2016 and CROHME 2019. Hong-Yu Guo, Chuang Wang 0007, Heng-Ye Liu, Jin-Wen Wu, Cheng-Lin Liu 0001 |
ICPR | 6 |
| 2022 | DPAM: A New Deep Parallel Attention Model for Multiple License Plate Number RecognitionabstractLicense plate number recognition is challenging for complex scenes containing multiple vehicles of different types, shapes, distances etc. To recognize multiple license plate numbers in an image, we propose a new model, called Deep Parallel Attention Model (DPAM), which simultaneously extracts unique features at character levels. The proposed model exploits the observation that the combination of alphanumeric characters does not have correlation at semantic level for extracting the features. This led to the introduction of parallelism for feature extraction at character levels to make it efficient in terms of time to fit in a real time environment. To test the proposed model, we consider our own dataset consisting of Indian license plate numbers and other standard datasets to show the superiority of the proposed model over the existing methods in terms of recognition rate. Furthermore, the proposed method is tested on scene text dataset to show its ability to detect text in natural scene images. Amish Kumar, Palaiahnakote Shivakumara, Pinaki Nath Chowdhury, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICPR | 5 |
| 2022 | Document Image Rectification in Complex Scene Using Stacked Siamese NetworksabstractWith the popularity of digital cameras and smart-phones, capturing document images of physical documents for electronic storage has become popular, but the captured document images suffer various deformations. Document image rectification has been studied intensively, but existing methods do not perform sufficiently for document images captured in complex scenes due to the various environmental factors. In this paper, we propose an end-to-end rectification model by stacking 3D and 2D Siamese networks. Three regularization terms are used to enforce 3D reconstruction consistency and 2D texture consistency, respectively. Experimental results on real world datasets demonstrate that the three regularization terms with Siamese networks can significantly improve the rectification performance, and our method performs superiorly compared to state-of-the-art methods. Peipei Yang, Cheng-Lin Liu 0001 |
ICPR | 4 |
| 2022 | Plane Geometry Diagram ParsingabstractGeometry diagram parsing plays a key role in geometry problem solving, wherein the primitive extraction and relation parsing remain challenging due to the complex layout and between-primitive relationship. In this paper, we propose a powerful diagram parser based on deep learning and graph reasoning. Specifically, a modified instance segmentation method is proposed to extract geometric primitives, and the graph neural network (GNN) is leveraged to realize relation parsing and primitive classification incorporating geometric features and prior knowledge. All the modules are integrated into an end-to-end model called PGDPNet to perform all the sub-tasks simultaneously. In addition, we build a new large-scale geometry diagram dataset named PGDP5K with primitive level annotations. Experiments on PGDP5K and an existing dataset IMP-Geometry3K show that our model outperforms state-of-the-art methods in four sub-tasks remarkably. Our code, dataset and appendix material are available at https://github.com/mingliangzhang2018/PGDP. Mingliang Zhang 0005, Yi-Han Hao, Cheng-Lin Liu 0001 |
IJCAI | 4 |
| 2022 | Convolutional Prototype Network for Open Set RecognitionabstractDespite the success of convolutional neural network (CNN) in conventional closed-set recognition (CSR), it still lacks robustness for dealing with unknowns (those out of known classes) in open environment. To improve the robustness of CNN in open-set recognition (OSR) and meanwhile maintain its high accuracy in CSR, we propose an alternative deep framework called convolutional prototype network (CPN), which keeps CNN for representation learning but replaces the closed-world assumed softmax with an open-world oriented and human-like prototype model. To equip CPN with discriminative ability for classifying known samples, we design several discriminative losses for training. Moreover, to increase the robustness of CPN for unknowns, we interpret CPN from the perspective of generative model and further propose a generative loss, which is essentially maximizing the log-likelihood of known samples and serves as a latent regularization for discriminative learning. The combination of discriminative and generative losses makes CPN a hybrid model with advantages for both CSR and OSR. Under the designed losses, the CPN is trained end-to-end for learning the convolutional network and prototypes jointly. For application of CPN in OSR, we propose two rejection rules for detecting different types of unknowns. Experiments on several datasets demonstrate the efficiency and effectiveness of CPN for both CSR and OSR tasks. Hong-Ming Yang, Xu-Yao Zhang, Qing Yang 0002, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Cross-modal prototype learning for zero-shot handwritten character recognition
Xiang Ao 0002, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2022 | Table Structure Recognition and Form Parsing by End-to-End Object Detection and Relation Parsing
Xiao-Hui Li 0012, He-Sen Dai, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2022 | Query Pixel Guided Stroke Extraction with Model-Based Matching for Offline Handwritten Chinese Characters
Tie-Qiang Wang, Xiaoyi Jiang 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2022 | A comprehensive scheme for tattoo text detectionabstractTattoo text detection provides a vital clue for person and crime identification. Due to the freestyle and unconstrained nature of handwritten tattoo text over skin regions, accurate tattoo text detection is very challenging. This paper proposes a comprehensive scheme for tattoo text detection which comprises (a) adaptive Deformable Convolutional Neural Network (DCNN) for skin region detection to reduce text detection complexity (b) a Decoupled Gradient Text Detector (DGTD) for tattoo text detection from skin region (c) a Deep Q-Network (DQN) to refine the bounding boxes detected by DGTD, and (d) a Term-Frequency-Inverse-Document-Frequency (TF-IDF) model to group the words into text lines based on semantic information to fix the bounding box for the line. To test the effectiveness, the proposed method is evaluated on different datasets, namely, (i) a newly developed tattoo text dataset, (ii) benchmark bib number dataset of the marathon, and (iii) person re-identification dataset. The proposed method achieves 91.2, 87.5, and 88.8 F-scores from these three respective datasets. To demonstrate its superior performance, the text detection module (without skin detection) is also compared with state-of-the-art scene text detection methods on benchmark datasets, namely, ICDAR 2019 ArT, Total-Text, and DAST1500 and the proposed method achieves 90.3, 88.5 and 89.8 F-score from these respective datasets. Ayan Banerjee 0002, Palaiahnakote Shivakumara, Umapada Pal 0001, Ramachandra Raghavendra, Cheng-Lin Liu 0001 |
Pattern Recognit. Lett. | 5 |
| 2022 | Decision-Based Adversarial Attack With Frequency MixupabstractIt has been widely observed that deep neural networks are highly vulnerable to adversarial examples. Decision-based attacks could generate adversarial examples based solely on top-1 labels returned by the target model. However, they typically make excessive queries and could not bypass detection effectively. To comprehensively assess a decision-based attack, besides its query efficiency, the performance against detection is also a concern. Considering that previous detections consume massive resources and always mistakenly recognize benign video frames as malicious attacks, we design a lightweight detection calledboundary detectionto overcome the above limitations, whose success reveals serious limitations of existing decision-based attacks. To develop more powerful attacks, we first presentf-mixupas a basic method to produce candidate adversarial examples in the frequency domain. Usingf-mixupas the building block, we proposef-attackas a complete decision-based attack. With the help of several natural images,f-attackcould both work well with limited (hundreds of) queries and bypass detection effectively. Nevertheless, if the attacker could make relatively adequate (thousands of) queries and the target model is not equipped with detection,f-attackwill lag behind existing decision-based attacks. We additionally introducefrequency binary searchbased onf-mixup, which serves as a plug-and-play module for existing decision-based attacks to further improve their query efficiency. Experimental results verify the effectiveness of our proposed methods. Xiu-Chuan Li, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Unsupervised Structure-Texture Separation Network for Oracle Character RecognitionabstractOracle bone script is the earliest-known Chinese writing system of the Shang dynasty and is precious to archeology and philology. However, real-world scanned oracle data are rare and few experts are available for annotation which make the automatic recognition of scanned oracle characters become a challenging task. Therefore, we aim to explore unsupervised domain adaptation to transfer knowledge from handprinted oracle data, which are easy to acquire, to scanned domain. We propose a structure-texture separation network (STSN), which is an end-to-end learning framework for joint disentanglement, transformation, adaptation and recognition. First, STSN disentangles features into structure (glyph) and texture (noise) components by generative models, and then aligns handprinted and scanned data in structure feature space such that the negative influence caused by serious noises can be avoided when adapting. Second, transformation is achieved via swapping the learned textures across domains and a classifier for final classification is trained to predict the labels of the transformed scanned characters. This not only guarantees the absolute separation, but also enhances the discriminative ability of the learned features. Extensive experiments on Oracle-241 dataset show that STSN outperforms other adaptation methods and successfully improves recognition performance on scanned data even when they are contaminated by long burial and careless excavation. Mei Wang 0001, Weihong Deng, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Mixed-Supervised Scene Text Detection With Expectation-Maximization AlgorithmabstractScene text detection is an important and challenging task in computer vision. For detecting arbitrarily-shaped texts, most existing methods require heavy data labeling efforts to produce polygon-level text region labels for supervised training. In order to reduce the cost in data labeling, we study mixed-supervised arbitrarily-shaped text detection by combining various weak supervision forms (e.g., image-level tags, coarse, loose and tight bounding boxes), which are far easier to annotate. Whereas the existing weakly-supervised learning methods (such as multiple instance learning) do not promote full object coverage, to approximate the performance of fully-supervised detection, we propose an Expectation-Maximization (EM) based mixed-supervised learning framework to train scene text detector using only a small amount of polygon-level annotated data combined with a large amount of weakly annotated data. The polygon-level labels are treated as latent variables and recovered from the weak labels by the EM algorithm. A new contour-based scene text detector is also proposed to facilitate the use of weak labels in our mixed-supervised learning framework. Extensive experiments on six scene text benchmarks show that (1) using only 10% strongly annotated data and 90% weakly annotated data, our method yields comparable performance to that of fully supervised methods, (2) with 100% strongly annotated data, our method achieves state-of-the-art performance on five scene text benchmarks (CTW1500, Total-Text, ICDAR-ArT, MSRA-TD500, and C-SVT), and competitive results on the ICDAR2015 Dataset. We will make our weakly annotated datasets publicly available. Mengbiao Zhao, Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | Instance GNN: A Learning Framework for Joint Symbol Segmentation and Recognition in Online Handwritten DiagramsabstractOnline handwritten diagram recognition (OHDR) has attracted considerable attention for its potential applications in many areas, but it is a challenging task due to the complex 2D structure, writing style variation, and lack of annotated data. Existing OHDR methods often have limitations in modeling and learning complex contextual relationships. To overcome these challenges, we propose an OHDR method based on graph neural networks (GNNs) in this paper. In particular, we formulate symbol segmentation and symbol recognition as node clustering and node classification problems on stroke graphs and solve the problems jointly under a unified learning framework with a GNN model. This GNN model is denoted as Instance GNN since it gives the symbol instance label as well as the semantic label. Extensive experiments on two flowchart datasets and a finite automata dataset show that our method consistently outperforms previous methods with large margins and achieves state-of-the-art performance. In addition, we release a large-scale annotated online handwritten flowchart dataset, CASIA-OHFC, and provide initial experimental results as a baseline. Xiao-Long Yun, Yan-Ming Zhang 0001, Cheng-Lin Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Meta-Prototypical Learning for Domain-Agnostic Few-Shot RecognitionabstractFew-shot learning (FSL) aims to classify novel images based on a few labeled samples with the help of meta-knowledge. Most previous works address this problem based on the hypothesis that the training set and testing set are from the same domain, which is not realistic for some real-world applications. Thus, we extend FSL to domain-agnostic few-shot recognition, where the domain of the testing task is unknown. In domain-agnostic few-shot recognition, the model is optimized on data from one domain and evaluated on tasks from different domains. Previous methods for FSL mostly focus on learning general features or adapting to few-shot tasks effectively. They suffer from inappropriate features or complex adaptation in domain-agnostic few-shot recognition. In this brief, we propose meta-prototypical learning to address this problem. In particular, a meta-encoder is optimized to learn the general features. Different from the traditional prototypical learning, the meta encoder can effectively adapt to few-shot tasks from different domains by the traces of the few labeled examples. Experiments on many datasets demonstrate that meta-prototypical learning performs competitively on traditional few-shot tasks, and on few-shot tasks from different domains, meta-prototypical learning outperforms related methods. Rui-Qi Wang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Deep Neural Network Self-Distillation Exploiting Data Representation InvarianceabstractTo harvest small networks with high accuracies, most existing methods mainly utilize compression techniques such as low-rank decomposition and pruning to compress a trained large model into a small network or transfer knowledge from a powerful large model (teacher) to a small network (student). Despite their success in generating small models of high performance, the dependence of accompanying assistive models complicates the training process and increases memory and time cost. In this article, we propose an elegant self-distillation (SD) mechanism to obtain high-accuracy models directly without going through an assistive model. Inspired by the invariant recognition in the human vision system, different distorted instances of the same input should possess similar high-level data representations. Thus, we can learn data representation invariance between different distorted versions of the same sample. Especially, in our learning algorithm based on SD, the single network utilizes the maximum mean discrepancy metric to learn the global feature consistency and the Kullback-Leibler divergence to constrain the posterior class probability consistency across the different distorted branches. Extensive experiments on MNIST, CIFAR-10/100, and ImageNet data sets demonstrate that the proposed method can effectively reduce the generalization error for various network architectures, such as AlexNet, VGGNet, ResNet, Wide ResNet, and DenseNet, and outperform existing model distillation methods with little extra training efforts. Ting-Bing Xu, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Proxy Graph Matching with Proximal Matching NetworksabstractEstimating feature point correspondence is a common technique in computer vision. A line of recent data-driven approaches utilizing the graph neural networks improved the matching accuracy by a large margin. However, these learning-based methods require a lot of labeled training data, which are expensive to collect. Moreover, we find most methods are sensitive to global transforms, for example, a random rotation. On the contrary, classical geometric approaches are immune to rotational transformation though their performance is generally inferior. To tackle these issues, we propose a new learning-based matching framework, which is designed to be rotationally invariant. The model only takes geometric information as input. It consists of three parts: a graph neural network to generate a high-level local feature, an attention-based module to normalize the rotational transform, and a global feature matching module based on proximal optimization. To justify our approach, we provide a convergence guarantee for the proximal method for graph matching. The overall performance is validated by numerical experiments. In particular, our approach is trained on the synthetic random graphs and then applied to several real-world datasets. The experimental results demonstrate that our method is robust to rotational transform and highlights its strong performance of matching accuracy. Haoru Tan, Chuang Wang 0007, Sitong Wu, Tie-Qiang Wang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
AAAI | 6 |
| 2021 | Graph-to-Graph: Towards Accurate and Interpretable Online Handwritten Mathematical Expression RecognitionabstractRecent handwritten mathematical expression recognition (HMER) approaches treat the problem as an image-to-markup generation task where the handwritten formula is translated into a sequence (e.g. LaTeX). The encoder-decoder framework is widely used to solve this image-to-sequence problem. However, (i) for structured mathematical formula, the hierarchical structure neither in the formula nor in the markup has been explored adequately. In addition, (ii) existing image-to-markup methods could not explicitly segment mathematical symbols in the formula corresponding to each target markup token. In this paper, we address the above issues by formulating the HMER as a graph-to-graph (G2G) learning problem. Graph is more flexible and general for structure representation and learning compared with image or sequence. At the core of our method lies the embedding of input formula and output markup into graphs on primitives, with Graph Neural Networks (GNN) to explore the structural information, and a novel sub-graph attention mechanism to match primitives in the input and output graphs. We conduct extensive experiments on CROHME datasets to demonstrate the benefits of the proposed G2G model. Our method yields significant improvements over previous SOTA image-to-markup systems. Moreover, it explicitly resolves the symbol segmentation problem while still being trained end-to-end, making the whole system much more accurate and interpretable. Jin-Wen Wu, Yan-Ming Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
AAAI | 5 |
| 2021 | Semantic-Aware Video Text DetectionabstractMost existing video text detection methods track texts with appearance features, which are easily influenced by the change of perspective and illumination. Compared with appearance features, semantic features are more robust cues for matching text instances. In this paper, we propose an end-to-end trainable video text detector that tracks texts based on semantic features. First, we introduce a new character center segmentation branch to extract semantic features, which encode the category and position of characters. Then we propose a novel appearance-semantic-geometry descriptor to track text instances, in which se-mantic features can improve the robustness against appearance changes. To overcome the lack of character-level an-notations, we propose a novel weakly-supervised character center detection module, which only uses word-level annotated real images to generate character-level labels. The proposed method achieves state-of-the-art performance on three video text benchmarks ICDAR 2013 Video, Minetto and RT-1K, and two Chinese scene text benchmarks CA-SIA10K and MSRA-TD500. Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 4 |
| 2021 | Prototype Augmentation and Self-Supervision for Incremental LearningabstractDespite the impressive performance in many individual tasks, deep neural networks suffer from catastrophic forgetting when learning new tasks incrementally. Recently, various incremental learning methods have been proposed, and some approaches achieved acceptable performance relying on stored data or complex generative models. However, storing data from previous tasks is limited by memory or privacy issues, and generative models are usually unstable and inefficient in training. In this paper, we propose a simple non-exemplar based method named PASS, to address the catastrophic forgetting problem in incremental learning. On the one hand, we propose to memorize one class-representative prototype for each old class and adopt prototype augmentation (protoAug) in the deep feature space to maintain the decision boundary of previous tasks. On the other hand, we employ self-supervised learning (SSL) to learn more generalizable and transferable features for other tasks, which demonstrates the effectiveness of SSL in incremental learning. Experimental results on benchmark datasets show that our approach significantly outperforms non-exemplar based methods, and achieves comparable performance compared to exemplar based approaches. Fei Zhu 0004, Xu-Yao Zhang, Chuang Wang 0007, Cheng-Lin Liu 0001 |
CVPR | 5 |
| 2021 | Adaptive Scaling for Archival Table Structure Recognition
Xiao-Hui Li 0012, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICDAR (1) | 4 |
| 2021 | Document Dewarping with Control Points
Guo-Wang Xie, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICDAR (1) | 4 |
| 2021 | Recognizing Handwritten Chinese Texts with Insertion and Swapping Using a Structural Attention Network
Jin-Wen Wu, Cheng-Lin Liu 0001 |
ICDAR (4) | 4 |
| 2021 | Handwriting Trajectory Recovery from Off-Line Multi-Stroke Characters by Deep Ordering Prediction and Heuristic SearchabstractStroke order recovery from off-line multi-stroke characters is a great challenge due to the ambiguity in intersection and connection among strokes. In this paper, we propose a novel framework for handwriting trajectory recovery from off-line handwritten characters by deep neural network (DNN) based ordering prediction and heuristic search, where several DNN modules are designed to extract stroke skeleton, ambiguous zones and starting points, respectively. Then, the ordering matrix Moamong all the stroke segments is calculated by a pointer network (Ptr-Net). Besides, a convolutional neural network (CNN) is used to measure the time adjacency between two arbitrary segments. Based on these necessary measurements, the final writing order is decided via searching for the optimal permutation by heuristic A∗search. Experiments on handwriting images synthesized from the public online handwriting datasets CASIA-OLHWDB1.1, ICDAR13-Online and UNIPEN show that the proposed method yields superior performance on Chinese and English/Arabic hand-writing. Tie-Qiang Wang, Cheng-Lin Liu 0001 |
ICME | 2 |
| 2021 | GSS: Graph-Based Subspace Learning with Shots Initialization for few-shot RecognitionabstractPrevious methods of few-shot Learning mostly solve different few-shot recognition tasks in an identical feature space. But identical features are hard to fit various tasks. Some works show that learning a unique subspace for each few-shot recognition task can improve the signal-noise ratio (SNR) of the features and boost the performance. However, there are still two problems remaining. First, in constructing the subspace for few-shot task, often some information (embeddings of queries or labels of shots) are discarded. Second, the eigendecomposition of covariance matrix is usually needed, which degrades the efficiency of the whole model. In this paper, we propose Graph-based Subspace learning with Shots initialization (GSS) for few-shot recognition to learn a better subspace efficiently. In GSS, the bases of the subspace are directly initialized with labels based on shots (given labeled samples) and iteratively updated for better discrimination based on a graph that connects bases and all samples. Extensive experiments on four few-shot benchmark datasets show that GSS reports better performance and higher efficient compared with previous subspace based methods and achieves state-of-the-art performance. Rui-Qi Wang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICME | 3 |
| 2021 | Class Forge: Boosting Feature Encoder for Few-Shot Learning with Synthesized ClassesabstractFew-shot Learning (FSL) aims to gain classification ability on novel classes with only a few labeled samples. Previous works explore meta-learning, metric learning, and graph based methods. Though data augmentation is important to enhance the generalizability of neural networks, it is not well exploited in the field of FSL. We investigate the augmentation in FSL and propose Class Forge to synthesize forged classes that help to learn an encoder with better generalization to novel classes. Specifically, Class Forge divides given base visual classes into parts and combines these parts to synthesize forged visual classes. Training with the additional forged classes forces the encoder to learn richer features that can embed different parts, so as to boost the generalization to novel classes. Intrinsically, Class Forge is a "class augmentation" method that provides a simple yet effective way to synthesize classes, other than synthesizing samples of given classes in previous works. Extensive experiments show that Class Forge yields consistent performance gain on different datasets for FSL. And the ablation studies validate that features learned with Class Forge demonstrate better generalization ability. Rui-Qi Wang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICME | 3 |
| 2021 | Calibration for Non-Exemplar Based Class-Incremental LearningabstractCatastrophic forgetting is the central challenge in incremental learning. Notable studies address the problem by using regularization or experience replay strategies. However, the performance is far from ideal without storing previous data, especially in the scenario of class-incremental learning (CIL). In CIL setting, an important factor causing catastrophic forgetting is the severe bias between the new and previously learned classes, in both classifier and feature extractor. In this paper, we propose calibrateCIL which contains two simple modifications to calibrate the bias in non-exemplar based CIL. Specifically, local softmax is proposed to calibrate the classifier, and cutout training is used to calibrate the feature extractor by learning richer, more generalizable and transferable features. Our method can give balance class scores without any post-processing technique. We show that our method outperforms state-of-the-art non-exemplar based methods on the challenging problem of CIL, and the ablation study demonstrates the effectiveness of the two modifications. Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICME | 3 |
| 2021 | Regularizing CTC in Expectation-Maximization Framework with Application to Handwritten Text RecognitionabstractConnectionist Temporal Classification (CTC) is an objective function for sequence learning and has shown promising results in speech and text recognition tasks. However, its inherent mechanism has not been investigated thoroughly. In this paper, we propose a theoretical explanation of CTC from the perspective of the Expectation-Maximization (EM) algorithm. Based on the EM analysis, we propose a pseudo-label-based L1 regularization and voting decoding algorithm to improve the performance of text recognition. The L1 regularization can reduce the pseudo-label estimation error, while the voting decoding algorithm modifies the built-in decoding logic of CTC and introduces a voting mechanism to the inference process. Experiments of handwritten text recognition show that the proposed method consistently improves over the CTC baseline and yields state-of-the-art results on three benchmark datasets. Likun Gao, Heng Zhang 0028, Cheng-Lin Liu 0001 |
IJCNN | 3 |
| 2021 | Region Ensemble Network for MCI Conversion Prediction with a Relation Regularized Loss
Yuan-Xing Zhao, Yan-Ming Zhang 0001, Cheng-Lin Liu 0001 |
MICCAI (5) | 4 |
| 2021 | Learning to Understand Traffic SignsabstractOne of the intelligent transportation system's critical tasks is to understand traffic signs and convey traffic information to humans. However, most related works are focused on the detection and recognition of traffic sign texts or symbols, which is not sufficient for understanding. Besides, there has been no public dataset for traffic sign understanding research. Our work takes the first step towards addressing this problem. First, we propose a "CASIA-Tencent Chinese Traffic Sign Understanding Dataset" (CTSU Dataset), which contains 5000 images of traffic signs with rich semantic descriptions. Second, we introduce a novel multi-task learning architecture that extracts text and symbol information from traffic signs, reasons the relationship between texts and symbols, classifies signs into different categories, and finally, composes the descriptions of the signs. Experiments show that the task of traffic sign understanding is achievable, and our architecture demonstrates state-of-the-art and superior performance. The CTSU Dataset is available at http://www.nlpr.ia.ac.cn/databases/CASIA-Tencent%20CTSU/index.html. Wei Feng 0016, Shuqi Mei, Cheng-Lin Liu 0001 |
ACM Multimedia | 6 |
| 2021 | Class-Incremental Learning via Dual AugmentationabstractDeep learning systems typically suffer from catastrophic forgetting of past knowledge when acquiring new skills continually. In this paper, we emphasize two dilemmas, representation bias and classifier bias in class-incremental learning, and present a simple and novel approach that employs explicit class augmentation (classAug) and implicit semantic augmentation (semanAug) to address the two biases, respectively. On the one hand, we propose to address the representation bias by learning transferable and diverse representations. Specifically, we investigate the feature representations in incremental learning based on spectral analysis and present a simple technique called classAug, to let the model see more classes during training for learning representations transferable across classes. On the other hand, to overcome the classifier bias, semanAug implicitly involves the simultaneous generating of an infinite number of instances of old classes in the deep feature space, which poses tighter constraints to maintain the decision boundary of previously learned classes. Without storing any old samples, our method can perform comparably with representative data replay based approaches. Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
NeurIPS | 4 |
| 2021 | Residual Dual Scale Scene Text Spotting by Fusing Bottom-Up and Top-Down Processing
Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 5 |
| 2021 | Multi-branch guided attention network for irregular text recognition
Cheng-Lin Liu 0001 |
Neurocomputing | 2 |
| 2021 | BlockQNN: Efficient Block-Wise Neural Network Architecture GenerationabstractConvolutional neural networks have gained a remarkable success in computer vision. However, most popular network architectures are hand-crafted and usually require expertise and elaborate design. In this paper, we provide a block-wise network generation pipeline called BlockQNN which automatically builds high-performance networks using the Q-Learning paradigm with epsilon-greedy exploration strategy. The optimal network block is constructed by the learning agent which is trained to choose component layers sequentially. We stack the block to construct the whole auto-generated network. To accelerate the generation process, we also propose a distributed asynchronous framework and an early stop strategy. The block-wise generation brings unique advantages: (1) it yields state-of-the-art results in comparison to the hand-crafted networks on image classification, particularly, the best network generated by BlockQNN achieves 2.35 percent top-1 error rate on CIFAR-10. (2) it offers tremendous reduction of the search space in designing networks, spending only 3 days with 32 GPUs. A faster version can yield a comparable result with only 1 GPU in 20 hours. (3) it has strong generalizability in that the network built on CIFAR also performs well on the larger-scale dataset. The best network achieves very competitive accuracy of 82.0 percent top-1 and 96.0 percent top-5 on ImageNet. Zhao Zhong, Zichen Yang, Boyang Deng, Wei Wu 0021, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2021 | SNAP: Shaping neural architectures progressively via information density criterion
Zhiqiang Chen 0002, Ting-Bing Xu, Weijian Liao, Zhengcheng Li, Jinpeng Li 0002, Cheng-Lin Liu 0001, Huiguang He |
Pattern Recognit. | 6 |
| 2021 | Joint stroke classification and text line grouping in online handwritten documents with edge pooling attention networks
Jun-Yu Ye, Yan-Ming Zhang 0001, Qing Yang 0002, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2021 | Deformable scene text detection using harmonic features and modified pixel aggregation network
Tanmay Jain, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. Lett. | 4 |
| 2021 | Dynamical Channel Pruning by Conditional Accuracy Change for Deep Neural NetworksabstractChannel pruning is an effective technique that has been widely applied to deep neural network compression. However, many existing methods prune from a pretrained model, thus resulting in repetitious pruning and fine-tuning processes. In this article, we propose a dynamical channel pruning method, which prunes unimportant channels at the early stage of training. Rather than utilizing some indirect criteria (e.g., weight norm, absolute weight sum, and reconstruction error) to guide connection or channel pruning, we design criteria directly related to the final accuracy of a network to evaluate the importance of each channel. Specifically, a channelwise gate is designed to randomly enable or disable each channel so that the conditional accuracy changes (CACs) can be estimated under the condition of each channel disabled. Practically, we construct two effective and efficient criteria to dynamically estimate CAC at each iteration of training; thus, unimportant channels can be gradually pruned during the training process. Finally, extensive experiments on multiple data sets (i.e., ImageNet, CIFAR, and MNIST) with various networks (i.e., ResNet, VGG, and MLP) demonstrate that the proposed method effectively reduces the parameters and computations of baseline network while yielding the higher or competitive accuracy. Interestingly, if we Double the initial Channels and then Prune Half (DCPH) of them to baseline's counterpart, it can enjoy a remarkable performance improvement by shaping a more desirable structure. Zhiqiang Chen 0002, Ting-Bing Xu, Changde Du, Cheng-Lin Liu 0001, Huiguang He |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Page Segmentation Using Convolutional Neural Network and Graphical Model
Xiao-Hui Li 0012, Cheng-Lin Liu 0001 |
DAS | 3 |
| 2020 | Dewarping Document Image by Displacement Flow Estimation with Fully Convolutional Network
Guo-Wang Xie, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
DAS | 4 |
| 2020 | Weakly Supervised Learning for Over-Segmentation Based Handwritten Chinese Text RecognitionabstractIn this paper, we proposed a weakly supervised learning method for string-level training of character classifier in over-segmentation based handwritten Chinese text recognition (HCTR). The over-segmentation based framework can easily integrate multiple context models and provide accurate character boundary and recognition confidence, but has not been implemented with string-level training for HCTR. We propose to optimize the character classifier by minimizing the marginal log-likelihood on a string-level annotated handwriting dataset, where the forward-backward algorithm is utilized in a segmentation-and-recognition lattice. Experimental results on the CASIA-HWDB and ICDAR-2013 competition datasets show that the proposed method improves the recognition performance significantly, which demonstrates its effectiveness. Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
ICFHR | 4 |
| 2020 | ICFHR 2020 Competition on Offline Recognition and Spotting of Handwritten Mathematical Expressions - OffRaSHMEabstractThis paper presents the competition on Offline Recognition and Spotting of Handwritten Mathematical Expressions (OffRaSHME) held at the 17th International Conference on Frontiers in Handwriting Recognition (ICFHR 2020). Handwritten Mathematical Expression Recognition (HMER) has wide potential applications and is capturing increasing attention in recent years. Previous HMER competitions mainly focused on online datasets or offline datasets that are converted from online data. In this competition, we have collected a dataset of offline handwritten mathematical expressions by scanning papers that contain expressions. Moreover, we labeled the offline dataset at symbol level, i.e., the bounding boxes of each symbol are also provided, to facilitate the research of HMER. At last, 19,749 offline handwritten mathematical expressions are collected for training, and 2,000 ones are provided for evaluating the participating systems. In the competition, 7 teams submitted 8 systems for the task of offline HMER, among which 5 systems only use the provided datasets without any extra data while 3 systems use extra data. The winner team achieved a recognition accuracy of 79.85% (without extra data) and 81.85% (with extra data) on the offline formula recognition task. Dahan Wang, Jin-Wen Wu, Yu-Pei Yan, Zhicai Huang, Gui-Yun Chen, Cheng-Lin Liu 0001 |
ICFHR | 8 |
| 2020 | SRR-GAN: Super-Resolution based Recognition with GAN for Low-Resolved Text ImagesabstractText images convey important information for various applications, while the recognition of low-resolution text images is a challenge. Most existing methods solve this problem using a cascaded scheme in two steps: image super-resolution and high-resolution text recognition. In this paper, we propose a novel framework, called SRR-GAN, which integrates text recognition with super-resolution via adversarial learning. By joint training of recognition and super-resolution models, more generic features for images of various quality can be learned, so as to yield high recognition performance for both high-resolution and low-resolution images. Experiments on natural scene and handwritten texts demonstrate that SRR-GAN outperforms the cascaded scheme on low-resolution images. The results show that SRR-GAN can improve recognition accuracies by 10%-20% relatively on five datasets of scene/handwritten texts. Meanwhile, SRR-GAN maintains high performance on high-resolution images. Ming-Chao Xu, Cheng-Lin Liu 0001 |
ICFHR | 3 |
| 2020 | Cross-Lingual Text Image Recognition via Multi-Task Sequence to Sequence LearningabstractThis paper considers recognizing texts shown in a source language and translating into a target language, without generating the intermediate source language text image recognition results. We call this problem Cross-Lingual Text Image Recognition (CLTIR). To solve this problem, we propose a multi-task system containing a main task of CLTIR and an auxiliary task of Mono-Lingual Text Image Recognition (MLTIR) simultaneously. Two different sequence to sequence learning methods, a convolution based attention model and a Bidirectional Long Short-Term Memory (BLSTM) model with Connectionist Temporal Classification (CTC), are adopted for these tasks respectively. We evaluate the system on a newly collected Chinese-English bilingual movie subtitle image dataset. Experimental results demonstrate the multi-task learning framework performs superiorly in both languages. Zhuo Chen 0051, Xu-Yao Zhang, Qing Yang 0002, Cheng-Lin Liu 0001 |
ICPR | 5 |
| 2020 | F-mixup: Attack CNNs From Fourier PerspectiveabstractRecent research has revealed that deep neural networks are highly vulnerable to adversarial examples. In this paper, different from most adversarial attacks which directly modify pixels in spatial domain, we propose a novel black-box attack in frequency domain, named as f-mixup, based on the property of natural images and perception disparity between human-visual system (HVS) and convolutional neural networks (CNNs): First, natural images tend to have the bulk of their Fourier spectrums concentrated on the low frequency domain; Second, HVS is much less sensitive to high frequencies while CNNs can utilize both low and high frequency information to make predictions. Extensive experiments are conducted and show that deeper CNNs tend to concentrate more on the higher frequency domain, which may explain the contradiction between robustness and accuracy. In addition, we compared f-mixup with existing attack methods and observed that our approach possesses great advantages. Finally, we show that f-mixup can be also incorporated in training to make deep CNNs defensible against a kind of perturbations effectively. Xiu-Chuan Li, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICPR | 4 |
| 2020 | Mutually Guided Dual-Task Network for Scene Text DetectionabstractScene text detection has been studied extensively. Existing methods detect either words or text lines and use either word-level or line-level annotated data for training. In this paper, we propose a dual-task network that can perform word-level and line-level text detection simultaneously and use training data of both levels of annotation to boost the performance. The dual-task network has two detection heads for word-level and line-level text detection, respectively. Then we propose a mutual guidance scheme for the joint training of the two tasks with two modules: line filtering module utilizes the output feature map of the text line detector to filter out the non-text regions for the word detector, and word enhancing module provides prior positions of words for the text line detector depending on the output feature map of the word detector. Experimental results of word-level and line-level text detection demonstrate the effectiveness of the proposed dual-task network and mutual guidance scheme, and the results of our method are competitive with state-of-the-art methods. Mengbiao Zhao, Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICPR | 5 |
| 2020 | Table detection and cell segmentation in online handwritten documents with graph attention networksabstractIn this paper, we propose a multi-task learning approach for table detection and cell segmentation with densely connected graph attention networks in free form online documents. Each online document is regarded as a graph, where nodes represent strokes and edges represent the relationships between strokes. Then we propose a graph attention network model to classify nodes and edges simultaneously. According to node classification results, tables can be detected in each document. By combining node and edge classification resutls, cells in each table can be segmented. To improve information flow in the network and enable efficient reuse of features among layers, dense connectivity among layers is used. Our proposed model has been experimentally validated on an online handwritten document dataset IAMOnDo and achieved encouraging results. Heng Zhang 0028, Xiao-Long Yun, Jun-Yu Ye, Cheng-Lin Liu 0001 |
MMAsia | 5 |
| 2020 | Teaching machines to write like humans using L-attributed grammar
Yunxue Shao, Cheng-Lin Liu 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2020 | Handwritten Mathematical Expression Recognition via Paired Adversarial Learning
Jin-Wen Wu, Yan-Ming Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 5 |
| 2020 | A benchmark for unconstrained online handwritten Uyghur word recognition
Wujiahemaiti Simayi, Mayire Ibrayim, Xu-Yao Zhang, Cheng-Lin Liu 0001, Askar Hamdulla |
Int. J. Document Anal. Recognit. | 4 |
| 2020 | Online semi-supervised learning with learning vector quantization
Yuan-Yuan Shen, Yan-Ming Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Neurocomputing | 4 |
| 2020 | DNA computing inspired deep networks design
Guoqiang Zhong 0001, Tao Li 0031, Wencong Jiao, Li-Na Wang, Junyu Dong, Cheng-Lin Liu 0001 |
Neurocomputing | 6 |
| 2020 | Towards Robust Pattern Recognition: A ReviewabstractThe accuracies for many pattern recognition tasks have increased rapidly year by year, achieving or even outperforming human performance. From the perspective of accuracy, pattern recognition seems to be a nearly solved problem. However, once launched in real applications, the high-accuracy pattern recognition systems may become unstable and unreliable due to the lack of robustness in open and changing environments. In this article, we present a comprehensive review of research toward robust pattern recognition from the perspective of breaking three basic and implicit assumptions: closed-world assumption, independent and identically distributed assumption, and clean and big data assumption, which form the foundation of most pattern recognition models. Actually, our brain is robust at learning concepts continually and incrementally, in complex, open, and changing environments, with different contexts, modalities, and tasks, by showing only a few examples, under weak or noisy supervision. These are the major differences between human intelligence and machine intelligence, which are closely related to the above three assumptions. After witnessing the significant progress in accuracy improvement nowadays, this review paper will enable us to analyze the shortcomings and limitations of current methods and identify future research directions for robust pattern recognition. Xu-Yao Zhang, Cheng-Lin Liu 0001, Ching Y. Suen |
Proc. IEEE | 2 |
| 2020 | MuLTReNets: Multilingual text recognition networks for simultaneous script identification and handwriting recognition
Zhuo Chen 0051, Xu-Yao Zhang, Qing Yang 0002, Cheng-Lin Liu 0001 |
Pattern Recognit. | 5 |
| 2020 | Realtime multi-scale scene text detection with scale-based region proposal network
Xu-Yao Zhang, Zhenbo Luo, Jean-Marc Ogier, Cheng-Lin Liu 0001 |
Pattern Recognit. | 6 |
| 2020 | Special issue on "Applications of graph-based techniques to pattern recognition"
Pasquale Foggia, Mario Vento, Cheng-Lin Liu 0001 |
Pattern Recognit. Lett. | 3 |
| 2020 | Multisource Transfer Learning for Cross-Subject EEG Emotion RecognitionabstractElectroencephalogram (EEG) has been widely used in emotion recognition due to its high temporal resolution and reliability. Since the individual differences of EEG are large, the emotion recognition models could not be shared across persons, and we need to collect new labeled data to train personal models for new users. In some applications, we hope to acquire models for new persons as fast as possible, and reduce the demand for the labeled data amount. To achieve this goal, we propose a multisource transfer learning method, where existing persons are sources, and the new person is the target. The target data are divided into calibration sessions for training and subsequent sessions for test. The first stage of the method is source selection aimed at locating appropriate sources. The second is style transfer mapping, which reduces the EEG differences between the target and each source. We use few labeled data in the calibration sessions to conduct source selection and style transfer. Finally, we integrate the source models to recognize emotions in the subsequent sessions. The experimental results show that the three-category classification accuracy on benchmark SEED improves by 12.72% comparing with the nontransfer method. Our method facilitates the fast deployment of emotion recognition models by reducing the reliance on the labeled data amount, which has practical significance especially in fast-deployment scenarios. Jinpeng Li 0002, Shuang Qiu 0002, Yuan-Yuan Shen, Cheng-Lin Liu 0001, Huiguang He |
IEEE Trans. Cybern. | 4 |
| 2019 | Data-Distortion Guided Self-Distillation for Deep Neural NetworksabstractKnowledge distillation is an effective technique that has been widely used for transferring knowledge from a network to another network. Despite its effective improvement of network performance, the dependence of accompanying assistive models complicates the training process of single network in the need of large memory and time cost. In this paper, we design a more elegant self-distillation mechanism to transfer knowledge between different distorted versions of same training data without the reliance on accompanying models. Specifically, the potential capacity of single network is excavated by learning consistent global feature distributions and posterior distributions (class probabilities) across these distorted versions of data. Extensive experiments on multiple datasets (i.e., CIFAR-10/100 and ImageNet) demonstrate that the proposed method can effectively improve the generalization performance of various network architectures (such as AlexNet, ResNet, Wide ResNet, and DenseNet), outperform existing distillation methods with little extra training efforts. Ting-Bing Xu, Cheng-Lin Liu 0001 |
AAAI | 2 |
| 2019 | TextDragon: An End-to-End Framework for Arbitrary Shaped Text SpottingabstractMost existing text spotting methods either focus on horizontal/oriented texts or perform arbitrary shaped text spotting with character-level annotations. In this paper, we propose a novel text spotting framework to detect and recognize text of arbitrary shapes in an end-to-end manner, using only word/line-level annotations for training. Motivated from the name of TextSnake, which is only a detection model, we call the proposed text spotting framework TextDragon. In TextDragon, a text detector is designed to describe the shape of text with a series of quadrangles, which can handle text of arbitrary shapes. To extract arbitrary text regions from feature maps, we propose a new differentiable operator named RoISlide, which is the key to connect arbitrary shaped text detection and recognition. Based on the extracted features through RoISlide, a CNN and CTC based text recognizer is introduced to make the framework free from labeling the location of characters. The proposed method achieves state-of-the-art performance on two curved text benchmarks CTW1500 and Total-Text, and competitive results on the ICDAR 2015 Dataset. Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICCV | 5 |
| 2019 | Cross-Modal Prototype Learning for Zero-Shot Handwriting RecognitionabstractIn contrast to machine recognizers that rely on training with large handwriting data, humans can recognize handwriting accurately on learning from few samples, and can even generalize to handwritten characters from printed samples. Simulating this ability in machine recognition is important to alleviate the burden of labeling large handwriting data, especially for large category set as in Chinese text. In this paper, inspired by human learning, we propose a cross-modal prototype learning (CMPL) method for zero-shot online handwritten character recognition: for unseen categories, handwritten characters can be recognized without learning from handwritten samples, but instead from printed characters. Particularly, the printed characters (one for each class) are embedded into a convolutional neural network (CNN) feature space to obtain prototypes representing each class, while the online handwriting trajectories are embedded with a recurrent neural network (RNN). Via cross-modal joint learning, handwritten characters can be recognized according to the printed prototypes. For unseen categories, handwritten characters can be recognized by only feeding a printed sample per category. Experiments on a benchmark Chinese handwriting database have shown the effectiveness and potential of the proposed method for zero-shot handwriting recognition. Xiang Ao 0002, Xu-Yao Zhang, Hong-Ming Yang, Cheng-Lin Liu 0001 |
ICDAR | 5 |
| 2019 | Instance Aware Document Image Segmentation using Label Pyramid Networks and Deep Watershed TransformationabstractSegmentation of complex document images remains a challenge due to the large variability of layout and image degradation. In this paper, we propose a method to segment complex document images based on Label Pyramid Network (LPN) and Deep Watershed Transform (DWT). The method can segment document images into instance aware regions including text lines, text regions, figures, tables, etc. The backbone of LPN can be any type of Fully Convolutional Networks (FCN), and in training, label map pyramids on training images are provided to exploit the hierarchical boundary information of regions efficiently through multi-task learning. The label map pyramid is transformed from region class label map by distance transformation and multi-level thresholding. In segmentation, the outputs of multiple tasks of LPN are summed into one single probability map, on which watershed transformation is carried out to segment the document image into instance aware regions. In experiments on four public databases, our method is demonstrated effective and superior, yielding state of the art performance for text line segmentation, baseline detection and region segmentation. Xiao-Hui Li 0012, Jean-Marc Ogier, Cheng-Lin Liu 0001 |
ICDAR | 6 |
| 2019 | A Robust Data Hiding Scheme Using Generated Content for Securing Genuine DocumentsabstractData hiding is an effective technique, compared to pervasive black-and-white code patterns such as barcode and quick response code, which can be used to secure document images against forgery or unauthorized intervention. In this work, we propose a robust digital watermarking scheme for securing genuine documents by leveraging generative adversarial networks (GAN). To begin with, the input document is adjusted to its right form by geometric correction. Next, the generated document is obtained from the input document by using the mentioned networks, and it is regarded as a reference for data hiding and detection. We then introduce an algorithm that hides a secret information into the document and produces a watermarked document whose content is minimally distorted in terms of normal observation. Furthermore, we also present a method that detects the hidden data from the watermarked document by measuring the distance of pixel values between the generated and watermarked document. For improving the security feature, we encode the secret information prior to hiding it by using pseudo random numbers. Lastly, we demonstrate that our approach gives high precision of data detection, and competitive performance compared to state-of-the-art approaches. Cu Vinh Loc, Jean-Christophe Burie, Jean-Marc Ogier, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2019 | Hiding Security Feature Into Text Content for Securing Documents Using Generated FontabstractMotivated by increasing possibility of the tampering of genuine documents during a transmission over digital channels, we focus on developing a watermarking framework for determining whether a given document is genuine or falsified. The proposed framework is performed by hiding a security feature or secret information within the document. In order to hide the security feature, we replace the appropriate characters of legal document by the equivalent characters coming from generated fonts, called hereafter the variations of characters. These variations are produced by training generative adversarial networks (GAN) with the features of character's skeleton and normal shape. Regarding the process of detecting hidden information, we make use of fully convolutional networks (FCN) to produce salient regions from the watermarked document. The salient regions mark positions of document where the characters are substituted by their variations, and these positions are used as a reference for extracting the hidden information. Lastly, we demonstrate that our approach gives high precision of data detection, and competitive performance compared to state-of-the-art approaches. Cu Vinh Loc, Jean-Christophe Burie, Jean-Marc Ogier, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2019 | ICDAR2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition - RRC-MLT-2019abstractWith the growing cosmopolitan culture of modern cities, the need of robust Multi-Lingual scene Text (MLT) detection and recognition systems has never been more immense. With the goal to systematically benchmark and push the state-of-the-art forward, the proposed competition builds on top of the RRC-MLT-2017 with an additional end-to-end task, an additional language in the real images dataset, a large scale multi-lingual synthetic dataset to assist the training, and a baseline End-to-End recognition method. The real dataset consists of 20,000 images containing text from 10 languages. The challenge has 4 tasks covering various aspects of multi-lingual scene text: (a) text detection, (b) cropped word script classification, (c) joint text detection and script classification and (d) end-to-end detection and recognition. In total, the competition received 60 submissions from the research and industrial communities. This paper presents the dataset, the tasks and the findings of the presented RRC-MLT-2019 challenge. Nibal Nayef, Cheng-Lin Liu 0001, Jean-Marc Ogier, Michal Busta, Pinaki Nath Chowdhury, Dimosthenis Karatzas, Wafa Khlif, Jiri Matas, Umapada Pal 0001, Jean-Christophe Burie |
ICDAR | 2 |
| 2019 | CASIA-AHCDB: A Large-Scale Chinese Ancient Handwritten Characters DatabaseabstractThis paper introduces a Chinese Ancient Handwritten Characters Database (CASIA-AHCDB) for character recognition research. The database was built by annotating 11,937 pages of Chinese ancient handwritten documents. It consists of more than 2.2 million annotated handwritten character samples of 10,350 categories. According to the source of these documents, the database is divided into two datasets of different styles: Complete Library in Four Sections (AHCDB-style1) and Ancient Buddhist Scriptures (AHCDB-style2). Each dataset can be divided into three parts based on its applications. The first part, called basic category set, contains samples of common categories in two datasets, and is suitable for basic character recognition task. The second part, called enhanced category set, is mainly used for open-set character recognition task based on the basic character recognition. The third part, called the reserved category set, can be used in many pattern recognition tasks in the future. Based on the large category set, the various writing styles and the imbalanced sample number per category, CASIA-AHCDB can also be used for various classification and learning tasks such as transfer learning, few-shot learning. We performed experiments of basic character recognition on the basic category set, and report the results for benchmark. More techniques can be evaluated on this challenging database in the future. Dahan Wang, Xu-Yao Zhang, Zhaoxiang Zhang 0001, Cheng-Lin Liu 0001 |
ICDAR | 6 |
| 2019 | Contextual Stroke Classification in Online Handwritten Documents with Graph Attention NetworksabstractClassifying strokes into different categories is an essential preprocessing step in the automatic document understanding process. To tackle this task, it is crucial to integrate different types of contextual information. Previous methods which are based on conditional random fields or recurrent neural networks have some limitations in model capacity or computational cost. In this paper, we propose a novel framework based on graph attention networks to solve this problem, which casts the stroke classification problem into the node classification problem in a document graph. In the graph, each node represents a stroke and the edges are built from temporal and spatial interactions between strokes. Combined graph convolution with attention mechanisms to dynamically aggregate features from the neighborhood, our model is very flexible to control the message passing routine between different nodes and therefore has strong capability learning context-aware features. We perform comparison experiments on the IAMonDo dataset and experimental results demonstrate the superiority of our approach. Jun-Yu Ye, Yan-Ming Zhang 0001, Qing Yang 0002, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2019 | Oracle Character Recognition by Nearest Neighbor Classification with Deep Metric LearningabstractOracle character is one kind of the earliest hieroglyphics, which can be dated back to Shang Dynasty in China. Oracle character recognition is important for modern archaeology, ancient text understanding, and historical chronology, etc. To overcome the limitation and class imbalance of training data in oracle character recognition, we propose a classification method based on deep metric learning. We use a convolutional neural network (CNN) to map the character images to an Euclidean space where the distance between different samples can measure their similarities such that classification can be performed by the Nearest Neighbor (NN) rule. Because new categories are still being discovered in reality, our model enables the rejection of unseen categories and the configuration of new categories. To accelerate NN classification, we also propose a prototype pruning method with little loss of accuracy. The proposed method exceeds the state of the art on the public dataset Oracle-20K and outperforms CNN with softmax layer on a new dataset Oracle-AYNU. Heng Zhang 0028, Yong-Ge Liu, Qing Yang 0002, Cheng-Lin Liu 0001 |
ICDAR | 5 |
| 2019 | Online Handwritten Diagram Recognition with Graph Attention Networks
Xiao-Long Yun, Yan-Ming Zhang 0001, Jun-Yu Ye, Cheng-Lin Liu 0001 |
ICIG (1) | 4 |
| 2019 | Multi-view Semi-supervised 3D Whole Brain Segmentation with a Self-ensemble Network
Yuan-Xing Zhao, Yan-Ming Zhang 0001, Cheng-Lin Liu 0001 |
MICCAI (3) | 4 |
| 2019 | Editorial for special issue on "Advanced Topics in Document Analysis and Recognition"
Cheng-Lin Liu 0001, Andreas Dengel 0001, Rafael Dueire Lins |
Int. J. Document Anal. Recognit. | 1 |
| 2019 | Special issue on advances in graph algorithm and applications
Zhiyong Liu 0001, Kaizhu Huang, Xu Yang 0004, Cheng-Lin Liu 0001 |
Neurocomputing | 4 |
| 2019 | LightweightNet: Toward fast and lightweight convolutional neural networks via architecture distillation
Ting-Bing Xu, Peipei Yang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2019 | Discriminative Feature Selection via Employing Smooth and Robust Hinge LossabstractA wide variety of sparsity-inducing feature selection methods have been developed in recent years. Most of the loss functions of these approaches are built upon regression since it is general and easy to optimize, but regression is not well suitable for classification. In contrast, the hinge loss (HL) of support vector machines has proved to be powerful to handle classification tasks, but a model with existing multiclass HL and sparsity regularization is difficult to optimize. In view of that, we propose a new loss, called smooth and robust HL, which gathers the merits of regression and HL but overcome their drawbacks, and apply it to our sparsity regularized feature selection model. To optimize the model, we present a new variant of accelerated proximal gradient (APG) algorithm, which boosts the discriminative margins among different classes, compared with standard APG algorithms. We further propose an efficient optimization technique to solve the proximal projection problem at each iteration step, which is a key component of the new APG algorithm. We theoretically prove that the new APG algorithm converges at rate O(1/k2) if it is convex (k is the iteration counter), which is the optimal convergence rate for smooth problems. Experimental results on nine publicly available data sets demonstrate the effectiveness of our method. Hanyang Peng, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Robust Classification With Convolutional Prototype LearningabstractConvolutional neural networks (CNNs) have been widely used for image classification. Despite its high accuracies, CNN has been shown to be easily fooled by some adversarial examples, indicating that CNN is not robust enough for pattern classification. In this paper, we argue that the lack of robustness for CNN is caused by the softmax layer, which is a totally discriminative model and based on the assumption of closed world (i.e., with a fixed number of categories). To improve the robustness, we propose a novel learning framework called convolutional prototype learning (CPL). The advantage of using prototypes is that it can well handle the open world recognition problem and therefore improve the robustness. Under the framework of CPL, we design multiple classification criteria to train the network. Moreover, a prototype loss (PL) is proposed as a regularization to improve the intra-class compactness of the feature representation, which can be viewed as a generative model based on the Gaussian assumption of different classes. Experiments on several datasets demonstrate that CPL can achieve comparable or even better results than traditional CNN, and from the robustness perspective, CPL shows great advantages for both the rejection and incremental category learning tasks. Hong-Ming Yang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 4 |
| 2018 | Practical Block-Wise Neural Network Architecture GenerationabstractConvolutional neural networks have gained a remarkable success in computer vision. However, most usable network architectures are hand-crafted and usually require expertise and elaborate design. In this paper, we provide a block-wise network generation pipeline called BlockQNN which automatically builds high-performance networks using the Q-Learning paradigm with epsilon-greedy exploration strategy. The optimal network block is constructed by the learning agent which is trained sequentially to choose component layers. We stack the block to construct the whole auto-generated network. To accelerate the generation process, we also propose a distributed asynchronous framework and an early stop strategy. The block-wise generation brings unique advantages: (1) it performs competitive results in comparison to the hand-crafted state-of-the-art networks on image classification, additionally, the best network generated by BlockQNN achieves 3.54% top-1 error rate on CIFAR-10 which beats all existing auto-generate networks. (2) in the meanwhile, it offers tremendous reduction of the search space in designing networks which only spends 3 days with 32 GPUs, and (3) moreover, it has strong generalizability that the network built on CIFAR also performs well on a larger-scale ImageNet dataset. Zhao Zhong, Wei Wu 0021, Cheng-Lin Liu 0001 |
CVPR | 5 |
| 2018 | Printed/Handwritten Texts and Graphics Separation in Complex Documents Using Conditional Random FieldsabstractIn this paper we propose a structured prediction based system for text/non-text classification and printed/handwritten texts separation at connected component (CC) level in complex documents. We formulate the separation of different elements as joint classification problems and use conditional random fields (CRFs) to integrate both local and contextual information for improving the classification accuracy. Both our unary and pairwise potentials are formulated as neural networks for better exploiting contextual information. Considering the different properties in text/non-text classification and printed/handwritten texts separation, we use multilayer perception (MLP) and convolutional neural network (CNN) for potentials, respectively. To evaluate the performance of the proposed method, we provide a test paper document database named TestPaper1.0, which can be used for many other tasks as well. Our method achieve impressive results for both tasks on TestPaper1.0 dataset. Moreover, even with very shallow CNNs as potentials, our method achieves state-of-the-art performance for writing type (printed/handwritten) separation on the highly heterogeneous Maurdor dataset, surpassing Maurdor2013 and Maurdor2014 campaign winners. This demonstrates the effectiveness and superiority of our method. Xiao-Hui Li 0012, Cheng-Lin Liu 0001 |
DAS | 3 |
| 2018 | Online Video Text Detection with Markov Decision ProcessabstractOnline video text detection is important in many applications, such as real-time translator and wearable camera system for visually-impaired. Existing methods for video text detection perform unsatisfactorily mainly because of the inferior text detection result and insufficient utilization of spatial and temporal information. Besides, the majority of them work in offline mode. In this paper, we propose an online video text detection method which works nearly in real time. We detect texts in each frame using a EAST based text detector, and formulate the online text tracking problem as decision making in Markov Decision Processes (MDPs). The similarity function in tracking stage can be learned by reinforcement learning. Besides, text detection and tracking are naturally unified by state transactions in the MDP. Extensive experiments on three benchmark datasets, ICDAR 2015, Minetto, and Youtube Video Text, verify the effectiveness of our method. Xue-Hang Yang, Cheng-Lin Liu 0001 |
DAS | 3 |
| 2018 | Memory-Augmented Attention Model for Scene Text RecognitionabstractNatural scene text recognition is a very challenging task. Attention-based encoder-decoder framework has achieved the state-of-the-art performance. However, for some complex and/or low-quality images, the alignments estimated by the content-based attention network are not accurate enough, and so, the generated glimpse vector is also not powerful enough to represent the predicted character at current time step. To solve this problem, in the paper we propose a memory-augmented attention model for scene text recognition. The proposed memory-augmented attention network (MAAN) feeds the part of character sequence already generated and all attended alignment history to the attention model when predicting the character at current time step. The whole network can be trained end-to-end. Experimental results on several challenging benchmark datasets demonstrate that the proposed memory-augmented attention model for scene text recognition can achieve a comparable or better performance compared with state-of-the-art methods. Cheng-Lin Liu 0001 |
ICFHR | 3 |
| 2018 | Deep Transfer Mapping for Unsupervised Writer AdaptationabstractConvolutional neural network (CNN) has achieved great success in handwriting recognition. However, it relies on large set of labeled data in training and its performance will deteriorate when the data distribution varies. To solve this problem, traditional methods usually consider adaptation of the single top layer of CNN. To better reduce the distribution discrepancy, in this paper, we consider adaptation of all layers of CNN including both convolutional and full layers. Four variations of transformations are designed based on different assumptions about the space relations for adaptation of convolutional layers. In order to make adaptation of multiple layers, we propose to cascade the transformations of different layers to conduct adaptation in a deep manner, and therefore this method is denoted as deep transfer mapping (DTM). DTM can capture the information from different layers and minimize the data divergence under different information abstract levels, thus it is more powerful and flexible for domain adaptation. Experiments on the online Chinese handwriting dataset (OLHWDB) demonstrate the efficiency and effectiveness of the proposed method for unsupervised writer adaptation. Hong-Ming Yang, Xu-Yao Zhang, Jun Sun 0004, Cheng-Lin Liu 0001 |
ICFHR | 5 |
| 2018 | Scene Text Detection with Recurrent Instance SegmentationabstractConvolutional Neural Network (CNN) based scene text detection methods mostly employ the semantic segmentation (text/non-text classification) task to localize the regions of texts. However, they cannot distinguish different text-lines like instance segmentation. In this paper, we propose a novel framework based on Fully Convolutional Networks (FCN) and Recurrent Neural Network (RNN) to achieve both scene text detection and instance segmentation. The FCN is used to classify text and non-text regions, and the RNN utilizes the features extracted by FCN to simultaneously detect and segment one text instance at each time step. Meanwhile, it also extracts bounding boxes by a much simpler way than the non-maximum suppression (NMS) method. The proposed method achieves competitive results on two public benchmarks including ICDAR 2015 Incidental Scene Text Dataset and ICDAR 2013 Focused Scene Text Dataset. Moreover, the benefits of adding regression task in the RNN module are manifested. Wei Feng 0016, Cheng-Lin Liu 0001 |
ICPR | 4 |
| 2018 | Page Object Detection from PDF Document Images by Deep Structured Prediction and Supervised ClusteringabstractPage object detection in document images remains a challenge because the page objects are diverse in scale and aspect ratio, and an object may contain largely apart components. In this paper, we propose a hybrid method combining deep structured prediction and supervised clustering to detect formulas, tables and figures in PDF document images within a unified framework. The primitive region proposals extracted from each column region are classified and clustered with conditional random field (CRF) based graphical models which can integrate both local and contextual information. Both the unary and pairwise potentials of CRFs are formulated as convolutional neural networks (CNNs) to better exploit spatial contextual information. The CRF for clustering predicts the linked/cut label of between-region links. After CRF inference, the line regions of same class within a cluster are grouped into a page object. The state-of-the-art performance obtained on the public available ICDAR2017 POD competition dataset demonstrates the effectiveness and superiority of the nronosed method. Xiao-Hui Li 0012, Cheng-Lin Liu 0001 |
ICPR | 3 |
| 2018 | Anomaly Detection via Minimum Likelihood Generative Adversarial NetworksabstractAnomaly detection aims to detect abnormal events by a model of normality. It plays an important role in many domains such as network intrusion detection, criminal activity identity and so on. With the rapidly growing size of accessible training data and high computation capacities, deep learning based anomaly detection has become more and more popular. In this paper, a new domain-based anomaly detection method based on generative adversarial networks (GAN) is proposed. Minimum likelihood regularization is proposed to make the generator produce more anomalies and prevent it from converging to normal data distribution. Proper ensemble of anomaly scores is shown to improve the stability of discriminator effectively. The proposed method has achieved significant improvement than other anomaly detection methods on Cifar10 and UCI datasets. Yan-Ming Zhang 0001, Cheng-Lin Liu 0001 |
ICPR | 3 |
| 2018 | Multi-task Layout Analysis for Historical Handwritten Documents Using Fully Convolutional NetworksabstractLayout analysis is a fundamental process in document image analysis and understanding. It consists of several sub-processes such as page segmentation, text line segmentation, baseline detection and so on. In this work, we propose a multi-task layout analysis method that use a single FCN model to solve the above three problems simultaneously. The FCN is trained to segment the document image into different regions and detect the center line of each text line by classifying pixels into different categories. By supervised learning on document images with pixel-wise labels, the FCN can extract discriminative features and perform pixel-wise classification accurately. After pixel-wise classification, post-processing steps are taken to reduce noises, correct wrong segmentations and find out overlapping regions. Experimental results on the public dataset DIVA-HisDB containing challenging medieval manuscripts demonstrate the effectiveness and superiority of the proposed method. Zhaoxiang Zhang 0001, Cheng-Lin Liu 0001 |
IJCAI | 4 |
| 2018 | Image-to-Markup Generation via Paired Adversarial Learning
Jin-Wen Wu, Yan-Ming Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ECML/PKDD (1) | 5 |
| 2018 | Special issue on deep learning for document analysis and recognition
Cheng-Lin Liu 0001, Gernot A. Fink, Venu Govindaraju |
Int. J. Document Anal. Recognit. | 1 |
| 2018 | Drawing and Recognizing Chinese Characters with Recurrent Neural NetworkabstractRecent deep learning based approaches have achieved great success on handwriting recognition. Chinese characters are among the most widely adopted writing systems in the world. Previous research has mainly focused on recognizing handwritten Chinese characters. However, recognition is only one aspect for understanding a language, another challenging and interesting task is to teach a machine to automatically write (pictographic) Chinese characters. In this paper, we propose a framework by using the recurrent neural network (RNN) as both a discriminative model for recognizing Chinese characters and a generative model for drawing (generating) Chinese characters. To recognize Chinese characters, previous methods usually adopt the convolutional neural network (CNN) models which require transforming the online handwriting trajectory into image-like representations. Instead, our RNN based approach is an end-to-end system which directly deals with the sequential structure and does not require any domain-specific knowledge. With the RNN system (combining an LSTM and GRU), state-of-the-art performance can be achieved on the ICDAR-2013 competition database. Furthermore, under the RNN framework, a conditional generative model with character embedding is proposed for automatically drawing recognizable Chinese characters. The generated characters (in vector format) are human-readable and also can be recognized by the discriminative RNN model with high accuracy. Experimental results verify the effectiveness of using RNNs as both generative and discriminative models for the tasks of drawing and recognizing Chinese characters. Xu-Yao Zhang, Yan-Ming Zhang 0001, Cheng-Lin Liu 0001, Yoshua Bengio |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Multi-Oriented and Multi-Lingual Scene Text Detection With Direct RegressionabstractMulti-oriented and multi-lingual scene text detection plays an important role in computer vision area and is challenging due to the wide variety of text and background. In this paper, firstly we point out the two key tasks when extending CNN based object detection frameworks to scene text detection. The first task is to localize the text region by a down-sampled segmentation based module, and the second task is to regress the boundaries of text region determined by the first task. Secondly, we propose a scene text detection framework based on fully convolutional network (FCN) with a bi-task prediction module in which one is pixel-wise classification between text and non-text, and the other is pixel-wise regression to determine the vertex coordinates of quadrilateral text boundaries. Post-processing for word-level detection is based on Non-Maximum Suppression (NMS), and for line-level detection we design a heuristic line segments grouping method to localize long text lines. We evaluated the proposed framework on various benchmarks including multi-oriented and multi-lingual scene text datasets, and achieved state-of-the-art performance on most of them. We also provide abundant ablation experiments to analyze several key factors in building high performance CNN based scene text detection systems. Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Deep Direct Regression for Multi-oriented Scene Text DetectionabstractIn this paper, we first provide a new perspective to divide existing high performance object detection methods into direct and indirect regressions. Direct regression performs boundary regression by predicting the offsets from a given point, while indirect regression predicts the offsets from some bounding box proposals. In the context of multioriented scene text detection, we analyze the drawbacks of indirect regression, which covers the state-of-the-art detection structures Faster-RCNN and SSD as instances, and point out the potential superiority of direct regression. To verify this point of view, we propose a deep direct regression based method for multi-oriented scene text detection. Our detection framework is simple and effective with a fully convolutional network and one-step post processing. The fully convolutional network is optimized in an end-to-end way and has bi-task outputs where one is pixel-wise classification between text and non-text, and the other is direct regression to determine the vertex coordinates of quadrilateral text boundaries. The proposed method is particularly beneficial to localize incidental scene texts. On the ICDAR2015 Incidental Scene Text benchmark, our method achieves the F-measure of 81%, which is a new state-ofthe-art and significantly outperforms previous approaches. On other standard datasets with focused scene texts, our method also reaches the state-of-the-art performance. Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICCV | 4 |
| 2017 | Simultaneous Script Identification and Handwriting Recognition via Multi-Task Learning of Recurrent Neural NetworksabstractIn this paper, we propose a method for simultaneous script identification and handwritten text line recognition in multi-task learning framework. Firstly, we use Separable Multi-Dimensional Long Short-Term Memory (SepMDLSTM) to encode the input text line images based on convolutional feature extraction. Then, the extracted features are fed into two classification modules for script identification and multi-script text recognition, respectively. All the network parameters are trained end-to-end by multi-task learning where the script identification task and the text recognition task are aimed to minimize the Negative Log Likelihood (NLL) loss and Connectionist Temporal Classification (CTC) loss, respectively. We evaluated the performance of the proposed method on handwritten text line datasets of three languages, namely, IAM (English), Rimes (French) and IFN/ENIT (Arabic). Experimental results demonstrate the multi-task learning framework performs superiorly for both script identification and text recognition. Particularly, the accuracy of script identification is higher than 99.9% and the character error rate (CER) of text recognition is even lower than that of some single-script text recognition systems. Zhuo Chen 0051, Yichao Wu, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2017 | ICDAR2017 Robust Reading Challenge on Multi-Lingual Scene Text Detection and Script Identification - RRC-MLTabstractText detection and recognition in a natural environment are key components of many applications, ranging from business card digitization to shop indexation in a street. This competition aims at assessing the ability of state-of-the-art methods to detect Multi-Lingual Text (MLT) in scene images, such as in contents gathered from the Internet media and in modern cities where multiple cultures live and communicate together. This competition is an extension of the Robust Reading Competition (RRC) which has been held since 2003 both in ICDAR and in an online context. The proposed competition is presented as a new challenge of the RRC. The dataset built for this challenge largely extends the previous RRC editions in many aspects: the multi-lingual text, the size of the dataset, the multi-oriented text, the wide variety of scenes. The dataset is comprised of 18,000 images which contain text belonging to 9 languages. The challenge is comprised of three tasks related to text detection and script classification. We have received a total of 16 participations from the research and industrial communities. This paper presents the dataset, the tasks and the findings of this RRC-MLT challenge. Nibal Nayef, Imen Bizid, Hyunsoo Choi, Dimosthenis Karatzas, Zhenbo Luo, Umapada Pal 0001, Christophe Rigaud, Joseph Chazalon, Wafa Khlif, Muhammad Muzzamil Luqman, Jean-Christophe Burie, Cheng-Lin Liu 0001, Jean-Marc Ogier |
ICDAR | 14 |
| 2017 | Radical-Based Chinese Character Recognition via Multi-Labeled Learning of Deep Residual NetworksabstractThe digitization of Chinese historical documents poses a new challenge that in the huge set of character categories, majority of characters are not in common use now and have few samples for training the character classifiers. To settle this problem, we consider the radical-level composition of Chinese characters, and propose to detect position-dependent radicals using a deep residual network with multi-labeled learning. This enables the recognition of novel characters without training samples if the characters are composed of radicals appearing in training samples. In multi-labeled learning, each training character sample is labeled as positive for each radical it contains, such that after training, all the radicals appearing in the character can be detected. Experimental results on a large-category-set database of printed Chinese characters demonstrate that the proposed method can detect radicals accurately. Moreover, according to radical configurations, our model can credibly recognize novel characters as well as trained characters. Tie-Qiang Wang, Cheng-Lin Liu 0001 |
ICDAR | 3 |
| 2017 | Scene Text Detection with Novel Superpixel Based Character Candidate ExtractionabstractMaximally stable extremal region (MSER) is popularly used for candidate character candidate extraction in scene text detection. Its requirement of maximum stability hinders high performance on images of high variability. In this paper, we propose a novel character candidate extraction method based on superpixel segmentation and hierarchical clustering. The proposed superpixel segmentation algorithm for scene text image takes advantage of the color consistency of characters and fuses color and edge information. Based on superpixel segmentation, character candidates are extracted by single-link clustering. To improve the accuracy of non-text candidate filtering, we use a deep convolutional neural networks (DCNN) classifier and double threshold strategy for classification. Experimental results on public datasets demonstrate that the proposed superpixel based method performs better than MSER in character candidate extraction, and the proposed system achieves competitive performance compared to state-of-the-art methods. Cheng-Lin Liu 0001 |
ICDAR | 3 |
| 2017 | Handwritten Chinese Text Recognition Using Separable Multi-Dimensional Recurrent Neural NetworkabstractThe Long Short-Term Memory Recurrent Neural Network (LSTM-RNN) has been demonstrated successful in handwritten text recognition of Western and Arabic scripts. It is totally segmentation free and can be trained directly from text line images. However, the application of LSTM-RNNs (including Multi-Dimensional LSTM-RNN (MDLSTM-RNN)) to Chinese text recognition has shown limited success, even when training them with large datasets and using pre-training on datasets of other languages. In this paper, we propose a handwritten Chinese text recognition method by using Separable MDLSTMRNN (SMDLSTM-RNN) modules, which extract contextual information in various directions, and consume much less computation efforts and resources compared with the traditional MDLSTMRNN. Experimental results on the ICDAR-2013 competition dataset show that the proposed method performs significantly better than the previous LSTM-based methods, and can compete with the state-of-the-art systems. Yi-Chao Wu, Zhuo Chen 0051, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2017 | Page Segmentation for Historical Handwritten Documents Using Fully Convolutional NetworksabstractPage segmentation is a fundamental and challenging task in document image analysis due to the layout diversity. In this work, we propose a pixel-wise segmentation method for historical handwritten documents using fully convolutional network (FCN). The document image is segmented into different regions by classifying pixels into different categories: background, main text body, comments, and decorations. By supervised learning on document images with pixel-wise labels, the FCN can extract discriminative features and perform pixel-wise segmentation accurately. After pixel-wise classification, post-processing steps are taken to reduce noises, correct wrong segmentations and find out overlapping regions. Experimental results on the public dataset DIVA-HisDB containing challenging medieval manuscripts demonstrate the effectiveness and superiority of the proposed method, which yields pixel-level accuracy of above 99%. Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2017 | A Unified Video Text Detection Method with Network FlowabstractScene text detection in videos has many application needs but has drawn less attention than that in images. Existing methods for video text detection perform unsatisfactorily because of the insufficient utilization of spatial and temporal information. In this paper, we propose a novel video text detection method with network flow based tracking. The system first applies a newly proposed Fully Convolutional Neural Network (FCN) based scene text detection method to detect texts in individual frames and then track proposals in adjacent frames with a motion-based method. Next, the text association problem is formulated into a cost-flow network and text trajectories are derived from the network with a min-cost flow algorithm. At last, the trajectories are post-processed to improve the precision accuracy. The method can detect multi-oriented scene text in videos and incorporate spatial and temporal information efficiently. Experimental results show that the method improves the detection performance remarkably on benchmark datasets, e.g., by a 15.66% increase of ATA Average Tracking Accuracy) on ICDAR video scene text dataset. Xue-Hang Yang, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2017 | Handwriting Style Mixture AdaptationabstractIn handwriting recognition, the test data usually come from multiple writers which are not shown in the training data. Therefore, adapting the base classifier towards the new style of each writer can significantly improve the generalization performance. Traditional writer adaptation methods usually assume that there is only one writer (one style) in the test data, and we call this situation as style-clear adaptation. However, a more common situation is that multiple handwriting styles exist in the test data, which is widely appeared in multi-font documents and handwriting data produced by the cooperation of multiple writers. We call the adaptation in this situation as style-mixture adaptation. To deal with this problem, in this paper, we propose a novel method called K-style mixture adaptation (K-SMA) with the assumption that there are totally K styles in the test data. Specifically, we first partition the test data into K groups (style clustering) according to their style consistency, which is measured by a newly designed style feature that can eliminate class (category) information and keep handwriting style information. After that, in each group, a style transfer mapping (STM) is used for writer adaptation. Since the initial style clustering may be not reliable, we repeat this process iteratively to improve the adaptation performance. The K-SMA model is fully unsupervised which do not require either the class label or the style index. Moreover, the K-SMA model can be effectively combined with the benchmark convolutional neural network (CNN) models. Experiments on the online Chinese handwriting database CASIA-OLHWDB demonstrate that K-SMA is an efficient and effective solution for style-mixture adaptation. Hong-Ming Yang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2017 | Chinese Handwriting Generation by Neural Network Based Style Transformation
Bi-Ren Tan, Yi-Chao Wu, Cheng-Lin Liu 0001 |
ICIG (1) | 4 |
| 2017 | Margin-Aware Binarized Weight Networks for Image Classification
Ting-Bing Xu, Peipei Yang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICIG (1) | 4 |
| 2017 | Diverse Neuron Type Selection for Convolutional Neural NetworksabstractThe activation function for neurons is a prominent element in the deep learning architecture for obtaining high performance. Inspired by neuroscience findings, we introduce and define two types of neurons with different activation functions for artificial neural networks: excitatory and inhibitory neurons, which can be adaptively selected by self-learning. Based on the definition of neurons, in the paper we not only unify the mainstream activation functions, but also discuss the complementariness among these types of neurons. In addition, through the cooperation of excitatory and inhibitory neurons, we present a compositional activation function that leads to new state-of-the-art performance comparing to rectifier linear units. Finally, we hope that our framework not only gives a basic unified framework of the existing activation neurons to provide guidance for future design, but also contributes neurobiological explanations which can be treated as a window to bridge the gap between biology and computer science. Guibo Zhu, Zhaoxiang Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IJCAI | 4 |
| 2017 | Keyword spotting in handwritten chinese documents using semi-markov conditional random fields
Heng Zhang 0028, Cheng-Lin Liu 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2017 | SDE: A Novel Selective, Discriminative and Equalizing Feature Representation for Visual Recognition
Guosen Xie, Xu-Yao Zhang, Shuicheng Yan, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 4 |
| 2017 | Building Regional Covariance Descriptors for Vehicle DetectionabstractWe study the question of building regional covariance descriptors (RCDs) for vehicle detection from high-resolution satellite images. A unified way is proposed to build RCD features by constant convolutional kernels in the forms of 2-D masks. Two novel formulas are designed to construct different RCD types based upon one or two convolutional masks, obtaining ten novel RCD features by four simple constant convolutional masks. Experiments show that such convolutional-mask-based RCDs outperform the previous image-derivative-based RCDs, the popular local binary patterns (LBPs), the histogram of oriented gradients (HOGs), and LBP+HOG. Furthermore, feeding to nonlinear support vector machines (SVMs) of two kernel types [L1kernel and radial basis function (RBF)], these RCDs outperform four known deep convolutional neural networks: AlexNet, GoogLeNet, CaffeNet, and LeNet, as well as their fine-tuned models by their well-trained weights of imageNet classification. Among three popular classic classifiers we have tested in the experiments, nonlinear SVMs outperform BP and Adaboost obviously, and L1kernel exceeds RBF slightly. Xueyun Chen, Ren-Xi Gong, Ling-Ling Xie, Shiming Xiang, Cheng-Lin Liu 0001, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2017 | Improving handwritten Chinese text recognition using neural network language models and convolutional neural network shape models
Yi-Chao Wu, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2017 | LG-CNN: From local parts to global discrimination for fine-grained recognition
Guosen Xie, Xu-Yao Zhang, Wenhan Yang, Mingliang Xu 0001, Shuicheng Yan, Cheng-Lin Liu 0001 |
Pattern Recognit. | 6 |
| 2017 | Online and offline handwritten Chinese character recognition: A comprehensive study and new benchmark
Xu-Yao Zhang, Yoshua Bengio, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2017 | Special issue "Advances in graph-based pattern recognition"
Cheng-Lin Liu 0001, Bin Luo 0001, Walter G. Kropatsch |
Pattern Recognit. Lett. | 1 |
| 2017 | Hybrid CNN and Dictionary-Based Models for Scene Recognition and Domain AdaptationabstractConvolutional neural network (CNN) has achieved the state-of-the-art performance in many different visual tasks. Learned from a large-scale training data set, CNN features are much more discriminative and accurate than the handcrafted features. Moreover, CNN features are also transferable among different domains. On the other hand, traditional dictionary-based features (such as BoW and spatial pyramid matching) contain much more local discriminative and structural information, which is implicitly embedded in the images. To further improve the performance, in this paper, we propose to combine CNN with dictionary-based models for scene recognition and visual domain adaptation (DA). Specifically, based on the well-tuned CNN models (e.g., AlexNet and VGG Net), two dictionary-based representations are further constructed, namely, mid-level local representation (MLR) and convolutional Fisher vector (CFV) representation. In MLR, an efficient two-stage clustering method, i.e., weighted spatial and feature space spectral clustering on the parts of a single image followed by clustering all representative parts of all images, is used to generate a class-mixture or a class-specific part dictionary. After that, the part dictionary is used to operate with the multiscale image inputs for generating mid-level representation. In CFV, a multiscale and scale-proportional Gaussian mixture model training strategy is utilized to generate Fisher vectors based on the last convolutional layer of CNN. By integrating the complementary information of MLR, CFV, and the CNN features of the fully connected layer, the state-of-the-art performance can be achieved on scene recognition and DA problems. An interested finding is that our proposed hybrid representation (from VGG net trained on ImageNet) is also complementary to GoogLeNet and/or VGG-11 (trained on Place205) greatly. Guosen Xie, Xu-Yao Zhang, Shuicheng Yan, Cheng-Lin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | End-to-End Online Writer Identification With Recurrent Neural NetworkabstractWriter identification is an important topic for pattern recognition and artificial intelligence. Traditional methods rely heavily on sophisticated hand-crafted features to represent the characteristics of different writers. In this paper, we propose an end-to-end framework for online text-independent writer identification by using a recurrent neural network (RNN). Specifically, the handwriting data of a particular writer are represented by a set of random hybrid strokes (RHSs). Each RHS is a randomly sampled short sequence representing pen tip movements ($xy$-coordinates) and pen-down or pen-up states. RHS is independent of the content and language involved in handwriting; therefore, writer identification at the RHS level is more general and convenient than the character level or the word level, which also requires character/word segmentation. The RNN model with bidirectional long short-term memory is used to encode each RHS into a fixed-length vector for final classification. All the RHSs of a writer are classified independently, and then, the posterior probabilities are averaged to make the final decision. The proposed framework is end-to-end and does not require any domain knowledge for handwriting data analysis. Experiments on both English (133 writers) and Chinese (186 writers) databases verify the advantages of our method compared with other state-of-the-art approaches. Xu-Yao Zhang, Guosen Xie, Cheng-Lin Liu 0001, Yoshua Bengio |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2017 | Traffic Sign Detection Using a Cascade Method With Fast Feature Extraction and Saliency TestabstractAutomatic traffic sign detection is challenging due to the complexity of scene images, and fast detection is required in real applications such as driver assistance systems. In this paper, we propose a fast traffic sign detection method based on a cascade method with saliency test and neighboring scale awareness. In the cascade method, feature maps of several channels are extracted efficiently using approximation techniques. Sliding windows are pruned hierarchically using coarse-to-fine classifiers and the correlation between neighboring scales. The cascade system has only one free parameter, while the multiple thresholds are selected by a data-driven approach. To further increase speed, we also use a novel saliency test based on mid-level features to pre-prune background windows. Experiments on two public traffic sign data sets show that the proposed method achieves competing performance and runs 2~7 times as fast as most of the state-of-the-art methods. Xinwen Hou, Jiawei Xu 0004, Shigang Yue, Cheng-Lin Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2016 | Large-Scale Graph-Based Semi-Supervised Learning via Tree Laplacian SolverabstractGraph-based Semi-Supervised learning is one of the most popular and successful semi-supervised learning methods. Typically, it predicts the labels of unlabeled data by minimizing a quadratic objective induced by the graph, which is unfortunately a procedure of polynomial complexity in the sample size $n$. In this paper, we address this scalability issue by proposing a method that approximately solves the quadratic objective in nearly linear time. The method consists of two steps: it first approximates a graph by a minimum spanning tree, and then solves the tree-induced quadratic objective function in O(n) time which is the main contribution of this work. Extensive experiments show the significant scalability improvement over existing scalable semi-supervised learning methods. Yan-Ming Zhang 0001, Xu-Yao Zhang, Xiao-Tong Yuan, Cheng-Lin Liu 0001 |
AAAI | 4 |
| 2016 | Effective Candidate Component Extraction for Text Localization in Born-Digital Images by Combining Text Contours and Stroke Interior RegionsabstractExtracting candidate text connected components (CCs) is critical for CC-based text localization. Based on the observation that text strokes in born-digital images mostly have complete contours and the text pixels have high contrast with the adjacent non-text pixels, we propose a method to extract candidate text CCs by combining text contours and stroke interior regions. After segmenting the image into non-smooth and smooth regions based on local contrast, text contour pixels in non-smooth regions are detached from adjacent non-text pixels by local binarization. Then, obvious non-text contours can be removed according to the spatial relationship of text and non-text contours. While smooth regions include stroke interior regions and non-text smooth regions, some non-text smooth regions can be easily removed because they are not surrounded by candidate text contours. At last, candidate text contours and stroke interior regions are combined to generate candidate text CCs. The CCs undergo CC filtering, text line grouping and line classification to give the text localization result. Experimental results on the born-digital dataset of ICDAR2013 robust reading competition demonstrate the efficiency and superiority of the proposed method. Cheng-Lin Liu 0001 |
DAS | 3 |
| 2016 | Natural Scene Character Recognition Using Robust PCA and Sparse RepresentationabstractNatural scene character recognition is challenging due to the cluttered background, which is hard to separate from text. In this paper, we propose a novel method for robust scene character recognition. Specifically, we first use robust principal component analysis (PCA) to denoise character image by recovering the missing low-rank component and filtering out the sparse noise term, and then use a simple Histogram of oriented Gradient (HOG) to perform image feature extraction, and finally, use a sparse representation based classifier for recognition. In experiments on four public datasets, namely the Char74K dataset, ICADAR 2003 robust reading dataset, Street View Text (SVT) dataset and IIIT5K-word dataset, our method was demonstrated to be competitive with the state-of-the-art methods. Zheng Zhang 0006, Yong Xu 0001, Cheng-Lin Liu 0001 |
DAS | 3 |
| 2016 | Unsupervised Adaptation of Neural Networks for Chinese Handwriting RecognitionabstractWriter adaptation is an important topic in handwriting recognition, which can further improve the performance of writer-independent recognizer. In this paper, we propose combining the neural network classifier with style transfer mapping (STM) for unsupervised writer adaptation, which only require writer-specific unlabeled data, and therefore is more common and efficient compared to supervised adaptation. We use some techniques like dropout, ReLU, momentum, and deeply supervised strategy to improve the performance of the neural network classifier. For a specific writer in the test data, an adaptation layer is added to the pre-trained neural network classifier. In adaptation process, only the parameters in adaptation layer are updated while other parameters of the neural network are kept unchanged. To train the adaptation layer, we use the same technology as STM learning but redefine the source point set, target point set and the corresponding confidence. Experiments on the online Chinese handwriting database CASIA-OLHWDB1.1 demonstrate that our method is very efficient and effective in improving classification accuracy. The experimental results also show that our proposed method outperforms the previous proposed learning vector quantization (LVQ) and modified quadratic discriminant function (MQDF) with STM methods for writer adaptation. Hong-Ming Yang, Xu-Yao Zhang, Zhenbo Luo, Cheng-Lin Liu 0001 |
ICFHR | 5 |
| 2016 | Exploiting coarse-to-fine mechanism for fine-grained recognitionabstractFine-grained object recognition is more challenging than generic categorization due to the subtle difference between subcategories under the large intra-class pose change and appearance variations. The state-of-the-art fine-grained recognition methods usually utilize part detection or pose alignment to alleviate the pose variation, and then use convolutional neural networks (CNNs) to extract local discriminative features. Although the hierarchical structure of deep CNNs enables rich and discriminative visual feature extraction, the recognition methods so far mostly use the features of only the last convolutional layer for classification. In this paper, by exploiting the correlation of the convolutional features of within-layer and between-layer, we propose a method to integrate multi-layer convolutional features based on coarse-to-fine mechanism for improving the discrimination capability. Experiments on a number of public datasets show that the proposed method, without part annotation or pose alignment, yields superior or comparable performance to the state-of-the-art methods. Yongzhong Wang, Xu-Yao Zhang, Yan-Ming Zhang 0001, Xinwen Hou, Cheng-Lin Liu 0001 |
ICIP | 5 |
| 2016 | Context-aware mathematical expression recognition: An end-to-end framework and a benchmarkabstractIn this paper we propose a novel end-to-end framework for mathematical expression (ME) recognition. The method uses a convolutional neural network (CNN) to perform mathematical symbol detection and recognition simultaneously incorporating spatial context, and can handle multi-part and touching symbols effectively. To evaluate the performance, we provide a benchmark that contains MEs both from real-life and synthetic data. Images in our dataset undergo multiple variations such as viewpoint, illumination and background. For training, we use pure synthetic data for saving human labeling effort. The proposed method achieved 87% accuracy of total correct for clear images and 45% for cluttered ones. Yuxuan Luo 0002, Junyu Han, Errui Ding, Cheng-Lin Liu 0001 |
ICPR | 7 |
| 2016 | Joint training of conditional random fields and neural networks for stroke classification in online handwritten documentsabstractThe task of text/non-text stroke classification in online handwritten documents is an essential preprocessing step in document analysis. It is also a challenging problem since in many cases local features are not enough to generate high accuracy results and contextual information, such as temporal information and spatial information, must be carefully considered. In this paper, we propose a novel method, which jointly trains a combined model of conditional random fields and neural networks, to solve this problem. Both our unary and pairwise potentials are formulated as neural networks. The parameters of conditional random fields and neural networks are learned together during the training process. With much fewer parameters and faster speed, our method achieves impressive performance on the IAMonDo database, a publicly available database of freely handwritten documents. Jun-Yu Ye, Yan-Ming Zhang 0001, Cheng-Lin Liu 0001 |
ICPR | 3 |
| 2016 | Handwritten Chinese character recognition with spatial transformer and deep residual networksabstractThis paper considers using deep neural networks for handwritten Chinese character recognition (HCCR) with arbitrary position, scale, and orientations. To solve this problem, we combine the recently proposed spatial transformer network (STN) with the deep residual network (DRN). The STN acts like a character shape normalization procedure. Different from the traditional heuristic shape normalization methods, STN is learned directly from the data. Furthermore, the DRN makes the training of very deep network to be both efficient and effective. With the combination of STN and DRN, the whole model can be trained jointly in an end-to-end manner. In this paper, new state-of-the-art performance has been achieved by our proposed model on the offline ICDAR-2013 Chinese handwriting competition database. Moreover, the experiment on randomly distorted samples shows that the STN is very effective for robust HCCR in rectifying the shape of distorted characters. Zhao Zhong, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICPR | 4 |
| 2016 | Adaptive spatial pooling for image classification
Yinglu Liu, Yan-Ming Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2016 | A fast projected fixed-point algorithm for large graph matching
Kaizhu Huang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2016 | Discriminative quadratic feature learning for handwritten Chinese character recognition
Ming-Ke Zhou, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 4 |
| 2016 | Text Detection, Tracking and Recognition in Video: A Comprehensive SurveyabstractThe intelligent analysis of video data is currently in wide demand because a video is a major source of sensory data in our lives. Text is a prominent and direct source of information in video, while the recent surveys of text detection and recognition in imagery focus mainly on text extraction from scene images. Here, this paper presents a comprehensive survey of text detection, tracking, and recognition in video with three major contributions. First, a generic framework is proposed for video text extraction that uniformly describes detection, tracking, recognition, and their relations and interactions. Second, within this framework, a variety of methods, systems, and evaluation protocols of video text extraction are summarized, compared, and analyzed. Existing text tracking techniques, tracking-based detection and recognition techniques are specifically highlighted. Third, related applications, prominent challenges, and future directions for video text extraction (especially from scene videos and web videos) are also thoroughly discussed. Xu-Cheng Yin, Ze-Yu Zuo, Shu Tian, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2015 | Task-Driven Feature Pooling for Image ClassificationabstractFeature pooling is an important strategy to achieve high performance in image classification. However, most pooling methods are unsupervised and heuristic. In this paper, we propose a novel task-driven pooling (TDP) model to directly learn the pooled representation from data in a discriminative manner. Different from the traditional methods (e.g., average and max pooling), TDP is an implicit pooling method which elegantly integrates the learning of representations into the given classification task. The optimization of TDP can equalize the similarities between the descriptors and the learned representation, and maximize the classification accuracy. TDP can be combined with the traditional BoW models (coding vectors) or the recent state-of-the-art CNN models (feature maps) to achieve a much better pooled representation. Furthermore, a self-training mechanism is used to generate the TDP representation for a new test image. A multi-task extension of TDP is also proposed to further improve the performance. Experiments on three databases (Flower-17, Indoor-67 and Caltech-101) well validate the effectiveness of our models. Guosen Xie, Xu-Yao Zhang, Xiangbo Shu, Shuicheng Yan, Cheng-Lin Liu 0001 |
ICCV | 5 |
| 2015 | Efficient text localization in born-digital images by local contrast-based segmentationabstractText localization in born-digital images is usually performed using methods designed for scene text images. Based on the observation that text strokes in born-digital images mostly have complete contours and the pixels on the contours have high contrast compared with the adjacent non-text pixels, we propose a method to extract candidate text components using local contrast. First, the image is segmented into smooth and non-smooth regions. After removing non-text smooth regions, the remaining smooth regions are merged with non-smooth regions to form a candidate text image, which is binarized into high-value and low-value connected components (CCs). The CCs undergo CC filtering, line grouping and line classification to give the text localization result. Experimental results on the born-digital dataset of ICDAR2013 robust reading competition demonstrate the efficiency and superiority of the proposed method. Amir Hussain 0001, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2015 | Evaluation of neural network language models in handwritten Chinese text recognitionabstractHandwritten Chinese text recognition based on over-segmentation and path search integrating contexts has been demonstrated successful, where language models play an important role. Recently, neural network language models (NNLMs) have shown superiority to back-off N-gram language models (BLMs) in handwriting recognition, but have not been studied in Chinese text recognition system. This paper investigates the effects of NNLMs in handwritten Chinese text recognition and compares the performance with BLMs. We trained character-level language models in 3-, 4- and 5- gram on large scale corpora and applied them in text line recognition system. Experimental results on the CASIA-HWDB database show that NNLM and BLM of the same order perform comparably, and the hybrid model by interpolating NNLM and BLM improves the recognition performance significantly. Yi-Chao Wu, Cheng-Lin Liu 0001 |
ICDAR | 3 |
| 2015 | Lexicon-driven recognition of one-stroke character strings in visual gestureabstractVisual gesture recognition enables natural human-machine interaction, and writing characters in gesture can convey rich information of intention. However, the recognition of character strings in gesture is challenging because multiple characters are in a single-stroke trajectory without pen lift information. We propose a lexicon-driven approach for gesture character string recognition. Using a lexicon of words to guide character segmentation and recognition, and meanwhile combining the geometric scores of characters and redundant segments with character classification score, we can achieve fairly high recognition accuracy on one-stroke character strings. For experiments, we collected 1,590 gesture strings in 100 word classes of television channel names, and achieved string-level recognition accuracy over 80% on the test set. Pai pai Liu, Linlin Huang 0001, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2015 | A saliency-based cascade method for fast traffic sign detectionabstractWe propose a cascade method for fast and accurate traffic sign detection. The main feature of the method is that mid-level saliency test is used to efficiently and reliably eliminate background windows. Fast feature extraction is adopted in the subsequent stages for rejecting more negatives. Combining with neighbor scales awareness in window search, the proposed method runs at 3~5 fps for high resolution (1360×800) images, 2~7 times as fast as most state-of-the-art methods. Compared with them, the proposed method yields competitive performance on prohibitory signs while sacrifices performance moderately on danger and mandatory signs. Shigang Yue, Jiawei Xu 0004, Xinwen Hou, Cheng-Lin Liu 0001 |
Intelligent Vehicles Symposium | 5 |
| 2015 | A Sparse Projection and Low-Rank Recovery Framework for Handwriting Representation and Salient Stroke Feature ExtractionabstractIn this article, we consider the problem of simultaneous low-rank recovery and sparse projection. More specifically, a new Robust Principal Component Analysis (RPCA)-based framework called Sparse Projection and Low-Rank Recovery (SPLRR) is proposed for handwriting representation and salient stroke feature extraction. In addition to achieving a low-rank component encoding principal features and identify errors or missing values from a given data matrix as RPCA, SPLRR also learns a similarity-preserving sparse projection for extracting salient stroke features and embedding new inputs for classification. These properties make SPLRR applicable for handwriting recognition and stroke correction and enable online computation. A cosine-similarity-style regularization term is incorporated into the SPLRR formulation for encoding the similarities of local handwriting features. The sparse projection and low-rank recovery are calculated from a convex minimization problem that can be efficiently solved in polynomial time. Besides, the supervised extension of SPLRR is also elaborated. The effectiveness of our SPLRR is examined by extensive handwritten digital repairing, stroke correction, and recognition based on benchmark problems. Compared with other related techniques, SPLRR delivers strong generalization capability and state-of-the-art performance for handwriting representation and recognition. Zhao Zhang 0001, Cheng-Lin Liu 0001, Ming-Bo Zhao |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2015 | MTC: A Fast and Robust Graph-Based Transductive Learning MethodabstractDespite the great success of graph-based transductive learning methods, most of them have serious problems in scalability and robustness. In this paper, we propose an efficient and robust graph-based transductive classification method, called minimum tree cut (MTC), which is suitable for large-scale data. Motivated from the sparse representation of graph, we approximate a graph by a spanning tree. Exploiting the simple structure, we develop a linear-time algorithm to label the tree such that the cut size of the tree is minimized. This significantly improves graph-based methods, which typically have a polynomial time complexity. Moreover, we theoretically and empirically show that the performance of MTC is robust to the graph construction, overcoming another big problem of traditional graph-based methods. Extensive experiments on public data sets and applications on web-spam detection and interactive image segmentation demonstrate our method's advantages in aspect of accuracy, speed, and robustness. Yan-Ming Zhang 0001, Kaizhu Huang, Guanggang Geng, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2015 | Retargeted Least Squares Regression AlgorithmabstractThis brief presents a framework of retargeted least squares regression (ReLSR) for multicategory classification. The core idea is to directly learn the regression targets from data other than using the traditional zero-one matrix as regression targets. The learned target matrix can guarantee a large margin constraint for the requirement of correct classification for each data point. Compared with the traditional least squares regression (LSR) and a recently proposed discriminative LSR models, ReLSR is much more accurate in measuring the classification error of the regression model. Furthermore, ReLSR is a single and compact model, hence there is no need to train two-class (binary) machines that are independent of each other. The convex optimization problem of ReLSR is solved elegantly and efficiently with an alternating procedure including regression and retargeting as substeps. The experimental evaluation over a range of databases identifies the validity of our method. Xu-Yao Zhang, Lingfeng Wang 0002, Shiming Xiang, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2014 | Efficient Feature Coding Based on Auto-encoder Network for Image Classification
Guosen Xie, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ACCV (1) | 3 |
| 2014 | Effective license plate detection using fast candidate region selection and covariance feature based filteringabstractThis paper presents a new real-time license plate detection method aiming for fast and accurate detection in live videos. Compared with the previous learning based detection schemes which scan multi-scale images with sliding window, our method takes a cascaded scheme. In the first stage, candidate plate regions are detected based on edge density in reduced image of very low resolution for guaranteeing high speed. In the second stage, the candidate regions are verified using a linear SVM classifier with covariance features for high accuracy. Experimental results on two datasets collected from practical traffic surveillance videos indicate the robustness of our method, which is relatively invariant to scaling, rotation, blurring and illumination. This method takes only 10 msec for detection on a 768 × 576 image. Bo-Yuan Feng, Mingwu Ren, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
AVSS | 4 |
| 2014 | A Seed-Based Segmentation Method for Scene Text ExtractionabstractScene text extraction, i.e., segmenting text pixels from background, is an important step before the text can be recognized. It is a challenging problem due to the cluttered background and the variation of lighting. In this paper, we propose a seed-based segmentation method that can automatically judge the text polarity, extract seed points of text and background, and segment texts by semi-supervised learning (SSL). First, we estimate the text polarity and the stroke width using gradient local correlation. Then, all the points in the middle of stroke edge pairs satisfying the width and polarity are taken as foreground seeds, and the points in the middle of the edge pairs with opposite polarity are taken as background seeds. The whole image is then segmented into text and background using an SSL algorithm. Owing to the accurate estimate of text polarity and extraction of seed points, the proposed method yields good segmentation performance. Experimental results on the KAIST dataset demonstrate the superiority of the method. Cheng-Lin Liu 0001 |
Document Analysis Systems | 3 |
| 2014 | Evaluation of Geometric Context Models for Handwritten Numeral String RecognitionabstractCharacter string recognition based on over segmentation by integrating character classifier and context models has been demonstrated successful. Geometric context models characterizing the candidate character likeliness and between character relationship have shown benefits in several scripts but have not been evaluated in numeral string recognition. Compared with Chinese scripts mixed with alphanumeric and punctuation marks, numeral strings are less variant in character outline and between-character relationship. This study, via evaluating geometric context models used in Chinese handwritten text recognition, shows that geometric context is beneficial to handwritten numeral string recognition as well. Particularly, we propose an improved binary geometric model that combines single-character and between-character features such that the model functions like a bi-character classifier. Combining this binary geometric model with unary geometric model and character classifier, we obtain significant improvement of numeral string recognition performance on the NIST SD-19. Yi-Chao Wu, Cheng-Lin Liu 0001 |
ICFHR | 3 |
| 2014 | Improving Handwritten Chinese Character Recognition with Discriminative Quadratic Feature ExtractionabstractDiscriminative feature extraction (DFE) is an effective linear dimensionality reduction method for pattern recognition. It improves the recognition performance via optimizing subspace projection axes and classifier parameters simultaneously. In this paper, we propose a nonlinear extension of DFE, called discriminative quadratic feature extraction (DQFE), for which feature vectors are firstly mapped to a high-dimensional nonlinear space and then projected to a low-dimensional subspace learned by DFE. The nonlinear mapping is obtained by adding quadratic (correlation or covariance) features computed directly on the original gradient feature maps with different region partition. In this way, both the structural information of the image and the correlation information of features are used to generate a nonlinear high-dimensional feature mapping (thousands of dimensions). Experimental results demonstrated that DQFE can improve the accuracy for different classifiers in handwritten Chinese character recognition. Ming-Ke Zhou, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICPR | 4 |
| 2014 | Integrating supervised subspace criteria with restricted Boltzmann Machine for feature extractionabstractRestricted Boltzmann Machine (RBM) is a widely used building-block in deep neural networks. However, RBM is an unsupervised model which can not exploit the rich supervised information of data. Therefore, we consider combining the descriptive (generative) ability of RBM with the discriminative ability of supervised subspace models, i.e., Fisher linear discriminant analysis (FDA), marginal Fisher analysis (MFA), and heat kernel MFA (hkMFA). Specifically, the hidden layer of RBM is regularized by the supervised subspace criteria, and the joint learning model can then be efficiently optimized by gradient descent and graph construction (used to define the scatter matrix in the subspace models) on mini-batch data. Compared with the traditional subspace models (FDA, MFA, hkMFA), the proposed hybrid models are essentially nonlinear and can be optimized by gradient descent instead of eigenvalue decomposition. More importantly, traditional subspace models can only reduce the dimensionality (because of linear transformation), while the proposed models can also increase the dimensionality for better class discrimination. Experiments on three databases demonstrate that the proposed hybrid models outperform both RBM and their counterpart subspace models (FDA, MFA, hkMFA) consistently. Guosen Xie, Xu-Yao Zhang, Yan-Ming Zhang 0001, Cheng-Lin Liu 0001 |
IJCNN | 4 |
| 2014 | Multi-class segmentation of free-form online documents with tree conditional random fields
Adrien Delaye, Cheng-Lin Liu 0001 |
Int. J. Document Anal. Recognit. | 2 |
| 2014 | Learning confidence transformation for handwritten Chinese text recognition
Dahan Wang, Cheng-Lin Liu 0001 |
Int. J. Document Anal. Recognit. | 2 |
| 2014 | An over-segmentation method for single-touching Chinese handwriting with learning-based filtering
Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
Int. J. Document Anal. Recognit. | 4 |
| 2014 | Graphical lasso quadratic discriminant function and its application to character recognition
Bo Xu 0008, Kaizhu Huang, Irwin King, Cheng-Lin Liu 0001, Jun Sun 0004, Satoshi Naoi |
Neurocomputing | 4 |
| 2014 | Unsupervised language model adaptation for handwritten Chinese text recognition
Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2014 | Character confidence based on N-best list for keyword spotting in online Chinese handwritten documents
Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2014 | Minimum-risk training for semi-Markov conditional random fields with application to handwritten Chinese/Japanese text recognition
Yan-Ming Zhang 0001, Feng Tian 0001, Hongan Wang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 5 |
| 2014 | Combination of Classification and Clustering Results with Label PropagationabstractThis letter considers the combination of multiple classification and clustering results to improve the prediction accuracy. First, an object-similarity graph is constructed from multiple clustering results. The labels predicted by the classification models are then propagated on this graph to adaptively satisfy the smoothness of the prediction over the graph. The convex learning problem is efficiently solved by the label propagation algorithm. A semi-supervised extension is also provided to further improve the performance. Experiments on 11 tasks identify the validity of the proposed models. Xu-Yao Zhang, Peipei Yang, Yan-Ming Zhang 0001, Kaizhu Huang, Cheng-Lin Liu 0001 |
IEEE Signal Process. Lett. | 5 |
| 2014 | Learning Locality Preserving Graph from DataabstractMachine learning based on graph representation, or manifold learning, has attracted great interest in recent years. As the discrete approximation of data manifold, the graph plays a crucial role in these kinds of learning approaches. In this paper, we propose a novel learning method for graph construction, which is distinct from previous methods in that it solves an optimization problem with the aim of directly preserving the local information of the original data set. We show that the proposed objective has close connections with the popular Laplacian Eigenmap problem, and is hence well justified. The optimization turns out to be a quadratic programming problem with n(n-1)/2 variables (n is the number of data points). Exploiting the sparsity of the graph, we further propose a more efficient cutting plane algorithm to solve the problem, making the method better scalable in practice. In the context of clustering and semi-supervised learning, we demonstrated the advantages of our proposed method by experiments. Yan-Ming Zhang 0001, Kaizhu Huang, Xinwen Hou, Cheng-Lin Liu 0001 |
IEEE Trans. Cybern. | 4 |
| 2013 | Scene Text Localization Using Gradient Local CorrelationabstractIn this paper, we propose an efficient scene text localization method using gradient local correlation, which can characterize the density of pair wise edges and stroke width consistency to get a text confidence map. Gradient local correlation is insensitive to the gradient direction and robust to noise, small character size and shadow. Based on the text confidence map, the regions with high confidence are segmented into connected components (CCs), which are classified to text CCs and non-text CCs using an SVM classifier. Then, the text CCs with similar color and stroke width are grouped into text lines, which are in turn partitioned into words. Experimental results on the ICDAR 2003 text locating competition dataset demonstrate the effectiveness of our method. Cheng-Lin Liu 0001 |
ICDAR | 3 |
| 2013 | Hybrid Page Segmentation with Efficient Whitespace Rectangles Extraction and GroupingabstractPage segmentation is still a challenging problem due to the large variety of document layouts. Methods examining both foreground and background regions are among the most effective to solve this problem. However, their performance is influenced by the implementation of two key steps: the extraction and selection of background regions, and the grouping of background regions into separators. This paper proposes an efficient hybrid method for page segmentation. The method extracts white space rectangles based on connected component analysis, and filters white space rectangles progressively incorporating foreground and background information such that the remaining rectangles are likely to form column separators. Experimental results on the ICDAR2009 page segmentation competition test set demonstrate the effectiveness and superiority of the proposed method. Cheng-Lin Liu 0001 |
ICDAR | 3 |
| 2013 | Learning-Based Candidate Segmentation Scoring for Real-Time Recognition of Online Overlaid Chinese HandwritingabstractIn overlaid handwriting, multiple characters are written sequentially in the same area. This needs special consideration for segmenting the stroke sequence into characters. We propose a learning-based model for scoring the candidate stroke cuts and segments for online overlaid Chinese handwriting recognition. Based on stroke cut classification using support vector machine (SVM), strokes are grouped into segments, and consecutive segments are concatenated into candidate characters. The likeliness of candidate characters (unary geometry) and the compatibility between adjacent characters (binary geometry) are measured by combining the stroke cut score and the between-segment geometric score, and are integrated with the character classification score and linguistic context for character string recognition. Experiments on a large database of online Chinese handwriting demonstrate the effectiveness of the proposed method. Yan-Fei Lv, Linlin Huang 0001, Dahan Wang, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2013 | ICDAR 2013 Chinese Handwriting Recognition CompetitionabstractThis paper describes the Chinese handwriting recognition competition held at the 12th International Conference on Document Analysis and Recognition (ICDAR 2013). This third competition in the series again used the CASIA-HWDB/OLHWDB databases as the training set, and all the submitted systems were evaluated on closed datasets to report character-level correct rates. This year, 10 groups submitted 27 systems for five tasks: classification on extracted features, online/offline isolated character recognition, online/offline handwritten text recognition. The best results (correct rates) are 93.89% for classification on extracted features, 94.77% for offline character recognition, 97.39% for online character recognition, 88.76% for offline text recognition, and 95.03% for online text recognition, respectively. In addition to the test results, we also provide short descriptions of the recognition methods and brief discussions on the results. Qiufeng Wang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2013 | Style Consistent Perturbation for Handwritten Chinese Character RecognitionabstractPerturbation-based recognition is effective to recover the deformation of handwritten characters and improve the recognition performance by generating multiple distortions and selecting a distortion that best restores character deformation. Considering that the characters in a field undergo similar deformation under a consistent style, we proposed style consistent perturbation for handwritten character recognition. By generating multiple distortions for the characters in a field, each distortion style is evaluated at the field level and the uniform distortion style of maximum recognition confidence is selected to give the final result. To overcome the slight deviation from uniform style, we also propose to search the neighborhood distortions from the optimal uniform distortion for higher confidence. The experiments of handwritten Chinese character recognition on multi-writer data show that style consistent perturbation in very short fields outperforms individual character recognition, and neighborhood distortion search yields further improvement. Ming-Ke Zhou, Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2013 | Locally Smoothed Modified Quadratic Discriminant FunctionabstractModified quadratic discriminant function (MQDF) is a state-of-the-art classifier for handwriting recognition. However, the big gap between accuracies on training and testing sets indicates that MQDF has a good capability to fit training data but the generalization performance is not promising. To solve this problem, we propose a new model called locally smoothed modified quadratic discriminant function (LSMQDF) by smoothing the covariance matrix of each class with its nearest neighbor classes. LSMQDF can be viewed as a regularization to avoid over-fitting. The covariance matrix estimated by local smoothing is more accurate and robust. LSMQDF can be also viewed as an extension of the global smoothing method, namely regularized discriminant analysis (RDA). Experiments on both offline and online Chinese handwriting databases demonstrate that: with local smoothing, the accuracy on training set is decreased (over-fitting avoided), and the accuracy on testing set is improved significantly and consistently (generalization improved). Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICDAR | 2 |
| 2013 | Minimum Risk Training for Handwritten Chinese/Japanese Text Recognition Using Semi-Markov Conditional Random FieldsabstractSemi-Markov conditional random fields (semi-CRFs) are usually trained with maximum a posteriori (MAP) criterion which adopts the 0/1 cost for measuring the loss of misclassification. In this paper, based on our previous work on handwritten Chinese/Japanese text recognition (HCTR) using semi-CRFs, we propose an alternative parameter learning method by minimizing the risk, in which the misclassification costs are not equal, but different depending on the hypothesis and the ground-truth. The proposed method is lattice-based, i.e., the hypothesis space is the entire lattice on which the semi-CRF is defined. Experimental results on two online handwriting databases: CASIA-OLHWDB and TUAT Kondate demonstrate that minimum-risk training can yield superior string recognition rates compared to MAP training. Feng Tian 0001, Cheng-Lin Liu 0001, Hongan Wang |
ICDAR | 3 |
| 2013 | GPU-Based Fast Training of Discriminative Learning Quadratic Discriminant Function for Handwritten Chinese Character RecognitionabstractThe discriminative training of classifiers for handwritten Chinese character recognition (HCCR) is highly demanding in computation due to the large number of categories. The inability of discriminative training with large sample set on personal computers has hindered the accuracy promotion for HCCR. To overcome this problem, we have implemented the training algorithm of discriminative learning quadratic discriminant function (DLQDF) on our graphics processing units (GPU) server, and have achieved 15 times speedup compared to single-core computation. By enlarging training sample set via distortion on a standard dataset of 3,755 classes, we could train the DLQDF on more than 50 million samples within 150min and get the test accuracy improved by 1.36%. Ming-Ke Zhou, Cheng-Lin Liu 0001 |
ICDAR | 3 |
| 2013 | Feature Transformation with Class Conditional DecorrelationabstractThe well-known feature transformation model of Fisher linear discriminant analysis (FDA) can be decomposed into an equivalent two-step approach: whitening followed by principal component analysis (PCA) in the whitened space. By proving that whitening is the optimal linear transformation to the Euclidean space in the sense of minimum log-determinant divergence, we propose a transformation model called class conditional decor relation (CCD). The objective of CCD is to diagonalize the covariance matrices of different classes simultaneously, which is efficiently optimized using a modified Jacobi method. CCD is effective to find the common principal components among multiple classes. After CCD, the variables become class conditionally uncorrelated, which will benefit the subsequent classification tasks. Combining CCD with the nearest class mean (NCM) classification model can significantly improve the classification accuracy. Experiments on 15 small-scale datasets and one large-scale dataset (with 3755 classes) demonstrate the scalability of CCD for different applications. We also discuss the potential applications of CCD for other problems such as Gaussian mixture models and classifier ensemble learning. Xu-Yao Zhang, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICDM | 3 |
| 2013 | Handwriting representation and recognition through a sparse projection and low-rank recovery frameworkabstractThis paper proposes a Robust Principal Component Analysis (RPCA) based framework called Sparse Projection and Low-Rank Recovery (SPLRR) for representing and recognizing handwritings. SPLRR calculates a similarity preserving sparse projection for salient feature extraction and processing new data for classification in addition to delivering a low-rank principal component and identifying errors or missing pixel values from a given data matrix. As a result, SPLRR will be applicable for handwritten recovery, recognition and the applications requiring online computation. To encode the similarity between features in the learning process, the Cosine similarity based regularization is incorporated to the SPLRR formulation. The sparse projection and the lowest-rank components are calculated from a scalable convex minimization problem that can be efficiently solved in polynomial time. The effectiveness of the proposed SPLRR is examined by handwritten digital repairing, stroke correction and recognition on two real problems. Results show that SPLRR can deliver state-of-the-art results in handwriting representation. Zhao Zhang 0001, Cheng-Lin Liu 0001, Ming-Bo Zhao |
IJCNN | 2 |
| 2013 | Fast kNN Graph Construction with Locality Sensitive Hashing
Yan-Ming Zhang 0001, Kaizhu Huang, Guanggang Geng, Cheng-Lin Liu 0001 |
ECML/PKDD (2) | 4 |
| 2013 | Dense Trajectories and Motion Boundary Descriptors for Action Recognition
Alexander Kläser, Cordelia Schmid, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 4 |
| 2013 | An evaluation of statistical methods in handwritten hangul recognition
Gyu-Ro Park, Injung Kim 0001, Cheng-Lin Liu 0001 |
Int. J. Document Anal. Recognit. | 3 |
| 2013 | Keyword Spotting from Online Chinese Handwritten Documents using One-versus-All Character Classification ModelabstractIn this paper, we propose a method for text-query-based keyword spotting from online Chinese handwritten documents using character classification model. The similarity between the query word and handwriting is obtained by combining the character classification scores. The classifier is trained by one-versus-all strategy so that it gives high similarity to the target class and low scores to the others. Using character classification-based word similarity also helps overcome the out-of-vocabulary (OOV) problem. We use a character-synchronous dynamic search algorithm to efficiently spot the query word in large database. The retrieval performance is further improved by using competing character confusion and writer-adaptive thresholds. Our experimental results on a large handwriting database CASIA-OLHWDB justify the superiority of one-versus-all trained classifiers and the benefits of confidence transformation, character confusion and adaptive thresholds. Particularly, a one-versus-all trained prototype classifier performs as well as a linear support vector machine (SVM) classifier, but consumes much less storage of index file. The experimental comparison with keyword spotting based on handwritten text recognition also demonstrates the effectiveness of the proposed method. Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001, Horst Bunke |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2013 | Keyword spotting in unconstrained handwritten Chinese documents using contextual word model
Qing-Hu Chen, Cheng-Lin Liu 0001 |
Image Vis. Comput. | 4 |
| 2013 | Geometry preserving multi-task metric learning
Peipei Yang, Kaizhu Huang, Cheng-Lin Liu 0001 |
Mach. Learn. | 3 |
| 2013 | A multi-task framework for metric learning with common subspace
Peipei Yang, Kaizhu Huang, Cheng-Lin Liu 0001 |
Neural Comput. Appl. | 3 |
| 2013 | Writer Adaptation with Style Transfer MappingabstractAdapting a writer-independent classifier toward the unique handwriting style of a particular writer has the potential to significantly increase accuracy for personalized handwriting recognition. This paper proposes a novel framework of style transfer mapping (STM) for writer adaptation. The STM is a writer-specific class-independent feature transformation which has a closed-form solution. After style transfer mapping, the data of different writers are projected onto a style-free space, where the writer-independent classifier needs no change to classify the transformed data and can achieve significantly higher accuracy. The framework of STM can be combined with different types of classifiers for supervised, unsupervised, and semi-supervised adaptation, where writer-specific data can be either labeled or unlabeled and need not cover all classes. In this paper, we combine STM with the state-of-the-art classifiers for large-category Chinese handwriting recognition: learning vector quantization (LVQ) and modified quadratic discriminant function (MQDF). Experiments on the online Chinese handwriting database CASIA-OLHWDB demonstrate that STM-based adaptation is very efficient and effective in improving classification accuracy. Semi-supervised adaptation achieves the best performance, while unsupervised adaptation is even better than supervised adaptation. On handwritten text data, semi-supervised adaptation achieves error reduction rates 31.95 and 25.00 percent by LVQ and MQDF, respectively. Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Handwritten Chinese/Japanese Text Recognition Using Semi-Markov Conditional Random FieldsabstractThis paper proposes a method for handwritten Chinese/Japanese text (character string) recognition based on semi-Markov conditional random fields (semi-CRFs). The high-order semi-CRF model is defined on a lattice containing all possible segmentation-recognition hypotheses of a string to elegantly fuse the scores of candidate character recognition and the compatibilities of geometric and linguistic contexts by representing them in the feature functions. Based on given models of character recognition and compatibilities, the fusion parameters are optimized by minimizing the negative log-likelihood loss with a margin term on a training string sample set. A forward-backward lattice pruning algorithm is proposed to reduce the computation in training when trigram language models are used, and beam search techniques are investigated to accelerate the decoding speed. We evaluate the performance of the proposed method on unconstrained online handwritten text lines of three databases. On the test sets of databases CASIA-OLHWDB (Chinese) and TUAT Kondate (Japanese), the character level correct rates are 95.20 and 95.44 percent, and the accurate rates are 94.54 and 94.55 percent, respectively. On the test set (online handwritten texts) of ICDAR 2011 Chinese handwriting recognition competition, the proposed method outperforms the best system in competition. Dahan Wang, Feng Tian 0001, Cheng-Lin Liu 0001, Masaki Nakagawa |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2013 | Online and offline handwritten Chinese character recognition: Benchmarking on new databases
Cheng-Lin Liu 0001, Dahan Wang, Qiufeng Wang 0001 |
Pattern Recognit. | 1 |
| 2013 | Transcript mapping for handwritten Chinese documents by integrating character recognition model and geometric context
Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2013 | Evaluation of weighted Fisher criteria for large category dimensionality reduction in application to Chinese handwriting recognition
Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2013 | Error-correcting output codes based ensemble feature extraction
Guoqiang Zhong 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2012 | A Fast Stroke-Based Method for Text Detection in VideoabstractTexts in video provide a rich clue for video indexing and retrieval, yet the detection and recognition of video text remains a challenge. This paper proposes an effective and real-time stroke-based method for text detection in video, which is robust to the change of stroke intensity and width. Particularly, we propose to characterize the text confidence using an edge orientation variance (EOV) and an opposite edge pair (OEP) feature. Based on the text confidence map, candidate text components are extracted and grouped into text lines by thresholding and connected component analysis. Our experimental results demonstrate that the proposed method can detect multilingual texts in video with fairly high accuracy. Cheng-Lin Liu 0001 |
Document Analysis Systems | 3 |
| 2012 | Arabic Handwritten Text Line Extraction by Applying an Adaptive Mask to Morphological DilationabstractThis paper presents a robust method for handwritten text line extraction. We use morphological dilation with a dynamic adaptive mask for line extraction. Line separation occurs because of the repulsion and attraction between connected components. The characteristics of the Arabic script are considered to ensure a high performance of the algorithm. Our method is evaluated on the CENPARMI Arabic handwritten documents database which contains multi-skewed and touching lines. With a matching score of 0.95, our method achieved precision and recall rates of 96:3% and 96:7% respectively, which demonstrate the effectiveness of our approach. Muna Khayyat, Louisa Lam, Ching Y. Suen, Cheng-Lin Liu 0001 |
Document Analysis Systems | 5 |
| 2012 | Improving Handwritten Chinese Text Recognition by Unsupervised Language Model AdaptationabstractThis paper investigates the effects of unsupervised language model adaptation (LMA) in handwritten Chinese text recognition. For no prior information of recognition text is available, we use a two-pass recognition strategy. In the first pass, the generic language model (LM) is used to get a preliminary result, which is used to choose the best matched LMs from a set of pre-defined domains, then the matched LMs are used in the second pass recognition. Each LM is compressed to a moderate size via the entropy-based pruning, tree-structure formatting and fewer-byte quantization. We evaluated the LMA for five LM types, including both character-level and word-level ones. Experiments on the CASIA-HWDB database show that language model adaptation improves the performance for each LM type in all domains. The documents of ancient domain gained the biggest improvement of character-level correct rate of 5.87 percent up and accurate rate of 6.05 percent up. Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
Document Analysis Systems | 3 |
| 2012 | A Touching Character Database from Chinese Handwriting for Assessing Segmentation AlgorithmsabstractFor assessing touching character segmentation algorithms, we present a database of touching characters collected from the Chinese handwriting database CASIA-HWDB, called CASIA-HWDB-T. It includes 56,469 two-character or multiple-character touching strings, among which 1,818 strings have multiple-touching characters. We also partition the touching strings into 50,157 all-Chinese strings, 2,788 all-digit ones, 328 all-letter ones, and 3,196 mixed-character ones. All the strings are annotated with the character classes, locations of touching points, and auxiliary values like string height and average stroke width. And last, we measure the segmentation performance of three existing algorithms on this database for reference. Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
ICFHR | 4 |
| 2012 | Confused Distance Maximization for Large Category Dimensionality ReductionabstractThe Fisher linear discriminant analysis (FDA) is the most well-known supervised dimensionality reduction model. However, when the number of classes is much larger than the reduced dimensionality, FDA suffers from the class separation problem in that it will preserve the distances of the already well-separated classes and cause a large overlap of neighboring classes. To cope with this problem, we propose a new model called confused distance maximization (CDM). The objective of CDM is to maximize the distance of the most confusable classes, according to the confusion matrix estimated from the training data with a pre-learned classifier. Compared with FDA that maximizes the sum of the distances of all class pairs, CDM is more relevant to the classification accuracy by weighting the pairwise distance according to the confusion matrix. Furthermore, CDM is computationally inexpensive which makes it indeed efficient and effective for large category problems. Experiments on two large-scale 3,755-class Chinese handwriting databases (offline and online) demonstrate that CDM can achieve the best performance compared with FDA and other competitive weighting based criteria. Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICFHR | 2 |
| 2012 | Multiple Outlooks Learning with Support Vector Machines
Yinglu Liu, Xu-Yao Zhang, Kaizhu Huang, Xinwen Hou, Cheng-Lin Liu 0001 |
ICONIP (3) | 5 |
| 2012 | Manifold Regularized Multi-Task Learning
Peipei Yang, Xu-Yao Zhang, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICONIP (3) | 4 |
| 2012 | Insect species recognition using discriminative local soft coding
An Lu, Xinwen Hou, Cheng-Lin Liu 0001 |
ICPR | 3 |
| 2012 | String-level learning of confidence transformation for Chinese handwritten text recognition
Dahan Wang, Cheng-Lin Liu 0001 |
ICPR | 2 |
| 2012 | A confidence-based method for keyword spotting in online Chinese handwritten documents
Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001 |
ICPR | 3 |
| 2012 | Geometry Preserving Multi-task Metric Learning
Peipei Yang, Kaizhu Huang, Cheng-Lin Liu 0001 |
ECML/PKDD (1) | 3 |
| 2012 | Joint learning of error-correcting output codes and dichotomizers from data
Guoqiang Zhong 0001, Kaizhu Huang, Cheng-Lin Liu 0001 |
Neural Comput. Appl. | 3 |
| 2012 | Maxi-Min discriminant analysis via online learning
Bo Xu 0008, Kaizhu Huang, Cheng-Lin Liu 0001 |
Neural Networks | 3 |
| 2012 | Handwritten Chinese Text Recognition by Integrating Multiple ContextsabstractThis paper presents an effective approach for the offline recognition of unconstrained handwritten Chinese texts. Under the general integrated segmentation-and-recognition framework with character oversegmentation, we investigate three important issues: candidate path evaluation, path search, and parameter estimation. For path evaluation, we combine multiple contexts (character recognition scores, geometric and linguistic contexts) from the Bayesian decision view, and convert the classifier outputs to posterior probabilities via confidence transformation. In path search, we use a refined beam search algorithm to improve the search efficiency and, meanwhile, use a candidate character augmentation strategy to improve the recognition accuracy. The combining weights of the path evaluation function are optimized by supervised learning using a Maximum Character Accuracy criterion. We evaluated the recognition performance on a Chinese handwriting database CASIA-HWDB, which contains nearly four million character samples of 7,356 classes and 5,091 pages of unconstrained handwritten texts. The experimental results show that confidence transformation and combining multiple contexts improve the text line recognition performance significantly. On a test set of 1,015 handwritten pages, the proposed approach achieved character-level accurate rate of 90.75 percent and correct rate of 91.39 percent, which are superior by far to the best results reported in the literature. Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | An approach for real-time recognition of online Chinese handwritten sentences
Dahan Wang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2011 | Action recognition by dense trajectoriesabstractFeature trajectories have shown to be efficient for representing videos. Typically, they are extracted using the KLT tracker or matching SIFT descriptors between frames. However, the quality as well as quantity of these trajectories is often not sufficient. Inspired by the recent success of dense sampling in image classification, we propose an approach to describe videos by dense trajectories. We sample dense points from each frame and track them based on displacement information from a dense optical flow field. Given a state-of-the-art optical flow algorithm, our trajectories are robust to fast irregular motions as well as shot boundaries. Additionally, dense trajectories cover the motion information in videos well. We, also, investigate how to design descriptors to encode the trajectory information. We introduce a novel descriptor based on motion boundary histograms, which is robust to camera motion. This descriptor consistently outperforms other state-of-the-art descriptors, in particular in uncontrolled realistic videos. We evaluate our video description in the context of action classification with a bag-of-features approach. Experimental results show a significant improvement over the state of the art on four datasets of varying difficulty, i.e. KTH, YouTube, Hollywood2 and UCF sports. Alexander Kläser, Cordelia Schmid, Cheng-Lin Liu 0001 |
CVPR | 4 |
| 2011 | Style transfer matrix learning for writer adaptationabstractIn this paper, we propose a novel framework of style transfer matrix (STM) learning to reduce the writing style variation in handwriting recognition. After writer-specific style transfer learning, the data of different writers is projected onto a style-free space, where a writer independent classifier can yield high accuracy. We combine STM learning with a specific nearest prototype classifier: learning vector quantization (LVQ) with discriminative feature extraction (DFE), where both the prototypes and the subspace transformation matrix are learned via online discriminative learning. To adapt the basic classifier (trained with writer-independent data) to particular writers, we first propose two supervised models, one based on incremental learning and the other based on supervised STM learning. To overcome the lack of labeled samples for particular writers, we propose an unsupervised model to learn the STM using the self-taught strategy (also known as self-training). Experiments on a large-scale Chinese online handwriting database demonstrate that STM learning can reduce recognition errors significantly, and the unsupervised adaptation model performs even better than the supervised models. Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 2 |
| 2011 | Keyword Spotting in Offline Chinese Handwritten Documents Using a Statistical ModelabstractThis paper proposes a method for keyword spotting in offline Chinese handwritten documents using a statistical model. On a text query word, the method measures the similarity between the query word and every candidate word in the document by combining a character classifier and four classifiers characterizing the geometric contexts. By over-segmenting text lines into primitive segments, candidate characters and words are generated by concatenating consecutive segments, and the beam search strategy is used to search all the candidate words. The character classifier and the model combining weights are trained by optimizing a one-vs-all discrimination objective so as to maximize the similarity of true words and minimize the similarity of imposters. In experiments on a test dataset containing 1,015 pages of 180 writers, the proposed methods yields promising performance. For retrieving four-characer words, the recall, precision and F-measure are 92.47%, 83.76% and 87.90%, respectively. Qing-Hu Chen, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2011 | CASIA Online and Offline Chinese Handwriting DatabasesabstractThis paper introduces a pair of online and offline Chinese handwriting databases, containing samples of isolated characters and handwritten texts. The samples were produced by 1,020 writers using Anoto pen on papers for obtaining both online trajectory data and offline images. Both the online samples and offline samples are divided into six datasets, three for isolated characters (DB1.0-C1.2) and three for handwritten texts (DB2.0-C2.2). The (either online or offline) datasets of isolated characters contain about 3.9 million samples of 7,356 classes (7,185 Chinese characters and 171 symbols), and the datasets of handwritten texts contain about 5,090 pages and 1.35 million character samples. Each dataset is segmented and annotated at character level, and is partitioned into standard training and test subsets. The online and offline databases can be used for the research of various handwritten document analysis tasks. Cheng-Lin Liu 0001, Dahan Wang, Qiufeng Wang 0001 |
ICDAR | 1 |
| 2011 | ICDAR 2011 Chinese Handwriting Recognition CompetitionabstractIn the Chinese handwriting recognition competition organized with the ICDAR 2011, four tasks were evaluated: offline and online isolated character recognition, offline and online handwritten text recognition. To enable the training of recognition systems, we announced the large databases CASIA-HWDB/OLHWDB. The submitted systems were evaluated on un-open datasets to report character-level correct rates. In total, we received 25 systems submitted by eight groups. On the test datasets, the best results (correct rates) are 92.18% for offline character recognition, 95.77% for online character recognition, 77.26% for offline text recognition, and 94.33% for online text recognition, respectively. In addition to the evaluation results, we provide short descriptions of the recognition methods and have brief discussions. Cheng-Lin Liu 0001, Qiufeng Wang 0001, Dahan Wang |
ICDAR | 1 |
| 2011 | Perceptron Learning of Modified Quadratic Discriminant FunctionabstractModified quadratic discriminant function (MQDF) is the state-of-the-art classifier in handwritten character recognition. Discriminative learning of MQDF can further improve its performance. Recent advances justify the efficacy of minimum classification error criteria in learning MQDF (MCE-MQDF). We provide an alternative choice to MCE-MQDF based on the Perceptron learning (PL-MQDF). For better generalization performance, we propose a new dynamic margin regularization. To relieve the heavy burden in training process, active set technique is employed, which can save most of the computation with negligible loss in accuracy. In experiments on handwritten digit datasets and a large-scale Chinese handwritten character database, the proposed PL-MQDF was demonstrated superior in both error reduction and training speedup. Tong-Hua Su, Cheng-Lin Liu 0001, Xu-Yao Zhang |
ICDAR | 2 |
| 2011 | Dynamic Text Line Segmentation for Real-Time Recognition of Chinese Handwritten SentencesabstractReal-time recognition of handwritten sentences enables fast text input but the dynamic nature of writing makes reliable text line segmentation difficult. This paper proposes a method for real-time dynamic text line segmentation of online Chinese handwriting. The core of the method is a statistical classifier for modeling the geometric relationship between an ongoing stroke and the previous text lines, to assign the stroke into a previous line or form a new line. The method can deal with delayed strokes and therefore enables robust real-time recognition. We evaluated the segmentation performance on a dataset of online Chinese handwriting by simulating the real-time writing and recognition process. The experimental results demonstrate the effectiveness and robustness of the proposed method. Dahan Wang, Cheng-Lin Liu 0001 |
ICDAR | 2 |
| 2011 | Improving Handwritten Chinese Text Recognition by Confidence TransformationabstractThis paper investigates the effects of confidence transformation (CT) of the character classifier outputs in handwritten Chinese text recognition. The classifier outputs are transformed to confidence values in three confidence types, namely, sigmoid, soft max and Dempster-Shafer theory of evidence (D-S evidence). The confidence parameters are optimized by minimizing the cross-entropy (CE) loss function (both binary and multi-class) on a validation dataset, where we add non-character samples to enhance the outlier rejection capability in text recognition. Experimental results on the CASIA-HWDB database show that confidence transformation improves the handwritten text recognition performance significantly and adding non-characters for confidence parameter estimation is beneficial. Among the confidence types, the D-S evidence performs best. Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
ICDAR | 3 |
| 2011 | Touching Character Separation in Chinese Handwriting Using Visibility-Based Foreground AnalysisabstractIn offline handwritten text recognition, the separation of touching characters remains a challenge due to the variability of touching structures. This paper proposes a new touching character separation method for Chinese handwriting based on skeleton analysis and contour analysis incorporating the visibility of separating points. Separating points are detected from strokes that are common in both upper and lower skeleton tracing, and the profile visibility of strokes and separating points is analyzed to adjust and verify separating points. Our experiments on two large handwriting databases demonstrate the effectiveness of the proposed method. Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2011 | Transcript Mapping for Handwritten Text Lines Using Conditional Random FieldsabstractThis paper presents a conditional random field (CRF) model for aligning online handwritten Chinese/Japanese text lines (character strings) with the corresponding transcripts. The CRF model is defined on a lattice which contains all possible segmentation hypotheses. The feature functions characterize the shape and context dependences of characters, including the scores of character recognition and the geometric compatibilities between characters. The combining parameters are optimized by energy minimization. Experimental results on two online databases: CASIA-OLHWDB and TUAT Kondate demonstrate the effectiveness of the proposed method. Dahan Wang, Qiufeng Wang 0001, Masaki Nakagawa, Cheng-Lin Liu 0001 |
ICDAR | 6 |
| 2011 | Fast and Robust Graph-based Transductive Learning via Minimum Tree CutabstractIn this paper, we propose an efficient and robust algorithm for graph-based transductive classification. After approximating a graph with a spanning tree, we develop a linear-time algorithm to label the tree such that the cut size of the tree is minimized. This significantly improves typical graph-based methods, which either have a cubic time complexity (for a dense graph) or O(kn2) (for a sparse graph with k denoting the node degree). Furthermore, our method shows great robustness to the graph construction both theoretically and empirically; this overcomes another big problem of traditional graph-based methods. In addition to its good scalability and robustness, the proposed algorithm demonstrates high accuracy. In particular, on a graph with 400,000 nodes (in which 10,000 nodes are labeled) and 10,455,545 edges, our algorithm achieves the highest accuracy of 99.6% but takes less than 10 seconds to label all the unlabeled data. Yan-Ming Zhang 0001, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICDM | 3 |
| 2011 | Low Rank Metric Learning with Manifold RegularizationabstractIn this paper, we present a semi-supervised method to learn a low rank Mahalanobis distance function. Based on an approximation to the projection distance from a manifold, we propose a novel parametric manifold regularizer. In contrast to previous approaches that usually exploit side information only, our proposed method can further take advantages of the intrinsic manifold information from data. In addition, we focus on learning a metric of low rank directly, this is different from traditional approaches that often enforce the l1norm on the metric. The resulting configuration is convex with respect to the manifold structure and the distance function, respectively. We solve it with an alternating optimization algorithm, which proves effective to find a satisfactory solution. For efficient implementation, we even present a fast algorithm, in which the manifold structure and the distance function are learned independently without alternating minimization. Experimental results over 12 standard UCI data sets demonstrate the advantages of our method. Guoqiang Zhong 0001, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICDM | 3 |
| 2011 | Graphical Lasso Quadratic Discriminant Function for Character Recognition
Bo Xu 0008, Kaizhu Huang, Irwin King, Cheng-Lin Liu 0001, Jun Sun 0004, Satoshi Naoi |
ICONIP (3) | 4 |
| 2011 | Multi-Task Low-Rank Metric Learning Based on Common Subspace
Peipei Yang, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICONIP (2) | 3 |
| 2011 | Pattern Field Classification with Style Normalized Transformation
Xu-Yao Zhang, Kaizhu Huang, Cheng-Lin Liu 0001 |
IJCAI | 3 |
| 2011 | A Hybrid Approach to Detect and Localize Texts in Natural Scene ImagesabstractText detection and localization in natural scene images is important for content-based image analysis. This problem is challenging due to the complex background, the non-uniform illumination, the variations of text font, size and line orientation. In this paper, we present a hybrid approach to robustly detect and localize texts in natural scene images. A text region detector is designed to estimate the text existing confidence and scale information in image pyramid, which help segment candidate text components by local binarization. To efficiently filter out the non-text components, a conditional random field (CRF) model considering unary component properties and binary contextual component relationships with supervised parameter learning is proposed. Finally, text components are grouped into text lines/words with a learning-based energy minimization method. Since all the three stages are learning-based, there are very few parameters requiring manual tuning. Experimental results evaluated on the ICDAR 2005 competition dataset show that our approach yields higher precision and recall performance compared with state-of-the-art methods. We also evaluated our approach on a multilingual image dataset with promising results. Yi-Feng Pan, Xinwen Hou, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2010 | Transductive Learning on Adaptive GraphsabstractGraph-based semi-supervised learning methods are based on some smoothness assumption about the data. As a discrete approximation of the data manifold, the graph plays a crucial role in the success of such graph-based methods. In most existing methods, graph construction makes use of a predefined weighting function without utilizing label information even when it is available. In this work, by incorporating label information, we seek to enhance the performance of graph-based semi-supervised learning by learning the graph and label inference simultaneously. In particular, we consider a particular setting of semi-supervised learning called transductive learning. Using the LogDet divergence to define the objective function, we propose an iterative algorithm to solve the optimization problem which has closed-form solution in each step. We perform experiments on both synthetic and real data to demonstrate improvement in the graph and in terms of classification accuracy. Yan-Ming Zhang 0001, Yu Zhang 0006, Dit-Yan Yeung, Cheng-Lin Liu 0001, Xinwen Hou |
AAAI | 4 |
| 2010 | Gaussian Process Latent Random FieldabstractIn this paper, we propose a novel supervised extension of GPLVM, called Gaussian process latent random field (GPLRF), by enforcing the latent variables to be a Gaussian Markov random field with respect to a graph constructed from the supervisory information. Guoqiang Zhong 0001, Wu-Jun Li, Dit-Yan Yeung, Xinwen Hou, Cheng-Lin Liu 0001 |
AAAI | 5 |
| 2010 | Similar Handwritten Chinese Characters Recognition by Critical Region Selection Based on Average Symmetric UncertaintyabstractWe consider the problem of similar Chinese character recognition in this paper. Engaging the Average Symmetric Uncertainty (ASU) criterion to measure the correlation between different image regions and the class label, we manage to detect the most critical regions for each pair of similar characters. These critical regions are proved to contain more discriminative information and hence can largely benefit the classification accuracy for similar characters. We conduct a series of experiments on the CASIA Chinese character data set. Experimental results show that our proposed method is superior to three competitive approaches in terms of both accuracy and efficiency. Bo Xu 0008, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICFHR | 3 |
| 2010 | Integrating Geometric Context for Text Alignment of Handwritten Chinese DocumentsabstractThe alignment of text line images with text transcript is a crucial step of handwritten document annotation. Handwritten text alignment is prone to errors due to the difficulty of character segmentation and the variability of character shape, size and position. In this paper, we propose to incorporate the geometric context of character strings to improve the alignment accuracy for offline handwritten Chinese documents. We use four statistical models to evaluate the geometric features of single characters and between-character relationships. By combining the geometric models with a character recognizer, we have achieved a large improvement of alignment accuracy in our experiments on unconstrained handwritten Chinese text lines. Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
ICFHR | 3 |
| 2010 | Keyword Spotting from Online Chinese Handwritten Documents Using One-vs-All Trained Character ClassifierabstractThis paper presents a text query-based method for keyword spotting from online Chinese handwritten documents. The similarity between a text word and handwriting is obtained by combining the character similiarity scores given by a character classifier. To overcome the ambiguity of character segmentation, multiple candidates of character patterns are generated by over-segmentation, and sequences of candidate characters are matched with the query word in beam search. The character classifier is trained by one-vs-all strategy so that it gives high similarity to the target class and low scores to the others. Particularly, we use a one-vs-all trained prototype classifier and a support vector machine (SVM) classifier for similarity scoring. The method yielded promising performance in experiments on a database containing 550 pages of 110 writers. For words of four characters, the recall, precision and F measure are 87.25%, 94.84% and 90.88%, respectively. Heng Zhang 0028, Dahan Wang, Cheng-Lin Liu 0001 |
ICFHR | 3 |
| 2010 | Error Reduction by Confusing Characters Discrimination for Online Handwritten Japanese Character RecognitionabstractTo reduce the classification errors of online handwritten Japanese character recognition, we propose a method for confusing characters discrimination with little additional costs. After building confusing sets by cross validation using a baseline quadratic classifier, a logistic regression (LR) classifier is trained to discriminate the characters in each set. The LR classifier uses subspace features selected from existing vectors of the baseline classifier, thus has no extra parameters except the weights, which consumes a small storage space compared to the baseline classifier. In experiments on the TUAT HANDS databases with the modified quadratic discriminant function (MQDF) as baseline classifier, the proposed method has largely reduced the confusion caused by non-Kanji characters. Dahan Wang, Masaki Nakagawa, Cheng-Lin Liu 0001 |
ICFHR | 4 |
| 2010 | Discriminative Training of Subspace Gaussian Mixture Model for Pattern Classification
Xiao-Hua Liu, Cheng-Lin Liu 0001 |
ICIC (1) | 2 |
| 2010 | Fast scene text localization by learning-based filtering and verificationabstractThis paper proposes a new method for fast text localization in natural scene images by combining learning-based region filtering and verification in a coarse-to-fine strategy. In each pyramid layer, a boosted region filter is used to extract candidate text regions, which are segmented into candidate text lines by multi-orientation projection analysis. A polynomial classifier with combined features is used to verify patches of candidate text lines for removing non-texts. The remaining text patches over all pyramid layers are grouped into text lines based on their spatial relationships. The text lines are further refined and partitioned into words by connected component analysis. Experimental results show that the proposed method provides competitive localization performance at high speed. Yi-Feng Pan, Cheng-Lin Liu 0001, Xinwen Hou |
ICIP | 2 |
| 2010 | Learning ECOC and Dichotomizers Jointly from Data
Guoqiang Zhong 0001, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICONIP (1) | 3 |
| 2010 | Multi-class AdaBoost with Hypothesis MarginabstractMost AdaBoost algorithms for multi-class problems have to decompose the multi-class classification into multiple binary problems, like the Adaboost.MH and the LogitBoost. This paper proposes a new multi-class AdaBoost algorithm based on hypothesis margin, called AdaBoost.HM, which directly combines multi-class weak classifiers. The hypothesis margin maximizes the output about the positive class meanwhile minimizes the maximal outputs about the negative classes. We discuss the upper bound of the training error about AdaBoost.HM and a previous multi-class learning algorithm AdaBoost.M1. Our experiments using feed forward neural networks as weak learners show that the proposed AdaBoost.HM yields higher classification accuracies than the AdaBoost.M1 and the AdaBoost.MH, and meanwhile, AdaBoost.HM is computationally efficient in training. Xiao-Bo Jin, Xinwen Hou, Cheng-Lin Liu 0001 |
ICPR | 3 |
| 2010 | Boosting Incremental Semi-supervised Discriminant Analysis for TrackingabstractTracking is recently formulated as a problem of discriminating the object from its nearby background, where the classifier is updated by new samples successively arriving during tracking. Depending on whether labeling the samples or not, the tracker can be designed in a supervised or semi-supervised manner. This paper proposes a novel semi-supervised algorithm for tracking by combining Semi-supervised Discriminant Analysis (SDA) with an online boosting framework. Using the local geometric structure information from the samples, the SDA-based weak classifier is made more robust to outliers. Meanwhile, we design an incremental updating mechanism for SDA so that it can adapt to appearance changes. We further propose an Extended SDA (ESDA) algorithm, which gives better discrimination ability. Results on several challenging video sequences demonstrate the effectiveness of the method. Xinwen Hou, Cheng-Lin Liu 0001 |
ICPR | 3 |
| 2010 | Dimensionality Reduction by Minimal Distance MaximizationabstractIn this paper, we propose a novel discriminant analysis method, called Minimal Distance Maximization (MDM). In contrast to the traditional LDA, which actually maximizes the average divergence among classes, MDM attempts to find a low-dimensional subspace that maximizes the minimal (worst-case) divergence among classes. This ``minimal" setting solves the problem caused by the ``average" setting of LDA that tends to merge similar classes with smaller divergence when used for multi-class data. Furthermore, we elegantly formulate the worst-case problem as a convex problem, making the algorithm solvable for larger data sets. Experimental results demonstrate the advantages of our proposed method against five other competitive approaches on one synthetic and six real-life data sets. Bo Xu 0008, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICPR | 3 |
| 2010 | Robust Metric Learning by Smooth Optimization
Kaizhu Huang, Rong Jin 0001, Zenglin Xu, Cheng-Lin Liu 0001 |
UAI | 4 |
| 2010 | Regularized margin-based conditional log-likelihood loss for prototype learning
Xiao-Bo Jin, Cheng-Lin Liu 0001, Xinwen Hou |
Pattern Recognit. | 2 |
| 2009 | Text Localization in Natural Scene Images Based on Conditional Random FieldabstractThis paper proposes a novel hybrid method to robustly and accurately localize texts in natural scene images. A text region detector is designed to generate a text confidence map, based on which text components can be segmented by local binarization approach. A conditional random field (CRF) model, considering the unary component property as well as binary neighboring component relationship, is then presented to label components as "text" or "non-text". Last, text components are grouped into text lines with an energy minimization approach. Experimental results show that the proposed method gives promising performance comparing with the existing methods on ICDAR 2003 competition dataset. Yi-Feng Pan, Xinwen Hou, Cheng-Lin Liu 0001 |
ICDAR | 3 |
| 2009 | CASIA-OLHWDB1: A Database of Online Handwritten Chinese CharactersabstractThis paper describes a publicly available database, CASIA-OLHWDB1, for research on online handwritten Chinese character recognition. This database is the first of our series of online/offline handwritten characters and texts, collected using Anoto pen on paper. It contains unconstrained handwritten characters of 4,037 categories (3,866 Chinese characters and 171 symbols) produced by 420 persons, and 1,694,741 samples in total. It can be used for design and evaluation of character recognition algorithms and classifier design for handwritten text recognition systems. We have partitioned the samples into three grades and into training and test sets. Preliminary experiments on the database using a state-of-the-art recognizer justify the challenge of recognition. Dahan Wang, Cheng-Lin Liu 0001, Jin-Lun Yu |
ICDAR | 2 |
| 2009 | Integrating Language Model in Handwritten Chinese Text RecognitionabstractThis paper describes a system for handwritten Chinese text recognition integrating language model. On a text line image, the system generates character segmentation and word segmentation candidates, and the candidate paths are evaluated by character recognition scores and language model. The optimal path, giving segmentation and recognition result, is found using a pruned dynamic programming search method. We evaluate various language models, including the character-based n-gram, word-based n-gram, and hybrid n-gram models. Experimental results on the HIT-HW database show that the language models improve the recognition performance remarkably. Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
ICDAR | 3 |
| 2009 | A Variational Bayes Method for Handwritten Text Line SegmentationabstractText line segmentation in unconstrained handwritten documents remains a challenge because handwritten text lines are multi-skewed and not obviously separated. This paper presents a new approach based on the variational Bayes (VB) framework for text line segmentation. Viewing the document image as a mixture density model, with each text line approximated by a Gaussian component, the VB method can automatically determine the number of components. We extend the VB method such that it can both eliminate and split components and control the orientation of text line lines. Experiments on Chinese handwritten documents demonstrated the effectiveness of the approach. Cheng-Lin Liu 0001 |
ICDAR | 2 |
| 2009 | A Tool for Ground-Truthing Text Lines and Characters in Off-Line Handwritten Chinese DocumentsabstractAnnotating the regions, text lines and characters of document images is an important, but tedious and expensive task. A ground-truthing tool may largely alleviate the human burden in this process. This paper describes an automated recognition-based tool GTLC for finding the best alignment between the text transcript and the connected components of unconstrained handwritten document image. The alignment process is formulated as an optimization problem involving candidate character segmentation and recognition. We have validated the effectiveness of this tool and have used it for annotating a large number of handwritten Chinese documents. Qiufeng Wang 0001, Cheng-Lin Liu 0001 |
ICDAR | 3 |
| 2009 | Object tracking by bidirectional learning with feature selectionabstractThis paper proposes a new tracking algorithm which combines object and background information, via building object and background appearance models simultaneously by non-parametric kernel density estimation. The major contribution is a novel bidirectional learning framework for discrimination between the object and background. It has the following advantages: 1) it embeds background information, unlike most other methods that focus on the object only, 2) it provides a mechanism to detect occlusion and distraction, which are two main causes of tracking failure, 3) it performs feature selection, making the tracker more robust to outliers. By this learning framework, we are able to embed discriminative information into the generative appearance model. Experimental results demonstrate that the tracker is able to model drastic appearance changes and robust to occlusion and distraction. Xinwen Hou, Cheng-Lin Liu 0001 |
ICIP | 3 |
| 2009 | Subspace Regularization: A New Semi-supervised Learning Method
Yan-Ming Zhang 0001, Xinwen Hou, Shiming Xiang, Cheng-Lin Liu 0001 |
ECML/PKDD (2) | 4 |
| 2009 | A new benchmark on the recognition of handwritten Bangla and Farsi numeral characters
Cheng-Lin Liu 0001, Ching Y. Suen |
Pattern Recognit. | 1 |
| 2009 | Handwritten Chinese text line segmentation by clustering with distance metric learning
Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2009 | A robust approach to text line grouping in online handwritten Japanese documents
Dahan Wang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2008 | A Robust System to Detect and Localize Texts in Natural Scene ImagesabstractIn this paper, we present a robust system to accurately detect and localize texts in natural scene images. For text detection, a region-based method utilizing multiple features and cascade AdaBoost classifier is adopted. For text localization, a window grouping method integrating text line competition analysis is used to generate text lines. Then within each text line, local binarization is used to extract candidate connected components (CCs) and non-text CCs are filtered out by Markov Random Fields (MRF) model, through which text line can be localized accurately. Experiments on the public benchmark ICDAR 2003 Robust Reading and Text Locating Dataset show that our system is comparable to the best existing methods both in accuracy and speed. Yi-Feng Pan, Xinwen Hou, Cheng-Lin Liu 0001 |
Document Analysis Systems | 3 |
| 2008 | Grouping Text Lines in Online Handwritten Japanese Documents by Combining Temporal and Spatial InformationabstractWe present an effective approach for grouping text lines in online handwritten Japanese documents by combining temporal and spatial information. Initially, strokes are grouped into text line strings according to off-stroke distances. Each text line string is segmented into text lines by dynamic programming (DP) optimizing a cost function trained by the minimum classification error (MCE) method. Over-segmented text lines are then merged with a support vector machine (SVM) classifier for making merge/non-merge decisions, and last, a spatial merge module corrects the segmentation errors caused by delayed strokes. In experiments on the TUAT Kondate database, the proposed approach achieves the Entity Detection Metric (EDM) rate of 0.8816, the Edit-Distance Rate (EDR) of 0.1234, which demonstrates the superiority of our approach. Dahan Wang, Cheng-Lin Liu 0001 |
Document Analysis Systems | 3 |
| 2008 | Prototype learning with margin-based conditional log-likelihood lossabstractThe classification performance of nearest prototype classifiers largely relies on the prototype learning algorithms, such as the learning vector quantization (LVQ) and the minimum classification error (MCE). This paper proposes a new prototype learning algorithm based on the minimization of a conditional log-likelihood loss (CLL), called log-likelihood of margin (LOGM). A regularization term is added to avoid over-fitting in training. The CLL loss in LOGM is a convex function of margin, and so, gives better convergence than the MCE algorithm. Our empirical study on a large suite of benchmark datasets demonstrates that the proposed algorithm yields higher accuracies than the MCE, the generalized LVQ (GLVQ), and the soft nearest prototype classifier (SNPC). Xiao-Bo Jin, Cheng-Lin Liu 0001, Xinwen Hou |
ICPR | 2 |
| 2008 | A pooled subspace mixture density model for pattern classification in high-dimensional spacesabstractDensity estimation in high-dimensional data spaces is a challenge due to the sparseness of data which is known as ldquothe curse of dimensionalityrdquo. Researchers often resort to low-dimensional subspaces for such tasks, while discard the distribution in the complementary subspace. In this paper, we propose a new mixture density model based on pooled subspace. In our method, the Gaussian components of each class share a subspace and the complementary subspace is incorporated in the density function. The subspace and Gaussian mixture density are estimated simultaneously in EM iteration steps. We apply the density model to pattern classification in experiments on UCI datasets and compare the proposed method with previous ones. The experimental results demonstrate the superiority of the proposed method. Xiao-Hua Liu, Cheng-Lin Liu 0001, Xinwen Hou |
IJCNN | 2 |
| 2006 | Learning Boosted Asymmetric Classifiers for Object DetectionabstractObject detection can be posted as those classification tasks where the rare positive patterns are to be distinguished from the enormous negative patterns. To avoid the danger of missing positive patterns, more attention should be payed on them. Therefore there should be different requirements for False Reject Rate (FRR) and False Accept Rate (FAR) , and learning a classifier should use an asymmetric factor to balance between FRR and FAR. In this paper, a normalized asymmetric classification error is proposed for the task of rejecting negative patterns. Minimizing it not only controls the ratio of FRR and FAR, but more importantly limits the upper-bound of FRR. The latter characteristic is advantageous for those tasks where there is a requirement for low FRR. Based on this normalized asymmetric classification error, we develop an asymmetric AdaBoost algorithm with variable asymmetric factor and apply it to the learning of cascade classifiers for face detection. Experiments demonstrate that the proposed method achieves less complex classifiers and better performance than some previous AdaBoost methods. Xinwen Hou, Cheng-Lin Liu 0001, Tieniu Tan |
CVPR (1) | 2 |