VLDB 2026 Research / reviewers in the wild / expert
Xu-Yao Zhang
dblp:58/9924 · also XuYao Zhang
· DBLP profile ↗
145ranked-venue papers
14as first author
84since 2021 · last 2026
0000-0001-9260-188XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 117 · 9 first-author · 69 since 2021Graphics, computer vision, multimedia, augmented reality and games · 62 · 4 first-author · 36 since 2021Databases, data management, data science and information retrieval · 14 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and UnderstandingabstractFor video anomaly detection, it's both important to detect when the event happens and what the event is. The tasks of temporal grounding and semantic understanding can benefit from joint learning, but no existing work support it. To address this problem, we introduce VAGU (Video Anomaly Grounding and Understanding), the first benchmark designed to jointly evaluate semantic understanding and precise temporal grounding of anomalies, with comprehensive annotations and objective multiple-choice Video QA. Besides, we propose Glance then Scrutinize (GtS), the first training-free framework that achieves the best balance performance in both accuracy and efficiency. GtS uniquely balances high temporal precision and semantic interpretability while meeting practical speed requirements, outperforming previous methods in real-world scenarios. Furthermore, we introduce the JeAUG metric for holistic evaluation of both speed and accuracy. Extensive experiments demonstrate the superior effectiveness and practicality of our benchmark, framework, and metric. Shibo Gao, Peipei Yang, Yi Chen 0027, Xu-Yao Zhang |
AAAI | 6 |
| 2026 | LBLLM: Lightweight Binarization of Large Language Models via Three-Stage DistillationabstractDeploying large language models (LLMs) in resource-constrained environments is hindered by heavy computational and memory requirements.We present LBLLM, a lightweight binarization framework that achieves effective W(1+1)A4 quantization through a novel threestage quantization strategy.The framework proceeds as follows: (1) initialize a high-quality quantized model via PTQ; (2) quantize binarized weights, group-wise bitmaps, and quantization parameters through layer-wise distillation while keeping activations in full precision; and (3) training learnable activation quantization factors to dynamically quantize activations to 4 bits.This decoupled design mitigates interference between weight and activation quantization, yielding greater training stability and better inference accuracy.LBLLM, trained only using 0.016B tokens with a single GPU, surpasses existing state-of-the-art binarization methods on W2A4 quantization settings across tasks of language modeling, commonsense QA, and language understanding.These results demonstrate that extreme low-bit quantization of LLMs can be both practical and highly effective without introducing any extra high-precision channels nor rotational matrices commonly used in recent PTQ-based works, offering a promising path toward efficient LLM deployment on resource-limited situations. Siqing Song, Chuang Wang 0007, Yong Lang, Xu-Yao Zhang |
ACL (1) | 5 |
| 2026 | Gradient Guided LoRA for Stable Fine-Tuning of LLMs
Peipei Yang, Hongjian Fang, Xu-Yao Zhang |
ICPR (12) | 4 |
| 2026 | Diverse feature generation for zero-shot Chinese character recognition
Song-Liang Pan, Kunchi Li, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu |
Expert Syst. Appl. | 4 |
| 2026 | Bayesian classifier calibration based on synthesized samples for zero-shot Chinese character recognition
Xiang Ao 0002, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2026 | Learning relationship-guided vision-language transformer for facial attribute recognition
Si Chen 0002, Mingxuan Lei, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001, Shunzhi Zhu |
Pattern Recognit. | 4 |
| 2026 | DUKAE: DUal-level Knowledge Accumulation and Ensemble for pre-trained model-based continual learning
Tonghua Su, Xu-Yao Zhang, Qixing Xu, Zhongjie Wang 0003 |
Pattern Recognit. | 3 |
| 2026 | MPRSurv: Multi-perspective prompted ranking for vision-language survival analysis on whole slide images
Ruofan Zhang, Mengjie Fang, Shaoli Zhao, Zipei Wang, Xin Feng 0010, Xu-Yao Zhang, Xuebin Xie, Jie Tian 0001, Di Dong |
Pattern Recognit. | 8 |
| 2026 | AFH-Net: An adaptive feature harmonization network for document image De-warping
Xinyue Zhou, Nanfeng Jiang, Wang Man, Xu-Yao Zhang, Shunzhou Wang, Dahan Wang |
Pattern Recognit. | 5 |
| 2026 | Video-Level Cross-Modal Temporal-Navigation for RGBT TrackingabstractRGBT tracking has recently garnered significant attention due to its all-weather tracking capability. Traditional RGBT tracking methods primarily concentrate on the fusion of cross-modal spatial information. However, these methods ignore contextual relationships between consecutive video frames and lack effective interactions between modalities, easily resulting in tracking drift gradually due to the accumulation of errors. To avoid this limitation, we propose a novel Video-Level Cross-Modal Temporal-Navigation method termed VCT for robust RGBT tracking, which fully leverages the complementary spatio-temporal information across modalities to improve cross-modal tracking accuracy. The VCT employs a simple, flexible, and effective video-level dual-stream architecture that accommodates video sequences of arbitrary length, enabling RGB and TIR streams to capture and synergize spatio-temporal features across frames. To achieve temporal consistency and adaptability, we design a Cross-Modal Temporal Prompt Navigator (CM-TPN) that dynamically aggregates and compresses the historical frame context to navigate predictions of subsequent frames through temporal prompts. In addition, we introduce a Modality-Specific Mixture of Adapters (MS-MoA) to promote the spatio-temporal interaction both within and between modalities, thereby dramatically adapting to appearance changes. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the four popular RGBT tracking benchmarks. Si Chen 0002, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text InformationabstractWith the advancement of large-scale language modeling techniques, large multimodal models combining visual encoders with large language models have demonstrated exceptional performance in various visual tasks. Most of the current large multimodal models achieve this by mapping visual features obtained from the visual encoder into a large language model and using them as inputs alongside text for downstream tasks. Therefore, the number of visual tokens directly affects the training and inference speed of the model. There has been significant work on token pruning for visual transformers, but for large multimodal models, only relying on visual information for token pruning or compression may lead to significant loss of important information. On the other hand, the textual input in the form of a question may contain valuable information that can aid in answering the question, providing additional knowledge to the model. To address the potential oversimplification and excessive pruning that can occur with most purely visual token pruning methods, we propose a text information-guided dynamic visual token recovery mechanism that does not require training. This mechanism leverages the similarity between the question text and visual tokens to recover visually meaningful tokens with important text information while merging other less important tokens, to achieve efficient computation for large multimodal models. Experimental results demonstrate that our proposed method achieves comparable performance to the original approach while compressing the visual tokens to an average of 10\% of the original quantity. Yi Chen 0027, Xu-Yao Zhang, Wen-Zhuo Liu, Yang-Yang Liu |
AAAI | 3 |
| 2025 | HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language ModelabstractHaiyang Guo, Fanhu Zeng, Ziwei Xiang, Fei Zhu, Da-Han Wang, Xu-Yao Zhang, Cheng-Lin Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haiyang Guo, Fanhu Zeng, Ziwei Xiang, Fei Zhu 0004, Dahan Wang, Xu-Yao Zhang |
ACL (1) | 6 |
| 2025 | ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided PromptabstractLarge Multimodal Models (LMMs) exhibit remarkable multi-tasking ability by learning mixed instruction datasets.However, novel tasks would be encountered sequentially in dynamic world, which urges for equipping LMMs with multimodal continual instruction learning (MCIT) ability especially for diverse and challenging generative tasks.Existing MCIT methods do not fully exploit the unique attribute of LMMs and often gain performance at the expense of efficiency.In this paper, we propose a novel prompt learning framework for MCIT to effectively alleviate forgetting of previous knowledge while managing computational complexity with natural image-text supervision.Concretely, we learn prompts for each task and exploit efficient prompt fusion for knowledge transfer and prompt selection for complexity management with dual-modality guidance.Extensive experiments demonstrate that our approach achieves substantial +14.26% performance gain on MCIT benchmarks with remarkable ×1.42 inference speed free from growing computation. Fanhu Zeng, Fei Zhu 0004, Haiyang Guo, Xu-Yao Zhang |
EMNLP | 4 |
| 2025 | Multi-Task Model Fusion via Adaptive MergingabstractMulti-task model fusion (MTMF) aims to integrate the capabilities of individual models into a unified model. Past approaches either require extensive training and fine-tuning or necessitate that models share the same pre-training and initialization. Recently, several fusion methods have been proposed that do not require extensive training or fine-tuning. These methods can merge multiple independently trained models with different task capabilities into a single multi-task model without increasing the number of parameters. In this work, we identify a common flaw in these fusion methods: they tend to focus on how well the modules of the individual models match before merging while neglecting the representation bias of the merged model. To address this problem, we propose a simple yet effective mitigation method called adaptive merging by representation alignment (AdMbRA). Specifically, we improve the method of weight matching by using representation bias as a constraint and optimize the merging process. Experiments demonstrate that our method can effectively mitigate the representation bias of the merged model, thus improving the performance of each task. Ziwei Xiang, Kai Lei, Xu-Yao Zhang |
ICASSP | 4 |
| 2025 | Federated Continual Instruction Tuning
Haiyang Guo, Fanhu Zeng, Fei Zhu 0004, Wenzhuo Liu, Dahan Wang, Jian Xu 0015, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICCV | 7 |
| 2025 | OracleGCD: Generalized Category Discovery for Oracle Bone Scripts
Hetao Wu, Kunchi Li, Xu-Yao Zhang, Dahan Wang |
ICDAR (3) | 3 |
| 2025 | Local-Prompt: Extensible Local Prompts for Few-Shot Out-of-Distribution DetectionabstractOut-of-Distribution (OOD) detection, aiming to distinguish outliers from known categories, has gained prominence in practical scenarios. Recently, the advent of vision-language models (VLM) has heightened interest in enhancing OOD detection for VLM through few-shot tuning. However, existing methods mainly focus on optimizing global prompts, ignoring refined utilization of local information with regard to outliers. Motivated by this, we freeze global prompts and introduce Local-Prompt, a novel coarse-to-fine tuning paradigm to emphasize regional enhancement with local prompts. Our method comprises two integral components: global prompt guided negative augmentation and local prompt enhanced regional regularization. The former utilizes frozen, coarse global prompts as guiding cues to incorporate negative augmentation, thereby leveraging local outlier knowledge. The latter employs trainable local prompts and a regional regularization to capture local information effectively, aiding in outlier identification. We also propose regional-related metric to empower the enrichment of OOD detection. Moreover, since our approach explores enhancing local prompts only, it can be seamlessly integrated with trained global prompts during inference to boost the performance. Comprehensive experiments demonstrate the effectiveness and potential of our method. Notably, our method reduces average FPR95 by 5.17% against state-of-the-art method in 4-shot tuning on challenging ImageNet-1k dataset, even outperforming 16-shot results of previous methods. Fanhu Zeng, Zhen Cheng 0003, Fei Zhu 0004, Hongxin Wei, Xu-Yao Zhang |
ICLR | 5 |
| 2025 | MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language ModelsabstractHallucinations pose a significant challenge in Large Vision Language Models (LVLMs), with misalignment between multimodal features identified as a key contributing factor. This paper reveals the negative impact of the long-term decay in Rotary Position Encoding (RoPE), used for positional modeling in LVLMs, on multimodal alignment. Concretely, under long-term decay, instruction tokens exhibit uneven perception of image tokens located at different positions within the two-dimensional space: prioritizing image tokens from the bottom-right region since in the one-dimensional sequence, these tokens are positionally closer to the instruction tokens. This biased perception leads to insufficient image-instruction interaction and suboptimal multimodal alignment. We refer to this phenomenon as ''image alignment bias.'' To enhance instruction's perception of image tokens at different spatial locations, we propose MCA-LLaVA, based on Manhattan distance, which extends the long-term decay to a two-dimensional, multi-directional spatial decay. MCA-LLaVA integrates the one-dimensional sequence order and two-dimensional spatial position of image tokens for positional modeling, mitigating hallucinations by alleviating image alignment bias. Experimental results of MCA-LLaVA across various hallucination and general benchmarks demonstrate its effectiveness and generality. The code can be accessed in https://github.com/ErikZ719/MCA-LLaVA. Qiyan Zhao, Xiaofeng Zhang 0006, Yun Xing 0001, Xiaosong Yuan, Sinan Fan, Xuhang Chen 0002, Dahan Wang, Xu-Yao Zhang |
ACM Multimedia | 10 |
| 2025 | Low-light image enhancement with quality-oriented pseudo labels via semi-supervised contrastive learning
Nanfeng Jiang, Yiwen Cao, Xu-Yao Zhang, Dahan Wang, Chiming Wang, Shunzhi Zhu |
Expert Syst. Appl. | 3 |
| 2025 | Breaking the Limits of Reliable Prediction via Generated Data
Zhen Cheng 0003, Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | ADR-Net: Attention-oriented detail recovery network for document image shadow removal
Nanfeng Jiang, Dahan Wang, Xu-Yao Zhang, Yun Wu 0001, Shunzhi Zhu |
Knowl. Based Syst. | 4 |
| 2025 | ProtoGCD: Unified and Unbiased Prototype Learning for Generalized Category DiscoveryabstractGeneralized category discovery (GCD) is a pragmatic but underexplored problem, which requires models to automatically cluster and discover novel categories by leveraging the labeled samples from old classes. The challenge is that unlabeled data contain both old and new classes. Early works leveraging pseudo-labeling with parametric classifiers handle old and new classes separately, which brings about imbalanced accuracy between them. Recent methods employing contrastive learning neglect potential positives and are decoupled from the clustering objective, leading to biased representations and sub-optimal results. To address these issues, we introduce a unified and unbiased prototype learning framework, namely ProtoGCD, wherein old and new classes are modeled with joint prototypes and unified learning objectives, enabling unified modeling between old and new classes. Specifically, we propose a dual-level adaptive pseudo-labeling mechanism to mitigate confirmation bias, together with two regularization terms to collectively help learn more suitable representations for GCD. Moreover, for practical considerations, we devise a criterion to estimate the number of new classes. Furthermore, we extend ProtoGCD to detect unseen outliers, achieving task-level unification. Comprehensive experiments show that ProtoGCD achieves state-of-the-art performance on both generic and fine-grained datasets. Shijie Ma, Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | PASS++: A Dual Bias Reduction Framework for Non-Exemplar Class-Incremental LearningabstractClass-incremental learning (CIL) aims to continually recognize new classes while preserving the discriminability of previously learned ones. Most existing CIL methods are exemplar-based, relying on the storage and replay of a subset of old data during training. Without access to such data, these methods typically suffer from catastrophic forgetting. In this paper, we identify two fundamental causes of forgetting in CIL: representation bias and classifier bias. To address these challenges, we propose a simple yet effective dual-bias reduction framework, which leverages self-supervised transformation (SST) in the input space and prototype augmentation (protoAug) in the feature space. On one hand, SST mitigates representation bias by encouraging the model to learn generic, diverse representations that generalize across tasks. On the other hand, protoAug tackles classifier bias by explicitly or implicitly augmenting the prototypes of old classes in the feature space, thereby imposing stronger constraints to preserve decision boundaries. We further enhance the framework with hardness-aware prototype augmentation and multi-view ensemble strategies, yielding significant performance gains. The proposed framework can be easily integrated with pre-trained models. Without storing any samples of old classes, our method performs comparably to state-of-the-art exemplar-based approaches that rely on extensive data storage. We hope to draw the attention of researchers back to non-exemplar CIL by rethinking the necessity of storing old samples. Fei Zhu 0004, Xu-Yao Zhang, Zhen Cheng 0003, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Joint radical embedding and detection for zero-shot Chinese character recognition
Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu |
Pattern Recognit. | 3 |
| 2025 | Towards trustworthy dataset distillation
Shijie Ma, Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang |
Pattern Recognit. | 4 |
| 2025 | Stain-adaptive self-supervised learning for histopathology image analysis
Haili Ye, Shunzhi Zhu, Dahan Wang, Xu-Yao Zhang, Heguang Huang |
Pattern Recognit. | 5 |
| 2025 | Towards reliable domain generalization: Insights from the PF2HC benchmark and dynamic evaluations
Xiang Ao 0002, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2025 | Average of Pruning: Improving Performance and Stability of Out-of-Distribution DetectionabstractDetecting out-of-distribution (OOD) inputs has been a critical issue for neural networks in the open world. However, the unstable behavior of OOD detection along the optimization trajectory during training has not been explored clearly. In this article, we first find the performance of OOD detection suffers from overfitting and instability during training: 1) the performance could decrease when the training error is near zero and 2) the performance would vary sharply in the final stage of training. Based on our findings, we propose an average of pruning (AoP), consisting of model averaging (MA) and pruning, to mitigate the unstable behaviors. Specifically, MA can help achieve a stable performance by smoothing the landscape, and pruning is theoretically and empirically verified to eliminate overfitting by avoiding redundant features. Comprehensive experiments on various datasets and architectures are conducted to verify the effectiveness of our method. Zhen Cheng 0003, Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Continual Learning With Knowledge Distillation: A SurveyabstractThe foremost challenge in continual learning is to mitigate catastrophic forgetting, allowing a model to retain knowledge of previous tasks while learning new tasks. Knowledge distillation (KD), a form of regularization, has gained significant attention for its ability to maintain a model's performance on previous tasks by mimicking the outputs of earlier models during the learning of new tasks, thus reducing forgetting. This article offers a comprehensive survey of continual learning methods employing KD within the realm of image classification. We provide a detailed analysis of how KD is utilized in continual learning methods, categorizing its application into three distinct paradigms. Besides, we classify these methods based on the type of knowledge source used and thoroughly examine how KD consolidates memory in continual learning from the perspective of loss functions. In addition, we have conducted extensive experiments on CIFAR-100, TinyImageNet, and ImageNet-100 across ten KD-integrated continual learning methods to analyze the role of KD in continual learning, and we have further discussed its effectiveness in other continual learning tasks. Our extensive experimental evidence demonstrates that KD plays a crucial role in mitigating forgetting in continual learning and substantiates that, when used with data replay, classification bias adversely affects the effectiveness of KD, whereas employing a separated softmax loss can significantly enhance its efficacy. Tonghua Su, Xu-Yao Zhang, Zhongjie Wang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Unified Entropy Optimization for Open-Set Test-Time AdaptationabstractTest-time adaptation (TTA) aims at adapting a model pretrained on the labeled source domain to the unlabeled target domain. Existing methods usually focus on improving TTA performance under covariate shifts, while neglecting semantic shifts. In this paper, we delve into a realistic open-set TTA setting where the target domain may contain samples from unknown classes. Many state-of-the-art closed-set TTA methods perform poorly when applied to open-set scenarios, which can be attributed to the inaccurate estimation of data distribution and model confidence. To address these issues, we propose a simple but effective framework called unified entropy optimization (UniEnt), which is capable of simultaneously adapting to covariate-shifted in-distribution (csID) data and detecting covariate-shifted out-of-distribution (csOOD) data. Specifically, UniEnt first mines pseudo-csID and pseudo-csOOD samples from test data, followed by entropy min-imization on the pseudo-csID data and entropy maximization on the pseudo-csOOD data. Furthermore, we introduce UniEnt+ to alleviate the noise caused by hard data partition leveraging sample-level confidence. Extensive experiments on CIFAR benchmarks and Tiny-ImageNet-C show the superiority of our framework. The code is available at https://github.com/gaozhengqing/UniEnt. Zhengqing Gao, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 2 |
| 2024 | Active Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) is a pragmatic and challenging open-world task, which endeavors to cluster unlabeled samples from both novel and old classes, leveraging some labeled data of old classes. Given that knowledge learned from old classes is not fully transferable to new classes, and that novel categories are fully unlabeled, GCD inherently faces intractable problems, including imbalanced classification performance and inconsistent confidence between old and new classes, especially in the low-labeling regime. Hence, some annotations of new classes are deemed necessary. However, labeling new classes is extremely costly. To address this issue, we take the spirit of active learning and propose a new setting called Active Generalized Category Discovery (AGCD). The goal is to improve the performance of GCD by actively selecting a limited amount of valuable samples for labeling from the oracle. To solve this problem, we devise an adaptive sampling strategy, which jointly considers novelty, informativeness and diversity to adaptively select novel samples with proper uncertainty. However, owing to the varied orderings of label indices caused by the clustering of novel classes, the queried labels are not directly applicable to subsequent training. To overcome this issue, we further propose a stable label mapping algorithm that transforms ground truth labels to the label space of the classifier, thereby ensuring consistent training across different active selection stages. Our method achieves state-of-the-art performance on both generic and fine-grained datasets. Our code is available at https://github.com/mashijie1028/ActiveGCD Shijie Ma, Fei Zhu 0004, Zhun Zhong, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 4 |
| 2024 | RCL: Reliable Continual Learning for Unified Failure DetectionabstractDeep neural networks are known to be overconfident for what they don't know in the wild, which is undesirable for decision-making in high-stakes applications. Despite quan-tities of existing works, most of them focus on detecting out-of-distribution (OOD) samples from unseen classes, while ignoring large parts of relevant failure sources like mis-classified samples from known classes. In particular, recent studies reveal that prevalent OOD detection methods are actually harmful for misclassification detection (MisD), indicating that there seems to be a tradeoff between those two tasks. In this paper, we study the critical yet under-explored problem of unified failure detection, which aims to detect both misclassified and OOD examples. Concretely, we identify the failure of simply integrating learning objectives of misclassification and OOD detection, and show the potential of sequence learning. Inspired by this, we propose a reliable continual learning paradigm, whose spirit is to equip the model with MisD ability first, and then improve the OOD detection ability without degrading the al-ready adequate MisD performance. Extensive experiments demonstrate that our method achieves strong unified failure detection performance. The code is available at https://github.com/Impression2805/RCL. Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang, Cheng-Lin Liu 0001, Zhaoxiang Zhang 0001 |
CVPR | 3 |
| 2024 | PILoRA: Prototype Guided Incremental LoRA for Federated Class-Incremental Learning
Haiyang Guo, Fei Zhu 0004, Wenzhuo Liu, Xu-Yao Zhang |
ECCV (65) | 4 |
| 2024 | Prototype Calibration with Synthesized Samples for Zero-Shot Chinese Character RecognitionabstractZero-shot Chinese character recognition aims to recognize unseen characters that have never appeared in training. Recently, many methods learn a cross-modal alignment between character samples and auxiliary semantic data like glyph templates in training, and directly employ it to recognize unseen characters by retrieving the class with most similar semantics. However, these approaches suffer from the domain shift problem, which means that the learned alignment shows a deviation on unseen characters. To alleviate this problem, we generate unseen character samples to calibrate the shifted prototypes in the feature space. Specifically, we train a cross-modal prototype classifier and a generator conditioned on glyph templates, then use the generator to synthesize unseen character samples to calibrate the prototypes of the classifier. The calibration process does not require any extra training. Experiments on a handwritten dataset and a nature scene dataset show the superiority of our method and the effectiveness of prototype calibration. Xiang Ao 0002, Xiao-Hui Li 0012, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICASSP | 3 |
| 2024 | Cross-Domain Few-Shot Learning with Equiangular Embedding and Dynamic Adversarial Augmentation
Ziwei Xiang, Kai Lei, Xu-Yao Zhang |
ICONIP (3) | 4 |
| 2024 | UAD-DPL: An Unknown Encrypted Attack Detection Method Based on Deep Prototype Learning
Liangchen Chen, Shu Gao, Baoxu Liu, Xu-Yao Zhang |
ICPR (5) | 4 |
| 2024 | ENS-RFMC: An Encrypted Network Traffic Sampling Method Based on Rule-Based Feature Extraction and Multi-hierarchical Clustering for Intrusion Detection
Liangchen Chen, Shu Gao, Zixuan Wei, Baoxu Liu, Xu-Yao Zhang |
ICPR (24) | 5 |
| 2024 | Delving into Feature Space: Improving Adversarial Robustness by Feature Spectral Regularization
Fei Zhu 0004, Xu-Yao Zhang |
ICPR (26) | 3 |
| 2024 | OCR4HSV: A Multi-task Learning Approach for Handwritten Signature Verification
Chao-Qun Lin, Dahan Wang, Yanfei Su, De-Wu Ge, Xu-Yao Zhang |
ICPR (31) | 5 |
| 2024 | SANS: Spatial-Aware Neural Solver for Plane Geometry Problem
Shun-Xin Xiao, Zirong Chen, Dahan Wang, Xu-Yao Zhang |
ICPR (31) | 6 |
| 2024 | Learning Explicit Radical Representations for Zero-Shot Chinese Character Recognition
Song-Liang Pan, Dahan Wang, Nanfeng Jiang, Xu-Yao Zhang, Shunzhi Zhu |
ICPR (31) | 4 |
| 2024 | Document Image Shadow Removal via Frequency Information-Oriented Network
Xinyue Zhou, Nanfeng Jiang, Dahan Wang, Xu-Yao Zhang, Guantin Li, Wang Man, Yun Wu 0001 |
ICPR (31) | 5 |
| 2024 | DocHFormer: Document Image Dewarping via Harmonized Modeling of Hierarchical Priors
Xinyue Zhou, Guanting Li, Nanfeng Jiang, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu |
ICPR (31) | 5 |
| 2024 | RCFormer : Interactive Image Segmentation via Reconstructing Click Vision TransformersabstractClick-based interactive image segmentation intends to segment an object from the background under user click guidance. Recently, Vision Transformer has made significant strides in interactive image segmentation. However, the previous studies 1) overlook the importance of different clicks in terms of their contribution to the segmentation results; and 2) suffer from inconsistency across different feature scales in the multi-scale structure. In this paper, we propose a new interactive segmentation framework, named RCFormer, with two novel components: reconstruct click patch embedding (RCPE) for encoding the importance of clicks, and multi-scale adaptive fusion (MSAF) for the adaptive fusion of feature maps across different scales. RCPE enhances the effectiveness of click interactions by spatially distinguishing the importance of clicks. MSAF adaptively fuses useful spatial information and filters the redundant feature at multi-scales. The experiments on several benchmarks show that our proposed approach achieves state-of-the-art performance. Notably, our method achieves 2.31 NoC@90 on the Berkeley dataset, improving by 8.6% over the previous best results. PanPan Chen, Dahan Wang, Yun Wu 0001, Xu-Yao Zhang, Shunzhi Zhu |
IJCNN | 6 |
| 2024 | Character Relationship Refinement Network for Handwritten Mathematical Expression RecognitionabstractMost current Handwritten Mathematical Expression Recognition (HMER) methods employ an attention-based encoder-decoder framework, which generates LaTeX sequences from the given images, following the paradigm of predicting "one-by-one". However, this paradigm may have some challenges: 1) without considering the connectivity between characters, the prior information in the prediction process will be ignored inadvertently, especially implicit information, such as " " and " ˆ ". 2) Some characters of high similarities, such as "6/b" and "o/O", will have negative effects on prediction results. To solve these issues, we propose a simple but effective Character Relationship Refinement Network (CRRN), which consists of Joint Character Learning (JCL) and Character Refinement Mask (CRM). Specifically, JCL calculates the relationship probability between characters and uses them to improve prediction accuracy. CRM takes the character confidence coefficient in a coarse-to-fine way that can reassign the weights of all characters to improve model discriminability on easily confused characters. With the collaboration of both modules, our proposed CRRN can outperform the state-of-the-art on popular datasets. LiWei Jiang, Nanfeng Jiang, Yun Wu 0001, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu |
IJCNN | 5 |
| 2024 | Happy: A Debiased Learning Framework for Continual Generalized Category DiscoveryabstractConstantly discovering novel concepts is crucial in evolving environments. This paper explores the underexplored task of Continual Generalized Category Discovery (C-GCD), which aims to incrementally discover new classes from *unlabeled* data while maintaining the ability to recognize previously learned classes. Although several settings are proposed to study the C-GCD task, they have limitations that do not reflect real-world scenarios. We thus study a more practical C-GCD setting, which includes more new classes to be discovered over a longer period, without storing samples of past classes. In C-GCD, the model is initially trained on labeled data of known classes, followed by multiple incremental stages where the model is fed with unlabeled data containing both old and new classes. The core challenge involves two conflicting objectives: discover new classes and prevent forgetting old ones. We delve into the conflicts and identify that models are susceptible to *prediction bias* and *hardness bias*. To address these issues, we introduce a debiased learning framework, namely **Happy**, characterized by **H**ardness-**a**ware **p**rototype sampling and soft entro**py** regularization. For the *prediction bias*, we first introduce clustering-guided initialization to provide robust features. In addition, we propose soft entropy regularization to assign appropriate probabilities to new classes, which can significantly enhance the clustering performance of new classes. For the *harness bias*, we present the hardness-aware prototype sampling, which can effectively reduce the forgetting issue for previously seen classes, especially for difficult classes. Experimental results demonstrate our method proficiently manages the conflicts of C-GCD and achieves remarkable performance across various datasets, e.g., 7.5% overall gains on ImageNet-100. Our code is publicly available at https://github.com/mashijie1028/Happy-CGCD. Shijie Ma, Fei Zhu 0004, Zhun Zhong, Wenzhuo Liu, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
NeurIPS | 5 |
| 2024 | Adapting Vision-Language Models to Open Classes via Test-Time Prompt Tuning
Zhengqing Gao, Xiang Ao 0002, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
PRCV (5) | 3 |
| 2024 | Ensemble Quadratic Assignment Network for Graph Matching
Haoru Tan, Chuang Wang 0007, Sitong Wu, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Revisiting Confidence Estimation: Towards Reliable Failure PredictionabstractReliable confidence estimation is a challenging yet fundamental requirement in many risk-sensitive applications. However, modern deep neural networks are often overconfident for their incorrect predictions, i.e., misclassified samples from known classes, and out-of-distribution (OOD) samples from unknown classes. In recent years, many confidence calibration and OOD detection methods have been developed. In this paper, we find a general, widely existing but actually-neglected phenomenon that most confidence estimation methods are harmful for detecting misclassification errors. We investigate this problem and reveal that popular calibration and OOD detection methods often lead to worse confidence separation between correctly classified and misclassified examples, making it difficult to decide whether to trust a prediction or not. Finally, we propose to enlarge the confidence gap by finding flat minima, which yields state-of-the-art failure prediction performance under various settings including balanced, long-tailed, and covariate-shift classification scenarios. Our study not only provides a strong baseline for reliable confidence estimation but also acts as a bridge between understanding calibration, OOD detection, and failure prediction. Fei Zhu 0004, Xu-Yao Zhang, Zhen Cheng 0003, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Large-scale continual learning for ancient Chinese character recognition
Xu-Yao Zhang, Zhaoxiang Zhang 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2024 | HASI: Hierarchical Attention-Aware Spatio-Temporal Interaction for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (re-ID) aims to match the same pedestrian of video sequences across non-overlapping cameras. Video re-ID methods generally adopt frame-level feature extraction for different video frames, but they still lack effective spatio-temporal interaction, easily leading to the multi-frame misalignment problem. In this paper, we propose a Hierarchical Attention-aware Spatio-temporal Interaction (HASI) network, including an Attention-aware Temporal Interaction (ATI) module and a Hierarchical Local-spatial Enhancement (HLE) module for video-based person re-ID. In order to avoid the spatial misalignment between video frames, the ATI module employs multiple Frame-to-Frame Temporal Interaction (2FTI) blocks with the Multi-head Inter-frame Alignment Attention (MIAA) to make the current frame iteratively interact with each rest frame of a video in a positive single-cycle manner, rather than only interacting with the adjacent frame or directly building the relationship of all frames at once. This module can not only obtain the long-range non-adjacent temporal information, but also learn the pairwise frame-to-frame relationships. Moreover, the HLE module is designed to enhance the local fine-grained features from multiple Transformer layers, whilst delivering low-level information to further enrich middle-level and high-level semantic knowledge. Thus, our method can learn multi-perspective pedestrian information, including inter-frame long-range interaction information and intra-frame multi-layer global and local information. Extensive experiments demonstrate the superiority of the proposed HASI method compared with the state-of-the-art methods on the three challenging video-based re-ID datasets, i.e., MARS, iLIDS-VID, and PRID-2011. Si Chen 0002, Hui Da, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001, Shunzhi Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Video Text Detection With Robust Feature RepresentationabstractExisting video text detection methods mostly track texts with appearance feature only, thus are easily influenced by the change of perspective and illumination. In this paper, we propose an end-to-end video text detector that tracks texts based on robust feature representation fusing multiple descriptors. First, we introduce a character center segmentation branch to extract semantic feature, which encodes the category and position information of characters. And for extracting the topology feature of each text instance, we propose a relative position awareness branch to encode the relative position information among texts. Then, an adaptive feature fusion network is proposed to dynamically fuse multiple descriptors to generate a robust feature representation for more robust tracking. In addition, to promote the research and evaluation in this field, we also construct a large Bilingual Road scene Video Text dataset, named BiRViT-1K, which contains 1000 videos of Chinese and English texts. Experimental results show the proposed semantic and topology features are beneficial to the text detection and tracking performance, and the proposed method achieves state-of-the-art performance on four public video text benchmarks ICDAR 2015 Video, YVT, RT-1K and BOVText, and two Chinese scene text benchmarks CASIA10K and MSRA-TD500. Wei Feng 0016, Mengbiao Zhao, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | OpenMix: Exploring Outlier Samples for Misclassification DetectionabstractReliable confidence estimation for deep neural classifiers is a challenging yet fundamental requirement in high-stakes applications. Unfortunately, modern deep neural networks are often overconfident for their erroneous predictions. In this work, we exploit the easily available outlier samples, i.e., unlabeled samples coming from non-target classes, for helping detect misclassification errors. Particularly, we find that the well-known Outlier Exposure, which is powerful in detecting out-of-distribution (OOD) samples from unknown classes, does not provide any gain in identifying misclassification errors. Based on these observations, we propose a novel method called OpenMix, which incorporates open-world knowledge by learning to reject uncertain pseudo-samples generated via outlier transformation. OpenMix significantly improves confidence reliability under various scenarios, establishing a strong and unified framework for detecting both misclassified samples from known classes and OOD samples from unknown classes. The code is publicly available at https://github.com/Impression2805/OpenMix. Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 3 |
| 2023 | Training with scaled logits to alleviate class-level over-fitting in few-shot learning
Rui-Qi Wang, Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Neurocomputing | 3 |
| 2023 | Imitating the oracle: Towards calibrated model for class incremental learning
Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Neural Networks | 3 |
| 2023 | Learning by Seeing More ClassesabstractTraditional pattern recognition models usually assume a fixed and identical number of classes during both training and inference stages. In this paper, we study an interesting but ignored question: can increasing the number of classes during training improve the generalization and reliability performance? For a k-class problem, instead of training with only these k classes, we propose to learn with k+m classes, where the additional m classes can be either real classes from other datasets or synthesized from known classes. Specifically, we propose two strategies for constructing new classes from known classes. By making the model see more classes during training, we can obtain several advantages. First, the added m classes serve as a regularization which is helpful to improve the generalization accuracy on the original k classes. Second, this will alleviate the overconfident phenomenon and produce more reliable confidence estimation for different tasks like misclassification detection, confidence calibration, and out-of-distribution detection. Lastly, the additional classes can also improve the learned feature representation, which is beneficial for new classes generalization in few-shot learning and class-incremental learning. Compared with the widely proved concept of data augmentation (dataAug), our method is driven from another dimension of augmentation based on additional classes (classAug). Comprehensive experiments demonstrated the superiority of our classAug under various open-environment metrics on benchmark datasets. Fei Zhu 0004, Xu-Yao Zhang, Rui-Qi Wang, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | A Survey on Learning to RejectabstractLearning to reject is a special kind of self-awareness (the ability to know what you do not know), which is an essential factor for humans to become smarter. Although machine intelligence has become very accurate nowadays, it lacks such kind of self-awareness and usually acts as omniscient, resulting in overconfident errors. This article presents a comprehensive overview of this topic from three perspectives: confidence, calibration, and discrimination. Confidence is an important measurement for the reliability of model predictions. Rejection can be realized by setting thresholds on confidence. However, most models, especially modern deep neural networks, are usually overconfident. Therefore, calibration is a process to ensure confidence matching the actual likelihood of correctness, including two approaches: post-calibration and self-calibration. Calibration reflects the global characteristic of confidence, and the local distinguishing property of confidence is also important. In light of this, discrimination focuses on the performance of accepting positive samples while rejecting negative samples. As a binary classification problem, the challenge of discrimination comes from the missing and nonrepresentativeness of the negative data. Three discrimination tasks are comprehensively analyzed and discussed: failure rejection, unknown rejection, and fake rejection. By rejecting failures, the risk could be controlled especially for mission-critical applications. By rejecting unknowns, the awareness of the knowledge blind zone would be enhanced. By rejecting fakes, security and privacy could be protected. We provide a general taxonomy, organization, and discussion of the methods for solving these problems, which are studied separately in the literature. The connections between different approaches and future directions that are worth further investigation are also presented. With a discriminative and calibrated confidence, learning to reject will let the decision-making process be more practical, reliable, and secure. Xu-Yao Zhang, Guosen Xie, Xiuli Li, Tao Mei 0001, Cheng-Lin Liu 0001 |
Proc. IEEE | 1 |
| 2023 | Adversarial training with distribution normalization and margin balance
Zhen Cheng 0003, Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2023 | Dynamics-aware loss for learning with label noise
Xiu-Chuan Li, Xiaobo Xia, Fei Zhu 0004, Tongliang Liu, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 5 |
| 2023 | Self-information of radicals: A new clue for zero-shot Chinese character recognition
Dahan Wang, Xia Du, Huayi Yin, Xu-Yao Zhang, Shunzhi Zhu |
Pattern Recognit. | 5 |
| 2023 | Towards prior gap and representation gap for long-tailed recognition
Mingliang Zhang 0005, Xu-Yao Zhang, Chuang Wang 0007, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2023 | Deep representation learning for domain generalization with information bottleneck principle
Xu-Yao Zhang, Chuang Wang 0007, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2023 | Leveraging Balanced Semantic Embedding for Generative Zero-Shot LearningabstractGenerative (generalized) zero-shot learning [(G)ZSL] models aim to synthesize unseen class features by using only seen class feature and attribute pairs as training data. However, the generated fake unseen features tend to be dominated by the seen class features and thus classified as seen classes, which can lead to inferior performances under zero-shot learning (ZSL), and unbalanced results under generalized ZSL (GZSL). To address this challenge, we tailor a novel balanced semantic embedding generative network (BSeGN), which incorporates balanced semantic embedding learning into generative learning scenarios in the pursuit of unbiased GZSL. Specifically, we first design a feature-to-semantic embedding module (FEM) to distinguish real seen and fake unseen features collaboratively with the generator in an online manner. We introduce the bidirectional contrastive and balance losses for the FEM learning, which can guarantee a balanced prediction for the interdomain features. In turn, the updated FEM can boost the learning of the generator. Next, we propose a multilevel feature integration module (mFIM) from the cycle-consistency branch of BSeGN, which can mitigate the domain bias through feature enhancement. To the best of our knowledge, this is the first work to explore embedding and generative learning jointly within the field of ZSL. Extensive evaluations on four benchmarks demonstrate the superiority of BSeGN over its state-of-the-art counterparts. Guosen Xie, Xu-Yao Zhang, Tian-Zhu Xiang, Fang Zhao 0006, Zheng Zhang 0006, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Rethinking Confidence Calibration for Failure Prediction
Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ECCV (25) | 3 |
| 2022 | Critical Radical Analysis Network for Chinese Character RecognitionabstractZero-shot learning is a challenging problem in many tasks due to the lack of training samples of the unseen classes. The radical-based zero-shot Chinese character recognition methods treat Chinese characters as a combination of radicals and structures, and recognize Chinese characters by identifying the radicals and structures contained in them. Current approaches generally treat the contribution of all radicals to Chinese character recognition as the same, and the recognition results rely on the network’s ability to recognize radicals and their corresponding position information, ignoring the potential value of radicals themselves in eliminating the uncertainty of Chinese characters. In this paper, we model the problem of radical-based Chinese character recognition as an uncertainty elimination problem and propose a Critical Radical Analysis Network (CRAN) to explore the Ideographic Description Sequence (IDS) information for zero-shot Chinese character recognition. Specifically, we propose a novel method to compute the critical values of radicals based on information theory using the predefined Chinese character IDS dictionary. In recognition, we use an iterative approach to translate the predicted radical sequence to target Chinese characters. That is, the radicals of the predicted sequence are sorted in descending order of the critical value, and then the radicals are continuously selected in this order as the information obtained to eliminate the uncertainty of the Chinese character until the character is recognized. We conduct experiments on the CTW, CASIA-AHCDB, and CASIA-HWDB datasets. The experimental results show that the proposed method improves the ability of recognizing unseen Chinese characters, demonstrating the effectiveness of the proposed method. Huayi Yin, Dahan Wang, Xu-Yao Zhang, Shunzhi Zhu |
ICPR | 4 |
| 2022 | Convolutional Prototype Network for Open Set RecognitionabstractDespite the success of convolutional neural network (CNN) in conventional closed-set recognition (CSR), it still lacks robustness for dealing with unknowns (those out of known classes) in open environment. To improve the robustness of CNN in open-set recognition (OSR) and meanwhile maintain its high accuracy in CSR, we propose an alternative deep framework called convolutional prototype network (CPN), which keeps CNN for representation learning but replaces the closed-world assumed softmax with an open-world oriented and human-like prototype model. To equip CPN with discriminative ability for classifying known samples, we design several discriminative losses for training. Moreover, to increase the robustness of CPN for unknowns, we interpret CPN from the perspective of generative model and further propose a generative loss, which is essentially maximizing the log-likelihood of known samples and serves as a latent regularization for discriminative learning. The combination of discriminative and generative losses makes CPN a hybrid model with advantages for both CSR and OSR. Under the designed losses, the CPN is trained end-to-end for learning the convolutional network and prototypes jointly. For application of CPN in OSR, we propose two rejection rules for detecting different types of unknowns. Experiments on several datasets demonstrate the efficiency and effectiveness of CPN for both CSR and OSR tasks. Hong-Ming Yang, Xu-Yao Zhang, Qing Yang 0002, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Cross-modal prototype learning for zero-shot handwritten character recognition
Xiang Ao 0002, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2022 | A survey of robust adversarial training in pattern recognition: Fundamental, theory, and methodologies
Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Xu-Yao Zhang |
Pattern Recognit. | 4 |
| 2022 | Decision-Based Adversarial Attack With Frequency MixupabstractIt has been widely observed that deep neural networks are highly vulnerable to adversarial examples. Decision-based attacks could generate adversarial examples based solely on top-1 labels returned by the target model. However, they typically make excessive queries and could not bypass detection effectively. To comprehensively assess a decision-based attack, besides its query efficiency, the performance against detection is also a concern. Considering that previous detections consume massive resources and always mistakenly recognize benign video frames as malicious attacks, we design a lightweight detection calledboundary detectionto overcome the above limitations, whose success reveals serious limitations of existing decision-based attacks. To develop more powerful attacks, we first presentf-mixupas a basic method to produce candidate adversarial examples in the frequency domain. Usingf-mixupas the building block, we proposef-attackas a complete decision-based attack. With the help of several natural images,f-attackcould both work well with limited (hundreds of) queries and bypass detection effectively. Nevertheless, if the attacker could make relatively adequate (thousands of) queries and the target model is not equipped with detection,f-attackwill lag behind existing decision-based attacks. We additionally introducefrequency binary searchbased onf-mixup, which serves as a plug-and-play module for existing decision-based attacks to further improve their query efficiency. Experimental results verify the effectiveness of our proposed methods. Xiu-Chuan Li, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Mixed-Supervised Scene Text Detection With Expectation-Maximization AlgorithmabstractScene text detection is an important and challenging task in computer vision. For detecting arbitrarily-shaped texts, most existing methods require heavy data labeling efforts to produce polygon-level text region labels for supervised training. In order to reduce the cost in data labeling, we study mixed-supervised arbitrarily-shaped text detection by combining various weak supervision forms (e.g., image-level tags, coarse, loose and tight bounding boxes), which are far easier to annotate. Whereas the existing weakly-supervised learning methods (such as multiple instance learning) do not promote full object coverage, to approximate the performance of fully-supervised detection, we propose an Expectation-Maximization (EM) based mixed-supervised learning framework to train scene text detector using only a small amount of polygon-level annotated data combined with a large amount of weakly annotated data. The polygon-level labels are treated as latent variables and recovered from the weak labels by the EM algorithm. A new contour-based scene text detector is also proposed to facilitate the use of weak labels in our mixed-supervised learning framework. Extensive experiments on six scene text benchmarks show that (1) using only 10% strongly annotated data and 90% weakly annotated data, our method yields comparable performance to that of fully supervised methods, (2) with 100% strongly annotated data, our method achieves state-of-the-art performance on five scene text benchmarks (CTW1500, Total-Text, ICDAR-ArT, MSRA-TD500, and C-SVT), and competitive results on the ICDAR2015 Dataset. We will make our weakly annotated datasets publicly available. Mengbiao Zhao, Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Meta-Prototypical Learning for Domain-Agnostic Few-Shot RecognitionabstractFew-shot learning (FSL) aims to classify novel images based on a few labeled samples with the help of meta-knowledge. Most previous works address this problem based on the hypothesis that the training set and testing set are from the same domain, which is not realistic for some real-world applications. Thus, we extend FSL to domain-agnostic few-shot recognition, where the domain of the testing task is unknown. In domain-agnostic few-shot recognition, the model is optimized on data from one domain and evaluated on tasks from different domains. Previous methods for FSL mostly focus on learning general features or adapting to few-shot tasks effectively. They suffer from inappropriate features or complex adaptation in domain-agnostic few-shot recognition. In this brief, we propose meta-prototypical learning to address this problem. In particular, a meta-encoder is optimized to learn the general features. Different from the traditional prototypical learning, the meta encoder can effectively adapt to few-shot tasks from different domains by the traces of the few labeled examples. Experiments on many datasets demonstrate that meta-prototypical learning performs competitively on traditional few-shot tasks, and on few-shot tasks from different domains, meta-prototypical learning outperforms related methods. Rui-Qi Wang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Proxy Graph Matching with Proximal Matching NetworksabstractEstimating feature point correspondence is a common technique in computer vision. A line of recent data-driven approaches utilizing the graph neural networks improved the matching accuracy by a large margin. However, these learning-based methods require a lot of labeled training data, which are expensive to collect. Moreover, we find most methods are sensitive to global transforms, for example, a random rotation. On the contrary, classical geometric approaches are immune to rotational transformation though their performance is generally inferior. To tackle these issues, we propose a new learning-based matching framework, which is designed to be rotationally invariant. The model only takes geometric information as input. It consists of three parts: a graph neural network to generate a high-level local feature, an attention-based module to normalize the rotational transform, and a global feature matching module based on proximal optimization. To justify our approach, we provide a convergence guarantee for the proximal method for graph matching. The overall performance is validated by numerical experiments. In particular, our approach is trained on the synthetic random graphs and then applied to several real-world datasets. The experimental results demonstrate that our method is robust to rotational transform and highlights its strong performance of matching accuracy. Haoru Tan, Chuang Wang 0007, Sitong Wu, Tie-Qiang Wang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
AAAI | 5 |
| 2021 | Graph-to-Graph: Towards Accurate and Interpretable Online Handwritten Mathematical Expression RecognitionabstractRecent handwritten mathematical expression recognition (HMER) approaches treat the problem as an image-to-markup generation task where the handwritten formula is translated into a sequence (e.g. LaTeX). The encoder-decoder framework is widely used to solve this image-to-sequence problem. However, (i) for structured mathematical formula, the hierarchical structure neither in the formula nor in the markup has been explored adequately. In addition, (ii) existing image-to-markup methods could not explicitly segment mathematical symbols in the formula corresponding to each target markup token. In this paper, we address the above issues by formulating the HMER as a graph-to-graph (G2G) learning problem. Graph is more flexible and general for structure representation and learning compared with image or sequence. At the core of our method lies the embedding of input formula and output markup into graphs on primitives, with Graph Neural Networks (GNN) to explore the structural information, and a novel sub-graph attention mechanism to match primitives in the input and output graphs. We conduct extensive experiments on CROHME datasets to demonstrate the benefits of the proposed G2G model. Our method yields significant improvements over previous SOTA image-to-markup systems. Moreover, it explicitly resolves the symbol segmentation problem while still being trained end-to-end, making the whole system much more accurate and interpretable. Jin-Wen Wu, Yan-Ming Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
AAAI | 4 |
| 2021 | Semantic-Aware Video Text DetectionabstractMost existing video text detection methods track texts with appearance features, which are easily influenced by the change of perspective and illumination. Compared with appearance features, semantic features are more robust cues for matching text instances. In this paper, we propose an end-to-end trainable video text detector that tracks texts based on semantic features. First, we introduce a new character center segmentation branch to extract semantic features, which encode the category and position of characters. Then we propose a novel appearance-semantic-geometry descriptor to track text instances, in which se-mantic features can improve the robustness against appearance changes. To overcome the lack of character-level an-notations, we propose a novel weakly-supervised character center detection module, which only uses word-level annotated real images to generate character-level labels. The proposed method achieves state-of-the-art performance on three video text benchmarks ICDAR 2013 Video, Minetto and RT-1K, and two Chinese scene text benchmarks CA-SIA10K and MSRA-TD500. Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 3 |
| 2021 | Prototype Augmentation and Self-Supervision for Incremental LearningabstractDespite the impressive performance in many individual tasks, deep neural networks suffer from catastrophic forgetting when learning new tasks incrementally. Recently, various incremental learning methods have been proposed, and some approaches achieved acceptable performance relying on stored data or complex generative models. However, storing data from previous tasks is limited by memory or privacy issues, and generative models are usually unstable and inefficient in training. In this paper, we propose a simple non-exemplar based method named PASS, to address the catastrophic forgetting problem in incremental learning. On the one hand, we propose to memorize one class-representative prototype for each old class and adopt prototype augmentation (protoAug) in the deep feature space to maintain the decision boundary of previous tasks. On the other hand, we employ self-supervised learning (SSL) to learn more generalizable and transferable features for other tasks, which demonstrates the effectiveness of SSL in incremental learning. Experimental results on benchmark datasets show that our approach significantly outperforms non-exemplar based methods, and achieves comparable performance compared to exemplar based approaches. Fei Zhu 0004, Xu-Yao Zhang, Chuang Wang 0007, Cheng-Lin Liu 0001 |
CVPR | 2 |
| 2021 | Adaptive Scaling for Archival Table Structure Recognition
Xiao-Hui Li 0012, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICDAR (1) | 3 |
| 2021 | Document Dewarping with Control Points
Guo-Wang Xie, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICDAR (1) | 3 |
| 2021 | GSS: Graph-Based Subspace Learning with Shots Initialization for few-shot RecognitionabstractPrevious methods of few-shot Learning mostly solve different few-shot recognition tasks in an identical feature space. But identical features are hard to fit various tasks. Some works show that learning a unique subspace for each few-shot recognition task can improve the signal-noise ratio (SNR) of the features and boost the performance. However, there are still two problems remaining. First, in constructing the subspace for few-shot task, often some information (embeddings of queries or labels of shots) are discarded. Second, the eigendecomposition of covariance matrix is usually needed, which degrades the efficiency of the whole model. In this paper, we propose Graph-based Subspace learning with Shots initialization (GSS) for few-shot recognition to learn a better subspace efficiently. In GSS, the bases of the subspace are directly initialized with labels based on shots (given labeled samples) and iteratively updated for better discrimination based on a graph that connects bases and all samples. Extensive experiments on four few-shot benchmark datasets show that GSS reports better performance and higher efficient compared with previous subspace based methods and achieves state-of-the-art performance. Rui-Qi Wang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICME | 2 |
| 2021 | Class Forge: Boosting Feature Encoder for Few-Shot Learning with Synthesized ClassesabstractFew-shot Learning (FSL) aims to gain classification ability on novel classes with only a few labeled samples. Previous works explore meta-learning, metric learning, and graph based methods. Though data augmentation is important to enhance the generalizability of neural networks, it is not well exploited in the field of FSL. We investigate the augmentation in FSL and propose Class Forge to synthesize forged classes that help to learn an encoder with better generalization to novel classes. Specifically, Class Forge divides given base visual classes into parts and combines these parts to synthesize forged visual classes. Training with the additional forged classes forces the encoder to learn richer features that can embed different parts, so as to boost the generalization to novel classes. Intrinsically, Class Forge is a "class augmentation" method that provides a simple yet effective way to synthesize classes, other than synthesizing samples of given classes in previous works. Extensive experiments show that Class Forge yields consistent performance gain on different datasets for FSL. And the ablation studies validate that features learned with Class Forge demonstrate better generalization ability. Rui-Qi Wang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICME | 2 |
| 2021 | Calibration for Non-Exemplar Based Class-Incremental LearningabstractCatastrophic forgetting is the central challenge in incremental learning. Notable studies address the problem by using regularization or experience replay strategies. However, the performance is far from ideal without storing previous data, especially in the scenario of class-incremental learning (CIL). In CIL setting, an important factor causing catastrophic forgetting is the severe bias between the new and previously learned classes, in both classifier and feature extractor. In this paper, we propose calibrateCIL which contains two simple modifications to calibrate the bias in non-exemplar based CIL. Specifically, local softmax is proposed to calibrate the classifier, and cutout training is used to calibrate the feature extractor by learning richer, more generalizable and transferable features. Our method can give balance class scores without any post-processing technique. We show that our method outperforms state-of-the-art non-exemplar based methods on the challenging problem of CIL, and the ablation study demonstrates the effectiveness of the two modifications. Fei Zhu 0004, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICME | 2 |
| 2021 | Class-Incremental Learning via Dual AugmentationabstractDeep learning systems typically suffer from catastrophic forgetting of past knowledge when acquiring new skills continually. In this paper, we emphasize two dilemmas, representation bias and classifier bias in class-incremental learning, and present a simple and novel approach that employs explicit class augmentation (classAug) and implicit semantic augmentation (semanAug) to address the two biases, respectively. On the one hand, we propose to address the representation bias by learning transferable and diverse representations. Specifically, we investigate the feature representations in incremental learning based on spectral analysis and present a simple technique called classAug, to let the model see more classes during training for learning representations transferable across classes. On the other hand, to overcome the classifier bias, semanAug implicitly involves the simultaneous generating of an infinite number of instances of old classes in the deep feature space, which poses tighter constraints to maintain the decision boundary of previously learned classes. Without storing any old samples, our method can perform comparably with representative data replay based approaches. Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
NeurIPS | 3 |
| 2021 | End-to-End Detection and Recognition of Arithmetic Expressions
Jiangpeng Wan, Mengbiao Zhao, Xu-Yao Zhang, Linlin Huang 0001 |
PRCV (1) | 4 |
| 2021 | Residual Dual Scale Scene Text Spotting by Fusing Bottom-Up and Top-Down Processing
Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 3 |
| 2021 | VMAN: A Virtual Mainstay Alignment Network for Transductive Zero-Shot LearningabstractTransductive zero-shot learning (TZSL) extends conventional ZSL by leveraging (unlabeled) unseen images for model training. A typical method for ZSL involves learning embedding weights from the feature space to the semantic space. However, the learned weights in most existing methods are dominated by seen images, and can thus not be adapted to unseen images very well. In this paper, to align the (embedding) weights for better knowledge transfer between seen/unseen classes, we propose the virtual mainstay alignment network (VMAN), which is tailored for the transductive ZSL task. Specifically, VMAN is casted as a tied encoder-decoder net, thus only one linear mapping weights need to be learned. To explicitly learn the weights in VMAN, for the first time in ZSL, we propose to generate virtual mainstay (VM) samples for each seen class, which serve as new training data and can prevent the weights from being shifted to seen images, to some extent. Moreover, a weighted reconstruction scheme is proposed and incorporated into the model training phase, in both the semantic/feature spaces. In this way, the manifold relationships of the VM samples are well preserved. To further align the weights to adapt to more unseen images, a novel instance-category matching regularization is proposed for model re-training. VMAN is thus modeled as a nested minimization problem and is solved by a Taylor approximate optimization paradigm. In comprehensive evaluations on four benchmark datasets, VMAN achieves superior performances under the (Generalized) TZSL setting. Guosen Xie, Xu-Yao Zhang, Yazhou Yao, Zheng Zhang 0006, Fang Zhao 0006, Ling Shao 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Dewarping Document Image by Displacement Flow Estimation with Fully Convolutional Network
Guo-Wang Xie, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
DAS | 3 |
| 2020 | Cross-Lingual Text Image Recognition via Multi-Task Sequence to Sequence LearningabstractThis paper considers recognizing texts shown in a source language and translating into a target language, without generating the intermediate source language text image recognition results. We call this problem Cross-Lingual Text Image Recognition (CLTIR). To solve this problem, we propose a multi-task system containing a main task of CLTIR and an auxiliary task of Mono-Lingual Text Image Recognition (MLTIR) simultaneously. Two different sequence to sequence learning methods, a convolution based attention model and a Bidirectional Long Short-Term Memory (BLSTM) model with Connectionist Temporal Classification (CTC), are adopted for these tasks respectively. We evaluate the system on a newly collected Chinese-English bilingual movie subtitle image dataset. Experimental results demonstrate the multi-task learning framework performs superiorly in both languages. Zhuo Chen 0051, Xu-Yao Zhang, Qing Yang 0002, Cheng-Lin Liu 0001 |
ICPR | 3 |
| 2020 | F-mixup: Attack CNNs From Fourier PerspectiveabstractRecent research has revealed that deep neural networks are highly vulnerable to adversarial examples. In this paper, different from most adversarial attacks which directly modify pixels in spatial domain, we propose a novel black-box attack in frequency domain, named as f-mixup, based on the property of natural images and perception disparity between human-visual system (HVS) and convolutional neural networks (CNNs): First, natural images tend to have the bulk of their Fourier spectrums concentrated on the low frequency domain; Second, HVS is much less sensitive to high frequencies while CNNs can utilize both low and high frequency information to make predictions. Extensive experiments are conducted and show that deeper CNNs tend to concentrate more on the higher frequency domain, which may explain the contradiction between robustness and accuracy. In addition, we compared f-mixup with existing attack methods and observed that our approach possesses great advantages. Finally, we show that f-mixup can be also incorporated in training to make deep CNNs defensible against a kind of perturbations effectively. Xiu-Chuan Li, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICPR | 2 |
| 2020 | Mutually Guided Dual-Task Network for Scene Text DetectionabstractScene text detection has been studied extensively. Existing methods detect either words or text lines and use either word-level or line-level annotated data for training. In this paper, we propose a dual-task network that can perform word-level and line-level text detection simultaneously and use training data of both levels of annotation to boost the performance. The dual-task network has two detection heads for word-level and line-level text detection, respectively. Then we propose a mutual guidance scheme for the joint training of the two tasks with two modules: line filtering module utilizes the output feature map of the text line detector to filter out the non-text regions for the word detector, and word enhancing module provides prior positions of words for the text line detector depending on the output feature map of the word detector. Experimental results of word-level and line-level text detection demonstrate the effectiveness of the proposed dual-task network and mutual guidance scheme, and the results of our method are competitive with state-of-the-art methods. Mengbiao Zhao, Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICPR | 4 |
| 2020 | Handwritten Mathematical Expression Recognition via Paired Adversarial Learning
Jin-Wen Wu, Yan-Ming Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 4 |
| 2020 | A benchmark for unconstrained online handwritten Uyghur word recognition
Wujiahemaiti Simayi, Mayire Ibrayim, Xu-Yao Zhang, Cheng-Lin Liu 0001, Askar Hamdulla |
Int. J. Document Anal. Recognit. | 3 |
| 2020 | Online semi-supervised learning with learning vector quantization
Yuan-Yuan Shen, Yan-Ming Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Neurocomputing | 3 |
| 2020 | Towards Robust Pattern Recognition: A ReviewabstractThe accuracies for many pattern recognition tasks have increased rapidly year by year, achieving or even outperforming human performance. From the perspective of accuracy, pattern recognition seems to be a nearly solved problem. However, once launched in real applications, the high-accuracy pattern recognition systems may become unstable and unreliable due to the lack of robustness in open and changing environments. In this article, we present a comprehensive review of research toward robust pattern recognition from the perspective of breaking three basic and implicit assumptions: closed-world assumption, independent and identically distributed assumption, and clean and big data assumption, which form the foundation of most pattern recognition models. Actually, our brain is robust at learning concepts continually and incrementally, in complex, open, and changing environments, with different contexts, modalities, and tasks, by showing only a few examples, under weak or noisy supervision. These are the major differences between human intelligence and machine intelligence, which are closely related to the above three assumptions. After witnessing the significant progress in accuracy improvement nowadays, this review paper will enable us to analyze the shortcomings and limitations of current methods and identify future research directions for robust pattern recognition. Xu-Yao Zhang, Cheng-Lin Liu 0001, Ching Y. Suen |
Proc. IEEE | 1 |
| 2020 | MuLTReNets: Multilingual text recognition networks for simultaneous script identification and handwriting recognition
Zhuo Chen 0051, Xu-Yao Zhang, Qing Yang 0002, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2020 | Realtime multi-scale scene text detection with scale-based region proposal network
Xu-Yao Zhang, Zhenbo Luo, Jean-Marc Ogier, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2020 | SRSC: Selective, Robust, and Supervised Constrained Feature Representation for Image ClassificationabstractFeature representation learning, an emerging topic in recent years, has achieved great progress. Powerful learned features can lead to excellent classification accuracy. In this article, a selective and robust feature representation framework with a supervised constraint (SRSC) is presented. SRSC seeks a selective, robust, and discriminative subspace by transforming the original feature space into the category space. Particularly, we add a selective constraint to the transformation matrix (or classifier parameter) that can select discriminative dimensions of the input samples. Moreover, a supervised regularization is tailored to further enhance the discriminability of the subspace. To relax the hard zero-one label matrix in the category space, an additional error term is also incorporated into the framework, which can lead to a more robust transformation matrix. SRSC is formulated as a constrained least square learning (feature transforming) problem. For the SRSC problem, an inexact augmented Lagrange multiplier method (ALM) is utilized to solve it. Extensive experiments on several benchmark data sets adequately demonstrate the effectiveness and superiority of the proposed method. The proposed SRSC approach has achieved better performances than the compared counterpart methods. Guosen Xie, Zheng Zhang 0006, Li Liu 0004, Fan Zhu 0001, Xu-Yao Zhang, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | TextDragon: An End-to-End Framework for Arbitrary Shaped Text SpottingabstractMost existing text spotting methods either focus on horizontal/oriented texts or perform arbitrary shaped text spotting with character-level annotations. In this paper, we propose a novel text spotting framework to detect and recognize text of arbitrary shapes in an end-to-end manner, using only word/line-level annotations for training. Motivated from the name of TextSnake, which is only a detection model, we call the proposed text spotting framework TextDragon. In TextDragon, a text detector is designed to describe the shape of text with a series of quadrangles, which can handle text of arbitrary shapes. To extract arbitrary text regions from feature maps, we propose a new differentiable operator named RoISlide, which is the key to connect arbitrary shaped text detection and recognition. Based on the extracted features through RoISlide, a CNN and CTC based text recognizer is introduced to make the framework free from labeling the location of characters. The proposed method achieves state-of-the-art performance on two curved text benchmarks CTW1500 and Total-Text, and competitive results on the ICDAR 2015 Dataset. Wei Feng 0016, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICCV | 4 |
| 2019 | Cross-Modal Prototype Learning for Zero-Shot Handwriting RecognitionabstractIn contrast to machine recognizers that rely on training with large handwriting data, humans can recognize handwriting accurately on learning from few samples, and can even generalize to handwritten characters from printed samples. Simulating this ability in machine recognition is important to alleviate the burden of labeling large handwriting data, especially for large category set as in Chinese text. In this paper, inspired by human learning, we propose a cross-modal prototype learning (CMPL) method for zero-shot online handwritten character recognition: for unseen categories, handwritten characters can be recognized without learning from handwritten samples, but instead from printed characters. Particularly, the printed characters (one for each class) are embedded into a convolutional neural network (CNN) feature space to obtain prototypes representing each class, while the online handwriting trajectories are embedded with a recurrent neural network (RNN). Via cross-modal joint learning, handwritten characters can be recognized according to the printed prototypes. For unseen categories, handwritten characters can be recognized by only feeding a printed sample per category. Experiments on a benchmark Chinese handwriting database have shown the effectiveness and potential of the proposed method for zero-shot handwriting recognition. Xiang Ao 0002, Xu-Yao Zhang, Hong-Ming Yang, Cheng-Lin Liu 0001 |
ICDAR | 2 |
| 2019 | CASIA-AHCDB: A Large-Scale Chinese Ancient Handwritten Characters DatabaseabstractThis paper introduces a Chinese Ancient Handwritten Characters Database (CASIA-AHCDB) for character recognition research. The database was built by annotating 11,937 pages of Chinese ancient handwritten documents. It consists of more than 2.2 million annotated handwritten character samples of 10,350 categories. According to the source of these documents, the database is divided into two datasets of different styles: Complete Library in Four Sections (AHCDB-style1) and Ancient Buddhist Scriptures (AHCDB-style2). Each dataset can be divided into three parts based on its applications. The first part, called basic category set, contains samples of common categories in two datasets, and is suitable for basic character recognition task. The second part, called enhanced category set, is mainly used for open-set character recognition task based on the basic character recognition. The third part, called the reserved category set, can be used in many pattern recognition tasks in the future. Based on the large category set, the various writing styles and the imbalanced sample number per category, CASIA-AHCDB can also be used for various classification and learning tasks such as transfer learning, few-shot learning. We performed experiments of basic character recognition on the basic category set, and report the results for benchmark. More techniques can be evaluated on this challenging database in the future. Dahan Wang, Xu-Yao Zhang, Zhaoxiang Zhang 0001, Cheng-Lin Liu 0001 |
ICDAR | 4 |
| 2019 | Fast Text/non-Text Image Classification with Knowledge DistillationabstractHow to efficiently judge whether a natural image contains texts or not is an important problem. Since text detection and recognition algorithms are usually time-consuming, and it is unnecessary to run them on images that do not contain any texts. In this paper, we investigate this problem from two perspectives: the speed and the accuracy. First, to achieve high speed for efficient filtering large number of images especially on CPU, we propose using small and shallow convolutional neural network, where the features from different layers are adaptively pooled into certain sizes to overcome difficulties caused by multiple scales and various locations. Although this can achieve high speed but its accuracy is not satisfactory due to limited capacity of small network. Therefore, our second contribution is using the knowledge distillation to improve the accuracy of the small network, by constructing a larger and deeper neural network as teacher network to instruct the learning process of the small network. With the above two strategies, we can achieve both high speed and high accuracy for filtering scene text images. Experimental results on a benchmark dataset have shown the effectiveness of our method: the teacher network yields state-of-the-art performance, and the distilled small network achieves high performance while maintaining high speed which is 176 times faster on CPU and 3.8 times faster on GPU than a compared benchmark method. Miao Zhao, Rui-Qi Wang, Xu-Yao Zhang, Linlin Huang 0001, Jean-Marc Ogier |
ICDAR | 4 |
| 2019 | LightweightNet: Toward fast and lightweight convolutional neural networks via architecture distillation
Ting-Bing Xu, Peipei Yang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2019 | Stochastic Conjugate Gradient Algorithm With Variance ReductionabstractConjugate gradient (CG) methods are a class of important methods for solving linear equations and nonlinear optimization problems. In this paper, we propose a new stochastic CG algorithm with variance reduction1and we prove its linear convergence with the Fletcher and Reeves method for strongly convex and smooth functions. We experimentally demonstrate that the CG with variance reduction algorithm converges faster than its counterparts for four learning models, which may be convex, nonconvex or nonsmooth. In addition, its area under the curve performance on six large-scale data sets is comparable to that of the LIBLINEAR solver for the L2-regularized L2-loss but with a significant improvement in computational efficiency. Xiao-Bo Jin, Xu-Yao Zhang, Kaizhu Huang, Guanggang Geng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Robust Classification With Convolutional Prototype LearningabstractConvolutional neural networks (CNNs) have been widely used for image classification. Despite its high accuracies, CNN has been shown to be easily fooled by some adversarial examples, indicating that CNN is not robust enough for pattern classification. In this paper, we argue that the lack of robustness for CNN is caused by the softmax layer, which is a totally discriminative model and based on the assumption of closed world (i.e., with a fixed number of categories). To improve the robustness, we propose a novel learning framework called convolutional prototype learning (CPL). The advantage of using prototypes is that it can well handle the open world recognition problem and therefore improve the robustness. Under the framework of CPL, we design multiple classification criteria to train the network. Moreover, a prototype loss (PL) is proposed as a regularization to improve the intra-class compactness of the feature representation, which can be viewed as a generative model based on the Gaussian assumption of different classes. Experiments on several datasets demonstrate that CPL can achieve comparable or even better results than traditional CNN, and from the robustness perspective, CPL shows great advantages for both the rejection and incremental category learning tasks. Hong-Ming Yang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 2 |
| 2018 | Deep Transfer Mapping for Unsupervised Writer AdaptationabstractConvolutional neural network (CNN) has achieved great success in handwriting recognition. However, it relies on large set of labeled data in training and its performance will deteriorate when the data distribution varies. To solve this problem, traditional methods usually consider adaptation of the single top layer of CNN. To better reduce the distribution discrepancy, in this paper, we consider adaptation of all layers of CNN including both convolutional and full layers. Four variations of transformations are designed based on different assumptions about the space relations for adaptation of convolutional layers. In order to make adaptation of multiple layers, we propose to cascade the transformations of different layers to conduct adaptation in a deep manner, and therefore this method is denoted as deep transfer mapping (DTM). DTM can capture the information from different layers and minimize the data divergence under different information abstract levels, thus it is more powerful and flexible for domain adaptation. Experiments on the online Chinese handwriting dataset (OLHWDB) demonstrate the efficiency and effectiveness of the proposed method for unsupervised writer adaptation. Hong-Ming Yang, Xu-Yao Zhang, Jun Sun 0004, Cheng-Lin Liu 0001 |
ICFHR | 2 |
| 2018 | Convolutional Discriminant AnalysisabstractSoftmax regressor is arguably the most commonly used classifier in convolutional neural networks (CNNs). However, the cross-entropy based softmax loss only supervises the deep neural networks to learn effective representations of data, but does not explicitly enforce the separability between the classes. In this paper, we propose a novel convolutional neural network model, called convolutional discriminative analysis (CDA). Beyond the softmax loss, CDA employs a convolutional discriminant loss (CD-Loss), which minimizes the distance between the sample and its class center while maximizes the distance between the sample and its adversarial class center in the space of the learned deep representations. Extensive experiments on two benchmark data sets, Fashion-MNIST and CIFAR-10, demonstrate the superiority of CDA over traditional deep CNNs on the image classification tasks. Guoqiang Zhong 0001, Xu-Yao Zhang, Hongxu Wei |
ICPR | 3 |
| 2018 | Image-to-Markup Generation via Paired Adversarial Learning
Jin-Wen Wu, Yan-Ming Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ECML/PKDD (1) | 4 |
| 2018 | Drawing and Recognizing Chinese Characters with Recurrent Neural NetworkabstractRecent deep learning based approaches have achieved great success on handwriting recognition. Chinese characters are among the most widely adopted writing systems in the world. Previous research has mainly focused on recognizing handwritten Chinese characters. However, recognition is only one aspect for understanding a language, another challenging and interesting task is to teach a machine to automatically write (pictographic) Chinese characters. In this paper, we propose a framework by using the recurrent neural network (RNN) as both a discriminative model for recognizing Chinese characters and a generative model for drawing (generating) Chinese characters. To recognize Chinese characters, previous methods usually adopt the convolutional neural network (CNN) models which require transforming the online handwriting trajectory into image-like representations. Instead, our RNN based approach is an end-to-end system which directly deals with the sequential structure and does not require any domain-specific knowledge. With the RNN system (combining an LSTM and GRU), state-of-the-art performance can be achieved on the ICDAR-2013 competition database. Furthermore, under the RNN framework, a conditional generative model with character embedding is proposed for automatically drawing recognizable Chinese characters. The generated characters (in vector format) are human-readable and also can be recognized by the discriminative RNN model with high accuracy. Experimental results verify the effectiveness of using RNNs as both generative and discriminative models for the tasks of drawing and recognizing Chinese characters. Xu-Yao Zhang, Yan-Ming Zhang 0001, Cheng-Lin Liu 0001, Yoshua Bengio |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Multi-Oriented and Multi-Lingual Scene Text Detection With Direct RegressionabstractMulti-oriented and multi-lingual scene text detection plays an important role in computer vision area and is challenging due to the wide variety of text and background. In this paper, firstly we point out the two key tasks when extending CNN based object detection frameworks to scene text detection. The first task is to localize the text region by a down-sampled segmentation based module, and the second task is to regress the boundaries of text region determined by the first task. Secondly, we propose a scene text detection framework based on fully convolutional network (FCN) with a bi-task prediction module in which one is pixel-wise classification between text and non-text, and the other is pixel-wise regression to determine the vertex coordinates of quadrilateral text boundaries. Post-processing for word-level detection is based on Non-Maximum Suppression (NMS), and for line-level detection we design a heuristic line segments grouping method to localize long text lines. We evaluated the proposed framework on various benchmarks including multi-oriented and multi-lingual scene text datasets, and achieved state-of-the-art performance on most of them. We also provide abundant ablation experiments to analyze several key factors in building high performance CNN based scene text detection systems. Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Deep Direct Regression for Multi-oriented Scene Text DetectionabstractIn this paper, we first provide a new perspective to divide existing high performance object detection methods into direct and indirect regressions. Direct regression performs boundary regression by predicting the offsets from a given point, while indirect regression predicts the offsets from some bounding box proposals. In the context of multioriented scene text detection, we analyze the drawbacks of indirect regression, which covers the state-of-the-art detection structures Faster-RCNN and SSD as instances, and point out the potential superiority of direct regression. To verify this point of view, we propose a deep direct regression based method for multi-oriented scene text detection. Our detection framework is simple and effective with a fully convolutional network and one-step post processing. The fully convolutional network is optimized in an end-to-end way and has bi-task outputs where one is pixel-wise classification between text and non-text, and the other is direct regression to determine the vertex coordinates of quadrilateral text boundaries. The proposed method is particularly beneficial to localize incidental scene texts. On the ICDAR2015 Incidental Scene Text benchmark, our method achieves the F-measure of 81%, which is a new state-ofthe-art and significantly outperforms previous approaches. On other standard datasets with focused scene texts, our method also reaches the state-of-the-art performance. Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICCV | 2 |
| 2017 | Handwriting Style Mixture AdaptationabstractIn handwriting recognition, the test data usually come from multiple writers which are not shown in the training data. Therefore, adapting the base classifier towards the new style of each writer can significantly improve the generalization performance. Traditional writer adaptation methods usually assume that there is only one writer (one style) in the test data, and we call this situation as style-clear adaptation. However, a more common situation is that multiple handwriting styles exist in the test data, which is widely appeared in multi-font documents and handwriting data produced by the cooperation of multiple writers. We call the adaptation in this situation as style-mixture adaptation. To deal with this problem, in this paper, we propose a novel method called K-style mixture adaptation (K-SMA) with the assumption that there are totally K styles in the test data. Specifically, we first partition the test data into K groups (style clustering) according to their style consistency, which is measured by a newly designed style feature that can eliminate class (category) information and keep handwriting style information. After that, in each group, a style transfer mapping (STM) is used for writer adaptation. Since the initial style clustering may be not reliable, we repeat this process iteratively to improve the adaptation performance. The K-SMA model is fully unsupervised which do not require either the class label or the style index. Moreover, the K-SMA model can be effectively combined with the benchmark convolutional neural network (CNN) models. Experiments on the online Chinese handwriting database CASIA-OLHWDB demonstrate that K-SMA is an efficient and effective solution for style-mixture adaptation. Hong-Ming Yang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICDAR | 2 |
| 2017 | Margin-Aware Binarized Weight Networks for Image Classification
Ting-Bing Xu, Peipei Yang, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICIG (1) | 3 |
| 2017 | Dynamic Multi-Task Learning with Convolutional Neural NetworkabstractMulti-task learning and deep convolutional neural network (CNN) have been successfully used in various fields. This paper considers the integration of CNN and multi-task learning in a novel way to further improve the performance of multiple related tasks. Existing multi-task CNN models usually empirically combine different tasks into a group which is then trained jointly with a strong assumption of model commonality. Furthermore, traditional approaches usually only consider small number of tasks with rigid structure, which is not suitable for large-scale applications. In light of this, we propose a dynamic multi-task CNN model to handle these problems. The proposed model directly learns the task relations from data instead of subjective task grouping. Due to its flexible structure, it supports task-wise incremental training, which is useful for efficient training of massive tasks. Specifically, we add a new task transfer connection (TTC) between the layers of each task. The learned TTC is able to reflect the correlation among different tasks guiding the model dynamically adjusting the multiplexing of the information among different tasks. With the help of TTC, multiple related tasks can further boost the whole performance for each other. Experiments demonstrate that the proposed dynamic multi-task CNN model outperforms traditional approaches. Yuchun Fang, Zhengyan Ma, Zhaoxiang Zhang 0001, Xu-Yao Zhang, Xiang Bai |
IJCAI | 4 |
| 2017 | Diverse Neuron Type Selection for Convolutional Neural NetworksabstractThe activation function for neurons is a prominent element in the deep learning architecture for obtaining high performance. Inspired by neuroscience findings, we introduce and define two types of neurons with different activation functions for artificial neural networks: excitatory and inhibitory neurons, which can be adaptively selected by self-learning. Based on the definition of neurons, in the paper we not only unify the mainstream activation functions, but also discuss the complementariness among these types of neurons. In addition, through the cooperation of excitatory and inhibitory neurons, we present a compositional activation function that leads to new state-of-the-art performance comparing to rectifier linear units. Finally, we hope that our framework not only gives a basic unified framework of the existing activation neurons to provide guidance for future design, but also contributes neurobiological explanations which can be treated as a window to bridge the gap between biology and computer science. Guibo Zhu, Zhaoxiang Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IJCAI | 3 |
| 2017 | SDE: A Novel Selective, Discriminative and Equalizing Feature Representation for Visual Recognition
Guosen Xie, Xu-Yao Zhang, Shuicheng Yan, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 2 |
| 2017 | LG-CNN: From local parts to global discrimination for fine-grained recognition
Guosen Xie, Xu-Yao Zhang, Wenhan Yang, Mingliang Xu 0001, Shuicheng Yan, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2017 | Online and offline handwritten Chinese character recognition: A comprehensive study and new benchmark
Xu-Yao Zhang, Yoshua Bengio, Cheng-Lin Liu 0001 |
Pattern Recognit. | 1 |
| 2017 | Hybrid CNN and Dictionary-Based Models for Scene Recognition and Domain AdaptationabstractConvolutional neural network (CNN) has achieved the state-of-the-art performance in many different visual tasks. Learned from a large-scale training data set, CNN features are much more discriminative and accurate than the handcrafted features. Moreover, CNN features are also transferable among different domains. On the other hand, traditional dictionary-based features (such as BoW and spatial pyramid matching) contain much more local discriminative and structural information, which is implicitly embedded in the images. To further improve the performance, in this paper, we propose to combine CNN with dictionary-based models for scene recognition and visual domain adaptation (DA). Specifically, based on the well-tuned CNN models (e.g., AlexNet and VGG Net), two dictionary-based representations are further constructed, namely, mid-level local representation (MLR) and convolutional Fisher vector (CFV) representation. In MLR, an efficient two-stage clustering method, i.e., weighted spatial and feature space spectral clustering on the parts of a single image followed by clustering all representative parts of all images, is used to generate a class-mixture or a class-specific part dictionary. After that, the part dictionary is used to operate with the multiscale image inputs for generating mid-level representation. In CFV, a multiscale and scale-proportional Gaussian mixture model training strategy is utilized to generate Fisher vectors based on the last convolutional layer of CNN. By integrating the complementary information of MLR, CFV, and the CNN features of the fully connected layer, the state-of-the-art performance can be achieved on scene recognition and DA problems. An interested finding is that our proposed hybrid representation (from VGG net trained on ImageNet) is also complementary to GoogLeNet and/or VGG-11 (trained on Place205) greatly. Guosen Xie, Xu-Yao Zhang, Shuicheng Yan, Cheng-Lin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | End-to-End Online Writer Identification With Recurrent Neural NetworkabstractWriter identification is an important topic for pattern recognition and artificial intelligence. Traditional methods rely heavily on sophisticated hand-crafted features to represent the characteristics of different writers. In this paper, we propose an end-to-end framework for online text-independent writer identification by using a recurrent neural network (RNN). Specifically, the handwriting data of a particular writer are represented by a set of random hybrid strokes (RHSs). Each RHS is a randomly sampled short sequence representing pen tip movements ($xy$-coordinates) and pen-down or pen-up states. RHS is independent of the content and language involved in handwriting; therefore, writer identification at the RHS level is more general and convenient than the character level or the word level, which also requires character/word segmentation. The RNN model with bidirectional long short-term memory is used to encode each RHS into a fixed-length vector for final classification. All the RHSs of a writer are classified independently, and then, the posterior probabilities are averaged to make the final decision. The proposed framework is end-to-end and does not require any domain knowledge for handwriting data analysis. Experiments on both English (133 writers) and Chinese (186 writers) databases verify the advantages of our method compared with other state-of-the-art approaches. Xu-Yao Zhang, Guosen Xie, Cheng-Lin Liu 0001, Yoshua Bengio |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2016 | Large-Scale Graph-Based Semi-Supervised Learning via Tree Laplacian SolverabstractGraph-based Semi-Supervised learning is one of the most popular and successful semi-supervised learning methods. Typically, it predicts the labels of unlabeled data by minimizing a quadratic objective induced by the graph, which is unfortunately a procedure of polynomial complexity in the sample size $n$. In this paper, we address this scalability issue by proposing a method that approximately solves the quadratic objective in nearly linear time. The method consists of two steps: it first approximates a graph by a minimum spanning tree, and then solves the tree-induced quadratic objective function in O(n) time which is the main contribution of this work. Extensive experiments show the significant scalability improvement over existing scalable semi-supervised learning methods. Yan-Ming Zhang 0001, Xu-Yao Zhang, Xiao-Tong Yuan, Cheng-Lin Liu 0001 |
AAAI | 2 |
| 2016 | Handwritten Chinese Character Recognition by Joint Classification and Similarity RankingabstractDeep convolutional neural networks (DCNN) have recently achieved state-of-the-art performance on handwritten Chinese character recognition (HCCR). However, most of DCNN models employ the softmax activation function and minimize cross-entropy loss, which may loss some inter-class information. To cope with this problem, we demonstrate a small but consistent advantage of using both classification and similarity ranking signals as supervision. Specifically, the presented method learns a DCNN model by maximizing the inter-class variations and minimizing the intra-class variations, and simultaneously minimizing the cross-entropy loss. In addition, we also review some loss functions for similarity ranking and evaluate their erformance. Our experiments demonstrate that the presented method achieves state-of-the-art accuracy on the well-known ICDAR 2013 offline HCCR competition dataset. Xu-Yao Zhang, Xiaohu Shao |
ICFHR | 2 |
| 2016 | Unsupervised Adaptation of Neural Networks for Chinese Handwriting RecognitionabstractWriter adaptation is an important topic in handwriting recognition, which can further improve the performance of writer-independent recognizer. In this paper, we propose combining the neural network classifier with style transfer mapping (STM) for unsupervised writer adaptation, which only require writer-specific unlabeled data, and therefore is more common and efficient compared to supervised adaptation. We use some techniques like dropout, ReLU, momentum, and deeply supervised strategy to improve the performance of the neural network classifier. For a specific writer in the test data, an adaptation layer is added to the pre-trained neural network classifier. In adaptation process, only the parameters in adaptation layer are updated while other parameters of the neural network are kept unchanged. To train the adaptation layer, we use the same technology as STM learning but redefine the source point set, target point set and the corresponding confidence. Experiments on the online Chinese handwriting database CASIA-OLHWDB1.1 demonstrate that our method is very efficient and effective in improving classification accuracy. The experimental results also show that our proposed method outperforms the previous proposed learning vector quantization (LVQ) and modified quadratic discriminant function (MQDF) with STM methods for writer adaptation. Hong-Ming Yang, Xu-Yao Zhang, Zhenbo Luo, Cheng-Lin Liu 0001 |
ICFHR | 2 |
| 2016 | Exploiting coarse-to-fine mechanism for fine-grained recognitionabstractFine-grained object recognition is more challenging than generic categorization due to the subtle difference between subcategories under the large intra-class pose change and appearance variations. The state-of-the-art fine-grained recognition methods usually utilize part detection or pose alignment to alleviate the pose variation, and then use convolutional neural networks (CNNs) to extract local discriminative features. Although the hierarchical structure of deep CNNs enables rich and discriminative visual feature extraction, the recognition methods so far mostly use the features of only the last convolutional layer for classification. In this paper, by exploiting the correlation of the convolutional features of within-layer and between-layer, we propose a method to integrate multi-layer convolutional features based on coarse-to-fine mechanism for improving the discrimination capability. Experiments on a number of public datasets show that the proposed method, without part annotation or pose alignment, yields superior or comparable performance to the state-of-the-art methods. Yongzhong Wang, Xu-Yao Zhang, Yan-Ming Zhang 0001, Xinwen Hou, Cheng-Lin Liu 0001 |
ICIP | 2 |
| 2016 | Handwritten Chinese character recognition with spatial transformer and deep residual networksabstractThis paper considers using deep neural networks for handwritten Chinese character recognition (HCCR) with arbitrary position, scale, and orientations. To solve this problem, we combine the recently proposed spatial transformer network (STN) with the deep residual network (DRN). The STN acts like a character shape normalization procedure. Different from the traditional heuristic shape normalization methods, STN is learned directly from the data. Furthermore, the DRN makes the training of very deep network to be both efficient and effective. With the combination of STN and DRN, the whole model can be trained jointly in an end-to-end manner. In this paper, new state-of-the-art performance has been achieved by our proposed model on the offline ICDAR-2013 Chinese handwriting competition database. Moreover, the experiment on randomly distorted samples shows that the STN is very effective for robust HCCR in rectifying the shape of distorted characters. Zhao Zhong, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICPR | 2 |
| 2016 | Adaptive spatial pooling for image classification
Yinglu Liu, Yan-Ming Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2016 | Discriminative quadratic feature learning for handwritten Chinese character recognition
Ming-Ke Zhou, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2016 | MSDLSR: Margin Scalable Discriminative Least Squares Regression for Multicategory ClassificationabstractIn this brief, we propose a new margin scalable discriminative least squares regression (MSDLSR) model for multicategory classification. The main motivation behind the MSDLSR is to explicitly control the margin of DLSR model. We first prove that the DLSR is a relaxation of the traditional L2-support vector machine. Based on this fact, we further provide a theorem on the margin of DLSR. With this theorem, we add an explicit constraint on DLSR to restrict the number of zeros of dragging values, so as to control the margin of DLSR. The new model is called MSDLSR. Theoretically, we analyze the determination of the margin and support vectors of MSDLSR. Extensive experiments illustrate that our method outperforms the current state-of-the-art approaches on various machine leaning and real-world data sets. Lingfeng Wang 0002, Xu-Yao Zhang, Chunhong Pan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Task-Driven Feature Pooling for Image ClassificationabstractFeature pooling is an important strategy to achieve high performance in image classification. However, most pooling methods are unsupervised and heuristic. In this paper, we propose a novel task-driven pooling (TDP) model to directly learn the pooled representation from data in a discriminative manner. Different from the traditional methods (e.g., average and max pooling), TDP is an implicit pooling method which elegantly integrates the learning of representations into the given classification task. The optimization of TDP can equalize the similarities between the descriptors and the learned representation, and maximize the classification accuracy. TDP can be combined with the traditional BoW models (coding vectors) or the recent state-of-the-art CNN models (feature maps) to achieve a much better pooled representation. Furthermore, a self-training mechanism is used to generate the TDP representation for a new test image. A multi-task extension of TDP is also proposed to further improve the performance. Experiments on three databases (Flower-17, Indoor-67 and Caltech-101) well validate the effectiveness of our models. Guosen Xie, Xu-Yao Zhang, Xiangbo Shu, Shuicheng Yan, Cheng-Lin Liu 0001 |
ICCV | 2 |
| 2015 | Retargeted Least Squares Regression AlgorithmabstractThis brief presents a framework of retargeted least squares regression (ReLSR) for multicategory classification. The core idea is to directly learn the regression targets from data other than using the traditional zero-one matrix as regression targets. The learned target matrix can guarantee a large margin constraint for the requirement of correct classification for each data point. Compared with the traditional least squares regression (LSR) and a recently proposed discriminative LSR models, ReLSR is much more accurate in measuring the classification error of the regression model. Furthermore, ReLSR is a single and compact model, hence there is no need to train two-class (binary) machines that are independent of each other. The convex optimization problem of ReLSR is solved elegantly and efficiently with an alternating procedure including regression and retargeting as substeps. The experimental evaluation over a range of databases identifies the validity of our method. Xu-Yao Zhang, Lingfeng Wang 0002, Shiming Xiang, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Efficient Feature Coding Based on Auto-encoder Network for Image Classification
Guosen Xie, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ACCV (1) | 2 |
| 2014 | Effective license plate detection using fast candidate region selection and covariance feature based filteringabstractThis paper presents a new real-time license plate detection method aiming for fast and accurate detection in live videos. Compared with the previous learning based detection schemes which scan multi-scale images with sliding window, our method takes a cascaded scheme. In the first stage, candidate plate regions are detected based on edge density in reduced image of very low resolution for guaranteeing high speed. In the second stage, the candidate regions are verified using a linear SVM classifier with covariance features for high accuracy. Experimental results on two datasets collected from practical traffic surveillance videos indicate the robustness of our method, which is relatively invariant to scaling, rotation, blurring and illumination. This method takes only 10 msec for detection on a 768 × 576 image. Bo-Yuan Feng, Mingwu Ren, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
AVSS | 3 |
| 2014 | Improving Handwritten Chinese Character Recognition with Discriminative Quadratic Feature ExtractionabstractDiscriminative feature extraction (DFE) is an effective linear dimensionality reduction method for pattern recognition. It improves the recognition performance via optimizing subspace projection axes and classifier parameters simultaneously. In this paper, we propose a nonlinear extension of DFE, called discriminative quadratic feature extraction (DQFE), for which feature vectors are firstly mapped to a high-dimensional nonlinear space and then projected to a low-dimensional subspace learned by DFE. The nonlinear mapping is obtained by adding quadratic (correlation or covariance) features computed directly on the original gradient feature maps with different region partition. In this way, both the structural information of the image and the correlation information of features are used to generate a nonlinear high-dimensional feature mapping (thousands of dimensions). Experimental results demonstrated that DQFE can improve the accuracy for different classifiers in handwritten Chinese character recognition. Ming-Ke Zhou, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICPR | 2 |
| 2014 | Integrating supervised subspace criteria with restricted Boltzmann Machine for feature extractionabstractRestricted Boltzmann Machine (RBM) is a widely used building-block in deep neural networks. However, RBM is an unsupervised model which can not exploit the rich supervised information of data. Therefore, we consider combining the descriptive (generative) ability of RBM with the discriminative ability of supervised subspace models, i.e., Fisher linear discriminant analysis (FDA), marginal Fisher analysis (MFA), and heat kernel MFA (hkMFA). Specifically, the hidden layer of RBM is regularized by the supervised subspace criteria, and the joint learning model can then be efficiently optimized by gradient descent and graph construction (used to define the scatter matrix in the subspace models) on mini-batch data. Compared with the traditional subspace models (FDA, MFA, hkMFA), the proposed hybrid models are essentially nonlinear and can be optimized by gradient descent instead of eigenvalue decomposition. More importantly, traditional subspace models can only reduce the dimensionality (because of linear transformation), while the proposed models can also increase the dimensionality for better class discrimination. Experiments on three databases demonstrate that the proposed hybrid models outperform both RBM and their counterpart subspace models (FDA, MFA, hkMFA) consistently. Guosen Xie, Xu-Yao Zhang, Yan-Ming Zhang 0001, Cheng-Lin Liu 0001 |
IJCNN | 2 |
| 2014 | Automatic recognition of serial numbers in bank notes
Bo-Yuan Feng, Mingwu Ren, Xu-Yao Zhang, Ching Y. Suen |
Pattern Recognit. | 3 |
| 2014 | Combination of Classification and Clustering Results with Label PropagationabstractThis letter considers the combination of multiple classification and clustering results to improve the prediction accuracy. First, an object-similarity graph is constructed from multiple clustering results. The labels predicted by the classification models are then propagated on this graph to adaptively satisfy the smoothness of the prediction over the graph. The convex learning problem is efficiently solved by the label propagation algorithm. A semi-supervised extension is also provided to further improve the performance. Experiments on 11 tasks identify the validity of the proposed models. Xu-Yao Zhang, Peipei Yang, Yan-Ming Zhang 0001, Kaizhu Huang, Cheng-Lin Liu 0001 |
IEEE Signal Process. Lett. | 1 |
| 2013 | Extraction of Serial Numbers on Bank NotesabstractThe study of RMB (renminbi bank note, the paper currency used in China) serial number recognition draws more and more attention in recent years, for reducing financial crime, improving financial market stability and social security. The accuracy of RMB recognition relies heavily on the extraction, which is a challenging problem due to background variations and uneven illumination. In this paper, we present a new system that extracts the RMB characters directly from scanned RMB images. First, two different techniques, namely skew correction and orientation identification are used to detect the region which contains RMB serial number. Then the detected text region is binarized by a combined thresholding technique. After that, a local contrast average method is introduced to extract the RMB characters from the binarization result. The experiments demonstrate that the proposed binarization method outperforms other well-known methods. For character extraction, we report an overlap-recall rate of 79.68% and an overlap-precision rate of 98.10% respectively. Bo-Yuan Feng, Mingwu Ren, Xu-Yao Zhang, Ching Y. Suen |
ICDAR | 3 |
| 2013 | ICDAR 2013 Chinese Handwriting Recognition CompetitionabstractThis paper describes the Chinese handwriting recognition competition held at the 12th International Conference on Document Analysis and Recognition (ICDAR 2013). This third competition in the series again used the CASIA-HWDB/OLHWDB databases as the training set, and all the submitted systems were evaluated on closed datasets to report character-level correct rates. This year, 10 groups submitted 27 systems for five tasks: classification on extracted features, online/offline isolated character recognition, online/offline handwritten text recognition. The best results (correct rates) are 93.89% for classification on extracted features, 94.77% for offline character recognition, 97.39% for online character recognition, 88.76% for offline text recognition, and 95.03% for online text recognition, respectively. In addition to the test results, we also provide short descriptions of the recognition methods and brief discussions on the results. Qiufeng Wang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICDAR | 3 |
| 2013 | Locally Smoothed Modified Quadratic Discriminant FunctionabstractModified quadratic discriminant function (MQDF) is a state-of-the-art classifier for handwriting recognition. However, the big gap between accuracies on training and testing sets indicates that MQDF has a good capability to fit training data but the generalization performance is not promising. To solve this problem, we propose a new model called locally smoothed modified quadratic discriminant function (LSMQDF) by smoothing the covariance matrix of each class with its nearest neighbor classes. LSMQDF can be viewed as a regularization to avoid over-fitting. The covariance matrix estimated by local smoothing is more accurate and robust. LSMQDF can be also viewed as an extension of the global smoothing method, namely regularized discriminant analysis (RDA). Experiments on both offline and online Chinese handwriting databases demonstrate that: with local smoothing, the accuracy on training set is decreased (over-fitting avoided), and the accuracy on testing set is improved significantly and consistently (generalization improved). Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICDAR | 1 |
| 2013 | Feature Transformation with Class Conditional DecorrelationabstractThe well-known feature transformation model of Fisher linear discriminant analysis (FDA) can be decomposed into an equivalent two-step approach: whitening followed by principal component analysis (PCA) in the whitened space. By proving that whitening is the optimal linear transformation to the Euclidean space in the sense of minimum log-determinant divergence, we propose a transformation model called class conditional decor relation (CCD). The objective of CCD is to diagonalize the covariance matrices of different classes simultaneously, which is efficiently optimized using a modified Jacobi method. CCD is effective to find the common principal components among multiple classes. After CCD, the variables become class conditionally uncorrelated, which will benefit the subsequent classification tasks. Combining CCD with the nearest class mean (NCM) classification model can significantly improve the classification accuracy. Experiments on 15 small-scale datasets and one large-scale dataset (with 3755 classes) demonstrate the scalability of CCD for different applications. We also discuss the potential applications of CCD for other problems such as Gaussian mixture models and classifier ensemble learning. Xu-Yao Zhang, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICDM | 1 |
| 2013 | Writer Adaptation with Style Transfer MappingabstractAdapting a writer-independent classifier toward the unique handwriting style of a particular writer has the potential to significantly increase accuracy for personalized handwriting recognition. This paper proposes a novel framework of style transfer mapping (STM) for writer adaptation. The STM is a writer-specific class-independent feature transformation which has a closed-form solution. After style transfer mapping, the data of different writers are projected onto a style-free space, where the writer-independent classifier needs no change to classify the transformed data and can achieve significantly higher accuracy. The framework of STM can be combined with different types of classifiers for supervised, unsupervised, and semi-supervised adaptation, where writer-specific data can be either labeled or unlabeled and need not cover all classes. In this paper, we combine STM with the state-of-the-art classifiers for large-category Chinese handwriting recognition: learning vector quantization (LVQ) and modified quadratic discriminant function (MQDF). Experiments on the online Chinese handwriting database CASIA-OLHWDB demonstrate that STM-based adaptation is very efficient and effective in improving classification accuracy. Semi-supervised adaptation achieves the best performance, while unsupervised adaptation is even better than supervised adaptation. On handwritten text data, semi-supervised adaptation achieves error reduction rates 31.95 and 25.00 percent by LVQ and MQDF, respectively. Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Evaluation of weighted Fisher criteria for large category dimensionality reduction in application to Chinese handwriting recognition
Xu-Yao Zhang, Cheng-Lin Liu 0001 |
Pattern Recognit. | 1 |
| 2012 | Confused Distance Maximization for Large Category Dimensionality ReductionabstractThe Fisher linear discriminant analysis (FDA) is the most well-known supervised dimensionality reduction model. However, when the number of classes is much larger than the reduced dimensionality, FDA suffers from the class separation problem in that it will preserve the distances of the already well-separated classes and cause a large overlap of neighboring classes. To cope with this problem, we propose a new model called confused distance maximization (CDM). The objective of CDM is to maximize the distance of the most confusable classes, according to the confusion matrix estimated from the training data with a pre-learned classifier. Compared with FDA that maximizes the sum of the distances of all class pairs, CDM is more relevant to the classification accuracy by weighting the pairwise distance according to the confusion matrix. Furthermore, CDM is computationally inexpensive which makes it indeed efficient and effective for large category problems. Experiments on two large-scale 3,755-class Chinese handwriting databases (offline and online) demonstrate that CDM can achieve the best performance compared with FDA and other competitive weighting based criteria. Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ICFHR | 1 |
| 2012 | Multiple Outlooks Learning with Support Vector Machines
Yinglu Liu, Xu-Yao Zhang, Kaizhu Huang, Xinwen Hou, Cheng-Lin Liu 0001 |
ICONIP (3) | 2 |
| 2012 | Manifold Regularized Multi-Task Learning
Peipei Yang, Xu-Yao Zhang, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICONIP (3) | 2 |
| 2011 | Style transfer matrix learning for writer adaptationabstractIn this paper, we propose a novel framework of style transfer matrix (STM) learning to reduce the writing style variation in handwriting recognition. After writer-specific style transfer learning, the data of different writers is projected onto a style-free space, where a writer independent classifier can yield high accuracy. We combine STM learning with a specific nearest prototype classifier: learning vector quantization (LVQ) with discriminative feature extraction (DFE), where both the prototypes and the subspace transformation matrix are learned via online discriminative learning. To adapt the basic classifier (trained with writer-independent data) to particular writers, we first propose two supervised models, one based on incremental learning and the other based on supervised STM learning. To overcome the lack of labeled samples for particular writers, we propose an unsupervised model to learn the STM using the self-taught strategy (also known as self-training). Experiments on a large-scale Chinese online handwriting database demonstrate that STM learning can reduce recognition errors significantly, and the unsupervised adaptation model performs even better than the supervised models. Xu-Yao Zhang, Cheng-Lin Liu 0001 |
CVPR | 1 |
| 2011 | Perceptron Learning of Modified Quadratic Discriminant FunctionabstractModified quadratic discriminant function (MQDF) is the state-of-the-art classifier in handwritten character recognition. Discriminative learning of MQDF can further improve its performance. Recent advances justify the efficacy of minimum classification error criteria in learning MQDF (MCE-MQDF). We provide an alternative choice to MCE-MQDF based on the Perceptron learning (PL-MQDF). For better generalization performance, we propose a new dynamic margin regularization. To relieve the heavy burden in training process, active set technique is employed, which can save most of the computation with negligible loss in accuracy. In experiments on handwritten digit datasets and a large-scale Chinese handwritten character database, the proposed PL-MQDF was demonstrated superior in both error reduction and training speedup. Tong-Hua Su, Cheng-Lin Liu 0001, Xu-Yao Zhang |
ICDAR | 3 |
| 2011 | Pattern Field Classification with Style Normalized Transformation
Xu-Yao Zhang, Kaizhu Huang, Cheng-Lin Liu 0001 |
IJCAI | 1 |