VLDB 2026 Research / reviewers in the wild / expert
Jianshu Zhang 0001
dblp:65/9878-1
· DBLP profile ↗
41ranked-venue papers
10as first author
25since 2021 · last 2026
0000-0002-2713-2535ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | See then tell: Enhancing key information extraction with vision grounding
Shuhang Liu, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Qing Wang 0008, Jianshu Zhang 0001 |
Neurocomputing | 7 |
| 2026 | Two-stage decomposition network for handwritten Chinese character error correction
Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001, Jianqing Gao, Qingfeng Liu |
Pattern Recognit. | 5 |
| 2026 | Reinforcement learning-powered co-optimization: Bridging critic model and multimodal LLM reasoning abilities
Qing Wang 0008, Shuhang Liu, Jun Du 0002, Jianshu Zhang 0001 |
Pattern Recognit. | 5 |
| 2025 | DocMamba: Efficient Document Pre-training with State Space ModelabstractIn recent years, visually-rich document understanding has attracted increasing attention. Transformer-based pre-trained models have become the mainstream approach, yielding significant performance gains in this field. However, the self-attention mechanism's quadratic computational complexity hinders their efficiency and ability to process long documents. In this paper, we present DocMamba, a novel framework based on the state space model. It is designed to reduce computational complexity to linear while preserving global modeling capabilities. To further enhance its effectiveness in document processing, we introduce the Segment-First Bidirectional Scan (SFBS) to capture contiguous semantic information. Experimental results demonstrate that DocMamba achieves new state-of-the-art results on downstream datasets such as FUNSD, CORD, and SORIE, while significantly improving speed and reducing memory usage. Notably, experiments on the HRDoc confirm DocMamba's potential for length extrapolation. Pengfei Hu 0006, Jiefeng Ma, Shuhang Liu, Jun Du 0002, Jianshu Zhang 0001 |
AAAI | 6 |
| 2025 | Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural IntegrationabstractRecent advances in Multimodal Large Language Models (MLLMs) have achieved remarkable progress in general domains and demonstrated promise in multimodal mathematical reasoning. However, applying MLLMs to geometry problem solving (GPS) remains challenging due to lack of accurate step-by-step solution data and severe hallucinations during reasoning. In this paper, we propose GeoGen, a pipeline that can automatically generates step-wise reasoning paths for geometry diagrams. By leveraging the precise symbolic reasoning, GeoGen produces large-scale, high-quality question-answer pairs. To further enhance the logical reasoning ability of MLLMs, we train GeoLogic, a Large Language Model (LLM) using synthetic data generated by GeoGen. Serving as a bridge between natural language and symbolic systems, GeoLogic enables symbolic tools to help verifying MLLM outputs, making the reasoning process more rigorous and alleviating hallucinations. Experimental results show that our approach consistently improves the performance of MLLMs, achieving remarkable results on benchmarks for geometric reasoning tasks. This improvement stems from our integration of the strengths of LLMs and symbolic systems, which enables a more reliable and interpretable approach for the GPS task. Codes are available at https://github.com/ycpNotFound/GeoGen. Yicheng Pan 0004, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001, Jianqing Gao |
ACM Multimedia | 6 |
| 2025 | Count, decompose and correct: A new approach to handwritten Chinese character error correction
Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001 |
Pattern Recognit. | 5 |
| 2024 | Maths: Multimodal Transformer-Based Human-Readable SolverabstractMultimodal mathematical reasoning has gained increasing attention in recent times. However, previous effective methods have not tried to reason in the form of natural language. In this paper, we introduce a model named MATHS (MultimodAl Transformer-based Human-readable Solver) for visual arithmetic and geometry problems in multimodal mathematical reasoning tasks. Drawing inspiration from Multimodal Large Language Models (MLLMs), our approach involves generating problem-solving processes expressed in natural language, in order to leverage the inherent reasoning capabilities embedded within language models. To address the challenge of precise calculations for language models, our work proposes a Math-Constrained Generation (MCG) method to impose hard constraints on generated outputs. Extensive experiments demonstrate our model excels in visual arithmetic task, and achieves results that are either better or comparable to existing methods in geometry problems. Code is available at https://github.com/ycpNotFound/MATHS. Yicheng Pan 0004, Jiefeng Ma, Pengfei Hu 0006, Jun Du 0002, Qing Wang 0008, Jianshu Zhang 0001, Dan Liu 0008, Si Wei |
ICME | 7 |
| 2024 | SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form UnderstandingabstractAccurately identifying and organizing textual content is crucial for the automation of document processing in the field of form understanding. Existing datasets, such as FUNSD and XFUND, support entity classification and relationship prediction tasks but are typically limited to local and entity-level annotations. This limitation overlooks the hierarchically structured representation of documents, constraining comprehensive understanding of complex forms. To address this issue, we present the SRFUND, a hierarchically structured multi-task form understanding benchmark. SRFUND provides refined annotations on top of the original FUNSD and XFUND datasets, encompassing five tasks: (1) word to text-line merging, (2) text-line to entity merging, (3) entity category classification, (4) item table localization, and (5) entity-based full-document hierarchical structure recovery. We meticulously supplemented the original dataset with missing annotations at various levels of granularity and added detailed annotations for multi-item table regions within the forms. Additionally, we introduce global hierarchical structure dependencies for entity relation prediction tasks, surpassing traditional local key-value associations. The SRFUND dataset includes eight languages including English, Chinese, Japanese, German, French, Spanish, Italian, and Portuguese, making it a powerful tool for cross-lingual form understanding. Extensive experimental results demonstrate that the SRFUND dataset presents new challenges and significant opportunities in handling diverse layouts and global hierarchical structures of forms, thus providing deep insights into the field of form understanding. The original dataset and implementations of baseline methods are available at https://sprateam-ustc.github.io/SRFUND. Jiefeng Ma, Jun Du 0002, Yu Hu 0003, Pengfei Hu 0006, Qing Wang 0008, Jianshu Zhang 0001 |
NeurIPS | 9 |
| 2024 | SEMv2: Table separation line detection based on instance segmentation
Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001, Cong Liu 0006 |
Pattern Recognit. | 5 |
| 2023 | HRDoc: Dataset and Baseline Method toward Hierarchical Reconstruction of Document StructuresabstractThe problem of document structure reconstruction refers to converting digital or scanned documents into corresponding semantic structures. Most existing works mainly focus on splitting the boundary of each element in a single document page, neglecting the reconstruction of semantic structure in multi-page documents. This paper introduces hierarchical reconstruction of document structures as a novel task suitable for NLP and CV fields. To better evaluate the system performance on the new task, we built a large-scale dataset named HRDoc, which consists of 2,500 multi-page documents with nearly 2 million semantic units. Every document in HRDoc has line-level annotations including categories and relations obtained from rule-based extractors and human annotators. Moreover, we proposed an encoder-decoder-based hierarchical document structure parsing system (DSPS) to tackle this problem. By adopting a multi-modal bidirectional encoder and a structure-aware GRU decoder with soft-mask operation, the DSPS model surpass the baseline method by a large margin. All scripts and datasets will be made publicly available at https://github.com/jfma-USTC/HRDoc. Jiefeng Ma, Jun Du 0002, Pengfei Hu 0006, Jianshu Zhang 0001, Cong Liu 0006 |
AAAI | 5 |
| 2023 | Group, Contrast and Recognize: A Self-supervised Method for Chinese Character Recognition
Xinzhe Jiang, Jun Du 0002, Pengfei Hu 0006, Mobai Xue, Jiefeng Ma, Jiajia Wu 0003, Jianshu Zhang 0001 |
ICDAR (4) | 7 |
| 2023 | A Tree-Structure Analysis Network on Handwritten Chinese Character Error CorrectionabstractExisting researches on handwritten Chinese characters are mainly based on recognition network designed to solve the complex structure and numerous amount characteristics of Chinese characters. In this paper, we investigate Chinese characters from the perspective of error correction, which is to diagnose a handwritten character to be right or wrong and provide a feedback on error analysis. For this handwritten Chinese character error correction task, we define a benchmark by unifying both the evaluation metrics and data splits for the first time. Then we design a diagnosis system that includes decomposition, judgement and correction stages. Specifically, a novel tree-structure analysis network (TAN) is proposed to model a Chinese character as a tree layout, which mainly consists of a CNN-based encoder and a tree-structure based decoder. Using the predicted tree layout for judgement, correction operation is performed for the wrongly written characters to do error analysis. The correction stage is composed of three steps: fetch the ideal character, correct the errors and locate the errors. Additionally, we propose a novel bucketing mining strategy to apply triplet loss at radical level to alleviate feature dispersion. Experiments on handwritten character dataset demonstrate that our proposed TAN shows great superiority on all three metrics comparing with other state-of-the-art recognition models. Through quantitative analysis, TAN is proved to capture more accurate spatial position information than regular encoder-decoder models, showing better generalization ability. Jun Du 0002, Jianshu Zhang 0001, Changjie Wu |
IEEE Trans. Multim. | 3 |
| 2023 | Multimodal Pre-Training Based on Graph Attention Network for Document UnderstandingabstractDocument intelligence as a relatively new research topic supports many business applications. Its main task is to automatically read, understand, and analyze documents. However, due to the diversity of formats (invoices, reports, forms, etc.) and layouts in documents, it is difficult to make machines understand documents. In this paper, we present the GraphDoc, a multimodal graph attention-based model for various document understanding tasks. GraphDoc is pre-trained in a multimodal framework by utilizing text, layout, and image information simultaneously. In a document, a text block relies heavily on its surrounding contexts, accordingly we inject the graph structure into the attention mechanism to form a graph attention layer so that each input node can only attend to its neighborhoods. The input nodes of each graph attention layer are composed of textual, visual, and positional features from semantically meaningful regions in a document image. We do the multimodal feature fusion of each node by the gate fusion layer. The contextualization between each node is modeled by the graph attention layer. GraphDoc learns a generic representation from only 320k unlabeled documents via the Masked Sentence Modeling task. Extensive experimental results on the publicly available datasets show that GraphDoc achieves state-of-the-art performance, which demonstrates the effectiveness of our proposed method. Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | TDv2: A Novel Tree-Structured Decoder for Offline Mathematical Expression RecognitionabstractIn recent years, tree decoders become more popular than LaTeX string decoders in the field of handwritten mathematical expression recognition (HMER) as they can capture the hierarchical tree structure of mathematical expressions. However previous tree decoders converted the tree structure labels into a fixed and ordered sequence, which could not make full use of the diversified expression of tree labels. In this study, we propose a novel tree decoder (TDv2) to fully utilize the tree structure labels. Compared with previous tree decoders, this new model does not require a fixed priority for different branches of a node during training and inference, which can effectively improve the model generalization capability. The input and output of the model make full use of the tree structure label, so that there is no need to find the parent node in the decoding process, which simplifies the decoding process and adds a prior information to help predict the node. We verified the effectiveness of each part of the model through comprehensive ablation experiments and attention visualization analysis. On the authoritative CROHME 14/16/19 datasets, our method achieves the state-of-the-art results. Changjie Wu, Jun Du 0002, Jianshu Zhang 0001, Bo Ren 0002, Yiqing Hu |
AAAI | 4 |
| 2022 | Improving Isolated Glyph Classification Task for Palm Leaf Manuscripts
Nimol Thuon, Jun Du 0002, Jianshu Zhang 0001 |
ICFHR | 3 |
| 2022 | Learning Contextually Fused Audio-Visual Representations For Audio-Visual Speech RecognitionabstractWith the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech recognition (AVSR) performance, as the multi-modal inputs contain more fruitful information in principle. In this paper, based on existing self-supervised representation learning methods for audio modality, we therefore propose an audio-visual representation learning approach. The proposed approach explores both the complementarity of audio-visual modalities and long-term context dependency using a transformer-based fusion module and a flexible masking strategy. After pre-training, the model is able to extract fused representations required by AVSR. Without loss of generality, it can be applied to single-modal tasks, e.g., audio/visual speech recognition by simply masking out one modality in the fusion module. The proposed pre-trained model is evaluated on speech recognition and lipreading tasks using one or two modalities, where the superiority is revealed. Jie Zhang 0042, Jianshu Zhang 0001, Ming-Hui Wu, Li-Rong Dai 0001 |
ICIP | 3 |
| 2022 | Multimodal Tree Decoder for Table of Contents Extraction in Document ImagesabstractTable of contents (ToC) extraction aims to extract headings of different levels in documents to better understand the outline of the contents, which can be widely used for document understanding and information retrieval. Existing works often use hand-crafted features and predefined rule-based functions to detect headings and resolve the hierarchical relationship between headings. Both the benchmark and research based on deep learning are still limited. Accordingly, in this paper, we first introduce a standard dataset, HierDoc, including image samples from 650 documents of scientific papers with their content labels. Then we propose a novel end-to-end model by using the multimodal tree decoder (MTD) for ToC as a benchmark for HierDoc. The MTD model is mainly composed of three parts, namely encoder, classifier, and decoder. The encoder fuses the multimodality features of vision, text, and layout information for each entity of the document. Then the classifier recognizes and selects the heading entities. Next, to parse the hierarchical relationship between the heading entities, a tree-structured decoder is designed. To evaluate the performance, both the metric of tree-edit-distance similarity (TEDS) and F1-Measure are adopted. Finally, our MTD approach achieves an average TEDS of 87.2% and an average F1-Measure of 88.1% on the test set of HierDoc. The code and dataset will be released at: https://github.com/Pengfei-Hu/MTD. Pengfei Hu 0006, Jianshu Zhang 0001, Jun Du 0002, Jiajia Wu 0003 |
ICPR | 3 |
| 2022 | Scene Text Recognition with Self-supervised Contrastive Predictive CodingabstractSelf-supervised visual pre-training has recently emerged in scene text recognition (STR), which designs the pretext tasks and takes unlabeled data as input to obtain useful representations for STR. However, most current self-supervised methods do not pay special attention to the importance of sequence awareness. Accordingly, we propose a novel self-supervised STR method based on contrastive predictive coding (STR-CPC), which regards a text instance as a sequence from left to right and captures the visual sequence correlation. Considering the information overlap problem within the feature map induced by the deep convolutional neural network (CNN) encoder, we design a widthwise causal convolution during model pre-training and a progressive recovery training strategy (PRTS) during model fine-tuning to improve the STR performance. Experiments on scene text show that our STR-CPC method outperforms the existing self-supervised methods, which testifies the advantage of visual sequence correlation for STR. Additionally, STR-CPC observably boosts performance compared with supervised training when the amount of labeled data decreases. Xinzhe Jiang, Jianshu Zhang 0001, Jun Du 0002, Jiajia Wu 0003 |
ICPR | 2 |
| 2022 | A multimodal attention fusion network with a dynamic vocabulary for TextVQA
Jiajia Wu 0003, Jun Du 0002, Fengren Wang, Xinzhe Jiang, Jinshui Hu, Jianshu Zhang 0001, Li-Rong Dai 0001 |
Pattern Recognit. | 8 |
| 2022 | Tree-based data augmentation and mutual learning for offline handwritten mathematical expression recognition
Jun Du 0002, Jianshu Zhang 0001, Changjie Wu, Mingjun Chen, Jiajia Wu 0003 |
Pattern Recognit. | 3 |
| 2022 | Split, Embed and Merge: An accurate table structure recognizer
Jianshu Zhang 0001, Jun Du 0002, Fengren Wang |
Pattern Recognit. | 2 |
| 2021 | MRD: A Memory Relation Decoder for Online Handwritten Mathematical Expression Recognition
Qing Wang 0008, Jun Du 0002, Jianshu Zhang 0001, Bin Wang 0070, Bo Ren 0002 |
ICDAR (3) | 4 |
| 2021 | Radical Composition Network for Chinese Character Generation
Mobai Xue, Jun Du 0002, Jianshu Zhang 0001, Zi-Rui Wang, Bin Wang 0070, Bo Ren 0002 |
ICDAR (1) | 3 |
| 2021 | Stroke constrained attention network for online handwritten mathematical expression recognition
Jun Du 0002, Jianshu Zhang 0001, Bin Wang 0070, Bo Ren 0002 |
Pattern Recognit. | 3 |
| 2021 | SRD: A Tree Structure Based Decoder for Online Handwritten Mathematical Expression RecognitionabstractRecently, recognition of online handwritten mathe- matical expression has been greatly improved by employing encoder-decoder based methods. Existing encoder-decoder models use string decoders to generate LaTeX strings for mathematical expression recognition. However, in this paper, we importantly argue that string representations might not be the most natural for mathematical expressions – mathematical expressions are inherently tree structures other than flat strings. For this purpose, we propose a novel sequential relation decoder (SRD) that aims to decode expressions into tree structures for online handwritten mathematical expression recognition. At each step of tree construction, a sub-tree structure composed of a relation node and two symbol nodes is computed based on previous sub-tree structures. This is the first work that builds a tree structure based decoder for encoder-decoder based mathematical expression recognition. Compared with string decoders, a decoder that better understands tree structures is crucial for mathematical expression recognition as it brings a more reasonable learning objective and improves overall generalization ability. We demonstrate how the proposed SRD outperforms state-of-the-art string decoders through a set of experiments on CROHME database, which is currently the largest benchmark for online handwritten mathematical expression recognition. Jianshu Zhang 0001, Jun Du 0002, Yongxin Yang, Yi-Zhe Song, Li-Rong Dai 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | A Tree-Structured Decoder for Image-to-Markup GenerationabstractRecent encoder-decoder approaches typically employ string decoders to convert images into serialized strings for image-to-markup. However, for tree-structured representational markup, string representations can hardly cope with the structural complexity. In this work, we first show via a set of toy problems that string decoders struggle to decode tree structures, especially as structural complexity increases, we then propose a tree-structured decoder that specifically aims at generating a tree-structured markup. Our decoders works sequentially, where at each step a child node and its parent node are simultaneously generated to form a sub-tree. This sub-tree is consequently used to construct the final tree structure in a recurrent manner. Key to the success of our tree decoder is twofold, (i) it strictly respects the parent-child relationship of trees, and (ii) it explicitly outputs trees as oppose to a linear string. Evaluated on both math formula recognition and chemical formula recognition, the proposed tree decoder is shown to greatly outperform strong string decoder baselines. Jianshu Zhang 0001, Jun Du 0002, Yongxin Yang, Yi-Zhe Song, Si Wei, Li-Rong Dai 0001 |
ICML | 1 |
| 2020 | Radical Counter Network for Robust Chinese Character RecognitionabstractChinese character recognition has attracted much interest due to its high challenge and various applications. The whole-character modeling method can recognize common characters well but unable to handle unseen situation. Some radical-based modeling methods have successfully achieved great performance in unseen condition but need RNN-based decoder for sequence decoding. Therefore, a compact model which can recognize unseen characters needs to be proposed. First, this paper introduces a novel radical counter network (RCN) to recognize Chinese characters by identifying radicals and spatial structures. The proposed RCN first extracts visual features from input by employing DenseNet as encoder. Then a decoder based on fully connected layer is employed, aiming at synchronously estimating the number of each caption in character. Additionally, we design a multi-task learning to combine global feature extraction capability of whole-character modeling and local feature extraction capability of radical-based modeling, which further improves the model generalization. Experiments on natural scene character dataset demonstrate that the proposed model significantly outperforms WCN by 5.48% and achieve comparable performance with RAN in lower model complexity. That shows great robustness and simplicity of our model. Yixing Zhu, Jun Du 0002, Changjie Wu, Jianshu Zhang 0001 |
ICPR | 5 |
| 2020 | Stroke Based Posterior Attention for Online Handwritten Mathematical Expression RecognitionabstractRecently, many researches propose to employ attention based encoder-decoder models to convert a sequence of trajectory points into a LaTeX string for online handwritten mathematical expression recognition (OHMER), and the recognition performance of these models critically relies on the accuracy of the attention. In this paper, unlike previous methods which basically employ a soft attention model, we propose to employ a posterior attention model, which modifies the attention probabilities after observing the output probabilities generated by the soft attention model. In order to further improve the posterior attention mechanism, we propose a stroke average pooling layer to aggregate point-level features obtained from the encoder into stroke-level features. We argue that posterior attention is better to be implemented on stroke-level features than point-level features as the output probabilities generated by stroke is more convincing than generated by point, and we prove that through experimental analysis. Validated on the CROHME competition task, we demonstrate that stroke based posterior attention achieves expression recognition rates of 54.26% on CROHME 2014 and 51.75% on CROHME 2016. According to attention visualization analysis, we empirically demonstrate that the posterior attention mechanism can achieve better alignment accuracy than the soft attention mechanism. Changjie Wu, Qing Wang 0008, Jianshu Zhang 0001, Jun Du 0002, Jiajia Wu 0003, Jin-Shui Hu |
ICPR | 3 |
| 2020 | A Transformer-based Radical Analysis Network for Chinese Character RecognitionabstractRecently, a novel radical analysis network (RAN) has the capability of effectively recognizing unseen Chinese character classes and largely reducing the requirement of training data by treating a Chinese character as a hierarchical composition of radicals rather than a single character class. However, when dealing with more challenging issues, such as the recognition of complicated characters, low-frequency character categories, and characters in natural scenes, RAN still has a lot of room for improvement. In this paper, we explore options to further improve the structure generalization and robustness capability of RAN with the Transformer architecture, which has achieved start-of-the-art results for many sequence-to-sequence tasks. More specifically, we propose to replace the original attention module in RAN with the transformer decoder, which is named as a transformer-based radical analysis network (RTN). The experimental results show that the proposed approach can significantly outperform the RAN on both printed Chinese character database and natural scene Chinese character database. Meanwhile, further analysis proves that RTN can be better generalized to complex samples and low-frequency characters, and has better robustness in recognizing Chinese characters with different attributes. Qing Wang 0008, Jun Du 0002, Jianshu Zhang 0001, Changjie Wu |
ICPR | 4 |
| 2020 | Semi-Supervised End-to-End ASR via Teacher-Student Learning with Conditional Posterior DistributionabstractEncoder-decoder based methods have become popular for automatic speech recognition (ASR), thanks to their simplified processing stages and low reliance on prior knowledge. However, large amounts of acoustic data with paired transcriptions is generally required to train an effective encoder-decoder model, which is expensive, time-consuming to be collected and not always readily available. However unpaired speech data is abundant, hence several semi-supervised learning methods, such as teacher-student (T/S) learning and pseudo-labeling, have recently been proposed to utilize this potentially valuable resource. In this paper, a novel T/S learning with conditional posterior distribution for encoder-decoder based ASR is proposed. Specifically, the 1-best hypotheses and the conditional posterior distribution from the teacher are exploited to provide more effective supervision. Combined with model perturbation techniques, the proposed method reduces WER by 19.2% relatively on the LibriSpeech benchmark, compared with a system trained using only paired data. This outperforms previous reported 1-best hypothesis results on the same task. Yan Song 0001, Jianshu Zhang 0001, Ian McLoughlin 0001, Li-Rong Dai 0001 |
INTERSPEECH | 3 |
| 2020 | Radical analysis network for learning hierarchies of Chinese characters
Jianshu Zhang 0001, Jun Du 0002, Li-Rong Dai 0001 |
Pattern Recognit. | 1 |
| 2019 | Episodic Training for Domain GeneralizationabstractDomain generalization (DG) is the challenging and topical problem of learning models that generalize to novel testing domains with different statistics than a set of known training domains. The simple approach of aggregating data from all source domains and training a single deep neural network end-to-end on all the data provides a surprisingly strong baseline that surpasses many prior published methods. In this paper we build on this strong baseline by designing an episodic training procedure that trains a single deep network in a way that exposes it to the domain shift that characterises a novel domain at runtime. Specifically, we decompose a deep network into feature extractor and classifier components, and then train each component by simulating it interacting with a partner who is badly tuned for the current domain. This makes both components more robust, ultimately leading to our networks producing state-of-the-art performance on three DG benchmarks. Furthermore, we consider the pervasive workflow of using an ImageNet trained CNN as a fixed feature extractor for downstream recognition tasks. Using the Visual Decathlon benchmark, we demonstrate that our episodic-DG training improves the performance of such a general purpose feature extractor by explicitly training a feature for robustness to novel problems. This shows that DG training can benefit standard practice in computer vision. Da Li 0001, Jianshu Zhang 0001, Yongxin Yang, Cong Liu 0006, Yi-Zhe Song, Timothy M. Hospedales |
ICCV | 2 |
| 2019 | Multi-modal Attention Network for Handwritten Mathematical Expression RecognitionabstractIn this paper, we propose a novel multi-modal attention network (MAN), which is based on encoder-decoder framework, for handwritten mathematical expression recognition (HMER). Here, multi-modal means two specific modalities: online and offline, where online modality employs dynamic trajectories as input and offline modality employs static images as input. More specifically, the proposed method first feeds dynamic trajectories and static images into online and offline channels of the multi-modal encoder respectively. The output of the encoder is then transferred to the multi-modal decoder to generate a LaTeX sequence as the mathematical expression recognition result. To make full use of the complementary information that comes from the two modalities, we propose a re-attention mechanism as an enhanced version of the multi-modal attention mechanism which can further improve the recognition performance. Evaluated on a benchmark published by CROHME competition, the proposed approach achieves an expression recognition accuracy of 54.05% on CROHME 2014 and 50.56% on CROHME 2016 which substantially outperforms the state-of-the-arts using the single online or offline modality. Jun Du 0002, Jianshu Zhang 0001, Zi-Rui Wang |
ICDAR | 3 |
| 2019 | Track, Attend, and Parse (TAP): An End-to-End Framework for Online Handwritten Mathematical Expression RecognitionabstractIn this paper, we introduce Track, Attend, and Parse (TAP), an end-to-end approach based on neural networks for online handwritten mathematical expression recognition (OHMER). The architecture of TAP consists of a tracker and a parser. The tracker employs a stack of bidirectional recurrent neural networks with gated recurrent units (GRU) to model the input handwritten traces, which can fully utilize the dynamic trajectory information in OHMER. Followed by the tracker, the parser adopts a GRU equipped with guided hybrid attention (GHA) to generate notations. The proposed GHA is composed of a coverage-based spatial attention, a temporal attention, and an attention guider. Moreover, we demonstrate the strong complementarity between offline information with static-image input and online information with ink-trajectory input by blending a fully convolutional networks-based watcher into TAP. Inherently, unlike traditional methods, this end-to-end framework does not require the explicit symbol segmentation and a predefined expression grammar for parsing. Validated on a benchmark published by the CROHME competition, the proposed approach outperforms the state-of-the-art methods and achieves the best reported results with an expression recognition accuracy of 61.16% on CROHME 2014 and 57.02% on CROHME 2016, using only official training dataset. Jianshu Zhang 0001, Jun Du 0002, Li-Rong Dai 0001 |
IEEE Trans. Multim. | 1 |
| 2018 | DenseRAN for Offline Handwritten Chinese Character RecognitionabstractRecently, great success has been achieved in offline handwritten Chinese character recognition by using deep learning methods. Chinese characters are mainly logographic and consist of basic radicals, however, previous research mostly treated each Chinese character as a whole without explicitly considering its internal two-dimensional structure and radicals. In this study, we propose a novel radical analysis network with densely connected architecture (DenseRAN) to analyze Chinese character radicals and its two-dimensional structures simultaneously. DenseRAN first encodes input image to high-level visual features by employing DenseNet as an encoder. Then a decoder based on recurrent neural networks is employed, aiming at generating captions of Chinese characters by detecting radicals and two-dimensional structures through attention mechanism. The manner of treating a Chinese character as a composition of two-dimensional structures and radicals can reduce the size of vocabulary and enable DenseRAN to possess the capability of recognizing unseen Chinese character classes, only if the corresponding radicals have been seen in training set. Evaluated on ICDAR-2013 competition database, the proposed approach significantly outperforms whole-character modeling approach with a relative character error rate (CER) reduction of 18.54%. Meanwhile, for the case of recognizing 3277 unseen Chinese characters in CASIA-HWDB1.2 database, DenseRAN can achieve a character accuracy of about 41% while the traditional whole-character method has no capability to handle them. Jianshu Zhang 0001, Jun Du 0002, Zi-Rui Wang, Yixing Zhu |
ICFHR | 2 |
| 2018 | Radical Analysis Network for Zero-Shot Learning in Printed Chinese Character RecognitionabstractChinese characters have a huge set of character categories, more than 20, 000 and the number is still increasing as more and more novel characters continue being created. However, the enormous characters can be decomposed into a compact set of about 500 fundamental and structural radicals. This paper introduces a novel radical analysis network (RAN) to recognize printed Chinese characters by identifying radicals and analyzing two-dimensional spatial structures among them. The proposed RAN first extracts visual features from input by employing convolutional neural networks as an encoder. Then a decoder based on recurrent neural networks is employed, aiming at generating captions of Chinese characters by detecting radicals and two-dimensional structures through a spatial attention mechanism. The manner of treating a Chinese character as a composition of radicals rather than a single character class largely reduces the size of vocabulary and enables RAN to possess the ability of recognizing unseen Chinese character classes, namely zero-shot learning. Jianshu Zhang 0001, Yixing Zhu, Jun Du 0002, Li-Rong Dai 0001 |
ICME | 1 |
| 2018 | Multi-Scale Attention with Dense Encoder for Handwritten Mathematical Expression RecognitionabstractHandwritten mathematical expression recognition is a challenging problem due to the complicated two-dimensional structures, ambiguous handwriting input and variant scales of handwritten math symbols. To settle this problem, recently we propose the attention based encoder-decoder model that recognizes mathematical expression images from two-dimensional layouts to one-dimensional LaTeX strings. In this study, we improve the encoder by employing densely connected convolutional networks as they can strengthen feature extraction and facilitate gradient propagation especially on a small training set. We also present a novel multi-scale attention model which is employed to deal with the recognition of math symbols in different scales and restore the fine-grained details dropped by pooling operations. Validated on the CROHME competition task, the proposed method significantly outperforms the state-of-the-art methods with an expression recognition accuracy of 52.8% on CROHME 2014 and 50.1% on CROHME 2016, by only using the official training dataset. Jianshu Zhang 0001, Jun Du 0002, Li-Rong Dai 0001 |
ICPR | 1 |
| 2018 | Trajectory-based Radical Analysis Network for Online Handwritten Chinese Character RecognitionabstractRecently, great progress has been made for online handwritten Chinese character recognition due to the emergence of deep learning techniques. However, previous research mostly treated each Chinese character as one class without explicitly considering its inherent structure, namely the radical components with complicated geometry. In this study, we propose a novel trajectory-based radical analysis network (TRAN) to firstly identify radicals and analyze two-dimensional structures among radicals simultaneously, then recognize Chinese characters by generating captions of them based on the analysis of their internal radicals. The proposed TRAN employs recurrent neural networks (RNNs) as both an encoder and a decoder. The RNN encoder makes full use of online information by directly transforming handwriting trajectory into high-level features. The RNN decoder aims at generating the caption by detecting radicals and spatial structures through an attention model. The manner of treating a Chinese character as a two-dimensional composition of radicals can reduce the size of vocabulary and enable TRAN to possess the capability of recognizing unseen Chinese character classes, only if the corresponding radicals have been seen. Evaluated on CASIA-OLHWDB database, the proposed approach significantly outperforms the state-of-the-art whole-character modeling approach with a relative character error rate (CER) reduction of 10%. Meanwhile, for the case of recognition of 500 unseen Chinese characters, TRAN can achieve a character accuracy of about 60 % while the traditional whole-character method has no capability to handle them. Jianshu Zhang 0001, Yixing Zhu, Jun Du 0002, Li-Rong Dai 0001 |
ICPR | 1 |
| 2017 | A GRU-Based Encoder-Decoder Approach with Attention for Online Handwritten Mathematical Expression RecognitionabstractIn this study, we present a novel end-to-end approach based on the encoder-decoder framework with the attention mechanism for online handwritten mathematical expression recognition (OHMER). First, the input two-dimensional ink trajectory information of handwritten expression is encoded via the gated recurrent unit based recurrent neural network (GRU-RNN). Then the decoder is also implemented by the GRU-RNN with a coverage-based attention model. The proposed approach can simultaneously accomplish the symbol recognition and structural analysis to output a character sequence in LaTeX format. Validated on the CROHME 2014 competition task, our approach significantly outperforms the state-of-the-art with an expression recognition accuracy of 52.43% by only using the official training dataset. Furthermore, the alignments between the input trajectories of handwritten expressions and the output LaTeX sequences are visualized by the attention mechanism to show the effectiveness of the proposed method. Jianshu Zhang 0001, Jun Du 0002, Li-Rong Dai 0001 |
ICDAR | 1 |
| 2017 | Watch, attend and parse: An end-to-end neural network based approach to handwritten mathematical expression recognition
Jianshu Zhang 0001, Jun Du 0002, Shiliang Zhang, Dan Liu 0008, Yulong Hu, Jin-Shui Hu, Si Wei, Li-Rong Dai 0001 |
Pattern Recognit. | 1 |
| 2016 | RNN-BLSTM Based Multi-Pitch Estimation
Jianshu Zhang 0001, Li-Rong Dai 0001 |
INTERSPEECH | 1 |