VLDB 2026 Research / reviewers in the wild / expert
Jiajia Wu 0003
dblp:142/3769-3
· DBLP profile ↗
11ranked-venue papers
4as first author
10since 2021 · last 2023
0000-0002-0951-4281ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Group, Contrast and Recognize: A Self-supervised Method for Chinese Character Recognition
Xinzhe Jiang, Jun Du 0002, Pengfei Hu 0006, Mobai Xue, Jiefeng Ma, Jiajia Wu 0003, Jianshu Zhang 0001 |
ICDAR (4) | 6 |
| 2023 | Vision-Language Adaptive Mutual Decoder for OOV-STR
Jinshui Hu, Qiandong Yan, Xuyang Zhu, Jiajia Wu 0003, Jun Du 0002, Li-Rong Dai 0001 |
ICIG (2) | 5 |
| 2023 | A Multimodal Text Block Segmentation Framework for Photo Translation
Jiajia Wu 0003, Anni Li, Zhengyan Yang, Cong Liu 0006, Li-Rong Dai 0001 |
ICIG (3) | 1 |
| 2023 | End-to-End Multilingual Text Recognition Based on Byte Modeling
Jiajia Wu 0003, Zhengyan Yang, Cong Liu 0006, Li-Rong Dai 0001 |
ICIG (3) | 1 |
| 2023 | Handwritten Chemical Structure Image to Structure-Specific Markup Using Random Conditional Guided DecoderabstractSatisfactory recognition performance has been achieved for simple and controllable printed molecular images. However, recognizing handwritten chemical structure images remains unresolved due to the inherent ambiguities in handwritten atoms and bonds, as well as the signifcant challenge of converting projected 2D molecular layouts into markup strings. Target to address these problems, this paper proposes an end-to-end framework for handwritten chemical structure images recognition, with novel structure-specific markup language (SSML) and random conditional guided decoder (RCGD). SSML alleviates ambiguity and complexity in Chemfig syntax by designing an innovative markup language to accurately depict molecular structures. Besides, we propose RCGD to address the issue of multiple path decoding of molecular structures, which is composed of conditional attention guidance, memory classification and path selection mechanisms. In order to fully confirm the effectiveness of the end-to-end method, a new database containing 50,000 handwritten chemical structure images (EDU-CHEMC) has been established. Experimental results demonstrate that compared to traditional SMILES sequences, our SSML can significantly reduces the semantic gap between chemical images and markup strings. It is worth noting that our method can also recognize invalid or non-existent organic molecular structures, making it highly applicable for tasks related to teaching evaluations in the fields of chemistry and biology education. The EDU-CHEMC will be released soon in https://github.com/iFLYTEK-CV/EDU-CHEMC. Jinshui Hu, Hao Wu 0090, Mingjun Chen, Jiajia Wu 0003, Cong Liu 0006, Jun Du 0002, Li-Rong Dai 0001 |
ACM Multimedia | 5 |
| 2022 | Multimodal Tree Decoder for Table of Contents Extraction in Document ImagesabstractTable of contents (ToC) extraction aims to extract headings of different levels in documents to better understand the outline of the contents, which can be widely used for document understanding and information retrieval. Existing works often use hand-crafted features and predefined rule-based functions to detect headings and resolve the hierarchical relationship between headings. Both the benchmark and research based on deep learning are still limited. Accordingly, in this paper, we first introduce a standard dataset, HierDoc, including image samples from 650 documents of scientific papers with their content labels. Then we propose a novel end-to-end model by using the multimodal tree decoder (MTD) for ToC as a benchmark for HierDoc. The MTD model is mainly composed of three parts, namely encoder, classifier, and decoder. The encoder fuses the multimodality features of vision, text, and layout information for each entity of the document. Then the classifier recognizes and selects the heading entities. Next, to parse the hierarchical relationship between the heading entities, a tree-structured decoder is designed. To evaluate the performance, both the metric of tree-edit-distance similarity (TEDS) and F1-Measure are adopted. Finally, our MTD approach achieves an average TEDS of 87.2% and an average F1-Measure of 88.1% on the test set of HierDoc. The code and dataset will be released at: https://github.com/Pengfei-Hu/MTD. Pengfei Hu 0006, Jianshu Zhang 0001, Jun Du 0002, Jiajia Wu 0003 |
ICPR | 5 |
| 2022 | Scene Text Recognition with Self-supervised Contrastive Predictive CodingabstractSelf-supervised visual pre-training has recently emerged in scene text recognition (STR), which designs the pretext tasks and takes unlabeled data as input to obtain useful representations for STR. However, most current self-supervised methods do not pay special attention to the importance of sequence awareness. Accordingly, we propose a novel self-supervised STR method based on contrastive predictive coding (STR-CPC), which regards a text instance as a sequence from left to right and captures the visual sequence correlation. Considering the information overlap problem within the feature map induced by the deep convolutional neural network (CNN) encoder, we design a widthwise causal convolution during model pre-training and a progressive recovery training strategy (PRTS) during model fine-tuning to improve the STR performance. Experiments on scene text show that our STR-CPC method outperforms the existing self-supervised methods, which testifies the advantage of visual sequence correlation for STR. Additionally, STR-CPC observably boosts performance compared with supervised training when the amount of labeled data decreases. Xinzhe Jiang, Jianshu Zhang 0001, Jun Du 0002, Jiajia Wu 0003 |
ICPR | 5 |
| 2022 | Structural String Decoder for Handwritten Mathematical Expression RecognitionabstractRecently, recognition of handwritten mathematical expression has been greatly improved by employing sequence modeling methods such as encoder-decoder based methods. Existing encoder-decoder models use string decoders or tree decoders to generate markup of mathematical expression recognition. String decoders directly generate LaTeX strings and tree decoders decode expressions into tree structures. The generalization of string decoders is poor on mathematical expressions with complex hierarchical structures, but its language model is better. Tree decoders can deal with the complex hierarchical structures, but its language model is weakened. In order to take advantage of the above two decoders, we propose a novel structural string decoder (SSD) which not only has good generalization but also can make good use of language model. We demonstrate how the proposed SSD outperforms state-of-the-art string decoders and tree decoders through a set of experiments on CROHME database, which is currently the largest benchmark for online handwritten mathematical expression recognition. Jiajia Wu 0003, Jinshui Hu, Mingjun Chen, Li-Rong Dai 0001, Xuejing Niu |
ICPR | 1 |
| 2022 | A multimodal attention fusion network with a dynamic vocabulary for TextVQA
Jiajia Wu 0003, Jun Du 0002, Fengren Wang, Xinzhe Jiang, Jinshui Hu, Jianshu Zhang 0001, Li-Rong Dai 0001 |
Pattern Recognit. | 1 |
| 2022 | Tree-based data augmentation and mutual learning for offline handwritten mathematical expression recognition
Jun Du 0002, Jianshu Zhang 0001, Changjie Wu, Mingjun Chen, Jiajia Wu 0003 |
Pattern Recognit. | 6 |
| 2020 | Stroke Based Posterior Attention for Online Handwritten Mathematical Expression RecognitionabstractRecently, many researches propose to employ attention based encoder-decoder models to convert a sequence of trajectory points into a LaTeX string for online handwritten mathematical expression recognition (OHMER), and the recognition performance of these models critically relies on the accuracy of the attention. In this paper, unlike previous methods which basically employ a soft attention model, we propose to employ a posterior attention model, which modifies the attention probabilities after observing the output probabilities generated by the soft attention model. In order to further improve the posterior attention mechanism, we propose a stroke average pooling layer to aggregate point-level features obtained from the encoder into stroke-level features. We argue that posterior attention is better to be implemented on stroke-level features than point-level features as the output probabilities generated by stroke is more convincing than generated by point, and we prove that through experimental analysis. Validated on the CROHME competition task, we demonstrate that stroke based posterior attention achieves expression recognition rates of 54.26% on CROHME 2014 and 51.75% on CROHME 2016. According to attention visualization analysis, we empirically demonstrate that the posterior attention mechanism can achieve better alignment accuracy than the soft attention mechanism. Changjie Wu, Qing Wang 0008, Jianshu Zhang 0001, Jun Du 0002, Jiajia Wu 0003, Jin-Shui Hu |
ICPR | 6 |