VLDB 2026 Research / reviewers in the wild / expert
Kai Ding 0009
dblp:44/2891-9
· DBLP profile ↗
25ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-9371-0751ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DocAligner: Automating the annotation of photographed documents through real-virtual alignment
Jiaxin Zhang 0003, Peirong Zhang 0001, Huiyi Cheng, Kai Ding 0009 |
Pattern Recognit. | 6 |
| 2025 | MCS-Bench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in Chinese Classical StudiesabstractWith the rapid development of Multimodal Large Language Models (MLLMs), their potential in Chinese Classical Studies (CCS), a field which plays a vital role in preserving and promoting China’s rich cultural heritage, remains largely unexplored due to the absence of specialized benchmarks. To bridge this gap, we propose MCS-Bench, the first-of-its-kind multimodal benchmark specifically designed for CCS across multiple subdomains. MCS-Bench spans seven core subdomains (Ancient Chinese Text, Calligraphy, Painting, Oracle Bone Script, Seal, Cultural Relic, and Illustration), with a total of 45 meticulously designed tasks. Through extensive evaluation of 37 representative MLLMs, we observe that even the top-performing model (InternVL2.5-78B) achieves an average score below 50, indicating substantial room for improvement. Our analysis reveals significant performance variations across different tasks and identifies critical challenges in areas such as Optical Character Recognition (OCR) and cultural context interpretation. MCS-Bench not only establishes a standardized baseline for CCS-focused MLLM research but also provides valuable insights for advancing cultural heritage preservation and innovation in the Artificial General Intelligence (AGI) era. Data and code will be publicly available. Yang Liu 0353, Jiahuan Cao, Hiuyi Cheng, Yongxin Shi, Kai Ding 0009 |
ACL (1) | 5 |
| 2025 | TongGu-VL: Advancing Visual-Language Understanding in Chinese Classical Studies through Parameter Sensitivity-Guided Instruction TuningabstractChinese Classical Studies (CCS) is a pivotal gateway to ancient Chinese culture. Spanning ancient texts, illustrations, paintings, and calligraphy, CCS presents significant challenges for non-specialists due to its language and visual complexity. While Large Language Models (LLMs) have been explored to facilitate CCS, current methods primarily focus on textual analysis, overlooking the rich visual information intrinsic to classical materials. To bridge this gap, we propose TongGu-VL, a pioneering specialized MLLM designed for CCS applications. Our contributions are threefold. First, we construct CCS358K, a comprehensive multimodal instruction dataset to enhance MLLMs' CCS capabilities. Second, we propose Parameter Sensitivity-Guided Instruction Tuning (PSG-IT), a novel method that mitigates catastrophic forgetting without data replay. It effectively preserves TongGu-VL's general skills, while optimizing its CCS performance. Third, we design a Visual-Text Early Fusion (VTEF) module, which harnesses MLLMs' modality alignment to generate instruction-aware visual representations, thus improving language modeling. Extensive experimental results demonstrate that our model outperforms existing MLLMs on a broad range of CCS tasks, while maintaining general capabilities that benefit other domains beyond CCS. Our model and dataset will be publicly available. Jiahuan Cao, Yang Liu 0353, Peirong Zhang 0001, Yongxin Shi, Kai Ding 0009 |
ACM Multimedia | 5 |
| 2025 | Capturing More: Learning Multi-Domain Representations for Robust Online Handwriting Verification
Peirong Zhang 0001, Kai Ding 0009 |
ACM Multimedia | 2 |
| 2024 | UPOCR: Towards Unified Pixel-Level OCR InterfaceabstractExisting optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies, which significantly increases the complexity of research and maintenance and hinders the fast deployment in applications. To this end, we propose UPOCR, a simple-yet-effective generalist model for Unified Pixel-level OCR interface. Specifically, the UPOCR unifies the paradigm of diverse OCR tasks as image-to-image transformation and the architecture as a vision Transformer (ViT)-based encoder-decoder with learnable task prompts. The prompts push the general feature representations extracted by the encoder towards task-specific spaces, endowing the decoder with task awareness. Moreover, the model training is uniformly aimed at minimizing the discrepancy between the predicted and ground-truth images regardless of the inhomogeneity among tasks. Experiments are conducted on three pixel-level OCR tasks including text removal, text segmentation, and tampered text detection. Without bells and whistles, the experimental results showcase that the proposed method can simultaneously achieve state-of-the-art performance on three tasks with a unified single model, which provides valuable strategies and insights for future research on generalist OCR models. Code is available at https://github.com/shannanyinxiang/UPOCR. Dezhi Peng, Zhenhua Yang, Jiaxin Zhang 0003, Chongyu Liu, Yongxin Shi, Kai Ding 0009, Fengjun Guo |
ICML | 6 |
| 2024 | WenMind: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Classical Literature and Language ArtsabstractLarge Language Models (LLMs) have made significant advancements across numerous domains, but their capabilities in Chinese Classical Literature and Language Arts (CCLLA) remain largely unexplored due to the limited scope and tasks of existing benchmarks. To fill this gap, we propose WenMind, a comprehensive benchmark dedicated for evaluating LLMs in CCLLA. WenMind covers the sub-domains of Ancient Prose, Ancient Poetry, and Ancient Literary Culture, comprising 4,875 question-answer pairs, spanning 42 fine-grained tasks, 3 question formats, and 2 evaluation scenarios: domain-oriented and capability-oriented. Based on WenMind, we conduct a thorough evaluation of 31 representative LLMs, including general-purpose models and ancient Chinese LLMs. The results reveal that even the best-performing model, ERNIE-4.0, only achieves a total score of 64.3, indicating significant room for improvement of LLMs in the CCLLA domain. We also provide insights into the strengths and weaknesses of different LLMs and highlight the importance of pre-training data in achieving better results.Overall, WenMind serves as a standardized and comprehensive baseline, providing valuable insights for future CCLLA research. Our benchmark and related code are available at \url{https://github.com/SCUT-DLVCLab/WenMind}. Jiahuan Cao, Yang Liu 0353, Yongxin Shi, Kai Ding 0009 |
NeurIPS | 4 |
| 2024 | A tree-based model with branch parallel decoding for handwritten mathematical expression recognition
Zhe Li 0046, Wentao Yang 0003, Hengnian Qi, Yichao Huang, Kai Ding 0009 |
Pattern Recognit. | 6 |
| 2024 | Improving Handwritten Mathematical Expression Recognition via Similar Symbol DistinguishingabstractHandwritten mathematical expression recognition (HMER) is an essential task in the OCR community, which consists of two sub-tasks, i.e., symbol recognition and structure parsing. Modern literature treats HMER as a LaTeX sequence predicting problem that simultaneously recognizes symbols and parses the structures of MEs. Although deep learning-based HMER methods have been achieving promising results on public benchmarks, it is admitted that the misclassification error between visually similar symbols still prevents these approaches from more generalized scenes. In this paper, we try to solve this issue from three aspects. 1) We enhanced the feature extraction progress by introducing path signature features, which incorporates local writing details and global spatial information. 2) We developed a language model that uses contextual information to correct the symbols misclassified by vision-only-based recognition models. 3) We solved the misalignment problem in existing ensemble method by designing a dynamic time warping (DTW) based algorithm. By combining the above improvements, our method achieved state-of-the-art results on three CROHME benchmarks, outperforming previous methods by a large margin. Zhe Li 0046, Xinyu Wang 0010, Yichao Huang, Kai Ding 0009 |
IEEE Trans. Multim. | 6 |
| 2023 | M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout AnalysisabstractDocument layout analysis is a crucial prerequisite for document understanding, including document retrieval and conversion. Most public datasets currently contain only PDF documents and lack realistic documents. Models trained on these datasets may not generalize well to real-world scenarios. Therefore, this paper introduces a large and diverse document layout analysis dataset called M6Doc. The M6 designation represents six properties: (1) Multi-Format (including scanned, photographed, and PDF documents); (2) Multi-Type (such as scientific articles, textbooks, books, test papers, magazines, newspapers, and notes); (3) Multi-Layout (rectangular, Manhattan, non-Manhattan, and multi-column Manhattan); (4) Multi-Language (Chinese and English); (5) Multi-Annotation Category (74 types of annotation labels with 237,116 annotation instances in 9,080 manually annotated pages); and (6) Modern documents. Additionally, we propose a transformer-based document layout analysis method called TransDLANet, which leverages an adaptive element matching mechanism that enables query embedding to better match ground truth to improve recall, and constructs a segmentation branch for more precise document image instance segmentation. We conduct a comprehensive evaluation of M6Doc with various layout analysis methods and demonstrate its effectiveness. TransDLANet achieves state-of-the-art performance on M6 Doc with 64.5% mAP. The M6Doc dataset will be available at https://github.com/HcIILAB/M6Doc. Hiuyi Cheng, Peirong Zhang 0001, Sihang Wu, Jiaxin Zhang 0003, Qiyuan Zhu, Zecheng Xie, Jing Li 0036, Kai Ding 0009 |
CVPR | 8 |
| 2022 | LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document UnderstandingabstractStructured document understanding has attracted considerable attention and made significant progress recently, owing to its crucial role in intelligent document processing.However, most existing related models can only deal with the document data of specific language(s) (typically English) included in the pre-training collection, which is extremely limited.To address this issue, we propose a simple yet effective Language-independent Layout Transformer (LiLT) for structured document understanding.LiLT can be pre-trained on the structured documents of a single language and then directly fine-tuned on other languages with the corresponding off-the-shelf monolingual/multilingual pre-trained textual models.Experimental results on eight languages have shown that LiLT can achieve competitive or even superior performance on diverse widely-used downstream benchmarks, which enables language-independent benefit from the pre-training of document layout structure.Code and model are publicly available at https://github.com/jpWang/LiLT. Jiapeng Wang 0003, Kai Ding 0009 |
ACL (1) | 3 |
| 2022 | SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionabstractEnd-to-end scene text spotting has attracted great attention in recent years due to the success of excavating the intrinsic synergy of the scene text detection and recognition. However, recent state-of-the-art methods usually incorporate detection and recognition simply by sharing the backbone, which does not directly take advantage of the feature interaction between the two tasks. In this paper, we propose a new end-to-end scene text spotting framework termed SwinTextSpotter. Using a transformer encoder with dynamic head as the detector, we unify the two tasks with a novel Recognition Conversion mechanism to explicitly guide text localization through recognition loss. The straightforward design results in a concise framework that requires neither additional rectification module nor character-level annotation for the arbitrarily-shaped text. Qualitative and quantitative experiments on multi-oriented datasets RoIC13 and ICDAR 2015, arbitrarily-shaped datasets Total-Text and CTW1500, and multi-lingual datasets ReCTS (Chinese) and VinText (Viet-namese) demonstrate SwinTextSpotter significantly outperforms existing methods. Code is available at https://github.com/mxin262/SwinTextSpotter. Mingxin Huang, Zhenghao Peng, Chongyu Liu, Dahua Lin, Shenggao Zhu, Nicholas Jing Yuan, Kai Ding 0009 |
CVPR | 8 |
| 2022 | Don't Forget Me: Accurate Background Recovery for Text Removal via Modeling Local-Global Context
Chongyu Liu, Canjie Luo, Bangdong Chen, Fengjun Guo, Kai Ding 0009 |
ECCV (28) | 7 |
| 2022 | Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the WildabstractCamera-captured document images usually suffer from perspective and geometric deformations. It is of great value to rectify them when considering poor visual aesthetics and the deteriorated performance of OCR systems. Recent learning-based methods intensively focus on the accurately cropped document image. However, this might not be sufficient for overcoming practical challenges, including document images either with large marginal regions or without margins. Due to this impracticality, users struggle to crop documents precisely when they encounter large marginal regions. Simultaneously, dewarping images without margins is still an insurmountable problem. To the best of our knowledge, there is still no complete and effective pipeline for rectifying document images in the wild. To address this issue, we propose a novel approach called Marior (Margin Removal and Iterative Content Rectification). Marior follows a progressive strategy to iteratively improve the dewarping quality and readability in a coarse-to-fine manner. Specifically, we divide the pipeline into two modules: margin removal module (MRM) and iterative content rectification module (ICRM). First, we predict the segmentation mask of the input image to remove the margin, thereby obtaining a preliminary result. Then we refine the image further by producing dense displacement flows to achieve content-aware rectification. We determine the number of refinement iterations adaptively. Experiments demonstrate the state-of-the-art performance of our method on public benchmarks. The resources are available at https://github.com/ZZZHANG-jx/Marior for further comparison. Jiaxin Zhang 0003, Canjie Luo, Fengjun Guo, Kai Ding 0009 |
ACM Multimedia | 5 |
| 2021 | Towards Fast, Accurate and Compact Online Handwritten Chinese Text Recognition
Dezhi Peng, Canyu Xie, Zecheng Xie, Kai Ding 0009, Yichao Huang, Yaqiang Wu |
ICDAR (3) | 6 |
| 2021 | Improving Machine Understanding of Human Intent in Charts
Sihang Wu, Canyu Xie, Guozhi Tang, Qianying Liao, Jiapeng Wang 0003, Bangdong Chen, Xinfeng Chang, Kai Ding 0009, Yichao Huang |
ICDAR (3) | 11 |
| 2021 | DeMatch: Towards Understanding the Panel of Chart Documents
Hesuo Zhang, Weihong Ma, Yichao Huang, Kai Ding 0009, Yaqiang Wu |
ICDAR (3) | 5 |
| 2021 | Tag, Copy or Predict: A Unified Weakly-Supervised Learning Framework for Visual Information Extraction using SequencesabstractVisual information extraction (VIE) has attracted increasing attention in recent years. The existing methods usually first organized optical character recognition (OCR) results in plain texts and then utilized token-level category annotations as supervision to train a sequence tagging model. However, it expends great annotation costs and may be exposed to label confusion, the OCR errors will also significantly affect the final performance. In this paper, we propose a unified weakly-supervised learning framework called TCPNet (Tag, Copy or Predict Network), which introduces 1) an efficient encoder to simultaneously model the semantic and layout information in 2D OCR results, 2) a weakly-supervised training method that utilizes only sequence-level supervision; and 3) a flexible and switchable decoder which contains two inference modes: one (Copy or Predict Mode) is to output key information sequences of different categories by copying a token from the input or predicting one in each time step, and the other (Tag Mode) is to directly tag the input sequence in a single forward pass. Our method shows new state-of-the-art performance on several public benchmarks, which fully proves its effectiveness. Jiapeng Wang 0003, Guozhi Tang, Weihong Ma, Kai Ding 0009, Yichao Huang |
IJCAI | 6 |
| 2011 | SCUT-COUCH2009 - a comprehensive online unconstrained Chinese handwriting database and benchmark evaluation
Yan Gao 0011, Yunyang Li, Kai Ding 0009 |
Int. J. Document Anal. Recognit. | 5 |
| 2010 | Incremental MQDF Learning for Writer Adaptive Handwriting RecognitionabstractWriter adaptation has been proved to be an effective approach to improve the recognition performance of the writer-independent recognizer for a particular writer. In this paper, we propose a writer adaptive handwriting recognition approach by incremental learning the Modified Quadratic Discriminant Function (MQDF) classifier. We derived the solution of Incremental MQDF (IMQDF) and then present a Discriminative IMQDF (DIMQDF) by deriving the solution of IMQDF in the updated discriminative feature space. Based on IMQDF or DIMQDF, the writer adaptation is finally performed by updating the MQDF recognizer adaptively. The experimental results for recognizing handwriting Chinese characters indicate that the proposed IMQDF and DIQMDF approaches can reduce as much as 52.71% and 45.38% error rate respectively on the writer-dependent dataset while only have less than 0.18% accuracy loss on the writer-independent dataset. In other words, the proposed IMQDF and DIMQDF based writer adaptation approaches can significantly increase the recognition accuracy on writer-dependent dataset while only have limited negative influence for general writer. Kai Ding 0009 |
ICFHR | 1 |
| 2010 | A New Approach for Synthesis and Recognition of Large Scale Handwritten Chinese WordsabstractLacking of dataset is still a serious problem for researchers who study on online handwriting word recognition (HWR). In this paper, a handwritten Chinese word synthesis method is proposed for the first time to generate a large scale handwritten Chinese word dataset. The distributions of shape and position characteristics, such as aspect radio, character interval and the angle of gravity center line in each word sample of the Word8888 dataset have been estimated respectively. Based on this, we synthesize as large as 44,208 categories of 8,311,104 unconstrained handwritten Chinese word samples. To verify the validity of the synthesized dataset, a practical rotation free handwriting Chinese word recognition system is presented based on a new holistic approach. Experimental results for randomly rotated word samples demonstrate that the holistic approach can achieve 91.96% recognition accuracy, which provides evidence for the effectiveness of our method. Kai Ding 0009, Hanyu Yan |
ICFHR | 3 |
| 2010 | Incremental learning of LDA model for Chinese writer adaptation
Kai Ding 0009, Zhibin Huang |
Neurocomputing | 2 |
| 2009 | An Investigation of Imaginary Stroke Techinique for Cursive Online Handwriting Chinese Character RecognitionabstractImaginary stroke technique has been proved to be an effective solution to the problem of the stroke connection in online handwritten character recognition. However, it may cause confusions among characters with similar but actually different trajectories after adding imaginary strokes. In this paper, we first investigate both the benefit and the defect of the imaginary stroke technique, and then two modified methods are proposed under the framework of feature fusion and local feature enhance respectively. With the proposed methods, the feature of imaginary strokes is employed to unify the writing styles and the feature of real strokes is enhanced to strengthen discriminability. Experimental results for handwritten Chinese character recognition indicate that comparing with feature without imaginary strokes and feature with imaginary strokes, our proposed methods provide about 3%~8% and 1%~4% recognition accuracy improvement respectively. Kai Ding 0009, Guoqiang Deng |
ICDAR | 1 |
| 2009 | A New Method for Rotation Free Method for Online Unconstrained Handwritten Chinese Word Recognition: A Holistic ApproachabstractMost online handwriting word recognition (HWR) approaches proceed by segmenting words into isolate characters which are recognized separately. Inspired by results in cognitive psychology, holistic word recognition approaches provides another effective way to deal the problem of HWR. In this paper, we propose a new method for rotation free online unconstrained Chinese word recognition through a holistic approach. By a gravity center balancing skew detection and correction method, the rotation ranging from 0deg to 360deg of a Chinese handwritten word can be detected. Through the process of preprocessing, feature extraction using elastic meshing technique and classification, the handwritten words with characters even connected or partially overlapped can be recognized through a holistic approach. Experiments were performed on 8888 categories of 1,137,664 unconstrained handwritten Chinese word samples. Experimental results for randomly rotated unconstrained cursive handwritten Chinese word data demonstrated that the proposed method can achieve about 96.58% recognition accuracy. Kai Ding 0009, Xue Gao |
ICDAR | 1 |
| 2009 | Writer Adaptive Online Handwriting Recognition Using Incremental Linear Discriminant AnalysisabstractWriter adaptive handwriting recognition, which has potential of increasing accuracies for a particular user, is the process of converting a writer-independent recognition system to a writer-dependent one. In this paper, we provide a general incremental learning solution for linear discriminant analysis (LDA) on the basis of previous researches, and propose an Incremental LDA (ILDA) based writer adaptive online handwriting recognition method. The adaptation is performed by modifying both the prototypes and the LDA transformation matrix through ILDA algorithm. It includes: (1) modifying prototypes in original feature space; (2) updating the LDA transformation matrix; (3) projecting the updated prototypes to LDA feature space. Experiments are performed on two datasets, the writer-dependent dataset, in which the writing style is consistent with the incremental training data, and the writer-independent dataset. The results demonstrated that our proposed method can reduce as much as 46.35% error rate on the writer-dependent dataset with only 0.20% accuracy loss on the writer-independent dataset. It indicates that our proposed method can significantly increase the recognition accuracy for a particular writer while has minor effects for general writers. Zhibin Huang, Kai Ding 0009, Xue Gao |
ICDAR | 2 |
| 2009 | Character-SIFT: A Novel Feature for Offline Handwritten Chinese Character RecognitionabstractSIFT descriptor has been widely applied in computer vision and object recognition, but has not been explored in the field of handwritten Chinese character recognition. In this paper we proposed a novel SIFT based feature for offline handwritten Chinese character recognition. The presented feature is a modification of SIFT descriptor taking into account of the characteristics of handwritten Chinese samples. In our approach, global elastic meshing is first constructed and then the related gradient code of each sub-region is accumulated dynamically. Experiments using MQDF classifier show our featurepsilas effectiveness with a recognition rate of 97.868%, which outperforms original SIFT feature and two traditional features, Gabor feature and gradient feature. Kai Ding 0009, Xue Gao |
ICDAR | 3 |