VLDB 2026 Research / reviewers in the wild / expert
Tongkun Guan
dblp:290/9351
· DBLP profile ↗
9ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0003-3346-8315ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Marten: Visual Question Answering with Mask Generation for Multi-modal Document UnderstandingabstractMulti-modal Large Language Models (MLLMs) have introduced a novel dimension to document understanding, i.e., they endow large language models with visual comprehension capabilities; however, how to design a suitable image-text pre-training task for bridging the visual and language modality in document-level MLLMs remains underexplored. In this study, we introduce a novel visuallanguage alignment method that casts the key issue as a Visual Question Answering with Mask generation (VQA-Mask) task, optimizing two tasks simultaneously: VQA-based text parsing and mask generation. The former allows the model to implicitly align images and text at the semantic level. The latter introduces an additional mask generator (discarded during inference) to explicitly ensure alignment between visual texts within images and their corresponding image regions at a spatially-aware level. Together, they can prevent model hallucinations when parsing visual text and effectively promote spatially-aware feature representation learning. To support the proposed VQAMask task, we construct a comprehensive image-mask generation pipeline and provide a large-scale dataset with 6M data (MTMask6M). Subsequently, we demonstrate that introducing the proposed mask generation task yields competitive document-level understanding performance. Leveraging the proposed VQAMask, we introduce Marten, a trainingefficient MLLM tailored for document-level understanding. Extensive experiments show that our Marten consistently achieves significant improvements among 8B-MLLMs in document-centric tasks. Code and datasets are available at https://github.com/PriNing/Marten. Tongkun Guan, Pei Fu, Chen Duan, Qianyi Jiang, Zhentao Guo, Shan Guo, Junfeng Luo, Wei Shen 0002, Xiaokang Yang 0001 |
CVPR | 2 |
| 2025 | A Token-Level Text Image Foundation Model for Document Understanding
Tongkun Guan, Pei Fu, Zhengtao Guo, Wei Shen 0002, Tiezhu Yue, Chen Duan, Qianyi Jiang, Junfeng Luo, Xiaokang Yang 0001 |
ICCV | 1 |
| 2025 | CCDPlus: Towards Accurate Character to Character Distillation for Text RecognitionabstractExisting scene text recognition methods leverage large-scale labeled synthetic data (LSD) to reduce reliance on labor-intensive annotation tasks and improve recognition capability in real-world scenarios. However, the emergence of a synth-to-real domain gap still limits their efficiency and robustness. Consequently, harvesting the meaningful intrinsic qualities of unlabeled real data (URD) is of great importance, given the prevalence of text-laden images. Toward the target, recent efforts have focused on pre-training on URD through sequence-to-sequence self-supervised learning, followed by fine-tuning on LSD via supervised learning. Nevertheless, they encounter three important issues: coarse representation learning units, inflexible data augmentation, and an emerging real-to-synth domain drift. To overcome these challenges, we propose CCDPlus, an accurate character-to-character distillation method for scene text recognition with a joint supervised and self-supervised learning framework. Specifically, tailored for text images, CCDPlus delineates the fine-grained character structures on URD as representation units by transferring knowledge learned from LSD online. Without requiring extra bounding box or pixel-level annotations, this process allows CCDPlus to enable character-to-character distillation flexibly with versatile data augmentation, which effectively extracts general real-world character-level feature representations. Meanwhile, the unified framework combines self-supervised learning on URD with supervised learning on LSD, effectively solving the domain inconsistency and enhancing the recognition performance. Extensive experiments demonstrate that CCDPlus outperforms previous state-of-the-art (SOTA) supervised, semi-supervised, and self-supervised methods by an average of 1.8%, 0.6%, and 1.1% on standard datasets, respectively. Additionally, it achieves a 6.1% improvement on the more challenging Union14M-L dataset. Tongkun Guan, Wei Shen 0002, Xiaokang Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | PosFormer: Recognizing Complex Handwritten Mathematical Expression with Position Forest Transformer
Tongkun Guan, Chengyu Lin 0002, Wei Shen 0002, Xiaokang Yang 0001 |
ECCV (22) | 1 |
| 2024 | Bridging Synthetic and Real Worlds for Pre-Training Scene Text Detectors
Tongkun Guan, Wei Shen 0002, Xue Yang 0005, Xiaokang Yang 0001 |
ECCV (44) | 1 |
| 2023 | Self-Supervised Implicit Glyph Attention for Text RecognitionabstractThe attention mechanism has become the de facto module in scene text recognition (STR) methods, due to its capability of extracting character-level representations. These methods can be summarized into implicit attention based and supervised attention based, depended on how the attention is computed, i.e., implicit attention and supervised attention are learned from sequence-level text annotations and or character-level bounding box annotations, respectively. Implicit attention, as it may extract coarse or even incorrect spatial regions as character attention, is prone to suffering from an alignment-drifted issue. Supervised attention can alleviate the above issue, but it is character category-specific, which requires extra laborious character-level bounding box annotations and would be memory-intensive when handling languages with larger character categories. To address the aforementioned issues, we propose a novel attention mechanism for STR, self-supervised implicit glyph attention (SICA). SICA delineates the glyph structures of text images by jointly self-supervised text seg-mentation and implicit attention alignment, which serve as the supervision to improve attention correctness without extra character-level annotations. Experimental results demonstrate that SIGA performs consistently and significantly better than previous attention-based STR methods, in terms of both attention correctness and final recognition performance on publicly available context benchmarks and our contributed contextless benchmarks. Tongkun Guan, Chaochen Gu, Jingzheng Tu, Xue Yang 0005, Yudi Zhao, Wei Shen 0002 |
CVPR | 1 |
| 2023 | Self-supervised Character-to-Character Distillation for Text RecognitionabstractWhen handling complicated text images (e.g., irregular structures, low resolution, heavy occlusion, and uneven illumination), existing supervised text recognition methods are data-hungry. Although these methods employ large-scale synthetic text images to reduce the dependence on annotated real images, the domain gap still limits the recognition performance. Therefore, exploring the robust text feature representations on unlabeled real images by self-supervised learning is a good solution. However, existing self-supervised text recognition methods conduct sequence-to-sequence representation learning by roughly splitting the visual features along the horizontal axis, which limits the flexibility of the augmentations, as large geometric-based augmentations may lead to sequence-to-sequence feature inconsistency. Motivated by this, we propose a novel self-supervised Character-to-Character Distillation method, CCD, which enables versatile augmentations to facilitate general text representation learning. Specifically, we delineate the character structures of unlabeled real images by designing a self-supervised character segmentation module. Following this, CCD easily enriches the diversity of local characters while keeping their pairwise alignment under flexible augmentations, using the transformation matrix between two augmented views from images. Experiments demonstrate that CCD achieves state-of-the-art results, with average performance gains of 1.38% in text recognition, 1.7% in text segmentation, 0.24 dB (PSNR) and 0.0321 (SSIM) in text super-resolution. Code is available at https://github.com/TongkunGuan/CCD. Tongkun Guan, Wei Shen 0002, Xue Yang 0005, Zekun Jiang, Xiaokang Yang 0001 |
ICCV | 1 |
| 2023 | An Oriented Object Detector towards DiatomsabstractAutomatic diatom detection refers to the task of identifying and characterizing diatoms based on artificial intelligence. It will replace traditional time-consuming and laborious manual microscopy method of diatom observation to greatly accelerate the process of diatom research and some diatomrelated studies, such as diatom abundance statistics, using diatom properties for environmental monitoring and paleoenvironmental reconstruction. However, complex background interference and the detection of slender diatoms with different integrity are two major challenges for automatic diatom detection. To solve the mentioned-above issues, we propose an oriented object detector for automatic diatom detection based on RepPoints, called OOD-RepPoints. Specifically, for encouraging the network to adaptively capture the feature of slender diatoms, we design a cascaded feature refinement head (CFRH) which consists of points generation stage and points refinement stage, to progressively optimize the extraction of slender diatom features. Furthermore, to fit the shape of diatoms well, especially for slender diatoms, we propose a tailored label assignment strategy for our CFRH, which contains a short side assigner (SSA) for points generation stage and an adaptive IoU thresholds assigner (AITA) for points refinement stage. Besides, we contribute a so called O-Diatom dataset for automatic diatom detection. The dataset has 1711 images which contains 3949 diatoms and provides finely manual oriented bounding box annotations. Extensive experiments demonstrated our method achieve state of the art performance and can reach mAP 89.9% which is highest on O-Diatom, and shows competitive results on slender categories of publicly available datasets (i.e., DOTA and HRSC2016). Song Gong, Kaijie Wu 0002, Zhiying Xia, Lihua Ran, Chaochen Gu, Changsheng Lu, Tongkun Guan, Yudi Zhao |
IJCNN | 7 |
| 2022 | Industrial Scene Text Detection With Refined Feature-Attentive NetworkabstractDetecting the marking characters of industrial metal parts remains challenging due to low visual contrast, uneven illumination, corroded surfaces, and cluttered background of metal part images. Affected by these factors, bounding boxes generated by most existing methods could not locate low-contrast text areas very well. In this paper, we propose a refined feature-attentive network (RFN) to solve the inaccurate localization problem. Specifically, we first design a parallel feature integration mechanism to construct an adaptive feature representation from multi-resolution features, which enhances the perception of multi-scale texts at each scale-specific level to generate a high-quality attention map. Then, an attentive proposal refinement module is developed by the attention map to rectify the location deviation of candidate boxes. Besides, a re-scoring mechanism is designed to select text boxes with the best rectified location. To promote the research towards industrial scene text detection, we contribute two industrial scene text datasets, including a total of 102156 images and 1948809 text instances with various character structures and metal parts. Extensive experiments on our dataset and four public datasets demonstrate that our proposed method achieves the state-of-the-art performance. Both code and dataset are available at:https://github.com/TongkunGuan/RFN. Tongkun Guan, Chaochen Gu, Changsheng Lu, Jingzheng Tu, Kaijie Wu 0002, Xin-Ping Guan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |