Chen Duan

dblp:173/6744 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 InstructOCR: Instruction Boosting Scene Text Spotting
abstract
In the field of scene text spotting, previous OCR methods primarily relied on image encoders and pre-trained text information, but they often overlooked the advantages of incorporating human language instructions. To address this gap, we propose InstructOCR, an innovative instruction-based scene text spotting model that leverages human language instructions to enhance the understanding of text within images. Our framework employs both text and image encoders during training and inference, along with instructions meticulously designed based on text attributes. This approach enables the model to interpret text more accurately and flexibly. Extensive experiments demonstrate the effectiveness of our model and we achieve state-of-the-art results on widely used benchmarks. Furthermore, the proposed framework can be seamlessly applied to scene text VQA tasks. By leveraging instruction strategies during pre-training, the performance on downstream VQA tasks can be significantly improved, with a 2.6% increase on the TextVQA dataset and a 2.1% increase on the ST-VQA dataset. These experimental results provide insights into the benefits of incorporating human language instructions for OCR-related tasks.
Chen Duan, Qianyi Jiang, Pei Fu, Shengxi Li, Shan Guo, Junfeng Luo
AAAI1
2025 Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
abstract
Multi-modal Large Language Models (MLLMs) have introduced a novel dimension to document understanding, i.e., they endow large language models with visual comprehension capabilities; however, how to design a suitable image-text pre-training task for bridging the visual and language modality in document-level MLLMs remains underexplored. In this study, we introduce a novel visuallanguage alignment method that casts the key issue as a Visual Question Answering with Mask generation (VQA-Mask) task, optimizing two tasks simultaneously: VQA-based text parsing and mask generation. The former allows the model to implicitly align images and text at the semantic level. The latter introduces an additional mask generator (discarded during inference) to explicitly ensure alignment between visual texts within images and their corresponding image regions at a spatially-aware level. Together, they can prevent model hallucinations when parsing visual text and effectively promote spatially-aware feature representation learning. To support the proposed VQAMask task, we construct a comprehensive image-mask generation pipeline and provide a large-scale dataset with 6M data (MTMask6M). Subsequently, we demonstrate that introducing the proposed mask generation task yields competitive document-level understanding performance. Leveraging the proposed VQAMask, we introduce Marten, a trainingefficient MLLM tailored for document-level understanding. Extensive experiments show that our Marten consistently achieves significant improvements among 8B-MLLMs in document-centric tasks. Code and datasets are available at https://github.com/PriNing/Marten.
Tongkun Guan, Pei Fu, Chen Duan, Qianyi Jiang, Zhentao Guo, Shan Guo, Junfeng Luo, Wei Shen 0002, Xiaokang Yang 0001
CVPR4
2025 A Token-Level Text Image Foundation Model for Document Understanding
Tongkun Guan, Pei Fu, Zhengtao Guo, Wei Shen 0002, Tiezhu Yue, Chen Duan, Qianyi Jiang, Junfeng Luo, Xiaokang Yang 0001
ICCV8
2024 ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and Spotting
abstract
In recent years, text-image joint pre-training techniques have shown promising results in various tasks. However, in Optical Character Recognition (OCR) tasks, aligning text instances with their corresponding text regions in images poses a challenge, as it requires effective alignment between text and OCR-Text (referring to the text in images as OCR-Text to distinguish from the text in natural language) rather than a holistic understanding of the overall image content. In this paper, we propose a new pre-training method called OCR-Text Destylization Modeling (ODM) that transfers Diverse styles of text found in images to a uniform style based on the text prompt. With ODM, we achieve better alignment between text and OCR-Text and enable pre-trained models to adapt to the complex and diverse styles of scene text detection and spotting tasks. Additionally, we have designed a new labeling generation method specifically for ODM and combined it with our proposed Text-Controller module to address the challenge of annotation costs in OCR tasks, allowing a larger amount of unlabeled data to participate in pre-training. Extensive experiments on multiple public datasets demonstrate that our method significantly improves performance and outperforms current pre-training methods in scene text detection and spotting tasks. Code is available at ODM.
Chen Duan, Pei Fu, Shan Guo, Qianyi Jiang, Xiaoming Wei
CVPR1
2022 Local Point Matching Network for Stabilized Crowd Counting and Localization
Lin Niu, Xinggang Wang, Chen Duan, Qiongxia Shen, Wenyu Liu 0001
PRCV (1)3
2021 CO-BPG: A Centralized Optimizer for Routing in BGP-Based Data Center Networks
Chen Duan, Wei Peng 0005
AINA (1)1
2021 Comprehensive characterization of alternative splicing in renal cell carcinoma
abstract
Irregular splicing was associated with tumor formation and progression in renal cell carcinoma (RCC) and many other cancers. By using splicing data in the TCGA SpliceSeq database, RCC subtype classification was performed and splicing features and their correlations with clinical course, genetic variants, splicing factors, pathways activation and immune heterogeneity were systemically analyzed. In this research, alternative splicing was found useful for classifying RCC subtypes. Splicing inefficiency with upregulated intron retention and cassette exon was associated with advanced conditions and unfavorable overall survival of patients with RCC. Splicing characteristics like splice site strength, guanine and cytosine content and exon length may be important factors disrupting splicing balance in RCC. Other than cis-acting and trans-acting regulation, alternative splicing also differed in races and tissue types and is also affected by mutation conditions, pathway settings and the response to environmental changes. Severe irregular splicing in tumor not only indicated terrible intra-cellular homeostasis, but also changed the activity of cancer-associated pathways by different splicing effects including isoforms switching and expression regulation. Moreover, irregular splicing and splicing-associated antigens were involved in immune reprograming and formation of immunosuppressive tumor microenvironment. Overall, we have described several clinical and molecular features in RCC splicing subtypes, which may be important for patient management and targeting treatment.
Jingzhen Li, Kui Sun, Libin Yan, Chen Duan, Zhangqun Ye, Mugen Liu
Briefings Bioinform.7
2019 Feature Refine Network for Text-Based CAPTCHA Recognition
Chen Duan, Rong Zhang 0004, Ke Qing
ICIG (2)1
2018 Nine-Switch Detroit Rectifier
abstract
This paper proposed a novel four-level rectifier with only nine power switches, called by nine-switch Detroit rectifier. Compared with the existing four-level rectifiers, the quantity of components and the voltage stresses across components are both reduced. Four output levels are achieved with small component voltage stress in the proposed four-level rectifier, which is suitable for medium voltage (<;10kV) applications and low voltage (<;1kV) applications as well. A carrier-based modulation scheme of the proposed nine-switch Detroit rectifier is also presented in this paper.
Jianfei Chen 0006, Caisheng Wang, Chen Duan, Chenguang Jiang
IECON3
2016 Location-Aware Image Classification
Xinggang Wang, Xin Yang 0008, Wenyu Liu 0001, Chen Duan, Longin Jan Latecki
MMM (1)4