Wenwen Yu

dblp:70/7773 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021
YearPublicationVenuePosition
2026 OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
abstract
Visually-situated text parsing (VsTP) has recently seen notable advancements, driven by the growing demand for automated document understanding and the emergence of large language models capable of processing document-based questions. While various methods have been proposed to tackle the complexities of VsTP, existing solutions often rely on task-specific architectures and objectives for individual tasks. This leads to modal isolation and complex workflows due to the diversified targets and heterogeneous schemas. In this paper, we introduce OmniParser V2, a universal model that unifies VsTP typical tasks, including text spotting, key information extraction, table recognition, and layout analysis, into a unified framework. Central to our approach is the proposed Structured-Points-of-Thought (SPOT) prompting schemas, which improves model performance across diverse scenarios by leveraging a unified encoder-decoder architecture, objective, and input&output representation. SPOT eliminates the need for task-specific architectures and loss functions, significantly simplifying the processing pipeline. Our extensive evaluations across four tasks on eight different datasets show that OmniParser V2 achieves state-of-the-art or competitive results in VsTP. Additionally, we explore the integration of SPOT within a multimodal large language model structure, further enhancing visual text parsing capabilities on four tasks, thereby confirming the generality of SPOT prompting technique.
Wenwen Yu, Zhibo Yang 0003, Jianqiang Wan, Sibo Song, Jun Tang 0008, Wenqing Cheng, Xiang Bai
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 DocThinker: Explainable Multimodal Large Language Models with Rule-Based Reinforcement Learning for Document Understanding
abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in document understanding. However, their reasoning processes remain largely black-box, making it difficult to ensure reliability and trustworthiness, especially in high-stakes domains such as legal, financial, and medical document analysis. Existing methods use fixed Chain-of-Thought (CoT) reasoning with supervised fine-tuning (SFT) but suffer from catastrophic forgetting, poor adaptability, and limited generalization across domain tasks. In this paper, we propose DocThinker, a rule-based Reinforcement Learning (RL) framework for dynamic inference-time reasoning. Instead of relying on static CoT templates, DocThinker autonomously refines reasoning strategies via policy learning, generating explainable intermediate results, including structured reasoning processes, rephrased questions, regions of interest (RoI) supporting the answer, and the final answer. By integrating multi-objective rule-based rewards and KL-constrained optimization, our method mitigates catastrophic forgetting and enhances both adaptability and transparency. Extensive experiments on multiple benchmarks demonstrate that DocThinker significantly improves generalization while producing more explainable and human-understandable reasoning steps. Our findings highlight RL as a powerful alternative for enhancing explainability and adaptability in MLLM-based document understanding. Code will be available at https://github.com/wenwenyu/DocThinker.
Wenwen Yu, Zhibo Yang 0003, Xiang Bai
ICCV1
2025 ClickTrack: Towards real-time interactive single object tracking
Kuiran Wang, Xuehui Yu, Wenwen Yu, Guorong Li, Xiangyuan Lan, Qixiang Ye, Jianbin Jiao, Zhenjun Han
Pattern Recognit.3
2024 OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
abstract
Recently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Generative Large Language Models (LLMs) capable of processing document-based questions. Various methods have been proposed to address the challenging problem of VsTP. However, due to the diversified targets and heterogeneous schemas, previous works usually design task-specific architectures and objectives for individual tasks, which in- advertently leads to modal isolation and complex workflow. In this paper, we propose a unified paradigm for parsing visually-situated text across diverse scenarios. Specifically, we devise a universal model, called OmniParser, which can simultaneously handle three typical visually-situated text parsing tasks: text spotting, key information extraction, and table recognition. In OmniParser, all tasks share the unified encoder-decoder architecture, the unified objective: point- conditioned text generation, and the unified input&output representation: prompt & structured sequences. Extensive experiments demonstrate that the proposed OmniParser achieves state-of-the-art (SOTA) or highly competitive performances on 7 datasets for the three visually-situated text parsing tasks, despite its unified, concise design. The code is available at AdvancedLiterateMachinery.
Jianqiang Wan, Sibo Song, Wenwen Yu, Wenqing Cheng, Fei Huang 0002, Xiang Bai, Cong Yao, Zhibo Yang 0003
CVPR3
2024 Knowledge Mining of Scene Text for Referring Expression Comprehension
Chenyang Gao, Wenwen Yu, Xiang Bai
ICDAR (5)3
2024 P2Seg: Pointly-supervised Segmentation via Mutual Distillation
abstract
Point-level Supervised Instance Segmentation (PSIS) aims to enhance the applicability and scalability of instance segmentation by utilizing low-cost yet instance-informative annotations. Existing PSIS methods usually rely on positional information to distinguish objects, but predicting precise boundaries remains challenging due to the lack of contour annotations. Nevertheless, weakly supervised semantic segmentation methods are proficient in utilizing intra-class feature consistency to capture the boundary contours of the same semantic regions. In this paper, we design a Mutual Distillation Module (MDM) to leverage the complementary strengths of both instance position and semantic information and achieve accurate instance-level object perception. The MDM consists of Semantic to Instance (S2I) and Istance to Semantic (I2S). S2I is guided by the precise boundaries of semantic regions to learn the association between annotated points and instance contours. I2S leverages discriminative relationships between instances to facilitate the differentiation of various objects within the semantic map. Extensive experiments substantiate the efficacy of MDM in fostering the synergy between instance and semantic information, consequently improving the quality of instance-level object representations. Our method achieves 55.7 mAP50 and 17.6 mAP on the PASCAL VOC and MS COCO datasets, significantly outperforming recent PSIS methods and several box-supervised instance segmentation competitors.
Xuehui Yu, Xumeng Han, Wenwen Yu, Zhixun Huang, Jianbin Jiao, Zhenjun Han
ICLR4
2024 Drug-target affinity prediction with extended graph learning-convolutional networks
abstract
BACKGROUND: High-performance computing plays a pivotal role in computer-aided drug design, a field that holds significant promise in pharmaceutical research. The prediction of drug-target affinity (DTA) is a crucial stage in this process, potentially accelerating drug development through rapid and extensive preliminary compound screening, while also minimizing resource utilization and costs. Recently, the incorporation of deep learning into DTA prediction and the enhancement of its accuracy have emerged as key areas of interest in the research community. Drugs and targets can be characterized through various methods, including structure-based, sequence-based, and graph-based representations. Despite the progress in structure and sequence-based techniques, they tend to provide limited feature information. Conversely, graph-based approaches have risen to prominence, attracting considerable attention for their comprehensive data representation capabilities. Recent studies have focused on constructing protein and drug molecular graphs using sequences and SMILES, subsequently deriving representations through graph neural networks. However, these graph-based approaches are limited by the use of a fixed adjacent matrix of protein and drug molecular graphs for graph convolution. This limitation restricts the learning of comprehensive feature representations from intricate compound and protein structures, consequently impeding the full potential of graph-based feature representation in DTA prediction. This, in turn, significantly impacts the models' generalization capabilities in the complex realm of drug discovery. RESULTS: To tackle these challenges, we introduce GLCN-DTA, a model specifically designed for proficiency in DTA tasks. GLCN-DTA innovatively integrates a graph learning module into the existing graph architecture. This module is designed to learn a soft adjacent matrix, which effectively and efficiently refines the contextual structure of protein and drug molecular graphs. This advancement allows for learning richer structural information from protein and drug molecular graphs via graph convolution, specifically tailored for DTA tasks, compared to the conventional fixed adjacent matrix approach. A series of experiments have been conducted to validate the efficacy of the proposed GLCN-DTA method across diverse scenarios. The results demonstrate that GLCN-DTA possesses advantages in terms of robustness and high accuracy. CONCLUSIONS: The proposed GLCN-DTA model enhances DTA prediction performance by introducing a novel framework that synergizes graph learning operations with graph convolution operations, thereby achieving richer representations. GLCN-DTA does not distinguish between different protein classifications, including structurally ordered and intrinsically disordered proteins, focusing instead on improving feature representation. Therefore, its applicability scope may be more effective in scenarios involving structurally ordered proteins, while potentially being limited in contexts with intrinsically disordered proteins.
Haiou Qi, Wenwen Yu
BMC Bioinform.3
2024 OCRBench: on the hidden mystery of OCR in large multimodal models
Mingxin Huang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Xiang Bai
Sci. China Inf. Sci.5
2024 Turning a CLIP Model Into a Scene Text Spotter
abstract
We exploit the potential of the large-scale Contrastive Language-Image Pretraining (CLIP) model to enhance scene text detection and spotting tasks, transforming it into a robust backbone, FastTCM-CR50. This backbone utilizes visual prompt learning and cross-attention in CLIP to extract image and text-based prior knowledge. Using predefined and learnable prompts, FastTCM-CR50 introduces an instance-language matching process to enhance the synergy between image and text embeddings, thereby refining text regions. Our Bimodal Similarity Matching (BSM) module facilitates dynamic language prompt generation, enabling offline computations and improving performance. FastTCM-CR50 offers several advantages: 1) It can enhance existing text detectors and spotters, improving performance by an average of 1.6% and 1.5%, respectively. 2) It outperforms the previous TCM-CR50 backbone, yielding an average improvement of 0.2% and 0.55% in text detection and spotting tasks, along with a 47.1% increase in inference speed. 3) It showcases robust few-shot training capabilities. Utilizing only 10% of the supervised data, FastTCM-CR50 improves performance by an average of 26.5% and 4.7% for text detection and spotting tasks, respectively. 4) It consistently enhances performance on out-of-distribution text detection and spotting datasets, particularly the NightTime-ArT subset from ICDAR2019-ArT and the DOTA dataset for oriented object detection.
Wenwen Yu, Xingkui Zhu, Haoyu Cao 0001, Xing Sun 0001, Xiang Bai
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Turning a CLIP Model into a Scene Text Detector
abstract
The recent large-scale Contrastive Language-Image Pretraining (CLIP) model has shown great potential in various downstream tasks via leveraging the pretrained vision and language knowledge. Scene text, which contains rich textual and visual information, has an inherent connection with a model like CLIP. Recently, pretraining approaches based on vision language models have made effective progresses in the field of text detection. In contrast to these works, this paper proposes a new method, termed TCM, focusing on Turning the CLIP Model directly for text detection without pretraining process. We demonstrate the advantages of the proposed TCM as follows: (1) The underlying principle of our framework can be applied to improve existing scene text detector. (2) It facilitates the few-shot training capability of existing methods, e.g., by using 10% of labeled data, we significantly improve the performance of the baseline method with an average of 22% in terms of the F-measure on 4 benchmarks. (3) By turning the CLIP model into existing scene text detection methods, we further achieve promising domain adaptation ability. The code will be publicly released at https://github.com/wenwenyu/TCM.
Wenwen Yu, Wei Hua 0005, Deqiang Jiang, Bo Ren 0002, Xiang Bai
CVPR1
2023 TextREC: A Dataset for Referring Expression Comprehension with Reading Comprehension
Chenyang Gao, Hao Wang 0207, Wenwen Yu, Xiang Bai
ICDAR (3)5
2023 ICDAR 2023 Competition on Reading the Seal Title
Wenwen Yu, Mingrui Chen 0001, Ning Lu 0003, Yinlong Wen, Dimosthenis Karatzas, Xiang Bai
ICDAR (2)1
2023 ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao 0001, Wei Hua 0005, Bohan Li 0010, Mingrui Chen 0001, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lv, Yuechen Yu, Wanxiang Che, Errui Ding, Cheng-Lin Liu 0001, Jiebo Luo 0001, Shuicheng Yan, Min Zhang 0005, Dimosthenis Karatzas, Xing Sun 0001, Jingdong Wang 0001, Xiang Bai
ICDAR (2)1
2022 CMS2-Net: Semi-Supervised Sleep Staging for Diverse Obstructive Sleep Apnea Severity
abstract
Although the development of computer-aided algorithms for sleep staging is integrated into automatic detection of sleep disorders, most supervised deep learning-based models might suffer from insufficient labeled data. While the adoption of semi-supervised learning (SSL) can mitigate the issue, the SSL models are still limited to the lack of discriminative feature extraction for diverse obstructive sleep apnea (OSA) severity. This model deterioration might be exacerbated during the domain adaptation. Such exploration on the alleviation of domain-shift of SSL model between different OSA conditions has attracted more and more attentions from the clinic. In this work, a co-attention meta sleep staging network (CMS2-net) is proposed to simultaneously deal with two issues: the inter-class disparity problem and the intra-class selection problem. Within CMS2-net, a co-attention module and a triple-classifier are designed to explicitly refine the coarse feature representations by identifying the class boundary inconsistency. Moreover, the mutual information with meta contrastive variance is introduced to supervise the gradient stream from a multi-scale view. The performance of the proposed framework is demonstrated on both public and local datasets. Furthermore, our approach achieves the state-of-the-art SSL results on both datasets.
Chuanhao Zhang, Wenwen Yu, Yamei Li, Hongqiang Sun, Yuan Zhang 0007, Maarten De Vos
IEEE J. Biomed. Health Informatics2
2021 MASTER: Multi-aspect non-local network for scene text recognition
Ning Lu 0003, Wenwen Yu, Xianbiao Qi, Ping Gong 0003, Rong Xiao 0003, Xiang Bai
Pattern Recognit.2
2020 Synthetic-to-Real Unsupervised Domain Adaptation for Scene Text Detection in the Wild
Weijia Wu 0001, Ning Lu 0003, Enze Xie, Wenwen Yu
ACCV (3)5
2020 PICK: Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional Networks
abstract
Computer vision with state-of-the-art deep learning models has achieved huge success in the field of Optical Character Recognition (OCR) including text detection and recognition tasks recently. However, Key Information Extraction (KIE) from documents as the downstream task of OCR, having a large number of use scenarios in real-world, remains a challenge because documents not only have textual features extracting from OCR systems but also have semantic visual features that are not fully exploited and play a critical role in KIE. Too little work has been devoted to efficiently make full use of both textual and visual features of the documents. In this paper, we introduce PICK, a framework that is effective and robust in handling complex documents layout for KIE by combining graph learning with graph convolution operation, yielding a richer semantic representation containing the textual and visual features and global layout without ambiguity. Extensive experiments on realworld datasets have been conducted to show that our method outperforms baselines methods by significant margins. Our code is available at https://github.com/wenwenyu/PICK-pytorch.
Wenwen Yu, Ning Lu 0003, Xianbiao Qi, Ping Gong 0003, Rong Xiao 0003
ICPR1
2015 Discriminative Structured Feature Engineering for Macroscale Brain Connectomes
abstract
Neuroimaging techniques can measure structural and functional brain connectivity with unprecedented detail in vivo. This so-called brain connectome can be represented as high dimensional matrices corresponding to edge weights in graphs. After measuring the matrices of two cohorts (i.e., patients and healthy controls), one is often required to formulate computational network models for effective feature engineering to draw discriminative distinctions between the cohorts, as well as estimate the associated statistical significance. We designed a novel method to reveal the intrinsic features of functional matrices of discriminative power for group comparison. More specifically, by encouraging co-selection of edges connected to the same node, we preserved the discriminative edges to maximum extent. To reduce the false positive rate of the extracted discriminative edges, an optimization procedure was developed to evaluate the significance of these edges and remove trivial ones. We validated the proposed method using both synthetic data and real benchmarks, and compared it to ℓ1 regularized logistic regression, univariate t-test and stability selection. The experimental results clearly showed that the proposed approach outperformed the three competing methods under various settings. In addition to increasing the F-measure of feature selection, our approach captured the endogenous, discriminative connectivity patterns consistent with recent findings in biomedical literature. This data-driven method paves a new avenue of enquiry into the inherent nature of network models for functional brain connectomes.
Jian Pu, Jun Wang 0006, Wenwen Yu, Zhuangming Shen, Kristina Zeljic, Bomin Sun, Zheng Wang 0035
IEEE Trans. Medical Imaging3