Yaqiang Wu

dblp:242/7854 · DBLP profile ↗
← Back
40ranked-venue papers
4as first author
37since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 3 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 16 since 2021Databases, data management, data science and information retrieval · 9 · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing Retrieval-Augmented Large Vision Language Models via Knowledge Conflict Mitigation
abstract
Multimodal Retrieval-Augmented Generation (MRAG) has recently been explored to empower Large Vision Language Models (LVLMs) with more comprehensive and up-to-date contextual knowledge, aiming to compensate for their limited and coarse-grained parametric knowledge in knowledge-intensive tasks. However, the retrieved contextual knowledge is usually not aligned with LVLMs’ internal parametric knowledge, leading to knowledge conflicts and further unreliable responses. To tackle this issue, we design KCM, a training-free and plug-and-play framework that can effectively mitigate knowledge conflicts while incorporating MRAG for more accurate LVLM responses. KCM enhances contextual knowledge utilization by modifying the LVLM architecture from three key perspectives. First, KCM adaptively adjusts attention distributions among multiple attention heads, encouraging LVLMs to focus on contextual knowledge with reduced distraction. Second, KCM identifies and prunes knowledge-centric LVLM neurons that encode coarse-grained parametric knowledge, thereby suppressing interferences and enabling more effective integration of contextual knowledge. Third, KCM amplifies the information flow from the input context by injecting supplementary context logits, reinforcing its contribution to the final output. Extensive experiments over multiple LVLMs and benchmarks show that KCM outperforms the state-of-the-art consistently by large margins, incurring neither extra training nor external tools.
Wenbin An, Jiahao Nie 0002, Feng Tian 0002, Mingxiang Cai, Yaqiang Wu, Shijian Lu
AAAI5
2026 ST-SAM: Multimodal Scene Text Segmentation with Dense Visual and Sparse Textual Prompts via SAM
abstract
Scene text segmentation is a critical preprocessing step in various text-based applications. Specialist text segmentation methods, often relying on a detect-then-segment paradigm, tend to exhibit reduced robustness and can lead to cascading errors. The introduction of the Segment Anything Model (SAM) has revolutionized general segmentation by leveraging vision foundation models. However, SAM still falls short when applied to domain-specific tasks such as scene text segmentation. To bridge this gap between SAM and specialized scene text segmentation approaches, we propose ST-SAM (Scene Text SAM), a parameter-efficient fine-tuning framework tailored to adapt SAM for high-quality scene text segmentation without relying on explicit text detection. ST-SAM incorporates a multimodal prompting mechanism: a lightweight visual encoder generates multi-scale spatial features to provide precise visual context; and textual prompts generated by a large language model offer high-level semantic guidance. We demonstrate the advantages of the proposed ST-SAM as follows: (1) ST-SAM achieves new state-of-the-art performance on multiple scene text segmentation benchmarks, including 85.30% fgIoU on Total-Text and 91.03% fgIoU on TextSeg, outperforming both specialist and generalist models. (2) ST-SAM enables effective domain adaptation by flexibly adapting the general SAM architecture to the domain of scene text. (3) By discarding the detect-then-segment pipeline, ST-SAM simplifies the inference process while still achieving robust performance on complex text cases.
Yaqiang Wu, Jiayi Yan, Yu Zhou 0015, Lingling Zhang 0005, Qianying Wang 0002
AAAI2
2026 Encode Geometric Diagram as Geo-Graph in Geometry Problem Solving
abstract
Geometry Problem Solving has become a hot topic these years due to its complexity of enabling the machine with geometric abstraction, multi-modal reasoning and mathematical capabilities. Majority of research works place their attention on the fusion of multi-modal data or the synergistic combination of neural and symbolic systems for performance improvement. However, their neglect of the unique characteristics of geometric diagrams, which distinguish them from natural images, impedes the further exploring of critical information in geometric diagrams. In this work, we introduce the novel concept of geo-graph and propose the Geo-Graph Geometry Problem Solving model which encodes the geometric diagram from a new perspective. The geo-graph is designed to include semantic, structural and spatial information in the diagram, which is crucial to subsequent problem reasoning stage. To facilitate the model's comprehension of the actual layout of geometric diagram, spatial and connecting attentions are devised to serve as intrinsic knowledge guidance for feature propagation. An extra cross-modal attention is used as external guidance to instruct the encoding of geo-graph to be related to specific problem target. Fused multi-modal features are then sent into a commonly used encoder-decoder framework for final solution generation. The model is first trained with three carefully designed pre-training tasks to establish its fundamental knowledge of geo-graph, leveraging numerous varied samples generated through a geo-graph-based augmentation method. Experiments on popular geometry problem solving datasets demonstrate the effectiveness and superiority of our model for geometric diagram encoding.
Lingling Zhang 0005, Xinyu Zhang 0021, Yaqiang Wu
AAAI6
2026 Programming knowledge tracing based on knowledge concept identification and hierarchical modeling
Junjiao Xiang, Yan Chen 0031, Qin Xia, Feng Tian 0002, Yaqiang Wu, Sibo Cai, Ping Chen 0001
Neurocomputing7
2026 Graph Mixture of Experts with Differential Cross-Attention Alignment for Multimodal Intent Recognition
Shilin Sun 0001, Wenbin An, Qidong Liu 0002, Jiahao Nie 0002, Zhi Zeng 0001, Xian-Sheng Hua, Yaqiang Wu, Feng Tian 0002
Knowl. Based Syst.8
2026 LFSRM: Few-Shot Diagram-Sentence Matching via Local-Feedback Self-Regulating Memory
abstract
Image-sentence matching that aims to understand the correspondence between vision and language, has achieved significant progress with various deep methods trained under large-scale supervision. Different from natural images taken by camera, diagrams in the textbooks contain more graphic objects, drawings, and natural objects, and the diagram-sentence matching plays an important role in textbook understanding and question answering. However, existing matching models are not suitable for the challenging task between diagrams and sentences, due to the more serious few-shot content and incomplete description problems. In this paper, we propose a novel local-feedback self-regulating memory framework (LFSRM) for diagram-sentence matching. On one hand, LFSRM includes an external memory to store the useful multi-modal information, especially uncommon ones, to overcome the few-shot content problem, where the memory is updated flexibly according to the local-feedback from visual-textual alignment scores. On the other hand, LFSRM designs an attention mechanism on local-level alignment scores and a strengthening factor impacted on sentence-to-diagram matching direction for alleviating the incomplete description problem. Extensive experiments on three datasets show that LFSRM achieves satisfactory results on conventional image-sentence matching, and outperforms SOTA methods on few-shot image/diagram-sentence matching by a large margin.
Lingling Zhang 0005, Jun Liu 0002, Xiaojun Chang, Yaqiang Wu
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 Unleashing the Potential of Model Bias for Generalized Category Discovery
abstract
Generalized Category Discovery is a significant and complex task that aims to identify both known and undefined novel categories from a set of unlabeled data, leveraging another labeled dataset containing only known categories. The primary challenges stem from model bias induced by pre-training on only known categories and the lack of precise supervision for novel ones, leading to category bias towards known categories and category confusion among different novel categories, which hinders models' ability to identify novel categories effectively. To address these challenges, we propose a novel framework named Self-Debiasing Calibration (SDC). Unlike prior methods that regard model bias towards known categories as an obstacle to novel category identification, SDC provides a novel insight into unleashing the potential of the bias to facilitate novel category learning. Specifically, we utilize the biased pre-trained model to guide the subsequent learning process on unlabeled data. The output of the biased model serves two key purposes. First, it provides an accurate modeling of category bias, which can be utilized to measure the degree of bias and debias the output of the current training model. Second, it offers valuable insights for distinguishing different novel categories by transferring knowledge between similar categories. Based on these insights, SDC dynamically adjusts the output logits of the current training model using the output of the biased model. This approach produces less biased logits to effectively address the issue of category bias towards known categories, and generates more accurate pseudo labels for unlabeled data, thereby mitigating category confusion for novel categories. Experiments on three benchmark datasets show that SDC outperforms SOTA methods, especially in the identification of novel categories.
Wenbin An, Haonan Lin, Jiahao Nie 0002, Feng Tian 0002, Wenkai Shi, Yaqiang Wu, Qianying Wang 0002, Ping Chen 0001
AAAI6
2025 Using Depth-Enhanced Spatial Transformation for Student Gaze Target Estimation in Dual-View Classroom Images
abstract
Dual-view gaze target estimation in classroom environments has not been thoroughly explored. Existing methods lack consideration of depth information, primarily focusing on 2D image information and neglecting the latent 3D spatial context, which could lead to suboptimal transformation and cause the gaze cone to intersect with an incorrect object. This paper introduces a novel dual-view gaze target estimation method tailored for classroom settings, leveraging depth-enhanced spatial transformations. By formulating a depth-enhanced 2D space, our method uses depth-enhanced spatial transformation to accurately project students’ gaze cones to the teacher-oriented image. Additionally, we collected a dataset named DVSGE, specifically for student gaze target estimation in dual-view classroom images. Experimental results demonstrate significant performance improvements of 9.8% in AUC and 19.9% in L2-Distance for our method, surpassing existing methods.
Haonan Miao, Peizheng Zhao, Yaqiang Wu, Feng Tian 0002
ICASSP6
2025 PACM: Position-Aware Cross-Modality Decoder for Handwritten Mathematical Expression Recognition
Zhijie Shen, Can Ma, Yaqiang Wu, Yu Zhou 0015
ICDAR (1)5
2025 PerturbCTC: Improving Alignment in Scene Text Recognition with Feature Perturbation Based CTC
Zhijie Shen, Yaqiang Wu, Gangyan Zeng, Dongbao Yang, Yu Zhou 0015
ICDAR (4)4
2025 Boosting Knowledge Utilization in Multimodal Large Language Models via Adaptive Logits Fusion and Attention Reallocation
abstract
Despite their recent progress, Multimodal Large Language Models (MLLMs) often struggle in knowledge-intensive tasks due to the limited and outdated parametric knowledge acquired during training. Multimodal Retrieval Augmented Generation addresses this issue by retrieving contextual knowledge from external databases, thereby enhancing MLLMs with expanded knowledge sources. However, existing MLLMs often fail to fully leverage the retrieved contextual knowledge for response generation. We examine representative MLLMs and identify two major causes, namely, attention bias toward different tokens and knowledge conflicts between parametric and contextual knowledge. To this end, we design Adaptive Logits Fusion and Attention Reallocation (ALFAR), a training-free and plug-and-play approach that improves MLLM responses by maximizing the utility of the retrieved knowledge. Specifically, ALFAR tackles the challenges from two perspectives. First, it alleviates attention bias by adaptively shifting attention from visual tokens to relevant context tokens according to query-context relevance. Second, it decouples and weights parametric and contextual knowledge at output logits, mitigating conflicts between the two types of knowledge. As a plug-and-play method, ALFAR achieves superior performance across diverse datasets without requiring additional training or external tools. Extensive experiments over multiple MLLMs and benchmarks show that ALFAR consistently outperforms the state-of-the-art by large margins. Our code and data are available at https://github.com/Lackel/ALFAR.
Wenbin An, Jiahao Nie 0002, Feng Tian 0002, Haonan Lin, Mingxiang Cai, Yaqiang Wu, Qianying Wang 0002, Shijian Lu
NeurIPS6
2025 ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
abstract
Charts are high-density visualization carriers for complex data, serving as a crucial medium for information extraction and analysis. Automated chart understanding poses significant challenges to existing multimodal large language models (MLLMs) due to the need for precise and complex visual reasoning. Current step-by-step reasoning models primarily focus on text-based logical reasoning for chart understanding. However, they struggle to refine or correct their reasoning when errors stem from flawed visual understanding, as they lack the ability to leverage multimodal interaction for deeper comprehension. Inspired by human cognitive behavior, we propose ChartSketcher, a multimodal feedback-driven step-by-step reasoning method designed to address these limitations. ChartSketcher is a chart understanding model that employs Sketch-CoT, enabling MLLMs to annotate intermediate reasoning steps directly onto charts using a programmatic sketching library, iteratively feeding these visual annotations back into the reasoning process. This mechanism enables the model to visually ground its reasoning and refine its understanding over multiple steps. We employ a two-stage training strategy: a cold start phase to learn sketch-based reasoning patterns, followed by off-policy reinforcement learning to enhance reflection and generalization. Experiments demonstrate that ChartSketcher achieves promising performance on chart understanding benchmarks and general vision tasks, providing an interactive and interpretable approach to chart comprehension.
Muye Huang, Lingling Zhang 0005, Jie Ma 0001, Han Lai, Fangzhi Xu, Yifei Li 0006, Yaqiang Wu, Jun Liu 0002
NeurIPS8
2025 Student gaze target estimation based on depth transformation on dual-view classroom images
Haonan Miao, Peizheng Zhao, Morteza Seberi, Yaqiang Wu, Feng Tian 0002
Comput. Vis. Image Underst.7
2024 Transfer and Alignment Network for Generalized Category Discovery
abstract
Generalized Category Discovery (GCD) is a crucial real-world task that aims to recognize both known and novel categories from an unlabeled dataset by leveraging another labeled dataset with only known categories. Despite the improved performance on known categories, current methods perform poorly on novel categories. We attribute the poor performance to two reasons: biased knowledge transfer between labeled and unlabeled data and noisy representation learning on the unlabeled data. The former leads to unreliable estimation of learning targets for novel categories and the latter hinders models from learning discriminative features. To mitigate these two issues, we propose a Transfer and Alignment Network (TAN), which incorporates two knowledge transfer mechanisms to calibrate the biased knowledge and two feature alignment mechanisms to learn discriminative features. Specifically, we model different categories with prototypes and transfer the prototypes in labeled data to correct model bias towards known categories. On the one hand, we pull instances with known categories in unlabeled data closer to these prototypes to form more compact clusters and avoid boundary overlap between known and novel categories. On the other hand, we use these prototypes to calibrate noisy prototypes estimated from unlabeled data based on category similarities, which allows for more accurate estimation of prototypes for novel categories that can be used as reliable learning targets later. After knowledge transfer, we further propose two feature alignment mechanisms to acquire both instance- and category-level knowledge from unlabeled data by aligning instance features with both augmented features and the calibrated prototypes, which can boost model performance on both known and novel categories with less noise. Experiments on three benchmark datasets show that our model outperforms SOTA methods, especially on novel categories. Theoretical analysis is provided for an in-depth understanding of our model in general. Our code and data are available at https://github.com/Lackel/TAN.
Wenbin An, Feng Tian 0002, Wenkai Shi, Yan Chen 0031, Yaqiang Wu, Qianying Wang 0002, Ping Chen 0001
AAAI5
2024 A Unified Knowledge Transfer Network for Generalized Category Discovery
abstract
Generalized Category Discovery (GCD) aims to recognize both known and novel categories in an unlabeled dataset by leveraging another labeled dataset with only known categories. Without considering knowledge transfer from known to novel categories, current methods usually perform poorly on novel categories due to the lack of corresponding supervision. To mitigate this issue, we propose a unified Knowledge Transfer Network (KTN), which solves two obstacles to knowledge transfer in GCD. First, the mixture of known and novel categories in unlabeled data makes it difficult to identify transfer candidates (i.e., samples with novel categories). For this, we propose an entropy-based method that leverages knowledge in the pre-trained classifier to differentiate known and novel categories without requiring extra data or parameters. Second, the lack of prior knowledge of novel categories presents challenges in quantifying semantic relationships between categories to decide the transfer weights. For this, we model different categories with prototypes and treat their similarities as transfer weights to measure the semantic similarities between categories. On the basis of two treatments, we transfer knowledge from known to novel categories by conducting pre-adjustment of logits and post-adjustment of labels for transfer candidates based on the transfer weights between different categories. With the weighted adjustment, KTN can generate more accurate pseudo-labels for unlabeled data, which helps to learn more discriminative features and boost model performance on novel categories. Extensive experiments show that our method outperforms state-of-the-art models on all evaluation metrics across multiple benchmark datasets. Furthermore, different from previous clustering-based methods that can only work offline with abundant data, KTN can be deployed online conveniently with faster inference speed. Code and data are available at https://github.com/yibai-shi/KTN.
Wenkai Shi, Wenbin An, Feng Tian 0002, Yan Chen 0031, Yaqiang Wu, Qianying Wang 0002, Ping Chen 0001
AAAI5
2024 A Multi-Scale Bimodal Fusion Network for Robust and Accurate Online Handwriting Recognition
abstract
Online handwriting recognition based on sensor trajectory information faces several unresolved challenges: 1) sensor signals lack sufficient global spatial context; 2) different recognition tasks have inconsistent requirements for feature receptive fields. This is due to the inconsistent scales of the input sequences and the different semantic complexity of different language units. In this paper, we propose an online handwritten text recognition method based on multi-scale bimodal feature fusion to address these challenges. First, we employ sequence-generated pseudo-images to supplement the two-dimensional spatial information, and then extract multi-scale features from both trajectories and images simultaneously. Subsequently, our designed bimodal embedding learning module jointly learns feature embeddings for trajectories and images at different scales. These embeddings are then fed into a novel position-aware multi-scale fusion module to extract features for text prediction. The proposed modules effectively mitigate the issues of scales and semantics misalignment. Experimental results demonstrate significant performance improvements on various handwriting recognition datasets using our approach.
Yaqiang Wu, Wanjun Lv, Qianying Wang 0002
ICASSP3
2024 MRCI: Multi-range Context Interaction for Boundary Refinement in Image Segmentation
Yaqiang Wu, Wanjun Lyu, Xianchen Liang
ICPR (33)1
2024 RDLNet: A Novel and Accurate Real-world Document Localization Method
abstract
The increasing use of smartphones for capturing documents in various real-world conditions has underscored the need for robust document localization technologies. Current challenges in this domain include handling diverse document types, complex backgrounds, and varying photographic conditions such as low contrast and occlusion. However, there currently are no publicly available datasets containing these complex scenarios and few methods demonstrate their capabilities on these complex scenes. To address these issues, we create a new comprehensive real-world document localization benchmark dataset which contains the complex scenarios mentioned above and propose a novel Real-world Document Localization Network (RDLNet) for locating targeted documents in the wild. The RDLNet consists of an innovative light-SAM encoder and a masked attention decoder. Utilizing light-SAM encoder, the RDLNet transfers the mighty generalization capability of SAM to the document localization task. In the decoding stage, the RDLNet exploits the masked attention and object query method to efficiently output the triple-branch predictions consisting of corner point coordinates, instance-level segmentation area and categories of different documents without extra post-processing. We compare the performance of RDLNet with other state-of-the-art approaches for real-world document localization on multiple benchmarks, the results of which reveal that the RDLNet remarkably outperforms contemporary methods, demonstrating its superiority in terms of both accuracy and practicability.
Yaqiang Wu, Yanlai Wu, Hui Li 0130
ACM Multimedia1
2024 DOWN: Dynamic Order Weighted Network for Fine-grained Category Discovery
Wenbin An, Feng Tian 0002, Wenkai Shi, Haonan Lin, Yaqiang Wu, Mingxiang Cai, Luyan Wang, Hua Wen, Ping Chen 0001
Knowl. Based Syst.5
2024 Programming knowledge tracing based on heterogeneous graph representation
Yaqiang Wu, Fujian Song, Yan Chen 0031, Feng Tian 0002
Knowl. Based Syst.1
2024 TGIN: Translation-Based Graph Inference Network for Few-Shot Relational Triplet Extraction
abstract
Extracting relational triplets aims at detecting entity pairs and their semantic relations. Compared with pipeline models, joint models can reduce error propagation and achieve better performance. However, all of these models require large amounts of training data, therefore performing poorly on many long-tail relations in reality with insufficient data. In this article, we propose a novel end-to-end model, called TGIN, for few-shot triplet extraction. The core of TGIN is a multilayer heterogeneous graph with two types of nodes (entity node and relation node) and three types of edges (relation-entity edge, entity-entity edge, and relation-relation edge). On the one hand, this heterogeneous graph with entities and relations as nodes can intuitively extract relational triplets jointly, thereby reducing error propagation. On the other hand, it enables the triplet information of limited labeled data to interact better, thus maximizing the advantage of this information for few-shot triplet extraction. Moreover, we devise a graph aggregation and update method that utilizes translation algebraic operations to mine semantic features while retaining structure features between entities and relations, thereby improving the robustness of the TGIN in a few-shot setting. After updating the node and edge features through layers, TGIN propagates the label information from a few labeled examples to unlabeled examples, thus inferring triplets from these unlabeled examples. Extensive experiments on three reconstructed datasets demonstrate that TGIN can significantly improve the accuracy of triplet extraction by 2.34%~10.74% compared with the state-of-the-art baselines. To the best of our knowledge, we are the first to introduce a heterogeneous graph for few-shot relational triplet extraction.
Jiaxin Wang 0002, Lingling Zhang 0005, Jun Liu 0002, Kunming Ma, Xiang Zhao 0002, Yaqiang Wu, Yi Huang 0017
IEEE Trans. Neural Networks Learn. Syst.7
2023 GPTR: Gestalt-Perception Transformer for Diagram Object Detection
abstract
Diagram object detection is the key basis of practical applications such as textbook question answering. Because the diagram mainly consists of simple lines and color blocks, its visual features are sparser than those of natural images. In addition, diagrams usually express diverse knowledge, in which there are many low-frequency object categories in diagrams. These lead to the fact that traditional data-driven detection model is not suitable for diagrams. In this work, we propose a gestalt-perception transformer model for diagram object detection, which is based on an encoder-decoder architecture. Gestalt perception contains a series of laws to explain human perception, that the human visual system tends to perceive patches in an image that are similar, close or connected without abrupt directional changes as a perceptual whole object. Inspired by these thoughts, we build a gestalt-perception graph in transformer encoder, which is composed of diagram patches as nodes and the relationships between patches as edges. This graph aims to group these patches into objects via laws of similarity, proximity, and smoothness implied in these edges, so that the meaningful objects can be effectively detected. The experimental results demonstrate that the proposed GPTR achieves the best results in the diagram object detection task. Our model also obtains comparable results over the competitors in natural image object detection.
Lingling Zhang 0005, Jun Liu 0002, Jinfu Fan, Yang You 0001, Yaqiang Wu
AAAI6
2023 SHGAE: Social Hypergraph AutoEncoder for Friendship Inference
Yan Chen 0031, Tianliang Qi, Feng Tian 0002, Yaqiang Wu, Qianying Wang 0002
ICANN (6)5
2023 EnsExam: A Dataset for Handwritten Text Erasure on Examination Papers
Liufeng Huang, Bangdong Chen, Chongyu Liu, Dezhi Peng, Weiying Zhou, Yaqiang Wu, Hao Ni 0001
ICDAR (3)6
2023 A prediction model of student performance based on self-attention mechanism
Yan Chen 0031, Ganglin Wei, Yunwei Chen, Feng Tian 0002, Qianying Wang 0002, Yaqiang Wu
Knowl. Inf. Syst.9
2023 Side-by-Side vs Face-to-Face: Evaluating Colocated Collaboration via a Transparent Wall-sized Display
abstract
Traditional wall-sized displays mostly only support side-by-side co-located collaboration, while transparent displays naturally support face-to-face interaction. Many previous works assume transparent displays support collaboration. Yet it is unknown how exactly its afforded face-to-face interaction can support loose or close collaboration, especially compared to the side-by-side configuration offered by traditional large displays. In this paper, we used an established experimental task that operationalizes different collaboration coupling and layout locality, to compare pairs of participants collaborating side-by-side versus face-to-face in each collaborative situation. We compared quantitative measures and collected interview and observation data to further illustrate and explain our observed user behavior patterns. The results showed that the unique face-to-face collaboration brought by transparent display can result in more efficient task performance, different territorial behavior, and both positive and negative collaborative factors. Our findings provided empirical understanding about the collaborative experience supported by wall-sized transparent displays and shed light on its future design.
Jiangtao Gong, Mengdi Chu, Minghao Luo, Liuxin Zhang, Yaqiang Wu, Qianying Wang 0002, Can Liu 0003
Proc. ACM Hum. Comput. Interact.8
2023 Improved prototypical network for active few-shot learning
Yaqiang Wu, Yifei Li 0006, Tianzhe Zhao, Lingling Zhang 0005, Bifan Wei, Jun Liu 0002
Pattern Recognit. Lett.1
2023 Spatial-Semantic Collaborative Graph Network for Textbook Question Answering
abstract
Textbook Question Answering (TQA) task requires answering questions by reasoning based on both the given diagrams and text context. There are mainly two challenges for the task. First, the diagrams are different from the natural images. Similar shapes or color blocks may express different semantics and there is also a large intra-topic variation for diagrams. Hence, the characteristics of visual semantic ambiguity and variable visual appearance make the diagram understanding more challenging. Second, for the text, the specific education domain with terminologies exists a great gap with the general domain. Therefore, it is difficult to represent the text semantics effectively using a text encoder pretrained in the general domain. In this paper, we propose a Spatial-Semantic Collaborative Graph Network (SSCGN) for TQA task, which can help enhance the diagram and text understanding and facilitate multimodal reasoning. Specifically, the Spatial-guided Semantic Enhancing (SSE) module fully exploits the spatial and semantic relationships between visual objects and OCR tokens to collaboratively enhance the diagram semantic understanding. Moreover, based on the semantically enhanced region representations of the SSE module, the Fine-grained Spatial-Aware Graph Network (FSA-GN) can help obtain richer relation-aware region representations for joint reasoning by capturing more fine-grained spatial relationships. We further propose multiple self-supervised auxiliary tasks to enhance the initial diagram and text semantic representations by pretraining the diagram encoder and text encoder. Extensive experiments and ablation studies are conducted to validate the effectiveness of SSCGN.
Yaxian Wang, Bifan Wei, Jun Liu 0002, Qika Lin, Lingling Zhang 0005, Yaqiang Wu
IEEE Trans. Circuits Syst. Video Technol.6
2023 MuL-GRN: Multi-Level Graph Relation Network for Few-Shot Node Classification
abstract
Few-shot learning (FSL) that acquires new knowledge with little supervision, attracts much attention due to expensive cost of data annotation. Various meta-learning methods have made a great progress for few-shot problem in image and text data. In reality, data samples are not independent but rich in link relations. Large amounts of data exists in the form of graph structure such as citation, social, and biological networks. However, FSL study on graph data is still in its infancy because of the obstacle on extracting meta-knowledge from a meta node classification task. Current research just simply combines the FSL methods experienced in computer vision with node representation models together, but ignores the effect of rich links among support and query nodes in few-shot meta-task. For this issue, we propose a novel Multi-Level Graph Relation Network (MuL-GRN) for the challenging few-shot node classification. MuL-GRN extracts node embeddings through the popular graph neural networks (GNNs). And it includes a relation learning module to mine the deep node relations from three views, namely node-level, global subgraph-level, and local subgraph-level relations. For any two nodes, the node-level relation is computed on their node embeddings, global subgraph-level relation is measured on their subgraph embeddings, and the local subgraph-level relation is mined according to the pairwise node comparison information in their subgraphs. The three-view relation vectors are fused together with an interesting relation fusion module, which measures the importance of relation vector for the current few-shot classification task automatically. Extensive experiments on five real datasets show that MuL-GRN significantly outperforms existing state-of-the-art methods by a large margin.
Lingling Zhang 0005, Jun Liu 0002, Xiaojun Chang, Qika Lin, Yaqiang Wu
IEEE Trans. Knowl. Data Eng.6
2022 MatchPrompt: Prompt-based Open Relation Extraction with Semantic Consistency Guided Clustering
abstract
Relation clustering is a general approach for open relation extraction (OpenRE).Current methods have two major problems.One is that their good performance relies on large amounts of labeled and pre-defined relational instances for pre-training, which are costly to acquire in reality.The other is that they only focus on learning a high-dimensional metric space to measure the similarity of novel relations and ignore the specific relational representations of clusters.In this work, we propose a new prompt-based framework named Match-Prompt, which can realize OpenRE with efficient knowledge transfer from only a few predefined relational instances as well as mine the specific meanings for cluster interpretability.To our best knowledge, we are the first to introduce a prompt-based framework for unlabeled clustering.Experimental results on different datasets show that MatchPrompt achieves the new SOTA results for OpenRE.
Jiaxin Wang 0002, Lingling Zhang 0005, Jun Liu 0002, Yaqiang Wu
EMNLP6
2022 ICPR 2022 Challenge on Multi-Modal Subtitle Recognition
abstract
Video subtitle recognition, as one of the basic elements of video editing, has received increasing attention recently. However, the misaligned between audio and subtitles as well as the costly manual annotation remain a demanding issue toward subsequent intelligent processing. In this paper, we introduce a multi-modal subtitle recognition challenge for ICPR 2022, in which we present a large-scale video dataset (215 hours in total for visual and audio annotations), and 3 tracks including: 1) extracting subtitles in visual modality with audio annotation (ESV); 2) extracting subtitles in audio modality with visual annotations (ESA); and 3) extracting subtitles with both visual and audio annotation (ESVA). The challenge attracts 376 participants, among which the methods of top 3 teams on each track have been elaborated.
Shen Huang, Pengfei Hu 0004, Jian Kang 0006, Weida Liang, Yaqiang Wu, Yong Liu 0027
ICPR11
2022 FPX-NIC: An FPGA-Accelerated 4K Ultra-High-Definition Neural Video Coding System
abstract
The recent trend in neural image compression (NIC) research could be generally grounded into two categories: analysis-synthesis transform network improvements and entropy estimation optimization. They promote the compression efficiency of NIC by leveraging more expressive network structures and advanced entropy models respectively. From a different but more systematic viewpoint, we extend the horizon of NIC from software- to hardware-based lossy compression using more resource-constrained platforms, such as field programmable gate array (FPGA) or deep-learning processor unit (DPU). In this paper, we propose a novel hardware-oriented NIC system for real-time edge-computing video services. We for the first time present FPX-NIC, an FPGA-accelerated NIC framework designed for hardware encoding, which consists of a novel NIC scheme and an energy-efficient neural network (NN) deployment method. The former contribution is a block-based adaptive NIC approach based on local content characteristics. Essential side-information is also signalled to realize adaptive patch representation. The critical advantage of our latter contribution lies in the network-reconfigurable framework plus fixed-precision weights quantization method that takes advantage of quantization-aware post training procedure to compensate the performance degradation caused by quantization error. Therefore it is able to improve both processing speed and energy efficiency. We finally establish an intelligent video coding system using the proposed scheme, enabling visual capturing, neural encoding, decoding, and display, realizing 4K ultra-high-definition (UHD) all intra neural video coding on edge-computing devices.
Chuanmin Jia, Xinyu Hang, Shanshe Wang, Yaqiang Wu, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution
abstract
Visual Information Extraction (VIE) has attracted considerable attention recently owing to its various advanced applications such as document understanding, automatic marking and intelligent education. Most existing works decoupled this problem into several independent sub-tasks of text spotting (text detection and recognition) and information extraction, which completely ignored the high correlation among them during optimization. In this paper, we propose a robust Visual Information Extraction System (VIES) towards real-world scenarios, which is an unified end-to-end trainable framework for simultaneous text detection, recognition and information extraction by taking a single document image as input and outputting the structured information. Specifically, the information extraction branch collects abundant visual and semantic representations from text spotting for multimodal feature fusion and conversely, provides higher-level semantic clues to contribute to the optimization of text spotting. Moreover, regarding the shortage of public benchmarks, we construct a fully-annotated dataset called EPHOIE (https://github.com/HCIILAB/EPHOIE), which is the first Chinese benchmark for both text spotting and visual information extraction. EPHOIE consists of 1,494 images of examination paper head with complex layouts and background, including a total of 15,771 Chinese handwritten or printed text instances. Compared with the state-of-the-art methods, our VIES shows significant superior performance on the EPHOIE dataset and achieves a 9.01% F-score gain on the widely used SROIE dataset under the end-to-end scenario.
Chongyu Liu, Guozhi Tang, Jiaxin Zhang 0003, Shuaitao Zhang, Qianying Wang 0002, Yaqiang Wu, Mingxiang Cai
AAAI8
2021 Towards an Efficient Framework for Data Extraction from Chart Images
Weihong Ma, Hesuo Zhang, Shuang Yan, Guangshun Yao, Yichao Huang, Yaqiang Wu
ICDAR (1)7
2021 Towards Fast, Accurate and Compact Online Handwritten Chinese Text Recognition
Dezhi Peng, Canyu Xie, Zecheng Xie, Kai Ding 0009, Yichao Huang, Yaqiang Wu
ICDAR (3)8
2021 DeMatch: Towards Understanding the Panel of Chart Documents
Hesuo Zhang, Weihong Ma, Yichao Huang, Kai Ding 0009, Yaqiang Wu
ICDAR (3)6
2021 MatchVIE: Exploiting Match Relevancy between Entities for Visual Information Extraction
abstract
Visual Information Extraction (VIE) task aims to extract key information from multifarious document images (e.g., invoices and purchase receipts). Most previous methods treat the VIE task simply as a sequence labeling problem or classification problem, which requires models to carefully identify each kind of semantics by introducing multimodal features, such as font, color, layout. But simply introducing multimodal features can't work well when faced with numeric semantic categories or some ambiguous texts. To address this issue, in this paper we propose a novel key-value matching model based on a graph neural network for VIE (MatchVIE). Through key-value matching based on relevancy evaluation, the proposed MatchVIE can bypass the recognitions to various semantics, and simply focuses on the strong relevancy between entities. Besides, we introduce a simple but effective operation, Num2Vec, to tackle the instability of encoded values, which helps model converge more smoothly. Comprehensive experiments demonstrate that the proposed MatchVIE can significantly outperform previous methods. Notably, to the best of our knowledge, MatchVIE may be the first attempt to tackle the VIE task by modeling the relevancy between keys and values and it is a good complement to the existing methods.
Guozhi Tang, Lele Xie, Jingdong Chen, Qianying Wang 0002, Yaqiang Wu
IJCAI8
2020 Decoupled Attention Network for Text Recognition
abstract
Text recognition has attracted considerable research interests because of its various applications. The cutting-edge text recognition methods are based on attention mechanisms. However, most of attention methods usually suffer from serious alignment problem due to its recurrency alignment operation, where the alignment relies on historical decoding results. To remedy this issue, we propose a decoupled attention network (DAN), which decouples the alignment operation from using historical decoding results. DAN is an effective, flexible and robust end-to-end text recognizer, which consists of three components: 1) a feature encoder that extracts visual features from the input image; 2) a convolutional alignment module that performs the alignment operation based on visual features from the encoder; and 3) a decoupled text decoder that makes final prediction by jointly using the feature map and attention maps. Experimental results show that DAN achieves state-of-the-art performance on multiple text recognition tasks, including offline handwritten text recognition and regular/irregular scene text recognition. Codes will be released.1
Canjie Luo, Xiaoxue Chen, Yaqiang Wu, Qianying Wang 0002, Mingxiang Cai
AAAI6
2019 A Fast and Accurate Fully Convolutional Network for End-to-End Handwritten Chinese Text Segmentation and Recognition
abstract
Handwritten Chinese Text Recognition (HCTR) is a challenging problem due to its high complexity. Previous methods based on over-segmentation, hidden Markov model (HMM) or long short-term memory recurrent neural network (LSTM-RNN) have achieved great success in recognition results. However, all of them, including over-segmentation based methods, are incompetent in accurate segmentation of single character. To solve this problem, we propose a fast and accurate fully convolutional network for end-to-end segmentation and recognition of handwritten Chinese text. Experiments on CASIA-HWDB datasets and ICDAR 2013 competition dataset show that our method achieves a competitive performance on recognition and produces great character segmentation results. Moreover, our model reaches a real-time speed of 70 fps, which is fast enough for various applications.
Dezhi Peng, Yaqiang Wu, Zhepeng Wang 0002, Mingxiang Cai
ICDAR3
2019 Omnidirectional Scene Text Detection with Sequential-free Box Discretization
abstract
Scene text in the wild is commonly presented with high variant characteristics. Using quadrilateral bounding box to localize the text instance is nearly indispensable for detection methods. However, recent researches reveal that introducing quadrilateral bounding box for scene text detection will bring a label confusion issue which is easily overlooked, and this issue may significantly undermine the detection performance. To address this issue, in this paper, we propose a novel method called Sequential-free Box Discretization (SBD) by discretizing the bounding box into key edges (KE) which can further derive more effective methods to improve detection performance. Experiments showed that the proposed method can outperform state-of-the-art methods in many popular scene text benchmarks, including ICDAR 2015, MLT, and MSRA-TD500. Ablation study also showed that simply integrating the SBD into Mask R-CNN framework, the detection performance can be substantially improved. Furthermore, an experiment on the general object dataset HRSC2016 (multi-oriented ships) showed that our method can outperform recent state-of-the-art methods by a large margin, demonstrating its powerful generalization ability.
Sheng Zhang 0024, Lele Xie, Yaqiang Wu, Zhepeng Wang 0002
IJCAI5