EDBT 2026 Demo / reviewers in the wild / expert
Deqiang Jiang
dblp:259/2591
· DBLP profile ↗
26ranked-venue papers
0as first author
25since 2021 · last 2026
0000-0003-3987-2431ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 19 since 2021Artificial intelligence and machine learning · 18 · 17 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating the Adversarial Robustness of Vision-Language Models via Internal Feature PerturbationsabstractVision-language models (VLMs), such as BLIP-2 and LLaVA, have significantly advanced multimodal understanding but exhibit critical vulnerabilities to visual adversarial perturbations. The high efficacy of untargeted attacks, in particular, poses significant concerns for their operational robustness. Conventional attack methods generate adversarial examples by backpropagating the language modeling loss from the final output to the input image. However, the deep architecture of the integrated large language models (LLMs) often diminishes this gradient flow, limiting the attack’s effectiveness in perturbing the visual domain. To address this limitation, we introduce a novel untargeted attack method based on Maximizing Information Entropy (MIE). Our approach enhances attack efficacy not only by maximizing the information entropy of the model’s final output but also by directly inducing uncertainty within its internal feature representations. This layer-wise perturbation strategy disrupts the model’s cognitive process more comprehensively than relying on the final output layer alone. We provide a theoretical analysis demonstrating that the uncertainty induced by MIE is greater than or equal to that of conventional output-only attacks. Comprehensive quantitative evaluations across multiple VLM architectures and datasets confirm that our method consistently outperforms existing techniques, thereby establishing a new, more rigorous benchmark for assessing adversarial robustness in vision-language models. Chaohu Liu, Yubo Wang 0010, Haoyu Cao 0001, Deqiang Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | DREAM: Document Reconstruction via End-to-end Autoregressive Model
Xin Li 0118, Mingming Gong, Jianxin Dai, Antai Guo, Xinghua Jiang, Haoyu Cao 0001, Yinsong Liu, Deqiang Jiang, Xing Sun 0001 |
ACM Multimedia | 9 |
| 2024 | Grab What You Need: Rethinking Complex Table Structure Recognition with Flexible Components DeliberationabstractRecently, Table Structure Recognition (TSR) task, aiming at identifying table structure into machine readable formats, has received increasing interest in the community. While impressive success, most single table component-based methods can not perform well on unregularized table cases distracted by not only complicated inner structure but also exterior capture distortion. In this paper, we raise it as Complex TSR problem, where the performance degeneration of existing methods is attributable to their inefficient component usage and redundant post-processing. To mitigate it, we shift our perspective from table component extraction towards the efficient multiple components leverage, which awaits further exploration in the field. Specifically, we propose a seminal method, termed GrabTab, equipped with newly proposed Component Deliberator, to handle various types of tables in a unified framework. Thanks to its progressive deliberation mechanism, our GrabTab can flexibly accommodate to most complex tables with reasonable components selected but without complicated post-processing involved. Quantitative experimental results on public benchmarks demonstrate that our method significantly outperforms the state-of-the-arts, especially under more challenging scenes. Hao Liu 0003, Xin Li 0118, Mingming Gong, Deqiang Jiang, Yinsong Liu, Xing Sun 0001 |
AAAI | 6 |
| 2024 | Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and ReactionabstractLarge Language Model (LLM)-enhanced agents become increasingly prevalent in Human-AI communication, offering vast potential from entertainment to professional domains.However, current multi-modal dialogue systems overlook the acoustic information present in speech, which is crucial for understanding human communication nuances.This oversight can lead to misinterpretations of speakers' intentions, resulting in inconsistent or even contradictory responses within dialogues.To bridge this gap, in this paper, we propose PerceptiveAgent, an empathetic multi-modal dialogue system designed to discern deeper or more subtle meanings beyond the literal interpretations of words through the integration of speech modality perception.Employing LLMs as a cognitive core, Per-ceptiveAgent perceives acoustic information from input speech and generates empathetic responses based on speaking styles described in natural language.Experimental results indicate that PerceptiveAgent excels in contextual understanding by accurately discerning the speakers' true intentions in scenarios where the linguistic meaning is either contrary to or inconsistent with the speaker's true feelings, producing more nuanced and expressive spoken dialogues.Code is publicly available at: https://github.com/Haoqiu-Yan/ PerceptiveAgent. Haoqiu Yan, Yongxin Zhu 0003, Haoyu Cao 0001, Deqiang Jiang, Linli Xu 0002 |
ACL (1) | 6 |
| 2024 | Few-shot Temporal Pruning Accelerates Diffusion Models for Text GenerationabstractDiffusion models have achieved significant success in computer vision and shown immense potential in natural language processing applications, particularly for text generation tasks. However, generating high-quality text using these models often necessitates thousands of iterations, leading to slow sampling rates. Existing acceleration methods either neglect the importance of the distribution of sampling steps, resulting in compromised performance with smaller number of iterations, or require additional training, introducing considerable computational overheads. In this paper, we present Few-shot Temporal Pruning, a novel technique designed to accelerate diffusion models for text generation without supplementary training while effectively leveraging limited data. Employing a Bayesian optimization approach, our method effectively eliminates redundant sampling steps during the sampling process, thereby enhancing the generation speed. A comprehensive evaluation of discrete and continuous diffusion models across various tasks, including machine translation, question generation, and paraphrasing, reveals that our approach achieves competitive performance even with minimal sampling steps after down to less than 1 minute of optimization, yielding a significant acceleration of up to 400x in text generation tasks. Bocheng Li, Zhujin Gao, Yongxin Zhu 0003, Kun Yin, Haoyu Cao 0001, Deqiang Jiang, Linli Xu 0002 |
LREC/COLING | 6 |
| 2024 | Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language ModelsabstractRecently, the advent of Large Visual-Language Models (LVLMs) has received increasing attention across various domains, particularly in the field of visual document understanding (VDU). Different from conventional vision-language tasks, VDU is specifically concerned with text-rich scenarios containing abundant document elements. Nevertheless, the importance of fine-grained features remains largely unexplored within the community of LVLMs, leading to suboptimal performance in text-rich scenarios. In this paper, we abbreviate it as the fine-grained feature collapse issue. With the aim of filling this gap, we propose a contrastive learning framework, termed Document Object COntrastive learning (DoCo), specifically tailored for the downstream tasks of VDU. DoCo leverages an auxiliary multimodal encoder to obtain the features of document objects and align them to the visual features generated by the vision encoder of LVLM, which enhances visual representation in text-rich scenarios. It can represent that the contrastive learning between the visual holistic representations and the multimodal fine-grained features of document objects can assist the vision encoder in acquiring more effective visual cues, thereby enhancing the comprehension of text-rich documents in LVLMs. We also demonstrate that the proposed DoCo serves as a plug-and-play pre-training method, which can be employed in the pre-training of various LVLMs without inducing any increase in computational complexity during the inference process. Extensive experimental results on multiple benchmarks of VDU reveal that LVLMs equipped with our proposed DoCo can achieve superior performance and mitigate the gap between VDU and generic vision-language tasks. Xin Li 0118, Xinghua Jiang, Mingming Gong, Haoyu Cao 0001, Yinsong Liu, Deqiang Jiang, Xing Sun 0001 |
CVPR | 8 |
| 2024 | HRVDA: High-Resolution Visual Document AssistantabstractLeveraging vast training data, multimodal large language models (MLLMs) have demonstrated formidable general visual comprehension capabilities and achieved remarkable performance across various tasks. However, their performance in visual document understanding still leaves much room for improvement. This discrepancy is primarily attributed to the fact that visual document understanding is a fine-grained prediction task. In natural scenes, MLLMs typically use low-resolution images, leading to a substantial loss of visual information. Furthermore, general-purpose MLLMs do not excel in handling document-oriented instructions. In this paper, we propose a High-Resolution Visual Document Assistant (HRVDA), which bridges the gap between MLLMs and visual document understanding. This model employs a content filtering mechanism and an instruction filtering module to separately filter out the content-agnostic visual tokens and instruction-agnostic visual tokens, thereby achieving efficient model training and inference for high-resolution images. In addition, we construct a document-oriented visual instruction tuning dataset and apply a multi-stage training strategy to enhance the model's document modeling capabilities. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple document understanding datasets, while maintaining training efficiency and inference speed comparable to low-resolution models. Chaohu Liu, Kun Yin, Haoyu Cao 0001, Xinghua Jiang, Xin Li 0118, Yinsong Liu, Deqiang Jiang, Xing Sun 0001, Linli Xu 0002 |
CVPR | 7 |
| 2024 | Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language ModelsabstractLarge vision-language models (LVLMs) integrate visual information into large language models, showcasing remarkable multi-modal conversational capabilities. However, the visual modules introduces new challenges in terms of robustness for LVLMs, as attackers can craft adversarial images that are visually clean but may mislead the model to generate incorrect answers. In general, LVLMs rely on vision encoders to transform images into visual tokens, which are crucial for the language models to perceive image contents effectively. Therefore, we are curious about one question: Can LVLMs still generate correct responses when the encoded visual tokens are attacked and disrupting the visual information? To this end, we propose a non-targeted attack method referred to as VT-Attack (Visual Tokens Attack), which constructs adversarial examples from multiple perspectives, with the goal of comprehensively disrupting feature representations and inherent relationships as well as the semantic properties of visual tokens output by image encoders. Using only access to the image encoder in the proposed attack, the generated adversarial examples exhibit transferability across diverse LVLMs utilizing the same image encoder and generality across different tasks. Extensive experiments validate the superior attack performance of the VT-Attack over baseline methods, demonstrating its effectiveness in attacking LVLMs with image encoders, which in turn can provide guidance on the robustness of LVLMs, particularly in terms of the stability of the visual feature space. Yubo Wang 0010, Chaohu Liu, Yanqiu Qu, Haoyu Cao 0001, Deqiang Jiang, Linli Xu 0002 |
ACM Multimedia | 5 |
| 2024 | Communication-efficient clustered federated learning via model distance
Mao Zhang 0002, Yifei Cheng 0002, Changcun Bao, Haoyu Cao 0001, Deqiang Jiang, Linli Xu 0002 |
Mach. Learn. | 6 |
| 2023 | The Devil Is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-trainingabstractThe self-supervised Masked Image Modeling (MIM) schema, following "mask-and-reconstruct" pipeline of recovering contents from masked image, has recently captured the increasing interest in the community, owing to the excellent ability of learning visual representation from unlabeled data. Aiming at learning representations with high semantics abstracted, a group of works attempts to reconstruct non-semantic pixels with large-ratio masking strategy, which may suffer from "over-smoothing" problem, while others directly infuse semantics into targets in off-line way requiring extra data. Different from them, we shift the perspective to the Fourier domain which naturally has global perspective and present a new Masked Image Modeling (MIM), termed Geminated Gestalt Autoencoder (Ge^2-AE) for visual pre-training. Specifically, we equip our model with geminated decoders in charge of reconstructing image contents from both pixel and frequency space, where each other serves as not only the complementation but also the reciprocal constraints. Through this way, more robust representations can be learned in the pre-trained encoders, of which the effectiveness is confirmed by the juxtaposing experimental results on downstream recognition tasks. We also conduct several quantitative and qualitative experiments to investigate the learning behavior of our method. To our best knowledge, this is the first MIM work to solve the visual pre-training through the lens of frequency domain. Hao Liu 0003, Xinghua Jiang, Xin Li 0118, Antai Guo, Yiqing Hu, Deqiang Jiang, Bo Ren 0002 |
AAAI | 6 |
| 2023 | TaCo: Textual Attribute Recognition via Contrastive LearningabstractAs textual attributes like font are core design elements of document format and page style, automatic attributes recognition favor comprehensive practical applications. Existing approaches already yield satisfactory performance in differentiating disparate attributes, but they still suffer in distinguishing similar attributes with only subtle difference. Moreover, their performance drop severely in real-world scenarios where unexpected and obvious imaging distortions appear. In this paper, we aim to tackle these problems by proposing TaCo, a contrastive framework for textual attribute recognition tailored toward the most common document scenes. Specifically, TaCo leverages contrastive learning to dispel the ambiguity trap arising from vague and open-ended attributes. To realize this goal, we design the learning paradigm from three perspectives: 1) generating attribute views, 2) extracting subtle but crucial details, and 3) exploiting valued view pairs for learning, to fully unlock the pre-training potential. Extensive experiments show that TaCo surpasses the supervised counterparts and advances the state-of-the-art remarkably on multiple attribute recognition tasks. Online services of TaCo will be made available. Chang Nie, Yiqing Hu, Yanqiu Qu, Hao Liu 0003, Deqiang Jiang, Bo Ren 0002 |
AAAI | 5 |
| 2023 | Turning a CLIP Model into a Scene Text DetectorabstractThe recent large-scale Contrastive Language-Image Pretraining (CLIP) model has shown great potential in various downstream tasks via leveraging the pretrained vision and language knowledge. Scene text, which contains rich textual and visual information, has an inherent connection with a model like CLIP. Recently, pretraining approaches based on vision language models have made effective progresses in the field of text detection. In contrast to these works, this paper proposes a new method, termed TCM, focusing on Turning the CLIP Model directly for text detection without pretraining process. We demonstrate the advantages of the proposed TCM as follows: (1) The underlying principle of our framework can be applied to improve existing scene text detector. (2) It facilitates the few-shot training capability of existing methods, e.g., by using 10% of labeled data, we significantly improve the performance of the baseline method with an average of 22% in terms of the F-measure on 4 benchmarks. (3) By turning the CLIP model into existing scene text detection methods, we further achieve promising domain adaptation ability. The code will be publicly released at https://github.com/wenwenyu/TCM. Wenwen Yu, Wei Hua 0005, Deqiang Jiang, Bo Ren 0002, Xiang Bai |
CVPR | 4 |
| 2023 | Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region ConcentrationabstractWe propose a novel end-to-end document understanding model called SeRum (SElective Region Understanding Model) for extracting meaningful information from document images, including document analysis, retrieval, and office automation. Unlike state-of-the-art approaches that rely on multi-stage technical schemes and are computationally expensive, SeRum converts document image understanding and recognition tasks into a local decoding process of the visual tokens of interest, using a content-aware token merge module. This mechanism enables the model to pay more attention to regions of interest generated by the query decoder, improving the model’s effectiveness and speeding up the decoding speed of the generative scheme. We also designed several pre-training tasks to enhance the understanding and local awareness of the model. Experimental results demonstrate that SeRum achieves state-of-the-art performance on document understanding tasks and competitive results on text spotting tasks. SeRum represents a substantial advancement towards enabling efficient and effective end-to-end document understanding. Haoyu Cao 0001, Changcun Bao, Chaohu Liu, Kun Yin, Hao Liu 0003, Yinsong Liu, Deqiang Jiang, Xing Sun 0001 |
ICCV | 8 |
| 2023 | Visual Information Extraction in the Wild: Practical Dataset and End-to-End Solution
Jianfeng Kuang, Wei Hua 0005, Dingkang Liang, Deqiang Jiang, Bo Ren 0002, Xiang Bai |
ICDAR (6) | 5 |
| 2022 | Sequence-to-Action: Grammatical Error Correction with Action Guided Sequence GenerationabstractThe task of Grammatical Error Correction (GEC) has received remarkable attention with wide applications in Natural Language Processing (NLP) in recent years. While one of the key principles of GEC is to keep the correct parts unchanged and avoid over-correction, previous sequence-to-sequence (seq2seq) models generate results from scratch, which are not guaranteed to follow the original sentence structure and may suffer from the over-correction problem. In the meantime, the recently proposed sequence tagging models can overcome the over-correction problem by only generating edit operations, but are conditioned on human designed language-specific tagging labels. In this paper, we combine the pros and alleviate the cons of both models by proposing a novel Sequence-to-Action (S2A) module. The S2A module jointly takes the source and target sentences as input, and is able to automatically generate a token-level action sequence before predicting each token, where each action is generated from three choices named SKIP, COPY and GENerate. Then the actions are fused with the basic seq2seq framework to provide final predictions. We conduct experiments on the benchmark datasets of both English and Chinese GEC tasks. Our model consistently outperforms the seq2seq baselines, while being able to significantly alleviate the over-correction problem as well as holding better generality and diversity in the generation results compared to the sequence tagging models. Jiquan Li, Junliang Guo, Yongxin Zhu 0003, Xin Sheng 0003, Deqiang Jiang, Bo Ren 0002, Linli Xu 0002 |
AAAI | 5 |
| 2022 | Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text RecognitionabstractWe introduce Perceiving Stroke-Semantic Context (PerSec), a new approach to self-supervised representation learning tailored for Scene Text Recognition (STR) task. Considering scene text images carry both visual and semantic properties, we equip our PerSec with dual context perceivers which can contrast and learn latent representations from low-level stroke and high-level semantic contextual spaces simultaneously via hierarchical contrastive learning on unlabeled text image data. Experiments in un- and semi-supervised learning settings on STR benchmarks demonstrate our proposed framework can yield a more robust representation for both CTC-based and attention-based decoders than other contrastive learning methods. To fully investigate the potential of our method, we also collect a dataset of 100 million unlabeled text images, named UTI-100M, covering 5 scenes and 4 languages. By leveraging hundred-million-level unlabeled data, our PerSec shows significant performance improvement when fine-tuning the learned representation on the labeled data. Furthermore, we observe that the representation learned by PerSec presents great generalization, especially under few labeled data scenes. Hao Liu 0003, Bin Wang 0070, Zhimin Bao, Mobai Xue, Sheng Kang, Deqiang Jiang, Yinsong Liu, Bo Ren 0002 |
AAAI | 6 |
| 2022 | NomMer: Nominate Synergistic Context in Vision Transformer for Visual RecognitionabstractRecently, Vision Transformers (ViT), with the self-attention (SA) as the de facto ingredients, have demon-strated great potential in the computer vision community. For the sake of trade-off between efficiency and performance, a group of works merely perform SA operation within local patches, whereas the global contextual information is abandoned, which would be indispensable for visual recognition tasks. To solve the issue, the subsequent global-local ViTs take a stab at marrying local SA with global one in parallel or alternative way in the model. Nevertheless, the exhaustively combined local and global context may exist redundancy for various visual data, and the receptive field within each layer is fixed. Alternatively, a more graceful way is that global and local context can adaptively contribute per se to accommodate different visual data. To achieve this goal, we in this paper propose a novel ViT architecture, termed NomMer, which can dynamically Nominate the synergistic global-local context in vision transforMer. By investigating the working pattern of NomMer, we further explore what context information is focused. Beneficial from this “dynamic nomination” mechanism, without bells and whistles, the NomMer can not only achieve 84.5% Top-1 classification accuracy on ImageNet with only 73M parameters, but also show promising performance on dense prediction tasks, i.e., object detection and semantic segmentation. The code and models are publicly available at https://github.com/TencentYoutuResearch/VisualRecognition-NomMer. Hao Liu 0003, Xinghua Jiang, Xin Li 0118, Zhimin Bao, Deqiang Jiang, Bo Ren 0002 |
CVPR | 5 |
| 2022 | Neural Collaborative Graph Machines for Table Structure RecognitionabstractRecently, table structure recognition has achieved impressive progress with the help of deep graph models. Most of them exploit single visual cues of tabular elements or simply combine visual cues with other modalities via early fusion to reason their graph relationships. However, neither early fusion nor individually reasoning in terms of multiple modalities can be appropriate for all varieties of table structures with great diversity. Instead, different modalities are expected to collaborate with each other in different patterns for different table cases. In the community, the importance of intrainter modality interactions for table structure reasoning is still unexplored. In this paper, we define it as heterogeneous table structure recognition (HeteroTSR) problem. With the aim offilling this gap, we present a novel Neural Collaborative Graph Machines (NCGM) equipped with stacked collaborative blocks, which alternatively extracts intramodality context and models inter-modality interactions in a hierarchical way. It can represent the intrainter modality relationships of tabular elements more robustly, which significantly improves the recognition performance. We also show that the proposed NCGM can modulate collaborative pattern of different modalities conditioned on the context of intramodality cues, which is vital for diversified table cases. Experimental results on benchmarks demonstrate our proposed NCGM achieves state-of-the-art performance and beats other contemporary methods by a large margin especially under challenging scenarios. Hao Liu 0003, Xin Li 0118, Deqiang Jiang, Yinsong Liu, Bo Ren 0002 |
CVPR | 4 |
| 2022 | Query-driven Generative Network for Document Information Extraction in the WildabstractThis paper focuses on solving Document Information Extraction (DIE) in the wild problem, which is rarely explored before. In contrast to existing studies mainly tailored for document cases in known templates with predefined layouts and keys under the ideal input without OCR errors involved, we aim to build up a more practical DIE paradigm for real-world scenarios where input document images may contain unknown layouts and keys in the scenes of the problematic OCR results. To achieve this goal, we propose a novel architecture, termed Query-driven Generative Network (QGN), which is equipped with two consecutive modules, i.e., Layout Context-aware Module (LCM) and Structured Generation Module (SGM). Given a document image with unseen layouts and fields, the former LCM yields the value prefix candidates serving as the query prompts for the SGM to generate the final key-value pairs even with OCR noise. To further investigate the potential of our method, we create a new large-scale dataset, named LArge-scale STructured Documents (LastDoc4000), containing 4,000 documents with 1,511 layouts and 3,500 different keys. In experiments, we demonstrate that our QGN consistently achieves the best F1-score on the new LastDoc4000 dataset by at most 30.32% absolute improvement. A more comprehensive experimental analysis and experiments on other public benchmarks also verify the effectiveness and robustness of our proposed method for the wild DIE task. Haoyu Cao 0001, Xin Li 0118, Jiefeng Ma, Deqiang Jiang, Antai Guo, Yiqing Hu, Hao Liu 0003, Yinsong Liu, Bo Ren 0002 |
ACM Multimedia | 4 |
| 2022 | Relational Representation Learning in Visually-Rich DocumentsabstractRelational understanding is critical for a number of visually-rich documents (VRDs) understanding tasks. Through multi-modal pre-training, recent studies provide comprehensive contextual representations and exploit them as prior knowledge for downstream tasks. In spite of their impressive results, we observe that the widespread relational hints (e.g., relation of key/value fields on receipts) built upon contextual knowledge are not excavated yet. To mitigate this gap, we propose DocReL, a Document Relational Representation Learning framework. The major challenge of DocReL roots in the variety of relations. From the simplest pairwise relation to the complex global structure, it is infeasible to conduct supervised training due to the definition of relation varies and even conflicts in different tasks. To deal with the unpredictable definition of relations, we propose a novel contrastive learning task named Relational Consistency Modeling (RCM), which harnesses the fact that existing relations should be consistent in differently augmented positive views. RCM provides relational representations which are more compatible to the urgent need of downstream tasks, even without any knowledge about the exact definition of relation. DocReL achieves better performance on a wide variety of VRD relational understanding tasks, including table structure recognition, key information extraction and reading order detection. Xin Li 0118, Yiqing Hu, Haoyu Cao 0001, Deqiang Jiang, Yinsong Liu, Bo Ren 0002 |
ACM Multimedia | 6 |
| 2022 | OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and ClassificationabstractScene segmentation and classification (SSC) serve as a critical step towards the field of video structuring analysis. Intuitively, jointly learning of these two tasks can promote each other by sharing common information. However, scene segmentation concerns more on the local difference between adjacent shots while classification needs the global representation of scene segments, which probably leads to the model dominated by one of the two tasks in the training phase. In this paper, from an alternate perspective to overcome the above challenges, we unite these two tasks into one task by a new form of predicting shots link: a link connects two adjacent shots, indicating that they belong to the same scene or category. To the end, we propose a general One Stage Multimodal Sequential Link Framework (OS-MSL) to both distinguish and leverage the two-fold semantics by reforming the two learning tasks into a unified one. Furthermore, we tailor a specific module called DiffCorrNet to explicitly extract the information of differences and correlations among shots. Extensive experiments on a brand-new large scale dataset collected from real-world applications, and MovieScenes are conducted. Both the results demonstrate the effectiveness of our proposed method against strong baselines. The code is made available. Ye Liu 0013, Lingfeng Qiao, Zhuoxuan Jiang, Xinghua Jiang, Deqiang Jiang, Bo Ren 0002 |
ACM Multimedia | 6 |
| 2022 | GMN: Generative Multi-modal Network for Practical Document Information ExtractionabstractHaoyu Cao, Jiefeng Ma, Antai Guo, Yiqing Hu, Hao Liu, Deqiang Jiang, Yinsong Liu, Bo Ren. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Haoyu Cao 0001, Jiefeng Ma, Antai Guo, Yiqing Hu, Hao Liu 0003, Deqiang Jiang, Yinsong Liu, Bo Ren 0002 |
NAACL-HLT | 6 |
| 2021 | Hierarchical Multi-label Text Classification with Horizontal and Vertical Category CorrelationsabstractHierarchical multi-label text classification (HMTC) deals with the challenging task where an instance can be assigned to multiple hierarchically structured categories at the same time.The majority of prior studies either focus on reducing the HMTC task into a flat multi-label problem ignoring the vertical category correlations or exploiting the dependencies across different hierarchical levels without considering the horizontal correlations among categories at the same level, which inevitably leads to fundamental information loss.In this paper, we propose a novel HMTC framework that considers both vertical and horizontal category correlations.Specifically, we first design a loosely coupled graph convolutional neural network as the representation extractor to obtain representations for words, documents, and, more importantly, level-wise representations for categories, which are not considered in previous works.Then, the learned category representations are adopted to capture the vertical dependencies among levels of category hierarchy and model the horizontal correlations.Finally, based on the document embeddings and category embeddings, we design a hybrid algorithm to predict the categories of the entire hierarchical structure.Extensive experiments conducted on real-world HMTC datasets validate the effectiveness of the proposed framework with significant improvements over the baselines. Linli Xu 0002, Sijie Teng, Junliang Guo, Deqiang Jiang, Bo Ren 0002 |
EMNLP (1) | 6 |
| 2021 | RecycleNet: An Overlapped Text Instance Recovery ApproachabstractText recognition is the key pillar for many real-world multimedia applications. Existing text recognition approaches focus on recognizing isolated instances, whose text fields are visually separated and have no interference with each other. Moreover, these approaches cannot handle overlapped instances that often appear in sheets like invoices, receipts and math exercises, where printed templates are generated beforehand and extra contents are added afterward on existing texts. In this paper, we aim to tackle this problem by proposing RecycleNet, which automatically extracts and reconstructs overlapped instances by fully recycling the intersecting pixels that used to be obstacles for recognition. RecycleNet parallels to existing recognition systems, and serves as a plug-and-play module to boost recognition performance with zero-effort. We also released an OverlapText-500 dataset, which helps to boost the design of better overlapped text recovery and recognition solutions. Yiqing Hu, Xinghua Jiang, Hao Liu 0003, Deqiang Jiang, Yinsong Liu, Bo Ren 0002, Rongrong Ji |
ACM Multimedia | 5 |
| 2021 | Show, Read and Reason: Table Structure Recognition with Flexible Context AggregatorabstractWe investigate the challenging problem of table structure recognition in this work. Many recent methods adopt graph-based context aggregator with strong inductive bias to reason sparse contextual relationships of table elements. However, the strong constraints may be too restrictive to represent the complicated table relationships. In order to learn more appropriate inductive bias from data, we try to introduce Transformer as context aggregator in this work. Nevertheless, Transformer taking dense context as input requires larger scale data and may suffer from unstable training procedure due to the weakening of inductive bias. To overcome the above limitations, we in this paper design a FLAG (FLexible context AGgregator), which marries Transformer with graph-based context aggregator in an adaptive way. Based on FLAG, an end-to-end framework requiring no extra meta-data or OCR information, termed FLAG-Net, is proposed to flexibly modulate the aggregation of dense context and sparse one for the relational reasoning of table elements. We investigate the modulation pattern in FLAG and show what contextual information is focused, which is vital for recognizing table structure. Extensive experimental results on benchmarks demonstrate the performance of our proposed FLAG-Net surpasses other compared methods by a large margin. Hao Liu 0003, Xin Li 0118, Deqiang Jiang, Yinsong Liu, Bo Ren 0002, Rongrong Ji |
ACM Multimedia | 4 |
| 2020 | Accurate Structured-Text Spotting for Arithmetical Exercise CorrectionabstractCorrecting arithmetical exercise is a labor intensive and time consuming task for primary school teachers all the time. To reduce their burdens, we propose Arithmetical Exercise Checker (AEC), which is the first system that automatically evaluates all arithmetical expressions (AEs) on exercise images. The major challenge is that AE is formed by printed and handwritten texts with particular arithmetical patterns (e.g., multi-line, fraction). Despite being part of AE, handwritten texts usually lead to zigzag boundaries and tangled rows. What's worse, AE may be arithmetical incorrect, which makes the contextual information less valuable for recognition. To tackle these problems, we introduce integrated detection, recognition and evaluation branches by leveraging AE's intrinsic features, namely 1) boundary indistinctive, 2) locally relevant patterns and 3) globally irrelevant symbols. Experimental results demonstrate that AEC yields a 93.72% correction accuracy on 40 kinds of mainstream primary arithmetical exercises. So far, the online service of AEC processes 75, 000 arbitrary exercises on average per day, and already reduced the burden of over 1, 000, 000 users. AEC shows the benefits for implementing an vision-based system as a way to aid teachers in reducing reduplicative tasks. Yiqing Hu, Hao Liu 0003, Deqiang Jiang, Yinsong Liu, Bo Ren 0002 |
AAAI | 4 |