Liang Qiao 0001

dblp:68/10765-1 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
17since 2021 · last 2025
0000-0003-4464-4644ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 GCSTG: Generating Class-Confusion-Aware Samples With a Tree-Structure Graph for Few-Shot Object Detection
abstract
Few-Shot Object Detection (FSOD) aims to detect the objects of novel classes using only a few manually annotated samples. With the few novel class samples, learning the inter-class relationships among foreground and constructing the corresponding class hierarchy in FSOD is a challenging task. The poor construction of the class hierarchy will result in the inter-class confusion problem, which has been identified as a primary cause of inferior performance in novel classes by recent FSOD methods. In this work, we further find that the intra-super-class confusion, where samples are misclassified as classes within their associated super-classes, is the main challenge in solving the confusion problem. To solve this issue, this work generates class-confusion-aware samples with a pre-defined tree-structure graph, for helping models to construct a precise class hierarchy. In precise, for generating class-confusion-aware samples, we add the noise into available samples and update the noise to maximize confidence scores on associated confusion categories of samples. Then, a confusion-aware curriculum learning strategy is proposed to make generated samples gradually participate in the training, which benefits the model convergence while learning the generated samples. Experimental results show that our method can be used as a plug-in in recent FSOD methods and consistently improve the model performance.
Longrong Yang, Hanbin Zhao, Hongliang Li 0001, Liang Qiao 0001, Xi Li 0001
IEEE Trans. Image Process.4
2024 Reading order detection in visually-rich documents with multi-modal layout-aware relation prediction
Liang Qiao 0001, Zhanzhan Cheng, Yunlu Xu, Xi Li 0001
Pattern Recognit.1
2023 Language Adaptive Weight Generation for Multi-Task Visual Grounding
abstract
Although the impressive performance in visual grounding, the prevailing approaches usually exploit the visual backbone in a passive way, i.e., the visual backbone extracts features with fixed weights without expression-related hints. The passive perception may lead to mismatches (e.g., redundant and missing), limiting further performance improvement. Ideally, the visual backbone should actively extract visual features since the expressions already provide the blueprint of desired visual features. The active perception can take expressions as priors to extract relevant visual features, which can effectively alleviate the mismatches. Inspired by this, we propose an active perception Visual Grounding framework based on Language Adaptive Weights, called VG-LAW. The visual backbone serves as an expression-specific feature extractor through dynamic weights generated for various expressions. Benefiting from the specific and relevant visual features extracted from the language-aware visual backbone, VG-LAW does not require additional modules for cross-modal interaction. Along with a neat multi-task head, VG-LAW can be competent in referring expression comprehension and segmentation jointly. Extensive experiments on four representative datasets, i.e., RefCOCO, RefCOCO+, RefCOCOg, and ReferItGame, validate the effectiveness of the proposed framework and demonstrate state-of-the-art performance.
Wei Su 0009, Peihan Miao 0002, Huanzhang Dou, Gaoang Wang, Liang Qiao 0001, Zheyang Li, Xi Li 0001
CVPR5
2023 Bridging Cross-task Protocol Inconsistency for Distillation in Dense Object Detection
abstract
Knowledge distillation (KD) has shown potential for learning compact models in dense object detection. However, the commonly used softmax-based distillation ignores the absolute classification scores for individual categories. Thus, the optimum of the distillation loss does not necessarily lead to the optimal student classification scores for dense object detectors. This cross-task protocol inconsistency is critical, especially for dense object detectors, since the foreground categories are extremely imbalanced. To address the issue of protocol differences between distillation and classification, we propose a novel distillation method with cross-task consistent protocols, tailored for the dense object detection. For classification distillation, we address the cross-task protocol inconsistency problem by formulating the classification logit maps in both teacher and student models as multiple binary-classification maps and applying a binary-classification distillation loss to each map. For localization distillation, we design an IoU-based Localization Distillation Loss that is free from specific network structures and can be compared with existing localization distillation losses. Our proposed method is simple but effective, and experimental results demonstrate its superiority over existing methods. Code is available at https://github.com/TinyTigerPan/BCKD.
Longrong Yang, Xianpan Zhou, Xuewei Li 0003, Liang Qiao 0001, Zheyang Li, Ziwei Yang 0004, Gaoang Wang, Xi Li 0001
ICCV4
2023 Generating Questions via Unexploited OCR Texts: Prompt-Based Data Augmentation for TextVQA
abstract
Text-based Visual Question Answering (TextVQA) tasks rely on Optical Character Recognition (OCR) text to answer. There have been many models successfully exploring multi-modal features fusing and knowledge reasoning. However, current TextVQA datasets are few and the cost of using manual annotation is too high. So generating pseudo-labeled data is a better choice. In this paper, a prompt-based data augmentation method is proposed. The problems of current data augmentation are solved: 1) the distribution of the number of answer words in the pseudo-labeled data is not consistent with the real dataset. 2) the question forms in the pseudo-labeled data are not diverse. Specifically, prompt words are first matched to the constraints in the questions by finding the same words in the vocabulary. So, our generating model can generate different types of questions when the different prompt words are input. Experiments show that our method is significantly better than other state-of-the-art methods on TextVQA.
Mingjie Han, Wancong Lin, Liang Qiao 0001
IJCNN5
2023 Finding Cycles in Graph: A Unified Approach for Various NER Tasks
abstract
Named Entity Recognition (NER) is the task of recognizing the entities' locations and types in text, which can be generally categorized into flat NER, overlapped NER, and discontinuous NER. Most previous methods are usually designed specifically for one of the tasks, such as sequence labeling approaches for flat NER and span-based models for overlapped NER. Recently, some new work has begun to propose the unified NER framework that can addresses all three scenarios simultaneously. However, there still has room for improvement in some complex scenarios (long/discontinuous entity). In this paper, we propose a concise framework that supports all types of NER tasks, where entities can be represented by unique cycles that are formed by the directed edges among tokens in the graph. The model integrates a Graph Feature Enhancement module to extract correlations at both the node-level and edge-level. At the node-level, the features are enhanced in the binary and ternary token relations. In edge-level, the model will go further to enhance the relations among token pairs using deformable convolutions. Furthermore, to benefit the completeness of cycle formation, we also propose a novel Cycle Loss that optimizes the independent edge classification in the group of cycles from a global perspective. Experimental results show that our model can achieve competitive and even new state-of-the-art performance on eight popular NER benchmarks, including flat NER, overlapped NER, and discontinuous NER.
Liang Qiao 0001, Xi Li 0001
IJCNN1
2022 Flooding-X: Improving BERT's Resistance to Adversarial Attacks via Loss-Restricted Fine-Tuning
abstract
Qin Liu, Rui Zheng, Bao Rong, Jingyi Liu, ZhiHua Liu, Zhanzhan Cheng, Liang Qiao, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Qin Liu 0010, Bao Rong, Zhanzhan Cheng, Liang Qiao 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)7
2022 MINER: Improving Out-of-Vocabulary Named Entity Recognition from an Information Theoretic Perspective
abstract
Xiao Wang, Shihan Dou, Limao Xiong, Yicheng Zou, Qi Zhang, Tao Gui, Liang Qiao, Zhanzhan Cheng, Xuanjing Huang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Xiao Wang 0001, Shihan Dou, Limao Xiong, Yicheng Zou, Qi Zhang 0001, Tao Gui, Liang Qiao 0001, Zhanzhan Cheng, Xuanjing Huang 0001
ACL (1)7
2022 Read Extensively, Focus Smartly: A Cross-document Semantic Enhancement Method for Visual Documents NER
abstract
The introduction of multimodal information and pretraining technique significantly improves entity recognition from visually-rich documents. However, most of the existing methods pay unnecessary attention to irrelevant regions of the current document while ignoring the potentially valuable information in related documents. To deal with this problem, this work proposes a cross-document semantic enhancement method, which consists of two modules: 1) To prevent distractions from irrelevant regions in the current document, we design a learnable attention mask mechanism, which is used to adaptively filter redundant information in the current document. 2) To further enrich the entity-related context, we propose a cross-document information awareness technique, which enables the model to collect more evidence across documents to assist in prediction. The experimental results on two documents understanding benchmarks covering eight languages demonstrate that our method outperforms the SOTA methods.
Jun Zhao 0019, Wenyu Zhan, Tao Gui, Qi Zhang 0001, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu
COLING6
2022 Dynamic Low-Resolution Distillation for Cost-Efficient End-to-End Text Spotting
Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xi Li 0001
ECCV (28)2
2022 End-to-End Compound Table Understanding with Multi-Modal Modeling
abstract
Table is a widely used data form in webpages, spreadsheets, or PDFs to organize and present structural data. Although studies on table structure recognition have been successfully used to convert image-based tables into digital structural formats, solving many real problems still relies on further understanding of the table, such as cell relationship extraction. The current datasets related to table understanding are all based on the digit format. To boost research development, we release a new benchmark named ComFinTab with rich annotations that support both table recognition and understanding tasks. Unlike previous datasets containing the basic tables, ComFinTab contains a large ratio of compound tables, which is much more challenging and requires methods using multiple information sources. Based on the dataset, we also propose a uniform, concise task form with the evaluation metric to better evaluate the model's performance on the table understanding task in compound tables. Finally, a framework named CTUNet is proposed to integrate the compromised visual, semantic, and position features with a graph attention network, which can solve the table recognition task and the challenging table understanding task as a whole. Experimental results compared with some previous advanced table understanding methods demonstrate the effectiveness of our proposed model. Code and dataset are available at \urlhttps://github.com/hikopensource/DAVAR-Lab-OCR.
Zaisheng Li, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xi Li 0001
ACM Multimedia3
2022 DavarOCR: A Toolbox for OCR and Multi-Modal Document Understanding
abstract
This paper presents DavarOCR, an open-source toolbox for OCR and document understanding tasks. DavarOCR currently implements 19 advanced algorithms, covering 9 different task forms. DavarOCR provides detailed usage instructions and the trained models for each algorithm. Compared with the previous open-source OCR toolbox, DavarOCR has relatively more complete support for the sub-tasks of the cutting-edge technology of document understanding. In order to promote the development and application of OCR technology in academia and industry, we pay more attention to the use of modules that different sub-domains of technology can share. DavarOCR is publicly released at https://github.com/hikopensource/Davar-Lab-OCR.
Liang Qiao 0001, Zaisheng Li, Baorui Zou, Dashan Guo, Yingda Xu, Yunlu Xu, Zhanzhan Cheng
ACM Multimedia1
2021 MANGO: A Mask Attention Guided One-Stage Scene Text Spotter
abstract
Recently end-to-end scene text spotting has become a popular research topic due to its advantages of global optimization and high maintainability in real applications. Most methods attempt to develop various region of interest (RoI) operations to concatenate the detection part and the sequence recognition part into a two-stage text spotting framework. However, in such framework, the recognition part is highly sensitive to the detected results (e.g., the compactness of text contours). To address this problem, in this paper, we propose a novel Mask AttentioN Guided One-stage text spotting framework named MANGO, in which character sequences can be directly recognized without RoI operation. Concretely, a position-aware mask attention module is developed to generate attention weights on each text instance and its characters. It allows different text instances in an image to be allocated on different feature map channels which are further grouped as a batch of instance features. Finally, a lightweight sequence decoder is applied to generate the character sequences. It is worth noting that MANGO inherently adapts to arbitrary-shaped text spotting and can be trained end-to-end with only coarse position information (e.g., rectangular bounding box) and text annotations. Experimental results show that the proposed method achieves competitive and even new state-of-the-art performance on both regular and irregular text spotting benchmarks, i.e., ICDAR 2013, ICDAR 2015, Total-Text, and SCUT-CTW1500.
Liang Qiao 0001, Zhanzhan Cheng, Yunlu Xu, Shiliang Pu, Fei Wu 0001
AAAI1
2021 A Strong Baseline for Semi-Supervised Incremental Few-Shot Learning
Linglan Zhao, Dashan Guo, Yunlu Xu, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xiangzhong Fang
BMVC4
2021 LGPMA: Complicated Table Structure Recognition with Local and Global Pyramid Mask Alignment
Liang Qiao 0001, Zaisheng Li, Zhanzhan Cheng, Peng Zhang 0075, Shiliang Pu, Wenqi Ren, Wenming Tan, Fei Wu 0001
ICDAR (1)1
2021 VSR: A Unified Framework for Document Layout Analysis Combining Vision, Semantics and Relations
Peng Zhang 0075, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Fei Wu 0001
ICDAR (1)3
2021 FREE: A Fast and Robust End-to-End Video Text Spotter
abstract
Currently, video text spotting tasks usually fall into the four-staged pipeline: detecting text regions in individual images, recognizing localized text regions frame-wisely, tracking text streams and post-processing to generate final results. However, they may suffer from the huge computational cost as well as sub-optimal results due to the interferences of low-quality text and the none-trainable pipeline strategy. In this article, we propose a fast and robust end-to-end video text spotting framework named FREE by only recognizing the localized text stream one-time instead of frame-wise recognition. Specifically, FREE first employs a well-designed spatial-temporal detector that learns text locations among video frames. Then a novel text recommender is developed to select the highest-quality text from text streams for recognizing. Here, the recommender is implemented by assembling text tracking, quality scoring and recognition into a trainable module. It not only avoids the interferences from the low-quality text but also dramatically speeds up the video text spotting. FREE unites the detector and recommender into a whole framework, and helps achieve global optimization. Besides, we collect a large scale video text dataset for promoting the video text spotting community, containing 100 videos from 21 real-life scenarios. Extensive experiments on public benchmarks show our method greatly speeds up the text spotting process, and also achieves the remarkable state-of-the-art.
Zhanzhan Cheng, Jing Lu 0004, Baorui Zou, Liang Qiao 0001, Yunlu Xu, Shiliang Pu, Fei Wu 0001, Shuigeng Zhou
IEEE Trans. Image Process.4
2020 Text Perceptron: Towards End-to-End Arbitrary-Shaped Text Spotting
abstract
Many approaches have recently been proposed to detect irregular scene text and achieved promising results. However, their localization results may not well satisfy the following text recognition part mainly because of two reasons: 1) recognizing arbitrary shaped text is still a challenging task, and 2) prevalent non-trainable pipeline strategies between text detection and text recognition will lead to suboptimal performances. To handle this incompatibility problem, in this paper we propose an end-to-end trainable text spotting approach named Text Perceptron. Concretely, Text Perceptron first employs an efficient segmentation-based text detector that learns the latent text reading order and boundary information. Then a novel Shape Transform Module (abbr. STM) is designed to transform the detected feature regions into regular morphologies without extra parameters. It unites text detection and the following recognition part into a whole framework, and helps the whole network achieve global optimization. Experiments show that our method achieves competitive performance on two standard text benchmarks, i.e., ICDAR 2013 and ICDAR 2015, and also obviously outperforms existing methods on irregular text benchmarks SCUT-CTW1500 and Total-Text.
Liang Qiao 0001, Sanli Tang, Zhanzhan Cheng, Yunlu Xu, Shiliang Pu, Fei Wu 0001
AAAI1
2020 TRIE: End-to-End Text Reading and Information Extraction for Document Understanding
abstract
Since real-world ubiquitous documents (e.g., invoices, tickets, resumes and leaflets) contain rich information, automatic document image understanding has become a hot topic. Most existing works decouple the problem into two separate tasks, (1) text reading for detecting and recognizing texts in images and (2) information extraction for analyzing and extracting key elements from previously extracted plain text.However, they mainly focus on improving information extraction task, while neglecting the fact that text reading and information extraction are mutually correlated. In this paper, we propose a unified end-to-end text reading and information extraction network, where the two tasks can reinforce each other. Specifically, the multimodal visual and textual features of text reading are fused for information extraction and in turn, the semantics in information extraction contribute to the optimization of text reading. On three real-world datasets with diverse document images (from fixed layout to variable layout, from structured text to semi-structured text), our proposed method significantly outperforms the state-of-the-art methods in both efficiency and accuracy.
Peng Zhang 0075, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu 0004, Liang Qiao 0001, Fei Wu 0001
ACM Multimedia6
2018 Queue State Based Dynamical Routing for Non-geostationary Satellite Networks
abstract
The actual queuing delay in satellite networks is hard to get due to long propagation. So, most existing routing algorithms take the expected queuing delay as the routing metrics so that links with short-time light traffic are often chosen when setting up routing tables, which results in that more packets could be sent to the nodes with short-time light traffic. In this paper, we propose a Queue State based Dynamical Routing (QSDR) mechanism for NGEO satellite networks. Instead of expected queuing delay, we model effective queuing delay through filtering short-time light traffic based on the proposed forgotten factor, which considers not only the traffic load but also their duration. To balance traffic load, we propose a dynamical route updating algorithm based on real-time queue states with route state model, which ensures that each satellite sends out packets as soon as possible and avoids congestion at current node. We develop a NS2-based simulation system to evaluate our QSDR. The results demonstrate that our QSDR outperforms related TLR and ELB in terms of packet drop rate, throughput and end-to-end delay.
Hezhong Li, Heteng Zhang, Liang Qiao 0001, Feilong Tang 0001, Wenchao Xu 0002, Long Chen 0025, Jie Li 0002
AINA3
2018 Feedback Based High-Quality Task Assignment in Collaborative Crowdsourcing
abstract
Collaborative Crowdsourcing focuses on tasks that need to be finished by a group of workers, usually is required to get information about worker affinities. However, there are few works study the problem how to get the accurate estimation of worker affinity along with worker skill. Under this situation, the task assignment problem in a cold-start collaborative crowdsourcing platform will be even harder. In this paper, we design a novel collaborative crowdsourcing framework that allows task requesters as well as workers to provide feedback to the task executors or co-workers. Using the information from feedbacks, we build the model to estimate worker affinity and worker skill more accurately and based on which we propose a measurement to measure the matching degree between task and a group of workers. Next, we propose a heuristic task assignment called FCC-SA which averagely distribute skillful workers to all collaborative tasks. Finally, we hire crowd workers to conduct real experiments to test our model and algorithm. Experimental results demonstrate that our FCC-SA algorithm significantly outperforms representative related proposals, and our novel measurement also has better precision than traditional measure method.
Liang Qiao 0001, Feilong Tang 0001, Jiacheng Liu 0001
AINA1