Yuanqiang Cai

dblp:96/5300 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Talk With Your Fingers: A Depth-Aware Benchmark for Air-Writing Recognition
abstract
Air-writing has emerged as a promising communication modality for AR/VR and metaverse environments, enabling quiet, non-contact text input by translating finger movements into natural language. However, existing approaches typically project in-air writing onto a virtual 2D plane and assume characters are formed with a single continuous stroke–an oversimplification that neglects the rich 3D structure inherent in natural handwriting. In this work, we challenge the “single-stroke 2D” paradigm and explore the role of depth cues in enhancing air-writing recognition. To this end, we present DAAWBench, the first large-scale Depth-Aware Air-Writing dataset, featuring 8.8 million RGB-D frames annotated with 3,755 Chinese characters from the GB2312-80 Level-1 set. Our analysis reveals consistent depth variations at stroke boundaries, indicating that stroke segmentation and character recognition can benefit from depth modeling. Based on these insights, we propose DARec, a novel 3D trajectory-based recognition model that effectively leverages depth-aware priors. Extensive experiments across in-domain and out-of-domain settings, including evaluations with vision-language models (e.g., GPT-4o, Qwen-VL) and human baselines, show that DARec significantly outperforms 2D-only counterparts, achieving 87.73% accuracy versus 9.05%. Our findings demonstrate the critical importance of depth modeling in human-computer co-creative interfaces, and we will publicly release our dataset and code at https://github.com/wmeiqi/DAAWBench.
Meiqi Wu, Yuzhong Zhao, Xuchen Li 0001, Yuanqiang Cai, Jiahong Wu 0005, Weiqiang Wang 0001, Kaiqi Huang
IEEE Trans. Circuits Syst. Video Technol.5
2025 A Bilingual, Open World Video Text Dataset and Real-Time Video Text Spotting With Contrastive Learning
abstract
Most existing video text spotting benchmarks focus on evaluating a single language and scenario with limited data. In this work, we introduce a large-scale, Bilingual, Open World Video text benchmark dataset (BOVText). There are four features for BOVText. Firstly, we provide 2,021 videos with more than 1,750,000 frames, 25 times larger than the existing largest dataset with incidental text in videos. Secondly, our dataset covers 32 open scenarios, including many virtual scenarios, e.g., Life Vlog, Driving, Movie, Game, etc. Thirdly, abundant text types annotation (i.e., title, caption or scene text) are provided for the different representational meanings in the video. Fourthly, the BOVText provides bilingual text annotation to promote multiple cultures’ lives and communication. Besides, we propose a real-time end-to-end video text spotting with Contrastive Learning of Semantic and Visual Representation (CoText), which includes two advantages: 1) With a lightweight architecture, CoText simultaneously addresses the three tasks (e.g., text detection, tracking, recognition) in a real-time end-to-end trainable framework. 2) CoText tracks texts by comprehending them and relating them to each other with visual and semantic representations. Extensive experiments show the superiority of our method. Especially, CoText achieves an video text spotting$\mathrm { ID_{F1}}$of 71.7% at 32.3 FPS on ICDAR2015video, with 10.2% and 23.3 FPS improvement the previous best method. The dataset and code of CoText can be found at: Dataset and CoText, respectively.
Weijia Wu 0001, Zhuang Li 0002, Yuanqiang Cai, Zheng Shou 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 PMRC: Prompt-Based Machine Reading Comprehension for Few-Shot Named Entity Recognition
abstract
The prompt-based method has been proven effective in improving the performance of pre-trained language models (PLMs) on sentence-level few-shot tasks. However, when applying prompting to token-level tasks such as Named Entity Recognition (NER), specific templates need to be designed, and all possible segments of the input text need to be enumerated. These methods have high computational complexity in both training and inference processes, making them difficult to apply in real-world scenarios. To address these issues, we redefine the NER task as a Machine Reading Comprehension (MRC) task and incorporate prompting into the MRC framework. Specifically, we sequentially insert boundary markers for various entity types into the templates and use these markers as anchors during the inference process to differentiate entity types. In contrast to the traditional multi-turn question-answering extraction in the MRC framework, our method can extract all spans of entity types in one round. Furthermore, we propose word-based template and example-based template that enhance the MRC framework's perception of entity start and end positions while significantly reducing the manual effort required for template design. It is worth noting that in cross-domain scenarios, PMRC does not require redesigning the model architecture and can continue training by simply replacing the templates to recognize entity types in the target domain. Experimental results demonstrate that our approach outperforms state-of-the-art models in low-resource settings, achieving an average performance improvement of +5.2% in settings where access to source domain data is limited. Particularly, on the ATIS dataset with a large number of entity types and 10-shot setting, PMRC achieves a performance improvement of +15.7%. Moreover, our method achieves a decoding speed 40.56 times faster than the template-based cloze-style approach.
Danfeng Yan, Yuanqiang Cai
AAAI3
2024 View-Category Interactive Sharing Transformer for Incomplete Multi-View Multi-Label Learning
abstract
As a problem often encountered in real-world scenarios, multi-view multi-label learning has attracted considerable research attention. However, due to oversights in data col-lection and uncertainties in manual annotation, real-world data often suffer from incompleteness. Regrettably, most existing multi-view multi-label learning methods sidestep missing views and labels. Furthermore, they often neglect the potential of harnessing complementary information be-tween views and labels, thus constraining their classification capabilities. To address these challenges, we propose a view-category interactive sharing transformer tailored for incomplete multi-view multi-label learning. Within this net-work, we incorporate a two-layer transformer module to characterize the interplay between views and labels. Additionally, to address view incompleteness, a KNN-style missing view generation module is employed. Finally, we in-troduce a view-category consistency guided embedding en-hancement module to align different views and improve the discriminating power of the embeddings. Collectively, these modules synergistically integrate to classify the incomplete multi-view multi-label data effectively. Extensive experi-ments substantiate that our approach outperforms the ex-isting state-of-the-art methods.
Shilong Ou, Zhe Xue, Yawen Li 0001, Meiyu Liang, Yuanqiang Cai, Junjiang Wu
CVPR5
2024 Restoring Real-World Images Affected by Varied Degradations Using a Semi-Supervised Domain Adaptation Network
abstract
Restoring real-world images suffering from complex degradations like haze, rain, and blur is a significant challenge. Existing models face difficulties when applied to these real-world images, mainly due to the domain gap between synthetically generated and authentic degradations. In this paper, we propose SDA-Net, a well-designed Semi-supervised Domain Adaptation Network that can effectively restore real-world images. Our method combines the advantages of two foundation models. The supervised knowledge transfer model helps a student network learn from diverse restoration networks, while the unsupervised domain adaptation models guide the student network to generalize from synthetic scenes to real-world scenes. Furthermore, we propose a grained-friendly contrastive learning loss to force our models to retain background details and clear representation. Extensive experiments demonstrate that our SDA-Net outperforms state-of-the-art algorithms on three common real-world datasets with various degradations, achieving a significant improvement of 2.7 scores on BRISQUE and 11.6 scores on PIQE.
Yongheng Zhang 0003, Yuanqiang Cai, Danfeng Yan
ICME2
2024 End-to-End Video Text Spotting with Transformer
Weijia Wu 0001, Yuanqiang Cai, Chunhua Shen, Debing Zhang, Ying Fu 0001, Ping Luo 0002
Int. J. Comput. Vis.2
2024 Spatio-temporal human action localization in indoor surveillances
Zihao Liu 0011, Danfeng Yan, Yuanqiang Cai
Pattern Recognit.3
2024 Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
abstract
Air-writing is a challenging task that combines the fields of computer vision and natural language processing, offering an intuitive and natural approach for human-computer interaction. However, current air-writing solutions face two primary challenges: (1) their dependency on complex sensors (e.g., Radar, EEGs and others) for capturing precise handwritten trajectories, and (2) the absence of a video-based air-writing dataset that covers a comprehensive vocabulary range. These limitations impede their practicality in various real-world scenarios, including the use on devices like iPhones and laptops. To tackle these challenges, we present the groundbreaking air-writing Chinese character video dataset (AWCV-100K), serving as a pioneering benchmark for video-based air-writing. This dataset captures handwritten trajectories in various real-world scenarios using commonly accessible RGB cameras, eliminating the need for complex sensors. AWCV-100K includes 8.8 million video frames, encompassing the complete set of 3,755 characters from the GB2312-80 level-1 set (GB1). Furthermore, we introduce our baseline approach, the video-based character recognizer (VCRec). VCRec adeptly extracts fingertip features from sparse visual cues and employs a spatio-temporal sequence module for analysis. Experimental results showcase the superior performance of VCRec compared to existing models in recognizing air-written characters, both quantitatively and qualitatively. This breakthrough paves the way for enhanced human-computer interaction in real-world contexts. Moreover, our approach leverages affordable RGB cameras, enabling its applicability in a diverse range of scenarios. The code and data examples will be made public at https://github.com/wmeiqi/AWCV.
Meiqi Wu, Kaiqi Huang, Yuanqiang Cai, Yuzhong Zhao, Weiqiang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Real-World Scene Image Enhancement with Contrastive Domain Adaptation Learning
abstract
Image enhancement methods leveraging learning-based approaches have demonstrated impressive results when trained on synthetic degraded-clear image pairs. However, when deployed in real-world scenarios, such models often suffer significant performance degradation due to the inherent domain gap between synthetic and real degradations. To bridge this gap, we propose a novel Two-stage Contrastive Domain Adaptation image Enhancement (TCDAE) framework consisting of two key strategies: (1) Synthetic-to-Real Domain Transfer Learning (S2R-DTL) that effectively translates images from the synthetic degraded domain to the real degraded domain, aligning the domains at the pixel level, and (2) Degraded-to-Clear Domain Transfer Learning (D2C-DTL) that further adapts the enhancement model from the synthetic to the real domain by translating images from the real degraded domain to the real clean domain in both supervised and unsupervised branches. A unique aspect of our approach is the integration of a Domain Noise Contrastive Estimation (DoNCE) loss in both learning strategies. This specialized loss formulation enables TCDAE to robustly translate images across domains, even in scenarios lacking strong positive examples. Consequently, our framework can generate enhanced images with natural, realistic appearances akin to real clear images. Comprehensive experiments on real-world degraded scenes across diverse tasks, including dehazing, deraining, and deblurring, demonstrate the superiority of TCDAE over state-of-the-art methods, achieving improved visual quality, quantitative metrics, and downstream task performance.
Yongheng Zhang 0003, Yuanqiang Cai, Danfeng Yan, Rongheng Lin
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Exploring highly concise and accurate text matching model with tiny weights
abstract
In this paper, we propose a simple and general lightweight approach named AL-RE2 for text matching models, and conduct experiments on three well-studied benchmark datasets across tasks of natural language inference and paraphrase identification. Firstly, we explore the feasibility of dimensional compression of word embedding vectors using principal component analysis, and then analyze the impact of the information retained in different dimensions on model accuracy. Considering the balance between compression efficiency and information loss, we choose 128 dimensions to represent each word and make the model params 1.6M. Finally, the feasibility of applying depthwise separable convolution instead of standard convolution in the field of text matching is analyzed in detail. The experimental results show that our model’s inference speed is at least 1.5 times faster and it has 42.76% fewer parameters compared to similarly performing models, while its accuracy on the SciTail dataset of is state-of-the-art among all lightweight models.
Yangchun Li, Danfeng Yan, Yuanqiang Cai, Zhihong Tian 0001
World Wide Web (WWW)4
2023 Semi-Swinderain: Semi-Supervised Image Deraining Network Using SWIN Transformer
abstract
Currently, single image deraining lacks paired rain/clean images in real world and most studies use synthetic data. Real rain image deraining is still a challenge. To solve this problem, we propose a semi-supervised image deraining network using Swin Transformer, which can both use features of synthetic data and real data to get a better result. Specifically, the network is divided into supervised branch and unsupervised branch. Supervised and unsupervised branches are trained using synthetic data and real data, respectively. The network architecture is based on Swin Transformer, which adds a self-supervised memory module between encoder and decoder to store rain information. In the unsupervised branch, contrastive loss is added to ensure restored real rain image in features space is close to clear image, away from real rain image. In addition, we propose a real rain dataset RealRain11k. Experiments show our method has better result in real rain image deraining. The source code and RealRain11k are available at https://github.com/imissrc/Semi-SwinDerain.
Chun Ren, Danfeng Yan, Yuanqiang Cai, Yangchun Li
ICASSP3
2023 Explore Faster Localization Learning For Scene Text Detection
abstract
Generally, pre-training and long-time training computation are necessary for obtaining a good-performance text detector based on deep networks. In this paper, we present a new scene text detection network (called FANet) with a Fast convergence speed and Accurate text localization. The proposed FANet is an end-to-end text detector based on transformer feature learning and normalized Fourier descriptor modeling, where the Fourier Descriptor Proposal Network and Iterative Text Decoding Network are designed to efficiently and accurately identify text proposals. Additionally, a Dense Matching Strategy and a well-designed loss function are also proposed for optimizing the network performance. Extensive experiments are carried out to demonstrate that the proposed FANet can achieve the SOTA performance with fewer training epochs and no pretraining. When we introduce additional data for pre-training, the proposed FANet can achieve SOTA performance on MSRA-TD500, CTW1500, and TotalText. The ablation experiments also verify the effectiveness of our contributions. Code is available at https://github.com/callsys/FANet.
Yuzhong Zhao, Yuanqiang Cai, Weijia Wu 0001, Weiqiang Wang 0001
ICME2
2022 Direct regression scene text detection with accuracy scoring
Peirui Cheng, Yuzhong Zhao, Yuanqiang Cai, Weiqiang Wang 0001
Neurocomputing3
2021 Rethinking Object Detection in Retail Stores
abstract
The conventional standard for object detection uses a bounding box to represent each individual object instance. However, it is not practical in the industry-relevant applications in the context of warehouses due to severe occlusions among groups of instances of the same categories. In this paper, we propose a new task, i.e., simultaneously object localization and counting, abbreviated as Locount, which requires algorithms to localize groups of objects of interest with the number of instances. However, there does not exist a dataset or benchmark designed for such a task. To this end, we collect a large-scale object localization and counting dataset with rich annotations in retail stores, which consists of 50,394 images with more than 1.9 million object instances in 140 categories. Together with this dataset, we provide a new evaluation protocol and divide the training and testing subsets to fairly evaluate the performance of algorithms for Locount, developing a new benchmark for the Locount task. Moreover, we present a cascaded localization and counting network as a strong baseline, which gradually classifies and regresses the bounding boxes of objects with the predicted numbers of instances enclosed in the bounding boxes, trained in an end-to-end manner. Extensive experiments are conducted on the proposed dataset to demonstrate its significance and the analysis is provided to indicate future directions. Dataset is available at https://isrc.iscas.ac.cn/gitlab/research/locount-dataset.
Yuanqiang Cai, Longyin Wen, Libo Zhang 0001, Dawei Du, Weiqiang Wang 0001
AAAI1
2021 Scale-Residual Learning Network for Scene Text Detection
abstract
Detecting incidentally captured text in the wild remains an open problem due to challenging factors including unconstrained scenarios and large scale variation. In this paper, we establish a large-scale scene text detection dataset (LS-Text), containing 36, 000 images and 270, 783 text instances with various scales and complex scenarios, to promote the research of text detection. We propose a Scale-residual Learning Network (SLN) to deal with the scale variation problem in a progressive optimization manner. Specifically, we integrate both learnable feature concatenation and feature up-sampling operator. It can effectively eliminate the residuals between the outputs of SLN and ground-truth text instances by processing both the Feature Fusion Residuals (FFR) and the Scale Transformation Residuals (STR), simultaneously. By stacking multi-scale feature maps in a deep-to-shallow manner, SLN continuously optimizes feature representation by accumulating strong semantic information and rich texture details in a scale-residual learning way. Extensive experimental results on five challenging datasets demonstrate the state-of-the-art performance of the proposed SLN model, and the challenging aspects related to real-world scenarios of the proposed LS-Text dataset. Both the source code of SLN and the LS-Text dataset are available athttps://github.com/SLN-Text-Detection.
Yuanqiang Cai, Chang Liu 0047, Peirui Cheng, Dawei Du, Libo Zhang 0001, Weiqiang Wang 0001, Qixiang Ye
IEEE Trans. Circuits Syst. Video Technol.1
2020 Guided Attention Network for Object Detection and Counting on Drones
abstract
Object detection and counting are related but challenging problems, especially for drone based scenes with small objects and cluttered background. In this paper, we propose a new Guided Attention network (GAnet) to deal with both object detection and counting tasks based on the feature pyramid. Different from the previous methods relying on unsupervised attention modules, we fuse different scales of feature maps by using the proposed weakly-supervised Background Attention (BA) between the background and objects for more semantic feature representation. Then, the Foreground Attention (FA) module is developed to consider both global and local appearance of the object to facilitate accurate localization. Moreover, the new data argumentation strategy is designed to train a robust model in the drone based scenes with various illumination conditions. Extensive experiments on three challenging benchmarks (i.e., UAVDT, CARPK and PUCPR+) show the state-of-the-art detection and counting performance of the proposed method compared with existing methods. Code can be found at https://isrc.iscas.ac.cn/gitlab/research/ganet.
Yuanqiang Cai, Dawei Du, Libo Zhang 0001, Longyin Wen, Weiqiang Wang 0001, Siwei Lyu
ACM Multimedia1
2020 Robustly detect different types of text in videos
Yuanqiang Cai, Weiqiang Wang 0001
Neural Comput. Appl.1
2020 SPN: short path network for scene text detection
Yuanqiang Cai, Weiqiang Wang 0001, Haiqing Ren, Ke Lu 0002
Neural Comput. Appl.1
2020 IOS-Net: An inside-to-outside supervision network for scale robust text detection in the wild
Yuanqiang Cai, Weiqiang Wang 0001, Qixiang Ye
Pattern Recognit.1
2020 A Direct Regression Scene Text Detector With Position-Sensitive Segmentation
abstract
Direct regression methods have demonstrated their success on various multi-oriented benchmarks for scene text detection due to the high recall rate for small targets and the direct regression for text boxes. However, too many false positive candidates and inaccurate position regression still limit the performance of these methods. In this paper, we propose an end-to-end method by introducing position-sensitive segmentation into the direct regression method to overcome these shortcomings. We generate the ground truth of position-sensitive segmentation maps based on the information of text boxes so that the position-sensitive segmentation module can be trained synchronously with the direct regression module. Besides, more information about the relative position of text is provided for the network through the training of position-sensitive segmentation maps, which improves the expressiveness of the network. We also introduce spatial pyramid of position-sensitive segmentation into the proposed method considering the huge differences in sizes and aspect ratios of scene texts and we propose position-sensitive COI(Corner area of Interest) pooling into the proposed method to speed up the inference. Experiments on datasets ICDAR2015, MLT-17 and COCO-Text demonstrate that the proposed method has a comparable performance with state-of-the-art methods while it is more efficient. We also provide abundant ablation experiments to demonstrate the effectiveness of these improvements in our proposed method.
Peirui Cheng, Yuanqiang Cai, Weiqiang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2019 Multi-scale Scene Text Detection via Resolution Transform
abstract
Scene text detection is a challenging task because there are many small text targets in the natural scene and the size of scene text varies greatly. For the current popular scene text detection methods, such as EAST and Textboxes, it is difficult to detect small text and text with large differences in size well at the same time. To solve the problem, we propose a novel multi-scale scene text detection method based on EAST. The proposed method extracts high-resolution feature maps at multiple scales via resolution transform and detects text on these feature maps. Through the resolution transform module, the proposed method can detect both kinds of text well. Besides, we use aggregated feature pyramid module to efficiently pass both low-level and high-level information to feature maps at each scale. Experiments on datasets ICDAR2015 and COCO-Text demonstrate that the proposed method has a comparable performance with state-of-the-art methods and it is more efficient. For ICDAR2015 and COCO-Text datasets, the proposed method achieves an F-score of 0.84 and 0.42 respectively.
Peirui Cheng, Weiqiang Wang 0001, Yuanqiang Cai
ICME3
2019 A new hybrid-parameter recurrent neural network for online handwritten chinese character recognition
Haiqing Ren, Weiqiang Wang 0001, Xiwen Qu, Yuanqiang Cai
Pattern Recognit. Lett.4
2018 Spatiotemporal text localization for videos
Yuanqiang Cai, Weiqiang Wang 0001, Shao Huang, Ke Lu 0002
Multim. Tools Appl.1
2016 MLPF algorithm for tracking fast moving target against light interference
abstract
In order to deal with the difficulty of tracking the fast moving aerial targets with light interference, we propose an improved particle tracking algorithm named multi-layers particle filter (MLPF). In MLPF, the particles are divided into three categories: the main particles (M-particles), the subordinate particles (S-particles) and the regenerate particles (R-particles). In the phase of resampling and state estimating, only M-particles are involved, then the R-particles are generated and considered as new S-particles in the next cycle. To a certain extent, our algorithm maintains the diversity of particles and reduces the computation time. Besides, MLPF has significant improvements on overcoming the tracing error after the sudden disappearance of the target and solving the degradation of particles. We demonstrate effectiveness of our proposed algorithm through systematic experiments. Experimental results show MLPF has better tracking effect compared to the traditional particle filter (PF) when the target is moving fast and affected by light interference. In the first experiment, the running time has been reduced from 47s to 21s while the precision increased from 64% to 96%. And for the second experiment, the running time has been reduced from 237s to 121s while precision increased from 46% to 89%.
Libo Zhang 0001, Yuanqiang Cai, Zakir Ullah, Tiejian Luo
ICPR2