Dongbao Yang

dblp:193/2602 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
26since 2021 · last 2026
0000-0001-8628-411XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 17 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Towards Breaking the Visual Perception Bottleneck for Geometry Problem Solving
Tianjiao Cao, Jiahao Lyu 0002, Dongbao Yang, Weimin Mu, Yu Zhou 0015
ICDAR (2)3
2026 EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration
Daiqing Wu, Dongbao Yang, Can Ma, Yu Zhou 0015
Pattern Recognit.2
2026 Resolving sentiment discrepancy for multimodal sentiment detection via semantics completion and decomposition
Daiqing Wu, Dongbao Yang, Huawen Shen, Can Ma, Yu Zhou 0015
Pattern Recognit.2
2025 Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
abstract
Scene text spotting has attracted the enthusiasm of relative researchers in recent years. Most existing scene text spotters follow the detection-then-recognition paradigm, where the vanilla detection module hardly determines the reading order and leads to failure recognition. After rethinking the auto-regressive scene text recognition method, we find that a well-trained recognizer can implicitly perceive the local semantics of all characters in a complete word or a sentence without a character-level detection module. Local semantic knowledge not only includes text content but also spatial information in the right reading order. Motivated by the above analysis, we propose the Local Semantics Guided scene text Spotter (LSGSpotter), which auto-regressively decodes the position and content of characters guided by the local semantics. Specifically, two effective modules are proposed in LSGSpotter. On the one hand, we design a Start Point Localization Module (SPLM) for locating text start points to determine the right reading order. On the other hand, a Multi-scale Adaptive Attention Module (MAAM) is proposed to adaptively aggregate text features in a local area. In conclusion, LSGSpotter achieves the arbitrary reading order spotting task without the limitation of sophisticated detection, while alleviating the cost of computational resources with the grid sampling strategy. Extensive experiment results show LSGSpotter achieves state-of-the-art performance on the InverseText benchmark. Moreover, our spotter demonstrates superior performance on English benchmarks for arbitrary-shaped text, achieving improvements of 0.7% and 2.5% on Total-Text and SCUT-CTW1500, respectively. These results validate our text spotter is effective for scene texts in arbitrary reading order and shape.
Jiahao Lyu 0002, Wei Wang 0315, Dongbao Yang, Jinwen Zhong, Yu Zhou 0015
AAAI3
2025 Specifying What You Know or Not for Multi-Label Class-Incremental Learning
abstract
Existing class incremental learning is mainly designed for single-label classification task, which is ill-equipped for multi-label scenarios due to the inherent contradiction of learning objectives for samples with incomplete labels. We argue that the main challenge to overcome this contradiction in multi-label class-incremental learning (MLCIL) lies in the model's inability to clearly distinguish between known and unknown knowledge. This ambiguity hinders the model's ability to retain historical knowledge, master current classes, and prepare for future learning simultaneously. In this paper, we target at specifying what is known or not to accommodate Historical, Current, and Prospective knowledge for MLCIL and propose a novel framework termed as HCP. Specifically, (i) we clarify the known classes by dynamic feature purification and recall enhancement with distribution prior, enhancing the precision and retention of known information. (ii) We design prospective knowledge mining to probe the unknown, preparing the model for future learning. Extensive experiments validate that our method effectively alleviates catastrophic forgetting in MLCIL, surpassing the previous state-of-the-art by 3.3% on average accuracy for MS-COCO B0-C10 setting without replay buffers.
Aoting Zhang, Dongbao Yang, Xiaopeng Hong, Yu Zhou 0015
AAAI2
2025 DCA: Dividing and Conquering Amnesia in Incremental Object Detection
abstract
Incremental object detection (IOD) aims to cultivate an object detector that can continuously localize and recognize novel classes while preserving its performance on previous classes. Existing methods achieve certain success by improving knowledge distillation and exemplar replay for transformer-based detection frameworks, but the intrinsic forgetting mechanisms remain underexplored. In this paper, we dive into the cause of forgetting and discover forgetting imbalance between localization and recognition in transformer-based IOD, which means that localization is less-forgetting and can generalize to future classes, whereas catastrophic forgetting occurs primarily on recognition. Based on these insights, we propose a Divide-and-Conquer Amnesia (DCA) strategy, which redesigns the transformer-based IOD into a localization-then-recognition process. DCA can well maintain and transfer the localization ability, leaving decoupled fragile recognition to be specially conquered. To reduce feature drift in recognition, we leverage semantic knowledge encoded in pre-trained language models to anchor class representations within a unified feature space across incremental tasks. This involves designing a duplex classifier fusion and embedding class semantic features into the recognition decoding process in the form of queries. Extensive experiments validate that our approach achieves state-of-the-art performance, especially for long-term incremental scenarios. For example, under the four-step setting on MS-COCO, our DCA strategy significantly improves the final AP by 6.9%.
Aoting Zhang, Dongbao Yang, Xiaopeng Hong, Miao Shang, Yu Zhou 0015
AAAI2
2025 TADoc: Robust Time-Aware Document Image Dewarping
abstract
Flattening curved, wrinkled, and rotated document images captured by portable photographing devices, termed document image dewarping, has become an increasingly important task with the rise of digital economy and online working. Although many methods have been proposed recently, they often struggle to achieve satisfactory results when confronted with intricate document structures and higher degrees of deformation in real-world scenarios. Our main insight is that, unlike other document restoration tasks (e.g., deblurring), dewarping in real physical scenes is a progressive motion rather than a one-step transformation. Based on this, we have undertaken two key initiatives. Firstly, we reformulate this task, modeling it for the first time as a dynamic process that encompasses a series of intermediate states. Secondly, we design a lightweight framework called TADoc (Time-Aware Document Dewarping Network) to address the geometric distortion of document images. In addition, due to the inadequacy of OCR metrics for document images containing sparse text, the comprehensiveness of evaluation is insufficient. To address this shortcoming, we propose a new metric – DLS (Document Layout Similarity) – to evaluate the effectiveness of document dewarping in downstream tasks. Extensive experiments and in-depth evaluations have been conducted and the results indicate that our model possesses strong robustness, achieving superiority on several benchmarks with different document types and degrees of distortion.
Fangmin Zhao, Weichao Zeng, Zhenhang Li, Dongbao Yang, Yu Zhou 0015
ECAI4
2025 PerturbCTC: Improving Alignment in Scene Text Recognition with Feature Perturbation Based CTC
Zhijie Shen, Yaqiang Wu, Gangyan Zeng, Dongbao Yang, Yu Zhou 0015
ICDAR (4)6
2025 An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
abstract
The advancements in Multimodal Large Language Models (MLLMs) have enabled various multimodal tasks to be addressed under a zero-shot paradigm. This paradigm sidesteps the cost of model fine-tuning, emerging as a dominant trend in practical application. Nevertheless, Multimodal Sentiment Analysis (MSA), a pivotal challenge in the quest for general artificial intelligence, fails to accommodate this convenience. The zero-shot paradigm exhibits undesirable performance on MSA, casting doubt on whether MLLMs can perceive sentiments as competent as supervised models. By extending the zero-shot paradigm to In-Context Learning (ICL) and conducting an in-depth study on configuring demonstrations, we validate that MLLMs indeed possess such capability. Specifically, three key factors that cover demonstrations' retrieval, presentation, and distribution are comprehensively investigated and optimized. A sentimental predictive bias inherent in MLLMs is also discovered and later effectively counteracted. By complementing each other, the devised strategies for three factors result in average accuracy improvements of 15.9% on six MSA datasets against the zero-shot paradigm and 11.2% against the random ICL baseline.
Daiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma, Yu Zhou 0015
ICML2
2025 The Role of Video Generation in Enhancing Data-Limited Action Understanding
abstract
Video action understanding tasks in real-world scenarios often suffer from data limitations. In this paper, we address the data-limited action understanding problem by bridging data scarcity. We propose a novel method that leverages a text-to-video diffusion transformer to generate annotated data for model training. This paradigm enables the generation of realistic annotated data on an infinite scale without human intervention. We proposed the Information Enhancement Strategy and the Uncertainty-Based Soft Target tailored to generate sample training. Through quantitative and qualitative analyzes, we discovered that real samples generally contain a richer level of information compared to generated samples. Based on this observation, the information enhancement strategy was designed to enhance the informational content of the generated samples from two perspectives: the environment and the character. Furthermore, we observed that a portion of low-quality generated samples might negatively affect model training. To address this, we devised an uncertainty-based label-smoothing strategy to increase the smoothing of these low-quality samples, thereby reducing their impact. We demonstrate the effectiveness of the proposed method on four datasets and five tasks, and achieve state-of-the-art performance for zero-shot action recognition.
Dezhao Luo, Dongbao Yang, Zhenhang Li, Weiping Wang 0005, Yu Zhou 0015
IJCAI3
2025 Uni-DocDiff: A Unified Document Restoration Model Based on Diffusion
abstract
Removing various degradations from damaged documents greatly benefits digitization, downstream document analysis, and readability. Previous methods often treat each restoration task independently with dedicated models, leading to a cumbersome and highly complex document processing system. Although recent studies attempt to unify multiple tasks, they often suffer from limited scalability due to handcrafted prompts and heavy preprocessing, and fail to fully exploit inter-task synergy within a shared architecture. To address the aforementioned challenges, we propose Uni-DocDiff, a Unified and highly scalable Doc ument restoration model based on Dif fusion. Uni-DocDiff develops a learnable task prompt design, ensuring exceptional scalability across diverse tasks. To further enhance its multi-task capabilities and address potential task interference, we devise a novel Prior Pool, a simple yet comprehensive mechanism that combines both local high-frequency features and global low-frequency features. Additionally, we design the Prior Fusion Module (PFM), which enables the model to adaptively select the most relevant prior information for each specific task. Extensive experiments show that the versatile Uni-DocDiff achieves performance comparable or even superior performance compared with task-specific expert models, and simultaneously holds the task scalability for seamless adaptation to new tasks.
Fangmin Zhao, Weichao Zeng, Zhenhang Li, Dongbao Yang, Binbin Li 0003, Xiaojun Bi 0002, Yu Zhou 0015
ACM Multimedia4
2024 First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
abstract
Diffusion models, known for their impressive image generation abilities, have played a pivotal role in the rise of visual text generation. Nevertheless, existing visual text generation methods often focus on generating entire images with text prompts, leading to imprecise control and limited practicality. A more promising direction is visual text blending, which focuses on seamlessly merging texts onto text-free backgrounds. However, existing visual text blending methods often struggle to generate high-fidelity and diverse images due to a shortage of backgrounds for synthesis and limited generalization capabilities. To overcome these challenges, we propose a new visual text blending paradigm including both creating backgrounds and rendering texts. Specifically, a background generator is developed to produce high-fidelity and text-free natural images. Moreover, a text renderer named GlyphOnly is designed for achieving visually plausible text-background integration. GlyphOnly, built on a Stable Diffusion framework, utilizes glyphs and backgrounds as conditions for accurate rendering and consistency control, as well as equipped with an adaptive text block exploration strategy for small-scale text rendering. We also explore several downstream applications based on our method, including scene text dataset synthesis for boosting scene text detectors, as well as text image customization and editing. Code and model will be available at https://github.com/Zhenhang-Li/GlyphOnly.
Zhenhang Li, Weichao Zeng, Dongbao Yang, Yu Zhou 0015
ECAI4
2024 Large Language Model for Action Anticipation
Dezhao Luo, Dongbao Yang
ICANN (3)3
2024 Accurate and Robust Scene Text Recognition via Adversarial Training
abstract
Adversarial training (AT) is a methodology that utilizes adversarial examples in the training process to enhance a model’s resistance to adversarial attacks and improve generalization. Despite its efficacy in several non-sequential computer vision tasks such as classification and object detection, its effects in the realm of Scene Text Recognition (STR) remain largely unexplored. This paper pioneers an investigation into the implications of AT on STR models and proposes a novel regularization-based AT method to develop an accurate and robust STR model, dynamically generating adversarial examples in the training procedure. Through extensive experiments across seven public real-world datasets, we find that AT not only bolsters the robustness of STR models but also improves overall recognition accuracy. This improvement is particularly significant in low-resolution images - a common challenge in STR. Furthermore, given the diverse nature of real-world text images, developing a robust STR model requires a large dataset. We propose viewing AT as a form of model-based data augmentation technique for STR, compatible with traditional augmentation methods. We hope these encouraging findings catalyze further research into the application of AT for scene text recognition.
Dongbao Yang, Yu Zhou 0015
ICASSP2
2024 Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs
abstract
Visual emotion recognition (VER) is a longstanding field that has garnered increasing attention with the advancement of deep neural networks. Although recent studies have achieved notable improvements by leveraging the knowledge embedded within pre-trained visual models, the lack of direct association between factual-level features and emotional categories, called the ''affective gap'', limits the applicability of pre-training knowledge for VER tasks. On the contrary, the explicit emotional expression and high information density in textual modality eliminate the ''affective gap''. Therefore, we propose borrowing the knowledge from the pre-trained textual model to enhance the emotional perception of pre-trained visual models. We focus on the factual and emotional connections between images and texts in noisy social media data, and propose Partitioned Adaptive Contrastive Learning (PACL) to leverage these connections. Specifically, we manage to separate different types of samples and devise distinct contrastive learning strategies for each type. By dynamically constructing negative and positive pairs, we fully exploit the potential of noisy samples. Through comprehensive experiments, we demonstrate that bridging the "affective gap'' significantly improves the performance of various pre-trained visual models in downstream emotion-related tasks. Our code is released on https://github.com/wdqqdw/PACL.
Daiqing Wu, Dongbao Yang, Yu Zhou 0015, Can Ma
ACM Multimedia2
2024 Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and Fusion
abstract
As posts on social media increase rapidly, analyzing the sentiments embedded in image-text pairs has become a popular research topic in recent years. Although existing works achieve impressive accomplishments in simultaneously harnessing image and text information, they lack the considerations of possible low-quality and missing modalities. In real-world applications, these issues might frequently occur, leading to urgent needs for models capable of predicting sentiment robustly. Therefore, we propose a Distribution-based feature Recovery and Fusion (DRF) method for robust multimodal sentiment analysis of image-text pairs. Specifically, we maintain a feature queue for each modality to approximate their feature distributions, through which we can simultaneously handle low-quality and missing modalities in a unified framework. For low-quality modalities, we reduce their contributions to the fusion by quantitatively estimating modality qualities based on the distributions. For missing modalities, we build inter-modal mapping relationships supervised by samples and distributions, thereby recovering the missing modalities from available ones. In experiments, two disruption strategies that corrupt and discard some modalities in samples are adopted to mimic the low-quality and missing modalities in various real-world scenarios. Through comprehensive experiments on three publicly available image-text datasets, we demonstrate the universal improvements of DRF compared to SOTA methods under both two strategies, validating its effectiveness in robust multimodal sentiment analysis.
Daiqing Wu, Dongbao Yang, Yu Zhou 0015, Can Ma
ACM Multimedia2
2024 Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
abstract
Scene text retrieval aims to find all images containing the query text from an image gallery. Current efforts tend to adopt an Optical Character Recognition (OCR) pipeline, which requires complicated text detection and/or recognition processes, resulting in inefficient and inflexible retrieval. Different from them, in this work we propose to explore the intrinsic potential of Contrastive Language-Image Pre-training (CLIP) for OCR-free scene text retrieval. Through empirical analysis, we observe that the main challenges of CLIP as a text retriever are: 1) limited text perceptual scale, and 2) entangled visual-semantic concepts. To this end, a novel model termed FDP (Focus, Distinguish, and Prompt) is developed. FDP first focuses on scene text via shifting the attention to the text area and probing the hidden text knowledge, and then divides the query text into content word and function word for processing, in which a semantic-aware prompting scheme and a distracted queries assistance module are utilized. Extensive experiments show that FDP significantly enhances the inference speed while achieving better or competitive retrieval accuracy compared to existing methods. Notably, on the IIIT-STR benchmark, FDP surpasses the state-of-the-art model by 4.37% with a 4 times faster speed. Furthermore, additional experiments under phrase-level and attribute-aware scene text retrieval settings validate FDP's particular advantages in handling diverse forms of query text. The source code will be available at https://github.com/Gyann-z/FDP.
Gangyan Zeng, Yuan Zhang 0013, Dongbao Yang, Peng Zhang 0044, Yiwen Gao 0001, Xugong Qin, Yu Zhou 0015
ACM Multimedia4
2024 TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
abstract
Centred on content modification and style preservation, Scene Text Editing (STE) remains a challenging task despite considerable progress in text-to-image synthesis and text-driven image manipulation recently. GAN-based STE methods generally encounter a common issue of model generalization, while Diffusion-based STE methods suffer from undesired style deviations. To address these problems, we propose TextCtrl, a diffusion-based method that edits text with prior guidance control. Our method consists of two key components: (i) By constructing fine-grained text style disentanglement and robust text glyph structure representation, TextCtrl explicitly incorporates Style-Structure guidance into model design and network training, significantly improving text style consistency and rendering accuracy. (ii) To further leverage the style prior, a Glyph-adaptive Mutual Self-attention mechanism is proposed which deconstructs the implicit fine-grained features of the source image to enhance style consistency and vision quality during inference. Furthermore, to fill the vacancy of the real-world STE evaluation benchmark, we create the first real-world image-pair dataset termed ScenePair for fair comparisons. Experiments demonstrate the effectiveness of TextCtrl compared with previous methods concerning both style fidelity and text accuracy. Project page: https://github.com/weichaozeng/TextCtrl.
Weichao Zeng, Zhenhang Li, Dongbao Yang, Yu Zhou 0015
NeurIPS4
2024 Masked and Permuted Implicit Context Learning for Scene Text Recognition
abstract
Scene Text Recognition (STR) is challenging because of various text styles, shapes, and backgrounds. Although the integration of linguistic information enhances models' performance, existing methods based on either permuted language modeling (PLM) or masked language modeling (MLM) have their drawbacks. PLM's autoregressive decoding lacks foresight into subsequent characters, while MLM overlooks intercharacter dependencies. To address these problems, we propose a masked and permuted implicit context learning network for STR, which unifies PLM and MLM within a single decoder, inheriting the advantages of both approaches. We utilize the training procedure of PLM and incorporate word length information into the decoding process to integrate MLM, substituting the undetermined characters with mask tokens. Besides, we employ the perturbation training technique to train a more robust model against potential length prediction errors. Our comprehensive evaluations demonstrate the performance of our model. It achieves superior performance on the popularly used benchmarks and outperforms previous state-of-the-art methods with a substantial improvement of 9.1% on the more challenging Union14M-Benchmark.
Dongbao Yang, Yu Zhou 0015
IEEE Signal Process. Lett.4
2023 One-Shot Replay: Boosting Incremental Object Detection via Retrospecting One Object
abstract
Modern object detectors are ill-equipped to incrementally learn new emerging object classes over time due to the well-known phenomenon of catastrophic forgetting. Due to data privacy or limited storage, few or no images of the old data can be stored for replay. In this paper, we design a novel One-Shot Replay (OSR) method for incremental object detection, which is an augmentation-based method. Rather than storing original images, only one object-level sample for each old class is stored to reduce memory usage significantly, and we find that copy-paste is a harmonious way to replay for incremental object detection. In the incremental learning procedure, diverse augmented samples with co-occurrence of old and new objects to existing training data are generated. To introduce more variants for objects of old classes, we propose two augmentation modules. The object augmentation module aims to enhance the ability of the detector to perceive potential unknown objects. The feature augmentation module explores the relations between old and new classes and augments the feature space via analogy. Extensive experimental results on VOC2007 and COCO demonstrate that OSR can outperform the state-of-the-art incremental object detection methods without using extra wild data.
Dongbao Yang, Yu Zhou 0015, Xiaopeng Hong, Aoting Zhang, Weiping Wang 0005
AAAI1
2023 Mask-Guided Stamp Erasure for Real Document Image
abstract
The application of text recognition in the automatic analysis of invoices, contracts and other documents has significantly raised office efficiency, but the stamps overlapping with the texts in these documents may seriously degrade the recognition accuracy. To mitigate the negative effect, we propose a stamp eraser, which can simultaneously remove the stamps and recover the occluded texts. To better distinguish stamps from the complex background, we propose a stamp localization module to generate fine-grained binary masks for the eraser to focus more on stamps. This module can also provide the background textual information for recovering text with a skip connection. We also propose the dilated mask to make the generated image look more natural by filling the stamp area with pixels around the stamp. In addition, to evaluate the effectiveness of our method in boosting text recognition, we propose a synthetic data set for training and a real dataset with complex background for testing. Experiments have shown that our method can effectively improve the recognition accuracy of the text occluded with the stamps.
Dongbao Yang, Yu Zhou 0015, Youhui Guo, Weiping Wang 0005
ICME2
2023 Cross-Architecture Relational Consistency for Point Cloud Self-Supervised Learning
abstract
In this paper, we present a novel Cross-architecture Relational Consistency (CRC) framework for point cloud self-supervised learning. Since the features produced by encoders of different architectures contain different information in point cloud processing, we design a cross-architecture framework to obtain complementary feature representations, which utilizes two parallel encoders, i.e., PointNet ++(point-based) and SparseConv (voxel-based). To better preserve the uniqueness of feature spaces of different encoders, the CRC framework learns the cross-architecture consistency of inter-instance relations in the feature space of two encoders through distribution alignment, which aligns the similarity distributions calculated from the two spaces using the KL divergence objective. The CRC framework not only learns complementary information from different encoders but also obtains two pre-trained models that can be used separately for downstream tasks. The distribution alignment objective enables the CRC framework to capture more inter-sample correlation information, thereby better preserving meaningful semantic structures. We validate the effectiveness of the CRC framework on the KITTI 3D object detection task, where it outperforms the previous self-supervised learning methods. Further, the ablation studies validate the potency of our approach for a better point cloud understanding. Code: https://github.com/Physu/CRCSSL
Yifei Zhang 0005, Dongbao Yang
ICTAI3
2023 Perceiving Ambiguity and Semantics without Recognition: An Efficient and Effective Ambiguous Scene Text Detector
abstract
Ambiguous scene text detection is an extremely challenging task. Existing text detectors that rely solely on visual cues often suffer from confusion due to being evenly distributed in rows/columns or incomplete detection owing to large character spacing. To overcome these challenges, the previous method recognizes a large number of proposals and utilizes semantic information predicted from recognition results to eliminate ambiguity. However, this method is inefficient, which limits their practical applications. In this paper, we propose a novel efficient and effective ambiguous text detector, which can Perceive Ambiguity and SEmantics without Recognition, termed PASER. On the one hand, PASER can perceive semantics without recognition with a light Perceiving Semantics (PerSem) module. In this way, proposals without reasonable semantics are filtered out, which largely speeds up the overall detection process. On the other hand, to detect both ambiguous and regular texts with a unified framework, PASER employs a Perceiving Ambiguity (PerAmb) module to distinguish ambiguous texts and regular texts, so that only the ambiguous proposals will be processed by PerSem while the regular texts are not, which further ensures the high efficiency. Extensive experiments show that our detector achieves state-of-the-art results on both ambiguous and regular scene text detection benchmarks. Notably, over 6 times faster speed and superior accuracy are achieved on TDA-ReCTS simultaneously.
Wei Wang 0315, Yu Zhou 0015, Shaohui Liu, Aoting Zhang, Dongbao Yang, Weiping Wang 0005
ACM Multimedia6
2023 Pseudo Object Replay and Mining for Incremental Object Detection
abstract
Incremental object detection (IOD) aims to mitigate catastrophic forgetting for object detectors when incrementally learning to detect new emerging object classes without using original training data. Most existing IOD methods benefit from the assumption that unlabeled old-class objects may co-occur with labeled new-class objects in the new training data. However, in practical scenarios, old-class objects may be absent, which is called non co-occurrence IOD. In this paper, we propose a pseudo object replay and mining method (PseudoRM) to handle the co-occurrence dependent problem, reducing the performance degradation caused by the absence of old-class objects. The new training data can be augmented by co-occurring fake (old-class) and real (new-class) objects with a patch-level data-free generation method in the pseudo object replay stage. To fully use existing training data, we propose pseudo object mining to explore false positives for transferring useful instance-level knowledge. In the incremental learning procedure, a generative distillation is introduced to distill image-level knowledge for balancing stability and plasticity. Experimental results on PASCAL VOC and COCO demonstrate that PseudoRM can effectively boost the performance on both co-occurrence and non co-occurrence scenarios without using old samples or extra wild data.
Dongbao Yang, Yu Zhou 0015, Xiaopeng Hong, Aoting Zhang, Linchengxi Zeng, Weiping Wang 0005
ACM Multimedia1
2022 Multi-View correlation distillation for incremental object detection
Dongbao Yang, Yu Zhou 0015, Aoting Zhang, Xurui Sun, Dayan Wu, Weiping Wang 0005, Qixiang Ye
Pattern Recognit.1
2022 RD-IOD: Two-Level Residual-Distillation-Based Triple-Network for Incremental Object Detection
abstract
As a basic component in multimedia applications, object detectors are generally trained on a fixed set of classes that are pre-defined. However, new object classes often emerge after the models are trained in practice. Modern object detectors based on Convolutional Neural Networks (CNN) suffer from catastrophic forgetting when fine-tuning on new classes without the original training data. Therefore, it is critical to improve the incremental learning capability on object detection. In this article, we propose a novel Residual-Distillation-based Incremental learning method on Object Detection (RD-IOD). Our approach rests on the creation of a triple-network based on Faster R-CNN. To enable continuous learning from new classes, we use the original model as well as a residual model to guide the learning of the incremental model on new classes while maintaining the previous learned knowledge. To better maintain the discrimination between the features of old and new classes, the residual model is jointly trained with the incremental model on new classes in the incremental learning procedure. In addition, a two-level distillation scheme is designed to guide the training process, which consists of (1) a general distillation for imitating the original model in feature space along with a residual distillation on the features in both image level and instance level, and (2) a joint classification distillation on the output layers. To well preserve the learned knowledge, we design a 2-threshold training strategy to guide the learning of a Region Proposal Network and a detection head. Extensive experiments conducted on VOC2007 and COCO demonstrate that the proposed method can effectively learn to incrementally detect objects of new classes, and the problem of catastrophic forgetting is mitigated. Our code is available at https://github.com/yangdb/RD-IOD.
Dongbao Yang, Yu Zhou 0015, Wei Shi 0001, Dayan Wu, Weiping Wang 0005
ACM Trans. Multim. Comput. Commun. Appl.1
2020 Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning
abstract
We propose a novel self-supervised method, referred to as Video Cloze Procedure (VCP), to learn rich spatial-temporal representations. VCP first generates “blanks” by withholding video clips and then creates “options” by applying spatio-temporal operations on the withheld clips. Finally, it fills the blanks with “options” and learns representations by predicting the categories of operations applied on the clips. VCP can act as either a proxy task or a target task in self-supervised learning. As a proxy task, it converts rich self-supervised representations into video clip operations (options), which enhances the flexibility and reduces the complexity of representation learning. As a target task, it can assess learned representation models in a uniform and interpretable manner. With VCP, we train spatial-temporal representation models (3D-CNNs) and apply such models on action recognition and video retrieval tasks. Experiments on commonly used benchmarks show that the trained models outperform the state-of-the-art self-supervised models with significant margins.
Dezhao Luo, Chang Liu 0042, Yu Zhou 0015, Dongbao Yang, Can Ma, Qixiang Ye, Weiping Wang 0005
AAAI4
2020 SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition
abstract
Scene text recognition is a hot research topic in computer vision. Recently, many recognition methods based on the encoder-decoder framework have been proposed, and they can handle scene texts of perspective distortion and curve shape. Nevertheless, they still face lots of challenges like image blur, uneven illumination, and incomplete characters. We argue that most encoder-decoder methods are based on local visual features without explicit global semantic information. In this work, we propose a semantics enhanced encoder-decoder framework to robustly recognize low-quality scene texts. The semantic information is used both in the encoder module for supervision and in the decoder module for initializing. In particular, the state-of-the-art ASTER method is integrated into the proposed framework as an exemplar. Extensive experiments demonstrate that the proposed framework is more robust for low-quality text images, and achieves state-of-the-art results on several benchmark datasets. The source code will be available.
Yu Zhou 0015, Dongbao Yang, Yucan Zhou, Weiping Wang 0005
CVPR3
2020 Self-Training for Domain Adaptive Scene Text Detection
abstract
Though deep learning based scene text detection has achieved great progress, well-trained detectors suffer from severe performance degradation for different domains. In general, a tremendous amount of data is indispensable to train the detector in the target domain. However, data collection and annotation are expensive and time-consuming. To address this problem, we propose a self-training framework to automatically mine hard examples with pseudo-labels from unannotated videos or images. To reduce the noise of hard examples, a novel text mining module is implemented based on the fusion of detection and tracking results. Then, an image-to-video generation method is designed for the tasks that videos are unavailable and only images can be used. Experimental results on standard benchmarks, including ICDAR2015, MSRA-TD500, ICDAR2017 MLT, demonstrate the effectiveness of our self-training method. The simple Mask R-CNN adapted with self-training and fine-tuned on real data can achieve comparable or even superior results with the state-of-the-art methods.
Yudi Chen, Wei Wang 0315, Yu Zhou 0015, Dongbao Yang, Weiping Wang 0005
ICPR5
2019 Curved Text Detection in Natural Scene Images with Semi- and Weakly-Supervised Learning
abstract
Detecting curved text in the wild is very challenging. Recently, most state-of-the-art methods are segmentation based and require pixel-level annotations. We propose a novel scheme to train an accurate text detector using only a small amount of pixel-level annotated data and a large amount of data annotated with rectangles or even unlabeled data. A light model is first obtained by training with the pixel-level annotated data and then used to annotate unlabeled or weakly labeled data. A novel strategy which utilizes ground-truth bounding boxes to generate pseudo mask annotations is proposed in weakly-supervised learning. Experimental results on CTW1500 and Total-Text demonstrate that our method can substantially reduce the requirement of pixel-level annotated data. Our method can also generalize well across the two datasets. The performance of the proposed method is comparable with the state-of-the-art methods with only 10% pixel-level annotated data and 90% rectangle-level weakly annotated data.
Xugong Qin, Yu Zhou 0015, Dongbao Yang, Weiping Wang 0005
ICDAR3
2019 Constrained Relation Network for Character Detection in Scene Images
Yudi Chen, Yu Zhou 0015, Dongbao Yang, Weiping Wang 0005
PRICAI (3)3
2019 Supervised deep hashing for image content security
Yanping Ma, Dongbao Yang, Hongtao Xie 0001, Jian Yin 0003
Multim. Tools Appl.2
2019 Automated pulmonary nodule detection in CT images using deep convolutional neural networks
Hongtao Xie 0001, Dongbao Yang, Nannan Sun, Zhineng Chen, Yongdong Zhang 0001
Pattern Recognit.2
2018 Deep Convolutional Nets for Pulmonary Nodule Detection and Classification
Nannan Sun, Dongbao Yang, Shancheng Fang, Hongtao Xie 0001
KSEM (2)2
2018 Supervised Hash Coding With Deep Neural Network for Environment Perception of Intelligent Vehicles
abstract
Image content analysis is an important surround perception modality of intelligent vehicles. In order to efficiently recognize the on-road environment based on image content analysis from the large-scale scene database, relevant images retrieval becomes one of the fundamental problems. To improve the efficiency of calculating similarities between images, hashing techniques have received increasing attentions. For most existing hash methods, the suboptimal binary codes are generated, as the hand-crafted feature representation is not optimally compatible with the binary codes. In this paper, a one-stage supervised deep hashing framework (SDHP) is proposed to learn high-quality binary codes. A deep convolutional neural network is implemented, and we enforce the learned codes to meet the following criterions: 1) similar images should be encoded into similar binary codes, and vice versa; 2) the quantization loss from Euclidean space to Hamming space should be minimized; and 3) the learned codes should be evenly distributed. The method is further extended into SDHP+ to improve the discriminative power of binary codes. Extensive experimental comparisons with state-of-the-art hashing algorithms are conducted on CIFAR-10 and NUS-WIDE, the MAP of SDHP reaches to 87.67% and 77.48% with 48 b, respectively, and the MAP of SDHP+ reaches to 91.16%, 81.08% with 12 b, 48 b on CIFAR-10 and NUS-WIDE, respectively. It illustrates that the proposed method can obviously improve the search accuracy.
Chenggang Yan 0001, Hongtao Xie 0001, Dongbao Yang, Jian Yin 0003, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Intell. Transp. Syst.3