Yu Zhou 0015

dblp:36/2728-15 · DBLP profile ↗
← Back
94ranked-venue papers
6as first author
70since 2021 · last 2026
0000-0003-4188-9953ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 66 · 3 first-author · 49 since 2021Artificial intelligence and machine learning · 49 · 2 first-author · 38 since 2021Databases, data management, data science and information retrieval · 8 · 6 since 2021Computer networks · 4 · 3 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement
abstract
Understanding 3D scene-level affordances from natural language instructions is essential for enabling embodied agents to interact meaningfully in complex environments. However, this task remains challenging due to the need for semantic reasoning and spatial grounding. Existing methods mainly focus on object-level affordances or merely lift 2D predictions to 3D, neglecting rich geometric structure information in point clouds and incurring high computational costs. To address these limitations, we introduce Task-Aware 3D Scene-level Affordance segmentation (TASA), a novel geometry-optimized framework that jointly leverages 2D semantic cues and 3D geometric reasoning in a coarse-to-fine manner. To improve the affordance detection efficiency, TASA features a task-aware 2D affordance detection module to identify manipulable points from language and visual inputs, guiding the selection of task-relevant views. To fully exploit 3D geometric information, a 3D affordance refinement module is proposed to integrate 2D semantic priors with local 3D geometry, resulting in accurate and spatially coherent 3D affordance masks. Experiments on SceneFun3D demonstrate that TASA significantly outperforms the baselines in both accuracy and efficiency in scene-level affordance segmentation.
Qilang Ye, Yu Zhou 0015
AAAI4
2026 ST-SAM: Multimodal Scene Text Segmentation with Dense Visual and Sparse Textual Prompts via SAM
abstract
Scene text segmentation is a critical preprocessing step in various text-based applications. Specialist text segmentation methods, often relying on a detect-then-segment paradigm, tend to exhibit reduced robustness and can lead to cascading errors. The introduction of the Segment Anything Model (SAM) has revolutionized general segmentation by leveraging vision foundation models. However, SAM still falls short when applied to domain-specific tasks such as scene text segmentation. To bridge this gap between SAM and specialized scene text segmentation approaches, we propose ST-SAM (Scene Text SAM), a parameter-efficient fine-tuning framework tailored to adapt SAM for high-quality scene text segmentation without relying on explicit text detection. ST-SAM incorporates a multimodal prompting mechanism: a lightweight visual encoder generates multi-scale spatial features to provide precise visual context; and textual prompts generated by a large language model offer high-level semantic guidance. We demonstrate the advantages of the proposed ST-SAM as follows: (1) ST-SAM achieves new state-of-the-art performance on multiple scene text segmentation benchmarks, including 85.30% fgIoU on Total-Text and 91.03% fgIoU on TextSeg, outperforming both specialist and generalist models. (2) ST-SAM enables effective domain adaptation by flexibly adapting the general SAM architecture to the domain of scene text. (3) By discarding the detect-then-segment pipeline, ST-SAM simplifies the inference process while still achieving robust performance on complex text cases.
Yaqiang Wu, Jiayi Yan, Yu Zhou 0015, Lingling Zhang 0005, Qianying Wang 0002
AAAI6
2026 SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
abstract
Large Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action classification and description. However, when treating LLM as a recognizer, two questions arise: 1) How can LLMs understand the skeleton? 2) How can LLMs distinguish among actions? To address these problems, we introduce a novel paradigm named learning Skeleton representation with visual-motion knowledge for Action Recognition (SUGAR). In our pipeline, we first utilize off-the-shelf large-scale video models as a knowledge base to generate visual, motion information related to actions. Then, we propose to supervise skeleton learning through this prior knowledge to yield discrete representations. Finally, we use the LLM with untouched pre-training weights to understand these representations and generate the desired action targets and descriptions. Notably, we present a Temporal Query Projection (TQP) module to continuously model the skeleton signals with long sequences. Experiments on several skeleton-based action classification benchmarks demonstrate the efficacy of our SUGAR. Moreover, experiments on zero-shot scenarios show that SUGAR is more versatile than linear-based methods.
Qilang Ye, Yu Zhou 0015, Jie Zhang 0081, Xuanming Guo, Mingkui Tan, Weicheng Xie 0001, Yue Sun 0001, Tao Tan 0002, Xiaochen Yuan, Ghada Khoriba, Zitong Yu
AAAI2
2026 When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?
abstract
Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, which simulates an “Audio-Visual Confusion” scene by modifying the corresponding sound of an object in the video, e.g., mute the sounding object and ask MLLMs “Is there a/an {muted-object} sound”. Experimental results reveal that MLLMs, such as Qwen2.5-Omni and Gemini 2.5, struggle to discriminate non-existent audio due to visually dominated reasoning. Motivated by this observation, we introduce RL-CoMM, a Reinforcement Learning-based Collaborative Multi-MLLM that is built upon the Qwen2.5-Omni foundation. RL-CoMM includes two stages: 1) To alleviate visually dominated ambiguities, we introduce an external model, a Large Audio Language Model (LALM), as the reference model to generate audio-only reasoning. Then, we design a Step-wise Reasoning Reward function that enables MLLMs to self-improve audio-visual reasoning with the audio-only reference. 2) To ensure an accurate answer prediction, we introduce Answer-centered Confidence Optimization to reduce the uncertainty of potential heterogeneous reasoning differences. Extensive experiments on audio-visual question answering and audio-visual hallucination show that RL-CoMM improves accuracy by 10~30% over the baseline model with limited training data.
Qilang Ye, Jie Zhang 0081, Zitong Yu, Yu Zhou 0015
AAAI7
2026 Towards Breaking the Visual Perception Bottleneck for Geometry Problem Solving
Tianjiao Cao, Jiahao Lyu 0002, Dongbao Yang, Weimin Mu, Yu Zhou 0015
ICDAR (2)5
2026 ComMark: Covert and Robust Black-Box Model Watermarking with Compressed Samples
abstract
The rapid advancement of deep learning has turned models into highly valuable assets due to their reliance on massive data and costly training processes. However, these models are increasingly vulnerable to leakage and theft, highlighting the critical need for robust intellectual property protection. Model watermarking has emerged as an effective solution, with black-box watermarking gaining significant attention for its practicality and flexibility. Nonetheless, existing black-box methods often fail to better balance covertness (hiding the watermark to prevent detection and forgery) and robustness (ensuring the watermark resists removal)—two essential properties for real-world copyright verification. In this paper, we propose ComMark, a novel black-box model watermarking framework that leverages frequency-domain transformations to generate compressed, covert, and attack-resistant watermark samples by filtering out high-frequency information. To further enhance watermark robustness, our method incorporates simulated attack scenarios and a similarity loss during training. Comprehensive evaluations across diverse datasets and architectures demonstrate that ComMark achieves state-of-the-art performance in both covertness and robustness.
Yunfei Yang 0001, Xiaojun Chen 0004, Zhendong Zhao, Yu Zhou 0015, Xiaoyan Gu 0001, Juan Cao 0001
ICMR4
2026 EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration
Daiqing Wu, Dongbao Yang, Can Ma, Yu Zhou 0015
Pattern Recognit.4
2026 Resolving sentiment discrepancy for multimodal sentiment detection via semantics completion and decomposition
Daiqing Wu, Dongbao Yang, Huawen Shen, Can Ma, Yu Zhou 0015
Pattern Recognit.5
2025 Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
abstract
Scene text spotting has attracted the enthusiasm of relative researchers in recent years. Most existing scene text spotters follow the detection-then-recognition paradigm, where the vanilla detection module hardly determines the reading order and leads to failure recognition. After rethinking the auto-regressive scene text recognition method, we find that a well-trained recognizer can implicitly perceive the local semantics of all characters in a complete word or a sentence without a character-level detection module. Local semantic knowledge not only includes text content but also spatial information in the right reading order. Motivated by the above analysis, we propose the Local Semantics Guided scene text Spotter (LSGSpotter), which auto-regressively decodes the position and content of characters guided by the local semantics. Specifically, two effective modules are proposed in LSGSpotter. On the one hand, we design a Start Point Localization Module (SPLM) for locating text start points to determine the right reading order. On the other hand, a Multi-scale Adaptive Attention Module (MAAM) is proposed to adaptively aggregate text features in a local area. In conclusion, LSGSpotter achieves the arbitrary reading order spotting task without the limitation of sophisticated detection, while alleviating the cost of computational resources with the grid sampling strategy. Extensive experiment results show LSGSpotter achieves state-of-the-art performance on the InverseText benchmark. Moreover, our spotter demonstrates superior performance on English benchmarks for arbitrary-shaped text, achieving improvements of 0.7% and 2.5% on Total-Text and SCUT-CTW1500, respectively. These results validate our text spotter is effective for scene texts in arbitrary reading order and shape.
Jiahao Lyu 0002, Wei Wang 0315, Dongbao Yang, Jinwen Zhong, Yu Zhou 0015
AAAI5
2025 LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
abstract
Visual Information Extraction (VIE) plays a crucial role in the comprehension of semi-structured documents, and several pre-trained models have been developed to enhance performance. However, most of these works are monolingual (usually English). Due to the extremely unbalanced quantity and quality of pre-training corpora between English and other languages, few works can extend to non-English scenarios. In this paper, we conduct systematic experiments to show that vision and layout modality hold invariance among images with different languages. If decoupling language bias from document images, a vision-layout-based model can achieve impressive cross-lingual generalization. Accordingly, we present a simple but effective multilingual training paradigm LDP (Language Decoupled Pre-training) for better utilization of monolingual pre-training data. Our proposed model LDM (Language Decoupled Model) is first pre-trained on the language-independent data, where the language knowledge is decoupled by a diffusion model, and then the LDM is fine-tuned on the downstream languages. Extensive experiments show that the LDM outperformed all SOTA multilingual pre-trained models, and also maintains competitiveness on downstream monolingual/English benchmarks.
Huawen Shen, Gengluo Li, Jinwen Zhong, Yu Zhou 0015
AAAI4
2025 Specifying What You Know or Not for Multi-Label Class-Incremental Learning
abstract
Existing class incremental learning is mainly designed for single-label classification task, which is ill-equipped for multi-label scenarios due to the inherent contradiction of learning objectives for samples with incomplete labels. We argue that the main challenge to overcome this contradiction in multi-label class-incremental learning (MLCIL) lies in the model's inability to clearly distinguish between known and unknown knowledge. This ambiguity hinders the model's ability to retain historical knowledge, master current classes, and prepare for future learning simultaneously. In this paper, we target at specifying what is known or not to accommodate Historical, Current, and Prospective knowledge for MLCIL and propose a novel framework termed as HCP. Specifically, (i) we clarify the known classes by dynamic feature purification and recall enhancement with distribution prior, enhancing the precision and retention of known information. (ii) We design prospective knowledge mining to probe the unknown, preparing the model for future learning. Extensive experiments validate that our method effectively alleviates catastrophic forgetting in MLCIL, surpassing the previous state-of-the-art by 3.3% on average accuracy for MS-COCO B0-C10 setting without replay buffers.
Aoting Zhang, Dongbao Yang, Xiaopeng Hong, Yu Zhou 0015
AAAI5
2025 DCA: Dividing and Conquering Amnesia in Incremental Object Detection
abstract
Incremental object detection (IOD) aims to cultivate an object detector that can continuously localize and recognize novel classes while preserving its performance on previous classes. Existing methods achieve certain success by improving knowledge distillation and exemplar replay for transformer-based detection frameworks, but the intrinsic forgetting mechanisms remain underexplored. In this paper, we dive into the cause of forgetting and discover forgetting imbalance between localization and recognition in transformer-based IOD, which means that localization is less-forgetting and can generalize to future classes, whereas catastrophic forgetting occurs primarily on recognition. Based on these insights, we propose a Divide-and-Conquer Amnesia (DCA) strategy, which redesigns the transformer-based IOD into a localization-then-recognition process. DCA can well maintain and transfer the localization ability, leaving decoupled fragile recognition to be specially conquered. To reduce feature drift in recognition, we leverage semantic knowledge encoded in pre-trained language models to anchor class representations within a unified feature space across incremental tasks. This involves designing a duplex classifier fusion and embedding class semantic features into the recognition decoding process in the form of queries. Extensive experiments validate that our approach achieves state-of-the-art performance, especially for long-term incremental scenarios. For example, under the four-step setting on MS-COCO, our DCA strategy significantly improves the final AP by 6.9%.
Aoting Zhang, Dongbao Yang, Xiaopeng Hong, Miao Shang, Yu Zhou 0015
AAAI6
2025 Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
abstract
Video text-based visual question answering (TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. Inspired by the development of TextVQA in image domain, existing Video TextVQA approaches leverage a language model (e.g. T5) to process text-rich multiple frames and generate answers auto-regressively. Nevertheless, the spatio-temporal relationships among visual entities (including scene text and objects) will be disrupted and models are susceptible to interference from unrelated information, resulting in irrational reasoning and inaccurate answering. To tackle these challenges, we propose the TEA (stands for "Track the Answer'') method that better extends the generative TextVQA framework from image to video. TEA recovers the spatio-temporal relationships in a complementary way and incorporates OCR-aware clues to enhance the quality of reasoning questions. Extensive experiments on several public Video TextVQA datasets validate the effectiveness and generalization of our framework. TEA outperforms existing TextVQA methods, video-language pretraining methods and video large language models by great margins. The code will be publicly released.
Gangyan Zeng, Huawen Shen, Daiqing Wu, Yu Zhou 0015, Can Ma
AAAI5
2025 Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
abstract
Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates contextual and semantic elements. In scenarios with degraded visual quality, linguistic patterns serve as crucial supplements for comprehension, highlighting the necessity of integrating both aspects for robust scene text recognition (STR). Contemporary STR approaches often use language models or semantic reasoning modules to capture linguistic features, typically requiring large-scale annotated datasets. Self-supervised learning, which lacks annotations, presents challenges in disentangling linguistic features related to the global context. Typically, sequence contrastive learning emphasizes the alignment of local features, while masked image modeling (MIM) tends to exploit local structures to reconstruct visual patterns, resulting in limited linguistic knowledge. In this paper, we propose a Linguistics-aware Masked Image Modeling (LMIM) approach, which channels the linguistic information into the decoding process of MIM through a separate branch. Specifically, we design a linguistics alignment module to extract vision-independent features as linguistic guidance using inputs with different visual appearances. As features extend beyond mere visual structures, LMIM must consider the global context to achieve reconstruction. Extensive experiments on various benchmarks quantitatively demonstrate our state-of-the-art performance, and attention visualizations qualitatively show the simultaneous capture of both visual and linguistic information. The code is available at https: //github.com/zhangyifei01/LMIM.
Yifei Zhang 0005, Yu Zhou 0015, Can Ma, Xiangyang Ji
CVPR5
2025 TADoc: Robust Time-Aware Document Image Dewarping
abstract
Flattening curved, wrinkled, and rotated document images captured by portable photographing devices, termed document image dewarping, has become an increasingly important task with the rise of digital economy and online working. Although many methods have been proposed recently, they often struggle to achieve satisfactory results when confronted with intricate document structures and higher degrees of deformation in real-world scenarios. Our main insight is that, unlike other document restoration tasks (e.g., deblurring), dewarping in real physical scenes is a progressive motion rather than a one-step transformation. Based on this, we have undertaken two key initiatives. Firstly, we reformulate this task, modeling it for the first time as a dynamic process that encompasses a series of intermediate states. Secondly, we design a lightweight framework called TADoc (Time-Aware Document Dewarping Network) to address the geometric distortion of document images. In addition, due to the inadequacy of OCR metrics for document images containing sparse text, the comprehensiveness of evaluation is insufficient. To address this shortcoming, we propose a new metric – DLS (Document Layout Similarity) – to evaluate the effectiveness of document dewarping in downstream tasks. Extensive experiments and in-depth evaluations have been conducted and the results indicate that our model possesses strong robustness, achieving superiority on several benchmarks with different document types and degrees of distortion.
Fangmin Zhao, Weichao Zeng, Zhenhang Li, Dongbao Yang, Yu Zhou 0015
ECAI5
2025 Char-SAM: Turning Segment Anything Model into Scene Text Segmentation Annotator with Character-level Visual Prompts
abstract
The recent emergence of the Segment Anything Model (SAM) enables various domain-specific segmentation tasks to be tackled cost-effectively by using bounding boxes as prompts. However, in scene text segmentation, SAM can not achieve desirable performance. The word-level bounding box as prompts is too coarse for characters, while the character-level bounding box as prompts suffers from over-segmentation and under-segmentation issues. In this paper, we propose an automatic annotation pipeline named Char-SAM, that turns SAM into a low-cost segmentation annotator with a Character-level visual prompt. Specifically, leveraging some existing text detection datasets with word-level bounding box annotations, we first generate finer-grained character-level bounding box prompts using the Character Bounding-box Refinement (CBR) module. Next, we employ glyph information corresponding to text character categories as a new prompt in the Character Glyph Refinement (CGR) module to guide SAM in producing more accurate segmentation masks, addressing issues of over-segmentation and under-segmentation. These modules fully utilize the bbox-to-mask capability of SAM to generate high-quality text segmentation annotations automatically. Extensive experiments on TextSeg validate the effectiveness of Char-SAM. Its training-free nature also enables the generation of high-quality scene text segmentation datasets from real-world datasets like COCO-Text and MLT17.
Enze Xie, Jiahao Lyu 0002, Daiqing Wu, Huawen Shen, Yu Zhou 0015
ICASSP5
2025 PACM: Position-Aware Cross-Modality Decoder for Handwritten Mathematical Expression Recognition
Zhijie Shen, Can Ma, Yaqiang Wu, Yu Zhou 0015
ICDAR (1)6
2025 PerturbCTC: Improving Alignment in Scene Text Recognition with Feature Perturbation Based CTC
Zhijie Shen, Yaqiang Wu, Gangyan Zeng, Dongbao Yang, Yu Zhou 0015
ICDAR (4)8
2025 Class-Agnostic Region-of-Interest Matching in Document Images
Demin Zhang, Jiahao Lyu 0002, Zhijie Shen, Yu Zhou 0015
ICDAR (4)4
2025 Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts
abstract
Chinese scene text retrieval is a practical task that aims to search for images containing visual instances of a Chinese query text. This task is extremely challenging because Chinese text often features complex and diverse layouts in real-world scenes. Current efforts tend to inherit the solution for English scene text retrieval, failing to achieve satisfactory performance. In this paper, we establish a Diversified Layout benchmark for Chinese Street View Text Retrieval (DL-CSVTR), which is specifically designed to evaluate retrieval performance across various text layouts, including vertical, cross-line, and partial alignments. To address the limitations in existing methods, we propose Chinese Scene Text Retrieval CLIP (CSTR-CLIP), a novel model that integrates global visual information with multi-granularity alignment training. CSTR-CLIP applies a two-stage training process to overcome previous limitations, such as the exclusion of visual features outside the text region and reliance on single-granularity alignment, thereby enabling the model to effectively handle diverse text layouts. Experiments on existing benchmark show that CSTR-CLIP outperforms the previous state-of-the-art model by 18.82% accuracy and also provides faster inference speed. Further analysis on DL-CSVTR confirms the superior performance of CSTR-CLIP in handling various text layouts. The dataset and code will be publicly available to facilitate research in Chinese scene text retrieval.
Gengluo Li, Huawen Shen, Yu Zhou 0015
ICML3
2025 An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
abstract
The advancements in Multimodal Large Language Models (MLLMs) have enabled various multimodal tasks to be addressed under a zero-shot paradigm. This paradigm sidesteps the cost of model fine-tuning, emerging as a dominant trend in practical application. Nevertheless, Multimodal Sentiment Analysis (MSA), a pivotal challenge in the quest for general artificial intelligence, fails to accommodate this convenience. The zero-shot paradigm exhibits undesirable performance on MSA, casting doubt on whether MLLMs can perceive sentiments as competent as supervised models. By extending the zero-shot paradigm to In-Context Learning (ICL) and conducting an in-depth study on configuring demonstrations, we validate that MLLMs indeed possess such capability. Specifically, three key factors that cover demonstrations' retrieval, presentation, and distribution are comprehensively investigated and optimized. A sentimental predictive bias inherent in MLLMs is also discovered and later effectively counteracted. By complementing each other, the devised strategies for three factors result in average accuracy improvements of 15.9% on six MSA datasets against the zero-shot paradigm and 11.2% against the random ICL baseline.
Daiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma, Yu Zhou 0015
ICML5
2025 The Devil is in Fine-tuning and Long-tailed Problems: A New Benchmark for Scene Text Detection
abstract
Scene text detection has seen the emergence of high-performing methods that excel on academic benchmarks. However, these detectors often fail to replicate such success in real-world scenarios. We uncover two key factors contributing to this discrepancy through extensive experiments. First, a Fine-tuning Gap, where models leverage Dataset-Specific Optimization (DSO) paradigm for one domain at the cost of reduced effectiveness in others, leads to inflated performances on academic benchmarks. Second, the suboptimal performance in practical settings is primarily attributed to the longtailed distribution of texts, where detectors struggle with rare and complex categories as artistic or overlapped text. Given that the DSO paradigm might undermine the generalization ability of models, we advocate for a Joint-Dataset Learning (JDL) protocol to alleviate the Fine-tuning Gap. Additionally, an error analysis is conducted to identify three major categories and 13 subcategories of challenges in long-tailed scene text, upon which we propose a Long-Tailed Benchmark (LTB). LTB facilitates a comprehensive evaluation of ability to handle a diverse range of long-tailed challenges. We further introduce MAEDet, a self-supervised learningbased method, as a strong baseline for LTB. The code is available at https://github.com/pd162/LTB.
Tianjiao Cao, Jiahao Lyu 0002, Weichao Zeng, Weimin Mu, Yu Zhou 0015
IJCAI5
2025 The Role of Video Generation in Enhancing Data-Limited Action Understanding
abstract
Video action understanding tasks in real-world scenarios often suffer from data limitations. In this paper, we address the data-limited action understanding problem by bridging data scarcity. We propose a novel method that leverages a text-to-video diffusion transformer to generate annotated data for model training. This paradigm enables the generation of realistic annotated data on an infinite scale without human intervention. We proposed the Information Enhancement Strategy and the Uncertainty-Based Soft Target tailored to generate sample training. Through quantitative and qualitative analyzes, we discovered that real samples generally contain a richer level of information compared to generated samples. Based on this observation, the information enhancement strategy was designed to enhance the informational content of the generated samples from two perspectives: the environment and the character. Furthermore, we observed that a portion of low-quality generated samples might negatively affect model training. To address this, we devised an uncertainty-based label-smoothing strategy to increase the smoothing of these low-quality samples, thereby reducing their impact. We demonstrate the effectiveness of the proposed method on four datasets and five tasks, and achieve state-of-the-art performance for zero-shot action recognition.
Dezhao Luo, Dongbao Yang, Zhenhang Li, Weiping Wang 0005, Yu Zhou 0015
IJCAI6
2025 Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
abstract
Video text-based visual question answering (Video TextVQA) aims to answer questions by explicitly reading and reasoning about the text involved in a video. Most works in this field follow a frame-level framework which suffers from redundant text entities and implicit relation modeling, resulting in limitations in both accuracy and efficiency. In this paper, we rethink the Video TextVQA task from an instance-oriented perspective and propose a novel model termed GAT (Gather and Trace). First, to obtain accurate reading result for each video text instance, a context-aggregated instance gathering module is designed to integrate the visual appearance, layout characteristics, and textual contents of the related entities into a unified textual representation. Then, to capture dynamic evolution of text in the video flow, an instance-focused trajectory tracing module is utilized to establish spatio-temporal relationships between instances and infer the final answer. Extensive experiments on several public Video TextVQA datasets validate the effectiveness and generalization of our framework. GAT outperforms existing Video TextVQA methods, video-language pretraining methods, and video large language models in both accuracy and inference speed. Notably, GAT surpasses the previous state-of-the-art Video TextVQA methods by 3.86% in accuracy and achieves ten times of faster inference speed than video large language models. The source code is available at https://github.com/zhangyan-ucas/GAT.
Gangyan Zeng, Daiqing Wu, Huawen Shen, Binbin Li 0003, Yu Zhou 0015, Can Ma, Xiaojun Bi 0002
ACM Multimedia6
2025 Uni-DocDiff: A Unified Document Restoration Model Based on Diffusion
abstract
Removing various degradations from damaged documents greatly benefits digitization, downstream document analysis, and readability. Previous methods often treat each restoration task independently with dedicated models, leading to a cumbersome and highly complex document processing system. Although recent studies attempt to unify multiple tasks, they often suffer from limited scalability due to handcrafted prompts and heavy preprocessing, and fail to fully exploit inter-task synergy within a shared architecture. To address the aforementioned challenges, we propose Uni-DocDiff, a Unified and highly scalable Doc ument restoration model based on Dif fusion. Uni-DocDiff develops a learnable task prompt design, ensuring exceptional scalability across diverse tasks. To further enhance its multi-task capabilities and address potential task interference, we devise a novel Prior Pool, a simple yet comprehensive mechanism that combines both local high-frequency features and global low-frequency features. Additionally, we design the Prior Fusion Module (PFM), which enables the model to adaptively select the most relevant prior information for each specific task. Extensive experiments show that the versatile Uni-DocDiff achieves performance comparable or even superior performance compared with task-specific expert models, and simultaneously holds the task scalability for seamless adaptation to new tasks.
Fangmin Zhao, Weichao Zeng, Zhenhang Li, Dongbao Yang, Binbin Li 0003, Xiaojun Bi 0002, Yu Zhou 0015
ACM Multimedia7
2025 When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
abstract
Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visually incorrect answers, which we refer to as semantic hallucination. In this work, we investigate the underlying causes of semantic hallucination and identify a key finding: Transformer layers in LLM with stronger attention focus on scene text regions are less prone to producing semantic hallucinations. Thus, we propose a training-free semantic hallucination mitigation framework comprising two key components: (1) ZoomText, a coarse-to-fine strategy that identifies potential text regions without external detectors; and (2) Grounded Layer Correction, which adaptively leverages the internal representations from layers less prone to hallucination to guide decoding, correcting hallucinated outputs for non-semantic samples while preserving the semantics of meaningful ones. To enable rigorous evaluation, we introduce TextHalu-Bench, a benchmark of 1,740 samples spanning both semantic and non-semantic cases, with manually curated question–answer pairs designed to probe model hallucinations. Extensive experiments demonstrate that our method not only effectively mitigates semantic hallucination but also achieves strong performance on public benchmarks for scene text spotting and understanding.
Hangui Lin, Yexin Liu, Gangyan Zeng, Yu Zhou 0015, Ser-Nam Lim, Harry Yang, Nicu Sebe
NeurIPS7
2025 IPAD: Iterative, Parallel, and Diffusion-Based Network for Scene Text Recognition
Yu Zhou 0015
Int. J. Comput. Vis.3
2025 TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
abstract
Existing scene text spotters are designed to locate and transcribe texts from images. However, it is challenging for a spotter to achieve precise detection and recognition of scene texts simultaneously. Inspired by the glimpse-focus spotting pipeline of human beings and impressive performances of Pre-trained Language Models (PLMs) on visual tasks, we ask: (1) “Can machines spot texts without precise detection just like human beings?”, and if yes, (2) “Is text block another alternative for scene text spotting other than word or character?” To this end, our proposed scene text spotter leverages advanced PLMs to enhance performance without fine-grained detection. Specifically, we first use a simple detector for block-level text detection to obtain rough positional information. Then, we fine-tune a PLM using a large-scale OCR dataset to achieve accurate recognition. Benefiting from the comprehensive language knowledge gained during the pre-training phase, the PLM-based recognition module effectively handles complex scenarios, including multi-line, reversed, occluded, and incomplete-detection texts. Taking advantage of the fine-tuned language model on scene recognition benchmarks and the paradigm of text block detection, extensive experiments demonstrate the superior performance of our scene text spotter across multiple public benchmarks. Additionally, we attempt to spot texts directly from an entire scene image to demonstrate the potential of PLMs, even Large Language Models (LLMs).
Jiahao Lyu 0002, Gangyan Zeng, Enze Xie, Wei Wang 0315, Can Ma, Yu Zhou 0015
ACM Trans. Multim. Comput. Commun. Appl.8
2024 First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
abstract
Diffusion models, known for their impressive image generation abilities, have played a pivotal role in the rise of visual text generation. Nevertheless, existing visual text generation methods often focus on generating entire images with text prompts, leading to imprecise control and limited practicality. A more promising direction is visual text blending, which focuses on seamlessly merging texts onto text-free backgrounds. However, existing visual text blending methods often struggle to generate high-fidelity and diverse images due to a shortage of backgrounds for synthesis and limited generalization capabilities. To overcome these challenges, we propose a new visual text blending paradigm including both creating backgrounds and rendering texts. Specifically, a background generator is developed to produce high-fidelity and text-free natural images. Moreover, a text renderer named GlyphOnly is designed for achieving visually plausible text-background integration. GlyphOnly, built on a Stable Diffusion framework, utilizes glyphs and backgrounds as conditions for accurate rendering and consistency control, as well as equipped with an adaptive text block exploration strategy for small-scale text rendering. We also explore several downstream applications based on our method, including scene text dataset synthesis for boosting scene text detectors, as well as text image customization and editing. Code and model will be available at https://github.com/Zhenhang-Li/GlyphOnly.
Zhenhang Li, Weichao Zeng, Dongbao Yang, Yu Zhou 0015
ECAI5
2024 Accurate and Robust Scene Text Recognition via Adversarial Training
abstract
Adversarial training (AT) is a methodology that utilizes adversarial examples in the training process to enhance a model’s resistance to adversarial attacks and improve generalization. Despite its efficacy in several non-sequential computer vision tasks such as classification and object detection, its effects in the realm of Scene Text Recognition (STR) remain largely unexplored. This paper pioneers an investigation into the implications of AT on STR models and proposes a novel regularization-based AT method to develop an accurate and robust STR model, dynamically generating adversarial examples in the training procedure. Through extensive experiments across seven public real-world datasets, we find that AT not only bolsters the robustness of STR models but also improves overall recognition accuracy. This improvement is particularly significant in low-resolution images - a common challenge in STR. Furthermore, given the diverse nature of real-world text images, developing a robust STR model requires a large dataset. We propose viewing AT as a form of model-based data augmentation technique for STR, compatible with traditional augmentation methods. We hope these encouraging findings catalyze further research into the application of AT for scene text recognition.
Dongbao Yang, Yu Zhou 0015
ICASSP4
2024 Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs
abstract
Visual emotion recognition (VER) is a longstanding field that has garnered increasing attention with the advancement of deep neural networks. Although recent studies have achieved notable improvements by leveraging the knowledge embedded within pre-trained visual models, the lack of direct association between factual-level features and emotional categories, called the ''affective gap'', limits the applicability of pre-training knowledge for VER tasks. On the contrary, the explicit emotional expression and high information density in textual modality eliminate the ''affective gap''. Therefore, we propose borrowing the knowledge from the pre-trained textual model to enhance the emotional perception of pre-trained visual models. We focus on the factual and emotional connections between images and texts in noisy social media data, and propose Partitioned Adaptive Contrastive Learning (PACL) to leverage these connections. Specifically, we manage to separate different types of samples and devise distinct contrastive learning strategies for each type. By dynamically constructing negative and positive pairs, we fully exploit the potential of noisy samples. Through comprehensive experiments, we demonstrate that bridging the "affective gap'' significantly improves the performance of various pre-trained visual models in downstream emotion-related tasks. Our code is released on https://github.com/wdqqdw/PACL.
Daiqing Wu, Dongbao Yang, Yu Zhou 0015, Can Ma
ACM Multimedia3
2024 Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and Fusion
abstract
As posts on social media increase rapidly, analyzing the sentiments embedded in image-text pairs has become a popular research topic in recent years. Although existing works achieve impressive accomplishments in simultaneously harnessing image and text information, they lack the considerations of possible low-quality and missing modalities. In real-world applications, these issues might frequently occur, leading to urgent needs for models capable of predicting sentiment robustly. Therefore, we propose a Distribution-based feature Recovery and Fusion (DRF) method for robust multimodal sentiment analysis of image-text pairs. Specifically, we maintain a feature queue for each modality to approximate their feature distributions, through which we can simultaneously handle low-quality and missing modalities in a unified framework. For low-quality modalities, we reduce their contributions to the fusion by quantitatively estimating modality qualities based on the distributions. For missing modalities, we build inter-modal mapping relationships supervised by samples and distributions, thereby recovering the missing modalities from available ones. In experiments, two disruption strategies that corrupt and discard some modalities in samples are adopted to mimic the low-quality and missing modalities in various real-world scenarios. Through comprehensive experiments on three publicly available image-text datasets, we demonstrate the universal improvements of DRF compared to SOTA methods under both two strategies, validating its effectiveness in robust multimodal sentiment analysis.
Daiqing Wu, Dongbao Yang, Yu Zhou 0015, Can Ma
ACM Multimedia3
2024 Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
abstract
Scene text retrieval aims to find all images containing the query text from an image gallery. Current efforts tend to adopt an Optical Character Recognition (OCR) pipeline, which requires complicated text detection and/or recognition processes, resulting in inefficient and inflexible retrieval. Different from them, in this work we propose to explore the intrinsic potential of Contrastive Language-Image Pre-training (CLIP) for OCR-free scene text retrieval. Through empirical analysis, we observe that the main challenges of CLIP as a text retriever are: 1) limited text perceptual scale, and 2) entangled visual-semantic concepts. To this end, a novel model termed FDP (Focus, Distinguish, and Prompt) is developed. FDP first focuses on scene text via shifting the attention to the text area and probing the hidden text knowledge, and then divides the query text into content word and function word for processing, in which a semantic-aware prompting scheme and a distracted queries assistance module are utilized. Extensive experiments show that FDP significantly enhances the inference speed while achieving better or competitive retrieval accuracy compared to existing methods. Notably, on the IIIT-STR benchmark, FDP surpasses the state-of-the-art model by 4.37% with a 4 times faster speed. Furthermore, additional experiments under phrase-level and attribute-aware scene text retrieval settings validate FDP's particular advantages in handling diverse forms of query text. The source code will be available at https://github.com/Gyann-z/FDP.
Gangyan Zeng, Yuan Zhang 0013, Dongbao Yang, Peng Zhang 0044, Yiwen Gao 0001, Xugong Qin, Yu Zhou 0015
ACM Multimedia8
2024 TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
abstract
Centred on content modification and style preservation, Scene Text Editing (STE) remains a challenging task despite considerable progress in text-to-image synthesis and text-driven image manipulation recently. GAN-based STE methods generally encounter a common issue of model generalization, while Diffusion-based STE methods suffer from undesired style deviations. To address these problems, we propose TextCtrl, a diffusion-based method that edits text with prior guidance control. Our method consists of two key components: (i) By constructing fine-grained text style disentanglement and robust text glyph structure representation, TextCtrl explicitly incorporates Style-Structure guidance into model design and network training, significantly improving text style consistency and rendering accuracy. (ii) To further leverage the style prior, a Glyph-adaptive Mutual Self-attention mechanism is proposed which deconstructs the implicit fine-grained features of the source image to enhance style consistency and vision quality during inference. Furthermore, to fill the vacancy of the real-world STE evaluation benchmark, we create the first real-world image-pair dataset termed ScenePair for fair comparisons. Experiments demonstrate the effectiveness of TextCtrl compared with previous methods concerning both style fidelity and text accuracy. Project page: https://github.com/weichaozeng/TextCtrl.
Weichao Zeng, Zhenhang Li, Dongbao Yang, Yu Zhou 0015
NeurIPS5
2024 Show Exemplars and Tell Me What You See: In-Context Learning with Frozen Large Language Models for TextVQA
Gangyan Zeng, Huawen Shen, Can Ma, Yu Zhou 0015
PRCV (7)5
2024 Masked and Permuted Implicit Context Learning for Scene Text Recognition
abstract
Scene Text Recognition (STR) is challenging because of various text styles, shapes, and backgrounds. Although the integration of linguistic information enhances models' performance, existing methods based on either permuted language modeling (PLM) or masked language modeling (MLM) have their drawbacks. PLM's autoregressive decoding lacks foresight into subsequent characters, while MLM overlooks intercharacter dependencies. To address these problems, we propose a masked and permuted implicit context learning network for STR, which unifies PLM and MLM within a single decoder, inheriting the advantages of both approaches. We utilize the training procedure of PLM and incorporate word length information into the decoding process to integrate MLM, substituting the undetermined characters with mask tokens. Besides, we employ the perturbation training technique to train a more robust model against potential length prediction errors. Our comprehensive evaluations demonstrate the performance of our model. It achieves superior performance on the popularly used benchmarks and outperforms previous state-of-the-art methods with a substantial improvement of 9.1% on the more challenging Union14M-Benchmark.
Dongbao Yang, Yu Zhou 0015
IEEE Signal Process. Lett.5
2024 Beyond Instance Discrimination: Relation-Aware Contrastive Self-Supervised Learning
abstract
Contrastive self-supervised learning (CSL) based on instance discrimination typically attracts positive samples while repelling negatives to learn representations with pre-defined binary self-supervision. However, vanilla CSL is inadequate in modeling sophisticated instance relations, limiting the learned model to retain fine semantic structure. On the one hand, samples with the same semantic category are inevitably pushed away as negatives. On the other hand, differences among samples cannot be captured. In this paper, we present relation-aware contrastive self-supervised learning (ReCo) to integrate instance relations, i.e., global distribution relation and local interpolation relation, into the CSL framework in a plug-and-play fashion. Specifically, we align similarity distributions calculated between the positive anchor views and the negatives at the global level to exploit diverse similarity relations among instances. Local-level interpolation consistency between the pixel space and the feature space is applied to quantitatively model the feature differences of samples with distinct apparent similarities. Through explicitly instance relation modeling, our ReCo avoids irrationally pushing away semantically identical samples and carves a well-structured feature space. Extensive experiments conducted on commonly used benchmarks justify that our ReCo consistently gains remarkable performance improvements.
Yifei Zhang 0005, Chang Liu 0047, Yu Zhou 0015, Weiping Wang 0005, Qixiang Ye, Xiangyang Ji
IEEE Trans. Multim.3
2023 One-Shot Replay: Boosting Incremental Object Detection via Retrospecting One Object
abstract
Modern object detectors are ill-equipped to incrementally learn new emerging object classes over time due to the well-known phenomenon of catastrophic forgetting. Due to data privacy or limited storage, few or no images of the old data can be stored for replay. In this paper, we design a novel One-Shot Replay (OSR) method for incremental object detection, which is an augmentation-based method. Rather than storing original images, only one object-level sample for each old class is stored to reduce memory usage significantly, and we find that copy-paste is a harmonious way to replay for incremental object detection. In the incremental learning procedure, diverse augmented samples with co-occurrence of old and new objects to existing training data are generated. To introduce more variants for objects of old classes, we propose two augmentation modules. The object augmentation module aims to enhance the ability of the detector to perceive potential unknown objects. The feature augmentation module explores the relations between old and new classes and augments the feature space via analogy. Extensive experimental results on VOC2007 and COCO demonstrate that OSR can outperform the state-of-the-art incremental object detection methods without using extra wild data.
Dongbao Yang, Yu Zhou 0015, Xiaopeng Hong, Aoting Zhang, Weiping Wang 0005
AAAI2
2023 EI2SR: Learning an Enhanced Intra-Instance Semantic Relationship for Arbitrary-Shaped Scene Text Detection
abstract
Text detection in natural scenarios, has made significant progress with the deep learning architecture. Towards arbitrary-shaped text detection, fracture detection is the major concern due to the lack of semantic relationship within an instance in existing methods. To circumvent this dilemma, we propose a novel network to learn an Enhanced Intra-Instance Semantic Relationship (EI2SR) which consists of Text-Specific Attention Mechanism (TAM) and Border Attraction Grouping (BAG). The former models the rich semantic information between different coarse-grained text regions to guide the fine-grained learning of corresponding text representations. The latter enhances the border-center semantic correlation by establishing high-dimension embedding space to attract and group the border at both ends to their corresponding center. Extensive experimental results show that the proposed EI2SR achieves state-of-the-art or competitive performance on existing benchmarks.
Shaohui Liu, Yu Zhou 0015, Feng Jiang 0001
ICASSP3
2023 UATVR: Uncertainty-Adaptive Text-Video Retrieval
abstract
With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted increasing attention. A common practice is to transfer text-video pairs to the same embedding space and craft cross-modal interactions with certain entities in specific granularities for semantic correspondence. Unfortunately, the intrinsic uncertainties of optimal entity combinations in appropriate granularities for cross-modal queries are understudied, which is especially critical for modalities with hierarchical semantics, e.g., video, text, etc. In this paper, we propose an Uncertainty-Adaptive Text-Video Retrieval approach, termed UATVR, which models each lookup as a distribution matching procedure. Concretely, we add additional learnable tokens in the encoders to adaptively aggregate multi-grained semantics for flexible high-level reasoning. In the refined embedding space, we represent text-video pairs as probabilistic distributions where prototypes are sampled for matching evaluation. Comprehensive experiments on four benchmarks justify the superiority of our UATVR, which achieves new state-of-the-art results on MSR-VTT (50.8%), VATEX (64.5%), MSVD (49.7%), and DiDeMo (45.8%). The code is available at https://github.com/bofang98/UATVR.
Bo Fang 0003, Yu Zhou 0015, YuXin Song 0001, Weiping Wang 0005, Xiangbo Shu, Xiangyang Ji, Jingdong Wang 0001
ICCV4
2023 Mask-Guided Stamp Erasure for Real Document Image
abstract
The application of text recognition in the automatic analysis of invoices, contracts and other documents has significantly raised office efficiency, but the stamps overlapping with the texts in these documents may seriously degrade the recognition accuracy. To mitigate the negative effect, we propose a stamp eraser, which can simultaneously remove the stamps and recover the occluded texts. To better distinguish stamps from the complex background, we propose a stamp localization module to generate fine-grained binary masks for the eraser to focus more on stamps. This module can also provide the background textual information for recovering text with a skip connection. We also propose the dilated mask to make the generated image look more natural by filling the stamp area with pixels around the stamp. In addition, to evaluate the effectiveness of our method in boosting text recognition, we propose a synthetic data set for training and a real dataset with complex background for testing. Experiments have shown that our method can effectively improve the recognition accuracy of the text occluded with the stamps.
Dongbao Yang, Yu Zhou 0015, Youhui Guo, Weiping Wang 0005
ICME3
2023 Divide Rows and Conquer Cells: Towards Structure Recognition for Large Tables
abstract
Recent advanced Table Structure Recognition (TSR) models adopt image-to-text solutions to parse table structure. These methods can be formulated as image caption problem, i.e., input a single-table image and output table structure description in a specific text format, e.g., HTML. With the impressive success of Transformer in text generation tasks, these methods use Transformer architecture to predict HTML table text in an autoregressive manner. However, tables always emerge with a large variety of shapes and sizes. Autoregressive models usually suffer from the error accumulation problem as the length of predicted text increases, which results in unsatisfactory performance for large tables. In this paper, we propose a novel image-to-text based TSR method that relieves error accumulation problems and improves performance noticeably. At the core of our method is a cascaded two-step decoder architecture with the former decoder predicting HTML table row tags non-autoregressively and the latter predicting HTML table cell tags of each row in a semi-autoregressive manner. Compared with existing methods that predict HTML text autoregressively, the superiority of our row-to-cell progressive table parsing is twofold: (1) it generates an HTML tag sequence with a vertical-and-horizontal two-step `scanning', which better fits the inherent 2D structure of image data, (2) it performs substantially better for large tables (long sequence prediction) since it alleviates error accumulation problem specific to autoregressive models. Extensive experiments demonstrate that our method achieves competitive performance on three public benchmarks.
Huawen Shen, Yu Zhou 0015, Zhanzhan Cheng
IJCAI5
2023 Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation Learning
abstract
Due to the flexible representation of arbitrary-shaped scene text and simple pipeline, bottom-up segmentation-based methods begin to be mainstream in real-time scene text detection. Despite great progress, these methods show deficiencies in robustness and still suffer from false positives and instance adhesion. Different from existing methods which integrate multiple-granularity features or multiple outputs, we resort to the perspective of representation learning in which auxiliary tasks are utilized to enable the encoder to jointly learn robust features with the main task of per-pixel classification during optimization. For semantic representation learning, we propose global-dense semantic contrast (GDSC), in which a vector is extracted for global semantic representation, then used to perform element-wise contrast with the dense grid features. To learn instance-aware representation, we propose to combine top-down modeling (TDM) with the bottom-up framework to provide implicit instance-level clues for the encoder. With the proposed GDSC and TDM, the encoder network learns stronger representation without introducing any parameters and computations during inference. Equipped with a very light decoder, the detector can achieve more robust real-time scene text detection. Experimental results on four public datasets show that the proposed method can outperform or be comparable to the state-of-the-art on both accuracy and speed. Specifically, the proposed method achieves 87.2% F-measure with 48.2 FPS on Total-Text and 89.6% F-measure with 36.9 FPS on MSRA-TD500 on a single GeForce RTX 2080 Ti GPU.
Xugong Qin, Pengyuan Lv, Chengquan Zhang, Yu Zhou 0015, Peng Zhang 0044, Hailun Lin, Weiping Wang 0005
ACM Multimedia4
2023 Perceiving Ambiguity and Semantics without Recognition: An Efficient and Effective Ambiguous Scene Text Detector
abstract
Ambiguous scene text detection is an extremely challenging task. Existing text detectors that rely solely on visual cues often suffer from confusion due to being evenly distributed in rows/columns or incomplete detection owing to large character spacing. To overcome these challenges, the previous method recognizes a large number of proposals and utilizes semantic information predicted from recognition results to eliminate ambiguity. However, this method is inefficient, which limits their practical applications. In this paper, we propose a novel efficient and effective ambiguous text detector, which can Perceive Ambiguity and SEmantics without Recognition, termed PASER. On the one hand, PASER can perceive semantics without recognition with a light Perceiving Semantics (PerSem) module. In this way, proposals without reasonable semantics are filtered out, which largely speeds up the overall detection process. On the other hand, to detect both ambiguous and regular texts with a unified framework, PASER employs a Perceiving Ambiguity (PerAmb) module to distinguish ambiguous texts and regular texts, so that only the ambiguous proposals will be processed by PerSem while the regular texts are not, which further ensures the high efficiency. Extensive experiments show that our detector achieves state-of-the-art results on both ambiguous and regular scene text detection benchmarks. Notably, over 6 times faster speed and superior accuracy are achieved on TDA-ReCTS simultaneously.
Wei Wang 0315, Yu Zhou 0015, Shaohui Liu, Aoting Zhang, Dongbao Yang, Weiping Wang 0005
ACM Multimedia3
2023 Pseudo Object Replay and Mining for Incremental Object Detection
abstract
Incremental object detection (IOD) aims to mitigate catastrophic forgetting for object detectors when incrementally learning to detect new emerging object classes without using original training data. Most existing IOD methods benefit from the assumption that unlabeled old-class objects may co-occur with labeled new-class objects in the new training data. However, in practical scenarios, old-class objects may be absent, which is called non co-occurrence IOD. In this paper, we propose a pseudo object replay and mining method (PseudoRM) to handle the co-occurrence dependent problem, reducing the performance degradation caused by the absence of old-class objects. The new training data can be augmented by co-occurring fake (old-class) and real (new-class) objects with a patch-level data-free generation method in the pseudo object replay stage. To fully use existing training data, we propose pseudo object mining to explore false positives for transferring useful instance-level knowledge. In the incremental learning procedure, a generative distillation is introduced to distill image-level knowledge for balancing stability and plasticity. Experimental results on PASCAL VOC and COCO demonstrate that PseudoRM can effectively boost the performance on both co-occurrence and non co-occurrence scenarios without using old samples or extra wild data.
Dongbao Yang, Yu Zhou 0015, Xiaopeng Hong, Aoting Zhang, Linchengxi Zeng, Weiping Wang 0005
ACM Multimedia2
2023 Filling in the Blank: Rationale-Augmented Prompt Tuning for TextVQA
abstract
Recently, generative Text-based visual question answering (TextVQA) methods, which are often based on language models, have exhibited impressive results and drawn increasing attention. However, due to the inconsistencies in both input forms and optimization objectives, the power of pretrained language models is not fully explored, resulting in the need for large amounts of training data. In this work, we rethink the characteristics of the TextVQA task and find that scene text is indeed a special kind of language embedded in images. To this end, we propose a text-centered generative framework FITB (stands for Filling In The Blank), in which multimodal information is mainly represented in textual form and rationale-augmented prompting is involved. Specifically, an infilling-based prompt strategy is utilized to formulate TextVQA as a novel problem of filling in the blank with proper scene text according to the language context. Furthermore, aiming to prevent the model from language bias overfitting, we design a rough answer grounding module to provide visual rationales for promoting multimodal reasoning. Extensive experiments verify the superiority of FITB in both fully-supervised and zero-shot/few-shot settings. Notably, even with a saving of about 64M data, FITB surpasses the state-of-the-art method by 3.00% and 1.99% on TextVQA and ST-VQA datasets, respectively.
Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015, Bo Fang 0003, Weiping Wang 0005
ACM Multimedia3
2023 Feature Enhancement with Text-Specific Region Contrast for Scene Text Detection
Xurui Sun, Jiahao Lyu 0002, Yifei Zhang 0005, Gangyan Zeng, Bo Fang 0003, Yu Zhou 0015, Enze Xie, Can Ma
PRCV (7)6
2023 Beyond OCR + VQA: Towards end-to-end reading and reasoning for robust and accurate textvqa
abstract
Text-based visual question answering (TextVQA), which answers a visual question by considering both visual contents and scene texts, has attracted increasing attention recently. Most existing methods employ an optical character recognition (OCR) module as a pre-processor to read texts, then combine it with a visual question answering (VQA) framework. However, inaccurate OCR results may lead to cumulative error propagation , and the correlation between text reading and text-based reasoning is not fully exploited. In this work, we integrate OCR into the flow of TextVQA, targeting the mutual reinforcement of OCR and VQA tasks. Specifically, a visually enhanced text embedding module is proposed to predict semantic features from the visual information of texts, by which texts can be reasonably understood even without accurate recognition. Further, two elaborate schemes are developed to leverage contextual information in VQA to modify OCR results. The first scheme is a reading modification module that adaptively selects the answer results according to the contexts. Second, we propose an efficient end-to-end text reading and reasoning network, where the downstream VQA signal contributes to the optimization of text reading. Extensive experiments show that our method outperforms existing alternatives in terms of accuracy and robustness, whether ground truth OCR annotations are used or not.
Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015, Weiping Wang 0005, Xu-Cheng Yin
Pattern Recognit.3
2023 Self-Supervised Motion Perception for Spatiotemporal Representation Learning
abstract
In this study, we propose a novel pretext task and a self-supervised motion perception (SMP) method for spatiotemporal representation learning. The pretext task is defined as video playback rate perception, which utilizes temporal dilated sampling to augment video clips to multiple duplicates of different temporal resolutions. The SMP method is built upon discriminative and generative motion perception models, which capture representations related to motion dynamics and appearance from video clips of multiple temporal resolutions in a collaborative fashion. To enhance the collaboration, we further propose difference and convolution motion attention (MA), which drives the generative model focusing on motion-related appearance, and leverage multiple granularity perception (MG) to extract accurate motion dynamics. Extensive experiments demonstrate SMP's effectiveness for video motion perception and state-of-the-art performance of self-supervised representation models upon target tasks, including action recognition and video retrieval. Code for SMP is available at github.com/yuanyao366/SMP.
Chang Liu 0047, Dezhao Luo, Yu Zhou 0015, Qixiang Ye
IEEE Trans. Neural Networks Learn. Syst.4
2022 Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification
abstract
Real-world data often follows a long-tailed distribution, which makes the performance of existing classification algorithms degrade heavily. A key issue is that the samples in tail categories fail to depict their intra-class diversity. Humans can imagine a sample in new poses, scenes and view angles with their prior knowledge even if it is the first time to see this category. Inspired by this, we propose a novel reasoning-based implicit semantic data augmentation method to borrow transformation directions from other classes. Since the covariance matrix of each category represents the feature transformation directions, we can sample new directions from similar categories to generate definitely different instances. Specifically, the long-tailed distributed data is first adopted to train a backbone and a classifier. Then, a covariance matrix for each category is estimated, and a knowledge graph is constructed to store the relations of any two categories. Finally, tail samples are adaptively enhanced via propagating information from all the similar categories in the knowledge graph. Experimental results on CIFAR-LT-100, ImageNet-LT, and iNaturalist 2018 have demonstrated the effectiveness of our proposed method compared with the state-of-the-art methods.
Xiaohua Chen 0002, Yucan Zhou, Dayan Wu, Wanqian Zhang, Yu Zhou 0015, Bo Li 0063, Weiping Wang 0005
AAAI5
2022 Video Motion Perception for Self-supervised Representation Learning
Dezhao Luo, Bo Fang 0003, Xiaoni Li, Yu Zhou 0015, Weiping Wang 0005
ICANN (4)5
2022 Towards Escaping from Language Bias and OCR Error: Semantics-Centered Text Visual Question Answering
abstract
Texts in scene images convey critical information for scene understanding and reasoning. The abilities of reading and rea-soning matter for the model in the text-based visual question answering (TextVQA) process. However, current TextVQA models do not center on the text and suffer from several limitations. The model is easily dominated by language biases and optical character recognition (OCR) errors due to the ab-sence of semantic guidance in the answer prediction process. In this paper, we propose a novel Semantics-Centered Net-work (SC-Net) that consists of an instance-level contrastive semantic prediction module (ICSP) and a semantics-centered transformer module (SCT). Equipped with the two modules, the semantics-centered model can resist the language biases and the accumulated errors from OCR. Extensive experiments on TextVQA and ST-VQA datasets show the effectiveness of our model. SC- Net surpasses previous works with a notice-able margin and is more reasonable for the TextVQA task.
Chengyang Fang, Gangyan Zeng, Yu Zhou 0015, Daiqing Wu, Can Ma, Dayong Hu, Weiping Wang 0005
ICME3
2022 UNITS: Unsupervised Intermediate Training Stage for Scene Text Detection
abstract
Recent scene text detection methods are almost based on deep learning and data-driven. Synthetic data is commonly adopted for pre-training due to expensive annotation cost. However, there are obvious domain discrepancies between synthetic data and real-world data. It may lead to suboptimal performance to directly adopt the model initialized by synthetic data in the fine-tuning stage. In this paper, we propose a new training paradigm for scene text detection, which introduces an UNsupervised Intermediate Training Stage (UNITS) that builds a buffer path to real-world data and can alleviate the gap between the pre-training stage and fine-tuning stage. Three training strategies are further explored to perceive information from real-world data in an unsupervised way. With UNITS, scene text detectors are improved without introducing any parameters and computations during inference. Extensive experimental results show consistent performance improvements on three public datasets.
Youhui Guo, Yu Zhou 0015, Xugong Qin, Enze Xie, Weiping Wang 0005
ICME2
2022 MaMiCo: Macro-to-Micro Semantic Correspondence for Self-supervised Video Representation Learning
abstract
Contrastive self-supervised learning (CSL) has remarkably promoted the progress of visual representation learning. However, existing video CSL methods mainly focus on clip-level temporal semantic consistency. The temporal and spatial semantic correspondence across different granularities, i.e., video, clip, and frame levels, is typically overlooked. To tackle this issue, we propose a self-supervised Macro-to-Micro Semantic Correspondence (MaMiCo) learning framework, pursuing fine-grained spatiotemporal representations from a macro-to-micro perspective. Specifically, MaMiCo constructs a multiple branch architecture of T-MaMiCo and S-MaMiCo on a temporally-nested clip pyramid (video-to-frame). On the pyramid, T-MaMiCo aims at temporal correspondence by simultaneously assimilating semantic invariance representations and retaining appearance dynamics in long temporal ranges. For spatial correspondence, S-MaMiCo perceives subtle motion cues via ameliorating dense CSL for videos where stationary clips are applied for stably dense contrasting reference to alleviate semantic inconsistency caused by ''mismatching''. Extensive experiments justify that MaMiCo learns rich general video representations and works well on various downstream tasks, e.g., (fine-grained) action recognition, action localization, and video retrieval.
Bo Fang 0003, Chang Liu 0042, Yu Zhou 0015, Dongliang He, Weiping Wang 0005
ACM Multimedia4
2022 TPSNet: Reverse Thinking of Thin Plate Splines for Arbitrary Shape Scene Text Representation
abstract
The research focus of scene text detection and recognition has shifted to arbitrary shape text in recent years, where the text shape representation is a fundamental problem. An ideal representation should be compact, complete, efficient, and reusable for subsequent recognition in our opinion. However, previous representations have flaws in one or more aspects. Thin-Plate-Spline (TPS) transformation has achieved great success in scene text recognition. Inspired by this, we reversely think of its usage and sophisticatedly take TPS as an exquisite representation for arbitrary shape text representation. The TPS representation is compact, complete, and efficient. With the predicted TPS parameters, the detected text region can be directly rectified to a near-horizontal one to assist the subsequent recognition. To further exploit the potential of the TPS representation, the Border Alignment Loss is proposed. Based on these designs, we implement the text detector TPSNet, which can be extended to a text spotter conveniently. Extensive evaluation and ablation of several public benchmarks demonstrate the effectiveness and superiority of the proposed method for text representation and spotting. Particularly, TPSNet achieves the detection F-Measure improvement of 4.4% (78.4% vs. 74.0%) on Art dataset and the end-to-end spotting F-Measure improvement of 5.0% (78.5% vs. 73.5%) on Total-Text, which are large margins with no bells and whistles. The source code will be available.
Wei Wang 0315, Yu Zhou 0015, Jiahao Lyu 0002, Dayan Wu, Weiping Wang 0005
ACM Multimedia2
2022 TextBlock: Towards Scene Text Spotting without Fine-grained Detection
abstract
Scene text spotting systems which integrate text detection and recognition modules have witnessed a lot of success in recent years. Existing works mostly follow the framework of word/character-level fine-grained detection and isolated-instance recognition, which overemphasize the role of detector and ignore the rich context information in recognition. After rethinking the conventional framework, and inspired by the glimpse-focus spotting pipeline of human beings, we ask:1) "can machine spot text without accurate detection just like human beings?", and if yes, 2) "is text block another alternative for scene text spotting other than word or character?". Based on these questions, we propose a new perspective of coarse-grained detection with multi-instance recognition for text spotting. Specifically, a pioneering network termed TextBlock is developed, and a heuristic text block generation method as well as a multi-instance block-level recognition module are proposed. In this way, the burden of detection is relieved, and the contextual semantic information is well explored for recognition. To train the block-level recognizer, a synthetic dataset including about 800K images is formed. As a by-product of attention, fine-grained detection can be recovered with the recognizer. Equipped with a detector without many bells and whistles (e.g., Faster R-CNN), TextBlock achieves competitive or even better performance compared with previous sophisticated text spotters on several public benchmarks. As a primary attempt, we expect this framework will have a potential impact on scene text spotting research in the future.
Yuan Zhang 0013, Yu Zhou 0015, Gangyan Zeng, Youhui Guo, Haiying Wu, Weiping Wang 0005
ACM Multimedia3
2022 Multi-View correlation distillation for incremental object detection
Dongbao Yang, Yu Zhou 0015, Aoting Zhang, Xurui Sun, Dayan Wu, Weiping Wang 0005, Qixiang Ye
Pattern Recognit.2
2022 Deep collaborative multi-task network: A human decision process inspired model for hierarchical image classification
Yu Zhou 0015, Xiaoni Li, Yucan Zhou, Yu Wang 0106, Qinghua Hu, Weiping Wang 0005
Pattern Recognit.1
2022 Exploring Relations in Untrimmed Videos for Self-Supervised Learning
abstract
Existing video self-supervised learning methods mainly rely on trimmed videos for model training. They apply their methods and verify the effectiveness on trimmed video datasets including UCF101 and Kinetics-400, among others. However, trimmed datasets are manually annotated from untrimmed videos. In this sense, these methods are not truly unsupervised. In this article, we propose a novel self-supervised method, referred to as Exploring Relations in Untrimmed Videos (ERUV), which can be straightforwardly applied to untrimmed videos (real unlabeled) to learn spatio-temporal features. ERUV first generates single-shot videos by shot change detection. After that, some designed sampling strategies are used to model relations for video clips. The strategies are saved as our self-supervision signals. Finally, the network learns representations by predicting the category of relations between the video clips. ERUV is able to compare the differences and similarities of video clips, which is also an essential procedure for video-related tasks. We validate our learned models with action recognition, video retrieval, and action similarity labeling tasks with four kinds of 3D convolutional neural networks. Experimental results show that ERUV is able to learn richer representations with untrimmed videos, and it outperforms state-of-the-art self-supervised methods with significant margins.
Dezhao Luo, Yu Zhou 0015, Bo Fang 0003, Yucan Zhou, Dayan Wu, Weiping Wang 0005
ACM Trans. Multim. Comput. Commun. Appl.2
2022 RD-IOD: Two-Level Residual-Distillation-Based Triple-Network for Incremental Object Detection
abstract
As a basic component in multimedia applications, object detectors are generally trained on a fixed set of classes that are pre-defined. However, new object classes often emerge after the models are trained in practice. Modern object detectors based on Convolutional Neural Networks (CNN) suffer from catastrophic forgetting when fine-tuning on new classes without the original training data. Therefore, it is critical to improve the incremental learning capability on object detection. In this article, we propose a novel Residual-Distillation-based Incremental learning method on Object Detection (RD-IOD). Our approach rests on the creation of a triple-network based on Faster R-CNN. To enable continuous learning from new classes, we use the original model as well as a residual model to guide the learning of the incremental model on new classes while maintaining the previous learned knowledge. To better maintain the discrimination between the features of old and new classes, the residual model is jointly trained with the incremental model on new classes in the incremental learning procedure. In addition, a two-level distillation scheme is designed to guide the training process, which consists of (1) a general distillation for imitating the original model in feature space along with a residual distillation on the features in both image level and instance level, and (2) a joint classification distillation on the output layers. To well preserve the learned knowledge, we design a 2-threshold training strategy to guide the learning of a Region Proposal Network and a detection head. Extensive experiments conducted on VOC2007 and COCO demonstrate that the proposed method can effectively learn to incrementally detect objects of new classes, and the problem of catastrophic forgetting is mitigated. Our code is available at https://github.com/yangdb/RD-IOD.
Dongbao Yang, Yu Zhou 0015, Wei Shi 0001, Dayan Wu, Weiping Wang 0005
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Which and Where to Focus: A Simple yet Accurate Framework for Arbitrary-Shaped Nearby Text Detection in Scene Images
Youhui Guo, Yu Zhou 0015, Xugong Qin, Weiping Wang 0005
ICANN (5)2
2021 MMF: Multi-task Multi-structure Fusion for Hierarchical Image Classification
Xiaoni Li, Yucan Zhou, Yu Zhou 0015, Weiping Wang 0005
ICANN (4)3
2021 FC2RN: A Fully Convolutional Corner Refinement Network for Accurate Multi-Oriented Scene Text Detection
abstract
Accurate detection of multi-oriented text that accounts for a large proportion in real practice is of great significance. The performance has improved rapidly on common benchmarks in recent years. However, dense long text case and the quality of detection are easy to be overlooked. Direct regression may produce low-quality and incomplete detections due to the constrain of the receptive field; proposal-based methods could alleviate this but might introduce redundant context due to RoI operation, degrading the performance. To address the dilemma, a novel proposed corner-aware convolution in which the sampling positions tightly cover the text area is utilized to encode an initial corner prediction into the feature maps, which can be further used to produce a refined corner prediction. We embed the proposed module into an anchor-free baseline model, leading to a simple and effective fully convolutional corner refinement network (FC2RN). Experimental results on four public datasets including MSRATD500, ICDAR2015, RCTW-17, and COCO-Text demonstrate that FC2RN can outperform state-of-the-art methods.
Xugong Qin, Yu Zhou 0015, Youhui Guo, Dayan Wu, Weiping Wang 0005
ICASSP2
2021 Density-Net: A Density-Aware Network for 3D Object Detection
abstract
To address the problem of density imbalance between nearby and faraway regions in point clouds, we propose a density-aware 3D single-stage object detector named Density-Net. First, a new 3D data augmentation method is designed to generate low-density areas by using the DBSCAN cluster method to acquire the cluster of point clouds, then by setting several different downsample ratios to stimulate sparse regions. By this way, there will be many sparse regions in point clouds. To restore the representative information, we propose Density-Set-Abstraction (Density-SA) to harmonize the high-level representative feature and low-level spatial feature, which boosts the final object detection performance. Moreover, we design a new Mask Sample branch to accomplish point sampling based on point-wise annotations, which increases the recall rate of meaningful points. We evaluate it on the widely used KITTI dataset. The experiments demonstrate that our method outperforms the well-established and highly-optimized 3DSSD (implemented by MMdetection3D) baseline 2.2% APs in moderate difficulty setting. At the same time, the runtime performance is also on par with 3DSSD. Code: https://github.com/Physu/density-net
Youhui Guo, Yu Zhou 0015, Weiping Wang 0005
ICTAI3
2021 Dense Semantic Contrast for Self-Supervised Visual Representation Learning
abstract
Self-supervised representation learning for visual pre-training has achieved remarkable success with sample (instance or pixel) discrimination and semantics discovery of instance, whereas there still exists a non-negligible gap between pre-trained model and downstream dense prediction tasks. Concretely, these downstream tasks require more accurate representation, in other words, the pixels from the same object must belong to a shared semantic category, which is lacking in the previous methods. In this work, we present Dense Semantic Contrast (DSC) for modeling semantic category decision boundaries at a dense level to meet the requirement of these tasks. Furthermore, we propose a dense cross-image semantic contrastive learning framework for multi-granularity representation learning. Specially, we explicitly explore the semantic structure of the dataset by mining relations among pixels from different perspectives. For intra-image relation modeling, we discover pixel neighbors from multiple views. And for inter-image relations, we enforce pixel representation from the same semantic class to be more similar than the representation from different classes in one mini-batch. Experimental results show that our DSC model outperforms state-of-the-art methods when transferring to downstream dense prediction tasks, including object detection, semantic segmentation, and instance segmentation. Code will be made available.
Xiaoni Li, Yu Zhou 0015, Yifei Zhang 0005, Aoting Zhang, Wei Wang 0315, Haiying Wu, Weiping Wang 0005
ACM Multimedia2
2021 PIMNet: A Parallel, Iterative and Mimicking Network for Scene Text Recognition
abstract
Nowadays, scene text recognition has attracted more and more attention due to its various applications. Most state-of-the-art methods adopt an encoder-decoder framework with attention mechanism, which generates text autoregressively from left to right. Despite the convincing performance, the speed is limited because of the one-by-one decoding strategy. As opposed to autoregressive models, non-autoregressive models predict the results in parallel with a much shorter inference time, but the accuracy falls behind the autoregressive counterpart considerably. In this paper, we propose a Parallel, Iterative and Mimicking Network (PIMNet) to balance accuracy and efficiency. Specifically, PIMNet adopts a parallel attention mechanism to predict the text faster and an iterative generation mechanism to make the predictions more accurate. In each iteration, the context information is fully explored. To improve learning of the hidden layer, we exploit the mimicking learning in the training phase, where an additional autoregressive decoder is adopted and the parallel decoder mimics the autoregressive decoder with fitting outputs of the hidden layer. With the shared backbone between the two decoders, the proposed PIMNet can be trained end-to-end without pre-training. During inference, the branch of the autoregressive decoder is removed for a faster speed. Extensive experiments on public benchmarks demonstrate the effectiveness and efficiency of PIMNet. Our code is available in the supplementary material.
Yu Zhou 0015, Wei Wang 0315, Yuan Zhang 0013, Weiping Wang 0005
ACM Multimedia2
2021 Mask is All You Need: Rethinking Mask R-CNN for Dense and Arbitrary-Shaped Scene Text Detection
abstract
Due to the large success in object detection and instance segmentation, Mask R-CNN attracts great attention and is widely adopted as a strong baseline for arbitrary-shaped scene text detection and spotting. However, two issues remain to be settled. The first is dense text case, which is easy to be neglected but quite practical. There may exist multiple instances in one proposal, which makes it difficult for the mask head to distinguish different instances and degrades the performance. In this work, we argue that the performance degradation results from the learning confusion issue in the mask head. We propose to use an MLP decoder instead of the "deconv-conv" decoder in the mask head, which alleviates the issue and promotes robustness significantly. And we propose instance-aware mask learning in which the mask head learns to predict the shape of the whole instance rather than classify each pixel to text or non-text. With instance-aware mask learning, the mask branch can learn separated and compact masks. The second is that due to large variations in scale and aspect ratio, RPN needs complicated anchor settings, making it hard to maintain and transfer across different datasets. To settle this issue, we propose an adaptive label assignment in which all instances especially those with extreme aspect ratios are guaranteed to be associated with enough anchors. Equipped with these components, the proposed method named MAYOR achieves state-of-the-art performance on five benchmarks including DAST1500, MSRA-TD500, ICDAR2015, CTW1500, and Total-Text.
Xugong Qin, Yu Zhou 0015, Youhui Guo, Dayan Wu, Zhihong Tian 0001, Weiping Wang 0005
ACM Multimedia2
2021 Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQA
abstract
Text-based visual question answering (TextVQA) requires analyzing both the visual contents and texts in an image to answer a question, which is more practical than general visual question answering (VQA). Existing efforts tend to regard optical character recognition (OCR) as a pre-processing and then combine it with a VQA framework. It makes the performance of multimodal reasoning and question answering highly depend on the accuracy of OCR. In this work, we address this issue with two perspectives. First, we take advantages of multimodal cues to complete the semantic information of texts. A visually enhanced text embedding is proposed to enable understanding of texts without accurately recognizing them. Second, we further leverage rich contextual information to modify the answer texts even if the OCR module does not correctly recognize them. In addition, the visual objects are endued with semantic representations to enable objects in the same semantic space as OCR tokens. Equipped with these techniques, the cumulative error propagation caused by poor OCR performance is effectively suppressed. Extensive experiments on TextVQA and ST-VQA datasets demonstrate that our approach achieves the state-of-the-art performance in terms of accuracy and robustness.
Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015
ACM Multimedia3
2021 A Cost-Efficient Framework for Scene Text Detection in the Wild
Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015
PRICAI (1)3
2021 Binary Neural Network Hashing for Image Retrieval
abstract
Hashing has become increasingly important for large-scale image retrieval, of which the low storage cost and fast searching are two key properties. However, existing methods adopt large neural networks, which are hard to be deployed in resource-limited devices due to the unacceptable memory and runtime overhead. We address that this huge overhead of neural networks somewhatviolates the appealing properties of hashing. In this paper, we propose a novel deep hashing method, called Binary Neural Network Hashing (BNNH) for fast image retrieval. Specifically, we construct an efficient binarized network architecture to provide lighter model and faster inference, which directly generates binary outputs as the desired hash codes without introducing the quantization loss. Besides, in order to circumvent the huge performance degradation caused by the extremely quantized activations, we introduce a simple yet effective activation-aware loss to explicitly guide the updating of activations in intermediate layers. Extensive experiments conducted on three benchmarks show that the proposed method outperforms the state-of-the-art binarization methods by large margins and validate the efficiency of BNNH.
Wanqian Zhang, Dayan Wu, Yu Zhou 0015, Bo Li 0063, Weiping Wang 0005, Dan Meng 0002
SIGIR3
2020 Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning
abstract
We propose a novel self-supervised method, referred to as Video Cloze Procedure (VCP), to learn rich spatial-temporal representations. VCP first generates “blanks” by withholding video clips and then creates “options” by applying spatio-temporal operations on the withheld clips. Finally, it fills the blanks with “options” and learns representations by predicting the categories of operations applied on the clips. VCP can act as either a proxy task or a target task in self-supervised learning. As a proxy task, it converts rich self-supervised representations into video clip operations (options), which enhances the flexibility and reduces the complexity of representation learning. As a target task, it can assess learned representation models in a uniform and interpretable manner. With VCP, we train spatial-temporal representation models (3D-CNNs) and apply such models on action recognition and video retrieval tasks. Experiments on commonly used benchmarks show that the trained models outperform the state-of-the-art self-supervised models with significant margins.
Dezhao Luo, Chang Liu 0042, Yu Zhou 0015, Dongbao Yang, Can Ma, Qixiang Ye, Weiping Wang 0005
AAAI3
2020 SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition
abstract
Scene text recognition is a hot research topic in computer vision. Recently, many recognition methods based on the encoder-decoder framework have been proposed, and they can handle scene texts of perspective distortion and curve shape. Nevertheless, they still face lots of challenges like image blur, uneven illumination, and incomplete characters. We argue that most encoder-decoder methods are based on local visual features without explicit global semantic information. In this work, we propose a semantics enhanced encoder-decoder framework to robustly recognize low-quality scene texts. The semantic information is used both in the encoder module for supervision and in the decoder module for initializing. In particular, the state-of-the-art ASTER method is integrated into the proposed framework as an exemplar. Extensive experiments demonstrate that the proposed framework is more robust for low-quality text images, and achieves state-of-the-art results on several benchmark datasets. The source code will be available.
Yu Zhou 0015, Dongbao Yang, Yucan Zhou, Weiping Wang 0005
CVPR2
2020 Video Playback Rate Perception for Self-Supervised Spatio-Temporal Representation Learning
abstract
In self-supervised spatio-temporal representation learning, the temporal resolution and long-short term characteristics are not yet fully explored, which limits representation capabilities of learned models. In this paper, we propose a novel self-supervised method, referred to as video Playback Rate Perception (PRP), to learn spatio-temporal representation in a simple-yet-effective way. PRP roots in a dilated sampling strategy, which produces self-supervision signals about video playback rates for representation model learning. PRP is implemented with a feature encoder, a classification module, and a reconstructing decoder, to achieve spatio-temporal semantic retention in a collaborative discrimination-generation manner. The discriminative perception model follows a feature encoder to prefer perceiving low temporal resolution and long-term representation by classifying fast-forward rates. The generative perception model acts as a feature decoder to focus on comprehending high temporal resolution and short-term representation by introducing a motion-attention mechanism. PRP is applied on typical video target tasks including action recognition and video retrieval. Experiments show that PRP outperforms state-of-the-art self-supervised models with significant margins. Code is available at github.com/yuanyao366/PRP.
Chang Liu 0042, Dezhao Luo, Yu Zhou 0015, Qixiang Ye
CVPR4
2020 Progressive Cluster Purification for Unsupervised Feature Learning
abstract
In unsupervised feature learning, sample specificity based methods ignore the inter-class information, which deteriorates the discriminative capability of representation models. Clustering based methods are error-prone to explore the complete class boundary information due to the inevitable class inconsistent samples in each cluster. In this work, we propose a novel clustering based method, which, by iteratively excluding class inconsistent samples during progressive cluster formation, alleviates the impact of noise samples in a simple-yet-effective manner. Our approach, referred to as Progressive Cluster Purification (PCP), implements progressive clustering by gradually reducing the number of clusters during training, while the sizes of clusters continuously expand consistently with the growth of model representation capability. With a well-designed cluster purification mechanism, it further purifies clusters by filtering noise samples which facilitate the subsequent feature learning by utilizing the refined clusters as pseudo-labels. Experiments on commonly used benchmarks demonstrate that the proposed PCP improves baseline method with significant margins. Our code will be available at https://github.com/zhangyifei0115/PCP.
Yifei Zhang 0005, Chang Liu 0042, Yu Zhou 0015, Wei Wang 0315, Weiping Wang 0005, Qixiang Ye
ICPR3
2020 Self-Training for Domain Adaptive Scene Text Detection
abstract
Though deep learning based scene text detection has achieved great progress, well-trained detectors suffer from severe performance degradation for different domains. In general, a tremendous amount of data is indispensable to train the detector in the target domain. However, data collection and annotation are expensive and time-consuming. To address this problem, we propose a self-training framework to automatically mine hard examples with pseudo-labels from unannotated videos or images. To reduce the noise of hard examples, a novel text mining module is implemented based on the fusion of detection and tracking results. Then, an image-to-video generation method is designed for the tasks that videos are unavailable and only images can be used. Experimental results on standard benchmarks, including ICDAR2015, MSRA-TD500, ICDAR2017 MLT, demonstrate the effectiveness of our self-training method. The simple Mask R-CNN adapted with self-training and fine-tuned on real data can achieve comparable or even superior results with the state-of-the-art methods.
Yudi Chen, Wei Wang 0315, Yu Zhou 0015, Dongbao Yang, Weiping Wang 0005
ICPR3
2020 Gaussian Constrained Attention Network for Scene Text Recognition
abstract
Scene text recognition has been a hot topic in computer vision. Recent methods adopt the attention mechanism for sequence prediction which achieve convincing results. However, we argue that the existing attention mechanism faces the problem of attention diffusion, in which the model may not focus on a certain character area. In this paper, we propose Gaussian Constrained Attention Network to deal with this problem. It is a 2D attention-based method integrated with a novel Gaussian Constrained Refinement Module, which predicts an additional Gaussian mask to refine the attention weights. Different from adopting an additional supervision on the attention weights simply, our proposed method introduces an explicit refinement. In this way, the attention weights will be more concentrated and the attention-based recognition network achieves better performance. The proposed Gaussian Constrained Refinement Module is flexible and can be applied to existing attention-based methods directly. The experiments on several benchmark datasets demonstrate the effectiveness of our proposed method. Our code has been available at https://github.com/Pay20Y/GCAN.
Xugong Qin, Yu Zhou 0015, Weiping Wang 0005
ICPR3
2020 Deep Unsupervised Hybrid-similarity Hadamard Hashing
abstract
Hashing has become increasingly important for large-scale image retrieval. Recently, deep supervised hashing has shown promising performance, yet little work has been done under the more realistic unsupervised setting. The most challenging problem in unsupervised hashing methods is the lack of supervised information. Besides, existing methods fail to distinguish image pairs with different similarity degrees, which leads to a suboptimal construction of similarity matrix. In this paper, we propose a simple yet effective unsupervised hashing method, dubbed Deep Unsupervised Hybrid-similarity Hadamard Hashing (DU3H), which tackles these issues in an end-to-end deep hashing framework. DU3H employs orthogonal Hadamard codes to provide auxiliary supervised information in unsupervised setting, which can maximally satisfy the independence and balance properties of hash codes. Moreover, DU3H utilizes both highly and normally confident image pairs to jointly construct a hybrid-similarity matrix, which can magnify the impacts of different pairs to better preserve the semantic relations between images. Extensive experiments conducted on three widely used benchmarks validate the superiority of DU3H.
Wanqian Zhang, Dayan Wu, Yu Zhou 0015, Bo Li 0063, Weiping Wang 0005, Dan Meng 0002
ACM Multimedia3
2020 Asymmetric Deep Hashing for Efficient Hash Code Compression
abstract
Benefiting from recent advances in deep learning, deep hashing methods have achieved promising performance in large-scale image retrieval. To improve storage and computational efficiency, existing hash codes need to be compressed accordingly. However, previous deep hashing methods have to retrain their models and then regenerate the whole database codes using the new models when code length changes, which is time consuming especially for large image databases. In this paper, we propose a novel deep hashing method, called Code Compression oriented Deep Hashing (CCDH), for efficiently compressing hash codes. CCDH learns deep hash functions for query images, while learning a one-hidden-layer Variational Autoencoder (VAE) from existing hash codes. With such asymmetric design, CCDH can efficiently compress database codes only using the learned encoder of VAE. Furthermore, CCDH is flexible enough to be used with a variety of deep hashing methods. Extensive experiments on three widely used image retrieval benchmarks demonstrate that CCDH can significantly reduce the cost for compressing database codes when code length changes while keeping the state-of-the-art retrieval accuracy.
Shu Zhao 0006, Dayan Wu, Wanqian Zhang, Yu Zhou 0015, Bo Li 0063, Weiping Wang 0005
ACM Multimedia4
2019 Curved Text Detection in Natural Scene Images with Semi- and Weakly-Supervised Learning
abstract
Detecting curved text in the wild is very challenging. Recently, most state-of-the-art methods are segmentation based and require pixel-level annotations. We propose a novel scheme to train an accurate text detector using only a small amount of pixel-level annotated data and a large amount of data annotated with rectangles or even unlabeled data. A light model is first obtained by training with the pixel-level annotated data and then used to annotate unlabeled or weakly labeled data. A novel strategy which utilizes ground-truth bounding boxes to generate pseudo mask annotations is proposed in weakly-supervised learning. Experimental results on CTW1500 and Total-Text demonstrate that our method can substantially reduce the requirement of pixel-level annotated data. Our method can also generalize well across the two datasets. The performance of the proposed method is comparable with the state-of-the-art methods with only 10% pixel-level annotated data and 90% rectangle-level weakly annotated data.
Xugong Qin, Yu Zhou 0015, Dongbao Yang, Weiping Wang 0005
ICDAR2
2019 Constrained Relation Network for Character Detection in Scene Images
Yudi Chen, Yu Zhou 0015, Dongbao Yang, Weiping Wang 0005
PRICAI (3)2
2016 Matching User Photos to Online Products with Robust Deep Features
abstract
This paper focuses on a practically very important problem of matching a real-world product photo to exactly the same item(s) in online shopping sites. The task is extremely challenging because the user photos (i.e., the queries in this scenario) are often captured in uncontrolled environments, while the product images in online shops are mostly taken by professionals with clean backgrounds and perfect lighting conditions. To tackle the problem, we study deep network architectures and training schemes, with the goal of learning a robust deep feature representation that is able to bridge the domain gap between the user photos and the online product images. Our contributions are two-fold. First, we propose an alternative of the popular contrastive loss used in siamese deep networks, namely robust contrastive loss, where we "relax" the penalty on positive pairs to alleviate over-fitting. Second, a multi-task fine-tuning approach is introduced to learn a better feature representation, which not only incorporates knowledge from the provided training photo pairs, but also explores additional information from the large ImageNet dataset to regularize the fine-tuning procedure. Experiments on two challenging real-world datasets demonstrate that both the robust contrastive loss and the multi-task fine-tuning approach are effective, leading to very promising results with a time cost suitable for real-time retrieval.
Xi Wang 0008, Zhenfeng Sun, Yu Zhou 0015, Yu-Gang Jiang 0001
ICMR4
2016 A Semantics-Aware Approach to the Automated Network Protocol Identification
abstract
Traffic classification, a mapping of traffic to network applications, is important for a variety of networking and security issues, such as network measurement, network monitoring, as well as the detection of malware activities. In this paper, we propose Securitas, a network trace-based protocol identification system, which exploits the semantic information in protocol message formats. Securitas requires no prior knowledge of protocol specifications. Deeming a protocol as a language between two processes, our approach is based upon the new insight that the n-grams of protocol traces, just like those of natural languages, exhibit highly skewed frequency-rank distribution that can be leveraged in the context of protocol identification. In Securitas, we first extract the statistical protocol message formats by clustering n-grams with the same semantics, and then use the corresponding statistical formats to classify raw network traces. Our tool involves the following key features: 1) applicable to both connection oriented protocols and connection less protocols; 2) suitable for both text and binary protocols; 3) no need to assemble IP packets into TCP or UDP flows; and 4) effective for both long-live flows and short-live flows. We implement Securitas and conduct extensive evaluations on real-world network traces containing both textual and binary protocols. Our experimental results on BitTorrent, CIFS/SMB, DNS, FTP, PPLIVE, SIP, and SMTP traces show that Securitas has the ability to accurately identify the network traces of the target application protocol with an average recall of about 97.4% and an average precision of about 98.4%. Our experimental results prove Securitas is a robust system, and meanwhile displaying a competitive performance in practice.
Xiao-chun Yun, Yipeng Wang 0001, Yongzheng Zhang 0002, Yu Zhou 0015
IEEE/ACM Trans. Netw.4
2015 Weakly Supervised Metric Learning towards Signer Adaptation for Sign Language Recognition
abstract
In this paper, we introduce metric learning into Sign Language Recognition(SLR) for the first time and propose a signer adaption framework to address signer-independent SLR. For adapting the general model to the new signer, both clustering and manifold constraints are considered in the adaptive distance metric optimization. The contribution of our work mainly lies in three-folds. Firstly, a Weakly Supervised Metric Learning(WSML) framework is proposed, which combines the clustering and manifold constraints simultaneously. Secondly, the general framework is applied to signer adaptation and achieves good performance. Thirdly, a fragment based feature is designed for sign language representation and the effectiveness is verified in large vocabulary datasets. Our proposed WSML framework can be decomposed into two key steps. The first one is to learn a generic metric from the given labeled data. Then the second step is to realize the distance metric adaptation by considering the clustering and manifold constraints with the unlabeled data. To learn a generic distance metric, the labeled data are used under clustering assumption with classical large margin hinge loss. Specifically, the distances between data points within the same cluster(with same label) should be minimized and the distances between data points from different clusters(with different labels) should be maximized. Here we define the index set with same labels as Sg = {(i, j)|yi = y j,xi,x j ∈ Xl} and the index triplet Bg = {(i, j,k)|yi = y j,yi 6= yk,xi,x j,xk ∈ Xl}. The objective function is
Fang Yin, Xiujuan Chai, Yu Zhou 0015, Xilin Chen 0001
BMVC3
2015 Semantics constrained dictionary learning for signer-independent sign language recognition
abstract
In this paper, a sparse coding based framework is proposed for sign language recognition (SLR), especially for the signer-independent case. To deal with the inter-signer variation, a dictionary capturing the common features among different signers is learnt by considering the semantic constraint. Thus for a given sign from an unknown signer, the sparse representation, which maintains more information of this specific sign class while neglecting the identity information as much as possible, can be generated. In our implementation, each sign is partitioned into a fixed number of fragments and the features fusing hand shape and moving trajectory are extracted from the fragments. The dictionary learnt from the training fragments can be taken as the basic subunits of signs and each fragment of sign video can be coded by these basis vectors. Finally, the recognition result is achieved through SVM with the concatenated sparse coding features of the fragments. The experiments and comparisons show that our method is more effective for the signer-independent recognition problem than other baseline methods. At the same time, it also performs well for the signer-dependent case.
Fang Yin, Xiujuan Chai, Yu Zhou 0015, Xilin Chen 0001
ICIP3
2015 Summarizing surveillance videos with local-patch-learning-based abnormality detection, blob sequence optimization, and type-based synopsis
Weiyao Lin, Jiwen Lu, Bing Zhou 0003, Jinjun Wang, Yu Zhou 0015
Neurocomputing6
2015 Unsupervised adaptive sign language recognition based on hypothesis comparison guided cross validation and linguistic prior filtering
Yu Zhou 0015, Xiaokang Yang 0001, Yongzheng Zhang 0002, Yipeng Wang 0001, Xiujuan Chai, Weiyao Lin
Neurocomputing1
2014 Representing And Recognizing Motion Trajectories: A Tube And Droplet Approach
abstract
This paper addresses the problem of representing and recognizing motion trajectories. We first propose to derive scene-related equipotential lines for points in a motion trajectory and concatenate them to construct a 3D tube for representing the trajectory. Based on this 3D tube, a droplet-based method is further proposed which derives a "water droplet" from the 3D tube and recognizes trajectory activities accordingly. Our proposed 3D tube can effectively embed both motion and scene-related information of a motion trajectory while the proposed droplet- based method can suitably catch the characteristics of the 3D tube for activity recognition. Experimental results demonstrate the effectiveness of our approach.
Weiyao Lin, Hang Su 0006, Jianxin Wu 0001, Jinjun Wang, Yu Zhou 0015
ACM Multimedia6
2014 A Segmentation Pattern Based Approach to Automated Protocol Identification
abstract
In-depth understanding of network traffic is important for a variety of applications, such as network management and network security. In this paper, we propose a novel protocol identification system PSKS, which relies on the statistical signatures of network packet payloads. The proposed approach is based on the key insight that message segmentation patterns can be leveraged for accurate application identification. Specifically, the segmentation possibility for every position of protocol messages exhibits highly skewed frequency distribution due to the reason that different protocols have different message formats (i.e., Distinct message segmentation patterns). Motivated by this observation, we want to extract statistical application fingerprints by exploiting the message segmentation patterns. In PSKS, we first extract the message segmentation patterns by scoring the segmentation possibility scale for each position of messages, and then extract statistical signatures by Kolmogorov-Smirnov test and feed the signatures to tri-training, a collaborative learning algorithm. The tri-training can improve the generalization ability of our final classifier. We implemented and evaluated PSKS, and the experimental results show that PSKS achieves an average precision and recall of approximately 98%.
Yafei Sang, Yongzheng Zhang 0002, Yipeng Wang 0001, Yu Zhou 0015
PDCAT4
2014 Visual Similarity Based Anti-phishing with the Combination of Local and Global Features
abstract
Phishing uses a fake Web page to steal personal sensitive information such as credit card numbers and passwords. Generally, the fake Web page is visually similar to the legitimate target Web page. The phishers can obtain financial benefits through these information. Anti-phishing is very important for a variety of applications such as phishing attacks, online transaction security, and user privacy protection. In this paper, we propose a novel and effective visual similarity based phishing detection approach that compares the snapshot image pair of the suspected Web page and the protected Web page. The proposed approach is based on the key insight that both the local and the global features of the Web page image can be used to represent the visual characteristics of the Web page together. This approach is purely on the image level, and thus can effectively deal with the non-text phishing tricks including images or Flashes objects in the HTML contents. For the local feature, the existence of the target logo is detected. For the global feature, the similarity of the visible part of the Web page is considered. We implemented and evaluated the proposed approach on a large scale dataset consisting of 2,129 real world phishing Web pages and 1,367 irrelevant legitimate Web pages. The experimental results show that the proposed approach can achieve over 90.00% true positive rate and 97.00% true negative rate. Our approach has been applied in the anti-phishing project of a major Internet Service Provider and gives a periodical reports to the potential users.
Yu Zhou 0015, Yongzheng Zhang 0002, Yipeng Wang 0001, Weiyao Lin
TrustCom1
2011 Hypothesis comparison guided cross validation for unsupervised signer adaptation
abstract
Signer adaptation is important to sign language recognition systems in that a one-size-fits-all model set can not perform well on all kinds of signers. Supervised signer adaptation must utilize the labeled adaptation data that are collected explicitly. To skip the data collecting process in signer adaptation, we propose an unsupervised adaptation method called hypothesis comparison guided cross validation (HC CV) algorithm. The algorithm not only addresses the problem of overlap between the data set to be labeled and the data set for adaptation, but also employs an additional hypothesis comparison step to decrease the noise rate of the adaptation data set. Experimental results show that the HC CV adaptation algorithm is superior to the CV adaptation algorithm and the conventional self-teaching algorithm. Though the algorithm is proposed for signer adaptation, it can also be applied to speaker adaptation and writer adaptation straightforwardly.
Yu Zhou 0015, Xiaokang Yang 0001, Weiyao Lin, Yi Xu 0001, Long Xu 0001
ICME1
2011 Priority pyramid based bit allocation for multiview video coding
abstract
In Multivew Video Coding (MVC), Hierarchial B Pictures (HBP) structure is adopted to remove the redundancy of multiview videos. It makes the rate control of MVC more difficult with respect to accurate bit control and good compression efficiency. The existing rate control algorithms of MVC are based on those of monoview video coding standards, which didn't exploit the correlations of multiview video frames fully. In this paper, a Group of pictures from multiple views (called GoGOP) are rearranged in the form of a pyramid. The top layer of the pyramid containing anchor frames receives the highest priority in bits consuming. And the next layer is inferior to it but superior to others in bits consuming. Secondly, a matrix of weighting factors is further introduced to perform the bit allocation for MVC. Thirdly, the parameter updating is handled in a vector operation, which is somewhat robust and has a low level of computational complexity. The experimental results demonstrate that our proposed bit allocation cooperating with the conventional rate-distortion (R-D) model in H.264/AVC is efficient in rate control of MVC. The coding performance of the proposed algorithm is comparable to that of hierarchical quantization scheme (HQS) of MVC. Meanwhile, a small bit control error is obtained by our algorithm.
Long Xu 0001, Sam Kwong, Tiesong Zhao, Yu Zhou 0015
VCIP4
2011 A new global-based video enhancement algorithm by fusing features of multiple region-of-interests
abstract
Video enhancement plays an important role in various video applications. It is desirable to achieve high visual quality of the entire picture where multiple region-of-interests (ROIs) within the frame can be adaptively and simultaneously enhanced. In this paper, a new global-based video enhancement algorithm is proposed. The proposed algorithm first analyzes features from different ROIs. Then, a 'global' tone mapping curve is created for the entire picture which can adaptively enhance different regions at the same time. According to the statistics of ROIs, two fusion strategies, i.e., piecewise-based and factor-based fusions, are proposed for creating the global tone mapping curve. Experimental results show that the proposed algorithm can obtain more appealing perceptual quality than the state-of-the-art algorithms.
Ning Xu 0007, Weiyao Lin, Yu Zhou 0015, Yuanzhe Chen, Zhenzhong Chen 0001, Hongxiang Li 0001
VCIP3
2010 Adaptive Sign Language Recognition With Exemplar Extraction and MAP/IVFS
abstract
Sign language recognition systems suffer from the problem of signer dependence. In this letter, we propose a novel method that adapts the original model set to a specific signer with his/her small amount of training data. First, affinity propagation is used to extract the exemplars of signer independent hidden Markov models; then the adaptive training vocabulary can be automatically formed. Based on the collected sign gestures of the new vocabulary, the combination of maximum a posteriori and iterative vector field smoothing is utilized to generate signer-adapted models. Experimental results on six signers demonstrate that the proposed method can reduce the amount of the adaptation data and still can achieve high recognition performance.
Yu Zhou 0015, Xilin Chen 0001, Debin Zhao, Hongxun Yao, Wen Gao 0001
IEEE Signal Process. Lett.1
2008 Mahalanobis distance based Polynomial Segment Model for Chinese Sign Language Recogniton
abstract
Sign Language Recognition (SLR) systems are mostly based on Hidden Markov Model (HMM) and have achieved excellent results. However, the assumption of frame independence in HMM makes it inconsistent with the characteristic of strong temporal correlation in sign language signals. Polynomial Segment Model (PSM) explicitly represents the temporal evolution of sign language features as a Gaussian process with time-varying parameters. In this paper PSM is first introduced to SLR framework to solve the temporal correlation problem. Considering the correlation among the coefficients of polynomial trajectorypsilas different orders, Mahalanobis distance is used as the classification criterion to evaluate the likelihood of test data. Experimental results show that our method outperform the conventional HMM methods by 6.81% in recognition accuracy.
Yu Zhou 0015, Xilin Chen 0001, Debin Zhao, Hongxun Yao, Wen Gao 0001
ICME1