VLDB 2026 Research / reviewers in the wild / expert
Pengwen Dai
dblp:213/8479
· DBLP profile ↗
31ranked-venue papers
6as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 18 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 11 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discretization Is Not Always Better: Rethinking Deep Quantization for Asymmetric Image RetrievalabstractAsymmetric image retrieval (AIR), which typically employs a compact model for the query side and a large model for the database server, has garnered significant attention in resource-constrained environments. While deep hashing methods have shown great potential in large-scale image retrieval, current attempts for the asymmetric image retrieval overlook the differences in quantization capabilities between query and gallery networks. In AIR, the conventional quantization scheme forces the outputs of small query models to approximate the discrete outputs of large models, imposing overly rigid and stringent constraints that severely limit the optimization of small query models. Furthermore, existing deep hashing methods for AIR necessitate labeled datasets from large models, which also limits their practical applicability. To this end, we reconsider the necessity of strict discretization in AIR and propose a novel asymmetric hashing method, named Deep Correlation Alignment Hashing (DCAH). Rather than explicitly quantizing continuous query features to match discrete gallery representations, we distill the correlation across both models and introduce a Correlation Alignment based Quantization (CAQ) scheme, thereby implicitly accomplishing quantization. To preserve the similarity consistency between the query and gallery models, we further employ a correlation alignment-based knowledge distillation strategy which is intrinsically compatible with the CAQ. Notably, the proposed quantization scheme can function as a plug-and-play module that seamlessly integrates with existing AIR methods. Comprehensive evaluations on three real-world benchmark datasets demonstrate the effectiveness of the proposed quantization scheme CAQ, and also show that DCAH achieves state-of-the-art performance in asymmetric image retrieval scenarios. Dayan Wu, Hengjie Zhu, Chenming Wu, Pengwen Dai |
AAAI | 5 |
| 2026 | One2Seq: One-Token Wise Decoder for Efficient Scene Text RecognitionabstractAuto-regressive (AR)-based decoders, owing to their flexibility in handling variable-length outputs and their strong capability in modeling character-level dependencies, have emerged as the predominant decoding paradigm in the field of scene text recognition (STR). However, AR-based decoders suffer from attention drift, slow decoding speed, and difficulty capturing global dependencies, restricting their performance in various scenarios. In this paper, we propose a novel paradigm for AR-based decoding, called One-Token to Sequence (One2Seq), to address the above issues. Unlike existing methods, we encode the semantic features into a single context token and design a One-Token Wise Decoder to perform the decoding, which alleviates the attention drift caused by the accumulation of semantic information. Moreover, we proposed Positioal-aware Hash Embedding to embed the decoded characters, ensuring the order information is obtained in the context token. By continuously updating this token, One2Seq fully leverages the decoded semantic information while avoiding the computational overhead associated with the growing query sequence. Furthermore, to leverage global information for decoding, we propose Dynamic Global Infusion to dynamically integrates global visual features into the context token. Equipped with the enriched context token, the model has an enhanced ability to extract discriminative local features under the guidance of global context, thereby enhancing recognition accuracy. Extensive experiments reveal that, with its ingenious design, One2Seq exhibits marked superiority on both accuracy and decoding speed compared to existing STR models. Zhibin Ma, Pengwen Dai, Xugong Qin |
AAAI | 2 |
| 2026 | Domain-Complementary Prior With Fine-Grained Feedback for Scene Text Image Super-ResolutionabstractEnhancing the resolution of scene text images is a critical preprocessing step that can substantially improve the accuracy of downstream text recognition in low-quality images. Existing methods primarily rely on auxiliary text features to guide the super-resolution process. However, these features often lack rich low-level information, making them insufficient for faithfully reconstructing both the global structure and fine-grained details of text. Moreover, previous methods often learn suboptimal feature representations from the original low-quality landmark images, which cannot provide precise guidance for super-resolution. In this study, we propose a Fine-Grained Feedback Domain-Complementary Network (FDNet) for scene text image super-resolution. Specifically, we first employ a fine-grained feedback mechanism to selectively refine landmark images, thereby enhancing feature representations. Then, we introduce a novel domain-trace prior interaction generator, which integrates domain-specific traces with a text prior to comprehensively complement the clear edges and structural coverage of the text. Finally, motivated by the limitations of existing datasets, which often exhibit limited scene scales and insufficient challenging scenarios, we introduce a new dataset, MDRText. The proposed dataset, MDRText, features multi-scale and diverse characteristics and is designed to support challenging text image recognition and super-resolution tasks. Extensive experiments on the MDRText and TextZoom datasets demonstrate that our method achieves superior performance in scene text image super-resolution and further improves the accuracy of subsequent recognition tasks. Yang Li 0093, Pengwen Dai, Guotao Xie |
IEEE Trans. Image Process. | 3 |
| 2025 | CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCRabstractScene Text Retrieval (STR) seeks to identify all images containing a given query string. Existing methods typically rely on an explicit Optical Character Recognition (OCR) process of text spotting or localization, which is susceptible to complex pipelines and accumulated errors. To settle this, we resort to the Contrastive Language-Image Pre-training (CLIP) models, which have demonstrated the capacity to perceive and understand scene text, making it possible to achieve strictly OCR-free STR. From the perspective of parameter-efficient transfer learning, a lightweight visual position adapter is proposed to provide a positional information complement for the visual encoder. Besides, we introduce a visual context dropout technique to improve the alignment of local visual features. A novel, parameter-free cross-attention mechanism transfers the contrastive relationship between images and text to that between visual tokens and text, producing a rich cross-modal representation, which can be utilized for efficient reranking with a linear classifier. The resulting model, CAYN, which proves that CLIP is Almost all You Need for STR with no more than 0.50M additional parameters required, achieves new state-of-the-art performance on the STR task, with 92.46%/89.49%/85.98% mAP on the SVT/IIIT-STR/TTR datasets. Our findings demonstrate that CLIP can serve as a reliable and efficient solution for OCR-free STR. Xugong Qin, Peng Zhang 0044, Jun Jie Ou Yang, Gangyan Zeng, Wanqian Zhang, Pengwen Dai |
CVPR | 8 |
| 2025 | Segmentation-Guided Sparse Transformer for Under-Display Camera Image RestorationabstractUnder-display Camera is an emerging technology for full-screen display with a camera under the display. However, the current implementation of UDC causes serious image degradation. Incident light required for camera imaging undergoes attenuation and diffraction when passing through the display. Current UDC image restoration methods predominantly utilize convolutional networks, whereas transformer-based methods with superior performance are lacking. This paper proposes a Segmentation-Guided Sparse Transformer method (SGSFormer) for restoring images from UDC degraded images. Specifically, we utilize sparse self-attention to filter out redundant information and noise, directing the model’s attention to focus on the features more relevant to the degraded regions in need of reconstruction. Moreover, we integrate an instance segmentation map as prior information to guide sparse self-attention in filtering and focusing on the correct regions. Extensive experiments exhibit the superior performance of our model over the state-of-the-art methods. Jingyun Xue, Tao Wang 0052, Pengwen Dai, Kaihao Zhang |
ICASSP | 3 |
| 2025 | Decoupled Graph Energy-based Model for Node Out-of-Distribution Detection on Heterophilic GraphsabstractDespite extensive research efforts focused on Out-of-Distribution (OOD) detection on images, OOD detection on nodes in graph learning remains underexplored. The dependence among graph nodes hinders the trivial adaptation of existing approaches on images that assume inputs to be i.i.d. sampled, since many unique features and challenges specific to graphs are not considered, such as the heterophily issue. Recently, GNNSafe, which considers node dependence, adapted energy-based detection to the graph domain with state-of-the-art performance, however, it has two serious issues: 1) it derives node energy from classification logits without specifically tailored training for modeling data distribution, making it less effective at recognizing OOD data; 2) it highly relies on energy propagation, which is based on homophily assumption and will cause significant performance degradation on heterophilic graphs, where the node tends to have dissimilar distribution with its neighbors. To address the above issues, we suggest training Energy-based Models (EBMs) by Maximum Likelihood Estimation (MLE) to enhance data distribution modeling and removing energy propagation to overcome the heterophily issues. However, training EBMs via MLE requires performing Markov Chain Monte Carlo (MCMC) sampling on both node feature and node neighbors, which is challenging due to the node interdependence and discrete graph topology. To tackle the sampling challenge, we introduce Decoupled Graph Energy-based Model (DeGEM), which decomposes the learning process into two parts—a graph encoder that leverages topology information for node representations and an energy head that operates in latent space. Additionally, we propose a Multi-Hop Graph encoder (MH) and Energy Readout (ERo) to enhance node representation learning, Conditional Energy (CE) for improved EBM training, and Recurrent Update for the graph encoder and energy head to promote each other. This approach avoids sampling adjacency matrices and removes the need for energy propagation to extract graph topology information. Extensive experiments validate that DeGEM, without OOD exposure during training, surpasses previous state-of-the-art methods, achieving an average AUROC improvement of 6.71% on *homophilic* graphs and 20.29% on *heterophilic* graphs, and even outperform methods trained with OOD exposure. Our code is available at: [https://github.com/draym28/DeGEM](https://github.com/draym28/DeGEM). Yuhan Chen 0007, Yihong Luo, Yifan Song 0006, Pengwen Dai, Jing Tang 0004, Xiaochun Cao |
ICLR | 4 |
| 2025 | Endogenous Recovery via Within-modality Prototypes for Incomplete Multimodal HashingabstractMultimodal hashing projects multimodal data into compact binary codes, enabling rapid and storage-efficient retrieval of large-scale multimedia content. In practical scenarios, the issue of missing modality frequently arises when dealing with multimodal data. Existing incomplete multimodal hashing techniques directly recover missing modalities by neural networks, resulting in a disjointed representation space between the recovered and true data. In this paper, we present a novel recovery paradigm, namely Prototype-based Modality Completion Hashing (PMCH). Instead of directly synthesizing it from available modalities, PMCH adaptively aggregates associated within-modality prototypes to recover missing modality data. Specifically, PMCH introduces an within-modality prototype learning module to optimize representative prototypes for each modality. These prototypes act as recovery anchors and reside within the same representation space as their corresponding modality data. Subsequently, PMCH adaptively aggregates the associated within-modality prototypes with coefficients derived from the modality-specific Weight-Net. By utilizing prototypes from the same modality, the semantic disparity between the reconstructed and authentic data can be substantially diminished. Extensive experiments on three widely used benchmark datasets demonstrate that PMCH can effectively recover the missing modality, and attain state-of-the-art performance in both complete and incomplete multimodal retrieval scenarios. Code is available at https://github.com/Sasa77777779/PMCH.git. Sa Zhu, Dayan Wu, Chenming Wu, Pengwen Dai, Bo Li 0063 |
IJCAI | 4 |
| 2025 | Towards Irreversible Attack: Fooling Scene Text Recognition via Multi-Population Coevolution SearchabstractRecent work has shown that scene text recognition (STR) models are vulnerable to adversarial examples.
Different from non-sequential vision tasks, the output sequence of STR models contains rich information.
However, existing adversarial attacks against STR models can only lead to a few incorrect characters in the predicted text.
These attack results still carry partial information about the original prediction and could be easily corrected by an external dictionary or a language model.
Therefore, we propose the Multi-Population Coevolution Search (MPCS) method to attack each character in the image.
We first decompose the global optimization objective into sub-objectives to solve the attack pixel concentration problem existing in previous attack methods.
While this distributed optimization paradigm brings a new joint perturbation shift problem, we propose a novel coevolution energy function to solve it.
Experiments on recent STR models show the superiority of our method.
The code is available at \url{https://github.com/Lee-Jingyu/MPCS}. Pengwen Dai, Mingqing Zhu, Chengwei Wang, Haolong Liu, Xiaochun Cao |
NeurIPS | 2 |
| 2025 | 4K-HAZE: A dehazing benchmark with 4K resolution hazy and haze-free images
Xin Su 0009, Pengwen Dai, Zhuoran Zheng |
Neurocomputing | 2 |
| 2025 | TextSafety: Visual Text Vanishing via Hierarchical Context-Aware Interaction ReconstructionabstractPrivacy information existing in the scene text will be leaked with the spread of images in cyberspace. Vanishing the scene text from the image is a simple yet effective method to prevent privacy disclosure to the machine and the human. Previous visual text vanishing methods have achieved promising results but the performance still fell short of expectations for complicated-shape scene texts with various scales. In this paper, we propose a novel hierarchical context-aware interaction reconstruction method to make the visual text vanish in the natural scene image. To avoid the interference of the non-text regions, we narrow down the reconstruction regions by the guidance of the hierarchical refined text region masks, helping provide accurate position information. Meanwhile, we propose to learn the long-range context-aware interaction in a lightweight way, which can ensure the smoothing of the artifacts that are easily generated by the convolutional layers. To be more specific, we first simultaneously generate the coarse text region mask and the initially vanishing scene text image. Then, we obtain more accurate refined masks to better capture the locations of complicated-shape texts via a hierarchical mask generation network. Next, based on the refined masks, we exploit a channel-wise context-aware interaction mechanism to model the long-range relationships between the reconstruction region and the backgrounds for better removing the artifacts. Finally, we fuse the reconstructed text regions with the non-masked regions to obtain the ultimate protected image. Experiments on two frequently-used benchmarks SCUT-EnsText and SCUT-Syn demonstrate that our proposed method outperforms previous related methods by a large margin. Pengwen Dai, Dayan Wu, Peijia Zheng, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Hybrid Matching Teacher Framework for Cross-Domain Visual Detection TransformerabstractObject detection is a critical component of autonomous vehicle perception systems. However, domain shifts between training environments and real-world scenarios often degrade detector performance. Cross-domain object detection aims to adapt detectors to unlabeled target domains utilizing only labeled source data. Recent popular cross-domain object detection methods employ the mean teacher framework, which uses pseudo-labels generated by the teacher model to guide training on unlabeled real-world data. Despite its effectiveness, continuous training with noisy pseudo-labels leads to abnormal performance degradation in the later stages of training. To address this issue, we propose a novel Hybrid Matching Teacher (HMT) framework for cross-domain visual detection transformers, which enhances cross-domain knowledge transfer across pseudo-label generation, filtering, and training processes. Specifically, we design a Feature Sparse Alignment (FSA) module to adapt DETR tokens and queries, generate domain-adaptive weights to initialize the teacher-student models, and mitigate the inherent initial source bias in the teacher model. Next, a Localization-aware Pseudo-label Filtering (LPF) module ensures high-quality pseudo-labels by considering the consistency between localization and classification tasks. Furthermore, to improve the efficiency of pseudo-label training, the Cross-view Hybrid Matching (CHM) module introduces an auxiliary matching branch to increase the number of positive queries that match with pseudo-labels. Extensive experiments demonstrate that our approach achieves state-of-the-art performance, outperforming previous benchmarks by 3.1%, 8.5%, and 4.4% in adverse weather, diverse scenes, and synthetic-to-real, respectively. Xiaowei Wang 0001, Jinhui Suo, Yang Li 0093, Ming Gao 0012, Peiwen Jiang, Pengwen Dai |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Rethinking Image Deraining via Text-guided Detail ReconstructionabstractImage deraining aims to recover clean images from degradation caused by rain streaks or raindrops of varying intensities. Recently many learning-based approaches have been proposed and achieved promising performance. However, these methods either focus on network architecture design or solely introduce image-level prior to the model. In this paper, we introduce text prior assisting the model in image deraining, as text descriptions have high flexibility and scalability. Text prior provides a wealth of semantic information to help the model achieve more detailed restoration, rather than blindly extrapolating details lost in the degraded image. To this end, we propose a novel image deraining framework based on the transformer, named TGDeraining. Specifically, to incorporate text prior into the framework, we design the Text Prior Embedded Transformer Block (TETB). TETB allows for dynamic guidance of the attention map guided by the text descriptions, thus emphasizing the restoration of critical missing details. The text prior is also fed to the feed-forward network to transform features in a controlled manner. Extensive experimental results demonstrate the effectiveness of our method in restoring a clear image using text as reference information. Chen Wu 0006, Zhuoran Zheng, Pengwen Dai, Chenggang Shan, Xiuyi Jia |
ICME | 3 |
| 2024 | Blind Face Video Restoration with Temporal Consistent Generative Prior and Degradation-Aware PromptabstractWithin the domain of blind face restoration (BFR), approaches lacking facial priors frequently result in excessively smoothed visual outputs. Exiting BFR methods predominantly utilize generative facial priors to achieve realistic and authentic details. However, these methods, primarily designed for images, encounter challenges in maintaining temporal consistency when applied to face video restoration. To tackle this issue, we introduce StableBFVR, an innovative Blind Face Video Restoration method based on Stable Diffusion that incorporates temporal information into the generative prior. This is achieved through the introduction of temporal layers in the diffusion process. These temporal layers consider both long-term and short-term information aggregation. Moreover, to improve generalizability, BFR methods employ complex, large-scale degradation during training, but it often sacrifices accuracy. Addressing this, StableBFVR features a novel mixed-degradation-aware prompt module, capable of encoding specific degradation information to dynamically steer the restoration process. Comprehensive experiments demonstrate that our proposed StableBFVR outperforms state-of-the-art methods. Jingfan Tan, Hyunhee Park, Tao Wang 0052, Kaihao Zhang, Pengwen Dai, Zikun Liu 0001, Wenhan Luo |
ACM Multimedia | 7 |
| 2024 | HIST: Hierarchical and sequential transformer for image captioningabstractAbstract Image captioning aims to automatically generate a natural language description of a given image, and most state‐of‐the‐art models have adopted an encoder–decoder transformer framework. Such transformer structures, however, show two main limitations in the task of image captioning. Firstly, the traditional transformer obtains high‐level fusion features to decode while ignoring other‐level features, resulting in losses of image content. Secondly, the transformer is weak in modelling the natural order characteristics of language. To address theseissues, the authors propose a HI erarchical and S equential T ransformer ( HIST ) structure, which forces each layer of the encoder and decoder to focus on features of different granularities, and strengthen the sequentially semantic information. Specifically, to capture the details of different levels of features in the image, the authors combine the visual features of multiple regions and divide them into multiple levels differently. In addition, to enhance the sequential information, the sequential enhancement module in each decoder layer block extracts different levels of features for sequentially semantic extraction and expression. Extensive experiments on the public datasets MS‐COCO and Flickr30k have demonstrated the effectiveness of our proposed method, and show that the authors’ method outperforms most of previous state of the arts. Feixiao Lv, Rui Wang 0032, Lihua Jing, Pengwen Dai |
IET Comput. Vis. | 4 |
| 2024 | Granularity-Aware Single-Point Scene Text Spotting With Sequential Recurrence Self-AttentionabstractScene text spotting, a unified framework between text detection and text recognition, has made great progress in recent years. Existing methods usually adopt the fully-supervised learning strategy, which relies on time-consuming location annotations, particularly for scene texts with arbitrary shapes. In this paper, we propose a weakly-supervised scene text spotting method via the location labels of single points with the corresponding text transcriptions. Due to the weak location annotations for challenging scene texts, previous weakly-supervised methods adopting the convolution neural network structure make it hard to model the different-scale text feature representations under blurring or nosing scenarios. In addition, as the single-point location can only cover part of the text instance, it will burden the confusion of sequential-like scene text recognition. To address these issues, we present a novel sequential recurrence self-attention for granularity-aware single-point scene text spotting. Specifically, we first enhance the scene text feature representations with different scales by integrating the global intra-interaction of high-level features with the low-level local features. Then, based on the granularity-aware text features, we decode them into text transcriptions in the sequential recurrence self-attention manner to capture the sequence-dependent relation in character-level semantics and locations. Extensive experiments show that our proposed method outperforms existing state-of-the-art weakly-supervised scene text spotters by a large margin. Xunquan Tong, Pengwen Dai, Xugong Qin, Rui Wang 0032, Wenqi Ren |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Explicitly-Decoupled Text Transfer With Minimized Background Reconstruction for Scene Text EditingabstractScene text editing aims to replace the source text with the target text while preserving the original background. Its practical applications span various domains, such as data generation and privacy protection, highlighting its increasing importance in recent years. In this study, we propose a novel Scene Text Editing network with Explicitly-decoupled text transfer and Minimized background reconstruction, called STEEM. Unlike existing methods that usually fuse text style, text content, and background, our approach focuses on decoupling text style and content from the background and utilizes the minimized background reconstruction to reduce the impact of text replacement on the background. Specifically, the text-background separation module predicts the text mask of the scene text image, separating the source text from the background. Subsequently, the style-guided text transfer decoding module transfers the geometric and stylistic attributes of the source text to the content text, resulting in the target text. Next, the background and target text are combined to determine the minimal reconstruction area. Finally, the context-focused background reconstruction module is applied to the reconstruction area, producing the editing result. Furthermore, to ensure stable joint optimization of the four modules, a task-adaptive training optimization strategy has been devised. Experimental evaluations conducted on two popular datasets demonstrate the effectiveness of our approach. STEEM outperforms state-of-the-art methods, as evidenced by a reduction in the FID index from 29.48 to 24.67 and an increase in text recognition accuracy from 76.8% to 78.8%. Jianqun Zhou, Pengwen Dai, Yang Li 0093, Manjiang Hu, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2023 | Hi-SIGIR: Hierachical Semantic-Guided Image-to-image Retrieval via Scene GraphabstractImage-to-image retrieval, a fundamental task, aims at matching similar images based on a query image. Existing methods with convolutional neural networks are usually sensitive to low-level visual features, and ignore high-level semantic relationship information. This makes retrieving complicated images with multiple objects and various relationships a significant challenge. Although some works introduce the scene graph to capture the global semantic features of the objects and their relations, they ignore the local visual representations. In addition, due to the fragility of individual modal representations, poisoning attacks in adversarial scenarios are easily achieved, hurting the robustness of the visual-guided foundation image retrieval model. To overcome these issues, we propose a novel hierarchical semantic-guided image-to-image retrieval method via scene graph, called Hi-SIGIR. Specifically, to begin with, our proposed method generates the scene graph of an image. Then, our model extracts and learns both the visual and semantic features of the nodes and relations within the scene graphs. Next, these features are fused to obtain local information and sent to the graph neural network to obtain global information. Using these information, the similarity between the scene graphs of several images is calculated at both the local and global levels to perform image retrieval. Finally, we introduce a surrogate that calculates relevance in a cross-modal manner to understand image content better. Experimental evaluations on several wildly-used benchmarks demonstrate the superiority of the proposed method. Yulu Wang, Pengwen Dai, Xiaojun Jia, Zhitao Zeng, Rui Li 0109, Xiaochun Cao |
ACM Multimedia | 2 |
| 2023 | Privacy-Enhancing Face Obfuscation Guided by Semantic-Aware Attribution MapsabstractFace recognition technology is increasingly being integrated into our daily life, e.g. Face ID. With the advancement of machine learning algorithms, the personal information such as age, gender, and race can be easily deduced from the recorded face images in these applications. This poses a serious privacy threat to individuals who do not want to be profiled, as face images are collected for biometric purposes. Existing methods mostly focus on adding the invisible adversarial perturbations into the images to make automatic inference infeasible. However, the application scenarios of these methods are limited due to the perturbations depending on the specific model. In this paper, we introduce a novel face privacy-enhancing framework by obfuscating the stored faces, which could maintain the data utility (face identity) while protecting the privacy of users (facial attributes). Specifically, we first develop a feature attribution module to discover the identity-related facial parts. Within this module, we introduce a pixel importance estimation model based on Shapley value to obtain a pixel-level attribution map, and then each pixel on the attribution map is aggregated into semantic facial parts, which are used to quantify the importance of different facial parts. Next, we design a privacy-enhancing module to generate the high-quality obfuscated images, which can modify the privacy semantic content and preserve the identity-related information. Using the proposed method, users can choose the single or multiple attributes to be obfuscated without affecting identity matching. Extensive experiments conducted on CelebA-HQ and VGGFace2-HQ benchmarks demonstrate the effectiveness and generalization ability of our method. Jingzhi Li 0002, Hua Zhang 0008, Siyuan Liang 0004, Pengwen Dai, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | The Best Protection is Attack: Fooling Scene Text Recognition With Minimal PixelsabstractScene text recognition (STR) has witnessed tremendous progress in the era of deep learning, but it also raises concerns about privacy infringement as scene texts usually contain valuable or sensitive information. Previous works in privacy protection of scene texts mainly focus on masking out the texts from the image/video. In this work, we learn from the idea of adversarial examples and use minimal pixel perturbation to protect the privacy of text information. Although there are well-established attacking methods on non-sequential vision tasks (e.g., classification), the attack on sequential tasks (e.g., scene text recognition) has not received sufficient attention yet. Moreover, existing works mainly focus on the white-box setting, which requires complete knowledge of the target model (e.g., architecture, parameters, or gradients). These requirements limit the scope of applications for the white-box adversarial attack. Therefore, we propose a novel black-box attacking approach for the STR models, only requiring prior knowledge of the model output. Besides, instead of disturbing most pixels as in existing STR attack methods, our proposed approach only manipulates a few pixels, meaning the perturbation is more inconspicuous. To determine the location and value of the manipulated pixels, we also provide an efficient Adaptive-Discrete Differential Evolution (AD$^{2}\text{E}$) by narrowing down the continuous searching space to a discrete space. It can greatly reduce the queries to the target model. Experiments on several real-world benchmarks show the effectiveness of our proposed approach. Especially, when attacking the commercial STR engine, Baidu-OCR, our method achieves higher attack success rates by a large margin than existing approaches. Our work establishes an important step towards using the black-box adversarial attack with minimal pixels to protect the privacy of text information from being easily obtained by STR models. Yikun Xu, Pengwen Dai, Zekun Li 0007, Hongjun Wang 0005, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Cognition Guided Human-Object Relationship DetectionabstractHuman-object relationship detection reveals the fine-grained relationship between humans and objects, helping the comprehensive understanding of videos. Previous human-object relationship detection approaches are mainly developed with object features and relation features without exploring the specific information of humans. In this paper, we propose a novel Relation-Pose Transformer (RPT) for human-object relationship detection. Inspired by the coordination of eye-head-body movements in cognitive science, we employ the head pose to find those crucial objects that humans focus on and use the body pose with skeleton information to represent multiple actions. Then, we utilize the spatial encoder to capture spatial contextualized information of the relation pair, which integrates the relation features and pose features. Next, the temporal decoder aims to model the temporal dependency of the relationship. Finally, we adopt multiple classifiers to predict different types of relationships. Extensive experiments on the benchmark Action Genome validate the effectiveness of our proposed method and show the state-of-the-art performance compared with related methods. Zhitao Zeng, Pengwen Dai, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2023 | A2SC: Adversarial Attacks on Subspace ClusteringabstractMany studies demonstrate that supervised learning techniques are vulnerable to adversarial examples. However, adversarial threats in unsupervised learning have not drawn sufficient scholarly attention. In this article, we formally address the unexplored adversarial attacks in the equally important unsupervised clustering field and propose the concept of the adversarial set and adversarial set attack for clustering. To illustrate the basic idea, we design a novel adversarial space-mapping attack algorithm to confuse subspace clustering, one of the mainstream branches of unsupervised clustering. It maps a sample into one wrong class by moving it towards the closest point on the linear subspace of the target class, that is, along the normal of the closest point. This simple single-step algorithm has the power to craft the adversarial set where the image samples can be wrongly clustered, even into the targeted labels. Empirical results on different image datasets verify the effectiveness and superiority of our algorithm. We further show that deep supervised learning algorithms (such as VGG and ResNet) are also vulnerable to our crafted adversarial set, which illustrates the good cross-task transferability of the adversarial set. Yikun Xu, Xingxing Wei 0001, Pengwen Dai, Xiaochun Cao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Electrical-STGCN: An Electrical Spatio-Temporal Graph Convolutional Network for Intelligent Predictive MaintenanceabstractWith the rapid improvement of Industrial Internet of Things and artificial intelligence, predictive maintenance (PdM) has attracted great attention from both academia and industrial practitioners. When equipment is running, the electrical attributes have intrinsic relations. Meanwhile, they are changing over time. However, existing PdM models are often limited as they lack considering both attribute interactions and temporal dependence of the dynamic working system. To address the problem, in this article, we propose an electrical spatio-temporal graph convolutional network (Electrical-STGCN) for PdM. First, it takes a sequence of electrical records as input. Next, both attribute interactions and temporal dependence are established to extract features. Then, the extracted features are fed into a prediction component. Finally, the output of the Electrical-STGCN (i.e., remaining useful life) can help the workers decide whether to carry out equipment maintenance. The effectiveness of the proposed method is verified in real-world cases. Our method achieves 85.2% Accuracy and 0.9 F1-Score, which are better than the other approaches. Yuchen Jiang 0004, Pengwen Dai, Pengcheng Fang, Ray Y. Zhong, Xiaochun Cao |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | ACE: Anchor-Free Corner Evolution for Real-Time Arbitrarily-Oriented Object DetectionabstractObjects with different orientations are ubiquitous in the real world (e.g., texts/hands in the scene image, objects in the aerial image, etc.), and the widely-used axis-aligned bounding box does not compactly enclose the oriented objects. Thus arbitrarily-oriented object detection has attracted rising attention in recent years. In this paper, we propose a novel and effective model to detect arbitrarily-oriented objects. Instead of directly predicting the angles of oriented bounding boxes like most existing methods, we evolve the axis-aligned bounding box to the oriented quadrilateral box with the assistance of dynamically gathering contour information. More specifically, we first obtain the axis-aligned bounding box in an anchor-free manner. After that, we set the key points based on the sampled contour points of the axis-aligned bounding box. To improve the localization performance, we enrich the feature representations of these key points by exploiting a dynamic information gathering mechanism. This technique propagates the geometrical and semantic information along the sampled contour points, and fuses the information from the semantic neighbors of each sampled point, which varies for different locations. Finally, we estimate the offsets between the axis-aligned bounding box key points and the oriented quadrilateral box corner points. Extensive experiments on two frequently-used aerial image benchmarks HRSC2016 and DOTA, as well as scene text/hand datasets ICDAR2015, TD500, and Oxford-Hand, demonstrate the effectiveness and advantage of our proposed model. Pengwen Dai, Siyuan Yao, Zekun Li 0007, Sanyi Zhang, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2022 | Accurate Scene Text Detection Via Scale-Aware Data Augmentation and Shape Similarity ConstraintabstractScene text detection has attracted increasing concerns with the rapid development of deep neural networks in recent years. However, existing scene text detectors may overfit on the public datasets due to the limited training data, or generate inaccurate localization for arbitrary-shape scene texts. This paper presents an arbitrary-shape scene text detection method that can achieve better generalization ability and more accurate localization. We first propose a Scale-Aware Data Augmentation (SADA) technique to increase the diversity of training samples. SADA considers the scale variations and local visual variations of scene texts, which can effectively relieve the dilemma of limited training data. At the same time, SADA can enrich the training minibatch, which contributes to accelerating the training process. Furthermore, a Shape Similarity Constraint (SSC) technique is exploited to model the global shape structure of arbitrary-shape scene texts and backgrounds from the perspective of the loss function. SSC encourages the segmentation of text or non-text in the candidate boxes to be similar to the corresponding ground truth, which is helpful to localize more accurate boundaries for arbitrary-shape scene texts. Extensive experiments have demonstrated the effectiveness of the proposed techniques, and state-of-the-art performances are achieved over public arbitrary-shape scene text benchmarks (e.g.,CTW1500,Total-TextandArT). Pengwen Dai, Yang Li 0093, Hua Zhang 0008, Jingzhi Li 0002, Xiaochun Cao |
IEEE Trans. Multim. | 1 |
| 2021 | Updated Paired Regions for Shadow Detection from Single Image
Xiao Wang 0017, Siyuan Yao, Pengwen Dai, Rui Wang 0032, Xiaochun Cao |
BMVC | 3 |
| 2021 | Progressive Contour Regression for Arbitrary-Shape Scene Text DetectionabstractState-of-the-art scene text detection methods usually model the text instance with local pixels or components from the bottom-up perspective and, therefore, are sensitive to noises and dependent on the complicated heuristic post-processing especially for arbitrary-shape texts. To relieve these two issues, instead, we propose to progressively evolve the initial text proposal to arbitrarily shaped text contours in a top-down manner. The initial horizontal text proposals are generated by estimating the center and size of texts. To reduce the range of regression, the first stage of the evolution predicts the corner points of oriented text proposals from the initial horizontal ones. In the second stage, the contours of the oriented text proposals are iteratively regressed to arbitrarily shaped ones. In the last iteration of this stage, we rescore the confidence of the final localized text by utilizing the cues from multiple contour points, rather than the single cue from the initial horizontal proposal center that may be out of arbitrary-shape text regions. Moreover, to facilitate the progressive contour evolution, we design a contour information aggregation mechanism to enrich the feature representation on text contours by considering both the circular topology and semantic context. Experiments conducted on CTW1500, Total-Text, ArT, and TD500 have demonstrated that the proposed method especially excels in line-level arbitrary-shape texts. Code is available at https://github.com/dpengwen/PCR. Pengwen Dai, Sanyi Zhang, Hua Zhang 0008, Xiaochun Cao |
CVPR | 1 |
| 2021 | Less Is Better: Fooling Scene Text Recognition with Minimal Perturbations
Yikun Xu, Pengwen Dai, Xiaochun Cao |
ICONIP (6) | 2 |
| 2021 | SLOAN: Scale-Adaptive Orientation Attention Network for Scene Text RecognitionabstractScene text recognition, the final step of the scene text reading system, has made impressive progress based on deep neural networks. However, existing recognition methods devote to dealing with the geometrically regular or irregular scene text. They are limited to the semantically arbitrary-orientation scene text. Meanwhile, previous scene text recognizers usually learn the single-scale feature representations for various-scale characters, which cannot model effective contexts for different characters. In this paper, we propose a novel scale-adaptive orientation attention network for arbitrary-orientation scene text recognition, which consists of a dynamic log-polar transformer and a sequence recognition network. Specifically, the dynamic log-polar transformer learns the log-polar origin to adaptively convert the arbitrary rotations and scales of scene texts into the shifts in the log-polar space, which is helpful to generate the rotation-aware and scale-aware visual representation. Next, the sequence recognition network is an encoder-decoder model, which incorporates a novel character-level receptive field attention module to encode more valid contexts for various-scale characters. The whole architecture can be trained in an end-to-end manner, only requiring the word image and its corresponding ground-truth text. Extensive experiments on several public datasets have demonstrated the effectiveness and superiority of our proposed method. Pengwen Dai, Hua Zhang 0008, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2020 | Deep Multi-Scale Context Aware Feature Aggregation for Curved Scene Text DetectionabstractScene text plays a significant role in image and video understanding, which has made great progress in recent years. Most existing models on text detection in the wild have the assumption that all the texts are surrounded by a rotated rectangle or quadrangle. While there also exist lots of curved texts in the wild, which would not be bounded by a regular bounding box. In this paper, we develop a novel architecture to localize the text regions, which can deal with curved-shape scene texts. Specifically, we first design a text-related feature enhancement module by incorporating the prior knowledge of the text shape to enhance the feature representations. After that, based on the enhanced features, we employ a region proposal network to generate the candidate boxes of scene texts. For each text candidate, a pyramid region-of-interest pooling attention module is utilized to extract the fixed-size features. Finally, we exploit the box-aware context-based text segmentation module and box refinement network to obtain the location of scene text. Experiments are conducted on four challenging benchmarks CTW1500, totalTEXT, ICDAR-2015 and MLT, and the experimental results have demonstrated the superiority of our model. Pengwen Dai, Hua Zhang 0008, Xiaochun Cao |
IEEE Trans. Multim. | 1 |
| 2019 | Pedestrian Trajectory Prediction with Learning-based Approaches: A Comparative StudyabstractTo enable safe and efficient navigations through the urban environment, autonomous vehicles need to anticipate the future motions of the walking pedestrians who might collide with them. The dynamic and stochastic behavior characteristics of the pedestrians make the trajectory prediction challengeable for most kinematics-based approaches. This paper presents a comparative study of six state-of-the-art learning-based methods for pedestrian trajectory prediction, including Gaussian Process (GP), LSTM, GP-LSTM, Character-based LSTM, Sequence-to-Sequence (Seq2Seq), and attention-based Seq2Seq. The trajectory prediction is formulated as the regression task or sequence generation problem that predicts future trajectories based on observed trajectories. We evaluate the performance of the learning-based methods on a public real-world pedestrian dataset. To address the concern of data scarcity, we employ three forms of data augmentation (i.e., translation, rotation, and stretch) to enlarge the dataset, which produce the transformed trajectories from the original trajectories. By comparison, those learning-based approaches are ranked based on prediction accuracy from high to low as Seq2Seq, attention-based Seq2Seq, C-LSTM, LSTM, GP, and GP-LSTM. Particularly, Seq2Seq model outperforms all baseline approaches, with the mean and final point errors less than 15cm in normal scenarios when predicting 1s ahead. Yang Li 0093, Long Xin, Dameng Yu, Pengwen Dai, Jianqiang Wang 0003, Shengbo Eben Li |
IV | 4 |
| 2017 | Deep Strip-Based Network with Cascade Learning for Scene Text LocalizationabstractScene text detection is currently a popular research topic in the computer vision community. However, it is a challenging task due to the variations of texts and clutter backgrounds. In this paper, we propose a novel framework for scene text localization. Based on the region proposal network, a Strip-based Text Detection Network (STDN) is developed with vertical anchor mechanism to predict the text/non-text strip-shaped proposals. Meanwhile, we incorporate the recurrent neural network layers in the proposed network to refine the predicted results. Specifically, hard example mining is performed to train the STDN with cascade learning, which has a remarkable improvement in precision. Besides, we exploit a clustering algorithm to generate anchor dimensions spontaneously without hand-picking, which is portable and time-saving. The text detection framework achieves the state-of-the-art performance on ICDAR2013 with 0.89 F-measure. Dao Wu, Rui Wang 0032, Pengwen Dai, Yueying Zhang, Xiaochun Cao |
ICDAR | 3 |