Anna Zhu

dblp:135/5039 · DBLP profile ↗
← Back
43ranked-venue papers
13as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Zero-shot Chinese Character Recognition with Radical Tree Positional Embedding and spatial alignment
Anna Zhu, Hongyi Cai
Eng. Appl. Artif. Intell.1
2026 FFTDiff: Tuning-free image texture transfer based on diffusion model
Shilin Li, Anna Zhu
J. Vis. Commun. Image Represent.3
2026 Promptable Selective Scene Text Removal
Anna Zhu, Seiichi Uchida
Pattern Recognit.2
2026 Cross-modal local and global alignment for Chinese Character Recognition
Anna Zhu, Hongyi Cai
Pattern Recognit.2
2025 Fusing Text Semantics for Text Image Inpainting
Conghao Han, Mengyang Duan, Anna Zhu
ICDAR (5)3
2025 FSTDiff: One-Shot Font Generation via Cross-Font Style Transformation Learning
Shilin Li, Anna Zhu
ICDAR (1)2
2025 MambaSTR: Scene Text Recognition with Masked State Space Model
Anna Zhu
ICIC (1)2
2025 High Resolution Wire Segmentation with Domain Adaption
abstract
Wires and power lines in images, despite occupying a very small fraction of the space, have a significant impact on aesthetics. To segment and remove this visual element automatically is challenging due to the unsatisfactory generalization of segmentation models and the untouchability of the real data. In this paper, we propose a novel wire segmentation network, named Global-Local-Residual Network (GLRNet), fusing multi-resolution features and edge-based residual information, which could capture both long-range context and preserve fine structural details for high-resolution images. Since the wire labeling data is particularly hard to acquire and the annotations by humans are often scarce, we synthesize a WireSyn dataset and propose to train the model in an unsupervised domain adaptation (UDA) manner. Extensive experiments demonstrate the effectiveness of our proposed strategies for precise wire segmentation leveraging both global and local information in high-resolution images. Our method achieves state-of-the-art performance on both synthesis and real datasets.
Anna Zhu
ICME3
2025 Edge guided and Fourier attention-based Dual Interaction Network for scene text erasing
Anna Zhu
Image Vis. Comput.2
2024 Controllable Text Layout Generation For Synthesizing Scene Text Image
Huen Chen, Jiangyang He, Anna Zhu
ICDAR (5)3
2024 Cross-Modal Alignment of Local and Global Features for Zero-Shot Chinese Character Recognition
abstract
Chinese character recognition (CCR) is a pivotal domain in computer vision due to its complexity and diverse applications, especially given the extensive character categories posing challenges in identifying unseen characters. Addressing the zero-shot hurdle, we propose a CLIP-style model, which independently extracts features from aligned Chinese character images and Ideographic Description Sequences (IDS), achieving cross-modal alignment. Our approach encompasses local and global feature alignment. Initially, we introduce learnable discrete tokens to represent shared embeddings for visual and textual modalities, capturing the local context of Chinese characters. Then, encoding each radical extracts local features, mapped to shared discrete tokens via attention mechanisms. Additionally, encoding the entire character obtains global features. Training utilizes contrastive loss to facilitate cross-modal alignment. Experimental results confirm our method’s superiority over conventional approaches, demonstrating remarkable performance on zero-shot Chinese character recognition benchmarks.
Hongyi Cai, Anna Zhu
ICIP2
2024 Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text Recognition
abstract
Inducing linguistic knowledge for scene text recognition (STR) is a new trend that could provide semantics for performance boost. However, most autoregressive STR models optimize one-step ahead prediction (i.e., 1-gram prediction) for character sequence, which only utilizes the previous semantic context. Most non-autoregressive models only apply linguistic knowledge individually on the output sequence to refine the results in parallel, which do not fully utilize the visual clues concurrently. In this paper, we propose a novel language-based STR model, called ProphetSTR. It adopts an n-stream attention mechanism in the decoder to simultaneously predict the next n characters based on the previous predictions at each time step. It behaves like a prophet, encouraging the model to predict more accurate results by utilizing the previous semantic information and the near future clues. If the prediction results for the same character at successive time steps are inconsistent, we should not trust any of them. Otherwise, they are reliable predictions. Therefore, we propose a multi-modality verification module, masking the unreliable semantic features and inputting with visual and trusted semantic ones simultaneously for masked prediction recovery in parallel. It learns to align different modalities implicitly and considers both visual context and linguistic knowledge, which could generate more reliable results. Furthermore, we propose a multi-scale weight-sharing encoder for multi-granularity image representation. Extensive experiments demonstrate that ProphetSTR achieves state-of-the-art performances on many benchmarks. Further ablative studies prove the effectiveness of our proposed components.
Anna Zhu
ACM Multimedia1
2024 Improving Scene Text Recognition with Counting-Aware Contrastive Learning and Attention Alignment
Anna Zhu
PRCV (7)3
2024 OCR-Aware Scene Graph Generation Via Multi-modal Object Representation Enhancement and Logical Bias Learning
Zihan Ji, Anna Zhu
PRCV (7)3
2024 Scene text recognition via dual character counting-aware visual and semantic modeling network
Anna Zhu, Brian Kenji Iwana
Sci. China Inf. Sci.2
2023 Few shot font generation via transferring similarity guided global style and quantization local style
abstract
Automatic few-shot font generation (AFFG), aiming at generating new fonts with only a few glyph references, reduces the labor cost of manually designing fonts. However, the traditional AFFG paradigm of style-content disentanglement cannot capture the diverse local details of different fonts. So, many component-based approaches are proposed to tackle this problem. The issue with component-based approaches is that they usually require special pre-defined glyph components, e.g., strokes and radicals, which is infeasible for AFFG of different languages. In this paper, we present a novel font generation approach by aggregating styles from character similarity-guided global features and stylized component-level representations. We calculate the similarity scores of the target character and the referenced samples by measuring the distance along the corresponding channels from the content features, and assigning them as the weights for aggregating the global style features. To better capture the local styles, a cross-attention-based style transfer module is adopted to transfer the styles of reference glyphs to the components, where the components are self-learned discrete latent codes through vector quantization without manual definition. With these designs, our AFFG method could obtain a complete set of component-level style representations, and also control the global glyph characteristics. The experimental results reflect the effectiveness and generalization of the proposed method on different linguistic scripts, and also show its superiority when compared with other state-of-the-art methods. The source code can be found at https://github.com/awei669/VQ-Font.
Anna Zhu, Brian Kenji Iwana, Shilin Li
ICCV2
2023 Scene Text Involved "Text"-to-Image Retrieval through Logically Hierarchical Matching
abstract
Text-to-image retrieval, one of the most important cross-modality tasks, aims to search the most relevant images through a given text query. Most recent approaches are based on large-scale models. The huge time costs make it impossible for real-time searching. They also ignore the fine-grained information, i.e., the scene text. To tackle these issues, we propose a novel matching method that considers the scene text in both modalities and adopts a fast matching way by aligning from objects to relations and finally to the global. This logically hierarchical process emulates the way humans understand information. To better implement our method, we relabel the TextCaps-OCR dataset, which contains 110K captions with word-level POS labeling and 22K corresponding scene text images with bounding boxes. Extensive experiments demonstrate the superiority and efficiency of our method, whose performance is significantly higher than the past SOTA on both OCR-contained and OCR-free datasets.
Anna Zhu, Huen Chen
ICME2
2023 HelixNet: Dual Helix Cooperative Decoders for Scene Text Removal
Guangtao Lyu, Anna Zhu
PRCV (7)3
2023 FETNet: Feature erasing and transferring network for scene text removal
Guangtao Lyu, Kun Liu 0027, Anna Zhu, Seiichi Uchida, Brian Kenji Iwana
Pattern Recognit.3
2023 Corrigendum to "FETNet: Feature Erasing and Transferring Network for Scene Text Removal": Pattern Recognition Volume 140 (2023) 109531
Guangtao Lyu, Kun Liu 0027, Anna Zhu, Seiichi Uchida, Brian Kenji Iwana
Pattern Recognit.3
2022 PSSTRNet: Progressive Segmentation-Guided Scene Text Removal Network
abstract
Scene text removal (STR) is a challenging task due to the complex text fonts, colors, sizes, and background textures in scene images. However, most previous methods learn both text location and background inpainting implicitly within a single network, which weakens the text localization mecha-nism and makes a lossy background. To tackle these prob-lems, we propose a simple Progressive Segmentation-guided Scene Text Removal Network(PSSTRNet) to remove the text in the image iteratively. It contains two decoder branches, a text segmentation branch, and a text removal branch, with a shared encoder. The text segmentation branch generates text mask maps as the guidance for the regional removal branch. In each iteration, the original image, previous text removal result, and text mask are input to the network to extract the rest part of the text segments and cleaner text removal result. To get a more accurate text mask map, an update module is developed to merge the mask map in the current and previous stages. The final text removal result is obtained by adaptive fusion of results from all previous stages. A sufficient number of experiments and ablation studies conducted on the real and synthetic public datasets demonstrate our proposed method achieves state-of-the-art performance.
Guangtao Lyu, Anna Zhu
ICME2
2022 Sketch Image Style Transfer Based on Sketch Density Controlling
Anna Zhu
ICONIP (7)3
2022 Find More Accurate Text Boundary for Scene Text Detection
abstract
Arbitrary shape text detection is still an open problem due to several challenges, e.g., dense text adhesion, various sizes, and noises. In this work, we propose a novel scene text detection network by extracting and connecting a set of points on the boundary of each text instance. Our method mainly consists of three branches, including text center line (TCL) prediction, text orientation (TO) prediction, and text boundary offset (TBO) prediction to get the text boundary point proposals. Utilize these proposals, and the final refinement results can be obtained by point sampling and graph attention network(GAT). The detector can overcome the text instance sticking problem with these text boundary representations. Additionally, we propose distance-based dice loss and instance-aware L1loss to remove false positives and get over various text sizes on text boundary offset prediction respectively. In this way, our method can directly and efficiently generate accurate text boundaries without any post-processing. Extensive experiments on publicly available datasets show the effectiveness of our design and training strategy, which also demonstrates our method’s state-of-the-art performance for arbitrary shape text detection.
Runqiu Pan, Zezhou Li, Anna Zhu
ICPR3
2022 Text Style Transfer based on Multi-factor Disentanglement and Mixture
abstract
Text style transfer aims to transfer the reference style of one text image to another text image. Previous works have only been able to transfer the style to a binary text image. In this paper, we propose a framework to disentangle the text images into three factors: text content, font, and style features, and then remix the factors of different images to transfer a new style. Both the reference and input text images have no style restrictions. Adversarial training through multi-factor cross recognition is adopted in the network for better feature disentanglement and representation. To decompose the input text images into a disentangled representation with swappable factors, the network is trained using similarity mining within pairs of exemplars. To train our model, we synthesized a new dataset with various text styles in both English and Chinese. Several ablation studies and extensive experiments on our designed and public datasets demonstrate the effectiveness of our approach for text style transfer.
Anna Zhu, Zhanhui Yin, Brian Kenji Iwana, Shengwu Xiong 0001
ACM Multimedia1
2022 Disentangled OCR: A More Granular Information for "Text"-to-Image Retrieval
Shilin Li, Huen Chen, Anna Zhu
PRCV (1)4
2021 Multi-document detection via corner localization and association
Runqiu Pan, Anna Zhu
Neurocomputing2
2020 Light Textspotter: An Extreme Light Scene Text Spotter
Jiazhi Guan, Anna Zhu
ICONIP (4)2
2020 Scene Text Detection with Selected Anchors
abstract
Object proposal technique with dense anchoring scheme for scene text detection were applied frequently to achieve high recall. It results in the significant improvement in accuracy but waste of computational searching, regression and classification. In this paper, we propose an anchor selection-based region proposal network (AS-RPN) using effective selected anchors instead of dense anchors to extract text proposals. The center, scales, aspect ratios and orientations of anchors are learnable instead of fixing, which leads to high recall and greatly reduced numbers of anchors. By replacing the anchor-based RPN in Faster RCNN, the AS-RPN-based Faster RCNN can achieve comparable performance with previous state-of-the-art text detecting approaches on standard benchmarks, including COCO-Text, ICDAR2013, ICDAR2015 and MSRA-TD500 when using single-scale and single model (ResNet50) testing only.
Anna Zhu, Shengwu Xiong 0001
ICPR1
2020 Zero-Shot Learning Based on Salient Region and Enhanced Semantics
Zongrong Pan, Anna Zhu
PRCV (2)2
2020 Few-Shot Text Style Transfer via Deep Feature Similarity
abstract
Generating text to have a consistent style with only a few observed highly-stylized text samples is a difficult task for image processing. The text style involving the typography, i.e., font, stroke, color, decoration, effects, etc., should be considered for transfer. In this paper, we propose a novel approach to stylize target text by decoding weighted deep features from only a few referenced samples. The deep features, including content and style features of each referenced text, are extracted from a Convolutional Neural Network (CNN) that is optimized for character recognition. Then, we calculate the similarity scores of the target text and the referenced samples by measuring the distance along the corresponding channels from the content features of the CNN when considering only the content, and assign them as the weights for aggregating the deep features. To enforce the stylized text to be realistic, a discriminative network with adversarial loss is employed. We demonstrate the effectiveness of our network by conducting experiments on three different datasets which have various styles, fonts, languages, etc. Additionally, the coefficients for character style transfer, including the character content, the effect of similarity matrix, the number of referenced characters, the similarity between characters, and performance evaluation by a new protocol are analyzed for better understanding our proposed framework.
Anna Zhu, Xiongbo Lu, Xiang Bai, Seiichi Uchida, Brian Kenji Iwana, Shengwu Xiong 0001
IEEE Trans. Image Process.1
2019 Scene Text Magnifier
abstract
Scene text magnifier aims to magnify text in natural scene images without recognition. It could help the special groups, who have myopia or dyslexia to better understand the scene. In this paper, we design the scene text magnifier through interacted four CNN-based networks: character erasing, character extraction, character magnify, and image synthesis. The architecture of the networks are extended based on the hourglass encoder-decoders. It inputs the original scene text image and outputs the text magnified image while keeps the background unchange. Intermediately, we can get the side-output results of text erasing and text extraction. The four sub-networks are first trained independently and fine-tuned in end-to-end mode. The training samples for each stage are processed through a flow with original image and text annotation in ICDAR2013 and Flickr dataset as input, and corresponding text erased image, magnified text annotation, and text magnified scene image as output. To evaluate the performance of text magnifier, the Structural Similarity is used to measure the regional changes in each character region. The experimental results demonstrate our method can magnify scene text effectively without effecting the background.
Toshiki Nakamura, Anna Zhu, Seiichi Uchida
ICDAR2
2019 Character Image Synthesis Based on Selected Content and Referenced Style Embedding
abstract
Arbitrary characters synthesis based on a few referenced examples poses a great challenge due to the diversity of characters category and style. We regard this problem as image translation problem and propose a character style transfer network consisting of content selector, style encoder, content encoder, feature embedding and embedded feature decoder to solve it. The content selector is used to select and match the most similar content (i.e., font) from our collected glyph dataset as content references. Then, we apply the style encoder and content encoder to extract the style and content representation separately and mix them for feature embedding. Finally, the embedded features are decoded to generate the target characters. We train them in an end-to-end manner and evaluate the proposed method on MC-GAN dataset and our collected dataset. The experimental results have demonstrated the effectiveness of the proposed model for character synthesis.
Anna Zhu, Xiongbo Lu, Shengwu Xiong 0001
ICME1
2019 Scene word recognition from pieces to whole
Anna Zhu, Seiichi Uchida
Frontiers Comput. Sci.1
2019 Coarse-to-fine document localization in natural scene image with regional attention and recursive corner refinement
Anna Zhu, Shengwu Xiong 0001
Int. J. Document Anal. Recognit.1
2018 Multi-frame Quantization of LSF Parameters Using a Deep Autoencoder and Pyramid Vector Quantizer
Yaxing Li, Eshete Derb Emiru, Shengwu Xiong 0001, Anna Zhu, Pengfei Duan 0005, Yichang Li
INTERSPEECH4
2018 Multi-frame Coding of LSF Parameters Using Block-Constrained Trellis Coded Vector Quantization
Yaxing Li, Shan Xu 0007, Shengwu Xiong 0001, Anna Zhu, Pengfei Duan 0005, Yueming Ding
INTERSPEECH4
2017 Scene Text Eraser
abstract
The character information in natural scene images contains various personal information, such as telephone numbers, home addresses, etc. It is a high risk of leakage the information if they are published. In this paper, we proposed a scene text erasing method to properly hide the information via an inpainting convolutional neural network (CNN) model. The input is a scene text image, and the output is expected to be text erased image with all the character regions filled up the colors of the surrounding background pixels. This work is accomplished by a CNN model through convolution to deconvolution with interconnection process. The training samples and the corresponding inpainting images are considered as teaching signals for training. To evaluate the text erasing performance, the output images are detected by a novel scene text detection method. Subsequently, the same measurement on text detection is utilized for testing the images in benchmark dataset ICDAR2013. Compared with direct text detection way, the scene text erasing process demonstrates a drastically decrease on the precision, recall and f-score. That proves the effectiveness of proposed method for erasing the text in natural scene images.
Toshiki Nakamura, Anna Zhu, Keiji Yanai, Seiichi Uchida
ICDAR2
2017 Scene Text Relocation with Guidance
abstract
Applying object proposal technique for scene text detection becomes popular for its significant improvement in speed and accuracy for object detection. However, some of the text regions after the proposal classification are overlapped and hard to remove or merge. In this paper, we present a scene text relocation system that refines the detection from text proposals to text. An object proposal-based deep neural network is employed to get the text proposals. To tackle the detection overlapping problem, a refinement deep neural network relocates the overlapped regions by estimating the text probability inside, and locating the accurate text regions by thresholding. Since the space between words indifferent text lines are various, a guidance mechanism is proposed in text relocation to guide where to extract the text regions in word level. This refinement procedure helps boost the precision after removing multiple overlapped text regions or joint cracked text regions. The experimental results on standard benchmark ICDAR 2013 demonstrate the effectiveness of the proposed approach.
Anna Zhu, Seiichi Uchida
ICDAR1
2016 A Further Step to Perfect Accuracy by Training CNN with Larger Data
abstract
Convolutional Neural Networks (CNN) are on the forefront of accurate character recognition. This paper explores CNNs at their maximum capacity by implementing the use of large datasets. We show a near-perfect performance by using a dataset of about 820,000 real samples of isolated handwritten digits, much larger than the conventional MNIST database. In addition, we report a near-perfect performance on the recognition of machine-printed digits and multi-font digital born digits. Also, in order to progress toward a universal OCR, we propose methods of combining the datasets into one classifier. This paper reveals the effects of combining the datasets prior to training and the effects of transfer learning during training. The results of the proposed methods also show an almost perfect accuracy suggesting the ability of the network to generalize all forms of text.
Seiichi Uchida, Shouta Ide, Brian Kenji Iwana, Anna Zhu
ICFHR4
2016 Could scene context be beneficial for scene text detection?
Anna Zhu, Renwu Gao, Seiichi Uchida
Pattern Recognit.1
2015 Recognizing perspective scene text with context feature
abstract
Text recognition has gained significant attention from the computer vision community. Correct character recognition is the premise of text recognition and affects the overall performance to large extent. This paper proposes a novel character representation for scene text recognition. First, a context-based feature that contains local information and relevant key points' feature is extracted from key points. The relativity is measured by the distance of vector that is generated by a trained Gaussian Mixture Model (GMM) between the target key point and other key points in each context bin. In order to recognize each individual character, we adopt a bag-of-words approach, in which the rotation-invariant context features are densely extracted from an individual character. All key points' context features are prone to build a vocabulary of visual words by using k-means clustering. Then we train a set of two-class linear Support Vector Machines in a one-vs-all schema for each category character. By using densely extracted context features that are rotation-invariant and efficient, our method is capable of recognizing perspective texts of arbitrary orientations. The evaluation results on benchmark datasets demonstrate that our proposed scheme of scene character recognition is highly efficient and achieves state-of-the-art performance on not only fontal character recognition but also perspective characters'.
Anna Zhu, Yangbo Dong, Guoyou Wang
ICDAR1
2015 Detecting natural scenes text via auto image partition, two-stage grouping and two-layer classification
Anna Zhu, Guoyou Wang, Yangbo Dong
Pattern Recognit. Lett.1
2013 Detection of wide linear structures by fusion of width and gray
abstract
Lines provide important information in images and line detection is crucial in many applications. Many line features can be used to detect line position while line width (i.e., thickness) is a more structured, higher-level feature compared to edge or other line features. Every point of the wide line structure has its own width in spite of the structure's asymmetry. In this paper, we use the parallel edges to get the line orientation and width. We do not need to know the actual width of the line, but rather recover it to get the region of interest (ROI). Then use a modified OTSU obtaining proper region threshold T to segment the wide line structures. By fusing feature of width and gray, a sequence of tests has been conducted on a variety of image samples obtained from simple natural scene and our experimental results demonstrate the practical and robust of the proposed method.
Anna Zhu, Guoyou Wang, Ran Wang 0005
ICME1