Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhaoyi Wan

dblp:227/2593 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0002-1994-260XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Image recognition and object detection · 50% Segmentation and scene understanding · 42% Language models and text generation · 8%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 67% Visual content generation and editing · 33%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
scene text recognition
1.232020
On Vocabulary Reliance in Scene Text Recognition · CVPR 2020
TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020
Scene Text Recognition from Two-Dimensional Perspective · AAAI 2019
Computer vision › Image recognition and object detection
scene text detection
1.122023
Real-Time Scene Text Detection With Differentiable Binarization and Adaptive Scale Fusion · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Real-Time Scene Text Detection with Differentiable Binarization · AAAI 2020
Computer vision › Segmentation and scene understanding
segmentation-based text detection
1.122023
Real-Time Scene Text Detection With Differentiable Binarization and Adaptive Scale Fusion · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Real-Time Scene Text Detection with Differentiable Binarization · AAAI 2020
Computer vision › Segmentation and scene understanding
semantic segmentation
0.822020
TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020
Scene Text Recognition from Two-Dimensional Perspective · AAAI 2019
Multimedia analysis and retrieval
image retrieval
0.612022
Cloud2Sketch: Augmenting Clouds with Imaginary Sketches · ACM Multimedia 2022
Multimedia analysis and retrieval › image retrieval
sketch-based retrieval
0.612022
Cloud2Sketch: Augmenting Clouds with Imaginary Sketches · ACM Multimedia 2022
Visual content generation and editing
sketch generation
0.612022
Cloud2Sketch: Augmenting Clouds with Imaginary Sketches · ACM Multimedia 2022
Natural language and speech › Language models and text generation › decoding
attention-based decoding
0.412020
On Vocabulary Reliance in Scene Text Recognition · CVPR 2020
Computer vision › Segmentation and scene understanding › image segmentation › document image segmentation
character segmentation
0.412020
TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020

Methods — techniques the papers use, named apart from their topics

attention mechanism · 0.8differentiable binarization · 0.7adaptive scale fusion · 0.7self-supervised learning · 0.6free-form deformation · 0.6segmentation network · 0.4recurrent neural network · 0.4parallel prediction · 0.4mutual learning · 0.4differentiable binarization module · 0.4fully convolutional network · 0.4
YearPublicationVenuePosition
2023 Real-Time Scene Text Detection With Differentiable Binarization and Adaptive Scale Fusion
abstract
Recently, segmentation-based scene text detection methods have drawn extensive attention in the scene text detection field, because of their superiority in detecting the text instances of arbitrary shapes and extreme aspect ratios, profiting from the pixel-level descriptions. However, the vast majority of the existing segmentation-based approaches are limited to their complex post-processing algorithms and the scale robustness of their segmentation models, where the post-processing algorithms are not only isolated to the model optimization but also time-consuming and the scale robustness is usually strengthened by fusing multi-scale feature maps directly. In this paper, we propose a Differentiable Binarization (DB) module that integrates the binarization process, one of the most important steps in the post-processing procedure, into a segmentation network. Optimized along with the proposed DB module, the segmentation network can produce more accurate results, which enhances the accuracy of text detection with a simple pipeline. Furthermore, an efficient Adaptive Scale Fusion (ASF) module is proposed to improve the scale robustness by fusing features of different scales adaptively. By incorporating the proposed DB and ASF with the segmentation network, our proposed scene text detector consistently achieves state-of-the-art results, in terms of both detection accuracy and speed, on five standard benchmarks.
Minghui Liao, Zhisheng Zou, Zhaoyi Wan, Cong Yao, Xiang Bai
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Cloud2Sketch: Augmenting Clouds with Imaginary Sketches
abstract
Have you ever looked up at the sky and imagined what the clouds look like? In this work, we present an interesting task that augments clouds in the sky with imagined sketches. Different from generic image-to-sketch translation tasks, unique challenges are introduced: real-world clouds have different levels of similarity to something; sketch generation without sketch retrieval could lead to something unrecognizable; a retrieved sketch from some dataset cannot be directly used because of the mismatch of the shape; an optimal sketch imagination is subjective. We propose Cloud2Sketch, a novel self-supervised pipeline to tackle the aforementioned challenges. First, we pre-process cloud images with a cloud detector and a thresholding algorithm to obtain cloud contours. Then, cloud contours are passed through a retrieval module to retrieve sketches with similar geometrical shapes. Finally, we adopt a novel sketch translation model with built-in free-form deformation for aligning the sketches to cloud contours. To facilitate training, an icon-based sketch collection named Sketchy Zoo is proposed. Extensive experiments validate the effectiveness of our method both qualitatively and quantitatively.
Zhaoyi Wan, Dejia Xu, Zhangyang Wang, Jiebo Luo 0001
ACM Multimedia1
2022 Facial Attribute Transformers for Precise and Robust Makeup Transfer
abstract
In this paper, we address the problem of makeup transfer, which aims at transplanting the makeup from the reference face to the source face while preserving the identity of the source. Existing makeup transfer methods have made notable progress in generating realistic makeup faces, but do not perform well in terms of color fidelity and spatial transformation. To tackle these issues, we propose a novel Facial Attribute Transformer (FAT) and its variant Spatial FAT for high-quality makeup transfer. Drawing inspirations from the Transformer in NLP, FAT is able to model the semantic correspondences and interactions between the source face and reference face, and then precisely estimate and transfer the facial attributes. To further facilitate shape deformation and transformation of facial parts, we also integrate thin plate splines (TPS) into FAT, thus creating Spatial FAT, which is the first method that can transfer geometric attributes in addition to color and texture. Extensive qualitative and quantitative experiments demonstrate the effectiveness and superiority of our proposed FATs in the following aspects: (1) ensuring high-fidelity color transfer; (2) allowing for geometric transformation of facial parts; (3) handling facial variations (such as poses and shadows) and (4) supporting high-resolution face generation.
Zhaoyi Wan, Jie An 0002, Cong Yao, Jiebo Luo 0001
WACV1
2020 Real-Time Scene Text Detection with Differentiable Binarization
abstract
Recently, segmentation-based methods are quite popular in scene text detection, as the segmentation results can more accurately describe scene text of various shapes such as curve text. However, the post-processing of binarization is essential for segmentation-based detection, which converts probability maps produced by a segmentation method into bounding boxes/regions of text. In this paper, we propose a module named Differentiable Binarization (DB), which can perform the binarization process in a segmentation network. Optimized along with a DB module, a segmentation network can adaptively set the thresholds for binarization, which not only simplifies the post-processing but also enhances the performance of text detection. Based on a simple segmentation network, we validate the performance improvements of DB on five benchmark datasets, which consistently achieves state-of-the-art results, in terms of both detection accuracy and speed. In particular, with a light-weight backbone, the performance improvements by DB are significant so that we can look for an ideal tradeoff between detection accuracy and efficiency. Specifically, with a backbone of ResNet-18, our detector achieves an F-measure of 82.8, running at 62 FPS, on the MSRA-TD500 dataset. Code is available at: https://github.com/MhLiao/DB.
Minghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen 0006, Xiang Bai
AAAI2
2020 TextScanner: Reading Characters in Order for Robust Scene Text Recognition
abstract
Driven by deep learning and a large volume of data, scene text recognition has evolved rapidly in recent years. Formerly, RNN-attention-based methods have dominated this field, but suffer from the problem of attention drift in certain situations. Lately, semantic segmentation based algorithms have proven effective at recognizing text of different forms (horizontal, oriented and curved). However, these methods may produce spurious characters or miss genuine characters, as they rely heavily on a thresholding procedure operated on segmentation maps. To tackle these challenges, we propose in this paper an alternative approach, called TextScanner, for scene text recognition. TextScanner bears three characteristics: (1) Basically, it belongs to the semantic segmentation family, as it generates pixel-wise, multi-channel segmentation maps for character class, position and order; (2) Meanwhile, akin to RNN-attention-based methods, it also adopts RNN for context modeling; (3) Moreover, it performs paralleled prediction for character position and class, and ensures that characters are transcripted in the correct order. The experiments on standard benchmark datasets demonstrate that TextScanner outperforms the state-of-the-art methods. Moreover, TextScanner shows its superiority in recognizing more difficult text such as Chinese transcripts and aligning with target characters.
Zhaoyi Wan, Minghang He, Xiang Bai, Cong Yao
AAAI1
2020 On Vocabulary Reliance in Scene Text Recognition
abstract
The pursuit of high performance on public benchmarks has been the driving force for research in scene text recognition, and notable progresses have been achieved. However, a close investigation reveals a startling fact that the state-of-the-art methods perform well on images with words within vocabulary but generalize poorly to images with words outside vocabulary. We call this phenomenon ``vocabulary reliance''. In this paper, we establish an analytical framework, in which different datasets, metrics and module combinations for quantitative comparisons are devised, to conduct an in-depth study on the problem of vocabulary reliance in scene text recognition. Key findings include: (1) Vocabulary reliance is ubiquitous, i.e., all existing algorithms more or less exhibit such characteristic; (2) Attention-based decoders prove weak in generalizing to words outside vocabulary and segmentation-based decoders perform well in utilizing visual features; (3) Context modeling is highly coupled with the prediction layers. These findings provide new insights and can benefit future research in scene text recognition. Furthermore, we propose a simple yet effective mutual learning strategy to allow models of two families (attention-based and segmentation-based) to learn collaboratively. This remedy alleviates the problem of vocabulary reliance and significantly improves the overall scene text recognition performance.
Zhaoyi Wan, Jielei Zhang, Jiebo Luo 0001, Cong Yao
CVPR1
2019 Scene Text Recognition from Two-Dimensional Perspective
abstract
Inspired by speech recognition, recent state-of-the-art algorithms mostly consider scene text recognition as a sequence prediction problem. Though achieving excellent performance, these methods usually neglect an important fact that text in images are actually distributed in two-dimensional space. It is a nature quite different from that of speech, which is essentially a one-dimensional signal. In principle, directly compressing features of text into a one-dimensional form may lose useful information and introduce extra noise. In this paper, we approach scene text recognition from a two-dimensional perspective. A simple yet effective model, called Character Attention Fully Convolutional Network (CA-FCN), is devised for recognizing the text of arbitrary shapes. Scene text recognition is realized with a semantic segmentation network, where an attention mechanism for characters is adopted. Combined with a word formation module, CA-FCN can simultaneously recognize the script and predict the position of each character. Experiments demonstrate that the proposed algorithm outperforms previous methods on both regular and irregular text datasets. Moreover, it is proven to be more robust to imprecise localizations in the text detection phase, which are very common in practice.
Minghui Liao, Zhaoyi Wan, Fengming Xie, Jiajun Liang, Pengyuan Lv, Cong Yao, Xiang Bai
AAAI3