EDBT 2026 Demo / reviewers in the wild / expert
Zhaoyi Wan
dblp:227/2593
· DBLP profile ↗
7ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0002-1994-260XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Image recognition and object detection · 50% Segmentation and scene understanding · 42% Language models and text generation · 8% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 67% Visual content generation and editing · 33% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
scene text recognition |
1.2 | 3 | 2020 | On Vocabulary Reliance in Scene Text Recognition · CVPR 2020 TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020 Scene Text Recognition from Two-Dimensional Perspective · AAAI 2019 |
Computer vision › Image recognition and object detection
scene text detection |
1.1 | 2 | 2023 | Real-Time Scene Text Detection With Differentiable Binarization and Adaptive Scale Fusion · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Real-Time Scene Text Detection with Differentiable Binarization · AAAI 2020 |
Computer vision › Segmentation and scene understanding
segmentation-based text detection |
1.1 | 2 | 2023 | Real-Time Scene Text Detection With Differentiable Binarization and Adaptive Scale Fusion · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Real-Time Scene Text Detection with Differentiable Binarization · AAAI 2020 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.8 | 2 | 2020 | TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020 Scene Text Recognition from Two-Dimensional Perspective · AAAI 2019 |
Multimedia analysis and retrieval
image retrieval |
0.6 | 1 | 2022 | Cloud2Sketch: Augmenting Clouds with Imaginary Sketches · ACM Multimedia 2022 |
Multimedia analysis and retrieval › image retrieval
sketch-based retrieval |
0.6 | 1 | 2022 | Cloud2Sketch: Augmenting Clouds with Imaginary Sketches · ACM Multimedia 2022 |
Visual content generation and editing
sketch generation |
0.6 | 1 | 2022 | Cloud2Sketch: Augmenting Clouds with Imaginary Sketches · ACM Multimedia 2022 |
Natural language and speech › Language models and text generation › decoding
attention-based decoding |
0.4 | 1 | 2020 | On Vocabulary Reliance in Scene Text Recognition · CVPR 2020 |
Computer vision › Segmentation and scene understanding › image segmentation › document image segmentation
character segmentation |
0.4 | 1 | 2020 | TextScanner: Reading Characters in Order for Robust Scene Text Recognition · AAAI 2020 |
Methods — techniques the papers use, named apart from their topics
attention mechanism · 0.8differentiable binarization · 0.7adaptive scale fusion · 0.7self-supervised learning · 0.6free-form deformation · 0.6segmentation network · 0.4recurrent neural network · 0.4parallel prediction · 0.4mutual learning · 0.4differentiable binarization module · 0.4fully convolutional network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Real-Time Scene Text Detection With Differentiable Binarization and Adaptive Scale FusionabstractRecently, segmentation-based scene text detection methods have drawn extensive attention in the scene text detection field, because of their superiority in detecting the text instances of arbitrary shapes and extreme aspect ratios, profiting from the pixel-level descriptions. However, the vast majority of the existing segmentation-based approaches are limited to their complex post-processing algorithms and the scale robustness of their segmentation models, where the post-processing algorithms are not only isolated to the model optimization but also time-consuming and the scale robustness is usually strengthened by fusing multi-scale feature maps directly. In this paper, we propose a Differentiable Binarization (DB) module that integrates the binarization process, one of the most important steps in the post-processing procedure, into a segmentation network. Optimized along with the proposed DB module, the segmentation network can produce more accurate results, which enhances the accuracy of text detection with a simple pipeline. Furthermore, an efficient Adaptive Scale Fusion (ASF) module is proposed to improve the scale robustness by fusing features of different scales adaptively. By incorporating the proposed DB and ASF with the segmentation network, our proposed scene text detector consistently achieves state-of-the-art results, in terms of both detection accuracy and speed, on five standard benchmarks. Minghui Liao, Zhisheng Zou, Zhaoyi Wan, Cong Yao, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Cloud2Sketch: Augmenting Clouds with Imaginary SketchesabstractHave you ever looked up at the sky and imagined what the clouds look like? In this work, we present an interesting task that augments clouds in the sky with imagined sketches. Different from generic image-to-sketch translation tasks, unique challenges are introduced: real-world clouds have different levels of similarity to something; sketch generation without sketch retrieval could lead to something unrecognizable; a retrieved sketch from some dataset cannot be directly used because of the mismatch of the shape; an optimal sketch imagination is subjective. We propose Cloud2Sketch, a novel self-supervised pipeline to tackle the aforementioned challenges. First, we pre-process cloud images with a cloud detector and a thresholding algorithm to obtain cloud contours. Then, cloud contours are passed through a retrieval module to retrieve sketches with similar geometrical shapes. Finally, we adopt a novel sketch translation model with built-in free-form deformation for aligning the sketches to cloud contours. To facilitate training, an icon-based sketch collection named Sketchy Zoo is proposed. Extensive experiments validate the effectiveness of our method both qualitatively and quantitatively. Zhaoyi Wan, Dejia Xu, Zhangyang Wang, Jiebo Luo 0001 |
ACM Multimedia | 1 |
| 2022 | Facial Attribute Transformers for Precise and Robust Makeup TransferabstractIn this paper, we address the problem of makeup transfer, which aims at transplanting the makeup from the reference face to the source face while preserving the identity of the source. Existing makeup transfer methods have made notable progress in generating realistic makeup faces, but do not perform well in terms of color fidelity and spatial transformation. To tackle these issues, we propose a novel Facial Attribute Transformer (FAT) and its variant Spatial FAT for high-quality makeup transfer. Drawing inspirations from the Transformer in NLP, FAT is able to model the semantic correspondences and interactions between the source face and reference face, and then precisely estimate and transfer the facial attributes. To further facilitate shape deformation and transformation of facial parts, we also integrate thin plate splines (TPS) into FAT, thus creating Spatial FAT, which is the first method that can transfer geometric attributes in addition to color and texture. Extensive qualitative and quantitative experiments demonstrate the effectiveness and superiority of our proposed FATs in the following aspects: (1) ensuring high-fidelity color transfer; (2) allowing for geometric transformation of facial parts; (3) handling facial variations (such as poses and shadows) and (4) supporting high-resolution face generation. Zhaoyi Wan, Jie An 0002, Cong Yao, Jiebo Luo 0001 |
WACV | 1 |
| 2020 | Real-Time Scene Text Detection with Differentiable BinarizationabstractRecently, segmentation-based methods are quite popular in scene text detection, as the segmentation results can more accurately describe scene text of various shapes such as curve text. However, the post-processing of binarization is essential for segmentation-based detection, which converts probability maps produced by a segmentation method into bounding boxes/regions of text. In this paper, we propose a module named Differentiable Binarization (DB), which can perform the binarization process in a segmentation network. Optimized along with a DB module, a segmentation network can adaptively set the thresholds for binarization, which not only simplifies the post-processing but also enhances the performance of text detection. Based on a simple segmentation network, we validate the performance improvements of DB on five benchmark datasets, which consistently achieves state-of-the-art results, in terms of both detection accuracy and speed. In particular, with a light-weight backbone, the performance improvements by DB are significant so that we can look for an ideal tradeoff between detection accuracy and efficiency. Specifically, with a backbone of ResNet-18, our detector achieves an F-measure of 82.8, running at 62 FPS, on the MSRA-TD500 dataset. Code is available at: https://github.com/MhLiao/DB. Minghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen 0006, Xiang Bai |
AAAI | 2 |
| 2020 | TextScanner: Reading Characters in Order for Robust Scene Text RecognitionabstractDriven by deep learning and a large volume of data, scene text recognition has evolved rapidly in recent years. Formerly, RNN-attention-based methods have dominated this field, but suffer from the problem of attention drift in certain situations. Lately, semantic segmentation based algorithms have proven effective at recognizing text of different forms (horizontal, oriented and curved). However, these methods may produce spurious characters or miss genuine characters, as they rely heavily on a thresholding procedure operated on segmentation maps. To tackle these challenges, we propose in this paper an alternative approach, called TextScanner, for scene text recognition. TextScanner bears three characteristics: (1) Basically, it belongs to the semantic segmentation family, as it generates pixel-wise, multi-channel segmentation maps for character class, position and order; (2) Meanwhile, akin to RNN-attention-based methods, it also adopts RNN for context modeling; (3) Moreover, it performs paralleled prediction for character position and class, and ensures that characters are transcripted in the correct order. The experiments on standard benchmark datasets demonstrate that TextScanner outperforms the state-of-the-art methods. Moreover, TextScanner shows its superiority in recognizing more difficult text such as Chinese transcripts and aligning with target characters. Zhaoyi Wan, Minghang He, Xiang Bai, Cong Yao |
AAAI | 1 |
| 2020 | On Vocabulary Reliance in Scene Text RecognitionabstractThe pursuit of high performance on public benchmarks has been the driving force for research in scene text recognition, and notable progresses have been achieved. However, a close investigation reveals a startling fact that the state-of-the-art methods perform well on images with words within vocabulary but generalize poorly to images with words outside vocabulary. We call this phenomenon ``vocabulary reliance''. In this paper, we establish an analytical framework, in which different datasets, metrics and module combinations for quantitative comparisons are devised, to conduct an in-depth study on the problem of vocabulary reliance in scene text recognition. Key findings include: (1) Vocabulary reliance is ubiquitous, i.e., all existing algorithms more or less exhibit such characteristic; (2) Attention-based decoders prove weak in generalizing to words outside vocabulary and segmentation-based decoders perform well in utilizing visual features; (3) Context modeling is highly coupled with the prediction layers. These findings provide new insights and can benefit future research in scene text recognition. Furthermore, we propose a simple yet effective mutual learning strategy to allow models of two families (attention-based and segmentation-based) to learn collaboratively. This remedy alleviates the problem of vocabulary reliance and significantly improves the overall scene text recognition performance. Zhaoyi Wan, Jielei Zhang, Jiebo Luo 0001, Cong Yao |
CVPR | 1 |
| 2019 | Scene Text Recognition from Two-Dimensional PerspectiveabstractInspired by speech recognition, recent state-of-the-art algorithms mostly consider scene text recognition as a sequence prediction problem. Though achieving excellent performance, these methods usually neglect an important fact that text in images are actually distributed in two-dimensional space. It is a nature quite different from that of speech, which is essentially a one-dimensional signal. In principle, directly compressing features of text into a one-dimensional form may lose useful information and introduce extra noise. In this paper, we approach scene text recognition from a two-dimensional perspective. A simple yet effective model, called Character Attention Fully Convolutional Network (CA-FCN), is devised for recognizing the text of arbitrary shapes. Scene text recognition is realized with a semantic segmentation network, where an attention mechanism for characters is adopted. Combined with a word formation module, CA-FCN can simultaneously recognize the script and predict the position of each character. Experiments demonstrate that the proposed algorithm outperforms previous methods on both regular and irregular text datasets. Moreover, it is proven to be more robust to imprecise localizations in the text detection phase, which are very common in practice. Minghui Liao, Zhaoyi Wan, Fengming Xie, Jiajun Liang, Pengyuan Lv, Cong Yao, Xiang Bai |
AAAI | 3 |