EDBT 2026 Demo / reviewers in the wild / expert
Fan Bai 0001
dblp:84/4809-1
· DBLP profile ↗
9ranked-venue papers
1as first author
4since 2021 · last 2022
0000-0002-4139-0653ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Image recognition and object detection · 74% Deep learning architectures and training · 23% Language models and text generation · 3% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 57% Image and video processing · 43% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 77% Recommender systems · 12% Web and social media mining · 12% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
scene text recognition |
1.5 | 4 | 2022 | C3-STISR: Scene Text Image Super-resolution with Triple Clues · IJCAI 2022 AON: Towards Arbitrarily-Oriented Text Recognition · CVPR 2018 Edit Probability for Scene Text Recognition · CVPR 2018 |
Computer vision › Image recognition and object detection
text recognition |
0.6 | 1 | 2022 | C3-STISR: Scene Text Image Super-resolution with Triple Clues · IJCAI 2022 |
Image and video processing › super-resolution › image super-resolution
scene text image super-resolution |
0.6 | 1 | 2022 | C3-STISR: Scene Text Image Super-resolution with Triple Clues · IJCAI 2022 |
Image and video processing
super-resolution |
0.6 | 1 | 2022 | C3-STISR: Scene Text Image Super-resolution with Triple Clues · IJCAI 2022 |
Data mining › structured data mining
graph mining |
0.5 | 1 | 2021 | Label Propagation on K-Partite Graphs with Heterophily · IEEE Trans. Knowl. Data Eng. 2021 |
Data mining › semi-supervised learning
label propagation |
0.5 | 1 | 2021 | Label Propagation on K-Partite Graphs with Heterophily · IEEE Trans. Knowl. Data Eng. 2021 |
Multimedia analysis and retrieval › video summarization
video highlight detection |
0.5 | 1 | 2021 | GIF Thumbnails: Attract More Clicks to Your Videos · AAAI 2021 |
Multimedia analysis and retrieval
video summarization |
0.5 | 1 | 2021 | GIF Thumbnails: Attract More Clicks to Your Videos · AAAI 2021 |
Multimedia analysis and retrieval › video summarization
video thumbnail generation |
0.5 | 1 | 2021 | GIF Thumbnails: Attract More Clicks to Your Videos · AAAI 2021 |
Computer vision › Image recognition and object detection › scene text recognition
arbitrarily-oriented text recognition |
0.3 | 1 | 2018 | AON: Towards Arbitrarily-Oriented Text Recognition · CVPR 2018 |
Machine learning › Deep learning architectures and training › encoder-decoder architecture
attention-based encoder-decoder |
0.3 | 1 | 2018 | Edit Probability for Scene Text Recognition · CVPR 2018 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning |
0.3 | 1 | 2018 | Edit Probability for Scene Text Recognition · CVPR 2018 |
Recommender systems
click-through rate prediction |
0.1 | 1 | 2021 | GIF Thumbnails: Attract More Clicks to Your Videos · AAAI 2021 |
Web and social media mining
social network analysis |
0.1 | 1 | 2021 | Label Propagation on K-Partite Graphs with Heterophily · IEEE Trans. Knowl. Data Eng. 2021 |
Natural language and speech › Language models and text generation › decoding
attention-based decoding |
0.1 | 1 | 2018 | AON: Towards Arbitrarily-Oriented Text Recognition · CVPR 2018 |
Machine learning › Deep learning architectures and training
encoder-decoder architecture |
0.1 | 1 | 2017 | Focusing Attention: Towards Accurate Text Recognition in Natural Images · ICCV 2017 |
Methods — techniques the papers use, named apart from their topics
recognizer feedback · 1.1cross-modal clue fusion · 1.1character-level language model · 1.1variational autoencoder · 1.0learning-based generation · 1.0k-partite graph model · 0.5incremental algorithm · 0.5dual-encoder · 0.5dual encoder · 0.5end-to-end training · 0.3edit probability · 0.3attention-based decoder · 0.3attention mechanism · 0.3resnet · 0.3focusing attention mechanism · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | C3-STISR: Scene Text Image Super-resolution with Triple CluesabstractScene text image super-resolution (STISR) has been regarded as an important pre-processing task for text recognition from low-resolution scene text images. Most recent approaches use the recognizer's feedback as clues to guide super-resolution. However, directly using recognition clue has two problems: 1) Compatibility. It is in the form of probability distribution, has an obvious modal gap with STISR - a pixel-level task; 2) Inaccuracy. it usually contains wrong information, thus will mislead the main task and degrade super-resolution performance. In this paper, we present a novel method C3-STISR that jointly exploits the recognizer's feedback, visual and linguistical information as clues to guide super-resolution. Here, visual clue is from the images of texts predicted by the recognizer, which is informative and more compatible with the STISR task; while linguistical clue is generated by a pre-trained character-level language model, which is able to correct the predicted texts. We design effective extraction and fusion mechanisms for the triple cross-modal clues to generate a comprehensive and unified guidance for super-resolution. Extensive experiments on TextZoom show that C3-STISR outperforms the SOTA methods in fidelity and recognition performance. Code is available in https://github.com/zhaominyiz/C3-STISR. Minyi Zhao, Fan Bai 0001, Bingjia Li, Shuigeng Zhou |
IJCAI | 3 |
| 2022 | Robustly Recognizing Irregular Scene Text by Rectifying Principle IrregularitiesabstractReading irregular scene text is a challenging problem in scene text recognition. Rectification is a popular measure to reduce irregularities of text in images. Existing rectification methods seek to rectify text images into a strictly regular form via free parametric transformation functions. However, they always suffer from information loss or severe deformation due to their poor constraints to the transformation functions. In our investigation, we found that CNN and attention are robust to many slight irregularities. What inspires us to propose a novel and effective rectification method that mainly rectifies the principle regularities, and leaves the slight irregularities to the CNN-LSTM-attention recognizer. Our rectification method first estimates the character densities and directions of the input image in a down-sampled map then finds a best fitting curve from a small predefined Bézier curve set, and finally rectifies the input image with a transformation function corresponding to the selected curve. Transformation functions are carefully designed so that they neither lose important visual information nor cause severe deformation. Extensive experiments on seven benchmark datasets show that our method achieves the state of the art performance in most cases, especially in curved text recognition. Changsheng Xu, Yang Wang 0100, Fan Bai 0001, Jihong Guan, Shuigeng Zhou |
WACV | 3 |
| 2021 | GIF Thumbnails: Attract More Clicks to Your VideosabstractWith the rapid increase of mobile devices and online media, more and more people prefer posting/viewing videos online. Generally, these videos are presented on video streaming sites with image thumbnails and text titles. While facing huge amounts of videos, a viewer clicks through a certain video with high probability because of its eye-catching thumbnail. However, current video thumbnails are created manually, which is time-consuming and quality-unguaranteed. And static image thumbnails contain very limited information of the corresponding videos, which prevents users from successfully clicking what they really want to view. In this paper, we address a novel problem, namely GIF thumbnail generation, which aims to automatically generate GIF thumbnails for videos and consequently boost their Click-Through-Rate (CTR). Here, a GIF thumbnail is an animated GIF file consisting of multiple segments from the video, containing more information of the target video than a static image thumbnail. To support this study, we build the first GIF thumbnails benchmark dataset that consists of 1070 videos covering a total duration of 69.1 hours, and 5394 corresponding manually-annotated GIFs. To solve this problem, we propose a learning-based automatic GIF thumbnail generation model, which is called Generative Variational Dual-Encoder (GEVADEN). As not relying on any user interaction information (e.g. time-sync comments and real-time view counts), this model is applicable to newly-uploaded/rarely-viewed videos. Experiments on our built dataset show that GEVADEN significantly outperforms several baselines, including video-summarization and highlight-detection based ones. Furthermore, we develop a pilot application of the proposed model on an online video platform with 9814 videos covering 1231 hours, which shows that our model achieves a 37.5% CTR improvement over traditional image thumbnails. This further validates the effectiveness of the proposed model and the promising application prospect of GIF thumbnails. Yi Xu 0003, Fan Bai 0001, Yingxuan Shi, Qiuyu Chen, Longwen Gao, Kai Tian 0001, Shuigeng Zhou, Huyang Sun |
AAAI | 2 |
| 2021 | Label Propagation on K-Partite Graphs with HeterophilyabstractIn this paper, for the first time, we study label propagation in heterogeneous graphs under heterophily assumption. Homophily label propagation (i.e., two connected nodes share similar labels) in homogeneous graph (with same types of vertices and relations) has been extensively studied before. Unfortunately, real-life networks (e.g., social networks) are heterogeneous, they contain different types of vertices (e.g., users, images, and texts) and relations (e.g., friendships and co-tagging) and allow for each node to propagate both the same and opposite copy of labels to its neighbors. We propose a IC-partite label propagation model to handle the mystifying combination of heterogeneous nodes/relations and heterophily propagation. With this model, we develop a novel label inference algorithm framework with update rules in near-linear time complexity. Since real networks change overtime, we devise an incremental approach, which supports fast updates for both new data and evidence (e.g., ground truth labels) with guaranteed efficiency. We further provide a utility function to automatically determine whether an incremental or a re-modeling approach is favored. Extensive experiments on real datasets have verified the effectiveness and efficiency of our approach, and its superiority over the state-of-the-art label propagation methods. Dingxiong Deng, Fan Bai 0001, Yiqi Tang, Shuigeng Zhou, Cyrus Shahabi, Linhong Zhu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Text Recognition in Real Scenarios with a Few Labeled SamplesabstractScene text recognition (STR) is still a hot research topic in computer vision field due to its various applications. Existing works mainly focus on learning a general model with a huge number of synthetic text images to recognize unconstrained scene texts, and have achieved substantial progress. However, these methods are not quite applicable in many real-world scenarios where 1) high recognition accuracy is required, while 2) labeled samples are lacked. To tackle this challenging problem, this paper proposes a few-shot adversarial sequence domain adaptation (FASDA) approach to build sequence adaptation between the synthetic source domain (with many synthetic labeled samples) and a specific target domain (with only some or a few real labeled samples). This is done by simultaneously learning each character's feature representation with an attention mechanism and establishing the corresponding character-level latent subspace with adversarial learning. Our approach can maximize the character-level confusion between the source domain and the target domain, thus achieves the sequence-level adaptation with even a small number of labeled samples in the target domain. Extensive experiments on various datasets show that our method significantly outperforms the finetuning scheme, and obtains comparable performance to the state-of-the-art STR methods. Jinghuang Lin, Zhanzhan Cheng, Fan Bai 0001, Shiliang Pu, Shuigeng Zhou |
ICPR | 3 |
| 2020 | Recognizing Multiple Text Sequences from an Image by Pure End-to-End Learning
Zhenlong Xu, Shuigeng Zhou, Fan Bai 0001, Zhanzhan Cheng, Shiliang Pu |
ICPR | 3 |
| 2018 | Edit Probability for Scene Text RecognitionabstractWe consider the scene text recognition problem under the attention-based encoder-decoder framework, which is the state of the art. The existing methods usually employ a frame-wise maximal likelihood loss to optimize the models. When we train the model, the misalignment between the ground truth strings and the attention's output sequences of probability distribution, which is caused by missing or superfluous characters, will confuse and mislead the training process, and consequently make the training costly and degrade the recognition accuracy. To handle this problem, we propose a novel method called edit probability (EP) for scene text recognition. EP tries to effectively estimate the probability of generating a string from the output sequence of probability distribution conditioned on the input image, while considering the possible occurrences of missing/superfluous characters. The advantage lies in that the training process can focus on the missing, superfluous and unrecognized characters, and thus the impact of the misalignment problem can be alleviated or even overcome. We conduct extensive experiments on standard benchmarks, including the IIIT-5K, Street View Text and ICDAR datasets. Experimental results show that the EP can substantially boost scene text recognition performance. Fan Bai 0001, Zhanzhan Cheng, Shiliang Pu, Shuigeng Zhou |
CVPR | 1 |
| 2018 | AON: Towards Arbitrarily-Oriented Text RecognitionabstractRecognizing text from natural images is a hot research topic in computer vision due to its various applications. Despite the enduring research of several decades on optical character recognition (OCR), recognizing texts from natural images is still a challenging task. This is because scene texts are often in irregular (e.g. curved, arbitrarily-oriented or seriously distorted) arrangements, which have not yet been well addressed in the literature. Existing methods on text recognition mainly work with regular (horizontal and frontal) texts and cannot be trivially generalized to handle irregular texts. In this paper, we develop the arbitrary orientation network (AON) to directly capture the deep features of irregular texts, which are combined into an attention-based decoder to generate character sequence. The whole network can be trained end-to-end by using only images and word-level annotations. Extensive experiments on various benchmarks, including the CUTE80, SVT-Perspective, IIIT5k, SVT and ICDAR datasets, show that the proposed AON-based method achieves the-state-of-the-art performance in irregular datasets, and is comparable to major existing methods in regular datasets. Zhanzhan Cheng, Yangliu Xu, Fan Bai 0001, Shiliang Pu, Shuigeng Zhou |
CVPR | 3 |
| 2017 | Focusing Attention: Towards Accurate Text Recognition in Natural ImagesabstractScene text recognition has been a hot research topic in computer vision due to its various applications. The state of the art is the attention-based encoder-decoder framework that learns the mapping between input images and output sequences in a purely data-driven way. However, we observe that existing attention-based methods perform poorly on complicated and/or low-quality images. One major reason is that existing methods cannot get accurate alignments between feature areas and targets for such images. We call this phenomenon “attention drift”. To tackle this problem, in this paper we propose the FAN (the abbreviation of Focusing Attention Network) method that employs a focusing attention mechanism to automatically draw back the drifted attention. FAN consists of two major components: an attention network (AN) that is responsible for recognizing character targets as in the existing methods, and a focusing network (FN) that is responsible for adjusting attention by evaluating whether AN pays attention properly on the target areas in the images. Furthermore, different from the existing methods, we adopt a ResNet-based network to enrich deep representations of scene text images. Extensive experiments on various benchmarks, including the IIIT5k, SVT and ICDAR datasets, show that the FAN method substantially outperforms the existing methods. Zhanzhan Cheng, Fan Bai 0001, Yunlu Xu, Shiliang Pu, Shuigeng Zhou |
ICCV | 2 |