VLDB 2026 Research / reviewers in the wild / expert
Kai Chen 0001
dblp:c/KaiChen1
· DBLP profile ↗
21ranked-venue papers
6as first author
4since 2021 · last 2023
0000-0001-6384-0355ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Improving Handwritten OCR with Training Samples Generated by Glyph Conditional Denoising Diffusion Probabilistic Model
Haisong Ding, Bozhi Luan, Dongnan Gui, Kai Chen 0001, Qiang Huo |
ICDAR (4) | 4 |
| 2023 | Zero-shot Generation of Training Data with Denoising Diffusion Probabilistic Model for Handwritten Chinese Character Recognition
Dongnan Gui, Kai Chen 0001, Haisong Ding, Qiang Huo |
ICDAR (2) | 2 |
| 2023 | GlyphControl: Glyph Conditional Control for Visual Text GenerationabstractRecently, there has been an increasing interest in developing diffusion-based text-to-image generative models capable of generating coherent and well-formed visual text. In this paper, we propose a novel and efficient approach called GlyphControl to address this task. Unlike existing methods that rely on character-aware text encoders like ByT5 and require retraining of text-to-image models, our approach leverages additional glyph conditional information to enhance the performance of the off-the-shelf Stable-Diffusion model in generating accurate visual text. By incorporating glyph instructions, users can customize the content, location, and size of the generated text according to their specific requirements. To facilitate further research in visual text generation, we construct a training benchmark dataset called LAION-Glyph. We evaluate the effectiveness of our approach by measuring OCR-based metrics, CLIP score, and FID of the generated visual text. Our empirical evaluations demonstrate that GlyphControl outperforms the recent DeepFloyd IF approach in terms of OCR accuracy, CLIP score, and FID, highlighting the efficacy of our method. Yukang Yang, Dongnan Gui, Yuhui Yuan, Weicong Liang, Haisong Ding, Han Hu 0001, Kai Chen 0001 |
NeurIPS | 7 |
| 2021 | An Encoder-Decoder Approach to Handwritten Mathematical Expression Recognition with Multi-head Attention and Stacked Decoder
Haisong Ding, Kai Chen 0001, Qiang Huo |
ICDAR (2) | 2 |
| 2020 | Parallelizing Adam Optimizer with Blockwise Model-Update FilteringabstractRecently Adam has become a popular stochastic optimization method in deep learning area. To parallelize Adam in a distributed system, synchronous stochastic gradient (SSG) technique is widely used, which is inefficient due to heavy communication cost. In this paper, we attempt to parallelize Adam with blockwise model-update filtering (BMUF) instead. BMUF synchronizes model-update periodically and introduces a block momentum to improve performance. We propose a novel way to modify the estimated moment buffers of Adam and figure out a simple yet effective trick for hyper-parameter setting under BMUF framework. Experimental results on large scale English optical character recognition (OCR) task and large vocabulary continuous speech recognition (LVCSR) task show that BMUF-Adam achieves almost a linear speedup without recognition accuracy degradation and outperforms SSG-based method in terms of speedup, scalability and recognition accuracy. Kai Chen 0001, Haisong Ding, Qiang Huo |
ICASSP | 1 |
| 2020 | Improving Handwritten OCR with Augmented Text Line Images Synthesized from Online Handwriting Samples by Style-Conditioned GANabstractBy leveraging large amounts of training data and deep learning technologies, performances of modern handwritten optical character recognition (OCR) systems have been greatly improved. However, collecting and labeling massive handwriting images are both time-consuming and expensive. In this paper, we propose to augment handwritten OCR training with online handwriting samples. To achieve this goal, we propose a style-conditioned generative adversarial network (SC-GAN) with a novel training data pair generation strategy. Then this network is used to transfer the styles of real handwriting images to skeleton images extracted from online handwriting samples to generate photo-realistic text line images. Experimental results on a large scale handwritten OCR task show that the recognition accuracy of our handwritten OCR system is improved by using the augmented synthetic training data. Mingyang Guan, Haisong Ding, Kai Chen 0001, Qiang Huo |
ICFHR | 3 |
| 2020 | Improving Knowledge Distillation of CTC-Trained Acoustic Models With Alignment-Consistent Ensemble and Target DelayabstractKnowledge distillation (KD) has been widely used to improve the performance of a simpler student model by imitating the outputs or intermediate representations of a more complex teacher model. The most commonly used KD technique is to minimize a Kullback-Leibler divergence between the output distributions of the teacher and student models. When it is applied to compressing acoustic models trained with a connectionist temporal classification (CTC) criterion, an assumption is made that the teacher and student share the same frame-level feature-transcription alignment. However, frame-level alignments learned by teachers can be inaccurate and unstable due to the lack of fine-grained frame-level guidance during CTC training. Forcing student to learn inaccurate alignments will lead to limited performance improvements. In this article, we investigate building powerful teacher models with more accurate and stable feature-transcription alignments. We achieve this goal by using a novel alignment-consistent ensemble (ACE) technique, where all models within an ensemble are jointly trained along with a regularization term to encourage consistent and stable alignments. With well-trained deep bidirectional LSTM (DBLSTM) ACE as a teacher, we can directly use the traditional frame-wise KD method to train DBLSTM students. When applying KD to transfer knowledge from a DBLSTM ACE to a deep unidirectional LSTM (DLSTM) student, a simple yet effective target delay technique is proposed to handle the alignment difference between bidirectional and unidirectional models. Experimental results on Switchboard-I speech recognition task show that, with DBLSTM ACE as a teacher, the simple frame-wise KD method can achieve competitive or better performance than other complex KD methods on DBLSTM students. When applying KD to build DLSTM students from DBLSTM teachers, our proposed target delay technique can achieve relative word error rate reductions of 14.2%$\sim$14.8% compared with the models trained from scratch, which outperforms other carefully-designed KD methods. Haisong Ding, Kai Chen 0001, Qiang Huo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Compression of CTC-Trained Acoustic Models by Dynamic Frame-Wise Distillation or Segment-Wise N-Best Hypotheses Imitation
Haisong Ding, Kai Chen 0001, Qiang Huo |
INTERSPEECH | 2 |
| 2019 | Compressing CNN-DBLSTM models for OCR with teacher-student learning and Tucker decomposition
Haisong Ding, Kai Chen 0001, Qiang Huo |
Pattern Recognit. | 2 |
| 2018 | Building Compact CNN-DBLSTM Based Character Models for Handwriting Recognition and OCR by Teacher-Student LearningabstractCharacter models based on convolutional neural network (CNN) and deep bidirectional long short-term memory (DBLSTM) have achieved high recognition accuracy on various handwriting recognition (HWR) and OCR tasks. To deploy CNN-DBLSTM models in products, it is necessary to reduce the footprint and runtime latency as much as possible. In this paper, we use a teacher-student learning approach to achieve this goal, where a new objective function is proposed to match the extracted CNN feature sequences of the teacher and student models under the guidance of the succeeding LSTM layer. Experimental results on large scale English HWR and OCR tasks show that the learned small student model can achieve about 14.6x footprint reduction and 9.6x speedup without recognition accuracy degradation against the big teacher model. Haisong Ding, Kai Chen 0001, Wenping Hu, Qiang Huo |
ICFHR | 2 |
| 2017 | A Compact CNN-DBLSTM Based Character Model for Online Handwritten Chinese Text RecognitionabstractRecently, character model based on integrated convolutional neural network (CNN) and deep bidirectional long short-term memory (DBLSTM) has been demonstrated to be effective for online handwritten Chinese text recognition (HCTR). However, the reported CNN-DBLSTM topologies are too complex to be practically useful. In this paper, we propose a compact CNN-DBLSTM which has small footprint and low computation cost yet be able to accommodate multiple receptive fields for CNN-based feature extraction. By using the training set of a popular benchmark database, namely CASIA-OLHWDB, we trained a compact CNN-DBLSTM by a connectionist temporal classification (CTC) criterion with a multi-step training strategy. Combined this character model with a character trigram language model, our online HCTR system with a WFSTbased decoder has achieved state-of-the-art performance on both CASIA and ICDAR-2013 Chinese handwriting recognition competition test sets. Kai Chen 0001, Haisong Ding, Lei Sun 0003, Sen Liang, Qiang Huo |
ICDAR | 1 |
| 2017 | An Open Vocabulary OCR System with Hybrid Word-Subword Language ModelsabstractThe accuracy of a typical state-of-the-art optical character recognition (OCR) system benefits greatly from using a language model (LM). However, a conventional LM has a limited vocabulary, resulting in out-of-vocabulary (OOV) words that cannot be recognized by the OCR system. In this paper, we present an open vocabulary OCR system based on a hybrid LM. The vocabulary of the hybrid LM consists of both words and subwords. OOV words can be generated by combinations of subwords. A refined hybrid LM training scheme is applied by interpolating a standard hybrid LM, a word-based LM and a subword-based LM. An efficient word combination method is performed by modeling optional space symbols in a decoding network. The overall system deals with OOV words in a general, data-driven and language-independent way. We conduct experiments on an English handwriting OCR task. Evaluations on three testing sets demonstrate that the OCR system with the proposed method achieves a word error rate of 33.4% on an OOV-only testing set, yet without degrading the recognition accuracies on the other two testing sets mainly consisting of in-vocabulary words. Wenping Hu, Kai Chen 0001, Lei Sun 0003, Sen Liang, Xiongjian Mo, Qiang Huo |
ICDAR | 3 |
| 2017 | A Compact CNN-DBLSTM Based Character Model for Offline Handwriting Recognition with Tucker DecompositionabstractRecently, character model based on integrated convolutional neural network (CNN) and deep bidirectional long short-term memory (DBLSTM) has achieved excellent performance for offline handwriting recognition (HWR). To deploy CNN-DBLSTM model in products, it is necessary to reduce the footprint and runtime latency as much as possible. In this paper, we study two methods to compress the CNN part: (1) Use Tucker decomposition to decompose pre-trained weights with low-rank approximation, followed by fine-tuning; (2) Use grouped convolution to construct sparse connections in channel domain. Experiments have been conducted on a large-scale offline English HWR task to compare the effectiveness of the above two techniques. Our results show that using Tucker decomposition alone offers a good solution to building a compact CNN-DBLSTM model which can reduce significantly both the footprint and latency yet without degrading recognition accuracy. Haisong Ding, Kai Chen 0001, Lei Sun 0003, Sen Liang, Qiang Huo |
ICDAR | 2 |
| 2017 | Sequence Discriminative Training for Offline Handwriting Recognition by an Interpolated CTC and Lattice-Free MMI Objective FunctionabstractWe study two sequence discriminative training criteria, i.e., Lattice-Free Maximum Mutual Information (LFMMI) and Connectionist Temporal Classification (CTC), for end-to-end training of Deep Bidirectional Long Short-Term Memory (DBLSTM) based character models of two offline English handwriting recognition systems with an input feature vector sequence extracted by Principal Component Analysis (PCA) and Convolutional Neural Network (CNN), respectively. We observe that refining CTC-trained PCA-DBLSTM model with an interpolated CTC and LFMMI objective function ("CTC+LFMMI") for several additional iterations achieves a relative Word Error Rate (WER) reduction of 24.6% and 13.9% on the public IAM test set and an in-house E2E test set, respectively. For a much better CTC-trained CNN-DBLSTM system, the proposed "CTC+LFMMI" method achieves a relative WER reduction of 19.6% and 8.3% on the above two test sets, respectively. Wenping Hu, Kai Chen 0001, Haisong Ding, Lei Sun 0003, Sen Liang, Xiongjian Mo, Qiang Huo |
ICDAR | 3 |
| 2016 | Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filteringabstractWe present a new approach to scalable training of deep learning machines by incremental block training with intra-block parallel optimization to leverage data parallelism and blockwise model-update filtering to stabilize learning process. By using an implementation on a distributed GPU cluster with an MPI-based HPC machine learning framework to coordinate parallel job scheduling and collective communication, we have trained successfully deep bidirectional long short-term memory (LSTM) recurrent neural networks (RNNs) and fully-connected feed-forward deep neural networks (DNNs) for large vocabulary continuous speech recognition on two benchmark tasks, namely 309-hour Switchboard-I task and 1,860-hour "Switch-board+Fisher" task. We achieve almost linear speedup up to 16 GPU cards on LSTM task and 64 GPU cards on DNN task, with either no degradation or improved recognition accuracy in comparison with that of running a traditional mini-batch based stochastic gradient descent training on a single GPU. Kai Chen 0001, Qiang Huo |
ICASSP | 1 |
| 2016 | Training Deep Bidirectional LSTM Acoustic Model for LVCSR by a Context-Sensitive-Chunk BPTT ApproachabstractThis paper presents a study of using deep bidirectional long short-term memory (DBLSTM) recurrent neural network as acoustic model for DBLSTM-HMM based large vocabulary continuous speech recognition (LVCSR), where a context-sensitive-chunk (CSC) back-propagation through time (BPTT) approach is used to train DBLSTM by splitting each training sequence into chunks with appended contextual observations, and a CSC-based decoding method with possibly overlapped CSCs is used for recognition. Our approach makes mini-batch based training on GPU more efficient and reduces the latency of DBLSTM-based LVCSR from a whole utterance to a short chunk. Evaluations have been made on Switchboard-I benchmark task. In comparison with epoch-wise BPTT training, our method can achieve more than three times speedup on a single GPU card without degrading recognition accuracy. In comparison with a highly optimized DNN-HMM system trained by a frame-level cross entropy (CE) criterion, our CE-trained DBLSTM-HMM system achieves relative word error rate reductions of 9% and 5% on Eval2000 and RT03S testing sets, respectively. Furthermore, by running model averaging based parallel training of DBLSTM on a cluster of GPUs, CSC-BPTT incurs less accuracy degradation than epoch-wise BPTT while achieves a linear speedup. Kai Chen 0001, Qiang Huo |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | A context-sensitive-chunk BPTT approach to training deep LSTM/BLSTM recurrent neural networks for offline handwriting recognitionabstractWe propose a context-sensitive-chunk based back-propagation through time (BPTT) approach to training deep (bidirectional) long short-term memory ((B)LSTM) recurrent neural networks (RNN) that splits each training sequence into chunks with appended contextual observations for character modeling of offline handwriting recognition. Using short context-sensitive chunks in both training and recognition brings following benefits: (1) the learned (B)LSTM will model mainly local character image dependency and the effect of long-range language model information reflected in training data is reduced; (2) mini-batch based training on GPU can be made more efficient; (3) low-latency BLSTM-based handwriting recognition is made possible by incurring only a delay of a short chunk rather than a whole sentence. Our approach is evaluated on IAM offline handwriting recognition benchmark task and performs better than the previous state-of-the-art BPTT-based approaches. Kai Chen 0001, Zhijie Yan, Qiang Huo |
ICDAR | 1 |
| 2015 | Training deep bidirectional LSTM acoustic model for LVCSR by a context-sensitive-chunk BPTT approach
Kai Chen 0001, Zhijie Yan, Qiang Huo |
INTERSPEECH | 1 |
| 2015 | A robust approach for text detection from natural scene images
Lei Sun 0003, Qiang Huo, Wei Jia 0003, Kai Chen 0001 |
Pattern Recognit. | 4 |
| 2014 | Robust Text Detection in Natural Scene Images by Generalized Color-Enhanced Contrasting Extremal Region and Neural NetworksabstractThis paper presents a robust text detection approach based on generalized color-enhanced contrasting extremal region (CER) and neural networks. Given a color natural scene image, six component-trees are built from its gray scale image, hue and saturation channel images in a perception-based illumination invariant color space, and their inverted images, respectively. From each component-tree, generalized color-enhanced CERs are extracted as character candidates. By using a "divide-and-conquer" strategy, each candidate image patch is labeled reliably by rules as one of five types, namely, Long, Thin, Fill, Square-large and Square-small, and classified as text or non-text by a corresponding neural network, which is trained by an ambiguity-free learning strategy. After pruning non-text components, repeating components in each component-tree are pruned by using color and area information to obtain a component graph, from which candidate text-lines are formed and verified by another set of neural networks. Finally, results from six component-trees are combined, and a post-processing step is used to recover lost characters and split text lines into words as appropriate. Our proposed method achieves 85.72% recall, 87.03% precision, and 86.37% F-score on ICDAR-2013 "Reading Text in Scene Images" test set. Lei Sun 0003, Qiang Huo, Wei Jia 0003, Kai Chen 0001 |
ICPR | 4 |
| 2012 | Designing compact classifiers for rotation-free recognition of large vocabulary online handwritten Chinese charactersabstractWe present a study of designing compact multiple-prototype based classifiers for rotation-free recognition of online handwritten Chinese characters. Several versions of Rprop algorithms are adopted to optimize a sample-separation-margin based minimum classification error objective function. Split vector quantization technique is used to compress classifier parameters and a fast-match tree is used for efficient recognition. A new preprocessing technique is proposed to achieve rotation-free recognition capability. Promising benchmark results are reported on an online handwritten character recognition task with a vocabulary of 27,720 characters. Jun Du 0002, Qiang Huo, Kai Chen 0001 |
ICASSP | 3 |