Qiang Huo

dblp:03/2617 · DBLP profile ↗
← Back
156ranked-venue papers
35as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 103 · 19 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 80 · 21 first-author · 2 since 2021Databases, data management, data science and information retrieval · 34 · 1 first-author · 9 since 2021Computer networks · 7 · 5 first-authorSystems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 UniHDSA: A unified relation prediction approach for hierarchical document structure analysis
Jiawei Wang 0026, Qiang Huo
Pattern Recognit.3
2025 SHMT: An SRAM and HBM Hybrid Computing-in-Memory Architecture With Optimized KV Cache for Multimodal Transformer
abstract
Multimodal Transformer (MMT) algorithms have become the state-of-the-art for multimodal tasks such as image captioning. The Encoder-Decoder (E-D) structure, consisting of Encoder, Decoder-causal, and Decoder-cross components, provides a flexible and effective framework for multimodal tasks. However, previous accelerators mainly focus on the dataflow and hardware optimization of the Encoder, which fails to accelerate the entire E-D structure efficiently. There remain three challenges: 1) the lack of pipeline and multicore optimization at the module, layer, and E-D level; 2) the Decoder-causal and Decoder-cross computations have lower arithmetic intensity compared to the Encoder, requiring a better solution for the varying arithmetic intensities; and 3) the autoregressive algorithm in Decoder-causal leads to redundant KV Cache accesses and considerable idle power. In this paper, SHMT, an SRAM and HBM hybrid computing-in-memory (CIM) architecture, is designed to efficiently support multimodal Transformers with three key contributions: 1) a multi-level pipelined multicore scheme, including pipeline optimization across E-D layer-head-module levels and a multicore network-on-chip (NoC) architecture, to reduce inference latency and off-chip accesses; 2) a heterogeneous SRAM-HBM architecture, utilizing high-density HBM-CIM for low-arithmetic-intensity (LAI) parts and high-performance SRAM-CIM for high-arithmetic-intensity (HAI) parts; and 3) by integrating KV Cache with zero-padding in SRAM-CIM, SHMT eliminates redundant read-write operations in KV Cache, reducing idle power consumption. Experiment results show that SHMT achieves 212× speedup, reduces energy consumption by 208×~2000× per token, and achieves 13.3× higher energy efficiency compared to NVIDIA A100 GPU.
Xiangqu Fu, Jinshan Yue, Muhammad Faizan, Zhi Li 0062, Qiang Huo, Feng Zhang 0014
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 UniVIE: A Unified Label Space Approach to Visual Information Extraction from Form-Like Documents
Jiawei Wang 0026, Weihong Lin, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICDAR (6)6
2024 DLAFormer: An End-to-End Transformer For Document Layout Analysis
Jiawei Wang 0026, Qiang Huo
ICDAR (4)3
2024 Dynamic Relation Transformer for Contextual Text Block Detection
Jiawei Wang 0026, Shunchi Zhang, Chixiang Ma, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICDAR (1)7
2024 Mathematical formula detection in document images: A new dataset and a new approach
Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
Pattern Recognit.4
2024 Detect-order-construct: A tree construction based approach for hierarchical document structure analysis
Jiawei Wang 0026, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
Pattern Recognit.5
2023 A Question-Answering Approach to Key Value Pair Extraction from Form-Like Document Images
abstract
In this paper, we present a new question-answering (QA) based key-value pair extraction approach, called KVPFormer, to robustly extracting key-value relationships between entities from form-like document images. Specifically, KVPFormer first identifies key entities from all entities in an image with a Transformer encoder, then takes these key entities as questions and feeds them into a Transformer decoder to predict their corresponding answers (i.e., value entities) in parallel. To achieve higher answer prediction accuracy, we propose a coarse-to-fine answer prediction approach further, which first extracts multiple answer candidates for each identified question in the coarse stage and then selects the most likely one among these candidates in the fine stage. In this way, the learning difficulty of answer prediction can be effectively reduced so that the prediction accuracy can be improved. Moreover, we introduce a spatial compatibility attention bias into the self-attention/cross-attention mechanism for KVPFormer to better model the spatial interactions between entities. With these new techniques, our proposed KVPFormer achieves state-of-the-art results on FUNSD and XFUND datasets, outperforming the previous best-performing method by 7.2% and 13.2% in F1 score, respectively.
Zhuoyuan Wu, Zhuoyao Zhong, Weihong Lin, Lei Sun 0003, Qiang Huo
AAAI6
2023 Improving Handwritten OCR with Training Samples Generated by Glyph Conditional Denoising Diffusion Probabilistic Model
Haisong Ding, Bozhi Luan, Dongnan Gui, Kai Chen 0001, Qiang Huo
ICDAR (4)5
2023 Zero-shot Generation of Training Data with Denoising Diffusion Probabilistic Model for Handwritten Chinese Character Recognition
Dongnan Gui, Kai Chen 0001, Haisong Ding, Qiang Huo
ICDAR (2)4
2023 DQ-DETR: Dynamic Queries Enhanced Detection Transformer for Arbitrary Shape Text Detection
Chixiang Ma, Lei Sun 0003, Jiawei Wang 0026, Qiang Huo
ICDAR (2)4
2023 A Hybrid Approach to Document Layout Analysis for Heterogeneous Document Images
Zhuoyao Zhong, Jiawei Wang 0026, Haiqing Sun, Erhan Zhang, Lei Sun 0003, Qiang Huo
ICDAR (5)7
2023 Robust Table Detection and Structure Recognition from Heterogeneous Document Images
abstract
We introduce a new table detection and structure recognition approach named RobusTabNet to detect the boundaries of tables and reconstruct the cellular structure of each table from heterogeneous document images. For table detection, we propose to use CornerNet as a new region proposal network to generate higher quality table proposals for Faster R-CNN, which has significantly improved the localization accuracy of Faster R-CNN for table detection. Consequently, our table detection approach achieves state-of-the-art performance on three public table detection benchmarks, namely cTDaR TrackA, PubLayNet and IIIT-AR-13K, by only using a lightweight ResNet-18 backbone network. Furthermore, we propose a new split-and-merge based table structure recognition approach, in which a novel spatial CNN based separation line prediction module is proposed to split each detected table into a grid of cells, and a Grid CNN based cell merging module is applied to recover the spanning cells. As the spatial CNN module can effectively propagate contextual information across the whole table image, our table structure recognizer can robustly recognize tables with large blank spaces and geometrically distorted (even curved) tables. Thanks to these two techniques, our table structure recognition approach achieves state-of-the-art performance on three public benchmarks, including SciTSR, PubTabNet and cTDaR TrackB2-Modern. Moreover, we have further demonstrated the advantages of our approach in recognizing tables with complex structures, large blank spaces, as well as geometrically distorted or even curved shapes on a more challenging in-house dataset.
Chixiang Ma, Weihong Lin, Lei Sun 0003, Qiang Huo
Pattern Recognit.4
2023 Robust table structure recognition with dynamic queries enhanced detection transformer
Jiawei Wang 0026, Weihong Lin, Chixiang Ma, Lei Sun 0003, Qiang Huo
Pattern Recognit.7
2023 A Security-Enhanced, Charge-Pump-Free, ISO14443-A-/ISO10373-6-Compliant RFID Tag With 16.2-μW Embedded RRAM and Reconfigurable Strong PUF
abstract
Radio frequency identification technology (RFID) has empowered a wide variety of automation industries, such as logistics and freight transportation. To further promote RFID tags adoption, security, power consumption, and cost have always been issues of general concern. This article presents the first synergy of the RFID tag with embedded resistive RAM (RRAM) array and RRAM-based reconfigurable strong physical unclonable function (R-SPUF). The RRAM not only meets the mass storage and technology downscaling but also renders the ultralow-cost “1-cent RFID tag” more feasible. Moreover, the R-SPUF facilitates multiple initializations until a satisfactory distribution and has strong secure keys benefiting from its reconfigurability that improves both safety and reliability. The complete system operates at 13.56 MHz and is compliant with the ISO14443-A and ISO10373-6 (test) protocols. The RFID tag was fabricated on a 1.1-mm2 die based on the 0.18-$\mu \text{m}$CMOS process. Without resorting to the charge pumps for RRAM read–write operations, the total power consumption is as low as 52.3$\mu \text{W}$, of which the RRAM dissipates$16.2~\mu \text{W}$under a wireless power supply.
Qirui Ren, Qiang Huo, Hao Wu 0084, Xiangqu Fu, Xiaoxin Xu, Jianfeng Gao 0005, Xiaojin Zhao, Dengyun Lei, Xinghua Wang 0005, Feng Zhang 0014, Yong Chen 0005, Pui-In Mak
IEEE Trans. Very Large Scale Integr. Syst.2
2022 TSRFormer: Table Structure Recognition with Transformers
abstract
We present a new table structure recognition (TSR) approach, called TSRFormer, to robustly recognizing the structures of complex tables with geometrical distortions from various table images. Unlike previous methods, we formulate table separation line prediction as a line regression problem instead of an image segmentation problem and propose a new two-stage DETR based separator prediction approach, dubbed Sep arator RE gression TR ansformer (SepRETR), to predict separation lines from table images directly. To make the two-stage DETR framework work efficiently and effectively for the separation line prediction task, we propose two improvements: 1) A prior-enhanced matching strategy to solve the slow convergence issue of DETR; 2) A new cross attention module to sample features from a high-resolution convolutional feature map directly so that high localization accuracy is achieved with low computational cost. After separation line prediction, a simple relation network based cell merging module is used to recover spanning cells. With these new techniques, our TSRFormer achieves state-of-the-art performance on several benchmark datasets, including SciTSR, PubTabNet and WTW. Furthermore, we have validated the robustness of our approach to tables with complex structures, borderless cells, large blank spaces, empty or spanning cells as well as distorted or even curved shapes on a more challenging real-world in-house dataset.
Weihong Lin, Chixiang Ma, Jiawei Wang 0026, Lei Sun 0003, Qiang Huo
ACM Multimedia7
2021 An Encoder-Decoder Approach to Handwritten Mathematical Expression Recognition with Multi-head Attention and Stacked Decoder
Haisong Ding, Kai Chen 0001, Qiang Huo
ICDAR (2)3
2021 ViBERTgrid: A Jointly Trained Multi-modal 2D Document Representation for Key Information Extraction from Documents
Weihong Lin, Qifang Gao, Lei Sun 0003, Zhuoyao Zhong, Qin Ren 0003, Qiang Huo
ICDAR (1)7
2021 ReLaText: Exploiting visual relationships for arbitrary-shaped scene text detection with graph convolutional networks
Chixiang Ma, Lei Sun 0003, Zhuoyao Zhong, Qiang Huo
Pattern Recognit.4
2020 Parallelizing Adam Optimizer with Blockwise Model-Update Filtering
abstract
Recently Adam has become a popular stochastic optimization method in deep learning area. To parallelize Adam in a distributed system, synchronous stochastic gradient (SSG) technique is widely used, which is inefficient due to heavy communication cost. In this paper, we attempt to parallelize Adam with blockwise model-update filtering (BMUF) instead. BMUF synchronizes model-update periodically and introduces a block momentum to improve performance. We propose a novel way to modify the estimated moment buffers of Adam and figure out a simple yet effective trick for hyper-parameter setting under BMUF framework. Experimental results on large scale English optical character recognition (OCR) task and large vocabulary continuous speech recognition (LVCSR) task show that BMUF-Adam achieves almost a linear speedup without recognition accuracy degradation and outperforms SSG-based method in terms of speedup, scalability and recognition accuracy.
Kai Chen 0001, Haisong Ding, Qiang Huo
ICASSP3
2020 Improving Handwritten OCR with Augmented Text Line Images Synthesized from Online Handwriting Samples by Style-Conditioned GAN
abstract
By leveraging large amounts of training data and deep learning technologies, performances of modern handwritten optical character recognition (OCR) systems have been greatly improved. However, collecting and labeling massive handwriting images are both time-consuming and expensive. In this paper, we propose to augment handwritten OCR training with online handwriting samples. To achieve this goal, we propose a style-conditioned generative adversarial network (SC-GAN) with a novel training data pair generation strategy. Then this network is used to transfer the styles of real handwriting images to skeleton images extracted from online handwriting samples to generate photo-realistic text line images. Experimental results on a large scale handwritten OCR task show that the recognition accuracy of our handwritten OCR system is improved by using the augmented synthetic training data.
Mingyang Guan, Haisong Ding, Kai Chen 0001, Qiang Huo
ICFHR4
2020 A Study of BPE-based Language Modeling for Open Vocabulary Latin Language OCR
abstract
We present a study of byte pair encoding (BPE) based language modeling for open vocabulary Latin language OCR. On a large-scale handwritten English OCR task, we demonstrate that a simple BPE-based n-gram language model (LM) can deal with out-of-vocabulary word problem effectively and achieve better accuracy-footprint tradeoff than a state-of-the-art hybrid word/subword n-gram LM interpolated by a standard hybrid LM, a word-based LM, and a subword-based LM. On another large-scale printed OCR task for six Latin languages, namely English, Spanish, French, German, Italian, and Portuguese, we discover that a unified OCR system with a single character-based optical model and a single BPE-based n-gram LM shared by six languages performs better than language-dependent OCR systems. BPE-based LM offers a good product solution for both monolingual and multilingual open-vocabulary Latin language OCR.
Wenping Hu, Yikang Luo, Ji Meng, Zifei Qian, Qiang Huo
ICFHR5
2020 Improving Knowledge Distillation of CTC-Trained Acoustic Models With Alignment-Consistent Ensemble and Target Delay
abstract
Knowledge distillation (KD) has been widely used to improve the performance of a simpler student model by imitating the outputs or intermediate representations of a more complex teacher model. The most commonly used KD technique is to minimize a Kullback-Leibler divergence between the output distributions of the teacher and student models. When it is applied to compressing acoustic models trained with a connectionist temporal classification (CTC) criterion, an assumption is made that the teacher and student share the same frame-level feature-transcription alignment. However, frame-level alignments learned by teachers can be inaccurate and unstable due to the lack of fine-grained frame-level guidance during CTC training. Forcing student to learn inaccurate alignments will lead to limited performance improvements. In this article, we investigate building powerful teacher models with more accurate and stable feature-transcription alignments. We achieve this goal by using a novel alignment-consistent ensemble (ACE) technique, where all models within an ensemble are jointly trained along with a regularization term to encourage consistent and stable alignments. With well-trained deep bidirectional LSTM (DBLSTM) ACE as a teacher, we can directly use the traditional frame-wise KD method to train DBLSTM students. When applying KD to transfer knowledge from a DBLSTM ACE to a deep unidirectional LSTM (DLSTM) student, a simple yet effective target delay technique is proposed to handle the alignment difference between bidirectional and unidirectional models. Experimental results on Switchboard-I speech recognition task show that, with DBLSTM ACE as a teacher, the simple frame-wise KD method can achieve competitive or better performance than other complex KD methods on DBLSTM students. When applying KD to build DLSTM students from DBLSTM teachers, our proposed target delay technique can achieve relative word error rate reductions of 14.2%$\sim$14.8% compared with the models trained from scratch, which outperforms other carefully-designed KD methods.
Haisong Ding, Kai Chen 0001, Qiang Huo
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 A Comparative Study of Attention-Based Encoder-Decoder Approaches to Natural Scene Text Recognition
abstract
Attention-based encoder-decoder approaches have shown promising results in scene text recognition. In the literature, models with different encoders, decoders and attention mechanisms have been proposed and compared on isolated word recognition tasks, where the models are trained on either synthetic word images or a small set of real-world images. In this paper, we investigate different components of the attention based framework and compare its performance with a CNN-DBLSTM-CTC based approach on large-scale real-world scene text sentence recognition tasks. We train character models by using more than 1.6M real-world text lines and compare their performance on test sets collected from a variety of real-world scenarios. Our results show that (1) attention on a two-dimensional feature map can yield better performance than one-dimensional one and an RNN based decoder performs better than CNN based one; (2) attention-based approaches can achieve higher recognition accuracy than CNN-DBLSTM-CTC based approaches on isolated word recognition tasks, but perform worse on sentence recognition tasks; (3) it is more effective and efficient for CNN-DBLSTM-CTC based approaches to leverage an explicit language model to boost recognition accuracy.
Fu'ze Cong, Wenping Hu, Qiang Huo, Li Guo 0004
ICDAR3
2019 A Relation Network Based Approach to Curved Text Detection
abstract
In this paper, a new relation network based approach to curved text detection is proposed by formulating it as a visual relationship detection problem. The key idea is to decompose curved text detection into two subproblems, namely detection of text primitives and prediction of link relationship for each nearby text primitive pair. Specifically, an anchor-free region proposal network based text detector is first used to detect text primitives of different scales from different feature maps of a feature pyramid network, from which a manageable number of text primitive pairs are selected. Then, a relation network is used to predict whether each text primitive pair belongs to a same text instance. Finally, isolated text primitives are grouped into curved text instances based on link relationships of text primitive pairs. Because pairwise link prediction has used features extracted from the bounding boxes of each text primitive and their union, the relation network can effectively leverage wider context information to improve link prediction accuracy. Furthermore, since the link relationships of relatively distant text primitives can be predicted robustly, our relation network based text detector is capable of detecting text instances with large inter-character spaces. Consequently, our proposed approach achieves superior performance on not only two public curved text detection datasets, namely Total-Text and SCUT-CTW1500, but also a multi-oriented text detection dataset, namely MSRA-TD500.
Chixiang Ma, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICDAR4
2019 A Teacher-Student Learning Based Born-Again Training Approach to Improving Scene Text Detection Accuracy
abstract
With the recent success of convolutional neural network (CNN) based text detection approaches, designing better CNN-based text detection frameworks has become a major research focus to improve text detection accuracy. In this paper, instead of following this direction, we propose to use a born-again training strategy, which is based on teacher-student learning (TSL), to improve the accuracy of the state-of-the-art CNN-based text detectors. More specifically, given a well-trained CNN-based text detector, we take it as a teacher model and train from scratch a new student model with the same topology under the supervision of both the teacher model and ground-truth labels. Furthermore, we propose a new proposal-free multi-level feature mimicking approach to making multi-level convolutional feature maps be effectively mimicked in a unified manner. Experiments demonstrate that the student models trained by the proposed approach can achieve substantially better results than their teacher models and have better generalization abilities.
Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICDAR3
2019 Compression of CTC-Trained Acoustic Models by Dynamic Frame-Wise Distillation or Segment-Wise N-Best Hypotheses Imitation
Haisong Ding, Kai Chen 0001, Qiang Huo
INTERSPEECH3
2019 Mask R-CNN With Pyramid Attention Network for Scene Text Detection
abstract
In this paper, we present a new Mask R-CNN based text detection approach which can robustly detect multi-oriented and curved text from natural scene images in a unified manner. To enhance the feature representation ability of Mask R-CNN for text detection tasks, we propose to use the Pyramid Attention Network (PAN) as a new backbone network of Mask R-CNN. Experiments demonstrate that PAN can suppress false alarms caused by text-like backgrounds more effectively. Our proposed approach has achieved superior performance on both multi-oriented (ICDAR-2015, ICDAR-2017 MLT) and curved (SCUT-CTW1500) text detection benchmark tasks by only using single-scale and single-model testing.
Zhida Huang, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
WACV4
2019 Flow-guided feature propagation with occlusion aware detail enhancement for hand segmentation in egocentric videos
Lei Sun 0003, Qiang Huo
Comput. Vis. Image Underst.3
2019 An anchor-free region proposal network for Faster R-CNN-based text detection approaches
Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
Int. J. Document Anal. Recognit.3
2019 Compressing CNN-DBLSTM models for OCR with teacher-student learning and Tucker decomposition
Haisong Ding, Kai Chen 0001, Qiang Huo
Pattern Recognit.3
2019 Improved localization accuracy by LocNet for Faster R-CNN based text detection in natural scene images
Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
Pattern Recognit.3
2018 A CNN-Based Approach to Detecting Text from Images of Whiteboards and Handwritten Notes
abstract
Detecting handwritten text from images of whiteboards and handwritten notes is an important yet under-researched topic. In this paper, we propose a convolutional neural network (CNN) based approach to address this problem. First, to detect text instances of different scales, a feature pyramid network is adopted as a backbone network to extract three feature maps of different scales from a given input image, where a scale-specific detection module is attached to each feature map. Then, for a pixel on each feature map, a detection module is used to predict whether there exists a text instance at its corresponding location in the input image. For positive prediction, the bounding box of the detected text segment and the links between the concerned pixel and its 8 neighbors on the feature map are predicted simultaneously. Based on the linkage information, text segments extracted from each feature map are grouped into text-lines respectively and wrongly grouped text-lines are separated by a graph-based text-line segmentation method. Finally, detection results from three different feature maps are aggregated by a skewed non-maximum suppression algorithm. Our proposed approach has achieved superior results on a testing set consisting of 285 natural scene images of whiteboards and handwritten notes.
Wei Jia 0003, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICFHR4
2018 Building Compact CNN-DBLSTM Based Character Models for Handwriting Recognition and OCR by Teacher-Student Learning
abstract
Character models based on convolutional neural network (CNN) and deep bidirectional long short-term memory (DBLSTM) have achieved high recognition accuracy on various handwriting recognition (HWR) and OCR tasks. To deploy CNN-DBLSTM models in products, it is necessary to reduce the footprint and runtime latency as much as possible. In this paper, we use a teacher-student learning approach to achieve this goal, where a new objective function is proposed to match the extracted CNN feature sequences of the teacher and student models under the guidance of the succeeding LSTM layer. Experimental results on large scale English HWR and OCR tasks show that the learned small student model can achieve about 14.6x footprint reduction and 9.6x speedup without recognition accuracy degradation against the big teacher model.
Haisong Ding, Kai Chen 0001, Wenping Hu, Qiang Huo
ICFHR5
2018 DFF-DEN: Deep Feature Flow with Detail Enhancement Network for Hand Segmentation in Depth Video
abstract
Recently, researches are emerging to extend CNN-based segmentation approaches from still image to video. Directly applying per-frame based image segmentation networks on video is not efficient. To address this issue, a promising direction is to explore the video continuity. A state-of-the-art approach called Deep Feature Flow (DFF) runs the segmentation network only on sparse key frames and propagates feature maps to other frames via cross-frame motion. However, such approach does not work well for hand segmentation in a video as it is not robust to hand posture change. In this paper, we propose to incorporate a light-weight detail enhancement network (DEN) into the DFF framework to achieve robustness against cross-frame motion and hand posture change. Experimental results on a public depth video dataset, FingerPaint, demonstrate that our approach achieves higher segmentation accuracy than the DFF-based approach with similar speedups against the per-frame based video hand segmentation approach.
Lei Sun 0003, Qiang Huo
ICIP3
2017 A Compact CNN-DBLSTM Based Character Model for Online Handwritten Chinese Text Recognition
abstract
Recently, character model based on integrated convolutional neural network (CNN) and deep bidirectional long short-term memory (DBLSTM) has been demonstrated to be effective for online handwritten Chinese text recognition (HCTR). However, the reported CNN-DBLSTM topologies are too complex to be practically useful. In this paper, we propose a compact CNN-DBLSTM which has small footprint and low computation cost yet be able to accommodate multiple receptive fields for CNN-based feature extraction. By using the training set of a popular benchmark database, namely CASIA-OLHWDB, we trained a compact CNN-DBLSTM by a connectionist temporal classification (CTC) criterion with a multi-step training strategy. Combined this character model with a character trigram language model, our online HCTR system with a WFSTbased decoder has achieved state-of-the-art performance on both CASIA and ICDAR-2013 Chinese handwriting recognition competition test sets.
Kai Chen 0001, Haisong Ding, Lei Sun 0003, Sen Liang, Qiang Huo
ICDAR7
2017 An Open Vocabulary OCR System with Hybrid Word-Subword Language Models
abstract
The accuracy of a typical state-of-the-art optical character recognition (OCR) system benefits greatly from using a language model (LM). However, a conventional LM has a limited vocabulary, resulting in out-of-vocabulary (OOV) words that cannot be recognized by the OCR system. In this paper, we present an open vocabulary OCR system based on a hybrid LM. The vocabulary of the hybrid LM consists of both words and subwords. OOV words can be generated by combinations of subwords. A refined hybrid LM training scheme is applied by interpolating a standard hybrid LM, a word-based LM and a subword-based LM. An efficient word combination method is performed by modeling optional space symbols in a decoding network. The overall system deals with OOV words in a general, data-driven and language-independent way. We conduct experiments on an English handwriting OCR task. Evaluations on three testing sets demonstrate that the OCR system with the proposed method achieves a word error rate of 33.4% on an OOV-only testing set, yet without degrading the recognition accuracies on the other two testing sets mainly consisting of in-vocabulary words.
Wenping Hu, Kai Chen 0001, Lei Sun 0003, Sen Liang, Xiongjian Mo, Qiang Huo
ICDAR7
2017 Compact and Efficient WFST-Based Decoders for Handwriting Recognition
abstract
We present two weighted finite-state transducer (WFST) based decoders for handwriting recognition. One decoder is a cloud-based solution that is both compact and efficient. The other is a device-based solution that has a small memory footprint. A compact WFST data structure is proposed for the cloud-based decoder. There are no output labels stored on transitions of the compact WFST. A decoder based on the compact WFST data structure produces the same result with significantly less footprint compared with a decoder based on the corresponding standard WFST. For the device-based decoder, on-the-fly language model rescoring is performed to reduce footprint. Careful engineering methods, such as WFST weight quantization, token and data type refinement, are also explored. When using a language model containing 600,000 n-grams, the cloud-based decoder achieves an average decoding time of 4.04 ms per text line with a peak footprint of 114.4 MB, while the device-based decoder achieves an average decoding time of 13.47 ms per text line with a peak footprint of 31.6 MB.
Qiang Huo
ICDAR2
2017 A Compact CNN-DBLSTM Based Character Model for Offline Handwriting Recognition with Tucker Decomposition
abstract
Recently, character model based on integrated convolutional neural network (CNN) and deep bidirectional long short-term memory (DBLSTM) has achieved excellent performance for offline handwriting recognition (HWR). To deploy CNN-DBLSTM model in products, it is necessary to reduce the footprint and runtime latency as much as possible. In this paper, we study two methods to compress the CNN part: (1) Use Tucker decomposition to decompose pre-trained weights with low-rank approximation, followed by fine-tuning; (2) Use grouped convolution to construct sparse connections in channel domain. Experiments have been conducted on a large-scale offline English HWR task to compare the effectiveness of the above two techniques. Our results show that using Tucker decomposition alone offers a good solution to building a compact CNN-DBLSTM model which can reduce significantly both the footprint and latency yet without degrading recognition accuracy.
Haisong Ding, Kai Chen 0001, Lei Sun 0003, Sen Liang, Qiang Huo
ICDAR7
2017 Sequence Discriminative Training for Offline Handwriting Recognition by an Interpolated CTC and Lattice-Free MMI Objective Function
abstract
We study two sequence discriminative training criteria, i.e., Lattice-Free Maximum Mutual Information (LFMMI) and Connectionist Temporal Classification (CTC), for end-to-end training of Deep Bidirectional Long Short-Term Memory (DBLSTM) based character models of two offline English handwriting recognition systems with an input feature vector sequence extracted by Principal Component Analysis (PCA) and Convolutional Neural Network (CNN), respectively. We observe that refining CTC-trained PCA-DBLSTM model with an interpolated CTC and LFMMI objective function ("CTC+LFMMI") for several additional iterations achieves a relative Word Error Rate (WER) reduction of 24.6% and 13.9% on the public IAM test set and an in-house E2E test set, respectively. For a much better CTC-trained CNN-DBLSTM system, the proposed "CTC+LFMMI" method achieves a relative WER reduction of 19.6% and 8.3% on the above two test sets, respectively.
Wenping Hu, Kai Chen 0001, Haisong Ding, Lei Sun 0003, Sen Liang, Xiongjian Mo, Qiang Huo
ICDAR8
2017 A Robust Approach to Detecting Text from Images of Whiteboards and Handwritten Notes
abstract
Detecting text from the images of whiteboards and handwritten notes is an important yet under-researched topic. In this paper, we present a robust approach to solving this challenging problem as follows. First, given a color image, colorenhanced Contrasting Extremal Regions (CERs) are extracted from its grayscale image as candidate text connected components (CCs). Second, four shallow neural networks are used to preprune efficiently most of unambiguous non-text CCs. Third, a Fast R-CNN based approach is proposed to filter out remaining nontext CCs by leveraging contextual information and to estimate the corresponding text-line orientation in the position of each remaining text CC. Fourth, each pair of the remaining text CCs within a certain distance and orientation constraint are connected to construct a directed graph. Finally, based on the estimated textline orientations, candidate text-lines are generated easily by pruning greedily redundant edges in the graph to make each vertex have at most one direct successor and one direct predecessor, respectively. Our proposed approach has achieved promising results on an in-house testing set consisting of 285 camera-captured images of whiteboards and handwritten notes.
Wei Jia 0003, Lei Sun 0003, Zhuoyao Zhong, Xiongjian Mo, Guoen Ma, Qiang Huo
ICDAR6
2017 Improved Localization Accuracy by LocNet for Faster R-CNN Based Text Detection
abstract
Although Faster R-CNN based approaches have achieved promising results for text detection, their localization accuracy is not satisfactory in certain cases. In this paper, we propose to use a LocNet to improve the localization accuracy of a Faster R-CNN based text detector. Given a proposal generated by region proposal network (RPN), instead of predicting directly the bounding box coordinates of the concerned text instance, the proposal is enlarged to create a search region so that conditional probabilities to each row and column of this search region can be assigned, which are then used to infer accurately the concerned bounding box. Experiments demonstrate that the proposed approach boosts the localization accuracy for Faster R-CNN based text detection significantly. Consequently, our new text detector has achieved superior performance on ICDAR-2011, ICDAR-2013 and MULTILIGUL text detection benchmark tasks.
Zhuoyao Zhong, Lei Sun 0003, Qiang Huo
ICDAR3
2017 Fingertip detection based on protuberant saliency from depth image
abstract
We propose a new approach for detecting a protuberant region from a depth image by leveraging a notion of protuberant saliency which describes how much the protuberant region stands out from its surroundings in the depth image. An intuitive and simple method is designed to calculate protuberant saliency from depth image, which can be used effectively together with the nearness information to detect the tip of a protuberant object such as a fingertip. We evaluate and compare our method with several state-of-the-art saliency methods for fingertip detection. Experimental results demonstrate that our method outperforms the comparing methods in terms of detection accuracy, being more robust against the rotation and isometric deformation of a fingertip region, the scale and depth noise issues, and the low resolution of a depth image.
Yuseok Ban, Lei Sun 0003, Qiang Huo
ICIP4
2016 Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering
abstract
We present a new approach to scalable training of deep learning machines by incremental block training with intra-block parallel optimization to leverage data parallelism and blockwise model-update filtering to stabilize learning process. By using an implementation on a distributed GPU cluster with an MPI-based HPC machine learning framework to coordinate parallel job scheduling and collective communication, we have trained successfully deep bidirectional long short-term memory (LSTM) recurrent neural networks (RNNs) and fully-connected feed-forward deep neural networks (DNNs) for large vocabulary continuous speech recognition on two benchmark tasks, namely 309-hour Switchboard-I task and 1,860-hour "Switch-board+Fisher" task. We achieve almost linear speedup up to 16 GPU cards on LSTM task and 64 GPU cards on DNN task, with either no degradation or improved recognition accuracy in comparison with that of running a traditional mini-batch based stochastic gradient descent training on a single GPU.
Kai Chen 0001, Qiang Huo
ICASSP2
2016 Precise hand segmentation from a single depth image
abstract
We propose a new approach to segmenting a hand accurately from a single depth image. Given a depth image, we extract first a rough hand region of interest (RoI) including a hand and a part of an arm. Then, the RoI is partitioned into triangles by using a constrained Delaunay triangulation (CDT) approach from which hand segmentation proposals are generated. Each segmentation proposal is evaluated by a shallow convolutional neural network (CNN) which is trained as a regression function to predict a confidence score for each proposal. Finally, the segmentation proposal with the highest confidence score is selected as our hand segmentation result. To evaluate the effectiveness of our approach, we use a set of real data containing more than 370,000 frames of hand depth images collected from 40 subjects with large variations in pose, orientation and sensing distance. Compared with segmentation results achieved by a random decision forest (RDF) based approach, our approach achieves much higher accuracy.
Lei Sun 0003, Qiang Huo
ICPR3
2016 Source and physical-layer network coding for correlated two-way relaying
abstract
In this paper, the authors study a half‐duplex two‐way relay channel with correlated sources exchanging bidirectional information. In the case, when both sources have the knowledge of correlation statistics, a source compression with physical‐layer network coding scheme is proposed to perform the distributed compression at each source node. When only the relay has the knowledge of correlation statistics, the authors propose a relay compression with physical‐layer network coding scheme to compress the bidirectional messages at the relay. The closed‐form block error rate expressions of both schemes are derived and verified through simulations. It is shown that the proposed schemes achieve considerable improvements in both error performance and throughput compared with the conventional non‐compression scheme in correlated two‐way relay networks.
Qiang Huo, Lingyang Song, Yonghui Li 0001, Bingli Jiao
IET Commun.1
2016 Training Deep Bidirectional LSTM Acoustic Model for LVCSR by a Context-Sensitive-Chunk BPTT Approach
abstract
This paper presents a study of using deep bidirectional long short-term memory (DBLSTM) recurrent neural network as acoustic model for DBLSTM-HMM based large vocabulary continuous speech recognition (LVCSR), where a context-sensitive-chunk (CSC) back-propagation through time (BPTT) approach is used to train DBLSTM by splitting each training sequence into chunks with appended contextual observations, and a CSC-based decoding method with possibly overlapped CSCs is used for recognition. Our approach makes mini-batch based training on GPU more efficient and reduces the latency of DBLSTM-based LVCSR from a whole utterance to a short chunk. Evaluations have been made on Switchboard-I benchmark task. In comparison with epoch-wise BPTT training, our method can achieve more than three times speedup on a single GPU card without degrading recognition accuracy. In comparison with a highly optimized DNN-HMM system trained by a frame-level cross entropy (CE) criterion, our CE-trained DBLSTM-HMM system achieves relative word error rate reductions of 9% and 5% on Eval2000 and RT03S testing sets, respectively. Furthermore, by running model averaging based parallel training of DBLSTM on a cluster of GPUs, CSC-BPTT incurs less accuracy degradation than epoch-wise BPTT while achieves a linear speedup.
Kai Chen 0001, Qiang Huo
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 A context-sensitive-chunk BPTT approach to training deep LSTM/BLSTM recurrent neural networks for offline handwriting recognition
abstract
We propose a context-sensitive-chunk based back-propagation through time (BPTT) approach to training deep (bidirectional) long short-term memory ((B)LSTM) recurrent neural networks (RNN) that splits each training sequence into chunks with appended contextual observations for character modeling of offline handwriting recognition. Using short context-sensitive chunks in both training and recognition brings following benefits: (1) the learned (B)LSTM will model mainly local character image dependency and the effect of long-range language model information reflected in training data is reduced; (2) mini-batch based training on GPU can be made more efficient; (3) low-latency BLSTM-based handwriting recognition is made possible by incurring only a delay of a short chunk rather than a whole sentence. Our approach is evaluated on IAM offline handwriting recognition benchmark task and performs better than the previous state-of-the-art BPTT-based approaches.
Kai Chen 0001, Zhijie Yan, Qiang Huo
ICDAR3
2015 A study on effects of implicit and explicit language model information for DBLSTM-CTC based handwriting recognition
abstract
Deep Bidirectional Long Short-Term Memory (DBLSTM) with a Connectionist Temporal Classification (CTC) output layer has been established as one of the state-of-the-art solutions for handwriting recognition. It is well-known that the DBLSTM trained by using a CTC objective function will learn both local character image dependency for character modeling and long-range contextual dependency for implicit language modeling. In this paper, we study the effects of implicit and explicit language model information for DBLSTM-CTC based handwriting recognition by comparing the performance of using or without using an explicit language model in decoding. It is observed that even using one million lines of training sentences to train the DBLSTM, using an explicit language model is still helpful. To deal with such a large-scale training problem, a GPU-based training tool has been developed for CTC training of DBLSTM by using a mini-batch based epochwise Back Propagation Through Time (BPTT) algorithm.
Qi Liu 0018, Qiang Huo
ICDAR3
2015 Building Handwriting Recognizers by Leveraging Skeletons of Both Offline and Online Samples
abstract
We present an approach to leveraging both offline and online handwriting samples to build a single recognizer for recognizing both offline and online handwritings. Given a training set of offline handwriting samples and another set of online handwriting samples, a skeleton is derived first from each offline handwriting sample via vectorization. Then both the skeleton samples and online handwriting samples are normalized and rendered by using the same method to generate a combined training set of skeleton images. Finally a handwriting recognizer based on Deep Bidirectional Long Short-Term Memory (DBLSTM) and Hidden Markov Model (HMM) is built from the skeleton images. In recognition, a preprocessing step consistent with that in training is applied to an unknown offline or online handwriting sample to derive a skeleton image, which is recognized by the hybrid DBLSTM-HMM handwriting recognition system accordingly. We have built such a recognizer by using IAM benchmark databases of offline and online English handwritings plus an internal online handwriting corpus, which outperforms the recognizers built from either offline or online handwriting samples only.
Qiang Huo
ICDAR4
2015 Training deep bidirectional LSTM acoustic model for LVCSR by a context-sensitive-chunk BPTT approach
Kai Chen 0001, Zhijie Yan, Qiang Huo
INTERSPEECH3
2015 A robust approach for text detection from natural scene images
Lei Sun 0003, Qiang Huo, Wei Jia 0003, Kai Chen 0001
Pattern Recognit.2
2015 Video-audio driven real-time facial animation
abstract
We present a real-time facial tracking and animation system based on a Kinect sensor with video and audio input. Our method requires no user-specific training and is robust to occlusions, large head rotations, and background noise. Given the color, depth and speech audio frames captured from an actor, our system first reconstructs 3D facial expressions and 3D mouth shapes from color and depth input with a multi-linear model. Concurrently a speaker-independent DNN acoustic model is applied to extract phoneme state posterior probabilities (PSPP) from the audio frames. After that, a lip motion regressor refines the 3D mouth shape based on both PSPP and expression weights of the 3D mouth shapes, as well as their confidences. Finally, the refined 3D mouth shape is combined with other parts of the 3D face to generate the final result. The whole process is fully automatic and executed in real time. The key component of our system is a data-driven regresor for modeling the correlation between speech data and mouth shapes. Based on a precaptured database of accurate 3D mouth shapes and associated speech audio from one speaker, the regressor jointly uses the input speech and visual features to refine the mouth shape of a new actor. We also present an improved DNN acoustic model. It not only preserves accuracy but also achieves real-time performance. Our method efficiently fuses visual and acoustic information for 3D facial performance capture. It generates more accurate 3D mouth motions than other approaches that are based on audio or video input only. It also supports video or audio only input for real-time facial animation. We evaluate the performance of our system with speech and facial expressions captured from different actors. Results demonstrate the efficiency and robustness of our method.
Feng Xu 0005, Jinxiang Chai, Xin Tong 0001, Qiang Huo
ACM Trans. Graph.6
2014 Synthesized stereo mapping via deep neural networks for noisy speech recognition
abstract
In our previous work, we extend the traditional stereo-based stochastic mapping by relaxing the constraint of stereo-data, which is not practical in real applications, via HMM-based speech synthesis to construct the “clean” channel data for noisy speech recognition. In this paper, we propose to use deep neural networks (DNNs) for stereo mapping compared with the joint Gaussian mixture model (GMM). The experimental results on Aurora3 databases show that our proposed DNN based synthesized stereo mapping can achieve consistently significant improvements of recognition performance over joint GMM based synthesized stereo mapping in the well-matched (WM) condition among four different European languages.
Jun Du 0002, Li-Rong Dai 0001, Qiang Huo
ICASSP3
2014 Robust Text Detection in Natural Scene Images by Generalized Color-Enhanced Contrasting Extremal Region and Neural Networks
abstract
This paper presents a robust text detection approach based on generalized color-enhanced contrasting extremal region (CER) and neural networks. Given a color natural scene image, six component-trees are built from its gray scale image, hue and saturation channel images in a perception-based illumination invariant color space, and their inverted images, respectively. From each component-tree, generalized color-enhanced CERs are extracted as character candidates. By using a "divide-and-conquer" strategy, each candidate image patch is labeled reliably by rules as one of five types, namely, Long, Thin, Fill, Square-large and Square-small, and classified as text or non-text by a corresponding neural network, which is trained by an ambiguity-free learning strategy. After pruning non-text components, repeating components in each component-tree are pruned by using color and area information to obtain a component graph, from which candidate text-lines are formed and verified by another set of neural networks. Finally, results from six component-trees are combined, and a post-processing step is used to recover lost characters and split text lines into words as appropriate. Our proposed method achieves 85.72% recall, 87.03% precision, and 86.37% F-score on ICDAR-2013 "Reading Text in Scene Images" test set.
Lei Sun 0003, Qiang Huo, Wei Jia 0003, Kai Chen 0001
ICPR2
2014 Selective combining for hybrid cooperative networks
abstract
In this study, we consider the selective combining in hybrid cooperative networks (SCHCNs scheme) with one source node, one destination node and N relay nodes. In the SCHCN scheme, each relay first adaptively chooses between amplify‐and‐forward protocol and decode‐and‐forward protocol on a per frame basis by examining the error‐detecting code result, and N c (1 ≤ N c ≤ N ) relays will be selected to forward their received signals to the destination. We first develop a signal‐to‐noise ratio (SNR) threshold‐based frame error rate (FER) approximation model. Then, the theoretical FER expressions for the SCHCN scheme are derived by utilising the proposed SNR threshold‐based FER approximation model. The analytical FER expressions are validated through simulation results.
Qiang Huo, Tianxi Liu, Shaohui Sun, Lingyang Song, Bingli Jiao
IET Commun.1
2014 Study on downlink spectral efficiency in orthogonal frequency division multiple access systems
abstract
In previous studies on the capacity of orthogonal frequency division multiple access (OFDMA) systems, it is usually assumed that co‐channel interference (CCI) from adjacent cells is a Gaussian‐distributed random variable. However, very‐little work shows that the Gaussian assumption does not hold true in OFDMA systems. In this study, the statistical property of CCI in downlink OFDMA systems is studied, and spectral efficiency of downlink OFDMA system is analysed based on the derived statistical model. First, the probability density function (PDF) of CCI in downlink OFDMA cellular systems is studied with the considerations of path loss, multipath fading and Gaussian‐like transmit signals. Moreover, some closed‐form expressions of the PDF are obtained for special cases. The derived results show that the PDFs of CCI are with a heavy tail, and significantly deviate from the Gaussian distribution. Then, based on the derived statistical properties of CCI, the downlink spectral efficiency is derived. Numerical and simulation results justify the derived statistical CCI model and spectral efficiency.
Qiang Huo, Bingli Jiao
IET Commun.1
2014 An irrelevant variability normalization approach to discriminative training of multi-prototype based classifiers and its applications for online handwritten Chinese character recognition
Jun Du 0002, Qiang Huo
Pattern Recognit.2
2014 An Improved VTS Feature Compensation using Mixture Models of Distortion and IVN Training for Noisy Speech Recognition
abstract
In our previous work, we proposed a feature compensation approach using high-order vector Taylor series (VTS) approximation for noisy speech recognition. In this paper, we report new progress on making it more powerful and practical in real applications. First, mixtures of densities are used to enhance the distortion models of both additive noise and convolutional distortion. New formulations for maximum likelihood (ML) estimation of distortion model parameters, and minimum mean squared error (MMSE) estimation of clean speech are derived and presented. Second, we improve the feature compensation in both efficiency and accuracy by applying higher order information of VTS approximation only to the noisy speech mean parameters, and a temporal smoothing operation for the posterior probability of Gaussian mixture components in clean speech estimation. Finally, we design a procedure to perform irrelevant variability normalization (IVN) based joint training of a reference Gaussian mixture model (GMM) for feature compensation and hidden Markov models (HMMs) for acoustic modeling using VTS-based feature compensation. The effectiveness of our proposed approach is confirmed by experiments on Aurora3 benchmark database for a real-world in-vehicle connected digits recognition task. Compared with ETSI advanced front-end, our approach achieves significant recognition accuracy improvement across three “training-testing” conditions for four languages.
Jun Du 0002, Qiang Huo
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 An Unsupervised Adaptation Approach to Leveraging Feedback Loop Data by Using i-Vector for Data Clustering and Selection
abstract
We present a study of using unsupervised adaptation approaches to improve speech recognition accuracy of a deployed speech service by leveraging large-scale untranscribed speech data collected from a feedback loop (FBL). For a regular user with lots of adaptation utterances, conventional CMLLR-based adaptation can be used for personalization directly. For a casual user with a few adaptation utterances, we propose to use CMLLR-based adaptation by augmenting his / her adaptation utterances with utterances acoustically close to the user, which are selected from the FBL data by an i-vector based approach. For a new user, we propose to perform a CMLLR-based recognition of an unknown utterance by selecting a set of CMLLR transforms from the most similar cluster, which are pre-trained by using the utterances from the corresponding cluster generated by an i-vector based utterance clustering method from the FBL data. The effectiveness of the above approaches are confirmed by our experiments on a short message dictation task on smart phones.
Jian Xu 0006, Zhijie Yan, Qiang Huo
IEEE ACM Trans. Audio Speech Lang. Process.3
2013 A VTS-based feature compensation approach to noisy speech recognition using mixture models of distortion
abstract
Recently, we proposed an approach to irrelevant variability normalization (IVN) based joint training of a reference Gaussian mixture model (GMM) for feature compensation and hidden Markov models (HMMs) for acoustic modeling by using a vector Taylor series (VTS) based feature compensation technique, where single-component densities are used to model additive noise and convolutional distortion respectively. In this paper, mixtures of densities are used to enhance the distortion model. New formulations for maximum likelihood (ML) estimation of distortion model parameters, and minimum mean squared error (MMSE) estimation of clean speech are derived and presented. A comparative study is conducted under three “training-testing” conditions on Aurora3 database. Experimental results confirm that the proposed mixture models of distortion can achieve significant performance gain compared with the traditional distortion modeling.
Jun Du 0002, Qiang Huo
ICASSP2
2013 Tied-state based discriminative training of context-expanded region-dependent feature transforms for LVCSR
abstract
We present a new discriminative feature transform approach to large vocabulary continuous speech recognition (LVCSR) using Gaussian mixture density hidden Markov models (GMM-HMMs) for acoustic modeling. The feature transform is formulated with a set of context-expanded region-dependent linear transforms (RDLTs) utilizing both long-span features and contextual weight expansion. The RDLTs are estimated by lattice-free, tied-state based discriminative training using maximum mutual information (MMI) criterion, while the GMM-HMMs are trained by conventional lattice-based, boosted MMI training. Compared with two baseline systems, which use RDLTs with either long-span features or weight expansion only and are trained using the conventional lattice-based discriminative training for both RDLTs and HMMs, the proposed approach achieves a relative word error rate reduction of 10% and 6% respectively on Switchboard-1 conversational telephone speech transcription task.
Zhijie Yan, Qiang Huo, Jian Xu 0006, Yu Zhang 0007
ICASSP2
2013 Novel multihop transmission schemes using selective network coding and differential modulation for two-way relay networks
abstract
In this paper, we propose a novel multihop transmission scheme using selective network coding (NC) and differential modulation (SNC-DM) for two-way relay networks (TWRNs) when neither the source nodes nor the relay nodes know the channel state information (CSI). We first develop a bidirectional transmission scheme using NC where the information exchange in a two-way multihop relay network with the arbitrary number of hops can be completed in four transmission phases. As a result, the maximum achievable throughput does not decrease as the number of hops increases. To overcome the error propagation in the multihop transmission with decode-and-forward (DF) protocol in wireless fading channels, a selective NC scheme is proposed. In addition, we apply differential modulation in the proposed scheme to avoid channel estimation in the multihop networks. The performance of the proposed scheme is analyzed, and a closed-form frame error rate (FER) expression is derived. It is shown that the proposed scheme achieves significant improvements in both FER performance and network throughput compared to the conventional multihop DF scheme in TWRNs. The analytical results are verified through numerical simulations.
Qiang Huo, Lingyang Song, Yonghui Li 0001, Bingli Jiao
ICC1
2013 An Irrelevant Variability Normalization Based Discriminative Training Approach for Online Handwritten Chinese Character Recognition
abstract
This paper presents a discriminative training approach to irrelevant variability normalization (IVN) based joint training of feature transforms and prototype-based classifier for recognition of online handwritten Chinese characters. A sample separation margin based minimum classification error criterion is adopted in IVN-based training, while an Rprop algorithm is used for optimizing the objective function. The IVN-trained recognizer can be made both compact and efficient by using a two-level fast-match tree whose internal nodes coincide with the labels of feature transforms. The effectiveness of the proposed approach is confirmed on an online handwritten character recognition task with a vocabulary of 9,306 characters.
Jun Du 0002, Qiang Huo
ICDAR2
2013 An Improved Component Tree Based Approach to User-Intention Guided Text Extraction from Natural Scene Images
abstract
We have proposed previously a component-tree based approach to user-intention guided text extraction from natural scene images. In this paper, in addition to improving the performance of text extraction algorithm for "swipe" gesture, the algorithm has also been extended to support a new mode of using "tap" gesture to indicate the intended text. Given a grayscale image, two component-trees are built and pre-pruned first by using a so-called contrasting extremal region (CER) criterion and simple rules of geometric features. The remaining nodes are enhanced by using color information in a perceptual color space. Then, a pre-trained neural network is used to classify a selected set of enhanced nodes as single-character or non-text objects. The remaining nodes are grouped into candidate text lines, where possible outliers are pruned in individual lines. Finally, the text line "swiped" or "tapped" by a user is selected as the target line and the intended text is extracted accordingly. The proposed algorithm has been evaluated on ICDAR-2003 benchmark dataset and a superior performance is achieved against the previous methods.
Lei Sun 0003, Qiang Huo
ICDAR2
2013 A scalable approach to using DNN-derived features in GMM-HMM based acoustic modeling for LVCSR
abstract
We present a new scalable approach to using deep neural network (DNN) derived features in Gaussian mixture density hidden Markov model (GMM-HMM) based acoustic modeling for large vocabulary continuous speech recognition (LVCSR). The DNN-based feature extractor is trained from a subset of training data to mitigate the scalability issue of DNN training, while GMM-HMMs are trained by using state-of-the-art scalable training methods and tools to leverage the whole training set. In a benchmark evaluation, we used 309-hour Switchboard-I (SWB) training data to train a DNN first, which achieves a word error rate (WER) of 15.4% on NIST-2000 Hub5 evaluation set by a traditional DNN-HMM based approach. When the same DNN is used as a feature extractor and 2,000-hour “SWB+Fisher” training data is used to train the GMM-HMMs, our DNN-GMM-HMM approach achieves a WER of 13.8%. If per-conversation-side based unsupervised adaptation is performed, a WER of 13.1% can be achieved.
Zhijie Yan, Qiang Huo, Jian Xu 0006
INTERSPEECH2
2013 A discriminative linear regression approach to adaptation of multi-prototype based classifiers and its applications for Chinese OCR
Jun Du 0002, Qiang Huo
Pattern Recognit.2
2012 A distributed differential space-time coding scheme with analog network coding in two-way relay networks
abstract
In this paper, we consider general two-way relay networks (TWRNs) with two source and N relay nodes when neither the source nodes nor the relay nodes have access to channel-state information (CSI). A distributed differential space time coding with analog network coding (DDSTC-ANC) scheme is proposed. A simple blind estimation and a differential signal detector are developed to recover the desired signal at each source. The pairwise error probability (PEP) and block error rate (BLER) of the DDSTC-ANC scheme are analyzed. Exact and simplified PEP expressions are derived, which can be used for power allocation between the source and relay nodes. The analytical results are verified through simulations.
Qiang Huo, Lingyang Song, Yonghui Li 0001, Bingli Jiao
GLOBECOM1
2012 Designing compact classifiers for rotation-free recognition of large vocabulary online handwritten Chinese characters
abstract
We present a study of designing compact multiple-prototype based classifiers for rotation-free recognition of online handwritten Chinese characters. Several versions of Rprop algorithms are adopted to optimize a sample-separation-margin based minimum classification error objective function. Split vector quantization technique is used to compress classifier parameters and a fast-match tree is used for efficient recognition. A new preprocessing technique is proposed to achieve rotation-free recognition capability. Promising benchmark results are reported on an online handwritten character recognition task with a vocabulary of 27,720 characters.
Jun Du 0002, Qiang Huo, Kai Chen 0001
ICASSP2
2012 A study of discriminative feature extraction for i-vector based acoustic sniffing in IVN acoustic model training
abstract
Recently, we proposed an i-vector approach to acoustic sniffing for irrelevant variability normalization based acoustic model training in large vocabulary continuous speech recognition (LVCSR). Its effectiveness has been confirmed by experimental results on Switchboard- 1 conversational telephone speech transcription task. In this paper, we study several discriminative feature extraction approaches in i-vector space to improve both recognition accuracy and run-time efficiency. New experimental results are reported on a much larger scale LVCSR task with about 2000 hours training data.
Yu Zhang 0007, Jian Xu 0006, Zhijie Yan, Qiang Huo
ICASSP4
2012 A discriminative linear regression approach to OCR adaptation
Jun Du 0002, Qiang Huo
ICPR2
2012 A component-tree based method for user-intention guided text extraction
Lei Sun 0003, Qiang Huo
ICPR2
2012 IVN-Based Joint Training Of GMM And HMMs Using An Improved VTS-Based Feature Compensation For Noisy Speech Recognition
Jun Du 0002, Qiang Huo
INTERSPEECH2
2012 Cooperative MIMO Channel Modeling and Multi-Link Spatial Correlation Properties
abstract
In this paper, a novel unified channel model framework is proposed for cooperative multiple-input multiple-output (MIMO) wireless channels. The proposed model framework is generic and adaptable to multiple cooperative MIMO scenarios by simply adjusting key model parameters. Based on the proposed model framework and using a typical cooperative MIMO communication environment as an example, we derive a novel geometry-based stochastic model (GBSM) applicable to multiple wireless propagation scenarios. The proposed GBSM is the first cooperative MIMO channel model that has the ability to investigate the impact of the local scattering density (LSD) on channel characteristics. From the derived GBSM, the corresponding multi-link spatial correlation functions are derived and numerically analyzed in detail.
Xiang Cheng 0001, Cheng-Xiang Wang 0001, Haiming Wang 0001, Xiqi Gao 0001, Xiaohu You 0001, Dongfeng Yuan, Bo Ai 0001, Qiang Huo, Lingyang Song, Bingli Jiao
IEEE J. Sel. Areas Commun.8
2012 Performance Analysis of Hybrid Relay Selection in Cooperative Wireless Systems
abstract
The hybrid relay selection (HRS) scheme, which adaptively chooses amplify-and-forward (AF) and decode-and-forward (DF) protocols based on the decoding results at the relay, is very effective to achieve robust performance in wireless relay networks. This paper analyzes the frame error rate (FER) of the HRS scheme in general wireless relay networks without and with utilizing error control coding at the source node. We first develop an improved signal-to-noise ratio (SNR) threshold-based FER approximation model. Then, we derive an analytical average FER expression as well as a high SNR asymptotic expression for the HRS scheme and generalize to other relaying schemes. Simulation results exhibit an excellent agreement with the theoretical analysis, which validates the derived FER expressions.
Tianxi Liu, Lingyang Song, Yonghui Li 0001, Qiang Huo, Bingli Jiao
IEEE Trans. Commun.4
2011 A study of an irrelevant variability normalization based discriminative training approach for LVCSR
abstract
This paper presents a discriminative training (DT) approach to irrelevant variability normalization (IVN) based training of feature transforms and hidden Markov models for large vocabulary continuous speech recognition. A speaker-clustering based method is used for acoustic sniffing and maximum mutual information (MMI) is used as a training criterion. Combined with unsupervised adaptation of feature transforms, the IVN-based DT approach achieves a 14.5% relative word error rate reduction over an MMI-trained baseline system on a Switchboard-1 conversational telephone speech transcription task.
Yu Zhang 0007, Jian Xu 0006, Zhijie Yan, Qiang Huo
ICASSP4
2011 Snap and Translate Using Windows Phone
abstract
We have developed a prototype of a mobile app called "Snap and Translate" on "Windows Phone 7". A person who is reading an English menu/sign and wants a Chinese translation of an English word or phrase or paragraph can use a Windows Phone to snap an image of the text, tap the word or swipe the phrase or circle the paragraph with a finger, and get a Chinese translation displayed on the screen of the phone. This is enabled by seamless integration of three Microsoft technologies: intelligent text extraction, OCR, and machine translation based on a client-plus-cloud architecture. The current prototype also supports Chinese OCR plus Chinese-to-English translation. In this paper, we highlight the UI design of the system and the corresponding user-intention guided text extraction approach to achieving a compelling user experience.
Jun Du 0002, Qiang Huo, Lei Sun 0003
ICDAR2
2011 Text Driven 3D Photo-Realistic Talking Head
Frank K. Soong, Qiang Huo
INTERSPEECH4
2011 An i-vector Based Approach to Acoustic Sniffing for Irrelevant Variability Normalization Based Acoustic Model Training and Speech Recognition
Jian Xu 0006, Yu Zhang 0007, Zhijie Yan, Qiang Huo
INTERSPEECH4
2011 An i-vector Based Approach to Training Data Clustering for Improved Speech Recognition
abstract
We present a new approach to clustering training data for improved speech recognition. Given a training corpus, a so-called i-vector is extracted from each training utterance. A hierarchical divisive clustering algorithm is then used to cluster the training i-vectors into multiple clusters. For each cluster, an acoustic model (AM) is trained accordingly. Such trained multiple AMs can then be used in recognition stage to improve recognition accuracy. The proposed approach is very efficient therefore can deal with very large scale training corpus on current mainstream computing platforms. We report experimental results on a voice search task with 7,500 hours of speech training data.
Yu Zhang 0007, Jian Xu 0006, Zhijie Yan, Qiang Huo
INTERSPEECH4
2011 Building compact recognizers of handwritten Chinese characters using precision constrained Gaussian model, minimum classification error training and parameter compression
Yongqiang Wang 0008, Qiang Huo
Int. J. Document Anal. Recognit.2
2011 A Feature Compensation Approach Using High-Order Vector Taylor Series Approximation of an Explicit Distortion Model for Noisy Speech Recognition
abstract
This paper presents a new feature compensation approach to noisy speech recognition by using high-order vector Taylor series (HOVTS) approximation of an explicit model of environmental distortions. Formulations for maximum-likelihood (ML) estimation of both additive noises and convolutional distortions, and minimum mean squared error (MMSE) estimation of clean speech are derived. Experimental results on Aurora2 and Aurora4 benchmark databases, where the modeling assumption of the distortion model is more accurate, demonstrate that the standard HOVTS-based feature compensation approaches achieve consistently significant improvement in recognition accuracy compared to traditional standard first-order VTS-based approach. For a real-world in-vehicle connected digits recognition task on Aurora3 benchmark database where the modeling assumption of the distortion model is less accurate, modifications are necessary to make VTS-based feature compensation approaches work. In this case, the second-order VTS-based approach performs only slightly better than the first-order VTS-based approach.
Jun Du 0002, Qiang Huo
IEEE Trans. Speech Audio Process.2
2010 Sample-separation-margin based minimum classification error training of pattern classifiers with quadratic discriminant functions
abstract
In this paper, we present a new approach to minimum classification error (MCE) training of pattern classifiers with quadratic discriminant functions. First, a so-called sample separation margin (SSM) is defined for each training sample and then used to define the misclassification measure in MCE formulation. The computation of SSM can be cast as a nonlinear constrained optimization problem and solved efficiently. Experimental results on a large-scale isolated online handwritten Chinese character recognition task demonstrate that SSM-based MCE training not only decreases the empirical classification error, but also pushes the training samples away from the decision boundaries, therefore a good generalization is achieved. Compared with conventional MCE training, an additional 7% to 18% relative error rate reduction is observed in our experiments.
Yongqiang Wang 0008, Qiang Huo
ICASSP2
2010 A Study of Discriminative Training for HMM-Based Online Handwritten Chinese/Japanese Character Recognition
abstract
We present a study of discriminative training of classifiers using both maximum mutual information (MMI) and minimum classification error (MCE) criteria for online handwritten Chinese/Japanese character recognition based on continuous-density hidden Markov models. It is observed that MCE-trained classifiers can achieve a much higher recognition accuracy than that of MMI-trained ones. Benchmark results of MCE-trained classifiers for simplified Chinese, traditional Chinese and Japanese characters are reported on three recognition tasks with a vocabulary of 9119, 20924, and 12333 characters respectively.
Yongqiang Wang 0008, Qiang Huo, Yu Shi 0001
ICFHR2
2010 A Study of Designing Compact Recognizers of Handwritten Chinese Characters Using Multiple-Prototype Based Classifiers
abstract
We present a study of designing compact recognizers of handwritten Chinese characters using multiple-prototype based classifiers. A modified Quick prop algorithm is proposed to optimize a sample-separation-margin based minimum classification error objective function. Split vector quantization technique is used to compress classifier parameters. Benchmark results are reported for classifiers with different footprints trained from about 10 million samples on a recognition task with a vocabulary of 9282 character classes which include 9119 Chinese characters, 62 alphanumeric characters, 101 punctuation marks and symbols.
Yongqiang Wang 0008, Qiang Huo
ICPR2
2010 A study of irrelevant variability normalization based training and unsupervised online adaptation for LVCSR
abstract
This paper presents an experimental study of a maximum likelihood (ML) approach to irrelevant variability normalization (IVN) based training and unsupervised online adaptation for large vocabulary continuous speech recognition. A moving window based frame labeling method is used for acoustic sniffing. The IVN-based approach achieves a 10% relative word error rate reduction over an ML-trained baseline system on a Switchboard-1 conversational telephone speech transcription task.
Guangchuan Shi, Yu Shi 0001, Qiang Huo
INTERSPEECH3
2009 Robust speech recognition based on structured modeling, irrelevant variability normalization and unsupervised online adaptation
abstract
We present a new approach to robust speech recognition based on structured modeling, irrelevant variability normalization (IVN) and unsupervised online adaptation (OLA). In offline training stage, a set of generic HMMs for basic speech units relevant to phonetic classification is trained along with several sets of feature transforms with different degrees of freedom by using a maximum likelihood (ML) IVN-based training strategy. In recognition stage, after a first-pass recognition, the most appropriate set of feature transforms is identified and adapted under ML criterion by using the unknown utterance itself, which is recognized again to achieve better performance by using the adapted feature transforms and the pre-trained generic HMMs. The effectiveness of the proposed approach is confirmed by evaluation experiments on Finnish Aurora3 database.
Qiang Huo, Donglai Zhu
ICASSP1
2009 A Character-Structure-Guided Approach to Estimating Possible Orientations of a Rotated Isolated Online Handwritten Chinese Character
abstract
This paper presents a character-structure-guided approach to estimating possible orientations of a rotated isolated online handwritten Chinese character. Using the estimated orientations, the original distorted sample can be transformed to a normal position, which can be recognized more accurately by using a classifier trained from normal-position samples. The effectiveness of this approach is demonstrated by recognizing rotated samples generated artificially from the popular Nakayosi and Kuchibue Japanese character databases, with average recognition accuracies of 96.05%, 97.35% and 99.13% on top-6, top-12, and top-100 candidates, respectively.
Qiang Huo
ICDAR2
2009 Affine Distortion Compensation for an Isolated Online Handwritten Chinese Character Using Combined Orientation Estimation and HMM-Based Minimax Classification
abstract
This paper presents a new approach to compensating affine distortion of an isolated online handwritten Chinese character. The input sample is first analyzed by using a character-structure-guided orientation estimation approach. If necessary, the orientation hypotheses are refined based on confidence evaluation of two pre-classifiers. Depending on the number of possible orientations, an HMM-based minimax classification approach is then used to estimate an affine transformation against either the original sample or the compensated sample with the previously identified orientation. The final compensated sample can be derived accordingly using the estimated affine transformation. The effectiveness of the proposed approach is demonstrated by recognition experiments using distorted samples generated artificially from the popular Nakayosi and Kuchibue Japanese character databases.
Qiang Huo
ICDAR2
2009 A Study of Feature Design for Online Handwritten Chinese Character Recognition Based on Continuous-Density Hidden Markov Models
abstract
We present a new feature extraction approach to online Chinese handwriting recognition based on continuous-density hidden Markov models (CDHMM). Given an online handwriting sample, a sequence of time-ordered dominant points are extracted first, which include stroke-endings, points corresponding to local extrema of curvature, and points with a large distance to the chords formed by pairs of previously identified neighboring dominant points. Then, at each dominant point, a 6-dimensional feature vector is extracted, which consists of two coordinate features, two delta features, and two double-delta features. Its effectiveness has been confirmed by experiments for a recognition task with a vocabulary of 9119 Chinese characters and CDHMMs trained from about 10 million samples using both maximum likelihood and discriminative training criteria.
Qiang Huo, Yu Shi 0001
ICDAR2
2009 Design Compact Recognizers of Handwritten Chinese Characters Using Precision Constrained Gaussian Models, Minimum Classification Error Training and Parameter Compression
abstract
In our previous work, a precision constrained Gaussian model (PCGM) was proposed for character modeling to design compact recognizers of handwritten Chinese characters. A maximum likelihood training procedure was developed to estimate model parameters from training data. In this paper, we extend the above work by using minimum classification error (MCE) training to improve recognition accuracy and split vector quantization technique to compress model parameters. Compared with the state-of-the-art MCE-trained and compressed classifiers based on modified quadratic discriminant function, PCGM-based classifiers can achieve much better memory-accuracy tradeoff, therefore offer a good solution to designing compact handwriting recognition systems for East Asian languages such as Chinese, Japanese, and Korean.
Yongqiang Wang 0008, Qiang Huo
ICDAR2
2009 Modeling inverse covariance matrices by expansion of tied basis matrices for online handwritten Chinese character recognition
Yongqiang Wang 0008, Qiang Huo
Pattern Recognit.2
2008 A feature compensation approach using piecewise linear approximation of an explicit distortion model for noisy speech recognition
abstract
This paper presents a new feature compensation approach to noisy speech recognition by using piecewise linear approximation (PLA) of an explicit model of environmental distortions. Two traditional approaches, namely vector Taylor series (VTS) and MAX approximations, are two special cases of our proposed approach. Formulations for maximum likelihood (ML) estimation of noise model parameters and minimum mean square error (MMSE) estimation of clean speech are derived. A hybrid approach of using different approximations for different types of noisy speech segments is also proposed. Experimental results on Aurora2 and Aurora3 databases demonstrate that the proposed approaches achieve consistently significant improvements in recognition accuracy compared to the traditional VTS-based feature compensation approach.
Jun Du 0002, Qiang Huo
ICASSP2
2008 Irrelevant variability normalization based HMM training using map estimation of feature transforms for robust speech recognition
abstract
In the past several years, we've been studying feature transformation (FT) approaches to robust automatic speech recognition (ASR) which can compensate for possible "distortions" caused by factors irrelevant to phonetic classification in both training and recognition stages. Several FT functions with different degrees of flexibility have been studied and the corresponding maximum likelihood (ML) training techniques developed. In this paper, we study yet another new FT function which takes the most flexible form of frame-dependent linear transformation. Maximum a posteriori (MAP) estimation is used for estimating FT function parameters to deal with the possible problem of insufficient training data caused by the increased number of model parameters. The effectiveness of the proposed approach is confirmed by evaluation experiments on Finnish Aurora3 database.
Donglai Zhu, Qiang Huo
ICASSP2
2008 A study of a new misclassification measure for minimum classification error training of prototype-based pattern classifiers
abstract
In this paper, we revisit the formulation of minimum classification error (MCE) training and propose a sample separation margin (SSM) based misclassification measure for MCE training of multiple-prototype-based pattern classifiers. Comparative experiments are conducted on the task of the recognition of isolated online handwritten Japanese Kanji characters using Nakayosi and Kuchibue databases. Experimental results demonstrate that MCE training with the new misclassification measure achieves significant character recognition error rate reduction compared with MCE training using two traditional misclassification measures.
Qiang Huo
ICPR2
2008 A study of semi-tied covariance modeling for online handwritten Chinese character recognition
abstract
This paper presents a new approach to large-vocabulary online handwritten Chinese character recognition based on semi-tied covariance (STC) modeling. Detailed procedures are described for estimating the STC model parameters under both maximum likelihood (ML) and minimum classification error (MCE) criteria. Compared with the state-of-the-art modified quadratic discriminant function (MQDF) based classifiers, STC-based classifiers can achieve a better memory-accuracy trade-off, thus provide more flexibility in designing compact online handwritten Chinese character recognizers. Its usefulness has been confirmed and demonstrated by comparative experiments on popular Nakayosi and Kuchibue Japanese character databases.
Yongqiang Wang 0008, Qiang Huo
ICPR2
2008 An ellipsoid constrained quadratic programming (ECQP) approach to MCE training of MQDF-based classifiers for handwriting recognition
abstract
In this study, we propose a novel optimization algorithm for minimum classification error (MCE) training of modified quadratic discriminant function (MQDF) models. An ellipsoid constrained quadratic programming (ECQP) problem is formulated with an efficient line search solution derived, and a subspace combination condition is proposed to simplify the problem in certain cases. We show that under the perspective of constrained optimization, the MCE training of MQDF models can be solved by ECQP with some reasonable approximation, and the hurdle of incomplete covariances can be handled by subspace combination. Experimental results on the Nakayosi/Kuchibue online handwritten Kanji character recognition task show that compared with the conventional generalized probabilistic descent (GPD) algorithm, the new approach achieves about 7% relative error rate reduction.
Yongqiang Wang 0008, Qiang Huo
ICPR3
2008 A speech enhancement approach using piecewise linear approximation of an explicit model of environmental distortions
abstract
This paper presents a speech enhancement approach derived by using a piecewise linear approximation (PLA) of an explicit model of environmental distortions. PLA is a generalization of two traditional approaches, namely vector Taylor series (VTS) and MAX approximations. Formulations are described for both maximum likelihood (ML) estimation of noise model parameters and minimum mean-squared error (MMSE) estimation of clean speech. Evaluation experiments are conducted to enhance speech signals corrupted by several types of additive noises. Compared to the traditional MAX-approximation based approach, our PLA-based speech enhancement approach achieves better performance in terms of two objective quality measures, namely segmental SNR and log-spectral distortion.
Jun Du 0002, Qiang Huo
INTERSPEECH2
2008 A feature compensation approach using high-order vector taylor series approximation of an explicit distortion model for noisy speech recognition
abstract
This paper presents a new feature compensation approach to noisy speech recognition by using high-order vector Taylor series (HOVTS) approximation of an explicit model of environmental distortions. Formulations for maximum-likelihood (ML) estimation of both additive noises and convolutional distortions, and minimum mean squared error (MMSE) estimation of clean speech are derived. Experimental results on Aurora2 and Aurora4 benchmark databases, where the modeling assumption of the distortion model is more accurate, demonstrate that the standard HOVTS-based feature compensation approaches achieve consistently significant improvement in recognition accuracy compared to traditional standard first-order VTS-based approach. For a real-world in-vehicle connected digits recognition task on Aurora3 benchmark database where the modeling assumption of the distortion model is less accurate, modifications are necessary to make VTS-based feature compensation approaches work. In this case, the second-order VTS-based approach performs only slightly better than the first-order VTS-based approach.
Jun Du 0002, Qiang Huo
INTERSPEECH2
2008 A rank-predicted pseudo-greedy approach to efficient text selection from large-scale corpus for maximum coverage of target units
Qiang Huo
INTERSPEECH2
2007 An Approach to Large Margin Design of Prototype-Based Pattern Classifiers
abstract
In this paper, we propose a maximum separation margin (MSM) training method for multiple-prototype (MP)-based pattern classifiers in which a sample separation margin defined as the distance from the training sample to the classification boundary can be calculated precisely. Similar to support vector machine (SVM) methodology, MSM training is formulated as a multicriteria optimization problem which aims at maximizing the separation margin and minimizing the empirical error rate on training data simultaneously. By making certain relaxation assumptions, MSM training can be reformulated as a semidefinite programming (SDP) problem that can be solved efficiently by some standard optimization algorithms designed for SDP. Evaluation experiments are conducted on the task of the recognition of most confusable Kanji character pairs identified from popular Nakayosi and Kuchibue handwritten Japanese character databases. It is observed that the MSM-trained MP-based classifier achieves a similar character recognition accuracy as that of the state-of-the-art SVM-based classifier, yet requires much fewer classifier parameters.
Qiang Huo
ICASSP (2)3
2007 A Maximum Likelihood Approach to Unsupervised Online Adaptation of Stochastic Vector Mapping Function for Robust Speech Recognition
abstract
In the past several years, we've been studying feature transformation approaches for robust automatic speech recognition (ASR) based on the concept of stochastic vector mapping (SVM) to compensating for possible "distortions" caused by factors irrelevant to phonetic classification in both training and recognition stages. Although we have demonstrated the usefulness of the SVM-based approaches for several robust ASR applications where diversified yet representative training data are available, the performance improvement of SVM-based approaches is less significant when there is a severe mismatch between training and testing conditions. In this paper, we present a maximum likelihood approach to unsupervised online adaptation (OLA) of SVM function parameters on an utterance-by-utterance basis for achieving further performance improvement. Its effectiveness is confirmed by evaluation experiments on Finnish AuroraS database.
Donglai Zhu, Qiang Huo
ICASSP (4)2
2007 A Generalized Feature Transformation Approach for Channel Robust Speaker Verification
abstract
In this paper we propose a generalized feature transformation approach to compensating for channel variation in speaker verification (SV) applications. Channel-dependent (CD) piecewise linear transformations are used for feature compensation. CD transformation parameters are estimated together with a channel-independent (CI) root Gaussian mixture model (GMM) from training data with a variety of channel conditions by using a maximum likelihood criterion. Experiments are conducted on the 2005 NIST Speaker Recognition Evaluation (SRE) corpus for several text-independent GMM-based SV systems. Experimental results show that the proposed approach achieves relative equal error rate (EER) reductions of 8.19% and 26.24% in comparison with a traditional feature mapping approach and a baseline system, respectively.
Donglai Zhu, Bin Ma 0001, Haizhou Li 0001, Qiang Huo
ICASSP (4)4
2007 Kernel Modified Quadratic Discriminant Function for Online Handwritten Chinese Characters Recognition
abstract
The modified quadratic discriminant function has been used successfully in handwriting recognition, which can be seen as a dot-product method by eigen- decomposition of the covariance matrix. Therefore, it is possible to expand MQDF to high dimension space by kernel trick. This paper presents a new kernel- based method, Kernel modified quadratic discriminant function (KMQDF) for online Chinese Characters Recognition. Experimental results show that the performance of MQDF is improved by the kernel approach.
Qiang Huo
ICDAR3
2007 Irrelevant variability normalization based HMM training using VTS approximation of an explicit model of environmental distortions
Qiang Huo
INTERSPEECH2
2007 An active approach to speaker and task adaptation based on automatic analysis of vocabulary confusability
abstract
In the field of automatic speech recognition (ASR), speaker and task adaptation has become a very important research topic in recent years. Although recognition performance obtained by a speaker independent (SI) ASR system can be sufficiently good for many tasks, there still exists a large performance gap between an SI system and a speaker dependent (SD) ASR system. Although a, number of successful adaptation techniques have been developed, which modify speaker independent model parameters to favor the speaker and task with given adaptation data, little attention has been given to designing the adaption script effectively and efficiently to make the adapted system benefit most from the data. I have therefore chosen this aspect as the research topic in this thesis. In recent years, the idea of active learning has been applied to several ASR applications. Given the large amount of unlabelled training data, manual annotation efforts can be reduced by intelligently selecting only a subset of training data for manual labelling most useful for learning purposes of building up the application at hand. In this thesis, the concept of active learning is extended to the scenario of supervised speaker and task adaptation, where the system takes the initiative of eliciting small (ideally minimum) amount of adaptation data from the user for achieving high (ideally maximum) performance improvement by adapting the HMMs using the elicited adaptation data. Based on the concept of active learning, the task vocabulary confusability, which is highly related to the difficulty of the given task, is analyzed by using a new DTW-based HMM dissimilarity measure. The adaptation script is then generated effectively according to the vocabulary confusability based information. In this thesis, the adaptation script generation problem is cast as two constrained optimization problems with the same constraints but different objective functions. The first problem is maximum coverage problem with a Knapsack constraint problem. The second is nonlinear binary optimization problem with linear constraints. Two new approaches, namely, rank predicted pseudo-greedy approach and variable-depth approach with gradient projection guidance, are proposed to resolve these two optimization problems respectively. The active approach with the efficient adaptation script generation by using these two new approaches can generate the adaptation script much faster than the traditional approach without sacrificing recognition performance. Comparative experiments are designed and conducted for a simple application scenario involving searching an item from a long list via voice. The experimental results demonstrate that the proposed active adaptation strategy performs much better than traditional passive adaptation strategies.
Qiang Huo
INTERSPEECH1
2007 A Study of Minimum Classification Error (MCE) Linear Regression for Supervised Adaptation of MCE-Trained Continuous-Density Hidden Markov Models
abstract
In this paper, we present a formulation of minimum classification error linear regression (MCELR) for the adaptation of Gaussian mixture continuous-density hidden Markov model (CDHMM) parameters. Two optimization approaches, namely generalized probabilistic descent (GPD) and Quickprop are studied and compared for the optimization of the MCELR objective function. The effectiveness of the proposed MCELR technique is confirmed via a series of supervised speaker adaptation experiments on a task of continuous Putonghua (Mandarin Chinese) speech recognition
Jian Wu 0029, Qiang Huo
IEEE Trans. Speech Audio Process.2
2006 Unsupervised Online Adaptation of Segmental Switching Linear Gaussian Hidden Markov Models for Robust Speech Recognition
abstract
In our previous works, a Segmental Switching Linear Gaussian Hidden Markov Model (SSLGHMM) was proposed to model "noisy" speech utterance for robust speech recognition. Both ML (maximum likelihood) and MCE (minimum classification error) training procedures were developed for training model parameters and their effectiveness was confirmed by evaluation experiments on Aurora2 and Aurora3 databases. In this paper, we present an ML approach to unsupervised online adaptation (OLA) of SSLGHMM parameters for achieving further performance improvement. An important implementation issue of how to initialize the switching linear Gaussian model parameters is also studied. Evaluation results on Finnish Aurora3 database show that in comparison with the performance of a baseline system based on ML-trained SSLGHMMs, unsupervised OLA yields a relative word error rate reduction of 4.3%, 9.1%, and 17.8% for well-matched, medium-mismatched, and high-mismatched conditions respectively.
Qiang Huo, Donglai Zhu, Jian Wu 0029
ICASSP (1)1
2006 A DTW-based dissimilarity measure for left-to-right hidden Markov models and its application to word confusability analysis
Qiang Huo
INTERSPEECH1
2006 A maximum likelihood training approach to irrelevant variability compensation based on piecewise linear transformations
Qiang Huo, Donglai Zhu
INTERSPEECH1
2006 An Environment-Compensated Minimum Classification Error Training Approach Based on Stochastic Vector Mapping
abstract
A conventional feature compensation module for robust automatic speech recognition is usually designed separately from the training of hidden Markov model (HMM) parameters of the recognizer, albeit a maximum-likelihood (ML) criterion might be used in both designs. In this paper, we present an environment-compensated minimum classification error (MCE) training approach for the joint design of the feature compensation module and the recognizer itself. The feature compensation module is based on a stochastic vector mapping function whose parameters have to be learned from stereo data in a previous approach called SPLICE. In our proposed MCE joint design approach, by initializing the parameters with an approximate ML training procedure, the requirement of stereo data can be removed. By evaluating the proposed approach on Auroral connected digits database, a digit recognition error rate, averaged on all three test sets, of 5.66% is achieved for multicondition training. In comparison with the performance achieved by the baseline system using ETSI advanced front-end, our approach achieves an additional overall error rate reduction of 12.4%
Jian Wu 0029, Qiang Huo
IEEE Trans. Speech Audio Process.2
2005 An Environment Compensated Maximum Likelihood Training Approach Based on Stochastic Vector Mapping
abstract
Several recent approaches for robust speech recognition are developed based on the concept of stochastic vector mapping (SVM) that perform a frame-dependent bias removal to compensate for environmental variabilities in both training and recognition stages. Some of them require stereo recordings of both clean and noisy speech for the estimation of SVM function parameters. In this paper, we present a detailed formulation of a maximum likelihood training approach for the joint design of SVM function parameters and HMM parameters of a speech recognizer that does not rely on the availability of stereo training data. Its learning behavior and effectiveness is demonstrated by using the experimental results on the Aurora3 Finnish connected digits database recorded by using both close-talking and hands-free microphones in cars.
Jian Wu 0029, Qiang Huo, Donglai Zhu
ICASSP (1)2
2005 A Study On the Use of 8-Directional Features For Online Handwritten Chinese Character Recognition
abstract
This paper presents a study of using 8-directional features for online handwritten Chinese character recognition. Given an online handwritten character sample, a series of processing steps, including linear size normalization, adding imaginary strokes, nonlinear shape normalization, equidistance resampling, and smoothing, are performed to derive a 64/spl times/64 normalized online character sample. Then, 8-directional features are extracted from each online trajectory point, and 8 directional pattern images are generated accordingly, from which blurred directional features are extracted at 8/spl times/8 uniformly sampled locations using a filter derived from the Gaussian envelope of a Gabor filter. Finally, a 512-dimensional vector of raw features is formed. Extensive experiments on the task of recognizing 3755 level-1 Chinese characters in GB2312-80 standard are performed to compare and discern the best setting for several algorithmic choices and control parameters. The effectiveness of the studied approach is confirmed.
Zhen-Long Bai, Qiang Huo
ICDAR2
2005 A Data Structure Using Hashing and Tries For Efficient Chinese Lexical Access
abstract
A lexicon is needed in many applications. In the past, different structures such as tries, hash tables and their variants have been investigated for lexicon organization and lexical access. In this paper, we propose a new data structure that combines the use of hash table and tries for storing a Chinese lexicon. The data structure facilitates an efficient lexical access yet requires less memory than that of a trie lexicon. Experiments are conducted to evaluate its performance for in-vocabulary lexical access, out-of-vocabulary word rejection, and substring matching. The effectiveness of the proposed approach is confirmed.
Yat-Kin Lam, Qiang Huo
ICDAR2
2004 A study of minimum classification error training for segmental switching linear Gaussian hidden Markov models
Jian Wu 0029, Donglai Zhu, Qiang Huo
INTERSPEECH3
2003 Modelling uncertainty in stochastic vector mapping with minimum classification error training for robust speech recognition
abstract
We have witness several works of considering the uncertainty of feature compensation module for robust speech recognition. In most of these studies, the modelling and the exploiting of the uncertainty are seldom treated in a unified way. In this paper, we present a new framework, which casts the problem of considering the uncertainty of feature compensation module as the one of designing a new discriminant function, thus the uncertainty parameters of the feature compensation module and other parameters of the discriminant function can be estimated jointly under a consistent criterion of minimum classification error (MCE). It is hoped that such MCE-trained discriminant function can improve the performance of a maximum discriminant function based speech recognition system. The preliminary experimental results on Aurora2 multi-condition tasks have confirmed the above conjecture.
Jian Wu 0029, Qiang Huo
ICASSP (2)2
2003 An Approach to Extracting the Target Text Line from a Document Image Captured by a Pen Scanner
abstract
In this paper, we present a new approach to extracting the target text line from a document image captured by a pen scanner. Given the binary image, a set of possible text lines are first formed by nearest-neighbor grouping of connected components (CC). They are then refined by text line merging and adding the missed CCs. The possible target text line is identified by using a geometric feature based score function and fed to an OCR engine for character recognition. If the recognition result is confident enough, the target text line is accepted. Otherwise, all the remaining text lines are fed to the OCR engine to verify whether an alternative target text line exists or the whole image should be rejected. The effectiveness of the above approach is confirmed by experiments on a testing database consisting of 117 document images captured by C-Pen and ScanEye pen scanners.
Zhen-Long Bai, Qiang Huo
ICDAR2
2003 Improving Chinese/English OCR Performance by Using MCE-based Character-Pair Modeling and Negative Training
abstract
In the past several years, we've been developing a high performance OCR engine for machine printed Chinese/ English documents. We have reported previously (1) how to use character modeling techniques based on MCE (minimum classification error) training to achieve the high recognition accuracy, and (2) how to use confidence-guided progressive search and fast match techniques to achieve the high recognition efficiency. In this paper, we present two more techniques that help reduce search errors and improve the robustness of our character recognizer. They are (1) to use MCE-trained character-pair models to avoid error-prone character-level segmentation for some trouble cases, and (2) to perform a MCE-based negative training to improve the rejection capability of the recognition models on the hypothesized garbage images during recognition process. The efficacy of the proposed techniques is confirmed by experiments in a benchmark test.
Qiang Huo, Zhi-Dan Feng
ICDAR1
2003 Several HKU approaches for robust speech recognition and their evaluation on Aurora connected digit recognition tasks
Jian Wu 0029, Qiang Huo
INTERSPEECH2
2003 A switching linear Gaussian hidden Markov model and its application to nonstationary noise compensation for robust speech recognition
abstract
In our previous works, a Switching Linear Gaussian Hidden Markov Model (SLGHMM) and its segmental derivative, SSLGHMM, were proposed to cast the problem of modelling a noisy speech utter-ance by a well-designed dynamic Bayesian network. We presented parameter learning procedures for both models with maximum likelihood (ML) criterion. The effectiveness of such models was confirmed by evaluation experiments on Aurora2 database. In this paper, we present a study of minimum classification error (MCE) training for SSLGHMM and discuss its relation to our ear-lier proposals based on stochastic vector mapping. An important implementation issue of SSLGHMM, namely the specification of switching states for a given utterance, is also studied. New evalua-tion results on Aurora3 database show that MCE-trained SSLGH-MMs achieve a relative error reduction of 21 % over a baseline sys-tem based on ML-trained continuous density HMMs (CDHMMs). 1.
Jian Wu 0029, Qiang Huo
INTERSPEECH2
2002 Offline recognition of handwritten Chinese characters using Gabor features, CDHMM modeling and MCE training
abstract
We've been developing a Chinese OCR engine for handwritten Chinese scripts. Currently, our OCR engine supports a vocabulary of 4616 characters which include 4516 simplified Chinese characters in GB2312-80, 62 alphanumeric characters, 38 punctuation marks and symbols. By using 1,384,800 character samples to train our recognizer, an averaged character recognition accuracy of 96.34% is achieved on a testing set of 1,025,535 character samples. An arguably best Chinese OCR product on the market achieves an accuracy of 94.07% for the recognizable Chinese characters in the above testing set. In this paper, we describe key techniques used in our recognizer that contribute to the high recognition accuracy, namely the use of Gabor features and their spatial derivatives as raw features, the use of LDA for feature extraction and dimension reduction, the use of CDHMMs for modeling Chinese characters along both horizontal and vertical directions, and the use of minimum classification error as a criterion for model training.
Qiang Huo, Zhi-Dan Feng
ICASSP2
2002 Supervised adaptation of MCE-trained CDHMMS using minimum classification error linear regression
abstract
In this paper, we present a formulation of minimum classification error linear regression (MCELR) for adaptation of Gaussian mixture continuous density HMM (CDHMM) parameters. We demonstrate that the MCELR can be used to adapt the MCE-trained HMM parameters under a consistent criterion. In a supervised speaker adaptation application, we observe that such adapted models perform better than the ones adapted using MLLR from the ML-trained seed models. We also observe that the MCELR performs consistently better than the MLLR for either sets of seed models.
Jian Wu 0029, Qiang Huo
ICASSP2
2002 Using time-stretched pulses for accurate splitting of speech utterances played back in noisy reverberant environments
Dorothea Kolossa, Qiang Huo
INTERSPEECH2
2002 An environment compensated minimum classification error training approach and its evaluation on Aurora2 database
Jian Wu 0029, Qiang Huo
INTERSPEECH2
2001 High performance Chinese OCR based on Gabor features, discriminative feature extraction and model training
abstract
We have developed a Chinese OCR engine for machine printed documents. Currently, our OCR engine can support a vocabulary of 6921 characters which include 6707 simplified Chinese characters in GB2312-80, 12 frequently used GBK Chinese characters, 62 alphanumeric characters, 140 punctuation marks and symbols. The supported font styles include Song, Fang Song, Kat, He, Yuan, LiShu, WeiBei, XingKai, etc. The averaged character recognition accuracy is above 99% for newspaper quality documents with a recognition speed of about 250 characters per second on a Pentium III-450 MHz PC yet only consuming less than 2 MB memory. We describe the key technologies we used to construct the above recognizer. Among them, we highlight three key techniques contributing to the high recognition accuracy, namely the use of Gabor features, the use of discriminative feature extraction, and the use of minimum classification error as a criterion for model training.
Qiang Huo, Zhi-Dan Feng
ICASSP1
2001 A Discrete Contextual Stochastic Model for the Offline Recognition of Handwritten Chinese Characters
abstract
We study a discrete contextual stochastic (CS) model for complex and variant patterns like handwritten Chinese characters. Three fundamental problems of using CS models for character recognition are discussed, and several practical techniques for solving these problems are investigated. A formulation for discriminative training of CS model parameters is also introduced and its practical usage investigated. To illustrate the characteristics of the various algorithms, comparative experiments are performed on a recognition task with a vocabulary consisting of 50 pairs of highly similar handwritten Chinese characters. The experimental results confirm the effectiveness of the discriminative training for improving recognition performance.
Qiang Huo, Chorkin Chan
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Robust speech recognition based on adaptive classification and decision strategies
Qiang Huo
Speech Commun.1
2001 Online adaptive learning of continuous-density hidden Markov models based on multiple-stream prior evolution and posterior pooling
abstract
We introduce a new adaptive Bayesian learning framework, called multiple-stream prior evolution and posterior pooling, for online adaptation of the continuous density hidden Markov model (CDHMM) parameters. Among three architectures we proposed for this framework, we study in detail a specific two stream system where linear transformations are applied to the mean vectors of the CDHMMs to control the evolution of their prior distribution. This new stream of prior distribution can be combined with another stream of prior distribution evolved without any constraints applied. In a series of speaker adaptation experiments on the task of continuous Mandarin speech recognition, we show that the new adaptation algorithm achieves a similar fast-adaptation performance as that of the incremental maximum likelihood linear regression (MLLR) in the case of small amount of adaptation data, while maintains the good asymptotic convergence property as that of our previously proposed quasi-Bayes adaptation algorithms.
Qiang Huo, Bin Ma 0001
IEEE Trans. Speech Audio Process.1
2000 Robust speech recognition based on off-line elicitation of multiple priors and on-line adaptive prior fusion
Qiang Huo, Bin Ma 0001
INTERSPEECH1
2000 On adaptive decision rules and decision parameter adaptation for automatic speech recognition
abstract
Recent advances in automatic speech recognition are accomplished by designing a plug-in maximum a posteriori decision rule such that the forms of the acoustic and language model distributions are specified and the parameters of the assumed distributions are estimated from a collection of speech and language training corpora. Maximum-likelihood point estimation is by far the most prevailing training method. However, due to the problems of unknown speech distributions, sparse training data, high spectral and temporal variabilities in speech, and possible mismatch between training and testing conditions, a dynamic training strategy is needed. To cope with the changing speakers and speaking conditions in real operational conditions for high-performance speech recognition, such paradigms incorporate a small amount of speaker and environment specific adaptation data into the training process. Bayesian adaptive learning is an optimal way to combine prior knowledge in an existing collection of general models with a new set of condition-specific adaptation data. In this paper, the mathematical framework for Bayesian adaptation of acoustic and language model parameters is first described. Maximum a posteriori point estimation is then developed for hidden Markov models and a number of useful parameters densities commonly used in automatic speech recognition and natural language processing.
Qiang Huo
Proc. IEEE2
2000 A Bayesian predictive classification approach to robust speech recognition
abstract
We introduce a new decision strategy called Bayesian predictive classification (BPC) for robust speech recognition where an unknown mismatch between the training and testing conditions exists. We then propose and focus on one of the approximate BPC approaches called quasi-Bayes predictive classification (QBPC). In a series of comparative experiments where the mismatch is caused by additive white Gaussian noise, we show that the proposed QBPC approach achieves a considerable improvement over the conventional plug-in MAP decision rule.
Qiang Huo
IEEE Trans. Speech Audio Process.1
1999 Irrelevant variability normalization in learning HMM state tying from data based on phonetic decision-tree
abstract
We propose to apply the concept of irrelevant variability normalization to the general problem of learning structure from data. Because of the problems of a diversified training data set and/or possible acoustic mismatches between training and testing conditions, the structure learned from the training data by using a maximum likelihood training method will not necessarily generalize well on mismatched tasks. We apply the above concept to the structural learning problem of phonetic decision-tree based hidden Markov model (HMM) state tying. We present a new method that integrates a linear-transformation based normalization mechanism into the decision-tree construction process to make the learned structure have a better modeling capability and generalizability. The viability and efficacy of the proposed method are confirmed in a series of experiments for continuous speech recognition of Mandarin Chinese.
Qiang Huo, Bin Ma 0001
ICASSP1
1999 On-line adaptive learning of CDHMM parameters based on multiple-stream prior evolution and posterior pooling
Qiang Huo, Bin Ma 0001
EUROSPEECH1
1999 Improving Viterbi Bayesian predictive classification via sequential bayesian learning in robust speech recognition
Hui Jiang 0001, Keikichi Hirose, Qiang Huo
Speech Commun.3
1999 Robust speech recognition based on a Bayesian prediction approach
abstract
We study a category of robust speech recognition problem in which mismatches exist between training and testing conditions, and no accurate knowledge of the mismatch mechanism is available. The only available information is the test data along with a set of pretrained Gaussian mixture continuous density hidden Markov models (CDHMMs). We investigate the problem from the viewpoint of Bayesian prediction. A simple prior distribution, namely constrained uniform distribution, is adopted to characterize the uncertainty of the mean vectors of the CDHMMs. Two methods, namely a model compensation technique based on Bayesian predictive density and a robust decision strategy called Viterbi Bayesian predictive classification are studied. The proposed methods are compared with the conventional Viterbi decoding algorithm in speaker-independent recognition experiments on isolated digits and TI connected digit strings (TIDTGITS), where the mismatches between training and testing conditions are caused by: (1) additive Gaussian white noise, (2) each of 25 types of actual additive ambient noises, and (3) gender difference. The experimental results show that the adopted prior distribution and the proposed techniques help to improve the performance robustness under the examined mismatch conditions.
Hui Jiang 0001, Keikichi Hirose, Qiang Huo
IEEE Trans. Speech Audio Process.3
1998 Improving Viterbi Bayesian predictive classification via sequential Bayesian learning in robust speech recognition
abstract
We extend our previously proposed Viterbi Bayesian predictive classification (VBPC) algorithm to accommodate a new class of prior probability density function (PDF) for continuous density hidden Markov model (CDHMM) based robust speech recognition. The initial prior PDF of CDHMM is assumed to be a finite mixture of natural conjugate prior PDF's of its complete-data density. With the new observation data, the true posterior PDF is approximated by the same type of finite mixture PDF's which retain the required most significant terms in the true posterior density according to their contribution to the corresponding predictive density. Then the updated mixture PDF is used to improve the VBPC performance. The experimental results on a speaker-independent recognition task of isolated Japanese digits confirm the viability and the usefulness of the proposed technique.
Hui Jiang 0001, Keikichi Hirose, Qiang Huo
ICASSP3
1998 A study of prior sensitivity for Bayesian predictive classification based robust speech recognition
abstract
We previously introduced a new Bayesian predictive classification (BPC) approach to robust speech recognition and showed that the BPC is capable of coping with many types of distortions. We also learned that the efficacy of the BPC algorithm is influenced by the appropriateness of the prior distribution for the mismatch being compensated. If the prior distribution fails to characterize the variability reflected in the model parameters, then the BPC will not help much. We show how the knowledge and/or experience of the interaction between the speech signal and the possible mismatch guide us to obtain a better prior distribution which improves the performance of the BPC approach.
Qiang Huo
ICASSP1
1998 A minimax search algorithm for CDHMM based robust continuous speech recognition
Hui Jiang 0001, Keikichi Hirose, Qiang Huo
ICSLP3
1998 On-line adaptive learning of the correlated continuous density hidden Markov models for speech recognition
abstract
We extend our previously proposed quasi-Bayes adaptive learning framework to cope with the correlated continuous density hidden Markov models (HMMs) with Gaussian mixture state observation densities in which all mean vectors are assumed to be correlated and have a joint prior distribution. A successive approximation algorithm is proposed to implement the correlated mean vectors' updating. As an example, by applying the method to an on-line speaker adaptation application, the algorithm is experimentally shown to be asymptotically convergent as well as being able to enhance the efficiency and the effectiveness of the Bayes learning by taking into account the correlation information between different model parameters. The technique can be used to cope with the time-varying nature of some acoustic and environmental variabilities, including mismatches caused by changing speakers, channels, transducers, environments, and so on.
Qiang Huo
IEEE Trans. Speech Audio Process.1
1997 Robust speech recognition based on Viterbi Bayesian predictive classification
abstract
In this paper, we investigate a new Bayesian predictive classification (BPC) approach to realize robust speech recognition when there exist mismatches between training and test conditions but no accurate knowledge of the mismatch mechanism is available. A specific approximate BPC algorithm called Viterbi BPC (VBPC) is proposed for both isolated word and continuous speech recognition. The proposed VBPC algorithm is compared with conventional Viterbi decoding algorithm on speaker-independent isolated digit and connected digit string (TIDIGITS) recognition tasks. The experimental results show that VBPC can considerably improve robustness when mismatches exist between training and testing conditions.
Hui Jiang 0001, Keikichi Hirose, Qiang Huo
ICASSP3
1997 A Bayesian predictive classification approach to robust speech recognition
abstract
We introduce a new Bayesian predictive classification (BPC) approach to robust speech recognition and apply the BPC framework to Gaussian mixture continuous density hidden Markov model based speech recognition. We propose and focus on one of the approximate BPC approaches called quasi-Bayesian predictive classification (QBPC). In comparison with the standard plug-in maximum a posteriori decoding, when the QBPC method is applied to speaker independent recognition of a confusable vocabulary namely 26 English letters, where a broad range of mismatches between training and testing conditions exist, the QBPC achieves around 14% relative recognition error rate reduction. While the QBPC method is applied to cross-gender testing on a less confusable vocabulary, namely 20 English digits and commands, the QBPC method achieves around 24% relative recognition error rate reduction.
Qiang Huo, Hui Jiang 0001
ICASSP1
1997 Combined on-line model adaptation and Bayesian predictive classification for robust speech recognition
Qiang Huo
EUROSPEECH1
1997 On-line adaptive learning of the continuous density hidden Markov model based on approximate recursive Bayes estimate
abstract
We present a framework of quasi-Bayes (QB) learning of the parameters of the continuous density hidden Markov model (CDHMM) with Gaussian mixture state observation densities. The QB formulation is based on the theory of recursive Bayesian inference. The QB algorithm is designed to incrementally update the hyperparameters of the approximate posterior distribution and the CDHMM parameters simultaneously. By further introducing a simple forgetting mechanism to adjust the contribution of previously observed sample utterances, the algorithm is adaptive in nature and capable of performing an online adaptive learning using only the current sample utterance. It can, thus, be used to cope with the time-varying nature of some acoustic and environmental variabilities, including mismatches caused by changing speakers, channels, and transducers. As an example, the QB learning framework is applied to on-line speaker adaptation and its viability is confirmed in a series of comparative experiments using a 26-letter English alphabet vocabulary.
Qiang Huo
IEEE Trans. Speech Audio Process.1
1996 A study of on-line quasi-Bayes adaptation for CDHMM-based speech recognition
abstract
We present a framework of quasi-Bayes (QB) learning of the parameters of the continuous density hidden Markov model (CDHMM) with Gaussian mixture state observation densities. Based on the theory of recursive Bayesian inference, the QB algorithm is designed to incrementally update the hyperparameters on the approximate posterior distribution and the CDHMM parameters simultaneously. By further introducing a simple forgetting mechanism to adjust the contribution of previously observed sample utterances, the algorithm is adaptive in nature and capable of performing an on-line adaptive learning using only the current sample utterance. It can thus be used to cope with the time-varying nature of some acoustic and environmental variabilities, including mismatches caused by changing speakers, channels, and transducers. As an example, the QB learning framework is applied to on-line speaker adaptation and its viability is confirmed in a series of comparative experiments using a 26-letter English alphabet vocabulary.
Qiang Huo
ICASSP1
1996 On-line adaptive learning of the correlated continuous density hidden Markov models for speech recognition
abstract
We extend our previously proposed quasi-Bayes adaptive learning framework to cope with the correlated continuous density hidden Markov models (HMM's) with Gaussian mixture state observation densities in which all mean vectors are assumed to be correlated and have a joint prior distribution.A successive approximation algorithm is proposed to implement the correlated mean vectors' updating.As an example, by applying the method to on-line speaker adaptation application, the algorithm is experimentally shown to be asymptotically convergent as well as being able to enhance the efficiency and the effectiveness of the Bayes learning by taking into account the correlation information between different model parameters.The technique can be used to cope with the time-varying nature of some acoustic and environmental variabilities, including mismatches caused by changing speakers, channels, transducers, environments, and so on.
Qiang Huo
ICSLP1
1996 A study on the use of bi-directional contextual dependence in Markov random field-based acoustic modelling for speech recognition
Qiang Huo, Chorkin Chan
Comput. Speech Lang.1
1996 On-line adaptation of the SCHMM parameters based on the segmental quasi-Bayes learning for speech recognition
abstract
On-line quasi-Bayes adaptation of the mixture coefficients and mean vectors in semicontinuous hidden Markov model (SCHMM) is studied. The viability of the proposed algorithm is confirmed and the related practical issues are addressed in a specific application of on-line speaker adaptation using a 26-word English alphabet vocabulary.
Qiang Huo, Chorkin Chan
IEEE Trans. Speech Audio Process.1
1995 On-line Bayes adaptation of SCHMM parameters for speech recognition
abstract
On-line adaptation of semi-continuous (or tied mixture) hidden Markov model (SCHMM) is studied. A theoretical formulation of the segmental quasi-Bayes learning of the mixture coefficients in SCHMM for speech recognition is presented. The practical issues related to the use of this algorithm for on-line speaker adaptation are addressed. A pragmatic on-line adaptation approach to combine the long-term adaptation of the mixture coefficients and the short-term adaptation of the mean vectors of the Gaussian mixture components are also proposed. The viability of these techniques are confirmed in a series of comparative experiments using a 26-word English alphabet vocabulary.
Qiang Huo, Chorkin Chan
ICASSP1
1995 Contextual vector quantization modeling of hand-printed Chinese character recognition
abstract
A hand-printed Chinese character recognizer based on contextual vector quantization (CVQ) has been built. The idea of CVQ is to quantize each pixel to a codeword by considering not just the pixel itself but its neighbors and their codeword identities as well. 100 samples of each character are collected from 100 writers, among them, 92 are used for training and 8 for testing. The characters are scanned by a 300 dpi scanner, which are then noise removed, thinned, segmented and size normalized. Stroke counts and segment strengths are adopted as observation features. For a vocabulary of 470 simplified Chinese characters, a recognition rate of 97% is achieved.
Sau-Lai Leung, Ping-Chong Chee, Chorkin Chan, Qiang Huo
ICIP (3)4
1995 Discriminative training of HMM based speech recognizer with gradient projection method
Qiang Huo, Chorkin Chan
EUROSPEECH1
1995 On the use of bi-directional contextual dependence in acoustic modeling for speech recognition
Qiang Huo, Chorkin Chan
EUROSPEECH1
1995 Contextual vector quantization for speech recognition with discrete hidden Markov model
Qiang Huo, Chorkin Chan
Pattern Recognit.1
1995 Bayesian adaptive learning of the parameters of hidden Markov model for speech recognition
abstract
A theoretical framework for Bayesian adaptive training of the parameters of a discrete hidden Markov model (DHMM) and of a semi-continuous HMM (SCHMM) with Gaussian mixture state observation densities is presented. In addition to formulating the forward-backward MAP (maximum a posteriori) and the segmental MAP algorithms for estimating the above HMM parameters, a computationally efficient segmental quasi-Bayes algorithm for estimating the state-specific mixture coefficients in SCHMM is developed. For estimating the parameters of the prior densities, a new empirical Bayes method based on the moment estimates is also proposed. The MAP algorithms and the prior parameter specification are directly applicable to training speaker adaptive HMMs. Practical issues related to the use of the proposed techniques for HMM-based speaker adaptation are studied. The proposed MAP algorithms are shown to be effective especially in the cases in which the training or adaptation data are limited.>
Qiang Huo, Chorkin Chan
IEEE Trans. Speech Audio Process.1
1994 Bayesian learning of the SCHMM parameters for speech recognition
abstract
A theoretical framework for Bayesian adaptive learning of semi-continuous HMM parameters is presented. Formulations of MAP estimation of SCHMM parameters are developed. An empirical Bayes method to estimate the hyperparameters of prior densities based on the moment estimate is proposed. Practical issues related to the use of the proposed technique for speaker adaptation application are studied. Effects of various adaptation schemes are examined and their viability is confirmed in a series of comparative experiments using a 26-word English alphabet vocabulary. The proposed method is applicable to other problems in HMM training for speech recognition such as sequential training, context adaptation and parameter smoothing.>
Qiang Huo, Chorkin Chan
ICASSP (1)1
1993 Bayesian learning of the parameters of discrete and tied mixture HMMs for speech recognition
Qiang Huo, Chorkin Chan
EUROSPEECH1
1993 The gradient projection method for the training of hidden Markov models
Qiang Huo, Chorkin Chan
Speech Commun.1