Hui Zhang 0031

dblp:z/HuiZhang31 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Construction and Evaluation of Large Language Models for the Mongolian Medicine Diagnostic and Treatment System
Jixieqi Bai, Feilong Bao, Hui Zhang 0031, Aruukhan Bai
KSEM (6)3
2025 Enhancing Multi-Channel Speech with Limited Microphones via Spherical Harmonic Transform
abstract
The performance of traditional beamforming algorithms is influenced by the number of microphones, with performance improving as the number increases. However, in practice, the number of microphones is often limited. In this paper, we propose a novel virtual microphone estimation method that combines the strengths of both traditional and neural network-based approaches using the spherical harmonic transform (SHT), effectively addressing their respective limitations. Our method predicts the SHT coefficients at virtual positions and inversely transforms them into virtual speech signals, leveraging spatial information in the spherical harmonic domain for more accurate and effective virtual microphone estimation. Evaluations on the open MS-SNSD dataset demonstrate that the proposed method outperforms established baselines.
Hui Zhang 0031, Xueliang Zhang 0001
ICASSP2
2024 Efficient Multi-Channel Speech Enhancement with Spherical Harmonics Injection for Directional Encoding
abstract
Multi-channel speech enhancement extracts speech using multiple microphones that capture spatial cues. Effectively utilizing directional information is therefore key. Deep learning shows great potential on multi-channel speech enhancement and often takes short-time Fourier Transform (STFT) as inputs directly. To fully leverage the spatial information, we introduce a method using spherical harmonics transform (SHT) coefficients as auxiliary model inputs. These coefficients concisely represent spatial distributions. Specifically, our model has two encoders, one for the STFT and another for the SHT. By fusing both encoders in the decoder to estimate the enhanced STFT, we effectively incorporate spatial context. Evaluations on TIMIT under varying noise and reverberation show our model outperforms established benchmarks. Remarkably, this is achieved with fewer computations and parameters. By leveraging spherical harmonics to incorporate directional cues, our model efficiently improves the performance of the multi-channel speech enhancement.
Pengjie Shen, Hui Zhang 0031, Xueliang Zhang 0001
ICASSP3
2024 Innovative Directional Encoding in Speech Processing: Leveraging Spherical Harmonics Injection for Multi-Channel Speech Enhancement
Pengjie Shen, Hui Zhang 0031, Xueliang Zhang 0001
IJCAI3
2024 The image and ground truth dataset of Mongolian movable-type newspapers for text recognition
Feilong Bao, Hui Zhang 0031, Guanglai Gao
Int. J. Document Anal. Recognit.3
2022 Alignment-Learning Based Single-Step Decoding for Accurate and Fast Non-Autoregressive Speech Recognition
abstract
Non-autoregressive transformer (NAT) based speech recognition models have gained more and more attention since they perform faster inference speed compared with autoregressive counterparts, especially when the single-step decoding is applied. However, the single-step decoding process with length prediction will suffer from the decoding stability problem and limited improvement for inference speed. To address this, in this paper, we propose an alignment learning based NAT model, named AL-NAT. Our idea is inspired by the fact that the encoder CTC output and the target sequence are monotonically related. Specifically, we design an alignment cost matrix between the CTC output tokens and the target tokens and define a novel alignment loss to minimize the distance between the alignment cost matrix and the ground truth monotonic alignment path. By eliminating the length prediction mechanism, our AL-NAT model achieves remarkable improvements in recognition accuracy and decoding speed. To learn the contextual knowledge to improve the decoding accuracy, we further add lightweight language model on both the encoder and decoder side. Our proposed method achieves WERs of 2.8%/6.3% and RTF of 0.011 on Librispeech test clean/other sets with a lightweight 3-gram LM, and a CER of 5.3% and RTF of 0.005 on Aishell1 without LM, respectively.
Yonghe Wang, Rui Liu 0008, Feilong Bao, Hui Zhang 0031, Guanglai Gao
ICASSP4
2022 Alleviating the Loss-Metric Mismatch in Supervised Single-Channel Speech Enhancement
abstract
In this paper, we study the loss-metric mismatch problem of supervised single-channel speech enhancement system. Most of the existing speech enhancement systems achieve unsatisfying performance since their empirically selected loss functions have semantic gaps with the non-differentiable evaluation metrics, a.k.a., the loss-metric mismatch problem. In this work, we propose a simple yet efficient method to generate suitable loss functions for the real front-end speech enhancement scenarios to alleviate the loss-metric mismatch problem. Specifically, we adopt the function smoothing technique and approximate the non-differentiable evaluation metrics by a set of basis functions and their linear combination. Experimental results demonstrate that the loss function generated by our method helps the speech enhancement system achieve remarkable performance in most evaluation metrics than the traditional empirically selected ones.
Yang Yang 0121, Hui Zhang 0031, Xueliang Zhang 0001, Huaiwen Zhang
ICASSP2
2022 Speaker recognition-assisted robust audio deepfake detection
Shuai Nie 0001, Hui Zhang 0031, Shulin He, Kanghao Zhang, Shan Liang 0007, Xueliang Zhang 0001, Jianhua Tao 0001
INTERSPEECH3
2021 Joint Alignment Learning-Attention Based Model for Grapheme-to-Phoneme Conversion
abstract
Sequence-to-sequence attention-based models for grapheme-to-phoneme (G2P) conversion have gained significant interests. The attention-based encoder-decoder framework learns the mapping of input to output tokens by selectively focusing on relevant information, and has been shown well performance. However, the attention mechanism can result in non-monotonic alignments, resulting in poor G2P conversion performance. In this paper, we present a novel approach to optimize the G2P conversion model directly alignment grapheme-phoneme sequence by using alignment learning (AL) as the loss function. Besides, we propose a multi-task learning method that uses a joint alignment learning model and attention model to predict the proper alignments and thus improve the accuracy of G2P conversion. Evaluations on Mongolian and CMUDict tasks show that alignment learning as the loss function can effectively train G2P conversion model. Further, our multi-task method can significantly outperform both the alignment learning-based model and attention-based model.
Yonghe Wang, Feilong Bao, Hui Zhang 0031, Guanglai Gao
ICASSP3
2021 Soft-BAC: Soft Bidirectional Alignment Cost for End-to-End Automatic Speech Recognition
Yonghe Wang, Hui Zhang 0031, Feilong Bao, Guanglai Gao
PRICAI (2)2
2020 An Efficient Joint Training Framework for Robust Small-Footprint Keyword Spotting
Zhihao Du, Hui Zhang 0031, Xueliang Zhang 0001
ICONIP (1)3
2020 Multi-Task Learning Based Traditional Mongolian Words Recognition
abstract
In this paper, a multi-task learning framework has been proposed for solving and improving traditional Mongolian words recognition. To be specific, a sequence-to-sequence model with attention mechanism was utilized to accomplish the task of recognition. Therein, the attention mechanism is designed to fulfill the task of glyph segmentation during the process of recognition. Although the glyph segmentation is an implicit operation, the information of glyph segmentation can be integrated into the process of recognition. After that, the two tasks can be accomplished simultaneously under the framework of multi-task learning. By this way, adjacent image frames can be decoded into a glyph more precisely, which results in improving not only the performance of words recognition but also the accuracy of character segmentation. Experimental results demonstrate that the proposed multi-task learning based scheme outperforms the conventional glyph segmentation-based method and various segmentation-free (i.e. holistic recognition) methods.
Hongxi Wei, Hui Zhang 0031
ICPR2
2020 Deep Features Representation of Word Image for Keyword Spotting in Historical Mongolian Document Images
abstract
Due to degradation of historical Mongolian documents, a task for retrieving them is challenging. In the field of document image retrieval, keyword spotting technology is an alternative when optical character recognition is infeasible. Representation of word images plays a very important role in keyword spotting. In this paper, various of convolutional neural networks have been used for representing word images of historical Mongolian documents. To be specific, activations of the fully-connected layer in convolutional neural network are extracted and taken as representation vectors of word images. And then, similarity can be calculated between their representation vectors of word images. Several classic structures of convolutional neural networks have been compared with each other and the best one has been determined. Furthermore, convolutional neural network has been also compared with several baselines and the state-of-the-art method on a dataset of historical Mongolian documents. Experimental results indicates that the performance of convolutional neural network is superior to these baseline and state-of-the-art methods.
Hongxi Wei, Hui Zhang 0031
ICTAI3
2020 Polishing the Classical Likelihood Ratio Test by Supervised Learning for Voice Activity Detection
Tianjiao Xu, Hui Zhang 0031, Xueliang Zhang 0001
INTERSPEECH2
2019 Supervised Speech Enhancement with Real Spectrum Approximation
abstract
Speech enhancement aims to separate a target speech from background noise. Recently, speech enhancement has been formulated as a supervised learning problem, in which a learning machine is trained to estimate the target spectrum denoted as mapping-based method or a time-frequency mask denoted as masking-based method. Signal approximation methods indirectly estimate the target spectrum via the mask estimation, which combines the advantages of both mapping based and masking based methods. Moreover, conventional methods usually ignore the phase which is also important to the speech quality. To consider the phase, the complex number spectrum needs to be modeled. However, modeling may be difficult. In this work, a pure real number spectrum is used as an alternative representation of the complex number spectrum, and a signal approximation method is used for speech enhancement. Experimental results show that the proposed method outperforms other commonly used methods.
Hui Zhang 0031, Xueliang Zhang 0001, Linju Yang
ICASSP2
2019 Woodblock-Printing Mongolian Words Recognition by Bi-LSTM with Attention Mechanism
abstract
Woodblock-printing Mongolian documents are seriously degraded due to aging. Therefore, it is difficult to segment woodblock-printing Mongolian words are into individual glyphs. In this paper, a holistic recognition approach based on sequence to sequence model has been proposed for the woodblock-printing Mongolian words. The input of the proposed model is the sequence of frames of a wood-block printing Mongolian word. In order to generating the corresponding sequence of frames, each word image should be normalized into the same sizes in advance. And then, each word image is segmented into several fragments with equal size along writing direction. The output of the proposed model is a sequence of letters. To be specific, the proposed model contains three parts: an encoder, a decoder and an attention network. The encoder consists of a deep neural network and a bi-directional Long Short-Term Memory (Bi-LSTM). The decoder consists of a Long Short-Term Memory (LSTM) with a softmax layer. The encoder and decoder are connected by an attention network, which can map multiple frames to one letter. Experimental results demonstrate that the proposed approach outperforms the segmentation based method.
Yanke Kang, Hongxi Wei, Hui Zhang 0031, Guanglai Gao
ICDAR3
2019 UNetGAN: A Robust Speech Enhancement Approach in Time Domain for Extremely Low Signal-to-Noise Ratio Condition
abstract
Speech enhancement at extremely low signal-to-noise ratio (SNR) condition is a very challenging problem and rarely investigated in previous works. This paper proposes a robust speech enhancement approach (UNetGAN) based on U-Net and generative adversarial learning to deal with this problem. This approach consists of a generator network and a discriminator network, which operate directly in the time domain. The generator network adopts a U-Net like structure and employs dilated convolution in the bottleneck of it. We evaluate the performance of the UNetGAN at low SNR conditions (up to -20dB) on the public benchmark. The result demonstrates that it significantly improves the speech quality and substantially outperforms the representative deep learning models, including SEGAN, cGAN fo SE, Bidirectional LSTM using phase-sensitive spectrum approximation cost function (PSA-BLSTM) and Wave-U-Net regarding Short-Time Objective Intelligibility (STOI) and Perceptual evaluation of speech quality (PESQ).
Xiangdong Su, Hui Zhang 0031, Batushiren
INTERSPEECH4
2019 Investigation of Cost Function for Supervised Monaural Speech Separation
Hui Zhang 0031, Xueliang Zhang 0001, Yuhang Cao
INTERSPEECH2
2019 An Automatic Spelling Correction Method for Classical Mongolian
Feilong Bao, Guanglai Gao, Weihua Wang 0006, Hui Zhang 0031
KSEM (2)5
2019 End-to-End Model for Offline Handwritten Mongolian Word Recognition
Hongxi Wei, Hui Zhang 0031, Feilong Bao, Guanglai Gao
NLPCC (2)3
2018 A LSTM Approach with Sub-Word Embeddings for Mongolian Phrase Break Prediction
abstract
In this paper, we first utilize the word embedding that focuses on sub-word units to the Mongolian Phrase Break (PB) prediction task by using Long-Short-Term-Memory (LSTM) model. Mongolian is an agglutinative language. Each root can be followed by several suffixes to form probably millions of words, but the existing Mongolian corpus is not enough to build a robust entire word embedding, thus it suffers a serious data sparse problem and brings a great difficulty for Mongolian PB prediction. To solve this problem, we look at sub-word units in Mongolian word, and encode their information to a meaningful representation, then fed it to LSTM to decode the best corresponding PB label. Experimental results show that the proposed model significantly outperforms traditional CRF model using manually features and obtains 7.49% F-Measure gain.
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang
COLING4
2018 Training Supervised Speech Separation System to Improve STOI and PESQ Directly
abstract
Supervised speech separation methods train learning machine to cast the noisy speech to the target clean speech. Most of them use mean-square error (MSE) as loss function. However, MSE is not the perfect choice because it doesn't match the human auditory perception. Short-time objective intelligibility (STOI) and perceptual evaluation of speech quality (PESQ) are closely related to the human auditory perception and widely used in speech separation research as evaluation criteria. Therefore, STOI and PESQ may be better choices for the loss function. However, they are nondifferentiable functions which cannot be optimized by the conventional gradient descent algorithm. In this work, a gradient approximation method is used to calculate the gradients of the STOI and PESQ. Then the calculated gradients are used in the gradient descent algorithm to optimize the STOI and PESQ directly. Experimental results show the speech separation performance can be improved by the proposed method.
Hui Zhang 0031, Xueliang Zhang 0001, Guanglai Gao
ICASSP1
2018 Word Image Representation Based on Visual Embeddings and Spatial Constraints for Keyword Spotting on Historical Documents
abstract
This paper proposed a visual embeddings approach to capturing semantic relatedness between visual words. To be specific, visual words are extracted and collected from a word image collection under the Bag-of-Visual-Words framework. And then, a deep learning procedure is used for mapping visual words into embedding vectors in a semantic space. To integrate spatial constraints into the representation of word images, one word image is segmented into several sub-regions with equal size along rows and columns. After that, each sub-region can be represented as an average of embedding vectors, which is the centroid of the embedding vectors of all visual words within the same sub-region. By this way, one word image can be converted into a fixed-length vector by concatenating the corresponding average embedding vectors from its all sub-regions. Euclidean distance can be calculated to measure similarity between word images. Experimental results demonstrate that the proposed representation approach outperforms Bag-of-Visual-Words, visual language model, spatial pyramid matching, latent Dirichlet allocation, average visual word embeddings and recurrent neural network.
Hongxi Wei, Hui Zhang 0031, Guanglai Gao
ICPR2
2018 Improving Mongolian Phrase Break Prediction by Using Syllable and Morphological Embeddings with BiLSTM Model
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang
INTERSPEECH4
2018 Using Shifted Real Spectrum Mask as Training Target for Supervised Speech Separation
Hui Zhang 0031, Xueliang Zhang 0001
INTERSPEECH2
2018 Phonologically Aware BiLSTM Model for Mongolian Phrase Break Prediction with Attention Mechanism
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang
PRICAI (1)4
2017 Segmentation-Free Printed Traditional Mongolian OCR Using Sequence to Sequence with Attention Model
abstract
Mongolian Optical Character Recognition (OCR) systems are required for printed document digitization and Mongolian cultural resources utilization. Existing Mongolian OCR systems are based on segmentation. But, the Mongolian segmentation is more difficult than other languages. So, these methods are highly costly and error suffering. In this study, a segmentation-free based traditional Mongolian word recognition method is proposed. Specifically, we formalize the OCR task as a sequence to sequence mapping problem, in which the input Mongolian word image and the output textual string are treated as a sequence of image frames and a sequence of letters, respectively. A sequence to sequence with attention model is adopted to solve this problem. Experimental results on a dataset show the effectiveness of the proposed method.
Hui Zhang 0031, Hongxi Wei, Feilong Bao, Guanglai Gao
ICDAR1
2017 Representing word image using visual word embeddings and RNN for keyword spotting on historical document images
abstract
Visual words of Bag-of-Visual-Words (BoVW) framework are independent each other, which results in not only discarding spatial orders between visual words but also lacking semantic information. This study is inspired by word embeddings that a similar embedding procedure is applied to a large number of visual words. By this way, the corresponding embedding vectors of the visual words can be formulated. For a word image, the average of embedding vectors of all visual words within the word image is taken as its embedding vector. Moreover, Recurrent Neural Network (RNN) is utilized to encode each word image into embeddings like an auto-encoder. The RNN embeddings and the visual word embeddings are complementary. In this study, all word images are represented by combining visual word embeddings and RNN embeddings. Experimental results show that the proposed representation approach is superior to the traditional BoVW, spatial pyramid matching and latent Dirichlet allocation.
Hongxi Wei, Hui Zhang 0031, Guanglai Gao
ICME2
2017 Using Word Mover's Distance with Spatial Constraints for Measuring Similarity Between Mongolian Word Images
Hongxi Wei, Hui Zhang 0031, Guanglai Gao, Xiangdong Su
ICONIP (4)2
2017 Multi-Target Ensemble Learning for Monaural Speech Separation
Hui Zhang 0031, Xueliang Zhang 0001, Guanglai Gao
INTERSPEECH1
2016 Convolutional neural network for robust pitch determination
abstract
Pitch is an important characteristic of speech and is useful for many applications. However, pitch determination in noisy conditions is difficult. In this paper, we propose a supervised learning algorithm to estimate pitch using a convolutional neural network (CNN). Specifically, we use a CNN for pitch candidate selection, and dynamic programming for pitch tracking. Our experimental results show that the proposed method can obtain accurate pitch estimation and they show good generalization ability to new speakers and noisy conditions. We credit the success to the use of CNN, which is suitable for modeling the shift-invariant spectral feature for pitch detection.
Hong Su, Hui Zhang 0031, Xueliang Zhang 0001, Guanglai Gao
ICASSP2
2016 Jointly Optimizing Activation Coefficients of Convolutive NMF Using DNN for Speech Separation
Hao Li 0046, Shuai Nie 0001, Xueliang Zhang 0001, Hui Zhang 0031
INTERSPEECH4
2016 A Pairwise Algorithm Using the Deep Stacking Network for Speech Separation and Pitch Estimation
abstract
Speech separation and pitch estimation in noisy conditions are considered to be a “chicken-and-egg” problem. On one hand, pitch information is an important cue for speech separation. On the other hand, speech separation makes pitch estimation easier when background noise is removed. In this paper, we propose a supervised learning architecture to solve these two problems iteratively. The proposed algorithm is based on the deep stacking network (DSN), which provides a method for stacking simple processing modules to build deep architectures. Each module is a classifier whose target is the ideal binary mask (IBM), and the input vector includes spectral features, pitch-based features and the output from the previous module. During the testing stage, we estimate the pitch using the separation results and update the pitch-based features to the next module. When embedded into the DSN, pitch estimation and speech separation each run several times. We obtain the final results from the last module. Systematic evaluations show that the proposed system results in both a high quality estimated binary mask and accurate pitch estimation and outperforms recent systems in its generalization ability.
Xueliang Zhang 0001, Hui Zhang 0031, Shuai Nie 0001, Guanglai Gao
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 A pairwise algorithm for pitch estimation and speech separation using deep stacking network
abstract
Pitch information is an important cue for speech separation. However, pitch estimation in noisy condition is also a task as challenging as speech separation. In this paper, we propose a supervised learning architecture which combines these two problems concisely. The proposed algorithm is based on deep stacking network (DSN) which provides a method of stacking simple processing modules in building deep architecture. In the training stage, an ideal binary mask is used as target. The input vector includes the outputs of lower module and frame-level features which consist of spectral and pitch-based features. In the testing stage, each module provides an estimated binary mask which is employed to re-estimate pitch. Then we update the pitch-based features to the next module. This procedure is embedded iteratively in DSN, and we obtain the final separation results from the last module of DSN. Systematic evaluations show that the proposed approach produces high quality estimated binary mask and outperforms recent systems in generalization.
Hui Zhang 0031, Xueliang Zhang 0001, Shuai Nie 0001, Guanglai Gao
ICASSP1
2014 Deep stacking networks with time series for speech separation
abstract
In many present speech separation approaches, the separation task is formulated as a binary classification problem. Several classification-based approaches have been proposed and performed satisfactorily. However, they do not explicitly model the correlation in time and each time-frequency (T-F) unit is still classified individually. As we know, the speech signal has a very rich time series and temporal dynamic information that can be exploited for speech separation. In this study, we incorporate the correlation in time into classification. Compared with the previous approaches, the proposed approach achieves better separation and generalization performance by using deep stacking networks (DSN) with time series and re-threshold method.
Shuai Nie 0001, Hui Zhang 0031, Xueliang Zhang 0001
ICASSP2