EDBT 2026 Demo / reviewers in the wild / expert
Haitong Zhang
dblp:33/3390
· DBLP profile ↗
17ranked-venue papers
5as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation LearningabstractRecent works of music representation learning mainly focus on learning acoustic music representations with unlabeled audios or further attempt to acquire multi-modal music representations with scarce annotated audio-text pairs. They either ignore the language semantics or rely on labeled audio datasets that are difficult and expensive to create. Moreover, merely modeling semantic space usually fails to achieve satisfactory performance on music recommendation tasks since the user preference space is ignored. In this paper, we propose a novel Hierarchical Two-stage Contrastive Learning (HTCL) method that models similarity from the semantic perspective to the user perspective hierarchically to learn a comprehensive music representation bridging the gap between semantic and user preference spaces. We devise a scalable audio encoder and leverage a pre-trained BERT model as the text encoder to learn audio-text semantics via large-scale contrastive pre-training. Further, we explore a simple yet effective way to exploit interaction data from our online music platform to adapt the semantic space to user preference space via contrastive fine-tuning, which differs from previous works that follow the idea of collaborative filtering. As a result, we obtain a powerful audio encoder that not only distills language semantics from the text encoder but also models similarity in user preference space with the integrity of semantic space preserved. Experimental results on both music semantic and recommendation tasks confirm the effectiveness of our method. Xiaofeng Pan, Jing Chen 0073, Haitong Zhang, Menglin Xing, Jiayi Wei, Xuefeng Mu, Zhongqian Xie |
ICMR | 3 |
| 2024 | Transformer Network Based Channel Prediction for CSI Feedback Enhancement in AI-Native Air InterfaceabstractWith the development of artificial intelligence (AI), wireless channel prediction based on deep learning (DL) has become a hot research issue. Channel prediction plays an important role in channel state information (CSI) feedback enhancement in AI-native air interface. To better predict the CSI, this paper investigates the Transformer network based channel prediction. Firstly, real channel data are obtained in Beijing-Tianjin railway line, and the channel prediction datasets are constructed through preprocessing. After formulating the channel prediction problem, a channel prediction model based on the Transformer network is newly proposed. The unique multi-head attention mechanism and position encoding of the Transformer network enable the proposed model to have more powerful parallel computation capability and better global information capture capability. Then, the hyper-parameters of the model are determined by autocorrelation analysis and cross-validation. Finally, the performance of the proposed model is evaluated in terms of prediction accuracy and space and time computational complexity using several evaluation metrics, and is compared with classical DL models. It is shown that the proposed model possesses higher prediction performance in the appropriate range of computational complexity. Tao Zhou 0004, Xiangping Liu, Zuowei Xiang, Haitong Zhang, Bo Ai 0001, Liu Liu 0001, Xiaorong Jing |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | NSV-TTS: Non-Speech Vocalization Modeling And Transfer In Emotional Text-To-SpeechabstractThis paper addresses the problem of non-speech vocalization (NSV) modeling and transfer in emotional TTS. We propose an emotion TTS system (NSV-TTS) to model NSV and emotional speech. The model utilizes self-supervised learning to extract unsupervised linguistic units (ULUs) for NSV labeling and zero-shot NSV transfer. Furthermore, we propose token mixing and random masking to boost the performance. We evaluate the proposed method on various NSV types and emotion classes. The experimental results reveal that the proposed method performs well in the zero-shot NSV transfer task. Lastly, we conduct ablation studies to investigate the proposed method further. Haitong Zhang, Xinyuan Yu |
ICASSP | 1 |
| 2022 | DGC-Vector: A New Speaker Embedding for Zero-Shot Voice ConversionabstractRecently, more and more zero-shot voice conversion algorithms have been proposed. As a fundamental part of zero-shot voice conversion, speaker embeddings are the key to improving the converted speech’s speaker similarity. In this paper, we study the impact of speaker embeddings on zero-shot voice conversion performance. To better represent the characteristics of the target speaker and improve the speaker similarity in zero-shot voice conversion, we propose a novel speaker representation method in this paper. Our method combines the advantages of D-vector, global style token (GST) based speaker representation and auxiliary supervision. Objective and subjective evaluations show that the proposed method achieves a decent performance on zero-shot voice conversion and significantly improves speaker similarity over D-vector and GST-based speaker embedding. Ruitong Xiao, Haitong Zhang, Yue Lin 0002 |
ICASSP | 2 |
| 2022 | Improve Few-Shot Voice Cloning Using Multi-Modal LearningabstractRecently, few-shot voice cloning has achieved a significant improvement. However, most models for few-shot voice cloning are single-modal, and multi-modal few-shot voice cloning has been understudied. In this paper, we propose to use multi-modal learning to improve the few-shot voice cloning performance. Inspired by the recent works on un-supervised speech representation, the proposed multi-modal system is built by extending Tacotron2 with an unsupervised speech representation module. We evaluate our proposed system in two few-shot voice cloning scenarios, namely few-shot text-to-speech (TTS) and voice conversion (VC). Experimental results demonstrate that the proposed multi-modal learning can significantly improve the few-shot voice cloning performance over their counterpart single-modal systems. Haitong Zhang, Yue Lin 0002 |
ICASSP | 1 |
| 2022 | Data Augmentation for Long-Tailed and Imbalanced Polyphone Disambiguation in MandarinabstractPolyphone disambiguation is an important module in Mandarin Chinese text-to-speech (TTS). Recently, neural-network-based (NN-based) models have achieved a great improvement on poly-phone disambiguation. However, a long-tailed and imbalanced distribution is usually observed in the training data of polyphone disambiguation, resulting in an unsatisfying performance on the low-frequent polyphone in the imbalanced pinyin set, and the least-frequent polyphonic characters and polyphones. In this paper, we proposed a simple data-augmentation method based on the pre-trained mask language model BERT to mitigate the long-tailed and imbalanced distribution problem. We incorporate a weighted sampling technique in the data augmentation method to balance the data distribution, and a useful filtering strategy to remove some noisy augmented data. Experimental results show that the proposed data-augmentation method can improve the prediction accuracy, especially for those low-frequent polyphone in the imbalanced pinyin set, and the least-frequent polyphonic characters and polyphones. Haitong Zhang, Yue Lin 0002 |
ICASSP | 2 |
| 2022 | Exploring Timbre Disentanglement in Non-Autoregressive Cross-Lingual Text-to-SpeechabstractIn this paper, we study the disentanglement of speaker and language representations in non-autoregressive cross-lingual TTS models from various aspects.We propose a phoneme length regulator that solves the length mismatch problem between IPA input sequence and monolingual alignment results.Using the phoneme length regulator, we present a FastPitch-based crosslingual model with IPA symbols as input representations.Our experiments show that language-independent input representations (e.g.IPA symbols), an increasing number of training speakers, and explicit modeling of speech variance information all encourage non-autoregressive cross-lingual TTS model to disentangle speaker and language representations.The subjective evaluation shows that our proposed model can achieve decent naturalness and speaker similarity in cross-language voice cloning. Haoyue Zhan, Xinyuan Yu, Haitong Zhang, Yue Lin 0002 |
INTERSPEECH | 3 |
| 2022 | Deep-Learning Based Scenario Identification for High-Speed Railway Propagation ChannelsabstractPropagation scenario identification is of vital significance for boosting the performance of future smart high-speed railway (HSR) communication networks. This paper investigates the HSR propagation scenario identification model, based on deep learning networks and feature fusion methods. With the assist of railway long-term evolution (LTE) networks, the channel impulse responses are collected in four typical HSR scenarios including unobstructed viaduct, obstructed viaduct, station and suburban. Four channel characteristics involving power delay profile, root mean square (RMS) delay spread, RMS angular spread and Rice K-factor form the datasets used for model training and testing. Then, a novel propagation scenario identification model is proposed by merging a weighted score based feature fusion method into the long short-term memory (LSTM) neural network. The hyperparameters such as time window length and numbers of hidden units and layers are determined by autocorrelation analysis and cross-validation. Finally, the model performance is evaluated by focusing on the impact of different feature fusion methods and computational complexity. Haitong Zhang, Tao Zhou 0004, Liu Liu 0001 |
VTC Spring | 1 |
| 2022 | Weighted Score Fusion Based LSTM Model for High-Speed Railway Propagation Scenario IdentificationabstractPropagation scenario identification is of vital significance for boosting the performance of future smart high-speed railway (HSR) communication networks. This paper investigates the HSR propagation scenario identification model, based on deep learning networks and feature fusion methods. With the assistance of railway long-term evolution (LTE) networks, we collected the channel impulse responses in four typical HSR scenarios including unobstructed viaduct, obstructed viaduct, station and suburban. Four channel characteristics involving power delay profile, root mean square (RMS) delay spread, RMS angular spread and Ricean K-factor form the datasets used for model training and testing. Then, a novel propagation scenario identification model is proposed by merging a weighted score based feature fusion method into the long short-term memory (LSTM) neural network. The hyper-parameters of the proposed model such as time window length and numbers of hidden units and layers are determined by autocorrelation analysis and cross-validation. Finally, the model performance is evaluated by focusing on the impact of feature selection, comparison of different feature fusion methods, and computational complexity. The evaluation results show that the proposed model has high identification accuracy but acceptable computational complexity. Tao Zhou 0004, Haitong Zhang, Bo Ai 0001, Liu Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Deep-Learning-Based Spatial-Temporal Channel Prediction for Smart High-Speed Railway Communication NetworksabstractIntelligent channel prediction plays a key role in artificial intelligence (AI)-optimized or AI-native communication networks for smart high-speed railways (HSRs). This paper investigates the spatial-temporal prediction of channel state information (CSI) and channel statistical characteristics (CSCs) based on deep-learning (DL) for the future smart HSR communication network. A propagation-graph simulation method is used to generate datasets of CSI and CSCs for massive multiple-input multiple-output (mMIMO) channels in a HSR cutting scenario, and realistic channel measurements are used to validate the datasets. Then, single-step ahead and multi-step ahead prediction problems are formulated with the consideration of both spatial and temporal information hidden in the datasets. By exploiting the temporal and spatial correlations of the HSR mMIMO channel, a novel spatial-temporal channel prediction model that combines the convolutional neural network (CNN) and convolutional long short-term memory (CLSTM) is proposed and called as Conv-CLSTM. Moreover, the hyper-parameters of the Conv-CLSTM model are determined by autocorrelation and similarity analysis and cross-validation. Finally, the performance of the Conv-CLSTM model is evaluated in terms of prediction accuracy and space and time computational complexity, and is compared with classical DL models. The evaluation results show that the proposed model has high prediction accuracy but acceptable computational complexity. Tao Zhou 0004, Haitong Zhang, Bo Ai 0001, Liu Liu 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2021 | Improve Cross-Lingual Text-To-Speech Synthesis on Monolingual Corpora with Pitch Contour Information
Haoyue Zhan, Haitong Zhang, Wenjie Ou, Yue Lin 0002 |
Interspeech | 2 |
| 2021 | Deep Learning Based Channel Prediction for Massive MIMO Systems in High-Speed Railway ScenariosabstractThis paper investigates the prediction model based on deep learning for the wireless channel characteristics of massive MIMO systems in high-speed railway (HSR) scenarios. Based on the propagation graph theory, we simulate the massive MIMO channel in a HSR cutting scenario. The datasets of spatial-temporal channel characteristics, involving channel state information, Ricean K-factor, delay spread, and angle spread, are generated for the model training and testing, and two kinds of prediction problem formulations, such as single-step and multi-steps, are designed. By considering both the spatial and temporal correlation properties in HSR massive MIMO channels, a novel channel prediction model that combines the convolutional long short-term memory (CLSTM) and convolutional neural network (CNN) is proposed and called as Conv-CLSTM. The hyperparameters of Conv-CLSTM are determined by comparative experiments and autocorrelation and similarity analysis. According to the performance evaluation, it is showed that the proposed Conv-CLSTM outperforms the other deep learning and machine learning models. Tao Zhou 0004, Haitong Zhang, Liu Liu 0001, Cheng Tao 0001 |
VTC Spring | 3 |
| 2021 | Physically based modeling and rendering of avalanches
Xincheng Liu, Haitong Zhang, Yuhong Zou, Zhangye Wang, Qunsheng Peng 0001 |
Vis. Comput. | 3 |
| 2020 | Unsupervised Learning for Sequence-to-Sequence Text-to-Speech for Low-Resource LanguagesabstractRecently, sequence-to-sequence models with attention have been successfully applied in Text-to-speech (TTS). These models can generate near-human speech with a large accurately-transcribed speech corpus. However, preparing such a large data-set is both expensive and laborious. To alleviate the problem of heavy data demand, we propose a novel unsupervised pre-training mechanism in this paper. Specifically, we first use Vector-quantization Variational-Autoencoder (VQ-VAE) to ex-tract the unsupervised linguistic units from large-scale, publicly found, and untranscribed speech. We then pre-train the sequence-to-sequence TTS model by using the pairs. Finally, we fine-tune the model with a small amount of paired data from the target speaker. As a result, both objective and subjective evaluations show that our proposed method can synthesize more intelligible and natural speech with the same amount of paired training data. Besides, we extend our proposed method to the hypothesized low-resource languages and verify the effectiveness of the method using objective evaluation. Haitong Zhang, Yue Lin 0002 |
INTERSPEECH | 1 |
| 2020 | Improving interpretability of word embeddings by generating definition and usage
Haitong Zhang, Yongping Du, Jiaxin Sun, Qingxiao Li |
Expert Syst. Appl. | 1 |
| 2020 | Wasserstein based transfer network for cross-domain sentiment classification
Yongping Du, Meng He 0009, Lulin Wang, Haitong Zhang |
Knowl. Based Syst. | 4 |
| 2019 | Physically based modeling and animation of landslides with MPM
Jianwang Zhao, Haitong Zhang, Zhangye Wang, Qunsheng Peng 0001 |
Vis. Comput. | 3 |