Tianrui Wang

dblp:276/7636 · DBLP profile ↗
← Back
30ranked-venue papers
11as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
abstract
Chunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang, Yuxin Guo, Teng Ma, Ziyu Zhang, Tianrui Wang, Cheng Gong, Yushen Chen, Ruibo Fu, Longbiao Wang, Jianwu Dang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chunyu Qiang, Yuzhe Liang, Tianrui Wang, Yushen Chen, Ruibo Fu, Longbiao Wang, Jianwu Dang 0001
ACL (1)8
2026 Evaluating the Expressive Appropriateness of Speech in Rich Contexts
abstract
Tianrui Wang, Ziyang Ma, Yizhou Peng, Haoyu Wang, Zhikang Niu, Zikang Huang, Yihao Wu, Yi-Wen Chao, Yu Jiang, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Cheng Gong, Yifan Yang, Tianchi Liu, Junyu Wang, Nana Hou, Meng Ge, Fuming You, Yang Wei, Zhongqian Sun, Hu Haifeng, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tianrui Wang, Ziyang Ma 0001, Yizhou Peng, Zhikang Niu, Zikang Huang, Yi-Wen Chao, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Yifan Yang 0005, Tianchi Liu 0004, Nana Hou, Meng Ge, Fuming You, Zhongqian Sun, Haifeng Hu 0009, Xiaobao Wang, Chng Eng Siong, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001
ACL (1)1
2026 Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
abstract
Yifan Yang, Bing Han, Hui Wang, Wei Wang, Ziyang Ma, Long Zhou, Zengrui Jin, Guanrou Yang, Tianrui Wang, Xu Tan, Xie Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yifan Yang 0005, Bing Han 0008, Hui Wang 0075, Wei Wang 0010, Ziyang Ma 0001, Zengrui Jin, Guanrou Yang, Tianrui Wang, Xu Tan 0003, Xie Chen 0001
ACL (1)9
2026 Improved ship-radiated noise recognition using a pre-trained noise reduction module combined with feature optimization network
Yunye Feng, Haonan Xing, Meng Ge, Tianrui Wang, Lin Gan 0003, Gaoyan Zhang
Knowl. Based Syst.5
2025 A Hybrid Algorithm for the Regular Syndrome Decoding Problem
Tianrui Wang, Anyu Wang 0001, Kang Yang 0002, Yu Yu 0001, Jun Zhang 0031, Xiaoyun Wang 0001
ASIACRYPT (4)1
2025 Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
abstract
The information loss or distortion caused by single-channel speech enhancement (SE) harms the performance of automatic speech recognition (ASR). Observation addition (OA) is an effective post-processing method to improve ASR performance by balancing noisy and enhanced speech. Determining the OA coefficient is crucial. However, the currently supervised OA coefficient module, called the bridging module, only utilizes simulated noisy speech for training, which has a severe mismatch with real noisy speech. In this paper, we propose training strategies to train the bridging module with real noisy speech. First, DNSMOS is selected to evaluate the perceptual quality of real noisy speech with no need for the corresponding clean label to train the bridging module. Additional constraints during training are introduced to enhance the robustness of the bridging module further. Each utterance is evaluated by the ASR back-end using various OA coefficients to obtain the word error rates (WERs). The WERs are used to construct a multidimensional vector. This vector is introduced into the bridging module with multi-task learning and is used to determine the optimal OA coefficients. The experimental results on the CHiME-4 dataset show that the proposed methods all had significant improvement compared with the simulated data trained bridging module, especially under real evaluation sets.
Zhongjian Cui, Chenrui Cui, Tianrui Wang, Mengnan He, Meng Ge, Caixia Gong, Longbiao Wang, Jianwu Dang 0001
ICASSP3
2025 Discrete Unit-based Low-latency Multi-lingual Speech Synthesis for LIMMITS'25 Challenge
abstract
In this paper, we present the system developed by our team, CCATTS, for the LIMMITS’25 challenge, focusing on few-shot and zero-shot TTS. We adopt a two-stage TTS strategy. In track 1, we fine-tune the pre-trained ZMM-TTS model and successfully achieve multilingual low-latency TTS. In track 2, we propose a token-based framework by modifying the first-stage model of ZMM-TTS, disentangling speech features into four types of discrete tokens—Content, Acoustic, Emotion, and Speaker—and integrating it with our designed token2wav module. This module consists of a HiFi-GAN-style decoder, an acoustic refiner, and U-Net flow matching, to generate high-quality speech. The official competition results demonstrate that our method achieves strong performance in both tracks.
Tianrui Wang, Chunyu Qiang, Qiuyu Liu, Yuheng Lu, Xiaobao Wang, Longbiao Wang, Jianwu Dang 0001
ICASSP3
2025 A Chinese Expressive Long-dialogue Speech Dataset with Scripts
abstract
With the advancement of large-scale models, the demand for emotionally rich, long-context, and highly natural communication in human-computer interaction increases. However, the exploration of long-context or script-level speech conversation tasks remains limited due to the lack of specific supervised data. To address this, we introduce a three-stage data processing pipeline for creating a Chinese expressive long-dialogue speech dataset with scripts (CELSDS). We collect videos from TV series, manually annotate speaker information for each character, apply Optical Character Recognition (OCR) to extract speech content, annotate episode summaries, and use a large language model (LLM) to generate sentence-level scenario descriptions. To our knowledge, this is the first Chinese long-context dialogue dataset that incorporates speaker and content annotations, script-level episode summaries, and sentence-level scenario details. Using this dataset, we develop a baseline model for both speech-to-script and script-to-speech generation tasks. The annotations and data production code are open-sourced at: https://github.com/lijin0120/CELSDS.
Tianrui Wang, Meng Ge, Chenrui Cui, Jianrong Wang, Longbiao Wang, Jianwu Dang 0001
ICASSP2
2025 Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
abstract
In recent speech enhancement (SE) research, transformer and its variants have emerged as the predominant methodologies. However, the quadratic complexity of the self-attention mechanism imposes certain limitations on practical deployment. Mamba, as a novel state-space model (SSM), has gained widespread application in natural language processing and computer vision due to its strong capabilities in modeling long sequences and relatively low computational complexity. In this work, we introduce Mamba-SEUNet, an innovative architecture that integrates Mamba with U-Net for SE tasks. By leveraging bidirectional Mamba to model forward and backward dependencies of speech signals at different resolutions, and incorporating skip connections to capture multi-scale information, our approach achieves state-of-the-art (SOTA) performance. Experimental results on the VCTK+DEMAND dataset indicate that Mamba-SEUNet attains a PESQ score of 3.59, while maintaining low computational complexity. When combined with the Perceptual Contrast Stretching technique, Mamba-SEUNet further improves the PESQ score to 3.73.
Zizhen Lin, Tianrui Wang, Meng Ge, Longbiao Wang, Jianwu Dang 0001
ICASSP3
2025 Time-Graph Frequency Representation with Singular Value Decomposition for Neural Speech Enhancement
abstract
Time-frequency (T-F) domain methods for monaural speech enhancement have benefited from the success of deep learning. Recently, focus has been put on designing two-stream network models to predict amplitude mask and phase separately, or, coupling the amplitude and phase into Cartesian coordinates and constructing real and imaginary pairs. However, most methods suffer from the alignment modeling of amplitude and phase (real and imaginary pairs) in a two-stream network framework, which inevitably incurs performance restrictions. In this paper, we introduce a graph Fourier transform defined with the singular value decomposition (GFT-SVD), resulting in real-valued time-graph representation for neural speech enhancement. This real-valued representation-based GFT-SVD provides an ability to align the modeling of amplitude and phase, leading to avoiding recovering the target speech phase information. Our findings demonstrate the effects of real-valued time-graph representation based on GFT-SVD for neutral speech enhancement. The extensive speech enhancement experiments establish that the combination of GFT-SVD and DNN outperforms the combination of GFT with the eigenvector decomposition (GFT-EVD) and magnitude estimation UNet, and outperforms the short-time Fourier transform (STFT) and DNN, regarding objective intelligibility and perceptual quality. We release our source code at: https://github.com/Wangfighting0015/GFTproject.
Tianrui Wang, Meng Ge, Qiquan Zhang, Zirui Ge, Zhen Yang 0001
ICASSP2
2025 Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
abstract
Code-switching (CS) automatic speech recognition (ASR) faces challenges due to the language confusion resulting from accents, auditory similarity, and seamless language switches. Adaptation on the pre-trained multi-lingual model has shown promising performance for CS-ASR. In this paper, we adapt Whisper, which is a large-scale multilingual pre-trained speech recognition model, to CS from both encoder and decoder parts. First, we propose an encoder refiner to enhance the encoder’s capacity of intra-sentence swithching. Second, we propose using two sets of language-aware adapters with different language prompt embeddings to achieve language-specific decoding information in each decoder layer. Then, a fusion module is added to fuse the language-aware decoding. The experimental results using the SEAME dataset show that, compared with the baseline model, the proposed approach achieves a relative MER reduction of 4.1% and 7.2% on the dev_man and dev_sge test sets, respectively, surpassing state-of-the-art methods. Through experiments, we found that the proposed method significantly improves the performance on non-native language in CS speech, indicating that our approach enables Whisper to better distinguish between the two languages.
Chenrui Cui, Tianrui Wang, Hexin Liu, Zhaoheng Ni, Lingxuan Ye, Longbiao Wang
ICASSP4
2025 A Progressive Generation Framework with Speech Pre-trained Model for Expressive Voice Conversion
abstract
Expressive voice conversion (EVC) aims to modify the speaker identity and emotional style of speech while preserving its content. Existing approaches often focus on disentangling speaker, emotion, and content information but overlook the progressive generation mechanisms in human speech production. To address this, we propose a three-stage framework that includes a speech disentanglement module, a progressive generator, and an acoustic refiner. This framework enables speech pre-trained models to parse linguistic content, emotional style, and speaker identity, which are then progressively integrated into the speech reconstruction branch to generate high-quality speech with replaceable emotional style and speaker identity. Experiments with six different pre-trained models show that our framework activates their disentanglement capabilities, surpassing baseline performance in EVC, and supports speaker and emotion control from different target samples. This framework also provides a valuable reference for evaluating the disentanglement capabilities of speech pre-training models.
Tianrui Wang, Meng Ge, Zhikang Niu, Chunyu Qiang, Zikang Huang, Ziyang Ma 0001, Xiaobao Wang, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001
ICME1
2025 A Three-Stage Beamforming with Harmonic Guidance for Multi-Channel Speech Enhancement
Nurali Alip, Tianrui Wang, Meng Ge, Jingru Lin, Longbiao Wang, Jianwu Dang 0001
INTERSPEECH2
2025 Augment Mandarin to Cantonese Speech Databases via Retrieval-Augmented Generation and Speech Synthesis
Boyu Zhu, Ruihao Jing, Chunyu Qiang, Tianrui Wang
INTERSPEECH6
2025 ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
Tianrui Wang, Meng Ge, Longbiao Wang, Jianwu Dang 0001
INTERSPEECH2
2025 EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
abstract
Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the emotional expression in the generated speech. In this work, we propose EmoVoice, a novel emotion-controllable TTS model that exploits large language models (LLMs) to enable fine-grained freestyle natural language emotion control, and a phoneme boost variant design that makes the model output phoneme tokens and audio tokens in parallel to enhance content consistency, inspired by chain-of-thought (CoT) and modality-of-thought (CoM) techniques. Besides, we introduce EmoVoice-DB, a high-quality 40-hour English emotion dataset featuring expressive speech and fine-grained emotion labels with natural language descriptions. EmoVoice achieves state-of-the-art performance on the English EmoVoice-DB test set using only synthetic training data, and on the Chinese Secap test set using our in-house data. We further investigate the reliability of existing emotion evaluation metrics and their alignment with human perceptual preferences, and explore using SOTA multimodal LLMs GPT-4o-audio and Gemini to assess emotional speech. Dataset, code, checkpoints and demo samples are available at https://github.com/yanghaha0908/EmoVoice.
Guanrou Yang, Qian Chen 0003, Ziyang Ma 0001, Wen Wang 0019, Tianrui Wang, Yifan Yang 0005, Zhikang Niu, Wenrui Liu 0003, Fan Yu 0002, Zhihao Du, Zhifu Gao, Shiliang Zhang, Xie Chen 0001
ACM Multimedia7
2025 MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
abstract
We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. These findings underscore the urgent need for greater research attention in audio-language reasoning, including both data and algorithm innovation. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area.
Ziyang Ma 0001, Yinghao Ma, Yanqiao Zhu 0003, Yi-Wen Chao, Yuanzhe Chen, Zhuo Chen 0006, Jian Cong, Keliang Li, Siyou Li, Xinfeng Li, Xiquan Li, Zheng Lian 0004, Yuzhe Liang, Minghao Liu 0003, Zhikang Niu, Tianrui Wang, Yuping Wang 0005, Yuxuan Wang 0002, Guanrou Yang, Jianwei Yu 0001, Ruibin Yuan, Zhisheng Zheng, Ziya Zhou, Haina Zhu, Wei Xue 0002, Emmanouil Benetos, Kai Yu 0004, Chng Eng Siong, Xie Chen 0001
NeurIPS20
2025 Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
abstract
While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modeling multi-emotion transitions and the scarcity of annotated datasets that capture intra-sentence emotional and prosodic variation. In this paper, we propose WeSCon, the first self-training framework that enables word-level control of both emotion and speaking rate in a pretrained zero-shot TTS model, without relying on datasets containing intra-sentence emotion or speed transitions. Our method introduces a transition-smoothing strategy and a dynamic speed control mechanism to guide the pretrained TTS model in performing word-level expressive synthesis through a multi-round inference process. To further simplify the inference, we incorporate a dynamic emotional attention bias mechanism and fine-tune the model via self-training, thereby activating its ability for word-level expressive control in an end-to-end manner. Experimental results show that WeSCon effectively overcomes data scarcity, achieving state-of-the-art performance in word-level emotional expression control while preserving the strong zero-shot synthesis capabilities of the original TTS model.
Tianrui Wang, Meng Ge, Chunyu Qiang, Ziyang Ma 0001, Zikang Huang, Guanrou Yang, Xiaobao Wang, Chng Eng Siong, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001
NeurIPS1
2025 LORT: Locally refined convolution and Taylor transformer for monaural speech enhancement
Zizhen Lin, Tianrui Wang, Meng Ge, Longbiao Wang, Jianwu Dang 0001
Speech Commun.3
2025 Emotional Style Transfer With Intensity Control in Zero-Shot TTS
abstract
Recent advancements in zero-shot text-to-speech have enabled the generated speech to preserve both the speaker's identity and emotion, guided by a reference speech. However, generating speech in emotional styles that the target speaker has never exhibited remains a central challenge in style transfer. Existing methods attempt to achieve this through disentangled modeling of speaker and style representations, yet insufficient disentanglement often leads to weak expressiveness and speaker leakage. To address these limitations, we propose a robust LM-based TTS framework that enables fine-grained control over speaker identity and emotional style. Specifically, we introduce a frame-level Style Transfer Pitch Predictor to capture fine-grained prosodic and speaker-related information, and a Speaker Emotional Style Alignment module that strengthens emotional expressiveness while preserving speaker identity. Our method enables precise control over the intensity of emotional style in the synthesized speech. Experimental results demonstrate that our approach outperforms existing methods, delivering more expressive and controllable emotional synthesized speech in the emotional style transfer task. Audio samples can be available athttps://whyrrrrun.github.io/ICST.github.io/.
Chunyu Qiang, Tianrui Wang, Longbiao Wang
IEEE Signal Process. Lett.3
2024 VoiCor: A Residual Iterative Voice Correction Framework for Monaural Speech Enhancement
Tianrui Wang, Meng Ge, Andong Li, Longbiao Wang, Jianwu Dang 0001, Yungang Jia
INTERSPEECH2
2024 VioLA: Conditional Language Models for Speech Recognition, Synthesis, and Translation
abstract
Recent research shows a big convergence in model architecture, training objectives, and inference methods across various tasks for different modalities. In this paper, we proposeVioLA, a single auto-regressive Transformer decoder-only network that unifies various cross-modal tasks involving speech and text, such as speech-to-text, text-to-text, text-to-speech, and speech-to-speech tasks, as a conditional language model task via multi-task learning framework. To accomplish this, we first convert the speech utterances to discrete tokens (similar to the textual data) using an offline neural codec encoder. In such a way, all these tasks are converted to token-based sequence prediction problems, which can be naturally handled with one conditional language model. We further integrate task IDs (TID), language IDs (LID), and LSTM-based acoustic embedding into the proposed model to enhance the modeling capability of handling different languages and tasks. Experimental results demonstrate that the proposedVioLAmodel can support both single-modal and cross-modal tasks well, and the decoder-only model achieves a comparable and even better performance than the strong baselines.
Tianrui Wang, Yu Wu 0012, Shujie Liu 0001, Yashesh Gaur, Zhuo Chen 0006, Jinyu Li 0001, Furu Wei
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 On Decoder-Only Architecture For Speech-to-Text and Large Language Model Integration
abstract
Large language models (LLMs) have achieved remarkable success in the field of natural language processing, enabling better human-computer interaction using natural language. However, the seamless integration of speech signals into LLMs has not been explored well. The “decoder-only“ architecture has also not been well studied for speech processing tasks. In this research, we introduce Speech-LLaMA, a novel approach that effectively incorporates acoustic information into text-based large language models. Our method leverages Connectionist Temporal Classification and a simple audio encoder to map the compressed acoustic features to the continuous semantic space of the LLM. In addition, we further probe the decoder-only architecture for speech-to-text tasks by training a smaller scale randomly initialized speech-LLaMA model from speech-text paired data alone. We conduct experiments on multilingual speech-to-text translation tasks and demonstrate a significant improvement over strong baselines, highlighting the potential advantages of decoder-only models for speech-to-text conversion.
Jian Wu 0027, Yashesh Gaur, Zhuo Chen 0006, Yimeng Zhu, Tianrui Wang, Jinyu Li 0001, Shujie Liu 0001, Linquan Liu, Yu Wu 0012
ASRU6
2023 Exploring Decryption Failures of BIKE: New Class of Weak Keys and Key Recovery Attacks
Tianrui Wang, Anyu Wang 0001, Xiaoyun Wang 0001
CRYPTO (3)1
2023 Black-box Word-level Textual Adversarial Attack Based On Discrete Harris Hawks Optimization
abstract
Neural network-based applications are prone to being fooled by adversarial examples due to the natural vulnerability of deep neural networks (DNNs). Textual adversarial attacks are particularly challenging due to the discreteness between texts. The adversarial examples crafted by word-level textual attacks which are typically treated as optimization problems in black-box scenarios perform better in human evaluation. Existing approaches have struggled to balance the success rate with the time consuming, mainly because the chosen optimization algorithm is not efficient enough. In this paper, we propose a method to generate textual adversarial examples called Discrete Harris Hawk Optimization (DHHO). We set up three operations for handling discrete data, which are applied to each stage of the Harris Hawk Optimization (HHO) to enable it to solve optimization problems in discrete space. By attacking BiLSTM and BERT on two benchmark data sets, we conduct extensive experiments to evaluate our attack method with a success rate of up to 98% and a reduction of time is at least 50%. Moreover, the experimental results also show that our adversarial examples can ensure high quality and transferability.
Tianrui Wang, Weina Niu, Kangyi Ding, Lingfeng Yao, Pengsen Cheng, Xiaosong Zhang 0001
CSCWD1
2023 An Adapter Based Multi-Label Pre-Training for Speech Separation and Enhancement
abstract
In recent years, self-supervised learning (SSL) has achieved tremendous success in various speech tasks due to its power to extract representations from massive unlabeled data. However, compared with tasks such as speech recognition (ASR), the improvements from SSL representation in speech separation (SS) and enhancement (SE) are considerably smaller. Based on HuBERT, this work investigates improving the SSL model for SS and SE. We first update HuBERT’s masked speech prediction (MSP) objective by integrating the separation and denoising terms, resulting in a multiple pseudo label pre-training scheme, which significantly improves HuBERT’s performance on SS and SE but degrades the performance on ASR. To maintain its performance gain on ASR, we further propose an adapter-based architecture for HuBERT’s Transformer encoder, where only a few parameters of each layer are adjusted to the multiple pseudo label MSP while other parameters remain frozen as default HuBERT. Experimental results show that our proposed adapter-based multiple pseudo label HuBERT yield consistent and significant performance improvements on SE, SS, and ASR tasks, with a faster pretraining speed, at only marginal parameters increase.
Tianrui Wang, Xie Chen 0001, Zhuo Chen 0006, Shu Yu 0001, Weibin Zhu
ICASSP1
2023 Harmonic Attention for Monaural Speech Enhancement
abstract
To further improve the quality of the enhanced speech, it is appealing that more profound articulatory and auditory knowledge should be introduced into the speech enhancement model. Among these, harmonics seriously affect speech timbre and play a crucial role in speech intelligibility. Especially in the frequency domain, harmonics appear as the local maximum peaks of energy, which could be expected to serve as anchors to recover the distorted speech. In this paper, an explicit modeling method, harmonic attention, is presented, patching the harmonics with the help of residual ones. In order to maintain the spectral structure of speech during the processing and to enable the network to support harmonic modeling, a harmonic attention-based progressive enhancement network (HAPNet) is applied, which gradually approaches clean speech with stacked modules of harmonic attention. In addition, to make enhanced speech more consistent with hearing, a loss function based on the loudness power compression (LC-SNR) is used, which measures both magnitude and phase values with appropriate auditory effects. The experimental visualization indicates that the harmonic attention can capture and recover the harmonics of speech. And the objective evaluations show that the presented HAPNet and LC-SNR outperform the referenced methods. Furthermore, the presented model trained on 100 hours of data achieves competitive results with the referenced models trained on 3000+ hours of data, and one trained on 500 hours of data yields the state-of-the-art performance.
Tianrui Wang, Weibin Zhu, Yingying Gao, Shilei Zhang, Junlan Feng
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Harmonic Gated Compensation Network Plus for ICASSP 2022 DNS Challenge
abstract
The harmonic structure of speech is resistant to noise, but the harmonics may still be partially masked by noise. Therefore, we previously proposed a harmonic gated compensation network (HGCN) to predict the full harmonic locations based on the unmasked harmonics and process the result of a coarse enhancement module to recover the masked harmonics. In addition, the auditory loudness loss function is used to train the network. For the DNS Challenge, we update HGCN with the following aspects, resulting in HGCN+. First, a high-band module is employed to help the model handle full-band signals. Second, cosine is used to model the harmonic structure more accurately. Then, the dual-path encoder and dual-path rnn (DPRNN) are introduced to take full advantage of the features. Finally, a gated residual linear structure replaces the gated convolution in the compensation module to increase the receptive field of frequency. The experimental results show that each updated module brings performance improvement to the model. HGCN+ also outperforms the referenced models on both wide-band and full-band test sets.
Tianrui Wang, Weibin Zhu, Yingying Gao, Junlan Feng, Shilei Zhang
ICASSP1
2022 HGCN: Harmonic Gated Compensation Network for Speech Enhancement
abstract
Mask processing in the time-frequency (T-F) domain through the neural network has been one of the mainstreams for single-channel speech enhancement. However, it is hard for most models to handle the situation when harmonics are partially masked by noise. To tackle this challenge, we propose a harmonic gated compensation network (HGCN). We design a high-resolution harmonic integral spectrum to improve the accuracy of harmonic locations prediction. Then we add voice activity detection (VAD) and voiced region detection (VRD) to the convolutional recurrent network (CRN) to filter harmonic locations. Finally, the harmonic gating mechanism is used to guide the compensation model to adjust the coarse results from CRN to obtain the refinedly enhanced results. Our experiments show HGCN achieves substantial gain over a number of advanced approaches in the community.
Tianrui Wang, Weibin Zhu, Yingying Gao, Junlan Feng, Shilei Zhang
ICASSP1
2022 Is Image Encoding Beneficial for Deep Learning in Finance?
abstract
In 2012, Securities and Exchange Commission (SEC) mandated all corporate filings for any company doing business in the U.S. be entered into the electronic data gathering, analysis, and retrieval (EDGAR) system. In this work, we are investigating ways to analyze the data available through the EDGAR database. This may serve portfolio managers (pension funds, mutual funds, insurance, and hedge funds) to get automated insights into companies they invest in, to better manage their portfolios. The analysis is based on artificial neural networks applied to the data. In particular, one of the most popular machine learning methods, the convolutional neural network (CNN) architecture, originally developed to interpret and classify images, is now being used to interpret financial data. This work investigates the best way to input data collected from the SEC filings into a CNN architecture. We incorporate accounting principles and mathematical methods into the design of three image encoding methods. Specifically, two methods are derived from accounting principles (sequential arrangement, category chunk arrangement) and one is using a purely mathematical technique [the Hilbert vector arrangement (HVA)]. In this work, we analyze fundamental financial data as well as financial ratio data and study companies from the financial, healthcare, and information technology sectors in the United States. We find that using imaging techniques to input data for CNN works better for financial ratio data but is not significantly better than simply using the 1-D input directly for fundamental data. We do not find the HVA technique to be significantly better than other imaging techniques.
Tianrui Wang, Ionut Florescu
IEEE Internet Things J.2