Tian-Hao Zhang

dblp:275/1100 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles
abstract
Humans can perceive speakers’ characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech ( TTS ) scholars grounded their investigations on real-person faces, thereby restricting effective speech synthesis from applying to vast potential usage scenarios with diverse characters and image styles. To solve this issue, we introduce a novel FaceSpeak approach. It extracts salient identity characteristics and emotional representations from a wide variety of image styles. Meanwhile, it mitigates the extraneous information (e.g., background, clothing, and hair color, etc.), resulting in synthesized speech closely aligned with a character’s persona. Furthermore, to overcome the scarcity of multi-modal TTS data, we have devised an innovative dataset, namely Expressive Multi-Modal TTS ( EM2TTS), which is diligently curated and annotated to facilitate research in this domain. The experimental results demonstrate our proposed FaceSpeak can generate portrait-aligned voice with satisfactory naturalness and quality.
Tian-Hao Zhang, Xinyuan Qian 0001, Xu-Cheng Yin
AAAI1
2025 Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
abstract
Recently, end-to-end automatic speech recognition has become the mainstream approach in both industry and academia. To optimize system performance in specific scenarios, the Weighted Finite-State Transducer (WFST) is extensively used to integrate acoustic and language models, leveraging its capacity to implicitly fuse language models within static graphs, thereby ensuring robust recognition while also facilitating rapid error correction. However, WFST necessitates a frame-by-frame search of CTC posterior probabilities through autoregression, which significantly hampers inference speed. In this work, we thoroughly investigate the spike property of CTC outputs and further propose the conjecture that adjacent frames to non-blank spikes carry semantic information beneficial to the model. Building on this, we propose the Spike Window Decoding algorithm, which greatly improves the inference speed by making the number of frames decoded in WFST linearly related to the number of spiking frames in the CTC output, while guaranteeing the recognition performance. Our method achieves SOTA recognition accuracy with significantly accelerates decoding speed, proven across both AISHELL-1 and large-scale In-House datasets, establishing a pioneering approach for integrating CTC output with WFST.
Tian-Hao Zhang, Xinyuan Qian 0001, Xu-Cheng Yin
ICASSP2
2025 Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
abstract
The kNN-CTC model has proven to be effective for monolingual automatic speech recognition (ASR). However, its direct application to multilingual scenarios like code-switching, presents challenges. Although there is potential for performance improvement, a kNN-CTC model utilizing a single bilingual datastore can inadvertently introduce undesirable noise from the alternative language. To address this, we propose a novel kNN-CTC-based code-switching ASR (CS-ASR) framework that employs dual monolingual datastores and a gated datastore selection mechanism to reduce noise interference. Our method selects the appropriate datastore for decoding each frame, ensuring the injection of language-specific information into the ASR process. We apply this framework to cutting-edge CTC-based models, developing an advanced CS-ASR system. Extensive experiments demonstrate the remarkable effectiveness of our gated datastore mechanism in enhancing the performance of zero-shot Chinese-English CS-ASR.
Jiaming Zhou 0001, Shiwan Zhao, Hui Wang 0075, Tian-Hao Zhang, Haoqin Sun, Xuechen Wang
ICASSP4
2024 CIF-T: A Novel CIF-Based Transducer Architecture for Automatic Speech Recognition
abstract
RNN-T models are widely used in ASR, which rely on the RNN-T loss to achieve length alignment between input audio and target sequence. However, the implementation complexity and the alignment-based optimization target of RNN-T loss lead to computational redundancy and a reduced role for predictor network, respectively. In this paper, we propose a novel model named CIF-Transducer (CIF-T) which incorporates the Continuous Integrate-and-Fire (CIF) mechanism with the RNN-T model to achieve efficient alignment. In this way, the RNN-T loss is abandoned, thus bringing a computational reduction and allowing the predictor network a more significant role. We also introduce Funnel-CIF, Context Blocks, Unified Gating and Bilinear Pooling joint network, and auxiliary training strategy to further improve performance. Experiments on the 178-hour AISHELL-1 and 10000-hour WenetSpeech datasets show that CIF-T achieves state-of-the-art results with lower computational overhead compared to RNN-T models.
Tian-Hao Zhang, Dinghao Zhou, Guiping Zhong, Jiaming Zhou 0001, Baoxiang Li
ICASSP1
2024 Transmitted and Aggregated Self-Attention for Automatic Speech Recognition
Tian-Hao Zhang, Xinyuan Qian 0001, Feng Chen 0040, Xu-Cheng Yin
INTERSPEECH1
2024 M3TTS: Multi-modal text-to-speech of multi-scale style control for dubbing
Li-Fang Wei, Xinyuan Qian 0001, Tian-Hao Zhang, Song-Lu Chen, Xu-Cheng Yin
Pattern Recognit. Lett.4
2024 Improving Multi-Type License Plate Recognition via Learning Globally and Contrastively
abstract
Previous license plate recognition (LPR) methods have achieved impressive performance on single-type license plates. However, multi-type license plate recognition is still challenging due to various character layouts and fonts. There are two main problems: one is that recognition models are prone to incorrectly perceive the location of characters due to diverse character layouts, and the other is that characters of different categories may have similar glyphs due to various fonts, causing character misidentification. Therefore, to solve the above problems, we propose two plug-and-play modules based on an attention-based framework for multi-type license plate recognition. First, we propose a global modeling module to integrate character layout information to precisely perceive the location of characters, thus generating accurate predictions. Second, a position-aware contrastive learning module is proposed to enhance the robustness and discriminability of features to alleviate character misidentification of similar glyphs. Finally, to verify the effectiveness and generality, we apply the proposed modules to six baseline models, and the results demonstrate that the proposed method can achieve state-of-the-art performance on three multi-type license plate datasets. Moreover, extensive experiments prove that our proposed modules can significantly improve performance by 6.8% on RODOSOL-ALPR with a small parameter increase.
Qi Liu 0041, Song-Lu Chen, Tian-Hao Zhang, Feng Chen 0040, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.4
2023 Stable Speech Emotion Recognition with Head-k-Pooling Loss
Chaoyue Ding, Jiakui Li, Daoming Zong, Baoxiang Li, Tian-Hao Zhang, Qunyan Zhou 0002
INTERSPEECH5
2023 InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition
Zhi-Hao Lai, Tian-Hao Zhang, Qi Liu 0041, Xinyuan Qian 0001, Li-Fang Wei, Feng Chen 0040, Song-Lu Chen, Xu-Cheng Yin
INTERSPEECH2
2023 Rethinking Speech Recognition with A Multimodal Perspective via Acoustic and Semantic Cooperative Decoding
Tian-Hao Zhang, Haibo Qin, Zhi-Hao Lai, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xinyuan Qian 0001, Xu-Cheng Yin
INTERSPEECH1
2023 Self-supervised contrastive speaker verification with nearest neighbor positive instances
Li-Fang Wei, Chuan-Fei Zhang, Tian-Hao Zhang, Song-Lu Chen, Xu-Cheng Yin
Pattern Recognit. Lett.4
2022 Non-Autoregressive Transformer with Unified Bidirectional Decoder for Automatic Speech Recognition
abstract
Non-autoregressive (NAR) transformer models have been studied intensively in automatic speech recognition (ASR), and many NAR transformer models is to use the causal mask to limit token dependencies. However, the causal mask is designed for the left-to-right decoding process of the non-parallel autoregressive (AR) transformer, which is inappropriate for the parallel NAR transformer since it ignores the right-to-left contexts. Some methods are proposed to utilize right-to-left contexts with an extra decoder, but these methods increase the model complexity. To tackle the above problems, we propose a new non-autoregressive transformer with a unified bidirectional decoder (NAT-UBD), which can simultaneously utilize left-to-right and right-to-left contexts for ASR. However, direct use of bidirectional contexts will cause information leakage, which means the decoder output can be affected by the character information of the input in the same position. To avoid information leakage, we propose a novel attention mask and modify vanilla queries, keys, and values matrices for NAT-UBD. Experimental results verify that NAT-UBD can achieve character error rates (CERs) of 5.0%/5.5% on the Aishell-1 dev/test sets, outperforming all previous NAR transformer models. Moreover, NAT-UBD can run 49.8× faster than the AR transformer baseline when decoding in a single step.
Chuan-Fei Zhang, Tian-Hao Zhang, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin
ICASSP3