VLDB 2026 Research / reviewers in the wild / expert
Yusheng Tian
dblp:251/6807
· DBLP profile ↗
7ranked-venue papers
6as first author
5since 2021 · last 2024
0009-0008-6534-4943ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction LossabstractThis research is about the creation of personalized synthetic voices for head and neck cancer survivors. It is focused particularly on tongue cancer patients whose speech might exhibit severe articulation impairment. Our goal is to restore normal articulation in the synthesized speech, while maximally preserving the target speaker’s individuality in terms of both the voice timbre and speaking style. This is formulated as a task of learning from noisy labels. We propose to augment the commonly used speech reconstruction loss with two additional terms. The first term constitutes a regularization loss that mitigates the impact of distorted articulation in the training speech. The second term is a consistency loss that encourages correct articulation in the generated speech. These additional loss terms are obtained from frame-level articulation scores of original and generated speech, which are derived using a separately trained phone classifier. Experimental results on a real case of tongue cancer patient confirm that the synthetic voice achieves comparable articulation quality to unimpaired natural speech, while effectively maintaining the target speaker’s individuality. Audio samples are available at https://myspeechproject.github.io/ArticulationRepair/. Yusheng Tian, Tan Lee |
ICASSP | 1 |
| 2023 | Diffusion-Based Mel-Spectrogram Enhancement for Personalized Speech Synthesis with Found DataabstractCreating synthetic voices with found data is challenging, as real-world recordings often contain various types of audio degradation. One way to address this problem is to pre-enhance the speech with an enhancement model and then use the enhanced data for text-to-speech (TTS) model training. This paper investigates the use of conditional diffusion models for generalized speech enhancement, which aims at addressing multiple types of audio degradation simultaneously. The enhancement is performed on the log Mel-spectrogram domain to align with the TTS training objective. Text information is introduced as an additional condition to improve the model robustness. Experiments on real-world recordings demonstrate that the synthetic voice built on data enhanced by the proposed model produces higher-quality synthetic speech, compared to those trained on data enhanced by strong baselines. Code and check-points of the proposed enhancement model are available at https://github.com/dmse4tts/DMSE4TTS. Yusheng Tian, Wei Liu 0147, Tan Lee |
ASRU | 1 |
| 2023 | Convolution-Based Channel-Frequency Attention for Text-Independent Speaker VerificationabstractDeep convolutional neural networks (CNNs) have been applied to extracting speaker embeddings with significant success in speaker verification. Incorporating the attention mechanism has shown to be effective in improving the model performance. This paper presents an efficient two-dimensional convolution-based attention module, namely C2D-Att. The interaction between the convolution channel and frequency is involved in the attention calculation by lightweight convolution layers. This requires only a small number of parameters. Fine-grained attention weights are produced to represent channel and frequency-specific information. The weights are imposed on the input features to improve the representation ability for speaker modeling. The C2D-Att is integrated into a modified version of ResNet for speaker embedding extraction. Experiments are conducted on VoxCeleb datasets. The results show that C2D-Att is effective in generating discriminative attention maps and outperforms other attention methods. The proposed model shows robust performance with different scales of model size and achieves state-of-the-art results. Yusheng Tian, Tan Lee |
ICASSP | 2 |
| 2023 | Creating Personalized Synthetic Voices from Post-Glossectomy Speech with Guided Diffusion Models
Yusheng Tian, Guangyan Zhang, Tan Lee |
INTERSPEECH | 1 |
| 2022 | Transport-Oriented Feature Aggregation for Speaker Embedding LearningabstractPooling is needed to aggregate frame-level features into utterance-level representations for speaker modeling.Given the success of statistics-based pooling methods, we hypothesize that speaker characteristics are well represented in the statistical distribution over the pre-aggregation layer's output, and propose to use transport-oriented feature aggregation for deriving speaker embeddings.The aggregated representation encodes the geometric structure of the underlying feature distribution, which is expected to contain valuable speaker-specific information that may not be represented by the commonly used statistical measures like mean and variance.The original transportoriented feature aggregation is also extended to a weightedframe version to incorporate the attention mechanism.Experiments on speaker verification with the Voxceleb dataset show improvement over statistics pooling and its attentive variant. Yusheng Tian, Tan Lee |
INTERSPEECH | 1 |
| 2020 | Improving End-to-End Speech-to-Intent Classification with ReptileabstractEnd-to-end spoken language understanding (SLU) systems have many advantages over conventional pipeline systems, but collecting in-domain speech data to train an end-to-end system is costly and time consuming. One question arises from this: how to train an end-to-end SLU with limited amounts of data? Many researchers have explored approaches that make use of other related data resources, typically by pre-training parts of the model on high-resource speech recognition. In this paper, we suggest improving the generalization performance of SLU models with a non-standard learning algorithm, Reptile. Though Reptile was originally proposed for model-agnostic meta learning, we argue that it can also be used to directly learn a target task and result in better generalization than conventional gradient descent. In this work, we employ Reptile to the task of end-to-end spoken intent classification. Experiments on four datasets of different languages and domains show improvement of intent prediction accuracy, both when Reptile is used alone and used in addition to pre-training. Yusheng Tian, Philip John Gorinski |
INTERSPEECH | 1 |
| 2013 | A dual-current-loop control method based on system current detection for LCL-filter-based Active Power FiltersabstractA dual-current-loop control method to solve the stability problem of LCL-filter-based shunt Active Power Filters (APF) with no need of additional sensors or damping resistors is proposed here. This method is based on system current detection, rather than load current detection. Making use of the inherent stability characteristic of converter-side current feedback control, the proportional (P) control of converter-side current, as inner loop, is adopted to guarantee the stability, which will not be affected by system parameter variations. The outer loop control of system current adopts the proportional resonant (PR) controller to improve steady-state accuracy. Discrete-time model of the proposed method is derived and analyzed, taking sampling, computation, and PWM output delay into consideration. Compared with the conventional method based on load current detection, the proposed one has no damping loss, higher steady-state accuracy and a much larger phase margin required for a better transient response. All these are verified by simulation results. Yusheng Tian, Qirong Jiang |
IECON | 1 |