Haowei Lou

dblp:287/7211 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
5since 2021 · last 2026
0009-0009-1359-872XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
abstract
Learning representative embeddings for different types of speaking styles, such as emotion, age, and gender, is critical for both recognition tasks (e.g., cognitive computing and human-computer interaction) and generative tasks (e.g., style-controllable speech generation). In this work, we introduce ParaMETA, a unified and flexible framework for learning and controlling speaking styles directly from speech. Unlike existing methods that rely on single-task models or cross-modal alignment, ParaMETA learns disentangled, task-specific embeddings by projecting speech into dedicated subspaces for each style type. This design reduces inter-task interference, mitigates negative transfer, and allows a single model to handle multiple paralinguistic tasks such as emotion, gender, age, and nationality classification. Beyond recognition, ParaMETA enables fine-grained style control in Text-To-Speech (TTS) generative models. It supports both speech- and text-based prompting and allows users to modify one speaking style while preserving others. Extensive experiments demonstrate that ParaMETA outperforms strong baselines in classification accuracy and generates more natural and expressive speech, while maintaining a lightweight and efficient model suitable for real-world applications.
Haowei Lou, Hye-Young Paik, Wen Hu 0001, Lina Yao 0001
AAAI1
2025 ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
abstract
Controlling speaking style in text-to-speech (TTS) systems has become a growing focus in both academia and industry. While many existing approaches rely on reference audio to guide style generation, such methods are often impractical due to privacy concerns and limited accessibility. More recently, large language models (LLMs) have been used to control speaking style through natural language prompts; however, their high computational cost, lack of interpretability, and sensitivity to prompt phrasing limit their applicability in real-time and resource-constrained environments. In this work, we propose ParaStyleTTS, a lightweight and interpretable TTS framework that enables expressive style control from text prompts alone. ParaStyleTTS features a novel two-level style adaptation architecture that separates prosodic and paralinguistic speech style modeling. It allows fine-grained and robust control over factors such as emotion, gender, and age. Unlike LLM-based methods, ParaStyleTTS maintains consistent style realization across varied prompt formulations and is well-suited for real-world applications, including on-device and low-resource deployment. Experimental results show that ParaStyleTTS generates high-quality speech with performance comparable to state-of-the-art LLM-based systems while being 30x faster, using 8x fewer parameters, and requiring 2.5x less CUDA memory. Moreover, ParaStyleTTS exhibits superior robustness and controllability over paralinguistic speaking styles, providing a practical and efficient solution for style-controllable text-to-speech generation. Demo can be found at https://parastyletts.github.io/ParaStyleTTS_Demo/. Code can be found at https://github.com/haoweilou/ParaStyleTTS.
Haowei Lou, Hye-Young Paik, Wen Hu 0001, Lina Yao 0001
CIKM1
2025 Achoio: A Skill-Aware Evaluation Management System for Text-To-Speech Research
abstract
Human subjective evaluation plays a crucial role in evaluating speech-related generative tasks such as text-to-speech (TTS) generation. However, current practices are often constrained by limited scalability, fragmented workflows, and inconsistent rating reliability. Researchers frequently rely on manual methods or general-purpose crowdsourcing systems, where recruiting appropriately skilled listeners is challenging, and result analysis is labor-intensive. In this work, we introduce Achoio, a dedicated end-to-end online system designed to streamline and scale human evaluation for the TTS research community. Achoio allows researchers to create and manage evaluation projects, upload synthesized speech samples, and automatically match them with qualified listeners based on linguistic proficiency and domain knowledge. The system provides built-in tools for project status tracking, result aggregation and visualization. In this demonstration, we will walk through the core features of Achoio, including intuitive project setup, skill-based listener matching algorithm, and automated analytics. By addressing the limitations of existing workflows, Achoio offers a scalable, domain-aware, and analysis-ready solution for conducting high-quality subjective TTS evaluations. Our system is live and can be found at https://www.achoio.com. Demo is available on YouTube at https://youtu.be/Ugjj3_YooSM.
Haowei Lou, Hye-Young Paik, Basem Suleiman, Wen Hu 0001, Lina Yao 0001
CIKM1
2025 LatentSpeech: Latent Diffusion for Text-To-Speech Generation
abstract
Text-To-Speech (TTS) generation plays a crucial role in human-robot interaction by allowing robots to communicate naturally with humans. Researchers have developed various TTS models to enhance speech generation. More recently, diffusion models have emerged as a powerful generative framework, achieving state-of-the-art performance in tasks such as image and video generation. However, their application in TTS has been limited by its slow inference speeds due to their iterative denoising process. Previous work has applied diffusion models to Mel-Spectrograms with an additional vocoder to convert them into waveforms. To address these limitations, we propose LatentSpeech, a novel diffusion-based TTS framework that operates directly in a latent space. This space is significantly more compact and information-rich than raw Mel-Spectrograms. Furthermore, we introduce an alternative latent space of Pseudo-Quadrature Mirror Filters (PQMF), which decomposes speech into multiple subbands. By leveraging PQMF’s near-perfect waveform reconstruction capability, LatentSpeech eliminates the need for a separate vocoder and reduces both model size and inference time. Our PQMF-based LatentSpeech model reduces inference time by 45% and model size by 77% compared to Mel-Spectrogram diffusion models. On benchmark datasets, it achieves 25% lower WER and 58% higher MOS using the same training data. These results highlight LatentSpeech as an efficient, high-quality TTS solution for real-time and human-robot interaction. Code and models are available here.
Haowei Lou, Hye-Young Paik, Pari Delir Haghighi, Sheng Li 0010, Wen Hu 0001, Lina Yao 0001
RO-MAN1
2024 StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
Haowei Lou, Hye-Young Paik, Wen Hu 0001, Lina Yao 0001
MMAsia1