EDBT 2026 Demo / reviewers in the wild / expert
Jeff Hwang
dblp:144/7320
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Speech recognition and synthesis · 80% Representation and self-supervised learning · 20% | |
| Computer networks
1 paper |
Physical-layer communications · 46% Wireless networking · 30% Network performance modeling · 23% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.9 | 1 | 2025 | Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement · ICLR 2025 |
Natural language and speech › Speech recognition and synthesis › speech synthesis
text-to-speech |
0.9 | 1 | 2025 | Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement · ICLR 2025 |
Natural language and speech › Speech recognition and synthesis
voice conversion |
0.9 | 1 | 2025 | Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement · ICLR 2025 |
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
zero-shot text-to-speech |
0.9 | 1 | 2025 | Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement · ICLR 2025 |
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
zero-shot voice cloning |
0.9 | 1 | 2025 | Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement · ICLR 2025 |
Physical-layer communications › receiver design › linear receivers
linear MMSE receivers |
0.2 | 1 | 2014 | Performance of Multiantenna Linear MMSE Receivers in Doubly Stochastic Networks · IEEE Trans. Commun. 2014 |
Physical-layer communications
multiple-antenna systems |
0.2 | 1 | 2014 | Performance of Multiantenna Linear MMSE Receivers in Doubly Stochastic Networks · IEEE Trans. Commun. 2014 |
Wireless networking
stochastic geometry |
0.2 | 1 | 2014 | Performance of Multiantenna Linear MMSE Receivers in Doubly Stochastic Networks · IEEE Trans. Commun. 2014 |
Network performance modeling
wireless network performance analysis |
0.2 | 1 | 2014 | Performance of Multiantenna Linear MMSE Receivers in Doubly Stochastic Networks · IEEE Trans. Commun. 2014 |
Wireless networking
interference modeling |
0.1 | 1 | 2014 | Performance of Multiantenna Linear MMSE Receivers in Doubly Stochastic Networks · IEEE Trans. Commun. 2014 |
Methods — techniques the papers use, named apart from their topics
flow matching · 0.9autoregressive transformer · 0.9VQ-VAE · 0.9HuBERT · 0.9stochastic geometry · 0.2rayleigh fading · 0.2poisson point process · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Long-Form Fuzzy Speech-to-Text Alignment for 1000+ LanguagesabstractConventional speech-to-text forced alignment typically operates at the utterance level. In practice, however, we do not usually have short segments (e.g., 10 seconds) of audio with exact, verbatim transcriptions (e.g., the LibriSpeech corpus) as in lab conditions. Instead, audio often comes in long-form (e.g., an hour-long lecture recording), and the available transcription may be non-verbatim or include unspoken annotations, making it misaligned with the actual speech. This motivates the need for long-form fuzzy speech-to-text alignment, which has practical applications - for example, preparing segmented supervised audio data for training machine learning models. We demonstrate the Torchaudio long-form aligner, which supports such use cases. Moreover, it can be equipped with any CTC model that predicts frame-wise labels, turning the model into a robust and powerful aligner. Ruizhe Huang, Xiaohui Zhang 0007, Zhaoheng Ni, Moto Hira, Jeff Hwang, Vineel Pratap, Ju Lin, Ming Sun 0013, Florian Metze |
ASRU | 5 |
| 2025 | Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised DisentanglementabstractThe imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotated data, and struggle with effectively disentangling timbre and style, leading to challenges in achieving controllable generation, especially in zero-shot scenarios. To address these issues, we propose Vevo, a versatile zero-shot voice imitation framework with controllable timbre and style. Vevo operates in two core stages: (1) Content-Style Modeling: Given either text or speech's content tokens as input, we utilize an autoregressive transformer to generate the content-style tokens, which is prompted by a style reference; (2) Acoustic Modeling: Given the content-style tokens as input, we employ a flow-matching transformer to produce acoustic representations, which is prompted by a timbre reference. To obtain the content and content-style tokens of speech, we design a fully self-supervised approach that progressively decouples the timbre, style, and linguistic content of speech. Specifically, we adopt VQ-VAE as the tokenizer for the continuous hidden features of HuBERT. We treat the vocabulary size of the VQ-VAE codebook as the information bottleneck, and adjust it carefully to obtain the disentangled speech representations. Solely self-supervised trained on 60K hours of audiobook speech data, without any fine-tuning on style-specific corpora, Vevo matches or surpasses existing methods in accent and emotion conversion tasks. Additionally, Vevo’s effectiveness in zero-shot voice conversion and text-to-speech tasks further demonstrates its strong generalization and versatility. Audio samples are available at https://versavoice.github.io/. Xueyao Zhang, Kainan Peng, Vimal Manohar, Yingru Liu, Jeff Hwang, Dangna Li, Julian Chan, Zhizheng Wu 0001, Mingbo Ma |
ICLR | 7 |
| 2024 | Less Peaky and More Accurate CTC Forced Alignment by Label PriorsabstractConnectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can cause inaccurate forced alignments (FA), especially at finer granularity, e.g., phoneme level. This paper aims at alleviating the peaky behavior for CTC and improve its suitability for forced alignment generation, by leveraging label priors, so that the scores of alignment paths containing fewer blanks are boosted and maximized during training. As a result, our CTC model produces less peaky posteriors and is able to more accurately predict the offset of the tokens besides their onset. It outperforms the standard CTC model and a heuristics-based approach for obtaining CTC’s token offset timestamps by 12 − 40% in phoneme and word boundary errors (PBE and WBE) measured on the Buckeye and TIMIT data. Compared with the most widely used FA toolkit Montreal Forced Aligner (MFA), our method performs similarly on PBE/WBE on Buckeye, yet falls behind MFA on TIMIT. Nevertheless, our method has a much simpler training pipeline and better runtime efficiency. Our training recipe and pretrained model are released in TorchAudio. Ruizhe Huang, Xiaohui Zhang 0007, Zhaoheng Ni, Li Sun 0010, Moto Hira, Jeff Hwang, Vimal Manohar, Vineel Pratap, Matthew Wiesner, Shinji Watanabe 0001, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 6 |
| 2023 | TorchAudio 2.1: Advancing Speech Recognition, Self-Supervised Learning, and Audio Processing Components for PytorchabstractTorchAudio is an open-source audio and speech processing library built for PyTorch. It aims to accelerate the research and development of audio and speech technologies by providing well-designed, easy-to-use, and performant PyTorch components. Its contributors routinely engage with users to understand their needs and fulfill them by developing impactful features. Here, we survey TorchAudio’s development principles and contents and highlight key features we include in its latest version (2.1): self-supervised learning pre-trained pipelines and training recipes, high-performance CTC decoders, speech recognition models and training recipes, advanced media I/O capabilities, and tools for performing forced alignment, multi-channel speech enhancement, and reference-less speech assessment. For a selection of these features, through empirical studies, we demonstrate their efficacy and show that they achieve competitive or state-of-the-art performance. Jeff Hwang, Moto Hira, Caroline Chen, Xiaohui Zhang 0007, Zhaoheng Ni, Guangzhi Sun, Pingchuan Ma 0001, Ruizhe Huang, Vineel Pratap, Yuekai Zhang, Anurag Kumar 0003, Chin-Yun Yu, Chuang Zhu, Chunxi Liu, Jacob Kahn, Mirco Ravanelli, Shinji Watanabe 0001, Yangyang Shi, Yumeng Tao |
ASRU | 1 |
| 2022 | Torchaudio: Building Blocks for Audio and Speech ProcessingabstractThis document describes version 0.10 of TorchAudio: building blocks for machine learning applications in the audio and speech processing domain. The objective of TorchAudio is to accelerate the development and deployment of machine learning applications for researchers and engineers by providing off-the-shelf building blocks. The building blocks are designed to be GPU-compatible, automatically differentiable, and production-ready. TorchAudio can be easily installed from Python Package Index repository and the source code is publicly available under a BSD-2-Clause License (as of September 2021) at https://github.com/pytorch/audio. In this document, we provide an overview of the design principles, functionalities, and benchmarks of TorchAudio. We also benchmark our implementation of several audio and speech operations and models. We verify through the benchmarks that our implementations of various operations and models are valid and perform similarly to other publicly available implementations. Yao-Yuan Yang, Moto Hira, Zhaoheng Ni, Artyom Astafurov, Caroline Chen, Christian Puhrsch, David Pollack, Dmitriy Genzel, Donny Greenberg, Edward Z. Yang, Jason Lian, Jeff Hwang, Peter Goldsborough, Sean Narenthiran, Shinji Watanabe 0001, Soumith Chintala, Vincent Quenneville-Bélair |
ICASSP | 12 |
| 2014 | Performance of Multiantenna Linear MMSE Receivers in Doubly Stochastic NetworksabstractA technique is presented to characterize the signal-to-interference-plus-noise ratio (SINR) of a representative link with a multiantenna linear minimum-mean-square-error receiver in a wireless network with transmitting nodes distributed according to a doubly stochastic process, which is a generalization of the Poisson point process. The cumulative distribution function of the SINR of the representative link is derived, assuming independent Rayleigh fading between antennas. Several representative spatial node distributions are considered, including networks with both deterministic and random clusters, strip networks (used to model roadways, for example), hardcore networks and networks with generalized path-loss models. In addition, it is shown that if the number of antennas at the representative receiver is linearly increased with the nominal node density, the signal-to-interference ratio converges in distribution to a random variable that is nonzero in general and a positive constant in certain cases. This result indicates that to the extent that the system assumptions hold, it is possible to scale such networks by increasing the number of receiver antennas linearly with the node density. The results presented here are useful in characterizing the performance of multiantenna wireless networks in more general network models than what are currently available. Siddhartan Govindasamy, Jeff Hwang |
IEEE Trans. Commun. | 3 |