VLDB 2026 Research / reviewers in the wild / expert
Michael Kuhlmann
dblp:133/0374
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-3664-6922ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speech Synthesis along Perceptual Voice Quality DimensionsabstractWhile expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we here address the control of perceptual voice qualities (PVQs) recognized by phonetic experts, which are speech properties at a lower level of abstraction. The ability to manipulate PVQs can be a valuable tool for teaching speech pathologists in training or voice actors. In this paper, we integrate a Conditional Continuous-Normalizing-Flow-based method into a Text-to-Speech system to modify perceptual voice attributes on a continuous scale. Unlike previous approaches, our system avoids direct manipulation of acoustic correlates and instead learns from examples. We demonstrate the system's capability by manipulating four voice qualities: Roughness, breathiness, resonance and weight. Phonetic experts evaluated these modifications, both for seen and unseen speaker conditions. The results highlight both the system's strengths and areas for improvement. Frederik Rautenberg, Michael Kuhlmann, Fritz Seebauer, Jana Wiechmann, Petra Wagner, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2025 | Towards Frame-level Quality Predictions of Synthetic SpeechabstractKuhlmann M, Seebauer FM, Wagner P, Haeb-Umbach R. Towards Frame-level Quality Predictions of Synthetic Speech. In: Interspeech 2025. Interspeech. Baixas: International Speech Communication Association; 2025: 2300-2304. Michael Kuhlmann, Fritz Seebauer, Petra Wagner, Reinhold Häb-Umbach |
INTERSPEECH | 1 |
| 2025 | Synthesizing Speech with Selected Perceptual Voice Qualities - A Case Study with Creaky VoiceabstractRautenberg F, Seebauer FM, Wiechmann J, Kuhlmann M, Wagner P, Haeb-Umbach R. Synthesizing Speech with Selected Perceptual Voice Qualities – A Case Study with Creaky Voice. In: Interspeech 2025. ISCA: ISCA; 2025: 1633-1637. Frederik Rautenberg, Fritz Seebauer, Jana Wiechmann, Michael Kuhlmann, Petra Wagner, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2022 | Investigation into Target Speaking Rate Adaptation for Voice ConversionabstractDisentangling speaker and content attributes of a speech signal into separate latent representations followed by decoding the content with an exchanged speaker representation is a popular approach for voice conversion, which can be trained with non-parallel and unlabeled speech data.However, previous approaches perform disentanglement only implicitly via some sort of information bottleneck or normalization, where it is usually hard to find a good trade-off between voice conversion and content reconstruction.Further, previous works usually do not consider an adaptation of the speaking rate to the target speaker or they put some major restrictions to the data or use case.Therefore, the contribution of this work is two-fold.First, we employ an explicit and fully unsupervised disentanglement approach, which has previously only been used for representation learning, and show that it allows to obtain both superior voice conversion and content reconstruction.Second, we investigate simple and generic approaches to linearly adapt the length of a speech signal, and hence the speaking rate, to a target speaker and show that the proposed adaptation allows to increase the speaking rate similarity with respect to the target speaker. Michael Kuhlmann, Fritz Seebauer, Janek Ebbers, Petra Wagner, Reinhold Häb-Umbach |
INTERSPEECH | 1 |
| 2021 | Contrastive Predictive Coding Supported Factorized Variational Autoencoder For Unsupervised Learning Of Disentangled Speech RepresentationsabstractIn this work we address disentanglement of style and content in speech signals. We propose a fully convolutional variational autoencoder employing two encoders: a content encoder and a style encoder. To foster disentanglement, we propose adversarial contrastive predictive coding. This new disentanglement method does neither need parallel data nor any supervision. We show that the proposed technique is capable of separating speaker and content traits into the two different representations and show competitive speaker-content disentanglement performance compared to other unsupervised approaches. We further demonstrate an increased robustness of the content representation against a train-test mismatch compared to spectral features, when used for phone recognition. Janek Ebbers, Michael Kuhlmann, Tobias Cord-Landwehr, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2013 | Comparative Visualization of Tracer Uptake in In Vivo Small Animal PET/CT Imaging of the Carotid ArteriesabstractAbstract Cardiovascular diseases are the main cause of death in the western world. Medical research on atherosclerosis is therefore of great interest and a very active research topic. We present a visualization system that supports scientists in exploring plaque development and evaluating the applicability of PET tracers for early diagnosis of cardiovascular diseases. In our application case a cone shaped cuff has been implanted around the carotid artery of ApoE knockout mice, fed with a high cholesterol western type diet. As a result, vascular lesions develop upstream and downstream from the cuff. Tracer uptake induced by these lesions needs to be analyzed in order to evaluate the effectiveness of different PET tracers. We discuss the approach previously utilized to perform this kind of analysis, the problems arising from in vivo image acquisition (in contrast to ex vivo) and the design process of our application. In close cooperation with domain experts we have developed new visualization techniques that display PET activity in the vessel wall and surrounding tissue in a single image. We use the vessel wall detected in the CT image to perform a normalized circular projection which allows the user to judge PET signal distribution in relation to the deformed vessel. Based on this projection a quantitative analysis of a defined region adjacent to the vessel wall can be performed and compared to the artery without the cuff. Stefan Diepenbrock, Sven Hermann, Michael Schäfers 0001, Michael Kuhlmann, Klaus H. Hinrichs |
Comput. Graph. Forum | 4 |