VLDB 2026 Research / reviewers in the wild / expert
Siegfried Kunzmann
dblp:64/3172
· DBLP profile ↗
27ranked-venue papers
2as first author
15since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 17 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MaZO: Masked Zeroth-Order Optimization for Multi-Task Fine-Tuning of Large Language ModelsabstractLarge language models have demonstrated exceptional capabilities across diverse tasks, but their fine-tuning demands significant memory, posing challenges for resource-constrained environments.Zeroth-order (ZO) optimization provides a memory-efficient alternative by eliminating the need for backpropagation.However, ZO optimization suffers from high gradient variance, and prior research has largely focused on single-task learning, leaving its application to multi-task learning unexplored.Multi-task learning is crucial for leveraging shared knowledge across tasks to improve generalization, yet it introduces unique challenges under ZO settings, such as amplified gradient variance and collinearity.In this paper, we present MaZO, the first framework specifically designed for multi-task LLM fine-tuning under ZO optimization.MaZO tackles these challenges at the parameter level through two key innovations: a weight importance metric to identify critical parameters and a multi-task weight update mask to selectively update these parameters, reducing the dimensionality of the parameter space and mitigating task conflicts.Experiments demonstrate that MaZO achieves state-of-the-art performance, surpassing even multi-task learning methods designed for firstorder optimization. Kai Zhen, Nathan Susanj, Athanasios Mouchtaris, Siegfried Kunzmann |
EMNLP | 6 |
| 2024 | Interleaved Audio/Audiovisual Transfer Learning for AV-ASR in Low-Resourced Languages
Patrick Blumenberg, Thomas Graave, Timo Lohrenz, Siegfried Kunzmann, Tim Fingscheidt |
INTERSPEECH | 6 |
| 2024 | CoMERA: Computing- and Memory-Efficient Training via Rank-Adaptive Tensor OptimizationabstractTraining large AI models such as LLMs and DLRMs costs massive GPUs and computing time. The high training cost has become only affordable to big tech companies, meanwhile also causing increasing concerns about the environmental impact. This paper presents CoMERA, a **Co**mputing- and **M**emory-**E**fficient training method via **R**ank-**A**daptive tensor optimization. CoMERA achieves end-to-end rank-adaptive tensor-compressed training via a multi-objective optimization formulation, and improves the training to provide both a high compression ratio and excellent accuracy in the training process. Our optimized numerical computation (e.g., optimized tensorized embedding and tensor-vector contractions) and GPU implementation eliminate part of the run-time overhead in the tensorized training on GPU. This leads to, for the first time, $2-3\times$ speedup per training epoch compared with standard training. CoMERA also outperforms the recent GaLore in terms of both memory and computing efficiency. Specifically, CoMERA is $2\times$ faster per training epoch and $9\times$ more memory-efficient than GaLore on a tested six-encoder transformer with single-batch training. Our method also shows $\sim 2\times$ speedup than standard pre-training on a BERT-like code-generation LLM while achieving $4.23\times$ compression ratio in pre-training.
With further HPC optimization, CoMERA may reduce the pre-training cost of many other LLMs. An implementation of CoMERA is available at <https://github.com/ziyangjoy/CoMERA>. Samridhi Choudhary, Xinfeng Xie, Cao Gao, Siegfried Kunzmann |
NeurIPS | 6 |
| 2023 | Parameter-Efficient Cross-Language Transfer Learning for a Language-Modular Audiovisual Speech RecognitionabstractIn audiovisual speech recognition (AV-ASR), for many languages only few audiovisual data is available. Building upon an English model, in this work, we first apply and analyze various adapters for cross-language transfer learning to build a parameter-efficient and easy-to-extend AV-ASR in multiple languages. Fine-tuning only the bottleneck adapter with 4% of encoder’s parameters and the decoder shows comparable performance to full fine-tuning in French and Spanish AV-ASR. Second, we investigate the effectiveness of various encoder components in cross-language transfer learning. Our proposed modular linguistic transfer learning approach excels the full fine-tuning method for German, French, and Spanish AV-ASR in almost all clean and noisy conditions (8/9). On low-resourced German AV data (13h), our proposed linguistic transfer learning achieves a 4.1% abs. WER reduction on average for clean and noisy speech, while fine-tuning only 50% of the encoder’s parameters. Our code is at GitHub.11https://github.com/ifnspaml/Cross_Language_Transfer_Learning_AVASR.git Thomas Graave, Timo Lohrenz, Siegfried Kunzmann, Tim Fingscheidt |
ASRU | 5 |
| 2023 | Multilingual End-To-End Spoken Language Understanding For Ultra-Low Footprint ApplicationsabstractTiny Signal-to-Interpretation (TinyS2I) has been recently introduced as an ultra low-footprint end-to-end spoken language understanding (SLU) model. This architecture is capable of running in ultra resource constrained environments like voice assistant devices, while at the same time reducing latency. In this work, we propose an extension to TinyS2I and train a multilingual system supporting several languages. Multilingual TinyS2I models show little to no degradation compared to their monolingual counterparts. Increasing the network size in width and depth improves the classification accuracy for mono- and multilingual setups, with the multilingual one improving beyond the monolingual accuracy. This enables users to interact with the device in the language of their choice and dynamically switch between languages without an explicit language setting or accuracy degradation. Anastasios Alexandridis, Zach Trozenski, Joel Whiteman, Grant P. Strimel, Nathan Susanj, Athanasios Mouchtaris, Siegfried Kunzmann |
ICASSP | 8 |
| 2023 | Dual-Attention Neural Transducers for Efficient Wake Word Spotting in Speech RecognitionabstractWe present dual-attention neural biasing, an architecture designed to boost Wake Words (WW) recognition and improve inference time latency on speech recognition tasks. This architecture enables a dynamic switch for its runtime compute paths by exploiting WW spotting to select which branch of its attention networks to execute for an input audio frame. With this approach, we effectively improve WW spotting accuracy while saving runtime compute cost as defined by floating point operations (FLOPs). Using an in-house de-identified dataset, we demonstrate that the proposed dual-attention network can reduce the compute cost by 90% for WW audio frames, with only 1% increase in the number of parameters. This architecture improves WW F1 score by 16% relative and improves generic rare word error rate by 3% relative compared to the baselines. Saumya Y. Sahai, Thejaswi Muniyappa, Kanthashree Mysore Sathyendra, Anastasios Alexandridis, Grant P. Strimel, Ross McGowan, Ariya Rastrow, Feng-Ju Chang, Athanasios Mouchtaris, Siegfried Kunzmann |
ICASSP | 11 |
| 2023 | Quantization-aware and Tensor-compressed Training of Transformers for Natural Language UnderstandingabstractFine-tuned transformer models have shown superior performances in many natural language tasks.However, the large model size prohibits deploying high-performance transformer models on resource-constrained devices.This paper proposes a quantization-aware tensor-compressed training approach to reduce the model size, arithmetic operations, and ultimately runtime latency of transformer-based models.We compress the embedding and linear layers of transformers into small low-rank tensor cores, which significantly reduces model parameters.A quantization-aware training with learnable scale factors is used to further obtain low-precision representations of the tensorcompressed models.The developed approach can be used for both end-to-end training and distillation-based training.To improve the convergence, a layer-by-layer distillation is applied to distill a quantized and tensor-compressed student model from a pre-trained transformer.The performance is demonstrated in two natural language understanding tasks, showing up to 63× compression ratio, little accuracy loss and remarkable inference and training speedup. Samridhi Choudhary, Siegfried Kunzmann |
INTERSPEECH | 3 |
| 2022 | Tie Your Embeddings Down: Cross-Modal Latent Spaces for End-to-end Spoken Language UnderstandingabstractEnd-to-end (E2E) spoken language understanding (SLU) systems can infer the semantics of a spoken utterance directly from an audio signal. However, training an E2E system remains a challenge, largely due to the scarcity of paired audio-semantics data. In this paper, we consider an E2E system as a multi-modal model, with audio and text functioning as its two modalities, and use a cross-modal latent space (CMLS) architecture, where a shared latent space is learned between the ‘acoustic’ and ‘text’ embeddings. We propose using different multi-modal losses to explicitly align the acoustic embedding to the text embeddings (obtained via a semantically powerful pre-trained BERT model) in the latent space. We train the CMLS model on two publicly available E2E datasets and one internal dataset, across different cross-modal losses. Our proposed triplet loss function achieves the best performance. It achieves a relative improvement of 22.1% over an E2E model without a cross-modal space and a relative improvement of 2.8% over a previously published CMLS model using L2loss on our internal dataset. Bhuvan Agrawal, Samridhi Choudhary, Martin Radfar, Athanasios Mouchtaris, Ross McGowan, Nathan Susanj, Siegfried Kunzmann |
ICASSP | 8 |
| 2022 | Caching Networks: Capitalizing on Common Speech for ASRabstractWe introduce Caching Networks (CachingNets), a speech recognition network architecture capable of delivering faster, more accurate decoding by leveraging common speech patterns. By explicitly incorporating select sentences unique to each user into the network’s design, we show how to train the model as an extension of the popular sequence transducer architecture through a multitask learning procedure. We further propose and experiment with different phrase caching policies, which are effective for virtual voice-assistant (VA) applications, to complement the architecture. Our results demonstrate that by pivoting between different inference strategies on the fly, CachingNets can deliver significant performance improvements. Specifically, on an industrial-scale, VA ASR task, we observe up to 7.4% relative word error rate (WER) and 11% sentence error rate (SER) improvements with accompanied latency gains. Anastasios Alexandridis, Grant P. Strimel, Ariya Rastrow, Pavel Kveton, Maurizio Omologo, Siegfried Kunzmann, Athanasios Mouchtaris |
ICASSP | 7 |
| 2022 | Contextual Adapters for Personalized Speech Recognition in Neural TransducersabstractPersonal rare word recognition in end-to-end Automatic Speech Recognition (E2E ASR) models is a challenge due to the lack of training data. A standard way to address this issue is with shallow fusion methods at inference time. However, due to their dependence on external language models and the deterministic approach to weight boosting, their performance is limited. In this paper, we propose training neural contextual adapters for personalization in neural transducer based ASR models. Our approach can not only bias towards user-defined words, but also has the flexibility to work with pretrained ASR models. Using an in-house dataset, we demonstrate that contextual adapters can be applied to any general purpose pretrained ASR model to improve personalization. Our method outperforms shallow fusion, while retaining functionality of the pretrained models by not altering any of the model weights. We further show that the adapter style training is superior to full-fine-tuning of the ASR models on datasets with user-defined content. Kanthashree Mysore Sathyendra, Thejaswi Muniyappa, Feng-Ju Chang, Jinru Su, Grant P. Strimel, Athanasios Mouchtaris, Siegfried Kunzmann |
ICASSP | 8 |
| 2021 | Context-Aware Transformer Transducer for Speech RecognitionabstractEnd-to-end (E2E) automatic speech recognition (ASR) systems often have difficulty recognizing uncommon words, that appear infrequently in the training data. One promising method, to improve the recognition accuracy on such rare words, is to latch onto personalized/contextual information at inference. In this work, we present a novel context-aware transformer transducer (CATT) network that improves the state-of-the-art transformer-based ASR system by taking advantage of such contextual signals. Specifically, we propose a multi-head attention-based context-biasing network, which is jointly trained with the rest of the ASR sub-networks. We explore different techniques to encode contextual data and to create the final attention context vectors. We also leverage both BLSTM and pretrained BERT based models to encode contextual data and guide the network training. Using an in-house far-field dataset, we show that CATT, using a BERT based context encoder, improves the word error rate of the baseline transformer transducer and outperforms an existing deep contextual model by 24.2% and 19.4% respectively. Feng-Ju Chang, Martin Radfar, Athanasios Mouchtaris, Maurizio Omologo, Ariya Rastrow, Siegfried Kunzmann |
ASRU | 7 |
| 2021 | In Pursuit of Babel - Multilingual End-to-End Spoken Language UnderstandingabstractEnd-to-end spoken language understanding (E2E SLU) systems predict the utterance semantics directly from speech. So far, to the best of our knowledge, E2E models have only been trained to recognize the semantics for a single language. In this work we introduce the first multilingual E2E SLU system and present results across three languages - English, Spanish and French. We propose a transformer-based, multilingual acoustic encoder to predict intents, that leverages pre-training for both acoustic and linguistic modalities of the SLU model. It learns a robust, cross-modal latent space using a pre-trained multilingual BERT as a semantic teacher. The best performing model achieves relative improvements of 7.2% in a single language setting, 5-6% in two, and 4-6% in three language settings. An intent-wise analysis shows that semantic supervision becomes more important for shorter utterances, while providing an explicit language identifier at the input leads to lower intent classification errors. Samridhi Choudhary, Clement Chung, Athanasios Mouchtaris, Siegfried Kunzmann |
ASRU | 5 |
| 2021 | End-to-End Multi-Channel Transformer for Speech RecognitionabstractTransformers are powerful neural architectures that allow integrating different modalities using attention mechanisms. In this paper, we leverage the neural transformer architectures for multi-channel speech recognition systems, where the spectral and spatial information collected from different microphones are integrated using attention layers. Our multi-channel transformer network mainly consists of three parts: channel-wise self attention layers (CSA), cross-channel attention layers (CCA), and multi-channel encoder-decoder attention layers (EDA). The CSA and CCA layers encode the contextual relationship "within" and "between" channels and across time, respectively. The channel-attended outputs from CSA and CCA are then fed into the EDA layers to help decode the next token given the preceding ones. The experiments show that in a far-field in-house dataset, our method outperforms the baseline single-channel transformer, as well as the super-directive and neural beamformers cascaded with the transformers. Feng-Ju Chang, Martin Radfar, Athanasios Mouchtaris, Brian John King, Siegfried Kunzmann |
ICASSP | 5 |
| 2021 | Joint ASR and Language Identification Using RNN-T: An Efficient Approach to Dynamic Language SwitchingabstractConventional dynamic language switching enables seamless multilingual interactions by running several monolingual ASR systems in parallel and triggering the appropriate downstream components using a standalone language identification (LID) service. Since this solution is neither scalable nor cost- and memory-efficient, especially for on-device applications, we propose end-to-end, streaming, joint ASR-LID architectures based on the recurrent neural network transducer framework. Two key formulations are explored: (1) joint training using a unified output space for ASR and LID vocabularies, and (2) joint training viewed as multi-task optimization. We also evaluate the benefit of using auxiliary language information obtained on-the-fly from an acoustic LID classifier. Experiments with the English-Hindi language pair show that: (a) multi-task architectures perform better overall, and (b) the best joint architecture surpasses monolingual ASR (6.4–9.2% word error rate reduction) and acoustic LID (53.9–56.1% error rate reduction) baselines while reducing the overall memory footprint by up to 46%. Surabhi Punjabi, Harish Arsikere, Zeynab Raeesy, Chander Chandak, Nikhil Bhave, Ankish Bansal, Sergio Murillo, Ariya Rastrow, Andreas Stolcke, Jasha Droppo, Sri Garimella, Roland Maas, Mathieu Hans, Athanasios Mouchtaris, Siegfried Kunzmann |
ICASSP | 16 |
| 2021 | FANS: Fusing ASR and NLU for On-Device SLUabstractSpoken language understanding (SLU) systems translate voice input commands to semantics which are encoded as an intent and pairs of slot tags and values.Most current SLU systems deploy a cascade of two neural models where the first one maps the input audio to a transcript (ASR) and the second predicts the intent and slots from the transcript (NLU).In this paper, we introduce FANS, a new end-to-end SLU model that fuses an ASR audio encoder to a multi-task NLU decoder to infer the intent, slot tags, and slot values directly from a given input audio, obviating the need for transcription.FANS consists of a shared audio encoder and three decoders, two of which are seq-to-seq decoders that predict non null slot tags and slot values in parallel and in an auto-regressive manner.FANS neural encoder and decoders architectures are flexible which allows us to leverage different combinations of LSTM, self-attention, and attenders.Our experiments show compared to the state-of-the-art end-toend SLU models, FANS reduces ICER and IRER errors relatively by 30% and 7%, respectively, when tested on an in-house SLU dataset and by 0.86% and 2% absolute when tested on a public SLU dataset. Martin Radfar, Athanasios Mouchtaris, Siegfried Kunzmann, Ariya Rastrow |
Interspeech | 3 |
| 2020 | Multilingual Grapheme-To-Phoneme Conversion with Byte RepresentationabstractGrapheme-to-phoneme (G2P) models convert a written word into its corresponding pronunciation and are essential components in automatic-speech-recognition and text-to-speech systems. Recently, the use of neural encoder-decoder architectures has substantially improved G2P accuracy for mono- and multi-lingual cases. However, most multilingual G2P studies focus on sets of languages that share similar graphemes, such as European languages. Multilingual G2P for languages from different writing systems, e.g. European and East Asian, remains an understudied area. In this work, we propose a multilingual G2P model with byte-level input representation to accommodate different grapheme systems, along with an attention-based Transformer architecture. We evaluate the performance of both character-level and byte-level G2P using data from multiple European and East Asian locales. Models using byte representation yield 16.2%– 50.2% relative word error rate improvement over character-based counterparts for mono- and multi-lingual use cases. In addition, byte-level models are 15.0%–20.1% smaller in size. Our results show that byte is an efficient representation for multilingual G2P with languages having large grapheme vocabularies. Mingzhi Yu, Hieu Duy Nguyen, Alex Sokolov, Jack Lepird, Kanthashree Mysore Sathyendra, Samridhi Choudhary, Athanasios Mouchtaris, Siegfried Kunzmann |
ICASSP | 8 |
| 2020 | End-to-End Neural Transformer Based Spoken Language UnderstandingabstractSpoken language understanding (SLU) refers to the process of inferring the semantic information from audio signals. While the neural transformers consistently deliver the best performance among the state-of-the-art neural architectures in field of natural language processing (NLP), their merits in a closely related field, i.e., spoken language understanding (SLU) have not beed investigated. In this paper, we introduce an end-to-end neural transformer-based SLU model that can predict the variable-length domain, intent, and slots vectors embedded in an audio signal with no intermediate token prediction architecture. This new architecture leverages the self-attention mechanism by which the audio signal is transformed to various sub-subspaces allowing to extract the semantic context implied by an utterance. Our end-to-end transformer SLU predicts the domains, intents and slots in the Fluent Speech Commands dataset with accuracy equal to 98.1 \%, 99.6 \%, and 99.6 \%, respectively and outperforms the SLU models that leverage a combination of recurrent and convolutional neural networks by 1.4 \% while the size of our model is 25\% smaller than that of these architectures. Additionally, due to independent sub-space projections in the self-attention layer, the model is highly parallelizable which makes it a good candidate for on-device SLU. Martin Radfar, Athanasios Mouchtaris, Siegfried Kunzmann |
INTERSPEECH | 3 |
| 2011 | Online Speaker Adaptation with Pre-Computed FMLLR Transformations
Volker Fischer 0002, Siegfried Kunzmann |
INTERSPEECH | 2 |
| 2006 | From pre-recorded prompts to corporate voices: on the migration of interactive voice response applicationsabstractThis paper describes our efforts towards the creation of corporate synthetic voices from low quality speech data, as it can typically be found on many Interactive Voice Response (IVR) units. In doing so, we first touch on several normalization techniques that aim on a better support of a highly automated voice construction process. Subsequently, we describe methods for the creation of enriched corporate voices which integrate speech recordings from different speakers in order to overcome problems arising from limited domain training data. Experiments are described which demonstrate the feasibility of the approach by comparing it to a less flexible solution that uses pre-recorded prompts in combination with a large footprint standard concatenative synthesizer. Results show that the enriched voices clearly outperform those voices build solely from IVR data, while achieving almost the same overall rating as the pre-recorded prompts solution. Volker Fischer 0002, Siegfried Kunzmann |
INTERSPEECH | 2 |
| 2004 | Multilingual acoustic models for speech recognition and synthesisabstractIn this paper, we review the design of a common phone alphabet for up to fifteen languages and describe its application in two important components of a seamless multilingual conversational system, namely speech recognition and synthesis. We report on experiments that demonstrate the advantages of multilingual acoustic models both for the recognition of foreign names and non-native speech, and describe the usefulness of a common phone alphabet for the construction of unit selection based mono- and bilingual speech synthesis systems. Siegfried Kunzmann, Volker Fischer 0002, Jorge Gonzalez, Ossama Emam, Carsten Günther, Eric Janke |
ICASSP (3) | 1 |
| 2004 | Domain adaptation methods in the IBM trainable text-to-speech systemabstractThis paper presents a comparison of domain adaptation techniques for a unit selection based text-to-speech system. The methods under investigation consider two different pre-requisites, namely the absence and the existence of addi-tional domain specific training prompts, spoken by the orig-inal voice talent. Whereas in the first case we employ do-main specific pre-selection, for the latter we compare a va-riety of methods that range from a simple extension of the segment inventory to a complete reconstruction of the sys-tem, which also includes the training of decision trees for the domain dependent prediction of prosody targets. An ex-perimental evaluation of the methods under consideration unveils significant improvements (up to 1.1 on a 5 point MOS scale) over the baseline system for sentences from the target domain, while showing no significant degradation when synthesizing sentences from other than the adaptation domain. 1. Volker Fischer 0002, Jaime Botella Ordinas, Siegfried Kunzmann |
INTERSPEECH | 3 |
| 2003 | Recent progress in the decoding of non-native speech with multilingual acoustic models
Volker Fischer 0002, Eric Janke, Siegfried Kunzmann |
INTERSPEECH | 3 |
| 2002 | Likelihood combination and recognition output voting for the decoding of non-native speech with multilingual HMMs
Volker Fischer 0002, Eric Janke, Siegfried Kunzmann |
INTERSPEECH | 3 |
| 2000 | A data-driven methodology for the production of multilingual conversational systems
Ossama Emam, Jorge Gonzalez, Carsten Günther, Eric Janke, Siegfried Kunzmann, Giulio Maltese, Claire Waast-Richard |
INTERSPEECH | 5 |
| 2000 | Acoustic language model classes for a large vocabulary continuous speech recognizer
Volker Fischer 0002, Siegfried Kunzmann |
INTERSPEECH | 2 |
| 2000 | SPEECON - Speech Data for Consumer Devices
Rainer Siemund, Harald Höge, Siegfried Kunzmann, Krzysztof Marasek |
LREC | 3 |
| 1988 | An experimental environment for the generation and verification of word hypotheses in continuous speech
Siegfried Kunzmann, Thomas Kuhn 0002, Heinrich Niemann |
Speech Commun. | 1 |