Jan Niehues

dblp:120/0365 · DBLP profile ↗
← Back
60ranked-venue papers
7as first author
30since 2021 · last 2026
0000-0002-4231-6543ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 7 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Sigmoid Head for Quality Estimation under Language Ambiguity
abstract
Language model (LM) probability is not a reliable quality estimator, as natural language is ambiguous.When multiple output options are valid, the model's probability distribution is spread across them, which can misleadingly indicate low output quality.This issue is caused by two reasons: (1) LMs' final output activation is softmax, which does not allow multiple correct options to receive high probabilities simultaneuously and (2) LMs' training data is single, one-hot encoded references, indicating that there is only one correct option at each output step.We propose training a module for Quality Estimation on top of pre-trained LMs to address these limitations.The module, called Sigmoid Head, is an extra unembedding head with sigmoid activation to tackle the first limitation.To tackle the second limitation, during the negative sampling process to train the Sigmoid Head, we use a heuristic to avoid selecting potentially alternative correct tokens.Our Sigmoid Head is computationally efficient during training and inference.The probability from Sigmoid Head is notably better quality signal compared to the original softmax head.As the Sigmoid Head does not rely on humanannotated quality data, it is more robust to outof-domain settings compared to supervised QE.
Tu Anh Dinh, Jan Niehues
ACL (1)2
2026 Beyond Transcripts: A Renewed Perspective on Audio Chaptering
abstract
Audio chaptering, the task of segmenting longform audio into coherent sections, is increasingly important for navigating podcasts, lectures, and videos.Despite its relevance, research remains limited and text-based, leaving key questions unresolved about leveraging audio information, handling ASR errors, and transcript-free evaluation.We address these gaps through three contributions: (1) a systematic comparison between text-based models with acoustic features, a novel audio-only architecture (AudioSeg) operating on learned audio representations, and multimodal LLMs;(2) empirical analysis of factors affecting performance, including transcript quality, acoustic features, duration, and speaker composition; and (3) formalized evaluation protocols contrasting transcript-dependent text-space protocols with transcript-invariant time-space protocols.Our experiments on YTSeg reveal that AudioSeg substantially outperforms text-based approaches, pauses provide the largest acoustic gains, and MLLMs remain limited by context length and weak instruction following, yet MLLMs are promising on shorter audio. 1
Fabian Retkowski, Maike Züfle, Thai-Binh Nguyen, Jan Niehues, Alex Waibel
ACL (1)4
2026 Talk2Ref: A Dataset for Reference Prediction from Scientific Talks
Frederik Yannick Broy, Maike Züfle, Jan Niehues
LREC3
2026 MuSaG: A Multimodal German Sarcasm Dataset with Full-Modal Annotations
Aaron Robert Scott, Maike Züfle, Jan Niehues
LREC3
2026 MUSCAT: MUltilingual, SCientific ConversATion Benchmark
Supriti Sinhamahapatra, Thai-Binh Nguyen, Yigit Oguz, Enes Yavuz Ugan, Jan Niehues, Alex Waibel
LREC5
2025 Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMs
abstract
While large language models demonstrate remarkable capabilities at task-specific applications through fine-tuning, extending these benefits across diverse languages is essential for broad accessibility.However, effective crosslingual transfer is hindered by LLM performance gaps across languages and the scarcity of fine-tuning data in many languages.Through analysis of LLM internal representations from over 1,000+ language pairs, we discover that middle layers exhibit the strongest potential for cross-lingual alignment.Building on this finding, we propose a middle-layer alignment objective integrated into task-specific training.Our experiments on slot filling, machine translation, and structured text generation show consistent improvements in cross-lingual transfer, especially to lower-resource languages.The method is robust to the choice of alignment languages and generalizes to languages unseen during alignment.Furthermore, we show that separately trained alignment modules can be merged with existing task-specific modules, improving cross-lingual capabilities without full re-training.Our code is publicly available 1 .0 4 8 12 16 20 24 28 32 Layer ID 0 50 100 Avg.retrieval accuracy (%) Llama 3 0 4 8 12 16 20 24 28 Layer ID Qwen 2.5 Overall Low-res.(a) Cross-lingual semantic alignment (measured by average retrieval accuracy over 35 languages and 1190 language directions) varies by layer, with the middle layer showing the highest score.Lower-resource languages are poorly aligned.
Jan Niehues
ACL (1)2
2025 Making Lecture Videos Accessible for Students who are Blind or have Low Vision through AI-Assisted Navigation and Visual Question Answering
abstract
Designing accessible lectures and lecture materials is crucial to promote inclusive higher education.We conducted need-finding interviews with 12 students who are blind or have low vision to learn their perspectives on how lectures and lecture material could become more accessible through Artificial Intelligence (AI) technologies.Key insights from the interviews reveal that students envision AI to automatically customize lecture material, connect disparate information sources, for example, to better keep track of the current lecture slide, and enhance interaction and engagement with lecture material.Based on these insights, we developed the LectureAssistant prototype, employing an iterative design process with visually impaired users that features AI-assisted video navigation and chatbot interaction.In a final evaluation with seven students, the participants expressed enthusiasm for features such as AI-powered video search and the possibility of asking questions about visual content in the current video frame.They provided valuable suggestions for future improvements, including notifications for lecture slide transitions and the provision of a short overview function for a slide.Insights from the study indicate great potential of the prototype to improve accessibility of lecture videos for students with visual impairments, although they also point to crucial areas for improvement, such as more reliable and personalized image descriptions.
Katharina Anderer, Karin Müller 0001, Lukas Strobel, Matthias Wölfel, Jan Niehues, Kathrin Maria Gerling
ASSETS5
2025 Are Generative Models Underconfident? Better Quality Estimation with Boosted Model Probability
abstract
Quality Estimation (QE) is estimating the quality of the model output during inference when the ground truth is not available.Deriving output quality from the models' output probability is the most trivial and low-effort way.However, we show that the output probability of text-generation models can appear underconfident.At each output step, there can be multiple correct options, making the probability distribution spread out more.Thus, lower probability does not necessarily mean lower output quality.Due to this observation, we propose a QE approach called BOOSTEDPROB 1 , which boosts the model's confidence in cases where there are multiple viable output options.With no increase in complexity, BOOSTEDPROB is notably better than raw model probability in different settings, achieving on average +0.194 improvement in Pearson correlation to groundtruth quality.It also comes close to or outperforms more costly approaches like supervised or ensemble-based QE in certain settings.
Tu Anh Dinh, Jan Niehues
EMNLP2
2025 Summarizing Speech: A Comprehensive Survey
abstract
Fabian Retkowski, Maike Züfle, Andreas Sudmann, Dinah Pfau, Shinji Watanabe, Jan Niehues, Alexander Waibel. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Fabian Retkowski, Maike Züfle, Andreas Sudmann, Dinah Pfau, Shinji Watanabe 0001, Jan Niehues, Alex Waibel
EMNLP6
2025 Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talks
abstract
State-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems primarily rely on acoustic information while disregarding additional multi-modal context.However, visual information are essential in disambiguation and adaptation.While most work focus on speaker images to handle noise conditions, this work also focuses on integrating presentation slides for the use cases of scientific presentation.In a first step, we create a benchmark for multimodal presentation including an automatic analysis of transcribing domain-specific terminology.Next, we explore methods for augmenting speech models with multi-modal information.We mitigate the lack of datasets with accompanying slides by a suitable approach of data augmentation.Finally, we train a model using the augmented dataset, resulting in a relative reduction in word error rate of approximately 34%, across all words and 35%, for domainspecific terms compared to the baseline model.Our implementation is available 1 .
Supriti Sinhamahapatra, Jan Niehues
EMNLP2
2025 In-context Language Learning for Endangered Languages in Speech Recognition
Zhaolin Li, Jan Niehues
INTERSPEECH2
2024 Speech Recognition Corpus of the Khinalug Language for Documenting Endangered Languages
abstract
Automatic Speech Recognition (ASR) can be a valuable tool to document endangered languages. However, building ASR tools for these languages poses several difficult research challenges, notably data scarcity. In this paper, we show the whole process of creating a useful ASR tool for language documentation scenarios. We publish the first speech corpus for Khinalug, an endangered language spoken in Northern Azerbaijan. The corpus consists of 2.67 hours of labeled data from recordings of spontaneous speech about various topics. As Khinalug is an extremely low-resource language, we investigate the benefits of multilingual models for self-supervised learning and supervised learning and achieve the performance of 6.65 Character Error Rate (CER) points and 25.53 Word Error Rate (WER) points. The benefits of multilingual models are further validated through experimentation with three additional under-resourced languages. Lastly, this work conducts quality assessments with linguists on new recordings to investigate the model’s usefulness in language documentation. We observe an evident degradation for new recordings, indicating the importance of enhancing model robustness. In addition, we find the inaudible content is the main cause of wrong ASR predictions, suggesting relating work on incorporating contextual information.
Zhaolin Li, Monika Rind-Pawlowski, Jan Niehues
LREC/COLING3
2024 Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
abstract
Human evaluation is a critical component in machine translation system development and has received much attention in text translation research. However, little prior work exists on the topic of human evaluation for speech translation, which adds additional challenges such as noisy data and segmentation mismatches. We take the first steps to fill this gap by conducting a comprehensive human evaluation of the results of several shared tasks from the last International Workshop on Spoken Language Translation (IWSLT 2023). We propose an effective evaluation strategy based on automatic resegmentation and direct assessment with segment context. Our analysis revealed that: 1) the proposed evaluation strategy is robust and scores well-correlated with other types of human judgements; 2) automatic metrics are usually, but not always, well-correlated with direct assessment scores; and 3) COMET as a slightly stronger automatic metric than chrF, despite the segmentation noise introduced by the resegmentation step systems. We release the collected human-annotated data in order to encourage further investigation.
Matthias Sperber, Ondrej Bojar, Barry Haddow, Dávid Javorský, Xutai Ma, Matteo Negri, Jan Niehues, Peter Polak, Elizabeth Salesky, Katsuhito Sudoh, Marco Turchi
LREC/COLING7
2024 How Transferable are Attribute Controllers on Pretrained Multilingual Translation Models?
abstract
Customizing machine translation models to comply with desired attributes (e.g., formality or grammatical gender) is a well-studied topic.However, most current approaches rely on (semi-)supervised data with attribute annotations.This data scarcity bottlenecks democratizing such customization possibilities to a wider range of languages, particularly lowerresource ones.This gap is out of sync with recent progress in pretrained massively multilingual translation models.In response, we transfer the attribute controlling capabilities to languages without attribute-annotated data with an NLLB-200 model as a foundation.Inspired by techniques from controllable generation, we employ a gradient-based inference-time controller to steer the pretrained model.The controller transfers well to zero-shot conditions, as it operates on pretrained multilingual representations and is attribute-rather than languagespecific.With a comprehensive comparison to finetuning-based control, we demonstrate that, despite finetuning's clear dominance in supervised settings, the gap to inference-time control closes when moving to zero-shot conditions, especially with new and distant target languages.The latter also shows stronger domain robustness.We further show that our inference-time control complements finetuning.A human evaluation on a real low-resource language, Bengali, confirms our findings.Our code is here.
Jan Niehues
EACL (1)2
2024 Quality Estimation with k-nearest Neighbors and Automatic Evaluation for Model-specific Quality Estimation
abstract
Providing quality scores along with Machine Translation (MT) output, so-called reference-free Quality Estimation (QE), is crucial to inform users about the reliability of the translation. We propose a model-specific, unsupervised QE approach, termed kNN-QE, that extracts information from the MT model’s training data using k-nearest neighbors. Measuring the performance of model-specific QE is not straightforward, since they provide quality scores on their own MT output, thus cannot be evaluated using benchmark QE test sets containing human quality scores on premade MT output. Therefore, we propose an automatic evaluation method that uses quality scores from reference-based metrics as gold standard instead of human-generated ones. We are the first to conduct detailed analyses and conclude that this automatic method is sufficient, and the reference-based MetricX-23 is best for the task.
Tu Anh Dinh, Tobias Palzer, Jan Niehues
EAMT (1)3
2024 SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
abstract
Tu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Danni Liu, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao, Fabian Peller-Konrad, Tobias Röddiger, Alexander Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Tu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao 0002, Fabian Tërnava, Tobias Röddiger, Alex Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues
EMNLP18
2024 Optimizing Rare Word Accuracy in Direct Speech Translation with a Retrieval-and-Demonstration Approach
abstract
Direct speech translation (ST) models often struggle with rare words.Incorrect translation of these words can have severe consequences, impacting translation quality and user trust.While rare word translation is inherently challenging for neural models due to sparse learning signals, real-world scenarios often allow access to translations of past recordings on similar topics.To leverage these valuable resources, we propose a retrieval-and-demonstration approach to enhance rare word translation accuracy in direct ST models.First, we adapt existing ST models to incorporate retrieved examples for rare word translation, which allows the model to benefit from prepended examples, similar to in-context learning.We then develop a cross-modal (speech-to-speech, speechto-text, text-to-text) retriever to locate suitable examples.We demonstrate that standard ST models can be effectively adapted to leverage examples for rare word translation, improving rare word translation accuracy over the baseline by 17.6% with gold examples and 8.5% with retrieved examples.Moreover, our speechto-speech retrieval approach outperforms other modalities and exhibits higher robustness to unseen speakers.Our code is publicly available 1 . Transformer Encoder (ST)
Jan Niehues
EMNLP3
2024 Contextual Refinement of Translations: Large Language Models for Sentence and Document-Level Post-Editing
abstract
Sai Koneru, Miriam Exel, Matthias Huck, Jan Niehues. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Sai Koneru, Miriam Exel, Matthias Huck, Jan Niehues
NAACL-HLT4
2024 Augmenting Automatic Speech Recognition Models With Disfluency Detection
abstract
Speech disfluency commonly occurs in conversational and spontaneous speech. However, standard Automatic Speech Recognition (ASR) models struggle to accurately recognize these disfluencies because they are typically trained on fluent transcripts. Current research mainly focuses on detecting disfluencies within transcripts, overlooking their exact location and duration in the speech. Additionally, previous work often requires model fine-tuning and addresses limited types of disfluencies. In this work, we present an inference-only approach to augment any ASR model with the ability to detect open-set disfluencies. We first demonstrate that ASR models have difficulty transcribing speech disfluencies. Next, this work proposes a modified Connectionist Temporal Classification(CTC)based forced alignment algorithm from [1] to predict wordlevel timestamps while effectively capturing disfluent speech. Additionally, we develop a model to classify alignment gaps between timestamps as either containing disfluent speech or silence. This model achieves an accuracy of 81.62% and an F1-score of 80.07%. We test the augmentation pipeline of alignment gap detection and classification on a disfluent dataset. Our results show that we captured 74.13% of the words that were initially missed by the transcription, demonstrating the potential of this pipeline for downstream tasks.
Robin Amann, Zhaolin Li, Barbara Bruno, Jan Niehues
SLT4
2024 Identifying the Information Gap for Visually Impaired Students during Lecture Talks
abstract
Visual slides have become a common tool in university classrooms, assisting students in following along with lectures. However, visually impaired individuals face an information gap when visual slides are the primary lecture material. This paper aims to identify which information is most likely to be withheld from visually impaired individuals. To analyse this information gap, the study first extracts textual elements from lecture slides and converts visual elements into textual descriptions on multiple levels that discriminate between different semantic depths of visuals. This information is compared to the transcribed audio of the lecture talk by a semantic similarity analysis. The study tests how a transformer model can be used for an automatic semantic similarity analysis, which is validated through human evaluation. The results indicate variations in the level of detail and semantic depth with which visuals are described across different lectures. As shown, direct references to visuals and elemental details of visuals are often omitted in speech. However, this information is important for enabling mental visualization processes for visually impaired individuals and for following along with the lecture. The identified information gap can serve as valuable feedback for lecturers and assistive systems, helping them to more effectively provide visually impaired students with the information that they may have otherwise missed during the lecture.
Katharina Anderer, Matthias Wölfel, Jan Niehues
VL/HCC3
2023 Analyzing Challenges in Neural Machine Translation for Software Localization
abstract
Advancements in Neural Machine Translation (NMT) greatly benefit the software localization industry by decreasing the post-editing time of human annotators.Although the volume of the software being localized is growing significantly, techniques for improving NMT for user interface (UI) texts are lacking.These UI texts have different properties than other collections of texts, presenting unique challenges for NMT.For example, they are often very short, causing them to be ambiguous and needing additional context (button, title text, a table item, etc.) for disambiguation.However, no such UI data sets are readily available with contextual information for NMT models to exploit.This work aims to provide a first step in improving UI translations and highlight its challenges.To achieve this, we provide a novel multilingual UI corpus collection (∼ 1.3M for English ↔ German) with a targeted test set and analyze the limitations of state-of-the-art methods on this challenging task.Specifically, we present a targeted test set for disambiguation from English to German to evaluate reliably and emphasize UI translation challenges.Furthermore, we evaluate several state-of-the-art NMT techniques from domain adaptation and document-level NMT on this challenging task.All the scripts to replicate the experiments and data sets are available here.1,2
Sai Koneru, Matthias Huck, Miriam Exel, Jan Niehues
EACL4
2023 Towards continually learning new languages
Ngoc-Quan Pham, Jan Niehues, Alex Waibel
INTERSPEECH2
2023 Perturbation-based QE: An Explainable, Unsupervised Word-level Quality Estimation Method for Blackbox Machine Translation
abstract
Quality Estimation (QE) is the task of predicting the quality of Machine Translation (MT) system output, without using any gold-standard translation references. State-of-the-art QE models are supervised: they require human-labeled quality of some MT system output on some datasets for training, making them domain-dependent and MT-system-dependent. There has been research on unsupervised QE, which requires glass-box access to the MT systems, or parallel MT data to generate synthetic errors for training QE models. In this paper, we present Perturbation-based QE - a word-level Quality Estimation approach that works simply by analyzing MT system output on perturbed input source sentences. Our approach is unsupervised, explainable, and can evaluate any type of blackbox MT systems, including the currently prominent large language models (LLMs) with opaque internal processes. For language directions with no labeled QE data, our approach has similar or better performance than the zero-shot supervised approach on the WMT21 shared task. Our approach is better at detecting gender bias and word-sense-disambiguation errors in translation than supervised QE, indicating its robustness to out-of-domain usage. The performance gap is larger when detecting errors on a nontraditional translation-prompting LLM, indicating that our approach is more generalizable to different MT systems. We give examples demonstrating our approach’s explainability power, where it shows which input source words have influence on a certain MT output word.
Tu Anh Dinh, Jan Niehues
MTSummit (1)2
2023 Joint modelling of audio-visual cues using attention mechanisms for emotion recognition
abstract
Abstract Emotions play a crucial role in human-human communications with complex socio-psychological nature. In order to enhance emotion communication in human-computer interaction, this paper studies emotion recognition from audio and visual signals in video clips, utilizing facial expressions and vocal utterances. Thereby, the study aims to exploit temporal information of audio-visual cues and detect their informative time segments. Attention mechanisms are used to exploit the importance of each modality over time. We propose a novel framework that consists of bi-modal time windows spanning short video clips labeled with discrete emotions. The framework employs two networks, with each one being dedicated to one modality. As input to a modality-specific network, we consider a time-dependent signal deriving from the embeddings of the video and audio modalities. We employ the encoder part of the Transformer on the visual embeddings and another one on the audio embeddings. The research in this paper introduces detailed studies and meta-analysis findings, linking the outputs of our proposition to research from psychology. Specifically, it presents a framework to understand underlying principles of emotion recognition as functions of three separate setups in terms of modalities: audio only, video only, and the fusion of audio and video. Experimental results on two datasets show that the proposed framework achieves improved accuracy in emotion recognition, compared to state-of-the-art techniques and baseline methods not using attention mechanisms. The proposed method improves the results over baseline methods by at least 5.4%. Our experiments show that attention mechanisms reduce the gap between the entropies of unimodal predictions, which increases the bimodal predictions’ certainty and, therefore, improves the bimodal recognition rates. Furthermore, evaluations with noisy data in different scenarios are presented during the training and testing processes to check the framework’s consistency and the attention mechanism’s behavior. The results demonstrate that attention mechanisms increase the framework’s robustness when exposed to similar conditions during the training and the testing phases. Finally, we present comprehensive evaluations of emotion recognition as a function of time. The study shows that the middle time segments of a video clip are essential in the case of using audio modality. However, in the case of video modality, the importance of time windows is distributed equally.
Esam Ghaleb, Jan Niehues, Stylianos Asteriadis
Multim. Tools Appl.2
2022 Tackling Data Scarcity in Speech Translation Using Zero-Shot Multilingual Machine Translation Techniques
abstract
Recently, end-to-end speech translation (ST) has gained significant attention as it avoids error propagation. However, the approach suffers from data scarcity. It heavily depends on direct ST data and is less efficient in making use of speech transcription and text translation data, which is often more easily available. In the related field of multilingual text translation, several techniques have been proposed for zero-shot translation. A main idea is to increase the similarity of semantically similar sentences in different languages. We investigate whether these ideas can be applied to speech translation, by building ST models trained on speech transcription and text translation data. We investigate the effects of data augmentation and auxiliary loss function. The techniques were successfully applied to few-shot ST using limited ST data, with improvements of up to +12.9 BLEU points compared to direct end-to-end ST and +3.1 BLEU points compared to ST models fine-tuned from ASR model.
Tu Anh Dinh, Jan Niehues
ICASSP3
2022 TED Talk Teaser Generation with Pre-Trained Models
abstract
While we have seen significant advances in automatic summarization for text, research on speech summarization is still limited. In this work, we address the challenge of automatically generating teasers for TED talks. In the first step, we create a corpus for automatic summarization of TED and TEDx talks consisting of the talks' recording, their transcripts and their descriptions. The corpus is used to build a speech summarization system for the task. We adapt and combine pre-trained models for automatic speech recognition (ASR) and text summarization using the collected data. This initial work shows that is more important to adapt the summarization model to the ASR transcripts than to adapt the ASR model to the talks.
Gianluca Vico, Jan Niehues
ICASSP2
2022 Adaptive multilingual speech recognition with pretrained models
abstract
Multilingual speech recognition with supervised learning has achieved great results as reflected in recent research.With the development of pretraining methods on audio and text data, it is imperative to transfer the knowledge from unsupervised multilingual models to facilitate recognition, especially in many languages with limited data.Our work investigated the effectiveness of using two pretrained models for two modalities: wav2vec 2.0 for audio and MBART50 for text, together with the adaptive weight techniques to massively improve the recognition quality on the public datasets containing CommonVoice and Europarl.Overall, we noticed an 44% improvement over purely supervised learning, and more importantly, each technique provides a different reinforcement in different languages.We also explore other possibilities to potentially obtain the best model by slightly adding either depth or relative attention to the architecture.
Ngoc-Quan Pham, Alex Waibel, Jan Niehues
INTERSPEECH3
2022 LibriS2S: A German-English Speech-to-Speech Translation Corpus
abstract
Recently, we have seen an increasing interest in the area of speech-to-text translation. This has led to astonishing improvements in this area. In contrast, the activities in the area of speech-to-speech translation is still limited, although it is essential to overcome the language barrier. We believe that one of the limiting factors is the availability of appropriate training data. We address this issue by creating LibriS2S, to our knowledge the first publicly available speech-to-speech training corpus between German and English. For this corpus, we used independently created audio for German and English leading to an unbiased pronunciation of the text in both languages. This allows the creation of a new text-to-speech and speech-to-speech translation model that directly learns to generate the speech signal based on the pronunciation of the source language. Using this created corpus, we propose Text-to-Speech models based on the example of the recently proposed FastSpeech 2 model that integrates source language information. We do this by adapting the model to take information such as the pitch, energy or transcript from the source speech as additional input.
Pedro Jeuris, Jan Niehues
LREC2
2021 Improving Zero-Shot Translation by Disentangling Positional Information
abstract
Danni Liu, Jan Niehues, James Cross, Francisco Guzmán, Xian Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jan Niehues, James Cross 0003, Francisco Guzmán, Xian Li 0003
ACL/IJCNLP (1)2
2021 Continuous Learning in Neural Machine Translation using Bilingual Dictionaries
abstract
While recent advances in deep learning led to significant improvements in machine translation, neural machine translation is often still not able to continuously adapt to the environment.For humans, as well as for machine translation, bilingual dictionaries are a promising knowledge source to continuously integrate new knowledge.However, their exploitation poses several challenges: The system needs to be able to perform one-shot learning as well as model the morphology of source and target language.In this work, we proposed an evaluation framework to assess the ability of neural machine translation to continuously learn new phrases.We integrate one-shot learning methods for neural machine translation with different word representations and show that it is important to address both in order to successfully make use of bilingual dictionaries.By addressing both challenges we are able to improve the ability to translate new, rare words and phrases from 30% to up to 70%.The correct lemma is even generated by more than 90%.
Jan Niehues
EACL1
2020 Improving Sequence-To-Sequence Speech Recognition Training with On-The-Fly Data Augmentation
abstract
Sequence-to-Sequence (S2S) models recently started to show state-of-the-art performance for automatic speech recognition (ASR). With these large and deep models overfitting remains the largest problem, outweighing performance improvements that can be obtained from better architectures. One solution to the overfitting problem is increasing the amount of available training data and the variety exhibited by the training data with the help of data augmentation. In this paper we examine the influence of three data augmentation methods on the performance of two S2S model architectures. One of the data augmentation method comes from literature, while two other methods are our own development - a time perturbation in the frequency domain and sub-sequence sampling. Our experiments on Switchboard and Fisher data show state-of-the-art performance for S2S models that are trained solely on the speech training data and do not use additional text data.
Thai Son Nguyen, Sebastian Stüker, Jan Niehues, Alex Waibel
ICASSP3
2020 Multimodal Attention-Mechanism For Temporal Emotion Recognition
abstract
Exploiting the multimodal and temporal interaction between audio-visual channels is essential for automatic audio-video emotion recognition (AVER). Modalities' strength in emotions and time-window of a video-clip could be further utilized through a weighting scheme such as attention mechanism to capture their complementary information. The attention mechanism is a powerful approach for sequence modeling, which can be employed to fuse audio-video cues overtime. We propose a novel framework which consists of biaudio-visual time-windows that span short video-clips labeled with discrete emotions. Attention is used to weigh these time windows for multimodal learning and fusion. Experimental results on two datasets show that the proposed methodology can achieve an enhanced multimodal emotion recognition.
Esam Ghaleb, Jan Niehues, Stylianos Asteriadis
ICIP2
2020 Low-Latency Sequence-to-Sequence Speech Recognition and Translation by Partial Hypothesis Selection
abstract
Encoder-decoder models provide a generic architecture for sequence-to-sequence tasks such as speech recognition and translation. While offline systems are often evaluated on quality metrics like word error rates (WER) and BLEU, latency is also a crucial factor in many practical use-cases. We propose three latency reduction techniques for chunk-based incremental inference and evaluate their efficiency in terms of accuracy-latency trade-off. On the 300-hour How2 dataset, we reduce latency by 83% to 0.8 second by sacrificing 1% WER (6% rel.) compared to offline transcription. Although our experiments use the Transformer, the hypothesis selection strategies are applicable to other encoder-decoder models. To avoid expensive re-computation, we use a unidirectionally-attending encoder. After an adaptation procedure to partial sequences, the unidirectional model performs on-par with the original model. We further show that our approach is also applicable to low-latency speech translation. On How2 English-Portuguese speech translation, we reduce latency to 0.7 second (-84% rel.) while incurring a loss of 2.4 BLEU points (5% rel.) compared to the offline system.
Gerasimos Spanakis, Jan Niehues
INTERSPEECH3
2020 Relative Positional Encoding for Speech Recognition and Direct Translation
abstract
Transformer models are powerful sequence-to-sequence architectures that are capable of directly mapping speech inputs to transcriptions or translations. However, the mechanism for modeling positions in this model was tailored for text modeling, and thus is less ideal for acoustic inputs. In this work, we adapt the relative position encoding scheme to the Speech Transformer, where the key addition is relative distance between input states in the self-attention network. As a result, the network can better adapt to the variable distributions present in speech data. Our experiments show that our resulting model achieves the best recognition result on the Switchboard benchmark in the non-augmentation condition, and the best published result in the MuST-C speech translation benchmark. We also show that this model is able to better utilize synthetic data than the Transformer, and adapts better to variable sentence segmentation quality for speech translation.
Ngoc-Quan Pham, Thanh-Le Ha, Tuan-Nam Nguyen, Thai Son Nguyen, Elizabeth Salesky, Sebastian Stüker, Jan Niehues, Alex Waibel
INTERSPEECH7
2020 Optimizing segmentation granularity for neural machine translation
Elizabeth Salesky, Andrew Runge, Alex Coda, Jan Niehues, Graham Neubig
Mach. Transl.4
2019 Modeling Confidence in Sequence-to-Sequence Models
abstract
Recently, significant improvements have been achieved in various natural language processing tasks using neural sequence-to-sequence models.While aiming for the best generation quality is important, ultimately it is also necessary to develop models that can assess the quality of their output.In this work, we propose to use the similarity between training and test conditions as a measure for models' confidence.We investigate methods solely using the similarity as well as methods combining it with the posterior probability.While traditionally only target tokens are annotated with confidence measures, we also investigate methods to annotate source tokens with confidence.By learning an internal alignment model, we can significantly improve confidence projection over using stateof-the-art external alignment tools.We evaluate the proposed methods on downstream confidence estimation for machine translation (MT).We show improvements on segmentlevel confidence estimation as well as on confidence estimation for source tokens.In addition, we show that the same methods can also be applied to other tasks using sequence-tosequence models.On the automatic speech recognition (ASR) task, we are able to find 60% of the errors by looking at 20% of the data.
Jan Niehues, Ngoc-Quan Pham
INLG1
2019 Survey Talk: A Survey on Speech Translation
Jan Niehues
INTERSPEECH1
2019 Very Deep Self-Attention Networks for End-to-End Speech Recognition
abstract
Recently, end-to-end sequence-to-sequence models for speech recognition have gained significant interest in the research community. While previous architecture choices revolve around time-delay neural networks (TDNN) and long short-term memory (LSTM) recurrent neural networks, we propose to use self-attention via the Transformer architecture as an alternative. Our analysis shows that deep Transformer networks with high learning capacity are able to exceed performance from previous end-to-end approaches and even match the conventional hybrid systems. Moreover, we trained very deep models with up to 48 Transformer layers for both encoder and decoders combined with stochastic residual connections, which greatly improve generalizability and training efficiency. The resulting models outperform all previous end-to-end ASR approaches on the Switchboard benchmark. An ensemble of these models achieve 9.9% and 17.7% WER on Switchboard and CallHome test sets respectively. This finding brings our end-to-end models to competitive levels with previous hybrid systems. Further, with model ensembling the Transformers can outperform certain hybrid systems, which are more complicated in terms of both structure and training procedure.
Ngoc-Quan Pham, Thai Son Nguyen, Jan Niehues, Markus Müller 0001, Alex Waibel
INTERSPEECH3
2019 Attention-Passing Models for Robust and Data-Efficient End-to-End Speech Translation
abstract
Speech translation has traditionally been approached through cascaded models consisting of a speech recognizer trained on a corpus of transcribed speech, and a machine translation system trained on parallel texts. Several recent works have shown the feasibility of collapsing the cascade into a single, direct model that can be trained in an end-to-end fashion on a corpus of translated speech. However, experiments are inconclusive on whether the cascade or the direct model is stronger, and have only been conducted under the unrealistic assumption that both are trained on equal amounts of data, ignoring other available speech recognition and machine translation corpora. In this paper, we demonstrate that direct speech translation models require more data to perform well than cascaded models, and although they allow including auxiliary data through multi-task training, they are poor at exploiting such data, putting them at a severe disadvantage. As a remedy, we propose the use of end- to-end trainable models with two attention mechanisms, the first establishing source speech to source text alignments, the second modeling source to target text alignment. We show that such models naturally decompose into multi-task–trainable recognition and translation tasks and propose an attention-passing technique that alleviates error propagation issues in a previous formulation of a model with two attention stages. Our proposed model outperforms all examined baselines and is able to exploit auxiliary training data much more effectively than direct attentional models.
Matthias Sperber, Graham Neubig, Jan Niehues, Alex Waibel
Trans. Assoc. Comput. Linguistics3
2018 Term Extraction via Neural Sequence Labeling a Comparative Evaluation of Strategies Using Recurrent Neural Networks
Maren Kucza, Jan Niehues, Thomas Zenkel, Alex Waibel, Sebastian Stüker
INTERSPEECH2
2018 Low-Latency Neural Speech Translation
abstract
Through the development of neural machine translation, the quality of machine translation systems has been improved significantly.By exploiting advancements in deep learning, systems are now able to better approximate the complex mapping from source sentences to target sentences.But with this ability, new challenges also arise.An example is the translation of partial sentences in low-latency speech translation.Since the model has only seen complete sentences in training, it will always try to generate a complete sentence, though the input may only be a partial sentence.We show that NMT systems can be adapted to scenarios where no task-specific training data is available.Furthermore, this is possible without losing performance on the original training data.We achieve this by creating artificial data and by using multi-task learning.After adaptation, we are able to reduce the number of corrections displayed during incremental output construction by 45%, without a decrease in translation quality.
Jan Niehues, Ngoc-Quan Pham, Thanh-Le Ha, Matthias Sperber, Alex Waibel
INTERSPEECH1
2018 Self-Attentional Acoustic Models
abstract
Self-attention is a method of encoding sequences of vectors by relating these vectors to each-other based on pairwise similarities. These models have recently shown promising results for modeling discrete sequences, but they are non-trivial to apply to acoustic modeling due to computational and modeling issues. In this paper, we apply self-attention to acoustic modeling, proposing several improvements to mitigate these issues: First, self-attention memory grows quadratically in the sequence length, which we address through a downsampling technique. Second, we find that previous approaches to incorporate position information into the model are unsuitable and explore other representations and hybrid models to this end. Third, to stress the importance of local context in the acoustic signal, we propose a Gaussian biasing approach that allows explicit control over the context range. Experiments find that our model approaches a strong baseline based on LSTMs with network-in-network connections while being much faster to compute. Besides speed, we find that interpretability is a strength of self-attentional acoustic models, and demonstrate that self-attention heads learn a linguistically plausible division of labor.
Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stüker, Alex Waibel
INTERSPEECH2
2018 KIT-Multi: A Translation-Oriented Multilingual Embedding Corpus
Thanh-Le Ha, Jan Niehues, Matthias Sperber, Ngoc-Quan Pham, Alex Waibel
LREC2
2018 Automated Evaluation of Out-of-Context Errors
Patrick Huber, Jan Niehues, Alex Waibel
LREC2
2018 Towards Fluent Translations From Disfluent Speech
abstract
When translating from speech, special consideration for conversational speech phenomena such as disfluencies is necessary. Most machine translation training data consists of well-formed written texts, causing issues when translating spontaneous speech. Previous work has introduced an intermediate step between speech recognition (ASR) and machine translation (MT) to remove disfluencies, making the data better-matched to typical translation text and significantly improving performance. However, with the rise of end-to-end speech translation systems, this intermediate step must be incorporated into the sequence-to-sequence architecture. Further, though translated speech datasets exist, they are typically news or rehearsed speech without many disfluencies (e.g. TED), or the disfluencies are translated into the references (e.g. Fisher). To generate clean translations from disfluent speech, cleaned references are necessary for evaluation. We introduce a corpus of cleaned target data for the Fisher Spanish-English dataset for this task. We compare how different architectures handle disfluencies and provide a baseline for removing disfluencies in end-to-end translation.
Elizabeth Salesky, Susanne Burger, Jan Niehues, Alex Waibel
SLT3
2017 Neural Lattice-to-Sequence Models for Uncertain Inputs
abstract
The input to a neural sequence-tosequence model is often determined by an up-stream system, e.g. a word segmenter, part of speech tagger, or speech recognizer.These up-stream models are potentially error-prone.Representing inputs through word lattices allows making this uncertainty explicit by capturing alternative sequences and their posterior probabilities in a compact form.In this work, we extend the TreeLSTM (Tai et al., 2015) into a LatticeLSTM that is able to consume word lattices, and can be used as encoder in an attentional encoderdecoder model.We integrate lattice posterior scores into this architecture by extending the TreeLSTM's child-sum and forget gates and introducing a bias term into the attention mechanism.We experiment with speech translation lattices and report consistent improvements over baselines that translate either the 1-best hypothesis or the lattice without posterior scores.
Matthias Sperber, Graham Neubig, Jan Niehues, Alex Waibel
EMNLP3
2017 NMT-Based Segmentation and Punctuation Insertion for Real-Time Spoken Language Translation
Eunah Cho, Jan Niehues, Alex Waibel
INTERSPEECH2
2017 Comparison of Decoding Strategies for CTC Acoustic Models
abstract
Connectionist Temporal Classification has recently attracted a lot of interest as it offers an elegant approach to building acoustic models (AMs) for speech recognition. The CTC loss function maps an input sequence of observable feature vectors to an output sequence of symbols. Output symbols are conditionally independent of each other under CTC loss, so a language model (LM) can be incorporated conveniently during decoding, retaining the traditional separation of acoustic and linguistic components in ASR. For fixed vocabularies, Weighted Finite State Transducers provide a strong baseline for efficient integration of CTC AMs with n-gram LMs. Character-based neural LMs provide a straight forward solution for open vocabulary speech recognition and all-neural models, and can be decoded with beam search. Finally, sequence-to-sequence models can be used to translate a sequence of individual sounds into a word string. We compare the performance of these three approaches, and analyze their error patterns, which provides insightful guidance for future research and development in this important area.
Thomas Zenkel, Ramon Sanabria, Florian Metze, Jan Niehues, Matthias Sperber, Sebastian Stüker, Alex Waibel
INTERSPEECH4
2017 Transcribing against time
Matthias Sperber, Graham Neubig, Jan Niehues, Satoshi Nakamura 0001, Alex Waibel
Speech Commun.3
2016 Pre-Translation for Neural Machine Translation
abstract
Recently, the development of neural machine translation (NMT) has significantly improved the translation quality of automatic machine translation. While most sentences are more accurate and fluent than translations by statistical machine translation (SMT)-based systems, in some cases, the NMT system produces translations that have a completely different meaning. This is especially the case when rare words occur. When using statistical machine translation, it has already been shown that significant gains can be achieved by simplifying the input in a preprocessing step. A commonly used example is the pre-reordering approach. In this work, we used phrase-based machine translation to pre-translate the input into the target language. Then a neural machine translation system generates the final hypothesis using the pre-translation. Thereby, we use either only the output of the phrase-based machine translation (PBMT) system or a combination of the PBMT output and the source sentence. We evaluate the technique on the English to German translation task. Using this approach we are able to outperform the PBMT system as well as the baseline neural MT system by up to 2 BLEU points. We analyzed the influence of the quality of the initial system on the final result.
Jan Niehues, Eunah Cho, Thanh-Le Ha, Alex Waibel
COLING1
2016 Lightly Supervised Quality Estimation
abstract
Evaluating the quality of output from language processing systems such as machine translation or speech recognition is an essential step in ensuring that they are sufficient for practical use. However, depending on the practical requirements, evaluation approaches can differ strongly. Often, reference-based evaluation measures (such as BLEU or WER) are appealing because they are cheap and allow rapid quantitative comparison. On the other hand, practitioners often focus on manual evaluation because they must deal with frequently changing domains and quality standards requested by customers, for which reference-based evaluation is insufficient or not possible due to missing in-domain reference data (Harris et al., 2016). In this paper, we attempt to bridge this gap by proposing a framework for lightly supervised quality estimation. We collect manually annotated scores for a small number of segments in a test corpus or document, and combine them with automatically predicted quality scores for the remaining segments to predict an overall quality estimate. An evaluation shows that our framework estimates quality more reliably than using fully automatic quality estimation approaches, while keeping annotation effort low by not requiring full references to be available for the particular domain.
Matthias Sperber, Graham Neubig, Jan Niehues, Sebastian Stüker, Alex Waibel
COLING3
2016 Dynamic Transcription for Low-Latency Speech Translation
Jan Niehues, Thai Son Nguyen, Eunah Cho, Thanh-Le Ha, Kevin Kilgour, Markus Müller 0001, Matthias Sperber, Sebastian Stüker, Alex Waibel
INTERSPEECH1
2015 Stripping Adjectives: Integration Techniques for Selective Stemming in SMT Systems
Isabel Slawik, Jan Niehues, Alex Waibel
EAMT2
2015 Combination of NN and CRF models for joint detection of punctuation and disfluencies
abstract
Inserting proper punctuation marks and deleting speech disfluencies are two of the most essential tasks in spoken language processing. This challenging task has prompted extensive research using various techniques, such as conditional random fields. Neural networks, however, are relatively under-explored for this task. Combining different modeling techniques with different advantages has the potential to lead to improvements. In this work, we first establish the performance of joint modeling of punctuation prediction and disfluency detection using neural networks. We then combine a conditional random fields based model and a neural networks based model log-linearly, and show that the combined approach outperforms both individual models, by 2.7% and 3.5% in F-score for speech disfluency and punctuation detection, respectively. When used as a preprocessing step to machine translation this also results in an improved translation quality of 2.5 BLEU points compared to the baseline and of 0.6 BLEU points compared to the non-combined model. Index Terms: speech disfluency detection, punctuation insertion, speech translation
Eunah Cho, Kevin Kilgour, Jan Niehues, Alex Waibel
INTERSPEECH3
2014 Tight Integration of Speech Disfluency Removal into SMT
abstract
Speech disfluencies are one of the main challenges of spoken language processing.Conventional disfluency detection systems deploy a hard decision, which can have a negative influence on subsequent applications such as machine translation.In this paper we suggest a novel approach in which disfluency detection is integrated into the translation process.We train a CRF model to obtain a disfluency probability for each word.The SMT decoder will then skip the potentially disfluent word based on its disfluency probability.Using the suggested scheme, the translation score of both the manual transcript and ASR output is improved by around 0.35 BLEU points compared to the CRF hard decision system.
Eunah Cho, Jan Niehues, Alex Waibel
EACL2
2014 Manual Analysis of Structurally Informed Reordering in German-English Machine Translation
Teresa Herrmann, Jan Niehues, Alex Waibel
LREC2
2013 A real-world system for simultaneous translation of German lectures
abstract
We present a real-time automatic speech translation system for university lectures that can interpret several lectures in parallel. University lectures are characterized by a multitude of diverse topics and a large amount of technical terms. This poses specific challenges, e.g., a very specific vocabulary and language model are needed. In addition, in order to be able to translate simultaneously, i.e., to interpret the lectures, the components of the systems need special modifications. The output of the system is delivered in the form or realtime subtitles via a web site that can be accessed by the students attending the lecture through mobile phones, tablet computers or laptops. We evaluated the system on our German to English lecture translation task at the Karlsruhe Institute of Technology. The system is now being installed in several lecture halls at KIT and is able to provide the translation to the students in several parallel sessions.
Eunah Cho, Christian Fügen, Teresa Herrmann, Kevin Kilgour, Mohammed Mediani, Christian Mohr, Jan Niehues, Kay Rottmann, Christian Saam, Sebastian Stüker, Alex Waibel
INTERSPEECH7
2012 The IWSLT 2011 Evaluation Campaign on Automatic Talk Translation
Marcello Federico, Sebastian Stüker, Luisa Bentivogli, Michael Paul, Mauro Cettolo, Teresa Herrmann, Jan Niehues, Giovanni Moretti
LREC7
2010 Domain Adaptation in Statistical Machine Translation using Factored Translation Models
Jan Niehues, Alex Waibel
EAMT1
2008 Simultaneous machine translation of german lectures into english: Investigating research challenges for the future
abstract
An increasingly globalized world fosters the exchange of students, researchers or employees. As a result, situations in which people of different native tongues are listening to the same lecture become more and more frequent. In many such situations, human interpreters are prohibitively expensive or simply not available. For this reason, and because first prototypes have already demonstrated the feasibility of such systems, automatic translation of lectures receives increasing attention. A large vocabulary and strong variations in speaking style make lecture translation a challenging, however not hopeless, task. The scope of this paper is to investigate a variety of challenges and to highlight possible solutions in building a system for simultaneous translation of lectures from German to English. While some of the investigated challenges are more general, e.g. environment robustness, other challenges are more specific for this particular task, e.g. pronunciation of foreign words or sentence segmentation. We also report our progress in building an end-to-end system and analyze its performance in terms of objective and subjective measures.
Matthias Wölfel, Muntsin Kolss, Florian Kraft, Jan Niehues, Matthias Paulik, Alex Waibel
SLT4