EDBT 2026 Demo / reviewers in the wild / expert
Gabriel Mittag
dblp:199/9540
· DBLP profile ↗
26ranked-venue papers
12as first author
11since 2021 · last 2026
0009-0005-2129-2414ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 12 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 2 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human-in-the-Loop Bandwidth Estimation for Quality of Experience Optimization in Real-Time Video CommunicationabstractThe quality of experience (QoE) delivered by video conferencing systems is significantly influenced by accurately estimating the time-varying available bandwidth between the sender and receiver. Bandwidth estimation for real-time communications remains an open challenge due to rapidly evolving network architectures, increasingly complex protocol stacks, and the difficulty of defining QoE metrics that reliably improve user experience. In this work, we propose a deployed, human-in-the-loop, data-driven framework for bandwidth estimation to address these challenges. Our approach begins with training objective QoE reward models derived from subjective user evaluations to measure audio and video quality in real-time video conferencing systems. Subsequently, we collect roughly 1M network traces with objective QoE rewards from real-world Microsoft Teams calls to curate a bandwidth estimation training dataset. We then introduce a novel distributional offline reinforcement learning (RL) algorithm to train a neural-network-based bandwidth estimator aimed at improving QoE for users. Our real-world A/B test demonstrates that the proposed approach reduces the subjective poor call ratio by 11.41% compared to the baseline bandwidth estimator. Furthermore, the proposed offline RL algorithm is benchmarked on D4RL tasks to demonstrate its generalization beyond bandwidth estimation. Sami Khairy, Gabriel Mittag, Vishak Gopal, Ross Cutler |
AAAI | 2 |
| 2026 | Offline Meta-learning for Real-time Bandwidth Estimation
Aashish Gottipati, Sami Khairy, Yasaman Hosseinkashi, Gabriel Mittag, Vishak Gopal, Francis Y. Yan, Ross Cutler |
ICC | 4 |
| 2025 | Offline to Online Learning for Real-Time Bandwidth EstimationabstractReal-time video applications require accurate bandwidth estimation (BWE) to maintain user experience across varying network conditions. However, increasing network heterogeneity challenges general-purpose BWE algorithms, necessitating solutions that adapt to end-user environments. While widely adopted, heuristic-based methods are difficult to individualize without extensive domain expertise. Conversely, online reinforcement learning (RL) offers ease of customization but neglects prior domain expertise and suffers from sample inefficiency. Thus, we present Merlin, an imitation learning-based solution that replaces the manual parameter tuning of heuristic-based methods with data-driven updates to streamline end-user personalization. Our key insight is that transforming heuristic-based BWE algorithms into neural networks facilitates data-driven personalization. Merlin utilizes Behavioral Cloning to efficiently learn from offline telemetry logs, capturing heuristic policies without live network interactions. The cloned policy can then be seamlessly tailored to end user network conditions through online finetuning. In real intercontinental videoconferencing calls, Merlin matches our heuristic's policy with no statistically significant differences in user quality of experience (QoE). Finetuning Merlin's control policy to end-user environments enables QoE improvements of up to 7.8 % compared to the heuristic policy. Lastly, our IL-based design performs competitively with current state-of-the-art online RL techniques but converges with 80 % fewer videoconferencing samples, facilitating practical end-user personalization. Aashish Gottipati, Sami Khairy, Gabriel Mittag, Vishak Gopal, Ross Cutler |
ICC | 3 |
| 2024 | ACM MMSys 2024 Bandwidth Estimation in Real Time Communications ChallengeabstractThe quality of experience (QoE) delivered by video conferencing systems to end users depends in part on correctly estimating the capacity of the bottleneck link between the sender and the receiver over time. Bandwidth estimation for real-time communications (RTC) remains a significant challenge, primarily due to the continuously evolving heterogeneous network architectures and technologies. From the first bandwidth estimation challenge which was hosted at ACM MMSys 2021, we learned that bandwidth estimation models trained with reinforcement learning (RL) in simulations to maximize network-based reward functions may not be optimal in reality due to the sim-to-real gap and the difficulty of aligning network-based rewards with user-perceived QoE. This grand challenge aims to advance bandwidth estimation model design by aligning reward maximization with user-perceived QoE optimization using offline RL and a real-world dataset with objective rewards which have high correlations with subjective audio/video quality in Microsoft Teams. All models submitted to the grand challenge underwent initial evaluation on our emulation platform. For a comprehensive evaluation under diverse network conditions with temporal fluctuations, top models were further evaluated on our geographically distributed testbed by using each model to conduct 600 calls within a 12-day period. The winning model is shown to deliver comparable performance to the top behavior policy in the released dataset. By leveraging real-world data and integrating objective audio/video quality scores as rewards, offline RL can therefore facilitate the development of competitive bandwidth estimators for RTC. Sami Khairy, Gabriel Mittag, Vishak Gopal, Francis Y. Yan, Zhixiong Niu, Ezra Ameri, Scott Inglis, Mehrsa Golestaneh, Ross Cutler |
MMSys | 2 |
| 2023 | LSTM-Based Video Quality Prediction Accounting for Temporal Distortions in Videoconferencing CallsabstractCurrent state-of-the-art video quality models, such as VMAF, give excellent prediction results by comparing the degraded video with its reference video. However, they do not consider temporal distortions (e.g., frame freezes or skips) that occur during videoconferencing calls. In this paper, we present a data-driven approach for modeling such distortions automatically by training an LSTM with subjective quality ratings labeled via crowdsourcing. The videos were collected from live videoconferencing calls in 83 different network conditions. We applied QR codes as markers on the source videos to create aligned references and compute temporal features based on the alignment vectors. Using these features together with VMAF core features, our proposed model achieves a PCC of 0.99 on the validation set. Furthermore, our model outputs per-frame quality that gives detailed insight into the cause of video quality impairments. The VCM model and dataset are open-sourced at https://github.com/microsoft/Video_Call_MOS. Gabriel Mittag, Babak Naderi, Vishak Gopal, Ross Cutler |
ICASSP | 1 |
| 2022 | ConferencingSpeech 2022 Challenge: Non-intrusive Objective Speech Quality Assessment (NISQA) Challenge for Online Conferencing ApplicationsabstractWith the advances in speech communication systems such as online conferencing applications, we can seamlessly work with people regardless of where they are. However, during online meetings, speech quality can be significantly affected by background noise, reverberation, packet loss, network jitter, etc. Because of its nature, speech quality is traditionally assessed in subjective tests in laboratories and lately also in crowdsourcing following the international standards from ITU-T Rec. P.800 series. However, those approaches are costly and cannot be applied to customer data. Therefore, an effective objective assessment approach is needed to evaluate or monitor the speech quality of the ongoing conversation. The ConferencingSpeech 2022 challenge targets the non-intrusive deep neural network models for the speech quality assessment task. We open-sourced a training corpus with more than 86K speech clips in different languages, with a wide range of synthesized and live degradations and their corresponding subjective quality scores through crowdsourcing. 18 teams submitted their models for evaluation in this challenge. The blind test sets included about 4300 clips from wide ranges of degradations. This paper describes the challenge, the datasets, and the evaluation methods and reports the final results. Gaoxiong Yi, Babak Naderi, Sebastian Möller 0001, Wafaa Wardah, Gabriel Mittag, Ross Cutler, Zhuohuang Zhang, Donald S. Williamson, Fei Chen 0011, Shidong Shang |
INTERSPEECH | 7 |
| 2021 | Effect of Language Proficiency on Subjective Evaluation of Noise Suppression AlgorithmsabstractSpeech communication systems based on Voice-over-IP technology are frequently used by native as well as non-native speakers of a target language, e.g. in international phone calls or telemeetings. Frequently, such calls also occur in a noisy environment, making noise suppression modules necessary to increase perceived quality of experience. Whereas standard tests for assessing perceived quality make use of native listeners, we assume that noise-reduced speech and residual noise may affect native and non-native listeners of a target language in different ways. To test this assumption, we report results of two subjective tests conducted with English and German native listeners who judge the quality of speech samples recorded by native English, German, and Mandarin speakers, which are degraded with different background noise levels and noise suppression effects. The experiments were conducted following the standardized ITU-T Rec. P.835 approach, however implemented in a crowdsourcing setting according to ITU-T Rec. P.808. Our results show a significant influence of language on speech signal ratings and, consequently, on the overall perceived quality in specific conditions. Babak Naderi, Gabriel Mittag, Rafael Zequeira Jiménez, Sebastian Möller 0001 |
ICASSP | 2 |
| 2021 | Extending the Fullband E-Model Towards Background Noise, Bursty Packet Loss, and Conversational Degradations
Thilo Michael, Gabriel Mittag, Andreas Bütow, Sebastian Möller 0001 |
Interspeech | 2 |
| 2021 | NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced DatasetsabstractIn this paper, we present an update to the NISQA speech quality prediction model that is focused on distortions that occur in communication networks. In contrast to the previous version, the model is trained end-to-end and the time-dependency modelling and time-pooling is achieved through a Self-Attention mechanism. Besides overall speech quality, the model also predicts the four speech quality dimensions Noisiness, Coloration, Discontinuity, and Loudness, and in this way gives more insight into the cause of a quality degradation. Furthermore, new datasets with over 13,000 speech files were created for training and validation of the model. The model was finally tested on a new, live-talking test dataset that contains recordings of real telephone calls. Overall, NISQA was trained and evaluated on 81 datasets from different sources and showed to provide reliable predictions also for unknown speech samples. The code, model weights, and datasets are open-sourced. Gabriel Mittag, Babak Naderi, Assmaa Chehadi, Sebastian Möller 0001 |
Interspeech | 1 |
| 2021 | Removing the Bias in Speech Quality Scores Collected in Noisy Crowdsourcing EnvironmentsabstractSubjective speech quality scores needed to train models for the automatic evaluation of telecommunication systems have generally been collected by conducting demanding laboratory tests. Alternatively, crowdsourcing has emerged as a valid method to address user-centered studies to a large pool of users over the Internet. However, crowdsourcing users often do not follow the instructions and may execute the assigned task in noisy environments. The validity of the data collected in a disturbed environment is questionable, especially in speech quality assessment and other audio-related tasks. This work investigates the applicability of different ensemble-based and non-linear models to correct the bias found in speech quality ratings given to a German speech dataset in noisy crowdsourcing environments. Such a model would help to avoid throwing away quality scores that could be used. The model was trained with data collected in a speech quality assessment study conducted in a simulated crowdsourcing environment in the laboratory. Two groups of listeners rated the quality of speech stimuli in the presence of environmental noise at different levels. The noise under test was street traffic, and the levels ranged from 36dBA to 65. 5dBA. A fine-tuned gradient boosting regressor yielded the best results with a R2score of 0.90 and RMSE of 0.416. Rafael Zequeira Jiménez, Gabriel Mittag, Sebastian Möller 0001 |
QoMEX | 2 |
| 2021 | Bias-Aware Loss for Training Image and Speech Quality Prediction Models from Multiple DatasetsabstractThe ground truth used for training image, video, or speech quality prediction models is based on the Mean Opinion Scores (MOS) obtained from subjective experiments. Usually, it is necessary to conduct multiple experiments, mostly with different test participants, to obtain enough data to train quality models based on machine learning. Each of these experiments is subject to an experiment-specific bias, where the rating of the same file may be substantially different in two experiments (e.g. depending on the overall quality distribution). These different ratings for the same distortion levels confuse neural networks during training and lead to lower performance. To overcome this problem, we propose a bias-aware loss function that estimates each dataset's biases during training with a linear function and considers it while optimising the network weights. We prove the efficiency of the proposed method by training and validating quality prediction models on synthetic and subjective image and speech quality datasets. Gabriel Mittag, Saman Zad Tootaghaj, Thilo Michael, Babak Naderi, Sebastian Möller 0001 |
QoMEX | 1 |
| 2020 | Full-Reference Speech Quality Estimation with Attentional Siamese Neural NetworksabstractIn this paper, we present a full-reference speech quality prediction model with a deep learning approach. The model determines a feature representation of the reference and the degraded signal through a Siamese recurrent convolutional network that shares the weights for both signals as input. The resulting features are then used to align the signals with an attention mechanism and are finally combined to estimate the overall speech quality. The proposed network architecture represents a simple solution for the time-alignment problem that occurs for speech signals transmitted through Voice-Over-IP networks and shows how the clean reference signal can be incorporated into speech quality models that are based on end-to-end trained neural networks. Gabriel Mittag, Sebastian Möller 0001 |
ICASSP | 1 |
| 2020 | Non-Intrusive Diagnostic Monitoring of Fullband Speech Quality
Sebastian Möller 0001, Tobias Hübschen, Thilo Michael, Gabriel Mittag, Gerhard Schmidt |
INTERSPEECH | 4 |
| 2020 | Deep Learning Based Assessment of Synthetic Speech NaturalnessabstractIn this paper, we present a new objective prediction model for synthetic speech naturalness. It can be used to evaluate Text-To-Speech or Voice Conversion systems and works language independently. The model is trained end-to-end and based on a CNN-LSTM network that previously showed to give good results for speech quality estimation. We trained and tested the model on 16 different datasets, such as from the Blizzard Challenge and the Voice Conversion Challenge. Further, we show that the reliability of deep learning-based naturalness prediction can be improved by transfer learning from speech quality prediction models that are trained on objective POLQA scores. The proposed model is made publicly available and can, for example, be used to evaluate different TTS system configurations. Gabriel Mittag, Sebastian Möller 0001 |
INTERSPEECH | 1 |
| 2020 | DNN No-Reference PSTN Speech Quality PredictionabstractClassic public switched telephone networks (PSTN) are often a black box for VoIP network providers, as they have no access to performance indicators, such as delay or packet loss. Only the degraded output speech signal can be used to monitor the speech quality of these networks. However, the current state-of-the-art speech quality models are not reliable enough to be used for live monitoring. One of the reasons for this is that PSTN distortions can be unique depending on the provider and country, which makes it difficult to train a model that generalizes well for different PSTN networks. In this paper, we present a new open-source PSTN speech quality test set with over 1000 crowdsourced real phone calls. Our proposed no-reference model outperforms the full-reference POLQA and no-reference P.563 on the validation and test set. Further, we analyzed the influence of file cropping on the perceived speech quality and the influence of the number of ratings and training size on the model accuracy. Gabriel Mittag, Ross Cutler, Yasaman Hosseinkashi, Michael Revow, Sriram Srinivasan 0003, Naglakshmi Chande, Robert Aichner |
INTERSPEECH | 1 |
| 2020 | Analyzing the Fullband E-Model and Extending it for Predicting Bursty Packet LossabstractThe E-model is the only recommended parametric tool for planning the quality of speech communication services, and its fullband version has recently been standardized by the International Telecommunication Union, ITU-T. In this paper, we analyze and extend the model by comparing its predictions for random and bursty packet loss as well as for delay to the results of a signal-based model, POLQA, as well as to the result of a subjective conversation test. The analysis shows that, by extending the fullband model to account for the burstiness, a reasonable prediction accuracy can be reached for random as well as bursty loss. The results are discussed and limitations of the current model are pointed out. Thilo Michael, Gabriel Mittag, Sebastian Möller 0001 |
QoMEX | 2 |
| 2019 | Non-intrusive Speech Quality Assessment for Super-wideband Speech Communication NetworksabstractThe quality of speech communication networks has recently improved significantly by extending the available audio bandwidth from narrowband, firstly to wideband, and then to super-wideband. This bandwidth extension marks the end of the typically muffled sound we know from plain old telephone services. Another reason for increased speech quality is the fully digitally packet-based transmission. However, so far, no speech quality prediction model is able to estimate super-wideband quality without a clean reference signal. In this paper, we present a non-intrusive speech quality assessment model NISQA, which - in contrast to current state-of-the-art models - can predict the quality of super-wideband speech transmission. Furthermore, it is able to accurately predict the quality impact of packet loss concealment of modern codecs, such as Opus and EVS. The model uses a novel approach, where a CNN firstly estimates the per-frame quality, and subsequently, an RNN aggregates the per-frame values over time, to estimate the overall speech quality. Averaged over a comprehensive test set, the model achieves an RMSE*3rd of 0.29 with subjective MOS. Gabriel Mittag, Sebastian Möller 0001 |
ICASSP | 1 |
| 2019 | Extending the E-Model Towards Super-Wideband and Fullband Speech Communication Scenarios
Sebastian Möller 0001, Gabriel Mittag, Thilo Michael, Vincent Barriac, Hitoshi Aoki |
INTERSPEECH | 2 |
| 2019 | Quality Degradation Diagnosis for Voice Networks - Estimating the Perceived Noisiness, Coloration, and Discontinuity of Transmitted Speech
Gabriel Mittag, Sebastian Möller 0001 |
INTERSPEECH | 1 |
| 2019 | Semantic Labeling of Quality Impairments in Speech Spectrograms with Deep Convolutional NetworksabstractThere are numerous instrumental tools available to monitor the perceived quality of speech communication networks. However, these tools give no insight into the cause of a quality degradation. In this paper, we present a new method for quality diagnosis of speech communication networks that builds upon recent developments in the field of semantic image segmentation. The proposed model works non-intrusively, without the need for a clean reference signal. We use the deep convolutional network architecture SegNet and label each pixel of a speech spectrogram image as either clean or with its corresponding distortion. This way, quality degradations can directly be located in the time and frequency domain. To train the model, we created a large database with four different distortion types: packet-loss, background noise, GSM buzz, and bandwidth limitation. While processing the speech files, we also generated corresponding ground-truth labels with which we trained SegNet. Our experiments show promising results of this new diagnostic approach with a mIoU of 0.75. Gabriel Mittag, Sebastian Möller 0001 |
QoMEX | 1 |
| 2018 | Detecting Packet-Loss Concealment Using Formant Features and Decision Tree Learning
Gabriel Mittag, Sebastian Möller 0001 |
INTERSPEECH | 1 |
| 2018 | Effect of Number of Stimuli on Users Perception of Different Speech Degradations. A Crowdsourcing Case StudyabstractCrowdsourcing (CS) has established as a powerful tool to collect human input for data acquisition and labeling. However, it remains the question about the validity of the data collected in a CS platform. Sometimes, the users work carelessly or they try to tweak the system to maximize their profits. This paper reports on whether the number of speech stimuli presented to the listeners has an impact on the user perception of certain degradation conditions applied to the speech signal. To this end, a crowdsourcing study has been conducted with 209 listeners that were divided in three non-overlapping user groups, each of which was presented with tasks containing a different number of stimuli: 10, 20, or 40. Listeners were asked to rate speech stimuli with respect to their overall quality and the ratings were collected on a 5-point scale in accordance with ITU-T Rec. P.800. Workers assessed the speech stimuli of the database 501 from ITU-T Rec. P.863. Additionally, the influence of certain speech signal characteristics, such as interruptions and bandwidth, on the quality perception of the workers was investigated. Rafael Zequeira Jiménez, Gabriel Mittag, Sebastian Möller 0001 |
ISM | 2 |
| 2018 | Non-intrusive Estimation of Packet Loss Rates in Speech Communication Systems Using Convolutional Neural NetworksabstractIn this paper, we analyze whether deep convolutional neural networks can be used to detect lost packets in speech communication systems. The speech quality of modern communication networks has significantly improved recently, for example through higher available audio bandwidth. This was, among other reasons, possible through the use of packet-based networks, which allow a fully digital transmission from the sender to the receiver terminal. However, these networks often suffer from frequent interruptions caused by lost packets due to transmission errors. Consequently, the packet loss rate is one of the main indicators for the quality of speech communication services. In spite of that, the information of how many packets are lost in a network is not always available. To estimate the amount of lost packets, we calculate spectrograms of the transmitted speech signals and use them as input of a convolutional neural network. This approach has recently gained popularity in the field of detection and recognition tasks for music and speech. The interruptions caused by lost packets can often clearly be seen in the spectrogram of the degraded signal. Therefore, it seems natural to interpret the spectrograms as images and use deep learning methods that are common for image classification. The proposed model allows for estimating the packet loss rate of a communication system by simply using the recorded speech file from the receiver side, without the need of the reference speech signal that was originally sent through the channel. Our results show that the model reduces the prediction error by more than 75% when compared to a model that is based on MFCC features. Gabriel Mittag, Sebastian Möller 0001 |
ISM | 1 |
| 2018 | Variable Voice Likability Affecting Subjective Speech Quality AssessmentsabstractIn telephone conversations, transmitted speech of good to excellent quality is desired for enhanced Quality of Experience and to sustain lasting customer loyalty. Subjective mean opinion scores account for perceived transmitted quality, while instrumental models, such as POLQA, are able to estimate the subjective judgments. To perform subjective or instrumental quality measurements, the International Telecommunication Union recommends to employ two sentences from both, male and female speakers as speech material. In this paper, we have examined whether subjective and instrumental MOS ratings are affected by perceptual voice likability. A listening test has been conducted over 8 degradations with 12 extremely likable and unlikable male and female speakers. Statistically significant effects of gender and of voice likability have been detected on subjective MOS, whereas instrumental MOS was only affected by gender differences. These results can contribute to further improvements needed in the POLQA perceptual modeling, as well as to the selection of speakers for speech quality assessment tests. Laura Fernández Gallardo, Gabriel Mittag, Sebastian Möller 0001, John Beerends |
QoMEX | 2 |
| 2018 | Quantifying Quality Degradation of the EVS Super-Wideband Speech CodecabstractVoice transmission networks are commonly planned with the help of computational quality models, which give an estimate of the expected quality that a user will experience. The most popular of these tools is the E-model. When certain parameters are known, such as the applied codec and its bitrate, the model is able to predict the perceived quality of a communication system. Up to now, the E-model is only available for narrowband telephony (300–3400 Hz) and limited also for wideband telephony (100–7000 Hz). With the extension of voice networks to super-wideband telephony (50–14000 Hz), and the introduction of the super-wideband codec EVS to mobile networks and state of the art smartphones, an update of the E-model has become necessary. To this end, we firstly examined the quality improvement of super-wideband over wideband with results from mixed-band listening-only tests, where we found that the quality is improved by 15%. Then, we calculated impairment factors for the EVS codec and analyzed its robustness towards packet loss, by using auditory and instrumental methods. Gabriel Mittag, Sebastian Möller 0001, Vincent Barriac, Stephane Ragor |
QoMEX | 1 |
| 2017 | Modeling the overall quality of experience on the basis of underlying quality dimensionsabstractIn several Quality of Experience (QoE) assessment disciplines it has become common practice to not only assess the overall QoE but also underlying quality dimensions. However, in most cases the relation between the underlying quality dimensions and the overall QoE is not clear. In addition, it is not known which method provides the best results when trying to model the overall QoE on the basis of its underlying quality dimensions. To provide new ideas and applicable approaches for the QoE community, four different approaches to model the overall QoE are presented in this paper. To this end, the use case QoE of transmitted speech with its underlying perceptual quality dimensions is used. Based on three available databases, linear regression, multivariate adaptive regression splines, peak rule, and a combination of linear regression and peak rule are presented and compared. The results and the discussion reveal new insights into the overall QoE modeling process of transmitted speech and allows for drawing conclusions regarding other QoE disciplines. Friedemann Köster, Gabriel Mittag, Sebastian Möller 0001 |
QoMEX | 2 |