VLDB 2026 Research / reviewers in the wild / expert
Babak Naderi
dblp:130/2871 · also Babak Nadari
· DBLP profile ↗
32ranked-venue papers
13as first author
18since 2021 · last 2025
0009-0006-4778-5417ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 13 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 15 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Multidimensional Measurement of Photorealistic Avatars Quality of ExperienceabstractPhotorealistic avatars are human avatars that look, move, and talk like real people. The performance of photorealistic avatars has significantly improved recently based on objective metrics such as PSNR, SSIM, LPIPS, FID, and FVD. However, recent photorealistic avatar publications do not provide subjective tests of the avatars to measure human usability factors. We provide an open source test framework to subjectively measure photorealistic avatar performance in ten dimensions: realism, trust, comfortableness using, comfortableness interacting with, appropriateness for work, creepiness, formality, affinity, resemblance to the person, and emotion accuracy. Using telecommunication scenarios, we show that the correlation of nine of these subjective metrics with PSNR, SSIM, LPIPS, FID, and FVD is weak, and moderate for emotion accuracy. The crowdsourced subjective test framework is highly reproducible and accurate when compared to a panel of experts. We analyze a wide range of avatars from photorealistic to cartoon-like and show that some photorealistic avatars are approaching real video performance based on these dimensions. We also find that for avatars above a certain level of realism, eight of these measured dimensions are strongly correlated. This means that avatars that are not as realistic as real video will have lower trust, comfortableness using, comfortableness interacting with, appropriateness for work, formality, and affinity, and higher creepiness compared to real video. In addition, because there is a strong linear relationship between avatar affinity and realism, there is no uncanny valley effect for photorealistic avatars in the telecommunication scenario. We suggest several extensions of this test framework for future work and discuss design implications for telecommunication systems. The test framework is available at https://github.com/microsoft/P.910. Ross Cutler, Babak Naderi, Vishak Gopal, Dharmendar Reddy Palle |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2024 | A Crowdsourcing Approach to Video Quality AssessmentabstractWe propose an open-source extension of the ITU-T Rec. P.910 subjective video quality test based on crowdsourcing principles. This extension addresses the speed, usage cost, and barrier to usage issues of P.910. We implement Absolute Category Rating (ACR), ACR with hidden reference (ACRHR), Degradation Category Rating (DCR), and Comparison Category Rating (CCR), and include rater, environment, hardware, and network qualifications, as well as gold and trapping questions to ensure quality. We have validated that the implementation is both accurate and highly reproducible. Babak Naderi, Ross Cutler |
ICASSP | 1 |
| 2024 | VCD: A Video Conferencing Dataset for Video CompressionabstractCommonly used datasets for evaluating video codecs are all very high quality and not representative of video typically used in video conferencing scenarios. We present the Video Conferencing Dataset (VCD) for evaluating video codecs for real-time communication, the first such dataset focused on video conferencing. VCD includes a wide variety of camera qualities and spatial and temporal information. It includes both desktop and mobile scenarios and two types of video background processing. We report the compression efficiency of H.264, H.265, H.266, and AV1 in low-delay settings on VCD and compare it with the non-video conferencing datasets UVC, MLC-JVC, and HEVC. The results show the source quality and the scenarios have a significant effect on the compression efficiency of all the codecs. VCD enables the evaluation and tuning of codecs for this important scenario. The VCD is publicly available as an open-source dataset at https://github.com/microsoft/VCD. Babak Naderi, Ross Cutler, Nabakumar Singh Khongbantabam, Yasaman Hosseinkashi, Henrik Turbell, Albert Sadovnikov |
ICASSP | 1 |
| 2024 | Multi-Dimensional Speech Quality Assessment in CrowdsourcingabstractSubjective speech quality assessment is the gold standard for evaluating speech enhancement processing and telecommunication systems. The commonly used standard ITU-T Rec. P.800 defines how to measure speech quality in lab environments, and ITU-T Rec. P.808 extended it for crowdsourcing. ITU-T Rec. P.835 extends P.800 to measure the quality of speech in the presence of noise. ITU-T Rec. P.804 targets the conversation test and introduces perceptual speech quality dimensions which are measured during the listening phase of the conversation. The perceptual dimensions are noisiness, coloration, discontinuity, and loudness. We create a crowd-sourcing implementation of a multi-dimensional subjective test following the scales from P.804 and extend it to include reverberation, the speech signal, and overall quality. We show the tool is both accurate and reproducible. The tool has been used in the ICASSP 2023 Speech Signal Improvement challenge and we show the utility of these speech quality dimensions in this challenge. The tool will be publicly available as open-source at https://github.com/microsoft/P.808. Babak Naderi, Ross Cutler, Nicolae-Catalin Ristea |
ICASSP | 1 |
| 2023 | LSTM-Based Video Quality Prediction Accounting for Temporal Distortions in Videoconferencing CallsabstractCurrent state-of-the-art video quality models, such as VMAF, give excellent prediction results by comparing the degraded video with its reference video. However, they do not consider temporal distortions (e.g., frame freezes or skips) that occur during videoconferencing calls. In this paper, we present a data-driven approach for modeling such distortions automatically by training an LSTM with subjective quality ratings labeled via crowdsourcing. The videos were collected from live videoconferencing calls in 83 different network conditions. We applied QR codes as markers on the source videos to create aligned references and compute temporal features based on the alignment vectors. Using these features together with VMAF core features, our proposed model achieves a PCC of 0.99 on the validation set. Furthermore, our model outputs per-frame quality that gives detailed insight into the cause of video quality impairments. The VCM model and dataset are open-sourced at https://github.com/microsoft/Video_Call_MOS. Gabriel Mittag, Babak Naderi, Vishak Gopal, Ross Cutler |
ICASSP | 2 |
| 2022 | ConferencingSpeech 2022 Challenge: Non-intrusive Objective Speech Quality Assessment (NISQA) Challenge for Online Conferencing ApplicationsabstractWith the advances in speech communication systems such as online conferencing applications, we can seamlessly work with people regardless of where they are. However, during online meetings, speech quality can be significantly affected by background noise, reverberation, packet loss, network jitter, etc. Because of its nature, speech quality is traditionally assessed in subjective tests in laboratories and lately also in crowdsourcing following the international standards from ITU-T Rec. P.800 series. However, those approaches are costly and cannot be applied to customer data. Therefore, an effective objective assessment approach is needed to evaluate or monitor the speech quality of the ongoing conversation. The ConferencingSpeech 2022 challenge targets the non-intrusive deep neural network models for the speech quality assessment task. We open-sourced a training corpus with more than 86K speech clips in different languages, with a wide range of synthesized and live degradations and their corresponding subjective quality scores through crowdsourcing. 18 teams submitted their models for evaluation in this challenge. The blind test sets included about 4300 clips from wide ranges of degradations. This paper describes the challenge, the datasets, and the evaluation methods and reports the final results. Gaoxiong Yi, Babak Naderi, Sebastian Möller 0001, Wafaa Wardah, Gabriel Mittag, Ross Cutler, Zhuohuang Zhang, Donald S. Williamson, Fei Chen 0011, Shidong Shang |
INTERSPEECH | 4 |
| 2022 | Subjective Text Complexity Assessment for GermanabstractFor different reasons, text can be difficult to read and understand for many people, especially if the text’s language is too complex. In order to provide suitable text for the target audience, it is necessary to measure its complexity. In this paper we describe subjective experiments to assess the readability of German text. We compile a new corpus of sentences provided by a German IT service provider. The sentences are annotated with the subjective complexity ratings by two groups of participants, namely experts and non-experts for that text domain. We then extract an extensive set of linguistically motivated features that are supposedly interacting with complexity perception. We show that a linear regression model with a subset of these features can be a very good predictor of text complexity. Laura Seiffe, Fares Kallel, Sebastian Möller 0001, Babak Naderi, Roland Roller |
LREC | 4 |
| 2022 | Evaluating the Robustness of Speech Evaluation Standards for the CrowdabstractSubjective assessments are a key component of speech quality research. Traditionally, the assessments are conducted in laboratories in controlled conditions and following international standards like ITU-T Rec.P.800. However, even before the current pandemic, more speech quality research used crowdsourcing-based approaches for collecting subjective ratings. Crowdsourcing allows researchers to collect data even without a dedicated test laboratory, to collect data from a huge and diverse group of participants, and to perform the assessment in various real-life settings. Still, this approach raises questions about the reliability and validity of the subjective ratings, especially when comparing the ratings with data collected in standardized procedures. One step to approach these challenges was the development of the ITU-T Rec.P.808 standard. This standard helps practitioners implement best practices from speech quality studies and crowdsourcing studies in their crowdsourced speech quality assessments. However, even with the ITU-T Rec.P.808 in action, it is unclear how much background knowledge is necessary to successfully “implement” this standard. Therefore, this paper aims to assess the data quality differences between two P.808 implementations. One implementation is from a co-author of the P.808 standard, and the other is a researcher with only a little background in crowdsourcing and speech quality assessments. Both implementations are used in a large-scale crowdsourcing study with about two hundred users from Amazon Mechanical Turk. The collected ratings are compared to gold-standard data from a certified laboratory. Also, the two implementations are compared to analyze whether they lead to the same conclusions. The results show that both implementations correlate strongly with the laboratory and with each other. Thus, suggesting that the ITU-T Rec.P.808 is robust enough to be implemented by non-experts in speech evaluation or crowdsourcing. Edwin Gamboa, Babak Naderi, Matthias Hirth, Sebastian Möller 0001 |
QoMEX | 2 |
| 2021 | Crowdsourcing Approach for Subjective Evaluation of Echo ImpairmentabstractThe quality of acoustic echo cancellers (AECs) in real-time communication systems is typically evaluated using objective metrics like ERLE [1] and PESQ [2], and less commonly with lab-based subjective tests like ITU-T Rec. P.831 [3]. We will show that these objective measures are not well correlated to subjective measures. We then introduce an open-source crowdsourcing approach for subjective evaluation of echo impairment which can be used to evaluate the performance of AECs. We provide a study that shows this tool is highly reproducible. This new tool has been recently used in the ICASSP 2021 AEC Challenge [4] which made the challenge possible to do quickly and cost effectively. Ross Cutler, Babak Naderi, Markus Loide, Sten Sootla, Ando Saabas |
ICASSP | 2 |
| 2021 | Effect of Language Proficiency on Subjective Evaluation of Noise Suppression AlgorithmsabstractSpeech communication systems based on Voice-over-IP technology are frequently used by native as well as non-native speakers of a target language, e.g. in international phone calls or telemeetings. Frequently, such calls also occur in a noisy environment, making noise suppression modules necessary to increase perceived quality of experience. Whereas standard tests for assessing perceived quality make use of native listeners, we assume that noise-reduced speech and residual noise may affect native and non-native listeners of a target language in different ways. To test this assumption, we report results of two subjective tests conducted with English and German native listeners who judge the quality of speech samples recorded by native English, German, and Mandarin speakers, which are degraded with different background noise levels and noise suppression effects. The experiments were conducted following the standardized ITU-T Rec. P.835 approach, however implemented in a crowdsourcing setting according to ITU-T Rec. P.808. Our results show a significant influence of language on speech signal ratings and, consequently, on the overall perceived quality in specific conditions. Babak Naderi, Gabriel Mittag, Rafael Zequeira Jiménez, Sebastian Möller 0001 |
ICASSP | 1 |
| 2021 | NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced DatasetsabstractIn this paper, we present an update to the NISQA speech quality prediction model that is focused on distortions that occur in communication networks. In contrast to the previous version, the model is trained end-to-end and the time-dependency modelling and time-pooling is achieved through a Self-Attention mechanism. Besides overall speech quality, the model also predicts the four speech quality dimensions Noisiness, Coloration, Discontinuity, and Loudness, and in this way gives more insight into the cause of a quality degradation. Furthermore, new datasets with over 13,000 speech files were created for training and validation of the model. The model was finally tested on a new, live-talking test dataset that contains recordings of real telephone calls. Overall, NISQA was trained and evaluated on 81 datasets from different sources and showed to provide reliable predictions also for unknown speech samples. The code, model weights, and datasets are open-sourced. Gabriel Mittag, Babak Naderi, Assmaa Chehadi, Sebastian Möller 0001 |
Interspeech | 2 |
| 2021 | Subjective Evaluation of Noise Suppression Algorithms in CrowdsourcingabstractThe quality of the speech communication systems, which include noise suppression algorithms, are typically evaluated in laboratory experiments according to the ITU-T Rec. P.835, in which participants rate background noise, speech signal, and overall quality separately. This paper introduces an open-source toolkit for conducting subjective quality evaluation of noise suppressed speech in crowdsourcing. We followed the ITU-T Rec. P.835, and P.808 and highly automate the process to prevent moderator's error. To assess the validity of our evaluation method, we compared the Mean Opinion Scores (MOS), calculate using ratings collected with our implementation, and the MOS values from a standard laboratory experiment conducted according to the ITU-T Rec P.835. Results show a high validity in all three scales namely background noise, speech signal and overall quality (average PCC = 0.961). Results of a round-robin test (N=5) showed that our implementation is also a highly reproducible evaluation method (PCC=0.99). Finally, we used our implementation in the INTERSPEECH 2021 Deep Noise Suppression Challenge as the primary evaluation metric, which demonstrates it is practical to use at scale. The results are analyzed to determine why the overall performance was the best in terms of background noise and speech quality. Babak Naderi, Ross Cutler |
Interspeech | 1 |
| 2021 | Perception of Social Speaker Characteristics in Synthetic Speech
Sai Sirisha Rallabandi, Abhinav Bharadwaj, Babak Naderi, Sebastian Möller 0001 |
Interspeech | 3 |
| 2021 | On the Impact of COVID-19 on Subjective Digital Media Quality AssessmentabstractThe COVID-19 pandemic has induced dramatic effects in all areas of society worldwide. It has led to drastic restrictions on academic exchanges and caused research projects to be put on hold. The need to maintain social distancing has severely affected the research communities that rely on in-person studies. In particular, Quality of Experience (QoE) research relies heavily on user studies and in-person subjective tests as means of obtaining ground truths for system designs. In this position paper, we focus on the impact of COVID-19 on conducting subjective tests for digital media quality assessment. An overview of the related discussions that take place in associated research communities is provided. A number of challenges are posed in terms of hygiene, ethics, standardization, and research quality issues. Opportunities for consolidating QoE research under COVID-19 and beyond are put forward, including alternative experimental designs and adaptation of research methodologies. Numerous enablers are suggested and discussed, revealing various avenues that can be followed to assist in conducting subjective tests under pandemic conditions. We believe that this position paper will be helpful for researchers and practitioners relying on in-person studies to keep abreast about the potential measures that support coming out of the COVID-19 pandemic strong and being prepared for future endeavors. Hans-Jürgen Zepernick, Kerstin Pieper, Robert P. Spang, Ulrich Engelke, Matthias Hirth, Babak Naderi |
MMSP | 6 |
| 2021 | On Inter-Rater Reliability for Crowdsourced QoEabstractCrowdsourcing offers a faster, cheaper, and more scalable approach than the traditional laboratory quality assessment tests. However, participants perform the test in their own working environment, using their own hardware and without direct supervision of a test moderator, leading to different types of biases on the ratings. In this paper, we compare several reliability metrics that are commonly applied to the subjective ratings in terms of their sensitivity to identify typical issues of crowdsourced media quality tests. Following the subject bias theory, we simulate the ratings of different user groups with different bias and various magnitudes of uncertainty, while also considering the presence of unreliable raters. We apply traditional reliability metrics on the ratings and compare their sensitivity in identifying the severity of the raters' biases and uncertainties. Our results show that the average Spearman's rank correlation coefficient between raters can serve as a strong indicator for issues with the crowdsourcing study. This means that scoring too low for this metric should encourage researchers to revisit their study design in order to eventually improve the reliability of results from crowdsourcing-based quality studies. Tobias Hoßfeld, Michael Seufert, Babak Naderi |
QoMEX | 3 |
| 2021 | Influence of Language Differences in Crowdsourcing Speech Quality Assessment StudiesabstractThe quality of the speech signal is essential as it influences the user experience of voiced interactive systems. Speech quality studies have traditionally been conducted in restricted laboratory rooms with professional audio equipment. Nowadays, crowd-sourcing represents a valid alternative for the rapid assessment of large speech databases at a fraction of the cost and time of traditional laboratory practices. However, crowd-sourcing users perform tasks in an unsupervised manner. Thus, it is challenging to control whether their skills match those of the study's intended audience. This is important in speech quality evaluations as some listeners may end up participating in a listening test of a target language other than their mother tongue. This paper investigates the influence of assessing the quality of a German speech dataset with native English and Spanish speakers. To this end, three studies were conducted in crowdsourcing where listeners evaluated the quality of speech stimuli following the ITU-T Rec. P.808. A strong Pearson correlation and low RMSE was found between the laboratory ratings and the scores collected in all crowdsourcing studies, despite the listeners' mother tongue. Still, a bias was seen between the mean opinion scores from the German crowd-workers and the native English and Spanish speakers. The non-German participants tended to overestimate the quality of the speech stimuli. Rafael Zequeira Jiménez, Babak Naderi, Sebastian Möller 0001 |
QoMEX | 2 |
| 2021 | Bias-Aware Loss for Training Image and Speech Quality Prediction Models from Multiple DatasetsabstractThe ground truth used for training image, video, or speech quality prediction models is based on the Mean Opinion Scores (MOS) obtained from subjective experiments. Usually, it is necessary to conduct multiple experiments, mostly with different test participants, to obtain enough data to train quality models based on machine learning. Each of these experiments is subject to an experiment-specific bias, where the rating of the same file may be substantially different in two experiments (e.g. depending on the overall quality distribution). These different ratings for the same distortion levels confuse neural networks during training and lead to lower performance. To overcome this problem, we propose a bias-aware loss function that estimates each dataset's biases during training with a linear function and considers it while optimising the network weights. We prove the efficiency of the proposed method by training and validating quality prediction models on synthetic and subjective image and speech quality datasets. Gabriel Mittag, Saman Zad Tootaghaj, Thilo Michael, Babak Naderi, Sebastian Möller 0001 |
QoMEX | 4 |
| 2021 | Speech Quality Assessment in Crowdsourcing: Comparison Category Rating MethodabstractTraditionally, Quality of Experience (QoE) for a communication system is evaluated through a subjective test. The most common test method for speech QoE is the Absolute Category Rating (ACR), in which participants listen to a set of stimuli, processed by the underlying test conditions, and rate their perceived quality for each stimulus on a specific scale. The Comparison Category Rating (CCR) is another standard approach in which participants listen to both reference and processed stimuli and rate their quality compared to the other one. The CCR method is particularly suitable for systems that improve the quality of input speech. This paper evaluates an adaptation of the CCR test procedure for assessing speech quality in the crowdsourcing set-up. The CCR method was introduced in the ITU-T Rec. P.800 for laboratory-based experiments. We adapted the test for the crowdsourcing approach following the guidelines from ITU-T Rec. P.800 and P.808. We show that the results of the CCR procedure via crowdsourcing are highly reproducible. We also compared the CCR test results with widely used ACR test procedures obtained in the laboratory and crowdsourcing. Our results show that the CCR procedure in crowdsourcing is a reliable and valid test method. Babak Naderi, Sebastian Möller 0001, Ross Cutler |
QoMEX | 1 |
| 2020 | An Open Source Implementation of ITU-T Recommendation P.808 with ValidationabstractThe ITU-T Recommendation P.808 provides a crowdsourcing approach for conducting a subjective assessment of speech quality using the Absolute Category Rating (ACR) method. We provide an open-source implementation of the ITU-T Rec. P.808 that runs on the Amazon Mechanical Turk platform. We extended our implementation to include Degradation Category Ratings (DCR) and Comparison Category Ratings (CCR) test methods. We also significantly speed up the test process by integrating the participant qualification step into the main rating task compared to a two-stage qualification and rating solution. We provide program scripts for creating and executing the subjective test, and data cleansing and analyzing the answers to avoid operational errors. To validate the implementation, we compare the Mean Opinion Scores (MOS) collected through our implementation with MOS values from a standard laboratory experiment conducted based on the ITU-T Rec. P.800. We also evaluate the reproducibility of the result of the subjective speech quality assessment through crowdsourcing using our implementation. Finally, we quantify the impact of parts of the system designed to improve the reliability: environmental tests, gold and trapping questions, rating patterns, and a headset usage test. Babak Naderi, Ross Cutler |
INTERSPEECH | 1 |
| 2020 | A latency compensation technique based on game characteristics to mitigate the influence of delay on cloud gaming quality of experienceabstractCloud Gaming (CG) is an immersive multimedia service that promises many benefits. In CG, the games are rendered in a cloud server, and the resulted scenes are streamed as a video sequence to the client. Using CG users are not forced to update their gaming hardware frequently, and available games can be played on any operating system or suitable device. However, cloud gaming requires a reliable and low-latency network, which makes it a very challenging service. Transmission latency strongly affects the playability of a cloud game and consequently reduces the users' Quality of Experience (QoE). In this paper, we propose a latency compensation technique using game adaptation that mitigates the influence of delay on QoE. This technique uses five game characteristics for the adaptation. These characteristics, in addition to an Aim-assistance technique, were implemented in four games for evaluation. A subjective study using 194 participants was conducted using a crowdsourcing approach. The results showed that the majority of the proposed adaptation techniques lead to significant improvements in the cloud gaming QoE. Saeed Shafiee Sabet, Steven Schmidt 0001, Saman Zad Tootaghaj, Babak Naderi, Carsten Griwodz, Sebastian Möller 0001 |
MMSys | 4 |
| 2020 | Assessing Interactive Gaming Quality of Experience using a Crowdsourcing ApproachabstractTraditionally, the Quality of Experience (QoE) is assessed in a controlled laboratory environment where participants give their opinion about the perceived quality of a stimulus on a standardized rating scale. Recently, the usage of crowdsourcing micro-task platforms for assessing the media quality is increasing. The crowdsourcing platforms provide access to a pool of geographically distributed, and demographically diverse group of workers who participate in the experiment in their own working environment and using their own hardware. The main challenge in crowdsourcing QoE tests is to control the effect of interfering influencing factors such as a user's environment and device on the subjective ratings. While in the past, the crowdsourcing approach was frequently used for speech and video quality assessment, research on a quality assessment for gaming services is rare. In this paper, we present a method to measure gaming QoE under typically considered system influence factors including delay, packet loss, and framerates as well as different game designs. The factors are artificially manipulated due to controlled changes in the implementation of games. The results of a total of five studies using a developed evaluation method based on a combination of the ITU-T Rec. P.809 on subjective evaluation methods for gaming quality and the ITU-T Rec. P.808 on subjective evaluation of speech quality with a crowdsourcing approach will be discussed. To evaluate the reliability and validity of results collected using this method, we finally compare subjective ratings regarding the effect of network delay on gaming QoE gathered from interactive crowdsourcing tests with those from equivalent laboratory experiments. Steven Schmidt 0001, Babak Naderi, Saeed Shafiee Sabet, Saman Zad Tootaghaj, Sebastian Möller 0001 |
QoMEX | 2 |
| 2020 | Effect of Environmental Noise in Speech Quality Assessment Studies using CrowdsourcingabstractCrowdsourcing is a valid approach to collect and annotate data efficiently and cost-effectively. This approach permits us to reach a large and diverse pool of users that usually work from home employing their computers and headphones. Still, there is insufficient information about the users' surroundings. Specifically, little knowledge about the background noise to which users might be exposed to when executing crowd-work. The validity of the data gathered in a disturbed environment is questionable, especially in speech quality assessment and other audio-related tasks. This work presents the results of a simulated crowdsourcing study conducted in the laboratory. We investigate the influence of environmental background noise in speech quality assessment tests. Three groups of listeners were recruited to rate the quality of speech files under the influence of background noise at different levels. Two types of noise were tested, i.e., street noises and TV-Show. Our findings suggest that the threshold at which an environmental background noise would significantly affect the speech quality ratings in crowdsourcing is between 43dB(A) and 50dB(A). Additionally, listeners tolerated more the TV-Show noise. They provided more accurate ratings while conducting the test under the influence of higher levels of the TV-Show noise, than at lower levels of the street noise. We also found that the presence of background noise does not cause a constant bias of the quality scores; instead, its impact depends on the speech degradation condition under test. Rafael Zequeira Jiménez, Babak Naderi, Sebastian Möller 0001 |
QoMEX | 2 |
| 2020 | Application of Just-Noticeable Difference in Quality as Environment Suitability Test for Crowdsourcing Speech Quality Assessment TaskabstractCrowdsourcing micro-task platforms facilitate subjective media quality assessment by providing access to a highly scaleable, geographically distributed and demographically diverse pool of crowd workers. Those workers participate in the experiment remotely from their own working environment, using their own hardware. In the case of speech quality assessment, preliminary work showed that environmental noise at the listener's side and the listening device (loudspeaker or headphone) significantly affect perceived quality, and consequently the reliability and validity of subjective ratings. As a consequence, ITU-T Rec. P.808 specifies requirements for the listening environment of crowd workers when assessing speech quality. In this paper, we propose a new Just Noticeable Difference of Quality (JNDQ) test as a remote screening method for assessing the suitability of the work environment for participating in speech quality assessment tasks. In a laboratory experiment, participants performed this JNDQ test with different listening devices in different listening environments, including a silent room according to ITU-T Rec. P.800 and a simulated background noise scenario. Results show a significant impact of the environment and the listening device on the JNDQ threshold. Thus, the combination of listening device and background noise needs to be screened in a crowdsourcing speech quality test. We propose a minimum threshold of our JNDQ test as an easily applicable screening method for this purpose. Babak Naderi, Sebastian Möller 0001 |
QoMEX | 1 |
| 2020 | Transformation of Mean Opinion Scores to Avoid Misleading of Ranked Based Statistical TechniquesabstractThe rank correlation coefficients and the ranked-based statistical tests (as a subset of non-parametric techniques) might be misleading when they are applied to subjectively collected opinion scores. Those techniques assume that the data is measured at least at an ordinal level and define a sequence of scores to represent a tied rank when they have precisely an equal numeric value. In this paper, we show that the definition of tied rank, as mentioned above, is not suitable for Mean Opinion Scores (MOS) and might be misleading conclusions of rank-based statistical techniques. Furthermore, we introduce a method to overcome this issue by transforming the MOS values considering their 95% Confidence Intervals. The rank correlation coefficients and ranked-based statistical tests can then be safely applied to the transformed values. We also provide open-source software packages in different programming languages to utilize the application of our transformation method in the quality of experience domain. Babak Naderi, Sebastian Möller 0001 |
QoMEX | 1 |
| 2020 | Impact of the Number of Votes on the Reliability and Validity of Subjective Speech Quality Assessment in the Crowdsourcing ApproachabstractThe subjective quality of transmitted speech is traditionally assessed in a controlled laboratory environment according to ITU-T Rec. P.800. In turn, with crowdsourcing, crowdworkers participate in a subjective online experiment using their own listening device, and in their own working environment. Despite such less controllable conditions, the increased use of crowdsourcing micro-task platforms for quality assessment tasks has pushed a high demand for standardized methods, resulting in ITU-T Rec. P.808. This work investigates the impact of the number of judgments on the reliability and the validity of quality ratings collected through crowdsourcing-based speech quality assessments, as an input to ITU-T Rec. P.808 . Three crowdsourcing experiments on different platforms were conducted to evaluate the overall quality of three different speech datasets, using the Absolute Category Rating procedure. For each dataset, the Mean Opinion Scores (MOS) are calculated using differing numbers of crowdsourcing judgements. Then the results are compared to MOS values collected in a standard laboratory experiment, to assess the validity of crowdsourcing approach as a function of number of votes. In addition, the reliability of the average scores is analyzed by checking inter-rater reliability, gain in certainty, and the confidence of the MOS. The results provide a suggestion on the required number of votes per condition, and allow to model its impact on validity and reliability. Babak Naderi, Tobias Hoßfeld, Matthias Hirth, Florian Metzger, Sebastian Möller 0001, Rafael Zequeira Jiménez |
QoMEX | 1 |
| 2019 | Modeling Worker Performance Based on Intra-rater Reliability in Crowdsourcing : A Case Study of Speech Quality AssessmentabstractCrowdsourcing has become a convenient instrument for addressing subjective user studies to a large amounts of users. Data from crowdsourcing can be corrupted due to users' neglect, and different mechanisms has been proposed to address the users' reliability and to ensure valid experiments' results. Users that are consistent in their answers or present a high intra-rater reliability score, are desired for subjective studies. This work investigates the relationship between the intra-rater reliability and the user performance in the context of a speech quality assessment task. To this end, a crowdsourcing study has been conducted in which users were requested to rate speech stimuli with respect to their overall quality. Ratings were collected on a 5-point scale in accordance with the ITU-T Rec. P.808. The speech stimuli were taken from the database ITU-T Rec. P.501 Annex D, and the results are to be contrasted with ratings collected in a laboratory experiment. Furthermore, a model as a function of intra-rater reliability, root-mean-squared-deviation between the listeners ratings and age, has been built to predict the listener performance. Such a model is intended to provide a measure of how valid the crowdsourcing results are, when there is no laboratory results to compare to. Rafael Zequeira Jiménez, Anna Llagostera Casanovas, Babak Naderi, Sebastian Möller 0001, Jens Berger |
QoMEX | 3 |
| 2019 | Background Environment Characteristics of Crowd-Workers from German Speaking Countries Experimental Survey on User Environment CharacteristicsabstractCrowdsourcing has been used extensively for gathering and annotating data cost efficiently. Nowadays, there are multiple platforms offering crowd-sourced workforce, still most of these users are from Asia or English speaking countries, and not so many native German speakers. Thus, there is a lack of information regarding the conditions in which German users execute tasks, neither about their habits when taking part in crowdsourcing campaigns. Which is of main importance to address properly user studies to German crowd-workers. This paper reports on a survey that investigated the environments' characteristics of users from German speaking countries. To this end, a study has been conducted in which users were asked to provide details about the surroundings in which they normally execute crowdsourcing tasks. Audio and visual data was collected per user which contributed to aggregate even more information on the users' input. We provide insights aimed at easing the decision making process when designing subjective user studies. Rafael Zequeira Jiménez, Babak Naderi, Sebastian Möller 0001 |
QoMEX | 2 |
| 2019 | Automated Text Readability Assessment for German Language: A Quality of Experience ApproachabstractData-driven approaches towards readability assessment, using automated linguistic analysis and machine learning methods, is a viable road forward for readability rankings. This paper describes the development of an automated readability assessment estimator based on supervised learning algorithms over German text corpora. For this purpose, natural language processing tools are used to extract 73 linguistic features grouped in traditional, lexical and morphological features. Feature engineering approaches are employed to select informative features. Different supervised learning models are implemented, with the top-ranked features fed as input. The results obtained depict that Random Forest Regressor yielding best result (0.847) for RMSE measure. Babak Naderi, Salar Mohtaj, Karan Karan, Sebastian Möller 0001 |
QoMEX | 1 |
| 2018 | Personalized Motivation-supportive Messages for Increasing Participation in Crowd-civic SystemsabstractIn crowd-civic systems, citizens form groups and work towards shared goals, such as discovering social issues or reforming official policies. Unfortunately, many real-world systems have been unsuccessful in continually motivating large numbers of citizens to participate voluntarily, despite various approaches such as gamification and persuasion techniques. In this paper, we examine the influence of personalized messages designed to support motivation as asserted by the Self-Determination Theory (SDT). We designed a crowd-civic platform for collecting community issues with personalized motivation-supportive messages and conducted two studies: a pair-comparison experiment with 150 participants on Amazon's Mechanical Turk and a live deployment study with 120 university members. Results of the pair-comparison study indicate applicability of SDT's perspective in crowd-civic systems. While applying it in the live system surfaced several challenges, including recruiting participants without interfering with general motivations, the collected data exhibited similar promising trends. Paul Grau, Babak Naderi, Juho Kim 0001 |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2015 | Effect of trapping questions on the reliability of speech quality judgments in a crowdsourcing paradigm
Babak Naderi, Tim Polzehl, Ina Wechsung, Friedemann Köster, Sebastian Möller 0001 |
INTERSPEECH | 1 |
| 2015 | Robustness in speech quality assessment and temporal training expiry in mobile crowdsourcing environments
Tim Polzehl, Babak Naderi, Friedemann Köster, Sebastian Möller 0001 |
INTERSPEECH | 2 |
| 2014 | Crowdee: mobile crowdsourcing micro-task platform for celebrating the diversity of languages
Babak Naderi, Tim Polzehl, André Beyer, Tibor Pilz, Sebastian Möller 0001 |
INTERSPEECH | 1 |