EDBT 2026 Demo / reviewers in the wild / expert
Ross Cutler
dblp:94/6224
· DBLP profile ↗
63ranked-venue papers
13as first author
35since 2021 · last 2026
0000-0002-2004-3003ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 10 first-author · 28 since 2021Artificial intelligence and machine learning · 24 · 6 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 4 since 2021Computer networks · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human-in-the-Loop Bandwidth Estimation for Quality of Experience Optimization in Real-Time Video CommunicationabstractThe quality of experience (QoE) delivered by video conferencing systems is significantly influenced by accurately estimating the time-varying available bandwidth between the sender and receiver. Bandwidth estimation for real-time communications remains an open challenge due to rapidly evolving network architectures, increasingly complex protocol stacks, and the difficulty of defining QoE metrics that reliably improve user experience. In this work, we propose a deployed, human-in-the-loop, data-driven framework for bandwidth estimation to address these challenges. Our approach begins with training objective QoE reward models derived from subjective user evaluations to measure audio and video quality in real-time video conferencing systems. Subsequently, we collect roughly 1M network traces with objective QoE rewards from real-world Microsoft Teams calls to curate a bandwidth estimation training dataset. We then introduce a novel distributional offline reinforcement learning (RL) algorithm to train a neural-network-based bandwidth estimator aimed at improving QoE for users. Our real-world A/B test demonstrates that the proposed approach reduces the subjective poor call ratio by 11.41% compared to the baseline bandwidth estimator. Furthermore, the proposed offline RL algorithm is benchmarked on D4RL tasks to demonstrate its generalization beyond bandwidth estimation. Sami Khairy, Gabriel Mittag, Vishak Gopal, Ross Cutler |
AAAI | 4 |
| 2026 | Offline Meta-learning for Real-time Bandwidth Estimation
Aashish Gottipati, Sami Khairy, Yasaman Hosseinkashi, Gabriel Mittag, Vishak Gopal, Francis Y. Yan, Ross Cutler |
ICC | 7 |
| 2025 | Offline to Online Learning for Real-Time Bandwidth EstimationabstractReal-time video applications require accurate bandwidth estimation (BWE) to maintain user experience across varying network conditions. However, increasing network heterogeneity challenges general-purpose BWE algorithms, necessitating solutions that adapt to end-user environments. While widely adopted, heuristic-based methods are difficult to individualize without extensive domain expertise. Conversely, online reinforcement learning (RL) offers ease of customization but neglects prior domain expertise and suffers from sample inefficiency. Thus, we present Merlin, an imitation learning-based solution that replaces the manual parameter tuning of heuristic-based methods with data-driven updates to streamline end-user personalization. Our key insight is that transforming heuristic-based BWE algorithms into neural networks facilitates data-driven personalization. Merlin utilizes Behavioral Cloning to efficiently learn from offline telemetry logs, capturing heuristic policies without live network interactions. The cloned policy can then be seamlessly tailored to end user network conditions through online finetuning. In real intercontinental videoconferencing calls, Merlin matches our heuristic's policy with no statistically significant differences in user quality of experience (QoE). Finetuning Merlin's control policy to end-user environments enables QoE improvements of up to 7.8 % compared to the heuristic policy. Lastly, our IL-based design performs competitively with current state-of-the-art online RL techniques but converges with 80 % fewer videoconferencing samples, facilitating practical end-user personalization. Aashish Gottipati, Sami Khairy, Gabriel Mittag, Vishak Gopal, Ross Cutler |
ICC | 5 |
| 2025 | A Multidimensional Measurement of Photorealistic Avatars Quality of ExperienceabstractPhotorealistic avatars are human avatars that look, move, and talk like real people. The performance of photorealistic avatars has significantly improved recently based on objective metrics such as PSNR, SSIM, LPIPS, FID, and FVD. However, recent photorealistic avatar publications do not provide subjective tests of the avatars to measure human usability factors. We provide an open source test framework to subjectively measure photorealistic avatar performance in ten dimensions: realism, trust, comfortableness using, comfortableness interacting with, appropriateness for work, creepiness, formality, affinity, resemblance to the person, and emotion accuracy. Using telecommunication scenarios, we show that the correlation of nine of these subjective metrics with PSNR, SSIM, LPIPS, FID, and FVD is weak, and moderate for emotion accuracy. The crowdsourced subjective test framework is highly reproducible and accurate when compared to a panel of experts. We analyze a wide range of avatars from photorealistic to cartoon-like and show that some photorealistic avatars are approaching real video performance based on these dimensions. We also find that for avatars above a certain level of realism, eight of these measured dimensions are strongly correlated. This means that avatars that are not as realistic as real video will have lower trust, comfortableness using, comfortableness interacting with, appropriateness for work, formality, and affinity, and higher creepiness compared to real video. In addition, because there is a strong linear relationship between avatar affinity and realism, there is no uncanny valley effect for photorealistic avatars in the telecommunication scenario. We suggest several extensions of this test framework for future work and discuss design implications for telecommunication systems. The test framework is available at https://github.com/microsoft/P.910. Ross Cutler, Babak Naderi, Vishak Gopal, Dharmendar Reddy Palle |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2024 | A Real-Time Active Speaker Detection System Integrating an Audio-Visual Signal with a Spatial Querying MechanismabstractWe introduce a distinctive real-time, causal, neural network-based active speaker detection system optimized for low-power edge computing. This system drives a virtual cinematography module and is deployed on a commercial device. The system uses data originating from a microphone array and a 360-degree camera. Our network requires only 127 MFLOPs per participant, for a meeting with 14 participants. Unlike previous work, we examine the error rate of our network when the computational budget is exhausted, and find that it exhibits graceful degradation, allowing the system to operate reasonably well even in this case. Departing from conventional DOA estimation approaches, our network learns to query the available acoustic data, considering the detected head locations. We train and evaluate our algorithm on a realistic meetings dataset featuring up to 14 participants in the same meeting, overlapped speech, and other challenging scenarios. Ilya Gurvich, Ido Leichter, Dharmendar Reddy Palle, Yossi Asher, Alon Vinnikov, Igor Abramovski, Vishak Gopal, Ross Cutler, Eyal Krupka |
ICASSP | 8 |
| 2024 | A Crowdsourcing Approach to Video Quality AssessmentabstractWe propose an open-source extension of the ITU-T Rec. P.910 subjective video quality test based on crowdsourcing principles. This extension addresses the speed, usage cost, and barrier to usage issues of P.910. We implement Absolute Category Rating (ACR), ACR with hidden reference (ACRHR), Degradation Category Rating (DCR), and Comparison Category Rating (CCR), and include rater, environment, hardware, and network qualifications, as well as gold and trapping questions to ensure quality. We have validated that the implementation is both accurate and highly reproducible. Babak Naderi, Ross Cutler |
ICASSP | 2 |
| 2024 | VCD: A Video Conferencing Dataset for Video CompressionabstractCommonly used datasets for evaluating video codecs are all very high quality and not representative of video typically used in video conferencing scenarios. We present the Video Conferencing Dataset (VCD) for evaluating video codecs for real-time communication, the first such dataset focused on video conferencing. VCD includes a wide variety of camera qualities and spatial and temporal information. It includes both desktop and mobile scenarios and two types of video background processing. We report the compression efficiency of H.264, H.265, H.266, and AV1 in low-delay settings on VCD and compare it with the non-video conferencing datasets UVC, MLC-JVC, and HEVC. The results show the source quality and the scenarios have a significant effect on the compression efficiency of all the codecs. VCD enables the evaluation and tuning of codecs for this important scenario. The VCD is publicly available as an open-source dataset at https://github.com/microsoft/VCD. Babak Naderi, Ross Cutler, Nabakumar Singh Khongbantabam, Yasaman Hosseinkashi, Henrik Turbell, Albert Sadovnikov |
ICASSP | 2 |
| 2024 | Multi-Dimensional Speech Quality Assessment in CrowdsourcingabstractSubjective speech quality assessment is the gold standard for evaluating speech enhancement processing and telecommunication systems. The commonly used standard ITU-T Rec. P.800 defines how to measure speech quality in lab environments, and ITU-T Rec. P.808 extended it for crowdsourcing. ITU-T Rec. P.835 extends P.800 to measure the quality of speech in the presence of noise. ITU-T Rec. P.804 targets the conversation test and introduces perceptual speech quality dimensions which are measured during the listening phase of the conversation. The perceptual dimensions are noisiness, coloration, discontinuity, and loudness. We create a crowd-sourcing implementation of a multi-dimensional subjective test following the scales from P.804 and extend it to include reverberation, the speech signal, and overall quality. We show the tool is both accurate and reproducible. The tool has been used in the ICASSP 2023 Speech Signal Improvement challenge and we show the utility of these speech quality dimensions in this challenge. The tool will be publicly available as open-source at https://github.com/microsoft/P.808. Babak Naderi, Ross Cutler, Nicolae-Catalin Ristea |
ICASSP | 2 |
| 2024 | ACM MMSys 2024 Bandwidth Estimation in Real Time Communications ChallengeabstractThe quality of experience (QoE) delivered by video conferencing systems to end users depends in part on correctly estimating the capacity of the bottleneck link between the sender and the receiver over time. Bandwidth estimation for real-time communications (RTC) remains a significant challenge, primarily due to the continuously evolving heterogeneous network architectures and technologies. From the first bandwidth estimation challenge which was hosted at ACM MMSys 2021, we learned that bandwidth estimation models trained with reinforcement learning (RL) in simulations to maximize network-based reward functions may not be optimal in reality due to the sim-to-real gap and the difficulty of aligning network-based rewards with user-perceived QoE. This grand challenge aims to advance bandwidth estimation model design by aligning reward maximization with user-perceived QoE optimization using offline RL and a real-world dataset with objective rewards which have high correlations with subjective audio/video quality in Microsoft Teams. All models submitted to the grand challenge underwent initial evaluation on our emulation platform. For a comprehensive evaluation under diverse network conditions with temporal fluctuations, top models were further evaluated on our geographically distributed testbed by using each model to conduct 600 calls within a 12-day period. The winning model is shown to deliver comparable performance to the top behavior policy in the released dataset. By leveraging real-world data and integrating objective audio/video quality scores as rewards, offline RL can therefore facilitate the development of competitive bandwidth estimators for RTC. Sami Khairy, Gabriel Mittag, Vishak Gopal, Francis Y. Yan, Zhixiong Niu, Ezra Ameri, Scott Inglis, Mehrsa Golestaneh, Ross Cutler |
MMSys | 9 |
| 2024 | Topic-Conversation Relevance (TCR) Dataset and BenchmarksabstractWorkplace meetings are vital to organizational collaboration, yet a large percentage of meetings are rated as ineffective. To help improve meeting effectiveness by understanding if the conversation is on topic, we create a comprehensive Topic-Conversation Relevance (TCR) dataset that covers a variety of domains and meeting styles. The TCR dataset includes 1,500 unique meetings, 22 million words in transcripts, and over 15,000 meeting topics, sourced from both newly collected Speech Interruption Meeting (SIM) data and existing public datasets. Along with the text data, we also open source scripts to generate synthetic meetings or create augmented meetings from the TCR dataset to enhance data diversity. For each data source, benchmarks are created using GPT-4 to evaluate the model accuracy in understanding transcription-topic relevance. Yaran Fan, Jamie Pool, Senja Filipi, Ross Cutler |
NeurIPS | 4 |
| 2024 | Meeting Effectiveness and Inclusiveness: Large-scale Measurement, Identification of Key Features, and Prediction in Real-world Remote MeetingsabstractWorkplace meetings are vital to organizational collaboration, yet relatively little progress has been made toward measuring meeting effectiveness and inclusiveness at scale. The recent rise in remote and hybrid meetings represents an opportunity to do so via computer-mediated communication (CMC) systems. Here, we share the results of an effective and inclusive meetings survey embedded within a CMC system in a diverse set of companies and organizations. We correlate the survey results with objective metrics available from the CMC system to identify the generalizable attributes that characterize perceived effectiveness and inclusiveness in meetings. Additionally, we explore a predictive model of meeting effectiveness and inclusiveness based solely on objective meeting attributes. Lastly, we show challenges and discuss solutions around the subjective measurement of meeting experiences. To our knowledge, this is the largest data-driven study conducted after the pandemic peak to measure, understand, and predict effectiveness and inclusiveness in real-world meetings at an organizational scale. Yasaman Hosseinkashi, Lev Tankelevitch, Jamie Pool, Ross Cutler, Chinmaya Madan |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2023 | Don't Forget the User: It's Time to Rethink Network MeasurementsabstractNetwork measurement has long focused on the bits and bytes --- low-level network metrics such as latency and throughput, which have the advantage of being objective and directly characterizing the performance of the network. We argue that users provide a rich and largely untapped source of implicit as well as explicit signals that could complement and expand the coverage of traditional methods. Implicit feedback leverages user actions to indirectly infer the network performance and the resulting quality of user experience. Explicit feedback leverages user input, typically provided offline, to expand the reach of network measurement, especially for newer ones. Aryan Taneja, Rahul Bothra, Debopam Bhattacherjee, Rohan Gandhi, Venkat N. Padmanabhan, Ranjita Bhagwan, Nagarajan Natarajan, Saikat Guha 0002, Ross Cutler |
HotNets | 9 |
| 2023 | Real-Time Speech Interruption Analysis: from Cloud to Client DeploymentabstractMeetings are an essential form of communication for all types of organizations, and remote collaboration systems have been much more widely used since the COVID-19 pandemic. One major issue with remote meetings is that it is challenging for remote participants to interrupt and speak. We have recently developed the first speech interruption analysis model WavLM_SI, which detects failed speech interruptions, shows very promising performance, and is being deployed in the cloud. To deliver this feature in a more cost-efficient and environment-friendly way, we reduced the model complexity and size to ship the WavLM_SI model in client devices. In this paper, we first describe how we successfully improved the True Positive Rate (TPR) at a 1% False Positive Rate (FPR) from 50.9% to 68.3% for the failed speech interruption detection model by training on a larger dataset and fine-tuning. We then shrank the model size from 222.7 MB to 9.3 MB with an acceptable loss in accuracy and reduced the complexity from 31.2 GMACS (Giga Multiply-Accumulate Operations per Second) to 4.3 GMACS. We also estimated the environmental impact of the complexity reduction, which can be used as a general guideline for large Transformer-based models, and thus make those models more accessible with less computation overhead. Quchen Fu, Szu-Wei Fu, Yaran Fan, Yu Wu 0012, Zhuo Chen 0006, Jayant Gupchup, Ross Cutler |
ICASSP | 7 |
| 2023 | AURA: Privacy-Preserving Augmentation to Improve Test Set Diversity in Speech EnhancementabstractSpeech enhancement models running in production environments are commonly trained on publicly available data. This approach leads to regressions due to the lack of training/testing on representative customer data. Moreover, due to privacy reasons, developers cannot listen to customer content. This ‘ears-off’ situation motivates Aura, an end-to-end solution to make existing speech enhancement train and test sets more challenging and diverse while being sample efficient. Aura is ‘ears-off’ because it relies on a feature extractor and metrics of speech quality, DNSMOS P.835, and AECMOS, that are pre-trained on data obtained from public sources. We evalaute Aura on two speech enhancement tasks: noise suppression (NS) and audio echo cancellation (AEC). Aura samples an NS test set 0.42 harder in terms of P.835 OVRL than random sampling; and, an AEC test set 1.93 harder in AECMOS. Moreover, Aura increases diversity by 30% for NS tasks and by 530% for AEC tasks compared to greedy sampling. Moreover, Aura achieves a 26% improvement in Spearman’s rank correlation coefficient (SRCC) compared to random sampling when used to stack rank NS models. Xavier Gitiaux, Aditya Khant, Ebrahim Beyrami, Chandan K. A. Reddy, Jayant Gupchup, Ross Cutler |
ICASSP | 6 |
| 2023 | LSTM-Based Video Quality Prediction Accounting for Temporal Distortions in Videoconferencing CallsabstractCurrent state-of-the-art video quality models, such as VMAF, give excellent prediction results by comparing the degraded video with its reference video. However, they do not consider temporal distortions (e.g., frame freezes or skips) that occur during videoconferencing calls. In this paper, we present a data-driven approach for modeling such distortions automatically by training an LSTM with subjective quality ratings labeled via crowdsourcing. The videos were collected from live videoconferencing calls in 83 different network conditions. We applied QR codes as markers on the source videos to create aligned references and compute temporal features based on the alignment vectors. Using these features together with VMAF core features, our proposed model achieves a PCC of 0.99 on the validation set. Furthermore, our model outputs per-frame quality that gives detailed insight into the cause of video quality impairments. The VCM model and dataset are open-sourced at https://github.com/microsoft/Video_Call_MOS. Gabriel Mittag, Babak Naderi, Vishak Gopal, Ross Cutler |
ICASSP | 4 |
| 2023 | PLCMOS - A Data-driven Non-intrusive Metric for The Evaluation of Packet Loss Concealment Algorithms
Lorenz Diener, Marju Purin, Sten Sootla, Ando Saabas, Robert Aichner, Ross Cutler |
INTERSPEECH | 6 |
| 2023 | DeepVQE: Real Time Deep Voice Quality Enhancement for Joint Acoustic Echo Cancellation, Noise Suppression and Dereverberation
Nicolae-Catalin Ristea, Evgenii Indenbom, Ando Saabas, Tanel Pärnamaa, Jegor Guzvin, Ross Cutler |
INTERSPEECH | 6 |
| 2022 | ICASSP 2022 Acoustic Echo Cancellation ChallengeabstractThe ICASSP 2022 Acoustic Echo Cancellation Challenge is intended to stimulate research in acoustic echo cancellation (AEC), which is an important area of speech enhancement and still a top issue in audio communication. This is the third AEC challenge and it is enhanced by including mobile scenarios, adding speech recognition word accuracy rate as a metric, and making the audio 48 kHz. We open source two large datasets to train AEC models under both single talk and double talk scenarios. These datasets consist of recordings from more than 10,000 real audio devices and human speakers in real environments, as well as a synthetic dataset. We open source an online subjective test framework and provide an online objective metric service for researchers to quickly test their results. The winners of this challenge were selected based on the average Mean Opinion Score (MOS) achieved across all scenarios and the word accuracy rate. Ross Cutler, Ando Saabas, Tanel Pärnamaa, Marju Purin, Hannes Gamper, Sebastian Braun, Karsten Sørensen, Robert Aichner |
ICASSP | 1 |
| 2022 | Icassp 2022 Deep Noise Suppression ChallengeabstractThe Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. This is the 4th DNS challenge, with the previous editions held at INTERSPEECH 2020 [1], ICASSP 2021 [2], and INTERSPEECH 2021 [3]. We open-source datasets and test sets for researchers to train their deep noise suppression models, as well as a subjective evaluation framework based on ITU-T P.835 to rate and rank-order the challenge entries. We provide access to DNS-MOS P.835 and word accuracy (WAcc) APIs to challenge participants to help with iterative model improvements. In this challenge, we introduced the following changes: (i) Included mobile device scenarios in the blind test set; (ii) Included a personalized noise suppression track with baseline; (iii) Added WAcc as an objective metric; (iv) Included DNSMOS P.835; (v) Made the training datasets and test sets fullband (48 kHz). We use an average of WAcc and subjective scores P.835 SIG, BAK, and OVRL to get the final score for ranking the DNS models. We believe that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-world scenarios. Harishchandra Dubey, Vishak Gopal, Ross Cutler, Ashkan Aazami, Sergiy Matusevych, Sebastian Braun, Sefik Emre Eskimez, Manthan Thakker, Takuya Yoshioka, Hannes Gamper, Robert Aichner |
ICASSP | 3 |
| 2022 | AECMOS: A Speech Quality Assessment Metric for Echo ImpairmentabstractTraditionally, the quality of acoustic echo cancellers is evaluated using intrusive speech quality assessment measures such as ERLE [1] and PESQ [2], or by carrying out subjective laboratory tests [3], [4]. Unfortunately, the former are not well correlated with human subjective measures, while the latter are time and resource consuming to carry out [5]. We provide a new tool for speech quality assessment for echo impairment which can be used to evaluate the performance of acoustic echo cancellers. More precisely, we develop a neural network model to evaluate call quality degradations in two separate categories: echo and degradations from other sources. We show that our model is accurate as measured by correlation with human subjective quality ratings. Our tool can be used effectively to stack rank echo cancellation models. AECMOS is being made publicly available as an Azure service. Marju Purin, Sten Sootla, Mateja Sponza, Ando Saabas, Ross Cutler |
ICASSP | 5 |
| 2022 | Dnsmos P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise SuppressorsabstractHuman subjective evaluation is the "gold standard" to evaluate speech quality optimized for human perception. Perceptual objective metrics serve as a proxy for subjective scores. We have recently developed a non-intrusive speech quality metric called Deep Noise Suppression Mean Opinion Score (DNSMOS) using the scores from ITU-T Rec. P.808 [1] subjective evaluation. The P.808 scores reflect the overall quality of the audio clip. ITU-T Rec. P.835 [2] subjective evaluation framework gives the standalone quality scores of speech and background noise in addition to the overall quality. In this work, we train an objective metric based on P.835 human ratings that output 3 scores: i) speech quality (SIG), ii) background noise quality (BAK), and iii) the overall quality (OVRL) of the audio. The developed metric is highly correlated with human ratings, with a Pearson’s Correlation Co-efficient (PCC)=0.94 for SIG and PCC=0.98 for BAK and OVRL. This is the first non-intrusive P.835 predictor we are aware of. DNSMOS P.835 is made publicly available as an Azure service. Chandan K. A. Reddy, Vishak Gopal, Ross Cutler |
ICASSP | 3 |
| 2022 | INTERSPEECH 2022 Audio Deep Packet Loss Concealment ChallengeabstractAudio Packet Loss Concealment (PLC) is the hiding of gaps in audio streams caused by data transmission failures in packet switched networks.This is a common problem, and of increasing importance as end-to-end VoIP telephony and teleconference systems become the default and ever more widely used form of communication in business as well as in personal usage.This paper presents the INTERSPEECH 2022 Audio Deep Packet Loss Concealment challenge.We first give an overview of the PLC problem, and introduce some classical approaches to PLC as well as recent work.We then present the open source dataset released as part of this challenge as well as the evaluation methods and metrics used to determine the winner.We also briefly introduce PLCMOS, a novel data-driven metric that can be used to quickly evaluate the performance PLC systems.Finally, we present the results of the INTERSPEECH 2022 Audio Deep PLC Challenge, and provide a summary of important takeaways. Lorenz Diener, Sten Sootla, Solomiya Branets, Ando Saabas, Robert Aichner, Ross Cutler |
INTERSPEECH | 6 |
| 2022 | MusicNet: Compact Convolutional Neural Network for Real-time Background Music DetectionabstractWith the recent growth of remote work, online meetings often encounter challenging audio contexts such as background noise, music, and echo.Accurate real-time detection of music events can help to improve the user experience.In this paper, we present MusicNet, a compact neural model for detecting background music in the real-time communications pipeline.In video meetings, music frequently co-occurs with speech and background noises, making the accurate classification quite challenging.We propose a compact convolutional neural network core preceded by an in-model featurization layer.MusicNet takes 9 seconds of raw audio as input and does not require any model-specific featurization in the product stack.We train our model on the balanced subset of the Audio Set [1] data and validate it on 1000 crowd-sourced real test clips.Finally, we compare MusicNet performance with 20 state-of-the-art models.MusicNet has a true positive rate (TPR) of 81.3% at a 0.1% false positive rate (FPR), which is significantly better than state-of-the-art models included in our study.MusicNet is also 10x smaller and has 4x faster inference than the best performing models we benchmarked. Chandan K. A. Reddy, Vishak Gopal, Harishchandra Dubey, Ross Cutler, Sergiy Matusevych, Robert Aichner |
INTERSPEECH | 4 |
| 2022 | ConferencingSpeech 2022 Challenge: Non-intrusive Objective Speech Quality Assessment (NISQA) Challenge for Online Conferencing ApplicationsabstractWith the advances in speech communication systems such as online conferencing applications, we can seamlessly work with people regardless of where they are. However, during online meetings, speech quality can be significantly affected by background noise, reverberation, packet loss, network jitter, etc. Because of its nature, speech quality is traditionally assessed in subjective tests in laboratories and lately also in crowdsourcing following the international standards from ITU-T Rec. P.800 series. However, those approaches are costly and cannot be applied to customer data. Therefore, an effective objective assessment approach is needed to evaluate or monitor the speech quality of the ongoing conversation. The ConferencingSpeech 2022 challenge targets the non-intrusive deep neural network models for the speech quality assessment task. We open-sourced a training corpus with more than 86K speech clips in different languages, with a wide range of synthesized and live degradations and their corresponding subjective quality scores through crowdsourcing. 18 teams submitted their models for evaluation in this challenge. The blind test sets included about 4300 clips from wide ranges of degradations. This paper describes the challenge, the datasets, and the evaluation methods and reports the final results. Gaoxiong Yi, Babak Naderi, Sebastian Möller 0001, Wafaa Wardah, Gabriel Mittag, Ross Cutler, Zhuohuang Zhang, Donald S. Williamson, Fei Chen 0011, Shidong Shang |
INTERSPEECH | 8 |
| 2022 | Improving Meeting Inclusiveness using Speech Interruption AnalysisabstractMeetings are a pervasive method of communication within all types of companies and organizations, and using remote collaboration systems to conduct meetings has increased dramatically since the COVID-19 pandemic. However, not all meetings are inclusive, especially in terms of the participation rates among attendees. In a recent large-scale survey conducted at Microsoft, the top suggestion given by meeting participants for improving inclusiveness is to improve the ability of remote participants to interrupt and acquire the floor during meetings. We show that the use of the virtual raise hand (VRH) feature can lead to an increase in predicted meeting inclusiveness at Microsoft. One challenge is that VRH is used in less than $1%$ of all meetings. In order to drive adoption of its usage to improve inclusiveness (and participation), we present a machine learning-based system that predicts when a meeting participant attempts to obtain the floor, but fails to interrupt (termed a 'failed interruption'). This prediction can be used to nudge the user to raise their virtual hand within the meeting. We believe this is the first failed speech interruption detector, and the performance on a realistic test set has an area under curve (AUC) of 0.95 with a true positive rate (TPR) of 50% at a false positive rate (FPR) of 1%. To our knowledge, this is also the first dataset of interruption categories (including the failed interruption category) for remote meetings. Finally, we believe this is the first such system designed to improve meeting inclusiveness through speech interruption analysis and active intervention. Szu-Wei Fu, Yaran Fan, Yasaman Hosseinkashi, Jayant Gupchup, Ross Cutler |
ACM Multimedia | 5 |
| 2022 | Performance optimizations on U-Net speech enhancement modelsabstractDeep learning approaches-while remarkably successful in audio enhancement-result in slower inference times which are prohibitive in real-time applications. We develop a practical strategy for compressing U-Net style deep neural network architectures. On deep noise suppression (DNS) models we achieve a state-of-the-art 7.25× inference speed up over the baseline CRUSE model, with a smooth model performance degradation. Our method is friendly to practitioners as it only requires setting a single compression parameter, while achieving non-uniform compression rates across layers. We report inference speed because a parameter or memory reduction does not necessitate speedup, and we measure model quality using an accurate non-intrusive objective speech quality metric. Jerry Chee, Sebastian Braun, Vishak Gopal, Ross Cutler |
MMSP | 4 |
| 2021 | Crowdsourcing Approach for Subjective Evaluation of Echo ImpairmentabstractThe quality of acoustic echo cancellers (AECs) in real-time communication systems is typically evaluated using objective metrics like ERLE [1] and PESQ [2], and less commonly with lab-based subjective tests like ITU-T Rec. P.831 [3]. We will show that these objective measures are not well correlated to subjective measures. We then introduce an open-source crowdsourcing approach for subjective evaluation of echo impairment which can be used to evaluate the performance of AECs. We provide a study that shows this tool is highly reproducible. This new tool has been recently used in the ICASSP 2021 AEC Challenge [4] which made the challenge possible to do quickly and cost effectively. Ross Cutler, Babak Naderi, Markus Loide, Sten Sootla, Ando Saabas |
ICASSP | 1 |
| 2021 | ICASSP 2021 Deep Noise Suppression ChallengeabstractThe Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH 2020 where we open-sourced training and test datasets for researchers to train their noise suppression models. We also open-sourced a subjective evaluation framework and used the tool to evaluate and select the final winners. Many researchers from academia and industry made significant contributions to push the field forward. We also learned that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-time conditions. In this challenge, we expanded both our training and test datasets. Clean speech in the training set has increased by 200% with the addition of singing voice, emotion data, and non-English languages. The test set has increased by 100% with the addition of singing, emotional, non-English (tonal and non-tonal) languages, and, personalized DNS test clips. There are two tracks with focus on (i) real-time denoising, and (ii) real-time personalized DNS. We present the challenge results at the end. Chandan K. A. Reddy, Harishchandra Dubey, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan 0003 |
ICASSP | 4 |
| 2021 | Dnsmos: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise SuppressorsabstractHuman subjective evaluation is the "gold standard" to evaluate speech quality optimized for human perception. Perceptual objective metrics serve as a proxy for subjective scores. The conventional and widely used metrics require a reference clean speech signal, which is unavailable in real recordings. Previous no-reference approaches correlate poorly with human ratings and are not widely adopted in the research community. One of the biggest use cases of these perceptual objective metrics is to evaluate noise suppression algorithms. This paper introduces a multi-stage self-teaching based perceptual objective metric that is designed to evaluate noise suppressors. The proposed method generalizes well in challenging test conditions with a high correlation to human ratings. Chandan K. A. Reddy, Vishak Gopal, Ross Cutler |
ICASSP | 3 |
| 2021 | ICASSP 2021 Acoustic Echo Cancellation Challenge: Datasets, Testing Framework, and ResultsabstractThe ICASSP 2021 Acoustic Echo Cancellation Challenge is intended to stimulate research in the area of acoustic echo cancellation (AEC), which is an important part of speech enhancement and still a top issue in audio communication and conferencing systems. Many recent AEC studies report good performance on synthetic datasets where the train and test samples come from the same underlying distribution. However, the AEC performance often degrades significantly on real recordings. Also, most of the conventional objective metrics such as echo return loss enhancement (ERLE) and perceptual evaluation of speech quality (PESQ) do not correlate well with subjective speech quality tests in the presence of background noise and reverberation found in realistic environments. In this challenge, we open source two large datasets to train AEC models under both single talk and double talk scenarios. These datasets consist of recordings from more than 2,500 real audio devices and human speakers in real environments, as well as a synthetic dataset. We open source two large test sets, and we open source an online subjective test framework for researchers to quickly test their results. The winners of this challenge will be selected based on the average Mean Opinion Score (MOS) achieved across all different single talk and double talk scenarios. Kusha Sridhar, Ross Cutler, Ando Saabas, Tanel Pärnamaa, Markus Loide, Hannes Gamper, Sebastian Braun, Robert Aichner, Sriram Srinivasan 0003 |
ICASSP | 2 |
| 2021 | INTERSPEECH 2021 Acoustic Echo Cancellation Challenge
Ross Cutler, Ando Saabas, Tanel Pärnamaa, Markus Loide, Sten Sootla, Marju Purin, Hannes Gamper, Sebastian Braun, Karsten Sørensen, Robert Aichner, Sriram Srinivasan 0003 |
Interspeech | 1 |
| 2021 | Subjective Evaluation of Noise Suppression Algorithms in CrowdsourcingabstractThe quality of the speech communication systems, which include noise suppression algorithms, are typically evaluated in laboratory experiments according to the ITU-T Rec. P.835, in which participants rate background noise, speech signal, and overall quality separately. This paper introduces an open-source toolkit for conducting subjective quality evaluation of noise suppressed speech in crowdsourcing. We followed the ITU-T Rec. P.835, and P.808 and highly automate the process to prevent moderator's error. To assess the validity of our evaluation method, we compared the Mean Opinion Scores (MOS), calculate using ratings collected with our implementation, and the MOS values from a standard laboratory experiment conducted according to the ITU-T Rec P.835. Results show a high validity in all three scales namely background noise, speech signal and overall quality (average PCC = 0.961). Results of a round-robin test (N=5) showed that our implementation is also a highly reproducible evaluation method (PCC=0.99). Finally, we used our implementation in the INTERSPEECH 2021 Deep Noise Suppression Challenge as the primary evaluation metric, which demonstrates it is practical to use at scale. The results are analyzed to determine why the overall performance was the best in terms of background noise and speech quality. Babak Naderi, Ross Cutler |
Interspeech | 2 |
| 2021 | INTERSPEECH 2021 Deep Noise Suppression ChallengeabstractThe Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH and ICASSP 2020. We open-sourced training and test datasets for the wideband scenario. We also open-sourced a subjective evaluation framework based on ITU-T standard P.808, which was also used to evaluate participants of the challenge. Many researchers from academia and industry made significant contributions to push the field forward, yet even the best noise suppressor was far from achieving superior speech quality in challenging scenarios. In this version of the challenge organized at INTERSPEECH 2021, we are expanding both our training and test datasets to accommodate full band scenarios. The two tracks in this challenge will focus on real-time denoising for (i) wide band, and(ii) full band scenarios. We are also making available a reliable non-intrusive objective speech quality metric called DNSMOS for the participants to use during their development phase. Chandan K. A. Reddy, Harishchandra Dubey, Kazuhito Koishida, Arun Asokan Nair, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan 0003 |
Interspeech | 6 |
| 2021 | Speech Quality Assessment in Crowdsourcing: Comparison Category Rating MethodabstractTraditionally, Quality of Experience (QoE) for a communication system is evaluated through a subjective test. The most common test method for speech QoE is the Absolute Category Rating (ACR), in which participants listen to a set of stimuli, processed by the underlying test conditions, and rate their perceived quality for each stimulus on a specific scale. The Comparison Category Rating (CCR) is another standard approach in which participants listen to both reference and processed stimuli and rate their quality compared to the other one. The CCR method is particularly suitable for systems that improve the quality of input speech. This paper evaluates an adaptation of the CCR test procedure for assessing speech quality in the crowdsourcing set-up. The CCR method was introduced in the ITU-T Rec. P.800 for laboratory-based experiments. We adapted the test for the crowdsourcing approach following the guidelines from ITU-T Rec. P.800 and P.808. We show that the results of the CCR procedure via crowdsourcing are highly reproducible. We also compared the CCR test results with widely used ACR test procedures obtained in the laboratory and crowdsourcing. Our results show that the CCR procedure in crowdsourcing is a reliable and valid test method. Babak Naderi, Sebastian Möller 0001, Ross Cutler |
QoMEX | 3 |
| 2021 | Meeting Effectiveness and Inclusiveness in Remote CollaborationabstractA primary goal of remote collaboration tools is to provide effective and inclusive meetings for all participants. To study meeting effectiveness and meeting inclusiveness, we first conducted a large-scale email survey (N=4,425; after filtering N=3,290) at a large technology company (pre-COVID-19); using this data we derived a multivariate model of meeting effectiveness and show how it correlates with meeting inclusiveness, participation, and feeling comfortable to contribute. We believe this is the first such model of meeting effectiveness and inclusiveness. The large size of the data provided the opportunity to analyze correlations that are specific to sub-populations such as the impact of video. The model shows the following factors are correlated with inclusiveness, effectiveness, participation, and feeling comfortable to contribute in meetings: sending a pre-meeting communication, sending a post-meeting summary, including a meeting agenda, attendee location, remote-only meeting, audio/video quality and reliability, video usage, and meeting size. The model and survey results give a quantitative understanding of how and where to improve meeting effectiveness and inclusiveness and what the potential returns are. Motivated by the email survey results, we implemented a post-meeting survey into a leading computer-mediated communication (CMC) system to directly measure meeting effectiveness and inclusiveness (during COVID-19). Using initial results based on internal flighting we created a similar model of effectiveness and inclusiveness, with many of the same findings as the email survey. This shows a method of measuring and understanding these metrics which are both practical and useful in a commercial CMC system. By improving meeting effectiveness, companies can save significant time and money. Improving meeting inclusiveness is hypothesized to improve meeting effectiveness, but also improves the working environment and employee retention at organizations. Ross Cutler, Yasaman Hosseinkashi, Jamie Pool, Senja Filipi, Robert Aichner, Yuan Tu, Johannes Gehrke |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2020 | Multimodal Active Speaker Detection and Virtual Cinematography for Video ConferencingabstractActive speaker detection (ASD) and virtual cinematography (VC) can significantly improve the experience of a video conference by automatically panning, tilting and zooming of a camera: subjectively users rate an expert video cinematographer significantly higher than the unedited video. We describe a new automated ASD and VC that performs within 0.3 MOS of an expert cinematographer based on subjective ratings with a 1-5 scale. This system uses a 4K wide-FOV camera, a depth camera, and a microphone array, extracts features from each modality and trains an ASD using an AdaBoost machine learning system that is very efficient and runs in real-time. A VC is similarly trained using machine learning. To avoid distracting the room participants the system has no moving parts - the VC works by cropping and zooming the 4K wide-FOV video stream. The system was tuned and evaluated using extensive crowdsourcing techniques and evaluated on a system with N=100 meetings, each 25 minutes in length. Ross Cutler, Ramin Mehran, Sam Johnson, Cha Zhang, Adam Kirk, Oliver Whyte, Adarsh Kowdle |
ICASSP | 1 |
| 2020 | Weighted Speech Distortion Losses for Neural-Network-Based Real-Time Speech EnhancementabstractThis paper investigates several aspects of training a RNN (recurrent neural network) that impact the objective and subjective quality of enhanced speech for real-time single-channel speech enhancement. Specifically, we focus on a RNN that enhances short-time speech spectra on a single-frame-in, single-frame-out basis, a framework adopted by most classical signal processing methods. We propose two novel mean-squared-error-based learning objectives that enable separate control over the importance of speech distortion versus noise reduction. The proposed loss functions are evaluated by widely accepted objective quality and intelligibility measures and compared to other competitive online methods. In addition, we study the impact of feature normalization and varying batch sequence lengths on the objective quality of enhanced speech. Finally, we show subjective ratings for the proposed approach and a state-of-the-art real-time RNN-based method. Yangyang Xia, Sebastian Braun, Chandan K. A. Reddy, Harishchandra Dubey, Ross Cutler, Ivan Tashev |
ICASSP | 5 |
| 2020 | DNN No-Reference PSTN Speech Quality PredictionabstractClassic public switched telephone networks (PSTN) are often a black box for VoIP network providers, as they have no access to performance indicators, such as delay or packet loss. Only the degraded output speech signal can be used to monitor the speech quality of these networks. However, the current state-of-the-art speech quality models are not reliable enough to be used for live monitoring. One of the reasons for this is that PSTN distortions can be unique depending on the provider and country, which makes it difficult to train a model that generalizes well for different PSTN networks. In this paper, we present a new open-source PSTN speech quality test set with over 1000 crowdsourced real phone calls. Our proposed no-reference model outperforms the full-reference POLQA and no-reference P.563 on the validation and test set. Further, we analyzed the influence of file cropping on the perceived speech quality and the influence of the number of ratings and training size on the model accuracy. Gabriel Mittag, Ross Cutler, Yasaman Hosseinkashi, Michael Revow, Sriram Srinivasan 0003, Naglakshmi Chande, Robert Aichner |
INTERSPEECH | 2 |
| 2020 | An Open Source Implementation of ITU-T Recommendation P.808 with ValidationabstractThe ITU-T Recommendation P.808 provides a crowdsourcing approach for conducting a subjective assessment of speech quality using the Absolute Category Rating (ACR) method. We provide an open-source implementation of the ITU-T Rec. P.808 that runs on the Amazon Mechanical Turk platform. We extended our implementation to include Degradation Category Ratings (DCR) and Comparison Category Ratings (CCR) test methods. We also significantly speed up the test process by integrating the participant qualification step into the main rating task compared to a two-stage qualification and rating solution. We provide program scripts for creating and executing the subjective test, and data cleansing and analyzing the answers to avoid operational errors. To validate the implementation, we compare the Mean Opinion Scores (MOS) collected through our implementation with MOS values from a standard laboratory experiment conducted based on the ITU-T Rec. P.800. We also evaluate the reproducibility of the result of the subjective speech quality assessment through crowdsourcing using our implementation. Finally, we quantify the impact of parts of the system designed to improve the reliability: environmental tests, gold and trapping questions, rating patterns, and a headset usage test. Babak Naderi, Ross Cutler |
INTERSPEECH | 2 |
| 2020 | The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge ResultsabstractThe INTERSPEECH 2020 Deep Noise Suppression (DNS) Challenge is intended to promote collaborative research in real-time single-channel Speech Enhancement aimed to maximize the subjective (perceptual) quality of the enhanced speech. A typical approach to evaluate the noise suppression methods is to use objective metrics on the test set obtained by splitting the original dataset. While the performance is good on the synthetic test set, often the model performance degrades significantly on real recordings. Also, most of the conventional objective metrics do not correlate well with subjective tests and lab subjective tests are not scalable for a large test set. In this challenge, we open-sourced a large clean speech and noise corpus for training the noise suppression models and a representative test set to real-world scenarios consisting of both synthetic and real recordings. We also open-sourced an online subjective test framework based on ITU-T P.808 for researchers to reliably test their developments. We evaluated the results using P.808 on a blind test set. The results and the key learnings from the challenge are discussed. The datasets and scripts can be found here for quick access https://github.com/microsoft/DNS-Challenge. Chandan K. A. Reddy, Vishak Gopal, Ross Cutler, Ebrahim Beyrami, Roger Cheng, Harishchandra Dubey, Sergiy Matusevych, Robert Aichner, Ashkan Aazami, Sebastian Braun, Puneet Rana, Sriram Srinivasan 0003, Johannes Gehrke |
INTERSPEECH | 3 |
| 2020 | Lumos: A Library for Diagnosing Metric Regressions in Web-Scale ApplicationsabstractWeb-scale applications can ship code on a daily to weekly cadence. These applications rely on online metrics to monitor the health of new releases. Regressions in metric values need to be detected and diagnosed as early as possible to reduce the disruption to users and product owners. Regressions in metrics can surface due to a variety of reasons: genuine product regressions, changes in user population and bias due to telemetry loss (or processing) are among the common causes. Diagnosing the cause of these metric regressions is costly for engineering teams as they need to invest time in finding the root cause of the issue as soon as possible. We presentLumos, a Python library built using the principles of A/B testing to systematically diagnose metric regressions to automate such analysis.Lumos has been deployed across the component teams in Microsoft's Real-Time Communication (RTC) applications Skype and Microsoft Teams. It has enabled engineering teams to detect 100s of real changes in metrics and reject 1000s of false alarms detected by anomaly detectors. The application ofLumos has resulted in freeing up as much as $95%$ of the time allocated to metric-based investigations. In this work, we open sourceLumos and present our results from applying it to two different components within the RTC group over millions of sessions. This general library can be coupled with any production system to manage the volume of alerting efficiently. Jamie Pool, Ebrahim Beyrami, Vishak Gopal, Ashkan Aazami, Jayant Gupchup, Jeff Rowland, Binlong Li, Pritesh Kanani, Ross Cutler, Johannes Gehrke |
KDD | 9 |
| 2019 | Non-intrusive Speech Quality Assessment Using Neural NetworksabstractEstimating the perceived quality of an audio signal is critical for many multimedia and audio processing systems. Providers strive to offer optimal and reliable services in order to increase the user quality of experience (QoE). In this work, we present an investigation of the applicability of neural networks for non-intrusive audio quality assessment. We propose three neural network-based approaches for mean opinion score (MOS) estimation. We compare our results to three instrumental measures: the perceptual evaluation of speech quality (PESQ), the ITU-T Recommendation P.563, and the speech-to-reverberation energy ratio. Our evaluation uses a speech dataset contaminated with convolutive and additive noise, labeled using a crowd-based QoE evaluation, evaluated with Pearson correlation with MOS labels, and mean-squared-error of the estimated MOS. Our proposed approaches outperform the aforementioned instrumental measures, with a fully connected deep neural network using Mel-frequency features providing the best correlation (0.87) and the lowest mean squared error (0.15). Anderson R. Avila, Hannes Gamper, Chandan K. A. Reddy, Ross Cutler, Ivan Tashev, Johannes Gehrke |
ICASSP | 4 |
| 2019 | A Scalable Noisy Speech Dataset and Online Subjective Test FrameworkabstractBackground noise is a major source of quality impairments in Voice over Internet Protocol (VoIP) and Public Switched Telephone Network (PSTN) calls.Recent work shows the efficacy of deep learning for noise suppression, but the datasets have been relatively small compared to those used in other domains (e.g., ImageNet) and the associated evaluations have been more focused.In order to better facilitate deep learning research in Speech Enhancement, we present a noisy speech dataset (MS-SNSD) that can scale to arbitrary sizes depending on the number of speakers, noise types, and Speech to Noise Ratio (SNR) levels desired.We show that increasing dataset sizes increases noise suppression performance as expected.In addition, we provide an open-source evaluation methodology to evaluate the results subjectively at scale using crowdsourcing, with a reference algorithm to normalize the results.To demonstrate the dataset and evaluation framework we apply it to several noise suppressors and compare the subjective Mean Opinion Score (MOS) with objective quality measures such as SNR, PESQ, POLQA, and VISQOL and show why MOS is still required.Our subjective MOS evaluation is the first large scale evaluation of Speech Enhancement algorithms that we are aware of. Chandan K. A. Reddy, Ebrahim Beyrami, Jamie Pool, Ross Cutler, Sriram Srinivasan 0003, Johannes Gehrke |
INTERSPEECH | 4 |
| 2019 | Supervised Classifiers for Audio Impairments with Noisy LabelsabstractVoice-over-Internet-Protocol (VoIP) calls are prone to various speech impairments due to environmental and network conditions resulting in bad user experience. A reliable audio impairment classifier helps to identify the cause for bad audio quality. The user feedback after the call can act as the ground truth labels for training a supervised classifier on a large audio dataset. However, the labels are noisy as most of the users lack the expertise to precisely articulate the impairment in the perceived speech. In this paper, we analyze the effects of massive noise in labels in training dense networks and Convolutional Neural Networks (CNN) using engineered features, spectrograms and raw audio samples as inputs. We demonstrate that CNN can generalize better on the training data with a large number of noisy labels and gives remarkably higher test performance. The classifiers were trained both on randomly generated label noise and the label noise introduced by human errors. We also show that training with noisy labels requires a significant increase in the training dataset size, which is in proportion to the amount of noise in the labels. Chandan K. A. Reddy, Ross Cutler, Johannes Gehrke |
INTERSPEECH | 2 |
| 2018 | Trustworthy Experimentation Under Telemetry LossabstractFailure to accurately measure the outcomes of an experiment can lead to bias and incorrect conclusions. Online controlled experiments (aka AB tests) are increasingly being used to make decisions to improve websites as well as mobile and desktop applications. We argue that loss of telemetry data (during upload or post-processing) can skew the results of experiments, leading to loss of statistical power and inaccurate or erroneous conclusions. By systematically investigating the causes of telemetry loss, we argue that it is not practical to entirely eliminate it. Consequently, experimentation systems need to be robust to its effects. Furthermore, we note that it is nontrivial to measure the absolute level of telemetry loss in an experimentation system. In this paper, we take a top-down approach towards solving this problem. We motivate the impact of loss qualitatively using experiments in real applications deployed at scale, and formalize the problem by presenting a theoretical breakdown of the bias introduced by loss. Based on this foundation, we present a general framework for quantitatively evaluating the impact of telemetry loss, and present two solutions to measure the absolute levels of loss. This framework is used by well-known applications at Microsoft, with millions of users and billions of sessions. These general principles can be adopted by any application to improve the overall trustworthiness of experimentation and data-driven decision making. Jayant Gupchup, Yasaman Hosseinkashi, Pavel A. Dmitriev, Ross Cutler, Andrei Jefremov, Martin Ellis |
CIKM | 5 |
| 2018 | On Design of Problem Token Questions in Quality of Experience SurveysabstractUser surveys for Quality of Experience (QoE) are a critical source of information for application developers. In addition to the common “star rating” used to estimate Mean Opinion Score (MOS), more detailed survey questions (problem tokens) about specific areas provide valuable insight into the factors impacting QoE. This paper explores two aspects of problem token questionnaire design. First, we study the bias introduced by fixed question order, and second, we provide a methodology to manage the size of the survey while keeping it informative. Based on 900,000 calls gathered using a randomized controlled experiment from Skype, we find that token selections can be strongly biased due to token positions and display design. This selection bias can be significantly reduced by randomizing the display order of tokens. It is worth noting that users respond to the randomized-order variant at levels that are comparable to the fixed-order variant. The effective selection of a subset of tokens is achieved by extracting tokens that provide the highest information gain over user ratings. This selection is known to be in the class of NP-hard problems. We apply a well-known greedy submodular maximization method on our dataset to capture 94% of the information using just 30% of thequestions. Jayant Gupchup, Ebrahim Beyrami, Martin Ellis, Yasaman Hosseinkashi, Sam Johnson, Ross Cutler |
QoMEX | 6 |
| 2017 | Analysis of problem tokens to rank factors impacting quality in VoIP applicationsabstractUser-perceived quality-of-experience (QoE) in internet telephony systems is commonly evaluated using subjective ratings computed as a Mean Opinion Score (MOS). In such systems, while user MOS can be tracked on an ongoing basis, it does not give insight into which factors of a call induced any perceived degradation in QoE - it does not tell us what caused a user to have a sub-optimal experience. For effective planning of product improvements, we are interested in understanding the impact of each of these degrading factors, allowing the estimation of the return (i.e., the improvement in user QoE) for a given investment. To obtain such insights, we advocate the use of an end-of-call “problem token questionnaire” (PTQ) which probes the user about common call quality issues (e.g., distorted audio or frozen video) which they may have experienced. In this paper, we show the efficacy of this questionnaire using data gathered from over 700,000 end-of-call surveys gathered from Skype (a large commercial VoIP application). We present a method to rank call quality and reliability issues and address the challenge of isolating independent factors impacting the QoE. Finally, we present representative examples of how these problem tokens have proven to be useful in practice. Jayant Gupchup, Yasaman Hosseinkashi, Martin Ellis, Sam Johnson, Ross Cutler |
QoMEX | 5 |
| 2010 | SureCall: Towards glitch-free real-time audio/video conferencingabstractGlobal enterprises are increasingly adopting unified communication solutions over traditional telephone systems. Such solutions provide integrated audio/video conferencing and messaging services, and enable flexible working environments by allowing mobile and dispersed users to communicate and collaborate easily and efficiently. The ultimate goal of unified communications is to ensure a smooth and best possible user experience across all scenarios. To address this challenge and understand the impact of various network scenarios on unified audio/video conferencing, we have developed a distributed experimental platform - SureCall - and deployed it on over 80 machines across a global enterprise and many residential networks. SureCall has collected worth of more than 6 months of packet-level audio/video conferencing traces. Through in-depth analysis of these traces, we have quantitatively compared how key performance metrics, such as packet loss and jitter, as well as the correlation between them, are affected by the enterprise and residential networks, by WiFi connections and VPN links, etc. In addition, we show how SureCall can serve as an ideal platform to design, experiment and validate new schemes and algorithms. We have developed a new audio quality classifier using the SureCall platform, which is being experimented with the recent release of Office Communicator solution for large-scale validation. Amit Mondal, Ross Cutler, Cheng Huang 0002, Jin Li 0001, Aleksandar Kuzmanovic |
IWQoS | 2 |
| 2008 | Boosting-Based Multimodal Speaker Detection for Distributed Meeting VideosabstractIdentifying the active speaker in a video of a distributed meeting can be very helpful for remote participants to understand the dynamics of the meeting. A straightforward application of such analysis is to stream a high resolution video of the speaker to the remote participants. In this paper, we present the challenges we met while designing a speaker detector for the Microsoft RoundTable distributed meeting device, and propose a novel boosting-based multimodal speaker detection (BMSD) algorithm. Instead of separately performing sound source localization (SSL) and multiperson detection (MPD) and subsequently fusing their individual results, the proposed algorithm fuses audio and visual information at feature level by using boosting to select features from a combined pool of both audio and visual features simultaneously. The result is a very accurate speaker detector with extremely high efficiency. In experiments that includes hundreds of real-world meetings, the proposed BMSD algorithm reduces the error rate of SSL-only approach by 24.6%, and the SSL and MPD fusion approach by 20.9%. To the best of our knowledge, this is the first real-time multimodal speaker detection algorithm that is deployed in commercial products. Cha Zhang, Pei Yin, Yong Rui, Ross Cutler, Paul A. Viola, Xinding Sun, Nelson Pinto, Zhengyou Zhang |
IEEE Trans. Multim. | 4 |
| 2007 | Head-Size Equalization for Improved Visual Perception in Video ConferencingabstractIn a video conferencing setting, people often use an elongated meeting table with the major axis along the camera direction. A standard wide-angle perspective image of this setting creates significant foreshortening, thus the people sitting at the far end of the table appear very small relative to those nearer the camera. This has two consequences. First, it is difficult for the remote participants to see the faces of those at the far end, thus affecting the experience of the video conferencing. Second, it is a waste of the screen space and network bandwidth because most of the pixels are used on the background instead of on the faces of the meeting participants. In this paper, we present a novel technique, called Spatially-Varying-Uniform scaling functions, to warp the images to equalize the head sizes of the meeting participants without causing undue distortion. This technique works for both the 180-degree views where the camera is placed at one end of the table and the 360-degree views where the camera is placed at the center of the table. We have implemented this algorithm on two types of camera arrays: one with 180-degree view, and the other with 360-degree view. On both hardware devices, image capturing, stitching, and head-size equalization are run in real time. In addition, we have conducted user study showing that people clearly prefer head-size equalized images. Zicheng Liu 0001, Michael F. Cohen, Deepti Bhatnagar, Ross Cutler, Zhengyou Zhang |
IEEE Trans. Multim. | 4 |
| 2006 | Boosting-Based Multimodal Speaker Detection for Distributed MeetingsabstractSpeaker detection is a very important task in distributed meeting applications. This paper discusses a number of challenges we met while designing a speaker detector for the Microsoft RoundTable distributed meeting device, and proposes a boosting-based multimodal speaker detection (BMSD) algorithm. Instead of performing sound source localization (SSL) and multi-person detection (MPD) separately and subsequently fusing their individual results, the proposed algorithm uses boosting to select features from a combined pool of both audio and visual features simultaneously. The result is a very accurate speaker detector with extremely high efficiency. The algorithm reduces the error rate of SSL-only approach by 47%, and the SSL and MPD fusion approach by 27% Cha Zhang, Pei Yin, Yong Rui, Ross Cutler, Paul A. Viola |
MMSP | 4 |
| 2005 | Automatic Head-size Equalization in Panorama Images for Video ConferencingabstractIn panorama images captured by omni-directional cameras during video conferencing, the image sizes of the people around the conference table are not uniform due to the varying distances to the camera. Spatially varying-uniform (SVU) scaling functions have been proposed to warp a panorama image smoothly such that the participants have similar sizes on the image. To generate the SVU function, one needs to segment the table boundaries, which was generated manually in the previous work. In this paper, we propose a robust algorithm to automatically segment the table boundaries. To ensure the robustness, we apply a symmetry voting scheme to filter out noisy points on the edge map. Trigonometry and quadratic fitting methods are developed to fit a continuous curve to the remaining edge points. We report experimental results on both synthetic and real images. Ya Chang, Ross Cutler, Zicheng Liu 0001, Zhengyou Zhang, Alex Acero, Matthew Turk 0001 |
ICME | 2 |
| 2005 | Quality Assessment of Panorama Video for Videoconferencing ApplicationsabstractNew video-conference devices based on omnidirectional multi-camera systems have been emerging in the last few years. These devices require innovative and automated video quality assessment in the earlier stages of their design in order to guarantee competitive product development and quality monitoring. Current quality assessment techniques are not adequate since they are mostly tailored to single video cameras. Even if these techniques are capable of assessing the quality of each video stream separately, the overall quality of a composite video stream generated with the outputs of multiple cameras stitched together presents strong deviations from the results of subjective quality tests. In this paper, we present new strategies for assessing the quality of composite video streams with specific emphasis on the following problems: noticeable calibration differences between adjacent cameras, concentration of motion in limited regions of the panoramic scene, combined vignetting problems, non-uniformity of the surfaces in the seam region between two adjacent cameras. Combining different features from high and low level vision, we evaluate the proposed perceptual quality metric using a set of customized test sequences and verify its correlation with subjective quality tests Simone Leorin, Luca Lucchese, Ross Cutler |
MMSP | 3 |
| 2004 | High-quality linear interpolation for demosaicing of Bayer-patterned color imagesabstractThis paper introduces a new interpolation technique for demosaicing of color images produced by single-CCD digital cameras. We show that the proposed simple linear filter can lead to an improvement in PSNR of over 5.5 dB when compared to bilinear demosaicing, and about 0.7 dB improvement in R and B interpolation when compared to a recently introduced linear interpolator. The proposed filter also outperforms most nonlinear demosaicing algorithms, without the artifacts due to nonlinear processing, and a much reduced computational complexity. Henrique S. Malvar, Li-wei He, Ross Cutler |
ICASSP (3) | 3 |
| 2003 | The distributed meetings systemabstractMeetings are an integral part of everyday life for most work-groups. However, due to travel, time, or other constraints, people are often not able to attend all the meetings they need to. Teleconferencing and recording of meetings can address this problem. We describe a system that provides these features, as well as a user study evaluation of the system. The system uses a variety of capture devices (a novel 360/spl deg/ camera, a whiteboard camera, an overview camera, PC graphics capture, and a microphone array) to provide a rich experience for people who want to participate in a meeting from a distance. The system is also combined with speaker clustering, spatial indexing, and time compression to provide a rich experience for people who miss a meeting and want to watch it afterward. Ross Cutler |
ICASSP (4) | 1 |
| 2002 | Distributed meetings: a meeting capture and broadcasting systemabstractThe common meeting is an integral part of everyday life for most workgroups. However, due to travel, time, or other constraints, people are often not able to attend all the meetings they need to. Teleconferencing and recording of meetings can address this problem. In this paper we describe a system that provides these features, as well as a user study evaluation of the system. The system uses a variety of capture devices (a novel 360° camera, a whiteboard camera, an overview camera, and a microphone array) to provide a rich experience for people who want to participate in a meeting from a distance. The system is also combined with speaker clustering, spatial indexing, and time compression to provide a rich experience for people who miss a meeting and want to watch it afterward. Ross Cutler, Yong Rui, Anoop Gupta, Jonathan J. Cadiz, Ivan Tashev, Li-wei He, Alex Colburn, Zhengyou Zhang, Zicheng Liu 0001, Steve Silverberg |
ACM Multimedia | 1 |
| 2001 | Backpack: Detection of People Carrying Objects Using Silhouettes
Ismail Haritaoglu, Ross Cutler, David Harwood, Larry Davis 0001 |
Comput. Vis. Image Underst. | 2 |
| 2000 | Robust Periodic Motion and Motion Symmetry DetectionabstractWe describe a robust technique for detecting nonstationary periodic motion from a moving and static camera. We also describe a robust technique for discriminating motion symmetries (periodic motion classification), which we apply to classifying running humans (bipeds) and canines (quadrupeds). The system has been implemented to run in real-time (30 Hz) on standard PC workstations. Ross Cutler, Larry Davis 0001 |
CVPR | 1 |
| 2000 | Robust Real-Time Periodic Motion Detection, Analysis, and ApplicationsabstractWe describe new techniques to detect and analyze periodic motion as seen from both a static and a moving camera. By tracking objects of interest, we compute an object's self-similarity as it evolves in time. For periodic motion, the self-similarity measure is also periodic and we apply time-frequency analysis to detect and characterize the periodic motion. The periodicity is also analyzed robustly using the 2D lattice structures inherent in similarity matrices. A real-time system has been implemented to track and classify objects using periodicity. Examples of object classification (people, running dogs, vehicles), person counting, and nonstationary periodicity are provided. Ross Cutler, Larry Davis 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1999 | Real-Time Periodic Motion Detection, Analysis, and ApplicationsabstractWe describe a new technique to detect and analyze periodic motion as seen from both a static and moving camera. By tracking objects of interest, we compute an object's self-similarity as it evolves in time. For periodic motion, the self-similarity measure is also periodic, and we apply time-frequency analysis to detect and characterize the periodic motion. A real-time system has been implemented to track and classify objects using periodicity. Examples of object classification, person counting, and non-stationary periodicity are provided. Ross Cutler, Larry Davis 0001 |
CVPR | 1 |
| 1999 | Backpack: Detection of People Carrying Objects using SilhouettesabstractWe described a video-rate surveillance algorithm to detect and track people from a stationary camera, and to determine if they are carrying objects or moving unencumbered. The contribution of the paper is the shape analysis algorithm that both determines if a person is carrying an object and segments the object from the person so that it can be tracked, e.g., during an exchange of objects between two people. As the object is segmented an appearance model of the object is constructed. The method combines periodic motion estimation with static symmetry analysis of the silhouettes of a person in each frame of the sequence. Experimental results demonstrate robustness and real-time performance of the proposed algorithm. Ismail Haritaoglu, Ross Cutler, David Harwood, Larry Davis 0001 |
ICCV | 2 |
| 1998 | View-Based Interpretation of Real-Time Optical Flow for Gesture Recognition
Ross Cutler, Matthew Turk 0001 |
FG | 1 |
| 1998 | View-based detection and analysis of periodic motionabstractWe describe a technique that detects periodic motion. Assuming a static camera, we first segment moving objects from the background. By tracking objects of interest, we compute the object's self-similarity as it evolves in time. For periodic motion, the self-similarity metric is periodic, and is Fourier analyzed to detect and characterize periodicity. Examples on real image sequences are given. Ross Cutler, Larry Davis 0001 |
ICPR | 1 |