Pranay Manocha

dblp:207/8256 · DBLP profile ↗
← Back
12ranked-venue papers
11as first author
10since 2021 · last 2024
0000-0003-3284-5908ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 6 first-author · 5 since 2021
YearPublicationVenuePosition
2024 Corn: Co-Trained Full- and No-Reference Speech Quality Assessment
abstract
Perceptual evaluation constitutes a crucial aspect of various audio-processing tasks. Full reference (FR) or similarity-based metrics rely on high-quality reference recordings, to which lower-quality or corrupted versions of the recording may be compared for evaluation. In contrast, no-reference (NR) metrics evaluate a recording without relying on a reference. Both the FR and NR approaches exhibit advantages and drawbacks relative to each other. In this paper, we present a novel framework called CORN that amalgamates these dual approaches, concurrently training both FR and NR models together. After training, the models can be applied independently. We evaluate CORN by predicting several common objective metrics and across two different architectures. The NR model trained using CORN has access to a reference recording during training, and thus, as one would expect, it consistently outperforms baseline NR models trained independently. Perhaps even more remarkable is that the CORN FR model also outperforms its baseline counterpart, even though it relies on the same training data and the same model architecture. Thus, a single training regime produces two independently useful models, each outperforming independently trained models.
Pranay Manocha, Donald Williamson, Adam Finkelstein
ICASSP1
2023 Torchaudio-Squim: Reference-Less Speech Quality and Intelligibility Measures in Torchaudio
abstract
Measuring quality and intelligibility of a speech signal is usually a critical step in development of speech processing systems. To enable this, a variety of metrics to measure quality and intelligibility under different assumptions have been developed. Through this paper, we introduce tools and a set of models to estimate such known metrics using deep neural networks. These models are made available in the well-established TorchAudio library, the core audio and speech processing library within the PyTorch deep learning framework. We refer to it as TorchAudio-Squim, TorchAudio-Speech QUality and Intelligibility Measures. More specifically, in the current version of TorchAudio-squim, we establish and release models for estimating PESQ, STOI and SI-SDR among objective metrics and MOS among subjective metrics. We develop a novel approach for objective metric estimation and use a recently developed approach for subjective metric estimation. These models operate in a "referenceless" manner, that is they do not require the corresponding clean speech as reference for speech assessment. Given the unavailability of clean speech and the effortful process of subjective evaluation in real-world situations, such easy-to-use tools would greatly benefit speech processing research and development.
Anurag Kumar 0003, Ke Tan 0001, Zhaoheng Ni, Pranay Manocha, Xiaohui Zhang 0007, Ethan Henderson, Buye Xu
ICASSP4
2023 Nord: Non-Matching Reference Based Relative Depth Estimation from Binaural Speech
abstract
We propose NORD: a novel framework for estimating the relative depth between two binaural speech recordings. In contrast to existing depth estimation techniques, ours only requires audio signals as input. We trained the framework to solve depth preference (i.e. which input perceptually sounds closer to the listener’s head), and quantification tasks (i.e. quantifying the depth difference between the inputs). In addition, training leverages recent advances in metric and multi-task learning, which allows the framework to be invariant to both signal content (i.e. non-matched reference) and directional cues (i.e. azimuth and elevation). Our framework has additional useful qualities that make it suitable for use as an objective metric to benchmark binaural audio systems, particularly depth perception and sound externalization, which we demonstrate through experiments. We also show that NORD generalizes well under different reverberation and environments. The results from preference and quantification tasks correlate well with measured results.
Pranay Manocha, Israel D. Gebru, Anurag Kumar 0003, Dejan Markovic, Alexander Richard
ICASSP1
2023 Spatialization Quality Metric for Binaural Speech
Pranay Manocha, Israel D. Gebru, Anurag Kumar 0003, Dejan Markovic, Alexander Richard
INTERSPEECH1
2022 SQAPP: No-Reference Speech Quality Assessment Via Pairwise Preference
abstract
Automatic speech quality assessment remains challenging, as we lack complete models of human auditory perception. Many existing full-reference models correlate well with human perception, but cannot be used in real-world scenarios where ground truth clean reference recordings are not available. On the other hand no-reference metrics typically suffer from several shortcomings, such as lack of robustness to unseen perturbations and reliance on (limited) labeled data for training. Moreover, noise or large variance among the labels makes it difficult to learn generalizable representations, especially for recordings with subtle differences. This paper proposes a learning framework for estimating the quality of a recording without any reference, and without any human judgments. The main component of this framework is a pairwise quality-preference strategy that reduces label noise, thereby making learning more robust. From pairwise preferences, we first learn a content invariant quality ordering; and then we re-target the model to predict quality on an absolute scale. We show that the resulting learned metric is well-calibrated with human judgments. Since it is a deep network, the metric is differentiable, making it suitable as a loss function for downstream tasks. For example, we show that adding this metric to an existing speech enhancement method yields significant improvement.
Pranay Manocha, Zeyu Jin, Adam Finkelstein
ICASSP1
2022 Speech Quality Assessment through MOS using Non-Matching References
abstract
Human judgments obtained through Mean Opinion Scores (MOS) are the most reliable way to assess the quality of speech signals.However, several recent attempts to automatically estimate MOS using deep learning approaches lack robustness and generalization capabilities, limiting their use in real-world applications.In this work, we present a novel framework, NORESQA-MOS, for estimating the MOS of a speech signal.Unlike prior works, our approach uses non-matching references as a form of conditioning to ground the MOS estimation by neural networks.We show that NORESQA-MOS provides better generalization and more robust MOS estimation than previous state-of-the-art methods such as DNSMOS [1] and NISQA [2], even though we use a smaller training set.Moreover, we also show that our generic framework can be combined with other learning methods such as self-supervised learning and can further supplement the benefits from these methods.
Pranay Manocha, Anurag Kumar 0003
INTERSPEECH1
2022 SAQAM: Spatial Audio Quality Assessment Metric
abstract
Audio quality assessment is critical for assessing the perceptual realism of sounds.However, the time and expense of obtaining "gold standard" human judgments limit the availability of such data.For AR&VR, good perceived sound quality and localizability of sources are among the key elements to ensure complete immersion of the user.Our work introduces SAQAM which uses a multi-task learning framework to assess listening quality (LQ) and spatialization quality (SQ) between any given pair of binaural signals without using any subjective data.We model LQ by training on a simulated dataset of triplet human judgments, and SQ by utilizing activation-level distances from networks trained for direction of arrival (DOA) estimation.We show that SAQAM correlates well with human responses across four diverse datasets.Since it is a deep network, the metric is differentiable, making it suitable as a loss function for other tasks.For example, simply replacing an existing loss with our metric yields improvement in a speech-enhancement network.
Pranay Manocha, Anurag Kumar 0003, Buye Xu, Anjali Menon, Israel D. Gebru, Vamsi K. Ithapu, Paul Calamia
INTERSPEECH1
2022 Audio Similarity is Unreliable as a Proxy for Audio Quality
abstract
Many audio processing tasks require perceptual assessment.However, the time and expense of obtaining "gold standard" human judgments limit the availability of such data.Most applications incorporate full reference or other similarity-based metrics (e.g.PESQ) that depend on a clean reference.Researchers have relied on such metrics to evaluate and compare various proposed methods, often concluding that small, measured differences imply one is more effective than another.This paper demonstrates several practical scenarios where similarity metrics fail to agree with human perception, because they: (1) vary with clean references; (2) rely on attributes that humans factor out when considering quality, and (3) are sensitive to imperceptible signal level differences.In those scenarios, we show that no-reference metrics do not suffer from such shortcomings and correlate better with human perception.We conclude therefore that similarity serves as an unreliable proxy for audio quality.
Pranay Manocha, Zeyu Jin, Adam Finkelstein
INTERSPEECH1
2021 CDPAM: Contrastive Learning for Perceptual Audio Similarity
abstract
Many speech processing methods based on deep learning require an automatic and differentiable audio metric for the loss function. The DPAM approach of Manocha et al. [1] learns a full-reference metric trained directly on human judgments, and thus correlates well with human perception. However, it requires a large number of human annotations and does not generalize well outside the range of perturbations on which it was trained. This paper introduces CDPAM –a metric that builds on and advances DPAM. The primary improvement is to combine contrastive learning and multi-dimensional representations to build robust models from limited data. In addition, we collect human judgments on triplet comparisons to improve generalization to a broader range of audio perturbations. CDPAM correlates well with human responses across nine varied datasets. We also show that adding this metric to existing speech synthesis and enhancement methods yields significant improvement, as measured by objective and subjective tests.
Pranay Manocha, Zeyu Jin, Richard Zhang 0001, Adam Finkelstein
ICASSP1
2021 NORESQA: A Framework for Speech Quality Assessment using Non-Matching References
abstract
The perceptual task of speech quality assessment (SQA) is a challenging task for machines to do. Objective SQA methods that rely on the availability of the corresponding clean reference have been the primary go-to approaches for SQA. Clearly, these methods fail in real-world scenarios where the ground truth clean references are not available. In recent years, non-intrusive methods that train neural networks to predict ratings or scores have attracted much attention, but they suffer from several shortcomings such as lack of robustness, reliance on labeled data for training and so on. In this work, we propose a new direction for speech quality assessment. Inspired by human's innate ability to compare and assess the quality of speech signals even when they have non-matching contents, we propose a novel framework that predicts a subjective relative quality score for the given speech signal with respect to any provided reference without using any subjective data. We show that neural networks trained using our framework produce scores that correlate well with subjective mean opinion scores (MOS) and are also competitive to methods such as DNSMOS, which explicitly relies on MOS from humans for training networks. Moreover, our method also provides a natural way to embed quality-related information in neural networks, which we show is helpful for downstream tasks such as speech enhancement.
Pranay Manocha, Buye Xu, Anurag Kumar 0003
NeurIPS1
2020 A Differentiable Perceptual Audio Metric Learned from Just Noticeable Differences
abstract
Many audio processing tasks require perceptual assessment.The "gold standard" of obtaining human judgments is timeconsuming, expensive, and cannot be used as an optimization criterion.On the other hand, automated metrics are efficient to compute but often correlate poorly with human judgment, particularly for audio differences at the threshold of human detection.In this work, we construct a metric by fitting a deep neural network to a new large dataset of crowdsourced human judgments.Subjects are prompted to answer a straightforward, objective question: are two recordings identical or not?These pairs are algorithmically generated under a variety of perturbations, including noise, reverb, and compression artifacts; the perturbation space is probed with the goal of efficiently identifying the just-noticeable difference (JND) level of the subject.We show that the resulting learned metric is well-calibrated with human judgments, outperforming baseline methods.Since it is a deep network, the metric is differentiable, making it suitable as a loss function for other tasks.Thus, simply replacing an existing loss (e.g., deep feature loss) with our metric yields significant improvement in a denoising network, as measured by subjective pairwise comparison.
Pranay Manocha, Adam Finkelstein, Richard Zhang 0001, Nicholas J. Bryan, Gautham J. Mysore, Zeyu Jin
INTERSPEECH1
2018 Content-Based Representations of Audio Using Siamese Neural Networks
abstract
In this paper, we focus on the problem of content-based retrieval for audio, which aims to retrieve all semantically similar audio recordings for a given audio clip query. This problem is similar to the problem of query by example of audio, which aims to retrieve media samples from a database, which are similar to the user-provided example. We propose a novel approach which encodes the audio into a vector representation using Siamese Neural Networks. The goal is to obtain an encoding similar for files belonging to the same audio class, thus allowing retrieval of semantically similar audio. Using simple similarity measures such as those based on simple euclidean distance and cosine similarity we show that these representations can be very effectively used for retrieving recordings similar in audio content.
Pranay Manocha, Rohan Badlani, Anurag Kumar 0003, Ankit Shah 0001, Benjamin Elizalde, Bhiksha Raj
ICASSP1