VLDB 2026 Research / reviewers in the wild / expert
Stephen D. Voran
dblp:76/8758 · also Stephen Voran
· DBLP profile ↗
17ranked-venue papers
15as first author
5since 2021 · last 2024
0000-0001-7840-8848ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 12 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Why Some Audio Signal Short-Time Fourier Transform Coefficients Have Nonuniform Phase DistributionsabstractThe short-time Fourier transform (STFT) represents a window of audio samples as a set of complex coefficients. These are advantageously viewed as magnitudes and phases and the overall distribution of phases is very often assumed to be uniform. We show that when audio signal STFT phase distributions are analyzed per-frequency or per-magnitude range, they can be far from uniform. That is, the uniform phase distribution assumption obscures significant important details. We explain the significance of the nonuniform phase distributions and how they might be exploited, derive their source, and explain why the choice of the STFT window shape influences the nonuniformity of the resulting phase distributions. Stephen D. Voran |
ICME | 1 |
| 2024 | AlignNet: Learning dataset score alignment functions to enable better training of speech quality estimators
Jaden Pieper, Stephen D. Voran |
INTERSPEECH | 2 |
| 2021 | Optimal Frame Duration for Oracle Audio Signal Separation is Determined by Joint Minimization of Two Antagonistic ArtifactsabstractWe demonstrate that the optimal audio signal processing frame duration in oracle binary masking and oracle magnitude restoration is determined by joint minimization of two antagonistic artifacts: temporal blurring (which increases with frame duration and log-spectral-error change per unit time (which decreases with frame duration). This is novel — the factors underlying the empirical optimization of frame duration have not been previously identified. Signal stationarity alone cannot explain the existence of an optimal frame duration. Stationarity can explain why a frame duration is too long, but it cannot explain why a frame duration is too short.We introduce a method for measuring the stationarity of an audio signal. We then use this essential tool along with measurements, modeling, and analysis in order to identify the two underlying factors that cause there to be an optimal frame duration. In addition we show that when recovering s from the mixture y = s + n with oracle binary masks or oracle magnitudes, the stationarity of s and the stationarity of n have opposite influences on the optimal frame duration. Increasing the stationarity of s increases optimal frame duration but increasing the stationarity of n decreases optimal frame duration. Stationarity alone cannot explain these opposing influences but our results do. Stephen D. Voran |
MMSP | 1 |
| 2021 | Full-Reference and No-Reference Objective Evaluation of Deep Neural Network SpeechabstractObjective speech quality and intelligibility estimators do not correctly assess speech generated by deep neural networks (DNNs). We use 256 speech files and subjective scores that cover 14 DNN speech conditions and 18 nonDNN speech conditions to show that 8 different full-reference (FR) estimators consistently underestimate subjective scores for the DNN conditions. Conversely, we find that five no-reference (NR) estimators consistently overestimate subjective scores for the DNN conditions. We show that a rudimentary but effective solution to these shortcomings is to simply average an FR result with an NR result. We also explore root causes and propose more fundamental solutions. It has been previously suggested that FR estimators over-penalize inaudible timing variations or jitter. We conduct several experiments that measure and remove jitter from spectral representations of DNN speech inside FR estimators. Jitter removal compensates for some of the underestimation, thus confirming that jitter is a part of the cause. In additional experiments we show that power mismatches on a syllabic time-scale also contribute to the underestimation issue in FR estimators. Regarding NR estimators, we suggest that they can be trained to accurately rate DNN speech when sufficient speech signals and corresponding subjective scores are available. Stephen D. Voran |
QoMEX | 1 |
| 2021 | Measuring Speech Quality of System Input while Observing only System OutputabstractWe present a set of relatively small-scale proof-of-concept experiments where we construct no-reference (NR) speech quality estimators that give reliable values of system-under-test (SUT) input speech quality in spite of the fact that NR estimators can only access SUT output speech. We then explain why this success is not as counter-intuitive as it might initially seem. Next we demonstrate that this advance can be used to adjust NR relative speech quality values to obtain the much more desirable and useful NR absolute speech quality values. The experiments start with over seven hours of studio-quality speech. A processor adds filtering, reverberation, and noise to simulate the somewhat lower quality speech that often must be used to test systems. Four different established full-reference speech quality estimators provide ground-truth values for these experiments. Stephen D. Voran |
QoMEX | 1 |
| 2020 | Wawenets: A No-Reference Convolutional Waveform-Based Approach to Estimating Narrowband and Wideband Speech QualityabstractBuilding on prior work we have developed a no-reference (NR) waveform-based convolutional neural network (CNN) architecture that can accurately estimate speech quality or intelligibility of narrowband and wideband speech segments. These Wideband Audio Waveform Evaluation Networks, or WAWEnets, achieve very high per-speech-segment correlation (ρseg≥ 0.92, RMSE ≤ 0.38) to established full-reference quality and intelligibility estimators (PESQ, POLQA, PEMO, STOI) based on over 17 hours of speech from 127 previously unseen talkers speaking in 13 different languages; just 10% of our total data. NR correlations at this level across such a broad scope are unprecedented. This achievement was made possible by using full-reference estimates as training targets so that WAWEnets could learn implicit undistorted speech models and exploit them to produce accurate NR estimates. Andrew Catellier, Stephen D. Voran |
ICASSP | 2 |
| 2017 | A multiple bandwidth objective speech intelligibility estimator based on articulation index band correlations and attentionabstractWe present ABC-MRT16-a new algorithm for objective estimation of speech intelligibility following the Modified Rhyme Test (MRT) paradigm. ABC-MRT16 is simple, effective and robust. When compared to subjective MRT data from 367 diverse conditions that include coding, noise, frame erasures, and much more, ABC-MRT16 (containing just one optimized parameter) yields a very high Pearson correlation (above 0.95) and a remarkably low RMS estimation error (below 7% of full scale.) We attribute these successes to concise modeling of core human processes in audition and forced-choice word selection. On each trial, ABC-MRT16 gathers word selection evidence in the form of articulation index band correlations and then uses a simple attention model to perform word selection using the best available evidence. Attending to best evidence allows ABC-MRT16 to work well for narrowband, wideband, superwideband, and fullband speech and noise without any bandwidth detection algorithm or side information. Stephen D. Voran |
ICASSP | 1 |
| 2013 | Lossless compression of G.711 speech using only look-up tablesabstractThe lossless compression algorithm specified in ITU-T Recommendation G.711.0 provides bit-exact G.711 speech coding at reduced bit-rates. We introduce two Look-Up Coders (LUCs) that also offer bit-exact G.711 speech coding at reduced rates but the LUCs do not use arithmetic operations and hence eliminate the need for a processor. Instead they read in eight G.711 symbols, reinterpret those 64 bits to form eight new symbols that carry temporal information, then look up Huffman codes for those new symbols. When compared to G.711.0, LUC rates are 9% to 40% higher and they require 2 to 8 kB additional ROM, but LUCs eliminate about one million weighted arithmetic operations per second. LUCs reduce the 8 b/smpl G.711 rate to 3.8 to 6.7 b/smpl, depending on speech and noise levels. Stephen D. Voran |
ICASSP | 1 |
| 2013 | When should a speech coding quality increase be allowed within a talk-spurt?abstractThe value or harm associated with an increase in speech coding quality depends on the type of the increase as well as the temporal location of the increase in an utterance. For example, some increases in speech coding bandwidth can be perceived as impairments. The higher quality associated with the wider bandwidth can offset the impairment, but only if the increase happens early enough in an utterance. We present a subjective speech-quality experiment that qualifies these relationships at the talk-spurt time-scale for six different combinations of AMR and SILK speech coders. If a quality increase does not include a bandwidth increase, then, on average, it is beneficial only if it occurs in the first 2.8 seconds of a talk-spurt. If a quality increase includes a bandwidth increase, then it is beneficial only if it occurs in the first 1.8 seconds of a talk-spurt. Stephen D. Voran, Andrew Catellier |
ICASSP | 1 |
| 2010 | Multiple-Description Speech Coding Using Speech-Polarity DecompositionabstractWe present and evaluate a new multiple-description coding extension to the international standard for pulse code modulation speech coding (ITU-T Rec. G.711). This extension is inserted between the G.711 encoder and decoder. It uses speech-polarity decomposition to spread the speech signal across two channels thus increasing robustness to channel losses. When both channels deliver their payloads the extension becomes transparent and bit-exact G.711 speech samples are produced-there is no quality penalty. Due to low inter-channel redundancy, block coding, and entropy coding, the average total speech payload bit-rate is no greater than the 64 kbps rate of conventional G.711-there is no rate penalty. When either channel fails to deliver, the remaining channel still produces intelligible speech with moderately reduced quality thanks to a compressed sine-pulse fill-in algorithm. We are not aware of any other viable multiple-description coding extension that simultaneously meets the opposing goals of no quality penalty and no rate penalty. Stephen D. Voran, Andrew Catellier |
GLOBECOM | 1 |
| 2010 | Subjective ratings of instantaneous and gradual transitions from narrowband to wideband active speechabstractIn advanced heterogeneous telecommunication networks, network resources can dynamically dictate the type of speech coding that is used. An increase in resources allows for lower coding distortion or it might also be used to provide wideband speech instead of narrowband speech. Existing studies have demonstrated that wideband speech is preferred to narrowband speech, but they have also demonstrated that an abrupt transition from narrowband to wideband is perceived as an impairment, even though it is a transition to a higher quality signal. We describe our recent work that resulted in subjective scores for abrupt and gradual transitions from narrowband to wideband at the midpoint of a six-second segment of active speech. On average, signals that start narrowband and end wideband are rated slightly lower than constant narrowband signals and results are nearly the same for abrupt and gradual (2.5 second) transitions. Scores from 20 listeners show a wide range of individual opinions so we conclude that studies of bandwidth transitions may be quite sensitive to the listener population sample. Stephen D. Voran |
ICASSP | 1 |
| 2008 | Listener detection of talker stress in low-rate coded speechabstractWe describe an experiment where listeners were asked to detect two specific forms of stress in talkers' recorded voices heard via six different simulated communication systems. Both task-induced stress and dramatized urgency were used. Communication systems included low-rate digital speech coding combined with bit errors, packet loss, and packet loss concealment. Twenty-four listeners participated in a total of 11,520 detection trials. A parallel investigation of word intelligibility in sentence context used 576 trials. Intelligibility results showed wide variance due to communication system and stress detection results showed less variance. More specifically, we found that listener detection of dramatized talker urgency was 4.7 times more robust to communication system degradations than word intelligibility in sentence context. Stephen D. Voran |
ICASSP | 1 |
| 2005 | A Multiple-Description PCM Speech Coder using Structured Dual Vector QuantizersabstractIn this paper, we describe a 2-channel multiple-description speech coder based on the ITU-T recommendation G.711 PCM speech coder. The new coder operates in the PCM code domain in order to exploit the companding gain of PCM. It applies a pair of 2D structured vector quantizers to each pair of PCM codes, thus exploiting the correlation between adjacent speech samples. If both quantizer outputs are received, they are combined to generate an approximation to the original pair of PCM codes. If only one quantizer output is received, a coarser approximation is still possible. When using 6 bits/sample/channel (for a total data rate of 96 kbps) the coder provides an equivalent PCM speech quality of 7.3 bits/sample when both channels are working and 6.4 bits/sample when one channel is working. Stephen D. Voran |
ICASSP (1) | 1 |
| 2004 | Compensating for gain in objective quality estimation algorithmsabstractWhen objectively estimating speech, audio, or video quality, it is often necessary to compensate for the system gain or to "gain match" two or more signals. One can take three views of a system, leading to three different definitions of gain, and three different gain compensation solutions: one that minimizes distortion, one that matches input-output power, and one that maximizes signal-to-distortion ratio. We derive these three solutions, describe the algebraic and geometric relationships between them, and provide a generalized result that subsumes all three. We provide examples showing that these three solutions do differ in practical quality estimation situations. We also report some of the gain compensation choices found in the quality estimation literature. Stephen D. Voran |
ICASSP (3) | 1 |
| 1999 | Objective estimation of perceived speech quality. I. Development of the measuring normalizing block techniqueabstractPerceived speech quality is most directly measured by subjective listening tests. These tests are often slow and expensive, and numerous attempts have been made to supplement them with objective estimators of perceived speech quality. These attempts have found limited success, primarily in analog and higher-rate, error-free digital environments where speech waveforms are preserved or nearly preserved. The objective estimation of the perceived quality of highly compressed digital speech, possibly with bit errors or frame erasures has remained an open question. We report our findings regarding two essential components of objective estimators of perceived speech quality: perceptual transformations and distance measures. A perceptual transformation modifies a representation of an audio signal in a way that is approximately equivalent to the human hearing process. A distance measure reflects the magnitude of a perceived distance between two perceptually transformed signals. We then describe a new objective estimation approach that uses a simple but effective perceptual transformation and a distance measure that consists of a hierarchy of measuring normalizing blocks. Each measuring normalizing block integrates two perceptually transformed signals over some time or frequency interval to determine the average difference across that interval. This difference is then normalized out of one signal, and is further processed to generate one or more measurements. Stephen D. Voran |
IEEE Trans. Speech Audio Process. | 1 |
| 1999 | Objective estimation of perceived speech quality .II. Evaluation of the measuring normalizing block techniqueabstractFor pt.I see ibid., vol.7, no.4, p.371-82. Part I of this paper describes a new approach to the objective estimation of perceived speech quality. This new approach uses a simple but effective perceptual transformation and a distance measure that consists of a hierarchy of measuring normalizing blocks. Each measuring normalizing block integrates two perceptually transformed signals over some time or frequency interval to determine the average difference across that interval. This difference is then normalized out of one signal, and is further processed to generate one or more measurements. In this part, the resulting estimates of the perceived speech quality are correlated with the results of nine subjective listening tests. Together, these tests include 219 4 kHz bandwidth speech codecs, transmission systems, and reference conditions, with bit rates ranging from 2.4 to 61 kb/s. When compared with six other estimators, significant improvements are seen in many cases, particularly at lower bit rates, and when bit errors or frame erasures are present. These hierarchical structures of measuring normalizing blocks, or other structures of measuring normalizing blocks may also address open issues in perceived audio quality estimation, layered speech or audio coding, automatic speech or speaker recognition, audio signal enhancement, and other areas. Stephen D. Voran |
IEEE Trans. Speech Audio Process. | 1 |
| 1998 | A simplified version of the ITU algorithm for objective measurement of speech codec qualityabstractITU-T Recommendation P.861 describes an objective speech duality assessment algorithm for speech codecs. This algorithm transforms codec input and output speech signals into a perceptual domain, compares them, and generates a noise disturbance value, which can be used to estimate perceived speech quality. The performance of this algorithm can be judged by the correlation between those estimates and actual listener opinions from formal subjective listening tests. We show that significant simplifications can be made to the P.861 algorithm with very minimal effect on its performance. Specifically, for the portions of the algorithm under study here, 64% of the floating point operations can be eliminated with only a 3.5% decrease in average correlation to listener opinions. The resulting simplified algorithm may offer a practical new objective function to drive parameter selections, excitation searches, and bit-allocations in speech and audio coders. Stephen D. Voran |
ICASSP | 1 |