EDBT 2026 Demo / reviewers in the wild / expert
Richard V. Cox
dblp:56/6024
· DBLP profile ↗
40ranked-venue papers
12as first author
0since 2021 · last 2003
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 8 first-authorComputer networks · 7 · 2 first-authorArtificial intelligence and machine learning · 4Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Speech recognition and synthesis · 90% Question answering and dialogue systems · 10% | |
| Computer graphics and multimedia
7 papers |
Audio and music processing · 69% Multimedia systems and quality of experience · 12% Multimedia analysis and retrieval · 12% | |
| Computer networks
4 papers |
Physical-layer communications · 77% Cellular and mobile networks · 14% Content delivery and video streaming · 5% |
Topics — the 27 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing
speech coding |
0.1 | 5 | 2003 | Improving the transcoding capability of speech coders · IEEE Trans. Multim. 2003 Joint source-channel (de-)coding for mobile communications · IEEE Trans. Commun. 2002 A Low-Delay CELP Coder for the CCITT 16 kb/s Speech Coding Standard · IEEE J. Sel. Areas Commun. 1992 |
Natural language and speech › Speech recognition and synthesis
speech coding |
0.1 | 2 | 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigm · IEEE Trans. Speech Audio Process. 2001 Speech and language processing for next-millennium communications services · Proc. IEEE 2000 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
robust speech recognition |
0.0 | 1 | 2002 | Performance improvement of a bitstream-based front-end for wireless speech recognition in adverse environments · IEEE Trans. Speech Audio Process. 2002 |
Physical-layer communications › coding theory
joint source-channel coding |
0.0 | 1 | 2002 | Joint source-channel (de-)coding for mobile communications · IEEE Trans. Commun. 2002 |
Natural language and speech › Speech recognition and synthesis › speech coding
low-bit-rate speech coding |
0.0 | 1 | 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigm · IEEE Trans. Speech Audio Process. 2001 |
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
unit selection |
0.0 | 1 | 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigm · IEEE Trans. Speech Audio Process. 2001 |
Audio and music processing
speech recognition |
0.0 | 1 | 2001 | A bitstream-based front-end for wireless speech recognition on IS-136 communications system · IEEE Trans. Speech Audio Process. 2001 |
Natural language and speech › Question answering and dialogue systems
spoken dialogue systems |
0.0 | 1 | 2000 | Speech and language processing for next-millennium communications services · Proc. IEEE 2000 |
Multimedia analysis and retrieval
multimodal signal processing |
0.0 | 1 | 1998 | On the applications of multimedia processing to communications · Proc. IEEE 1998 |
Physical-layer communications › channel modeling
fading channel modeling |
0.0 | 1 | 2002 | Joint source-channel (de-)coding for mobile communications · IEEE Trans. Commun. 2002 |
Physical-layer communications › fading channels
rayleigh fading |
0.0 | 1 | 2002 | Joint source-channel (de-)coding for mobile communications · IEEE Trans. Commun. 2002 |
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
concatenative speech synthesis |
0.0 | 1 | 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigm · IEEE Trans. Speech Audio Process. 2001 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.0 | 1 | 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigm · IEEE Trans. Speech Audio Process. 2001 |
Biometric security
speaker verification |
0.0 | 1 | 2000 | Speech and language processing for next-millennium communications services · Proc. IEEE 2000 |
Audio and music processing › speech enhancement
post-filtering |
0.0 | 1 | 1988 | Enhancement of ADPCM speech coding with backward-adaptive algorithms for postfiltering and noise feedback · IEEE J. Sel. Areas Commun. 1988 |
Audio and music processing
speech enhancement |
0.0 | 1 | 1988 | Enhancement of ADPCM speech coding with backward-adaptive algorithms for postfiltering and noise feedback · IEEE J. Sel. Areas Commun. 1988 |
Image and video coding › transform coding
subband coding |
0.0 | 1 | 1988 | New directions in subband coding · IEEE J. Sel. Areas Commun. 1988 |
Image and video coding › quantization
vector quantization |
0.0 | 1 | 1988 | New directions in subband coding · IEEE J. Sel. Areas Commun. 1988 |
Content delivery and video streaming › source coding
speech coding |
0.0 | 2 | 1982 | Real-Time Speech Coding · IEEE Trans. Commun. 1982 Multiple User Variable Rate Coding for TASI and Packet Transmission Systems · IEEE Trans. Commun. 1980 |
Physical-layer communications
signal processing for communications |
0.0 | 1 | 1982 | Real-Time Speech Coding · IEEE Trans. Commun. 1982 |
Internet architecture and protocols › packet voice
packet speech transmission |
0.0 | 1 | 1980 | Multiple User Variable Rate Coding for TASI and Packet Transmission Systems · IEEE Trans. Commun. 1980 |
Network performance modeling › statistical multiplexing
TASI |
0.0 | 1 | 1980 | Multiple User Variable Rate Coding for TASI and Packet Transmission Systems · IEEE Trans. Commun. 1980 |
Physical-layer communications › channel coding › adaptive coding
variable rate coding |
0.0 | 1 | 1980 | Multiple User Variable Rate Coding for TASI and Packet Transmission Systems · IEEE Trans. Commun. 1980 |
Audio and music processing › speech coding
linear predictive coding |
0.0 | 1 | 1988 | Enhancement of ADPCM speech coding with backward-adaptive algorithms for postfiltering and noise feedback · IEEE J. Sel. Areas Commun. 1988 |
Embedded and real-time systems
real-time signal processing |
0.0 | 1 | 1982 | Real-Time Speech Coding · IEEE Trans. Commun. 1982 |
Performance modeling and evaluation
queueing models |
0.0 | 1 | 1980 | Multiple User Variable Rate Coding for TASI and Packet Transmission Systems · IEEE Trans. Commun. 1980 |
Performance modeling and evaluation
stability analysis |
0.0 | 1 | 1980 | Multiple User Variable Rate Coding for TASI and Packet Transmission Systems · IEEE Trans. Commun. 1980 |
Methods — techniques the papers use, named apart from their topics
hidden markov model · 0.1speech enhancement · 0.1soft-decision decoding · 0.1low-dimensional quantization · 0.1codebook gain re-estimation · 0.1speech synthesis · 0.1speech recognition · 0.1natural language understanding · 0.1multimedia signal processing · 0.0perceptual subjective quality measure · 0.0bitstream mapping · 0.0rate-distortion piecewise linear approximation · 0.0harmonic+noise model · 0.0cepstrum · 0.0code-excited linear prediction · 0.0backward-adaptive linear prediction · 0.0adaptive postfiltering · 0.0subband coding · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2003 | An intelligibility enhancement for the mixed excitation linear prediction speech coderabstractWe recently proposed a technique to increase the robustness of perceptually important acoustic cues in speech, based on modification of the phoneme timing structure. We now apply the technique as a preprocessor to improve the intelligibility of the mixed excitation linear prediction (MELP) coder. The enhancement strategy, however, requires adaptations to minimize the introduction of unnatural speech characteristics that are introduced and intensified by parametric coding schemes. The enhanced 2.4 kb/s fixed-point MELP coder achieves a level of speech intelligibility comparable to the 8 kb/s G.729 Annex A coder. Nicola R. Chong-White, Richard V. Cox |
IEEE Signal Process. Lett. | 2 |
| 2003 | Improving the transcoding capability of speech codersabstractWith the trend of merging various communication networks, a need arises to provide transcoding between different speech coding formats. Presently this means a cross tandem between the two coders in each case. This results in both quality loss and extra delay. A possible alternative is using a bitstream mapping approach that directly converts parameter values. For several standard coders having a similar coding structure, it should be possible to generate comparable or better quality without adding much delay or complexity. This paper proposes a bitstream mapping method between ITU-T Recommendation G.729 and TIA IS-641. Informal listening tests and the perceptual subjective quality measure (PSQM) scores show that the proposed method has better quality than the cross tandem method, while it has at least 5 ms less delay and six times less computation. Hong-Goo Kang, Hong Kook Kim, Richard V. Cox |
IEEE Trans. Multim. | 3 |
| 2002 | A segmental speech coder based on a concatenative TTS
Ki-Seung Lee, Richard V. Cox |
Speech Commun. | 2 |
| 2002 | Performance improvement of a bitstream-based front-end for wireless speech recognition in adverse environmentsabstractWe propose a feature enhancement algorithm for wireless speech recognition in adverse acoustic environments. A speech recognition system is realized at the network side of a wireless communications system and feature parameters are extracted directly from the bitstream of the speech coder employed in the system, where the feature parameters are composed of spectral envelope information and coder-specific information. The coder-specific information is apt to be affected by environmental noise because the speech coder fails to generate high quality speech in noisy environments. We first found that enhancing noisy speech prior to speech coding improves the recognizer's performance. However, our aim was to develop a robust front-end operating at the network side of a wireless communications system without regard to whether speech enhancement was applied at the sender side. We investigated the effect of a speech enhancement algorithm on the bitstream-based feature parameters. Consequently, a feature enhancement algorithm is proposed which incorporates feature parameters obtained from the decoded speech and a noise suppressed version of the decoded speech. The coder-specific information can also be improved by re-estimating the codebook gains and residual energy from the enhanced residual signal. HMM-based connected digit recognition experiments show that the proposed feature enhancement algorithm significantly improves recognition performance at low signal-to-noise ratio (SNR) without causing poorer performance at high SNR. From large vocabulary speech recognition experiments with far-field microphone speech signals recorded in an office environment, we show that the feature enhancement algorithm greatly improves word recognition accuracy. Hong Kook Kim, Richard V. Cox, Richard C. Rose |
IEEE Trans. Speech Audio Process. | 2 |
| 2002 | Joint source-channel (de-)coding for mobile communicationsabstractReal world source coding algorithms usually leave a certain amount of redundancy within the coded bit stream. Shannon (1948) already mentioned that this redundancy can be exploited at the receiver side to achieve a higher robustness against channel errors. We show how joint source-channel decoding can be performed in a way that is applicable to any mobile communication system standard. Considerable gains in terms of bit error rate or signal-to-noise ratio (SNR) are possible dependent on the amount of redundancy. However, an even better performance can be achieved by changing also the transmitter sided source and channel encoders. We propose an encoding concept employing low-dimensional quantization. Keeping the gross bit rate as well as the clean channel quality the same, it decreases the complexity of the source encoder and the decoder significantly. Finally, we give an application of our methods to spectral coefficient coding in speech transmission over a Rayleigh fading channel resulting in channel SNR gains of about 2 dB as compared to state-of-the-art (de-)coding and bad frame handling methods. Tim Fingscheidt, Thomas Hindelang, Richard V. Cox, Nambi Seshadri |
IEEE Trans. Commun. | 3 |
| 2001 | Feature enhancement for a bitstream-based front-end in wireless speech recognitionabstractWe propose a feature enhancement algorithm for wireless speech recognition in adverse acoustic environments. A speech recognition system is realized at the receiver side of a wireless communications system and feature parameters are extracted directly from the bitstream of the speech coder employed in the system. The feature parameters are composed of spectral envelope and coder-specific information. The proposed feature enhancement algorithm incorporates feature parameters obtained from the decoded speech and an enhanced version into the bitstream-based feature parameters. Moreover, the coder-specific parameters are improved by reestimating the codebook gains and residual energy from the enhanced residual signal. HMM-based connected digit recognition experiments show that the proposed feature enhancement algorithm significantly improves recognition accuracy at low SNR without causing poorer performance at high SNR. Hong Kook Kim, Richard V. Cox |
ICASSP | 2 |
| 2001 | A bitstream-based front-end for wireless speech recognition on IS-136 communications systemabstractWe propose a feature extraction method for a speech recognizer that operates in digital communication networks. The feature parameters are basically extracted by converting the quantized spectral information of a speech coder into a cepstrum. We also include the voiced/unvoiced information obtained from the bitstream of the speech coder in the recognition feature set. We performed speaker-independent connected digit HMM recognition experiments under clean, background noise, and channel impairment conditions. From these results, we found that the speech recognition system employing the proposed bitstream-based front-end gives superior word and string accuracies over a recognizer constructed from decoded speech signals. Its performance is comparable to that of a wireline recognition system that uses the cepstrum as a feature set. Next, we extended the evaluation of the proposed bitstream-based front-end to large vocabulary speech recognition with a name database. The recognition results proved that the proposed bitstream-based front-end also gives a comparable performance to the conventional wireline front-end. Hong Kook Kim, Richard V. Cox |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | A very low bit rate speech coder based on a recognition/synthesis paradigmabstractPrevious studies have shown that a concatenative speech synthesis system with a large database produces more natural sounding speech. We apply this paradigm to the design of improved very low bit rate speech coders (sub 1000 b/s). The proposed speech coder consists of unit selection, prosody coding, prosody modification and waveform concatenation. The encoder selects the best unit sequence from a large database and compresses the prosody information. The transmitted parameters include unit indices and the prosody information. To increase naturalness as well as intelligibility, two costs are considered in the unit selection process: an acoustic target cost and a concatenation cost. A rate-distortion-based piecewise linear approximation is proposed to compress the pitch contour. The decoder concatenates the set of units, and then synthesizes the resultant sequence of speech frames using the harmonic+noise model (HNM) scheme. Before concatenating units, prosody modification which includes pitch shifting and gain modification is applied to match those of the input speech. With single speaker stimuli, a comparison category rating (CCR) test shows that the performance of the proposed coder is close to that of the 2400-b/s MELP coder at an average bit rate of about 800-b/s during talk spurts. Ki-Seung Lee, Richard V. Cox |
IEEE Trans. Speech Audio Process. | 2 |
| 2000 | Bitstream-based feature extraction for wireless speech recognitionabstractIn this paper, we propose a feature extraction method for a speech recognizer that operates in digital communication networks. The feature parameters are basically extracted by converting the quantized spectral information of a speech coder into a cepstrum. We also combine the voiced/unvoiced information obtained from the bitstream of the speech coder into the recognition feature set. From speaker-independent connected digit HMM recognition, we find that the speech recognition system employing the proposed bitstream-based front-end gives superior word and string accuracies over a recognizer constructed from decoded speech signals. Its performance is comparable to that of the wireline recognition system that uses only the cepstrum as a feature set. Hong Kook Kim, Richard V. Cox |
ICASSP | 2 |
| 2000 | A combined WI and MELP coder at 5.2 kbpsabstractThis paper presents a low bit rate speech coder that encodes speech using two parallel coders into one embedded bit frame structure of 5.2 kbps. The two coders, a mixed excitation linear predictive coder (MELP) and a waveform interpolation (WI) coder, share the same pitch and linear prediction analysis and quantization. The combined coder is structured to provide a higher performance option of WI while maintaining compatibility with standard 2.4 kbps MELP via an embedded MELP mode. Jan Skoglund, Richard V. Cox, John S. Collura |
ICASSP | 2 |
| 2000 | Combined Source/Channel (De-)Coding: Can a Priori Information be Used Twice?abstractIn digital transmission of speech, audio, images and video signals residual redundancy is often left after source coding due to the complexity and delay constraints. This redundancy remains both inside one block or frame but also in a time correlation of subsequent frames. We describe an approach to improve channel and source decoding by using both kinds of correlation. Further on we consider the bit-mapping and the multiplexing in coding and its effect on decoding. The mutual information as measurement for the gain in decoding through a priori information is explained. In this work on combined source and channel decoding, we try to answer the following question: can a priori information that models the source parameters be used twice; first at the channel decoder and then at the source decoder. The channel decoder uses the a priori information that models the bit stream generated by the source coder. This does not capture all the details of the source parameter level statistics. By exploiting the a priori knowledge of parameters (once more) at the source decoder, we show that it is possible to achieve better reconstruction than if this information was used at either of the decoders. Thomas Hindelang, Tim Fingscheidt, Nambi Seshadri, Richard V. Cox |
ICC (3) | 4 |
| 2000 | Speech and language processing for next-millennium communications servicesabstractIn the future, the world of telecommunications will be vastly different than it is today. The driving force will be the seamless integration of real time communications (e.g. voice, video, music, etc.) and data into a single network, with ubiquitous access to that network anywhere, anytime, and by a wide range of devices. The only currently available ubiquitous access device to the network is the telephone, and the only ubiquitous user access technology mode is spoken voice commands and natural language dialogues with machines. In the future, new access devices and modes will augment speech in this role, but are unlikely to supplant the telephone and access by speech anytime soon. Speech technologies have progressed to the point where they are now viable for a broad range of communications services, including: compression of speech for use over wired and wireless networks; speech synthesis, recognition, and understanding for dialogue access to information, people, and messaging; and speaker verification for secure access to information and services. The paper provides brief overviews of these technologies, discusses some of the unique properties of wireless, plain old telephone service, and Internet protocol networks that make voice communication and control problematic, and describes the types of voice services available in the past and today, and those that we foresee becoming available over the next several years. Richard V. Cox, Candace A. Kamm, Lawrence R. Rabiner, Juergen Schroeter, Jay G. Wilpon |
Proc. IEEE | 1 |
| 1999 | A modular approach to speech enhancement with an application to speech codingabstractEphraim and Malah's (1984, 1985) MMSE-LSA speech enhancement algorithm, while robust and effective, is difficult to tune and adjust for the tradeoff between noise reduction and distortion. We suggest a means of generalizing this design, which allows for other estimators besides the MMSE-LSA to be used within the same supporting framework. When a modified version of Ephraim and Van Trees's (see IEEE Trans. Speech and Audio Proc., vol.3, p.251-66, 1995) spectral domain constrained signal subspace estimator is used in this manner, we obtain a system with greater flexibility and similar performance. We also explore the possibility of using different speech enhancement techniques as pre-processors for different parameter extraction modules of the IS-641 speech coder (a 7.4 kbit/s ACELP codec). We show that such a strategy can increase the quality of the coded speech and lead to a system that is more robust to differing noise types. Anthony J. Accardi, Richard V. Cox |
ICASSP | 2 |
| 1999 | TTS based very low bit rate speech coderabstractThis paper addresses a speech coder which uses a text-to-speech (TTS) synthesis system to achieve very low bit rates (sub 1 kbps). The main issue of the work is the accurate coding of the pitch (f/sub 0/) and gain contours which are principle components of prosody. This is of paramount interest since the correct prosody will increase naturalness and an efficient coding scheme will provide high coding gain. Together with the phonetic transcription, the f/sub 0/ and gain contour constitute the parameters that are necessary for the TTS system to synthesize the speech signal. Piecewise linear approximation is used to code the f/sub 0/ parameter. A technique which minimizes the bit rate while maintaining f/sub 0/ error below a given threshold are described. To obtain both high compression and smoothly changing gain contours, the variance of the signal is averaged over each half phoneme length is transmitted as gain information. With single speaker stimuli, and a priori text transcription information, we obtained natural sounding speech at an average bit rate of about 300 bps. Ki-Seung Lee, Richard V. Cox |
ICASSP | 2 |
| 1999 | Tracking speech-presence uncertainty to improve speech enhancement in non-stationary noise environmentsabstractSpeech enhancement algorithms which are based on estimating the short-time spectral amplitude of the clean speech have better performance when a soft-decision gain modification, depending on the a priori probability of speech absence, is used. In reported works a fixed probability, q, is assumed. Since speech is non-stationary and may not be present in every frequency bin when voiced, we propose a method for estimating distinct values of q for different bins which are tracked in time. The estimation is based on a decision-theoretic approach for setting a threshold in each bin followed by short-time averaging. The estimated q's are used to control both the gain and the update of the estimated noise spectrum during speech presence in a modified MMSE log-spectral amplitude estimator. Subjective tests resulted in higher scores than for the IS-127 standard enhancement algorithm, when pre-processing noisy speech for a coding application. David Malah, Richard V. Cox, Anthony J. Accardi |
ICASSP | 2 |
| 1999 | Low delay analysis/synthesis schemes for joint speech enhancement and low bit rate speech codingabstractA corpus of spontaneous route descriptions was collected from 8 speakers (5 males and 3 females). The corpus was labelled according to the ToBI standard and a discourse analysis was completed. Four discourse tour acts were identified and these were found to occur mainly in non-embedded linear sequences. In general, the intonation of each route description was characterised by a single intonational phrase containing many intermediate phrases. There was a tendency for boundaries between tour acts and intermediate phrases to coincide, but there is usually more than one intermediate phrase to one tour act. Rainer Martin 0001, Hong-Goo Kang, Richard V. Cox |
EUROSPEECH | 3 |
| 1998 | On the applications of multimedia processing to communicationsabstractThe challenge of multimedia processing is to provide services that seamlessly integrate text, sound, image, and video information and to do it in a way that preserves the ease of use and interactivity of conventional plain old telephone service (POTS) telephony. To achieve this goal, there are a number of technological problems that must be considered, including: compression and coding of multimedia signals, including algorithmic issues, standards issues, and transmission issues; synthesis and recognition of multimedia signals, including speech, images, handwriting, and text; organization, storage, and retrieval of multimedia signals, including the appropriate method and speed of delivery, resolution, and quality of service; access methods to the multimedia signal, including spoken natural language interfaces, agent interfaces, and media conversion tools; searching by text, speech, and image queries; browsing by accessing the text, by voice, or by indexed images. In each of these areas, a great deal of progress has been made in the past few years, driven in part by the relentless growth in multimedia personal computers and in part by the promise of broad-band access from the home and from wireless connections. Standards have also played a key role in driving new multimedia services, both on the POTS network and on the Internet. It is the purpose of this paper to review the status of the technology in each of the areas listed above and to illustrate current capabilities by describing several multimedia applications that have been implemented at AT&T Labs over the past several years. Richard V. Cox, Barry G. Haskell, Yann LeCun, Behzad Shahraray, Lawrence R. Rabiner |
Proc. IEEE | 1 |
| 1997 | On the Applications of Multimedia Processing to TelecommunicationsabstractThe challenge of multimedia processing is to seamlessly integrate text, sound, image, and video information into a single communications channel, and to do it in a way that provides high quality communications while preserving the ease-of-use and interactivity of conventional telephony. There are a number of technology drivers that are pushing the technology forward, as well as a number of technological problems that must be overcome before multimedia becomes as ubiquitous as voiceband telephony. A key issue with any practical multimedia system has to do with standards that insure connectivity between customers and a range of service providers. Multimedia processing is an area of communications that is rapidly evolving. However, a number of interesting and important multimedia communications applications have evolved over the past several years, and some of these applications are described. Richard V. Cox, Barry G. Haskell, Yann LeCun, Behzad Shahraray, Lawrence R. Rabiner |
ICIP (1) | 1 |
| 1993 | The creation and evolution of 16 kbit/s LD-CELP: From concept to standard
Juin-Hwey Chen, Richard V. Cox |
Speech Commun. | 2 |
| 1992 | Improving the performance of the 16 kb/s LD-CELP speech coderabstractThe CCITT is in the process of standardizing a 16 kb/s speech coder submitted by AT&T called low-delay code excited linear prediction (LD-CELP). In the first phase of CCITT testing, the coder met all performance requirements except for the tandeming condition. To improve the tandeming performance, the authors first tuned the perceptual weighting filter to optimize the coder's performance after three tandems. They then added an adaptive postfilter which was also tuned for three tandems. The excitation codebook was also reoptimized using a multiple-language IRS-weighted training database. With these changes, in the second phase of CCITT testing, the coder significantly exceeded the tandeming performance requirement and also met all other requirements. In fact, the speech quality of LD-CELP was equivalent to or better than that of 32 kb/s ADPCM for all conditions tested. It is anticipated that the version of LD-CELP will be formally ratified as a CCITT standard in 1992. The author describes the modifications made to improve the performance and discusses the CCITT test results.> Juin-Hwey Chen, Nikil Jayant, Richard V. Cox |
ICASSP | 3 |
| 1992 | A Low-Delay CELP Coder for the CCITT 16 kb/s Speech Coding StandardabstractA low-delay code-excited linear prediction (LD-CELP) speech coder which is expected to be standardized in 1992 as a CCITT G Series Recommendation for universal applications of speech coding at 16 kb/s is presented. The coder achieves a one-way coding delay of less than 2 ms by making both the LPC predictor and the excitation gain backward-adaptive and by using a small excitation vector size of five samples. The official CCITT laboratory tests revealed that the speech quality of this 16 kb/s LD-CELP coder is either equivalent to or better than that of the CCITT G.721 standard 32-kb/s ADPCM coder for almost all conditions tested. A description of the LD-CELP algorithm, its implementation on the DSP32C for CCITT testing, and performance results from these tests are presented.> Juin-Hwey Chen, Richard V. Cox, Yen-Chun Lin, Nikil Jayant, Melvin J. Melchner |
IEEE J. Sel. Areas Commun. | 2 |
| 1991 | A fixed-point 16 kb/s LD-CELP algorithmabstractThe authors describe the algorithm modifications they have made in order to make the LD-CELP (low-delay code-excited linear prediction) algorithm suitable for 16 bit fixed-point implementation. They replaced T.P Barnwell's recursive windowing method (1981) by a novel hybrid which is partially recursive and partially nonrecursive. This method avoided the dynamic range problem and the double-precision arithmetic that would otherwise have been required. A fourth-order LPC filter was cascaded at the output of the original 50th-order LPC filter. This filter effectively reduced the spectral dynamic range of the output of the 50th-order LPC filter, therefore alleviating the ill-conditioning problem of the 50th-order LPC analysis. Although the algorithm of J. LeRoux and C. Gueguen (1979) was believed to be better suited for fixed-point arithmetic than Durbin's recursion (see L.R. Rabiner and R.W. Schafer, 1978), this was not necessarily true for the two-stage cascaded LPC filter used. Fixed-point simulation showed that the speech quality of the modified fixed-point LD-CELP algorithm was essentially the same as that of the original floating-point LD-CELP algorithm.> Juin-Hwey Chen, Yen-Chun Lin, Richard V. Cox |
ICASSP | 3 |
| 1990 | Real-time implementation and performance of a 16 kb/s low-delay CELP speech coderabstractA real-time implementation of the low-delay code-excited linear prediction (LD-CELP) speech coder based on the AT&T DSP32C is described. The performance of the coder is described. Recursive windowing is used to calculate the autocorrelation coefficients by dividing the Levinson-Durbin recursion into smaller parts and spreading the computation over several speech vectors. The LD-CELP encoder uses about 90% of the processor time of an 80-ns DSP32C, while the decoder uses about 40%. Hence, the coder is implemented on two DSP32Cs. For a single encoding with clear or noisy channels, the speech quality produced by this coder is either equivalent to or better than the CCITT G.721 standard 32-kb/s ADPCM. However, for multiple asynchronous encodings the ADPCM coders gives slightly better performance. The LD-CELP coder passes dual-tone multifrequency tones and 300-, 1200-, and 2400-b/s modem signals.> Juin-Hwey Chen, Melvin J. Melchner, Richard V. Cox, Duane O. Bowker |
ICASSP | 3 |
| 1989 | Spectral quantization and interpolation for CELP codersabstractThe authors present results on the comparative performance of nonuniform scalar quantizers using three different LPC (linear predictive coding) representations: the arcsine of reflection coefficients, the log area ratios, and the line spectral frequencies. On comparing the spectral distortion introduced by quantizers based on these representations, it was found that the average distortion was very similar for all three, with the arcsine showing fewer large spectral errors. In a parallel study, the performance of the above LPC representations and the autocorrelation coefficients for interpolating the spectrum between adjacent time frames was investigated and revealed only small differences between the different representations. Informal listening tests with a complete 8 kb/s code-excited linear predictive (CELP) coder, incorporating both quantization and interpolation, showed no significant differences between the various LPC representations, suggesting that the random codebook for the excitation is able to compensate for small spectral deviations.> Bishnu S. Atal, Richard V. Cox, Peter Kroon |
ICASSP | 2 |
| 1989 | Robust CELP coders for noisy backgrounds and noisy channelsabstractThe authors examine the robustness of the code-excited linear predictive (CELP) coder, such as its ability to cope with nonspeech and corrupted speech inputs or to survive errors in the transmission of the coder parameters. They describe how they determined the error sensitivity of each coder parameter and identified the error propagation mechanisms. They find that the coder is robust to many kinds of input signals but is sensitive to channel errors. They show that the effect of bit errors can be reduced significantly with relatively simple measures such as repetition of the coder parameters from the most recent error-free frame. Incorporating these techniques into CELP results in a robust system with an acceptable performance for both burst and random bit errors rates of up to 1%. The clear channel performance degrades slightly (0.2 dB) as a result of these techniques.> Richard V. Cox, W. Bastiaan Kleijn, Peter Kroon |
ICASSP | 1 |
| 1988 | A sub-band coder designed for combined source and channel coding [speech coding]abstractThe coder is a dynamic bit allocation sub-band coder of the type first proposed by Ramstad (1982). It is structured in such a way that the relative importance of all bits is established as a byproduct of the dynamic bit allocation. It is shown that there is a difference in error sensitivity of four orders of magnitude between the most and the least important bits of the bit stream on average. A very flexible unequal error protection scheme is used to match the error sensitivity of the sub-band coder bit stream. This is accomplished using the concept of rate compatible punctured convolutional coding. The resulting coder gives robust performance over a simulated noisy channel. Its nominal bit rate is 16 kb/s, with 12 kb/s assigned for speech coding and 4 kb/s assigned for error correction. The entire encoder and decoder have been implemented on a single AT&T DSP-32 digital signal processor.> Richard V. Cox, Joachim Hagenauer, Nambi Seshadri, Carl-Erik W. Sundberg |
ICASSP | 1 |
| 1988 | New directions in subband codingabstractTwo very different subband coders are described. The first is a modified dynamic bit-allocation-subband coder (D-SBC) designed for variable rate coding situations and easily adaptable to noisy channel environments. It can operate at rates as low as 12 kb/s and still give good quality speech. The second coder is a 16-kb/s waveform coder, based on a combination of subband coding and vector quantization (VQ-SBC). The key feature of this coder is its short coding delay, which makes it suitable for real-time communication networks. The speech quality of both coders has been enhanced by adaptive postfiltering. The coders have been implemented on a single AT&T DSP32 signal processor.> Richard V. Cox, Steven L. Gay, Yair Shoham, Schuyler R. Quackenbush, Nambi Seshadri, Nikil Jayant |
IEEE J. Sel. Areas Commun. | 1 |
| 1988 | Enhancement of ADPCM speech coding with backward-adaptive algorithms for postfiltering and noise feedbackabstractIt is shown that postfiltering circuits based on higher order LPC (linear predictive coding) models can provide very low distortion in terms of special tilt. Thus, they can provide better speech enhancement than circuits based on the backward-adaptive pole-zero predictor in ADPCM (adaptive digital pulse code modulation). Quantitative criteria for designing postfiltering circuits based on higher-order LPC models are discussed. These postfilters are particularly attractive for systems where high-order LPC analysis is an integral part of the coding algorithm. In a subjective test that used a computer-simulated version of these circuits, enhanced ADPCM obtained a mean opinion score of 3.6 at 16 kb/s.> Venkatasubbarao Ramamoorthy, Nikil Jayant, Richard V. Cox, Man Mohan Sondhi |
IEEE J. Sel. Areas Commun. | 3 |
| 1987 | A family of ADPCM coders implemented on real-time hardwareabstractRecently, work in ADPCM speech coding has focused on using the adaptive predictors, already included in the coder algorithm, as the basis for noise shaping and post-filtering. This technique is based on improving the perceived quality of the signal -- no actual improvement in signal-to-noise ratio is possible. The original work was performed with an intended application of improving the performance of 16 and 24 kbps ADPCM. We have made extensions of this work for lower bit rates. Specifically we have looked at reducing the sampling rate by use of digital interpolation and at extending the technique to quantization rates as low as one bit per sample. The algorithms we have implemented run at rates in the range of 6 to 16 kbps. The coders below 9 kbps are probably not useful because their quality is too poor, but they do demonstrate the power of the noise shaping and post-filtering concept. The 9, 12 and 16 kbps coders would be useful in applications requiring a low complexity coder with only a single encoding. This paper describes the algorithms, their implementation using the WE (R) DSP32 signal processor, and their performance. Richard V. Cox |
ICASSP | 1 |
| 1986 | The analog voice privacy systemabstractThe Analog Voice Privacy System is based on individual sample permutation of the output samples of a sub-band coder analysis filterbank. The system has a large number of digital keys, giving it the strength of a digital encryption system, but also retains the good quality characteristics of analog scramblers. It has been implemented in a real-time hardware prototype designed for evaluation in the field. The units work with any modular telephone and standard 120 volts AC electricity. The device contains two circuitry boards, one for analog and one for digital processing which contain four digital signal processors. There are 125! possible permutation keys. These prototypes were designed to be tested in real telephone environments. To date, the device has been successfully tested over long distance telephone connections, several different analog and digital PBXs and telephone switches, and a channel simulator. The quality of the decrypted speech is considered very natural, and in particular, speaker recognition is retained. This is a significant advantage over digital vocoders. This paper describes the underlying principles of the algorithm, the details of its implementation and laboratory test results. Richard V. Cox, D. Bock, K. B. Bauer, James D. Johnston, Jeffrey Snyder |
ICASSP | 1 |
| 1986 | A high quality subband speech coder with backward adaptive predictor and optimal time-frequency bit assignmentabstractIn this paper we propose a hybrid speech coder which utilizes properties of APC, ATC, and subband coding. Subband splitting is used to reduce the dynamic range of a full-band APC. As a result, the prediction or power gain is reduced and the instability problem associated with the APC coder is alleviated to a large extent. The optimal bit allocations used in ATC is extended in the new coder where the available coding bits are optimally allocated, both in the time and frequency domains. Frame boundary artifacts (such as those found in ATC) are not present due to the time-domain processing nature of SBC. In order to improve the efficiency in quantizing the side information (time and frequency domain signal power and the corresponding bit allocations), vector quantization is used. Adaptive backward predictors are used to further reduce the number of bits allocated to side information, leaving more bits to encode the prediction residual signals. At 16 kbps the coder achieves better than 20 dB segmental signal-to-noise ratio and sounds transparent for most speakers. Frank K. Soong, Richard V. Cox, Nikil Jayant |
ICASSP | 2 |
| 1985 | Subband coding of speech using backward adaptive prediction and bit allocationabstractIn this study it is our goal to improve the performance of ADPCM and subband speech coders at medium bit rates (9.6∼16 kb/s) without increasing the coder complexity substantially. Various major building blocks including the predictors (both fixed and adaptive), subband quadrature mirror filter (QMF) and bit assignment strategy (both static and dynamic) are investigated in detail. We have found that (1) the Least-Squares (LS) adaptive lattice predictor outperforms both the pole-zero adaptive predictor recommended by CCITT and a first-order fixed predictor. (2) more subbands can improve the coder performance (3) longer QMF can reduce the interband aliasing and improve the subjective performance of a subband coder (4) an optimal dynamic bit allocation scheme with an improvement of SNR as high as 5 dB is much more favorable than a fixed bit allocation. With all of the above finding we propose a 4-band hybrid subband coder with an LS adaptive lattice predictor and an optimal dynamic bit allocation strategy. Frank K. Soong, Richard V. Cox, Nikil Jayant |
ICASSP | 2 |
| 1984 | Testing of wideband digital codersabstractDevelopment of Circuit Switched Digital Capability (CSDC) in the United States and the general evolution toward an integrated digital services network worldwide has created interest in the development of 56 or 64 kbps digital coders for audio signals with 6 or 7 kHz bandwidth. To date there have been a number of coders proposed but no formal testing of them. Furthermore, a method for testing them has not been proposed. In this paper we discuss the problem of testing such coders and describe a series of both objective and subjective tests which can be performed using a real-time coding system. We discuss the results of these tests when performed on a selection of coders which we have implemented. Richard V. Cox, Jeffrey Snyder, Ronald E. Crochiere, D. Bock, James D. Johnston |
ICASSP | 1 |
| 1983 | A 32-band sub-band/Transform coder incorporating vector quantization for dynamic bit allocationabstractIn this paper we report on a study of a technique for 32-band subband/transform coding at 16 kb/s. This approach occupies the middle range of algorithm complexities and frequency resolution between that of Sub-Band Coding (SBC) and Adaptive Transform Coding (ATC). Two designs for 16 kb/s 32-band coders have been simulated on a laboratory computer. The results of informal listening tests indicate that the new designs offer performance comparable to existing ATC techniques while having complexities roughly three times that of existing 4 and 5 band sub-band coders. C. D. Heron, Ronald E. Crochiere, Richard V. Cox |
ICASSP | 3 |
| 1982 | A single chip speech periodicity detectorabstractA real-time pitch/periodicity detector has been implemented on the Bell Laboratories Digital Signal Processing (DSP) integrated circuit. The design is based on a novel modification of the autocorrelation type pitch detector. The modification allows the computation and peak-picking of the autocorrelation function of an 8 kHz sampled input signal to be performed within the memory and real-time constraints of a single DSP. The algorithm is intended for applications in speech processing and coding where pitch or a measure of long-term periodicity of a signal is required. Specifically it has been applied in the design of a 9.6 kb/s speech coding system discussed in a companion paper. The pitch detector measures pitch over the normal range of most speakers. It does not compute a voiced/unvoiced decision. Richard V. Cox, Ronald E. Crochiere |
ICASSP | 1 |
| 1982 | A 9.6 kb/s speech coder using the Bell laboratories DSP integrated circuitabstractA digital speech coder has been designed for real-time operation for a data rate of 9.6 kb/s. The design is based on a combination of two speech compression techniques: Time-Domain Harmonic Scaling (TDHS) and Sub-Band Coding (SBC). It is a highly modularized hardware implementation using five Bell Laboratories Digital Signal Processor (DSP) integrated circuits as the key processing elements. Three DSPs are used in the encoder for pitch detection, TDHS compression and sub-band encoding. Another two DSPs are used in the receiver for sub-band decoding and TDHS expansion. Ronald E. Crochiere, Richard V. Cox, James D. Johnston, Linda Seltzer |
ICASSP | 2 |
| 1982 | A generalized comb filtering technique for speech enhancementabstractBecause of speech signals nonstationarity, usual comb filtering of noisy speech signals results only in a modest improvement in signal to noise ratio, and only in a small perceptual reduction of structured noise or interference. A generalized comb filtering technique, which applies a time-varying weighting to each pitch period, is mathematically analyzed and shown to be capable of breaking up the noise structure, in addition to comb filtering. This is found to provide a meaningful perceptual improvement when the noise or interference are structured. The mathematical analysis is facilitated by using a polyphase network model of the generalized comb filter. Design constraints and rules are developed and several filter families are proposed. Computer simulation results are discussed. David Malah, Richard V. Cox |
ICASSP | 2 |
| 1982 | Real-Time Speech CodingabstractThis paper reviews our recent efforts in the design and implementation of real-time speech coders. We discuss our approach and methodology for real-time hardware for coder techniques ranging from low to high complexity. Examples of realizations are given for each approach. They include adaptive differential PCM coding, subband coding, harmonic scaling with subband coding, and adaptive transform coding. Low to medium complexity techniques are based on the use of the Bell Laboratories digital signal processing (DSP) integrated circuit. High complexity block processing techniques are based on the use of an array processing computer. We conclude with an assessment of the performance versus complexity tradeoffs involved in these coding methods. Ronald E. Crochiere, Richard V. Cox, James D. Johnston |
IEEE Trans. Commun. | 2 |
| 1981 | A technique for perceptually reducing periodically structured noise in speechabstractPeriodically structured noise is noise which occurs randomly but with a fixed or slowly varying period. The noise periodicity is usually due to some underlying process, such as block processing of the speech where discontinuities between successive blocks result. This type of noise permeates the entire speech spectrum and is not removable by standard filtering techniques. The recently developed time domain harmonic sealing (TDHS) algorithm has been found to be the basis for an effective enhancement technique. In this paper we discuss the underlying theory of this technique and establish a class of windows for its implementation. As an example the frame rate noise of adaptive transform coding was perceptually reduced using this technique. Results from a subjective testing experiment using ATC coded speech with bit rates of 7.2 to 16 Kb/s indicated an improvement in quality equivalent to an increase in code rate of 2.4 to 3 Kb/s for speech originally coded at 7.2 to 12 Kb/s. Richard V. Cox, David Malah |
ICASSP | 1 |
| 1980 | Multiple User Variable Rate Coding for TASI and Packet Transmission SystemsabstractIn this paper we examine the use of variable rate coding concepts for TASI and packet speech transmission systems. The paper is divided into three major parts. In the first part, the theoretical performance of variable rate coding is analyzed for multiple user (TASI) applications. Potential gains are experimentally determined from twoparty telephone conversation data for up to 12 shared conversations on a channel. In the second part of the paper, the buffer control mechanism for a dynamic buffer scheme for coupling a variable rate coder to a fixed rate (or slowly varying rate) channel is analyzed. It is shown that the buffer control can be modeled as a second-order control system and, under adverse parameter settings, the system can be unstable. By an appropriate design and parameter setting, the buffer control can be stabilized, The insight developed from this particular buffer control mechanism may also lead to a better understanding of other buffer control problems in variable rate transmission or packet systems. In the third part of the paper, a practical method is analyzed for implementing a variable rate ADPCM system for multiple user applications. Examples of computer simulations of the system are presented. Richard V. Cox, Ronald E. Crochiere |
IEEE Trans. Commun. | 1 |