Jerry D. Gibson

dblp:14/5109 · DBLP profile ↗
← Back
105ranked-venue papers
16as first author
0since 2021 · last 2015
0000-0002-9827-1196ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 49 · 6 first-authorComputer networks · 31 · 6 first-authorTheory of computation · 10 · 3 first-authorArtificial intelligence and machine learning · 6Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 2Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
9 papers
Wireless networking · 42% Content delivery and video streaming · 37% Routing and switching · 20%
Computer graphics and multimedia
23 papers
Audio and music processing · 50% Image and video coding · 44% Multimedia systems and quality of experience · 6%
Theoretical computer science
13 papers
Coding theory · 85% Information theory · 14% Mathematical optimization · 1%
Artificial intelligence
4 papers
Speech recognition and synthesis · 100%

Topics — the 30 heaviest of 68, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
speech coding
0.2142007
Multiple Descriptions and Path Diversity for Voice Communications Over Wireless Mesh Networks · IEEE Trans. Multim. 2007
Structures for SNR scalable speech coding · IEEE Trans. Speech Audio Process. 2006
Low delay tree coding of speech at 8 kbit/s · IEEE Trans. Speech Audio Process. 1994
Wireless networking
mobile ad hoc networks
0.112011
Routing-Aware Multiple Description Video Coding Over Mobile Ad-Hoc Networks · IEEE Trans. Multim. 2011
Routing and switching
multipath routing
0.112011
Routing-Aware Multiple Description Video Coding Over Mobile Ad-Hoc Networks · IEEE Trans. Multim. 2011
Image and video coding
multiple description coding
0.122007
Multiple Descriptions and Path Diversity for Voice Communications Over Wireless Mesh Networks · IEEE Trans. Multim. 2007
Permuted smoothed descriptions and refinement coding for images · IEEE J. Sel. Areas Commun. 2000
Coding theory › source coding
rate-distortion theory
0.172005
Coefficient rate and lossy source coding · IEEE Trans. Inf. Theory 2005
Lossy Source Coding · IEEE Trans. Inf. Theory 1998
A bound on the rate of a system for encoding an unknown Gaussian autoregressive source · IEEE Trans. Inf. Theory 1994
Content delivery and video streaming
quality of experience
0.112008
Video Capacity of WLANs With a Multiuser Perceptual Quality Constraint · IEEE Trans. Multim. 2008
Wireless networking
WLAN
0.112008
Video Capacity of WLANs With a Multiuser Perceptual Quality Constraint · IEEE Trans. Multim. 2008
Coding theory › source coding
lossy source coding
0.122005
Coefficient rate and lossy source coding · IEEE Trans. Inf. Theory 2005
Lossy Source Coding · IEEE Trans. Inf. Theory 1998
Wireless networking
link adaptation
0.112007
Payload Length and Rate Adaptation for Multimedia Communications in Wireless LANs · IEEE J. Sel. Areas Commun. 2007
Audio and music processing › speech coding
code-excited linear prediction
0.112006
Structures for SNR scalable speech coding · IEEE Trans. Speech Audio Process. 2006
Audio and music processing › speech coding
scalable speech coding
0.112006
Structures for SNR scalable speech coding · IEEE Trans. Speech Audio Process. 2006
Natural language and speech › Speech recognition and synthesis
speech coding
0.131999
Efficient pitch filter encoding for variable rate speech processing · IEEE Trans. Speech Audio Process. 1999
Speech coding in mobile radio communications · Proc. IEEE 1998
Variable-rate CELP based on subband flatness · IEEE Trans. Speech Audio Process. 1997
Coding theory
source coding
0.071994
A bound on the rate of a system for encoding an unknown Gaussian autoregressive source · IEEE Trans. Inf. Theory 1994
A constrained joint source/channel coder design · IEEE J. Sel. Areas Commun. 1994
Uniform and piecewise uniform lattice vector quantization for memoryless Gaussian and Laplacian sources · IEEE Trans. Inf. Theory 1993
Image and video coding › error resilience
error-resilient video coding
0.022011
Routing-Aware Multiple Description Video Coding Over Mobile Ad-Hoc Networks · IEEE Trans. Multim. 2011
Multiframe video coding for improved performance over wireless channels · IEEE Trans. Image Process. 2001
Image and video coding
video compression
0.012001
Multiframe video coding for improved performance over wireless channels · IEEE Trans. Image Process. 2001
Image and video coding › error resilience
error concealment
0.012000
Permuted smoothed descriptions and refinement coding for images · IEEE J. Sel. Areas Commun. 2000
Multimedia systems and quality of experience
video quality assessment
0.012008
Video Capacity of WLANs With a Multiuser Perceptual Quality Constraint · IEEE Trans. Multim. 2008
Content delivery and video streaming
multimedia transmission
0.012007
Payload Length and Rate Adaptation for Multimedia Communications in Wireless LANs · IEEE J. Sel. Areas Commun. 2007
Content delivery and video streaming › error resilience
packet loss concealment
0.012007
Payload Length and Rate Adaptation for Multimedia Communications in Wireless LANs · IEEE J. Sel. Areas Commun. 2007
Routing and switching › multipath routing
path diversity
0.012007
Multiple Descriptions and Path Diversity for Voice Communications Over Wireless Mesh Networks · IEEE Trans. Multim. 2007
Wireless networking
wireless mesh network
0.012007
Multiple Descriptions and Path Diversity for Voice Communications Over Wireless Mesh Networks · IEEE Trans. Multim. 2007
Image and video coding
tree coding
0.021994
Low delay tree coding of speech at 8 kbit/s · IEEE Trans. Speech Audio Process. 1994
Fractional rate multitree speech coding · IEEE Trans. Commun. 1991
Natural language and speech › Speech recognition and synthesis › speech analysis
voice activity detection
0.011997
Variable-rate CELP based on subband flatness · IEEE Trans. Speech Audio Process. 1997
Coding theory › source coding › universal coding
adaptive coding
0.012005
Coefficient rate and lossy source coding · IEEE Trans. Inf. Theory 2005
Natural language and speech › Speech recognition and synthesis
speech analysis
0.011996
Speech analysis and segmentation by parametric filtering · IEEE Trans. Speech Audio Process. 1996
Natural language and speech › Speech recognition and synthesis › speech analysis
speech segmentation
0.011996
Speech analysis and segmentation by parametric filtering · IEEE Trans. Speech Audio Process. 1996
Image and video coding › predictive coding
adaptive prediction
0.041994
Analysis of the smoothed residual-driven algorithm for speech coders · IEEE Trans. Speech Audio Process. 1994
Sequentially Adaptive Backward Prediction in ADPCM Speech Coders · IEEE Trans. Commun. 1978
Experimental Comparison of All-Pole, All-Zero, and Pole-Zero Predictors for ADPCM Speech Coding · IEEE Trans. Commun. 1986
Coding theory
joint source-channel coding
0.021994
A constrained joint source/channel coder design · IEEE J. Sel. Areas Commun. 1994
Alphabet-constrained data compression · IEEE Trans. Inf. Theory 1982
Image and video coding › quantization › vector quantization
lattice vector quantization
0.011995
Image coding with uniform and piecewise-uniform vector quantizers · IEEE Trans. Image Process. 1995
Image and video coding › quantization
vector quantization
0.011995
Image coding with uniform and piecewise-uniform vector quantizers · IEEE Trans. Image Process. 1995

Methods — techniques the papers use, named apart from their topics

statistical packet loss model · 0.2routing-aware estimation · 0.2IEEE 802.11a modeling · 0.2AVC/H.264 video coding · 0.2rate-distortion analysis · 0.1PESQ-MOS · 0.1link adaptation · 0.1analytical framework · 0.1source coding structures · 0.1frequency-weighted distortion measures · 0.1reverse waterfilling · 0.1realization-adaptive compression · 0.1source coding · 0.0markov chain analysis · 0.0error propagation modeling · 0.0vector quantization · 0.0long-term prediction · 0.0analysis-by-synthesis · 0.0
YearPublicationVenuePosition
2015 High-order regularization for stereo color editing
abstract
This paper pioneers a method for local color editing on stereo image pairs. We generalize the conventional edit propagation framework to stereo views by introducing recent advances in the field of image segmentation, thus allowing a user's edits in one view to be simultaneously performed in the other view. This new formulation maintains consistent editing quality in both views and avoids singularities by solving a well regularized linear system for edit propagation.
Kuo-Chin Lien, Jerry D. Gibson, Matthew Turk 0001
ICIP2
2013 Stereo random field for bi-layer image segmentation
abstract
Stereo image segmentation usually incorporates depth cues to achieve high quality. However, previous methods that pointwise propagate information within stereo pairs could suffer from a poorly estimated depth map. In this paper, we introduce a novel graphical model where a greater amount of reliable messages can be conveyed during two-view joint segmentation. This model leads to a strongly coupled stereo pair, thus improving robustness, accuracy and consistency of stereo segmentation. Additionally, we augment a depth map to a novel correspondence matrix which is suitable for the proposed stereo segmentation model. Our experiments on a public stereo dataset show that the proposed correspondence method and stereo model outperforms state-of-the-art stereo segmentation algorithms.
Kuo-Chin Lien, Jerry D. Gibson
ICME2
2013 Bi-layer disparity remapping for handheld 3D video communications
abstract
Handheld devices with “glasses-free” autostereoscopic displays present a new opportunity for 3D video communications. 3D can enhance realism and enrich the user experience, yet it must be employed without visual discomfort. A simple shift-convergence disparity remapping technique can align a user's face throughout a 3D video call, eliminating uncomfortable crossed disparities. However, this can produce large disparities in the background that the viewer is unable to fuse. Furthermore, reducing camera separation and thus all disparities may lead to a flat appearance that does not aid realism. Using foreground/background segmentation, we propose a novel bi-layer disparity remapping algorithm to limit uncomfortable background disparities during handheld 3D video communications. A user study with the HTC Evo 3D handheld device shows that this method improves visual comfort while preserving the critical depths within the face.
Stephen Mangiat, Kuo-Chin Lien, Jerry D. Gibson
ICME3
2013 Low Complexity Video Encoding and High Complexity Decoding for UAV Reconnaissance and Surveillance
abstract
Conventional video compression schemes such as H.264/AVC use a high complexity encoder with block motion estimation (ME) and a low complexity, low latency decoder. However, unmanned aerial vehicle (UAV) reconnaissance and surveillance applications require low complexity encoders but can accommodate high complexity decoders. Moreover, the video sequences in these applications often primarily have global motion due to the known movement of the UAV and camera mounts. Motivated by this scenario, we propose and investigate a low complexity encoder with global motion based frame prediction and no block ME. For fly-over videos, our encoder achieves more than a 40% bit rate savings over a H.264 encoder with ME block size restricted to 8 × 8 and at lower complexity. We also develop a high complexity decoder based on Kalman filtering along motion trajectories and show average PSNR improvements of up to 0.5 dB with respect to a classic low complexity decoder.
Malavika Bhaskaranand, Jerry D. Gibson
ISM2
2012 Global motion compensation and spectral entropy bit allocation for low complexity video coding
abstract
Most standard video compression schemes such as H.264/AVC involve a high complexity encoder with block motion estimation (ME) engine. However, applications such as video reconnaissance and surveillance using unmanned aerial vehicles (UAVs) require a low complexity video encoder. Additionally, in such applications, the motion in the video is primarily global and due to the known movement of the camera platform. Therefore in this work, we propose and investigate a low complexity encoder with global motion based frame prediction and no block ME. We show that for videos with mostly global motion, this encoder performs better than a baseline H.264 encoder with ME block size restricted to 8×8. Furthermore, the quality degradation of this encoder with decreasing bit rate is more gradual than that of the baseline H.264 encoder since it does not need to allocate bits across motion vectors (MVs) and residue data. We also incorporate a spectral entropy based coefficient selection and quantizer design scheme that entails latency and demonstrate that it helps achieve more consistent frame quality across the video sequence.
Malavika Bhaskaranand, Jerry D. Gibson
ICC2
2011 Spatially adaptive filtering for registration artifact removal in HDR video
abstract
One method to extend the dynamic range of video captured with inexpensive cameras is to alternate the exposure time between frames and combine the information in adjacent frames using post-processing. This method requires no hardware modification, yet traditionally there is a quality tradeoff. Dynamic range expansion corresponds to an increased number of saturated pixels in individual frames, which along with occlusions contributes to registration artifacts. Therefore, we describe a “High Dynamic Range (HDR) Filter” that can mitigate these artifacts to produce a pleasing HDR video without exact frame registration. This filter builds upon the bilateral filter to smooth frames while maintaining important edges. Additionally, the filter strength locally adapts to corresponding motion vectors. Since regions with poor registration generally correspond to higher motion, smoothing here can reduce artifacts without degrading perceptual quality. Results show a significant improvement for HDR videos with fast local motion within saturated regions.
Stephen Mangiat, Jerry D. Gibson
ICIP2
2011 Rate distortion bounds for speech coding based on a perceptual distortion measure (PESQ-MOS)
abstract
We develop practical rate distortion bounds for speech coding based on composite source models and the PESQ-MOS distortion measure. Specifically, the bounds and formulated using composite source models for speech, the rate distortion function for Gaussian autoregressive sources, the classical reverse water-filling result, and conditional rate distortion theory, along with a recently devised MSE-to-PESQ_MOS mapping. The resulting rate distortion bounds are shown to lower bound the performance of the AMR, G.729, and G.718 standardized codecs, and based on the tightness of these bounds, to indicate how the performance of voice codecs might be improved.
Ying-Yi Li, Jerry D. Gibson
ICME2
2011 Routing-Aware Multiple Description Video Coding Over Mobile Ad-Hoc Networks
abstract
Supporting video transmission over error-prone mobile ad-hoc networks is becoming increasingly important as these networks become more widely deployed. We propose a routing-aware multiple description video coding approach to support video transmission over mobile ad-hoc networks with multiple path transport. We build a statistical model to estimate the packet loss probability of each packet transmitted over the network based on the standard ad-hoc routing messages and network parameters. We then estimate the frame loss probability and dynamically select reference frames in order to alleviate error propagation caused by the packet losses. We conduct experiments using the QualNet simulator that accounts for node mobility, channel properties, MAC operation, multipath routing, and traffic type. The results demonstrate that our proposed method provides 0.7-2.3 dB gains in PSNR for different video sequences under different network settings and guarantees better video quality for a selectably high number of users of the network. Furthermore, we examine the estimation accuracy of our proposed estimation model and show that our model works effectively under various network settings.
Yiting Liao, Jerry D. Gibson
IEEE Trans. Multim.2
2010 Routing-aware multiple description video coding over wireless ad-hoc networks using multiple paths
abstract
Supporting video transmission over error-prone wireless ad-hoc networks is becoming increasingly important as these networks become more widely deployed. In this paper, we propose a routing-aware multiple description video coding approach to support video transmission over wireless ad-hoc networks with path diversity. Our method uses the standard ad-hoc routing messages to estimate the possible packet losses in the networks and dynamically selects reference frames in order to alleviate error propagation caused by packet losses. We conducted experiments using the QualNet simulator that accounts for node mobility, channel properties, MAC operation, multipath routing, and traffic type. The results demonstrate that our proposed method provides up to 2.3 dB gains in PSNR and significantly improves the perceptual video quality for multiple users.
Yiting Liao, Jerry D. Gibson
ICIP2
2010 Spectral entropy-based bit allocation
abstract
In transform-based compression schemes, the task of choosing, quantizing, and coding the coefficients that best represent a signal is of prime importance. As a step in this direction, Yang and Gibson have designed a coefficient selection scheme based on Campbell's coefficient rate and spectral entropy. Building on the spectral entropy-based coefficient selection mechanism, we develop a scheme to allocate bits amongst the chosen coefficients. We show that the proposed scheme can outperform the classical method under certain conditions. We then design quantization matrices (QMs) based on the proposed bit allocation method and show that the newly designed QMs perform better than the default QMs for H.264/AVC encoding in terms of both peak signal to noise ratio (PSNR) and structural similarity (SSIM).
Malavika Bhaskaranand, Jerry D. Gibson
ISITA2
2009 Distributions of 3D DCT coefficients for video
abstract
The three-dimensional discrete cosine transform (3D DCT) has been proposed as an alternative to motion-compensated transform coding for video content. However, so far no definitive study has been done on the distribution of 3D DCT coefficients of video sequences. This study performs two goodness-of-fit tests, the Kolmogorov-Smirnov (KS) test and the x2-test, to determine the distribution that best fits the 3D DCT coefficients of the luminance components of video sequences with low motion or structured motion. The results indicate that the DC coefficient can be well approximated by a Gaussian distribution and a majority of the high-energy AC coefficients can be approximated by a Gamma distribution. Knowledge of the coefficient distributions can be used to design quantizers optimized for 3D DCT coefficients and hence achieve better coding efficiency.
Malavika Bhaskaranand, Jerry D. Gibson
ICASSP2
2009 Enhanced error resilience of video communications for burst losses using an extended ROPE algorithm
abstract
Video communications over wireless networks suffers various patterns of losses, including burst losses that cause great degradation in video quality. In this paper, we propose an algorithm based on recursive optimal per-pixel estimate (ROPE) to accurately estimate the overall distortion accounting for the loss pattern. The estimated distortion is applied to the rate-distortion (RD)-based mode selection to provide the optimal tradeoff between intra and inter coding. Simulation results show that in lossy networks, the proposed extended ROPE algorithm achieves gains in average PSNR and PSNRr,fby up to 0.8 dB and 3.6 dB respectively. This shows that the proposed algorithm can enhance error resilience for video communications over networks with burst losses.
Yiting Liao, Jerry D. Gibson
ICASSP2
2009 Automatic scene relighting for video conferencing
abstract
This paper describes a new method to automatically improve scene lighting for video conferencing by learning the photometric mapping between a lower exposure and desired exposure created using High Dynamic Range (HDR) imaging techniques. Once the mapping is learned in a calibration step, it can be used to transform all subsequent images, effectively producing higher dynamic range video without ghosting artifacts. A stereo algorithm is also described that allows multiple exposures to be taken at every frame, useful if the lighting of a scene changes significantly. Results show that this is an effective way to improve face lighting and therefore the overall experience of video conferencing.
Stephen Mangiat, Jerry D. Gibson
ICIP2
2009 Error correction scheme for uncompressed HD video over wireless
abstract
Digital transmission of uncompressed high-definition video is challenging because of its high data rate and its extreme sensitivity to bit errors. In this paper we propose a simple error correction scheme to reduce the bit error effects in the video at the receiver end of a wireless channel. Our scheme uses the large amount of spatial redundancy already present in uncompressed HD video data to provide an extra layer of protection in addition to that provided by channel coding. Thus, our method requires no change to the video signal being transmitted and is compatible with any existing solution for HD video transmission. Using simulations over a range of byte error rates, we show that our scheme effectively reduces the number of visible artifacts because of uncorrected channel errors, and provides approximately 7 dB improvement in peak signal to noise ratio.
Megha Manohara, Raghuraman Mudumbai, Jerry D. Gibson, Upamanyu Madhow
ICME3
2009 New Rate Distortion Bounds for Natural Videos Based on a Texture-Dependent Correlation Model
abstract
We revisit the classic problem of developing a spatial correlation model for natural images and videos by proposing a conditional correlation model for relatively nearby pixels that is dependent upon five parameters. The conditioning is on local texture and the optimal parameters can be calculated for a specific image or video with a mean absolute error usually smaller than 5%. We use this conditional correlation model to calculate the conditional rate distortion function when universal side information on local texture is available at both the encoder and the decoder. We demonstrate that this side information, when available, can save as much as 1 bit per pixel for selected videos at low distortions. We further study the scenario when the video frame is processed in macroblocks (MBs) or smaller blocks and calculate the rate distortion bound when the texture information is coded losslessly and optimal predictive coding is utilized to partially incorporate the correlation between the neighboring MBs or blocks. These rate distortion bounds are compared to the operational rate distortion functions generated in intra-frame coding using the H.264/AVC video coding standard.
Jerry D. Gibson
IEEE Trans. Circuits Syst. Video Technol.2
2008 On Strategies for Source Information Transmission over MIMO Systems
abstract
We consider strategies for the lossy transmission of a zero mean Gaussian source over a 2times2 MIMO channel with Rayleigh fading. The source is represented either using a single description or a multiple description code, depending on each strategy characteristic. Performance is evaluated using normalized expected distortion at the receiver, as a function of outage probability. The first strategy employs repetition coding over the two transmit antennas for the transmission of a single description representation of the source. The second strategy uses a time-shared approach to the two transmit antennas, allowing for the transmission of a multiple description representation of the source. The third and fourth strategies are based on, respectively, the Alamouti scheme and spatial multiplexing, and both of these strategies are used for the transmission of a single description representation of the source. The results show that the spatial multiplexing strategy is able to achieve the lowest distortion, and also that it is possible, with the Alamouti strategy, to obtain similar performance at a lower complexity. We finally consider the outage rates of the different strategies and observe that if a system is designed to maximize the outage rate, the corresponding distortion observed at the receiver will not be minimized.
Marco Zoffoli, Jerry D. Gibson, Marco Chiani
GLOBECOM2
2008 Perceptually weighted distortion measures and the tandem connection of speech codecs
abstract
Tandem connections of voice codecs can occur today in mobile- to-mobile calls and for certain VoIP connections. While post-filtering in tandem encodings is well-understood, the effects of the perceptual weighting filter in CELP codecs in tandem encodings has not been investigated. We study the impact of perceptual weighting filters on tandem coding using the rate distortion theory of discrete-time autoregressive (AR) sources with a frequency weighted error criterion and by examining tandem connections involving the AMR-NB codec. We show that for the usual method of calculating the perceptual weighting based upon the codec input, the perceptual weighting has a cumulative effect that is more pronounced at lower bit rates.
Niranjan Shetty, Jerry D. Gibson
ICASSP2
2008 A Transcoding-Free Multiple Description Coder for Voice over Mobile Ad-Hoc Networks
abstract
We propose a new multiple description (MD) coder design based on the Adaptive Multi-Rate Wideband (AMR- WB) coder that can support transcoding-free communication between an ad-hoc network and another network that supports the AMR-WB codec. The encoder of the MD coder consists of the standard AMR-WB coder and a bit-stream splitting block that splits the AMR-WB bit-stream into two balanced descriptions. The decoder consists of a bit-stream substitution block that substitutes the missing bits, when only one description is received, to construct a valid AMR-WB frame that can be decoded using the standard AMR-WB decoder. We show that the performance of the new MD coder is better than a previous non-transcoding- free MD coder based on AMR-WB when transcoding is required and it is significantly better than using a single description of AMR-WB over the ad-hoc network supporting transcoding-free communication.
Jagadeesh Balam, Jerry D. Gibson
WCNC2
2008 Video Capacity of WLANs With a Multiuser Perceptual Quality Constraint
abstract
As wireless local area networks (WLANs) become a part of our network infrastructure, it is critical that we understand both the performance provided to the end users and the capacity of these WLANs in terms of the number of supported flows (calls). Since it is clear that video traffic, as well as voice and data, will be carried by these networks, it is particularly important that we investigate these issues for packetized video. In this paper, we investigate the video user capacity of wireless networks subject to a multiuser perceptual quality constraint. As a particular example, we study the transmission of AVC/H.264 coded video streams over an IEEE 802.11a WLAN subject to a constraint on the quality of the delivered video experienced by r% (75%, for example) of the users of the WLAN. This work appears to be the first such effort to address this difficult but important problem. Furthermore, the methodology employed is perfectly general and can be used for different networks, video codecs, transmission channels, protocols, and perceptual quality measures.
Sayantan Choudhury, Jerry D. Gibson
IEEE Trans. Multim.3
2007 Two-Hop Two-Path Voice Communications Over a Mobile Ad-Hoc Network
abstract
We consider two-hop communication of a delay- sensitive, memoryless Gaussian source over two independent paths in an ad-hoc network. To capture the behavior an ad-hoc network we combine a path availability model and a physical layer packet loss model. The path availability model includes the effect of path failures due to node mobility and route switching delays while the physical layer model accounts for the losses in the wireless channel. An analysis using the path availability model reveals potentially long connection down times due to path failures, suggesting that path diversity may be essential to support voice communications over a mobile ad-hoc network. We compare the performance of a few path diversity based communication methods involving multiple description coding and single description coding in an ad-hoc network with packet losses due to path failures and the physical channel.
Jagadeesh Balam, Jerry D. Gibson
GLOBECOM2
2007 Information Transmission Over Fading Channels
abstract
We consider the lossy transmission of source information over Rayleigh fading channels. We investigate the utility of two channel capacity definitions, ergodic capacity, where it is assumed that the channel cycles through all fading states, and outage capacity, where the source is transmitted at a constant rate with a specified outage probability. We also study the outage rate and expected source distortion for different outage probabilities. It is observed that the outage probabilities required to maximize outage rate and minimize expected distortion are quite different. This implies that schemes based on maximizing capacity might not lead to the most efficient design for lossy transmission of source information over wireless networks. We show that minimizing expected distortion over a wireless link does not necessarily minimize the variance of the distortion, and hence parameter selection based on minimizing expected distortion can lead to a high distortion for a specific realization. We also introduce different performance measures that take into account the variance of source distortion and might be more suitable for source transmission over fading channels. Finally, we observe that in a Rayleigh fading channel in the presence of channel state information (CSI) at both the transmitter and receiver, in addition to a capacity distribution, there is a source distortion distribution at the receiver for a memoryless Gaussian source. A careful investigation of the capacity and source distortion distributions reveal that the probability of achieving the average source distortion increases with an increase in average signal to noise ratio (SNR) while the probability of achieving the average capacity does not change significantly with SNR.
Sayantan Choudhury, Jerry D. Gibson
GLOBECOM2
2007 H.264 Video over 802.11a WLANs with Multipath Fading: Parameter Interplay and Delivered Quality
abstract
Emerging as the method of choice for compressing video over WLANs, the AVC/H.264 standard is a suite of coding options and parameters whose values are to be chosen for specific videos and channel conditions. We investigate the tradeoffs in performance across the quantization parameter (QP), the group of picture size (GOPS), the payload size (PS), physical layer (PHY) data rate in 802.11a, and the average channel SNR for multipath fading channels. The interesting and sometimes surprising results are: the delivered video quality is very sensitive to the change in PSs and in many cases the PS optimizing PHY throughput cannot yield acceptable video quality; although lower QP yields better video quality when there is no packet loss, when the quality is corrupted heavily due to packet losses, there is no advantage gained for finer quantization and higher bit rate; lower GOPS, i.e., to refresh I-frames more frequently does not necessarily improve the delivered video quality and in many cases medium values of GOPS should be chosen; different types of videos have dramatically different delivered quality in the same channel which should be taken into account in network evaluation and design.
Sayantan Choudhury, Jerry D. Gibson
ICME3
2007 New Rate Distortion Bounds for Natural Videos Based on a Texture Dependent Correlation Model
abstract
We revisit the classic problem of developing a spatial correlation model for natural images and videos by proposing a conditional correlation model for relatively nearby pixels that is dependent upon five parameters. The conditioning is on local texture and the optimal parameters can be calculated for a specific image or video with a mean absolute error (MAE) usually smaller than 5%. We use this conditional correlation model to calculate the conditional rate distortion function when universal side information is available at both the encoder and the decoder. We demonstrate that this side information, when available, can save as much as 1 bit per pixel for selected videos at low distortions. We further study the scenario when the video frame is processed in macroblocks (MBs) or smaller blocks and calculate the rate distortion bound when the texture information is coded losslessly and optimal predictive coding is utilized to partially incorporate the correlation between the neighboring MBs or blocks.
Jerry D. Gibson
ISIT2
2007 Perceptual quality constrained video user capacity of 802.11a WLANs with multipath fading
abstract
The fundamental tradeoff in a WLAN with multiple video users is between the number of video users the WLAN can support and the perceptual quality each user experiences. Our previous work has investigated the delivered AVC/H.264 video quality over 802.11a WLANs with multipath fading. In this paper we study two extreme scenarios to formulate the upper and lower boundaries for video user capacity of an 802.11a WLAN operated under the contention based distributed coordination function (DCF): the worst scenario is when all users happen to be sending/receiving their intra-coded (I) frames at the same time and the best scenario is when all users happen to be coordinated in I frame refreshing. To calculate the video user capacity boundaries we study the dependence of coded video data rate on the properties of individual videos as well as the major video coding parameters: group of picture size (GOPS), quantization (QP) and video packet payload size (PS). Two heuristic formulas are proposed to calculate approximately the I frame and inter-coded (P) frame sizes respectively. The video user capacity boundaries are further explored in the context of the delivered video quality constraints investigated in our previous work.
Sayantan Choudhury, Jerry D. Gibson
IWCMC3
2007 Payload Length and Rate Adaptation for Multimedia Communications in Wireless LANs
abstract
We provide a theoretical framework for cross-layer design in multimedia communications to optimize single-user throughput by selecting the transmitted bit rate and payload size as a function of channel conditions for both additive white Gaussian noise (AWGN) and Nakagami-m fading channels. Numerical results reveal that careful payload length adaptation significantly improves the throughput performance at low signal to noise ratios (SNRs), while at higher SNRs, rate adaptation with higher payload lengths provides better throughput performance. Since we are interested in multimedia applications, we do not allow retransmissions in order to minimize latency and to reduce congestion on the wireless link and we assume that packet loss concealment will be used to compensate for lost packets. We also investigate the throughput and packet error rate performance over multipath frequency selective fading channels for typical payload sizes used in voice and video applications. We explore the difference in link adaptation thresholds for these payload sizes using the Nafteli Chayat multipath fading channel model, and we present a link adaptation scheme to maximize the throughput subject to a packet error rate constraint.
Sayantan Choudhury, Jerry D. Gibson
IEEE J. Sel. Areas Commun.2
2007 Multiple Descriptions and Path Diversity for Voice Communications Over Wireless Mesh Networks
abstract
A key feature of wireless mesh networks is that multiple independent paths through the network are available. Multiple descriptions coding is often suggested as a source coding scheme to take advantage of this path diversity. We compare multiple description (MD) coding with path diversity (PD) against a full-rate single description (SD) coder without PD, and two simple PD methods of 1) repeating a half-rate SD coder over both paths and 2) repeating the full-rate parent SD coder over the two paths. We first present a theoretical analysis comparing the average distortion per symbol in packetized communication using the above mentioned MD and PD methods to transmit a memoryless Gaussian source over additive white Gaussian noise channels. Next, using two new MD speech coders with balanced side descriptions derived from the AMR-WB and G.729 standards, we evaluate delivered voice quality using PESQ-MOS and compare MD coding against the PD methods for random and bursty packet losses. Both the theoretical analyses and the speech coding experiments show that with packet overheads, the simple PD methods may be preferable to MD coding. A new performance measure that incorporates both quality and bit rate is shown to account for the tradeoffs more explicitly.
Jagadeesh Balam, Jerry D. Gibson
IEEE Trans. Multim.2
2006 Improving the Robustness of the G.722 Wideband Speech Codec to Packet Losses for Voice Over Wlans
abstract
Since the G.722 wideband speech codec offers higher quality and naturalness than G.711, is low in complexity, has low delay, and tandems well with other codecs, it is an attractive codec for voice over IP and voice over wireless LANs. However, packet losses in G.722 not only require good concealment of the lost frame, but a lost frame results in a mismatch of the encoder/decoder states for the next correctly received frame following the lost frame. Although proprietary schemes exist, the G.722 codec has no standardized packet loss concealment (PLC) method. We present a new PLC method for G.722 and propose an efficient approach to sending the side information to resynchronize the encoder and decoder that greatly improves the robustness of the G.722 codec to packet losses
Niranjan Shetty, Jerry D. Gibson
ICASSP (5)2
2006 Matching Pursuit Decompositions of Non-Noisy Speech Signals Using Several Dictionaries
abstract
Matching pursuit (MP) provides a way to expand signals in terms of any set of time-limited functions, or atoms, called a dictionary. These decompositions are finding use in signal analysis and coding. It has been shown that a dictionary should be designed carefully, but its effects on decomposition have not been studied in detail. We look at the effects of dictionaries on the decomposition of non-noisy speech signals using MP, by five dictionaries. It is found that Gabor atoms work sufficiently well, and have fewer adverse effects in reconstruction compared to the other dictionaries. For a reconstruction to sound perceptually close to the original, a rate of 3000 atoms per second (aps) on average is required. At rates as low as 400 aps the speech remains intelligible. Finally, the use of decompositions to visualize time-frequency distributions of speech is explored.
Bob L. T. Sturm, Jerry D. Gibson
ICASSP (3)2
2006 The Interpretation of Spectral Entropy Based Upon Rate Distortion Functions
abstract
In 1960 Campbell derived a quantity that he called coefficient rate which is expressible in terms of the entropy of the process power spectral density. Later, Yang, et al showed that the spectral entropy is proportional to the logarithm of the equivalent bandwidth of the smallest frequency band containing most of the energy. Gibson, et al also showed that for discrete time AR(1) sequences, Campbell's coefficient rate and Shannon's entropy rate power are equal but that the equality does not hold for higher order AR processes. In this paper, we derive a new expression for Campbell's coefficient rate in terms of the parametrized version of the rate distortion function of a Gaussian random process with a given power spectral density subject to the MSE fidelity criterion. We also derive expressions for the entropy rate power and coefficient rate in terms of the slope of the rate distortion function for the given source and for a source with flat power spectral density
Jaewoo Jung, Jerry D. Gibson
ISIT2
2006 Multiple descriptions and path diversity using the AMR-WB speech codec for voice communication over MANETs
abstract
We compare different source diversity methods for converstional voice communication over multiple routes in a mobile ad-hoc network (MANET). A new multiple description (MD) codec based on the AMR-WB codec, with two balanced side descriptions (6.9 kbps each) is presented. We compare the performance of the MD codec against two other diversity methods, 1) duplicating speech encoded with AMR-WB at 6.6 kbps and 2) duplicating speech encoded with AMR-WB at 12.65 kbps. We show that because of the large packet headers added to each packet by typical MANET protocols, the overhead of sending the simple path diversity methods is not much larger than the overhead for sending MD streams over different paths, and the gain in speech quality we get from duplicating AMR-WB at 12.65 kbps over sending MD codec streams is significant. We compare the speech quality delivered by each of the methods under random and bursty packet loss conditions. The quality of decoded speech is evaluated using WPESQ, a wideband extension to the PESQ algorithm.
Jagadeesh Balam, Jerry D. Gibson
IWCMC2
2006 Effect of payload length variation and retransmissions on multimedia in 802.11a WLANs
abstract
Multimedia transmission over wireless local area networks is challenging due to the varying nature of the wireless channel as well as the inherent difference between multimedia and data traffic. In the MAC layer, a single bit error in the packet can lead to the entire packet being discarded. This results in a higher packet error rate for larger payload sizes. Retransmission due to packet errors causes the contention window to double, and this leads to a decrease in throughput if the wireless channel does not improve for the retransmitted packets. Hence, throughput is a function of packet payload length as well as the maximum number of allowable retransmissions. In this paper, we investigate the effect of payload length adaptation and retransmissions on the throughput and capacity of multimedia users. Numerical results and simulations reveal that careful payload adaptation significantly improves the throughput performance at low signal to noise ratios (SNRs). It is also observed that excessive retransmissions reduce the effective throughput, thereby decreasing the capacity of multimedia users in the presence of data users. Since multimedia traffic is more latency constrained and less error constrained, by carefully selecting the payload length and maximum number of allowable retransmissions based on the channel conditions, a greater number of multimedia users can be supported.
Sayantan Choudhury, Irfan Sheriff, Jerry D. Gibson, Elizabeth M. Belding
IWCMC3
2006 Voice capacity under quality constraints for IEEE 802.11a based WLANs
abstract
The communication of voice over wireless local area networks (WLANs) is influenced by the choice of speech codec, packetization interval and PHY layer bit rates. These choices affect the number of voice users that can be supported on the WLAN as well as the speech quality experienced by each user. We investigate the effect of different combinations of these parameters for a 802.11a WLAN in different channel conditions with the objective of maximizing the number of voice users supported on the WLAN subject to a quality constraint. We use an indicator for assessing the speech quality experienced by a single user in a WLAN, based on a Perceptual Evaluation of Speech Quality (PESQ) Mean Opinion Score (MOS) constraint and the probability of a voice user achieving this constraint. The contributions of this paper are three-fold. First, a PHY layer rate adaptation scheme is proposed, in which the operating rate for each Signal to Noise Ratio (SNR) is chosen as the one that maximizes the capacity given a quality constraint. Second, based on the PHY layer rate adaptation scheme, we evaluate the effect of the choice of codec and voice payload size on the capacity values obtainable at different SNRs. Finally, we show the effect of channel conditions and the tightening of quality constraints on the capacity values at different SNRs.
Niranjan Shetty, Sayantan Choudhury, Jerry D. Gibson
IWCMC3
2006 Payload Length and Rate Adaptation for Throughput Optimization in Wireless LANs
abstract
Wireless local area networks offer a range of transmitted data rates that are to be selected according to estimated channel conditions. However, due to packet overheads and contention times introduced by the CSMA/CA multiple access protocol, effective throughput is much less than the transmitted bit rates. Furthermore, if there is even a single bit error in the packet, the entire packet is discarded and the packet is retransmitted. This causes the effective throughput to be a function of the packet payload length. We provide a theoretical framework to optimize single-user throughput by selecting the transmitted bit rate and payload size as a function of channel conditions for both additive white Gaussian noise (AWGN) and Nakagami-m fading channels. Numerical results reveal that careful payload adaptation significantly improves the throughput performance at low signal to noise ratios (SNRs) while at higher SNRs, rate adaptation with higher payload lengths provides better performance. We compare the range of SNRs over which payload length adaptation is crucial for AWGN and different fading channels based on the m-parameter of Nakagami fading realization. We then specify SNR values for switching between transmitted bit rates and payload lengths such that the effective throughput is maximized.
Sayantan Choudhury, Jerry D. Gibson
VTC Spring2
2006 Joint PHY/MAC based link adaptation for wireless LANs with multipath fading
abstract
Wireless local area networks offer a range of transmitted data rates that are to be selected according to estimated channel conditions. However, due to packet overheads and contention times introduced by the CSMA/CA multiple access protocol, effective throughput is much less than the nominal data rates. Most multimedia applications use small payload sizes in order to ensure reliable, low latency transmission. This results in a further loss in effective throughput thereby reducing network capacity drastically. Thus, a cross-layer based design approach along with link adaptation is required to improve the network performance under different channel conditions. We investigate the effect of payload size variations on single-user throughput for both non-fading and multipath fading environments. We explore the difference in link adaptation thresholds for different payload sizes with varying channel characteristics. A link adaptation scheme to maximize the throughput with a packet error constraint is presented
Sayantan Choudhury, Jerry D. Gibson
WCNC2
2006 Structures for SNR scalable speech coding
abstract
SNR scalable speech coding is desirable for a number of network multimedia applications, but relatively few SNR-scalable speech coders exist for operation at rates below 16 kb/s. We investigate several SNR scalable source coding structures and define the new concepts of dependent and independent SNR scalability, where independent SNR scalable coders depend on the core layer coder only through the core layer output. Independent SNR scalable structures offer the possibility of providing bit rate scalable functionality to existing nonscalable coders and standards. We show that the MPEG-4 scalable coders are examples of dependent SNR scalable coders, and we introduce a new independent SNR scalable coder called CELPTree, which has the additional advantage of being low delay. We compare the performance of the MPEG-4 coders and CELPTree for both clean and noisy speech, and we examine the effects of frequency-weighted distortion measures in the enhancement layers of SNR scalable speech coders.
Jerry D. Gibson
IEEE Trans. Speech Audio Process.2
2005 Allowing bit errors in speech over wireless LANs
Ian D. Chakeres, Elizabeth M. Belding, Allen Gersho, Jerry D. Gibson
Comput. Commun.5
2005 Coefficient rate and lossy source coding
abstract
Campbell derived and defined a quantity called the coefficient rate of a random process that involves the process spectral entropy. In this correspondence, his interpretation is substantiated with two new derivations. One derivation tightens the connection to source bandwidth, while the second derivation implies a specific approach to adaptive coefficient selection in realization-adaptive approaches to compression. After a discussion on the role the coefficient rate plays in adaptive source coding, a quantity called Campbell bandwidth is defined based on its connection to source bandwidth and is contrasted with Fourier bandwidth and Shannon bandwidth. The connection between coefficient rate and reverse water-filling from rate distortion theory is also demonstrated.
Wenye Yang, Jerry D. Gibson
IEEE Trans. Inf. Theory2
2005 A QRD-M/Kalman filter-based detection and channel estimation algorithm for MIMO-OFDM systems
abstract
The use of multiple transmit/receive antennas forming a multiple-input multiple-output (MIMO) system can significantly enhance channel capacity. This paper considers a V-BLAST-type combination of orthogonal frequency-division multiplexing (OFDM) with MIMO (MIMO-OFDM) for enhanced spectral efficiency and multiuser downlink throughput. A new joint data detection and channel estimation algorithm for MIMO-OFDM is proposed which combines the QRD-M algorithm and Kalman filter. The individual channels between antenna elements are tracked using a Kalman filter, and the QRD-M algorithm uses a limited tree search to approximate the maximum-likelihood detector. A closed-form symbol-error rate, conditioned on a static channel realization, is presented for the M=1 case with QPSK modulation. An adaptive complexity QRD-M algorithm (AC-QRD-M) is also considered which assigns different values of M to each subcarrier according to its estimated received power. A rule for choosing M using subcarrier powers is obtained using a kernel density estimate combined with the Lloyd-Max algorithm.
Kyeong Jin Kim, Jiang Yue 0002, Ronald A. Iltis, Jerry D. Gibson
IEEE Trans. Wirel. Commun.4
2004 Tandem voice communications: digital cellular, VoIP, and voice over Wi-Fi
abstract
We consider the problems of voice over wireless LANs and voice communications over heterogeneous asynchronous tandem networks, including digital cellular, VoIP, and voice over Wi-Fi. For voice over Wi-Fi, we minimize retransmissions by combining new packetization methods and packet loss concealment approaches. We demonstrate that tandem network connections can suffer significant loss in voice quality, even with ideal channels. We present a particular example of a VoIP network in tandem with voice over Wi-Fi and study end-to-end performance for different voice codecs, bit error rates, packet loss rates, and packet loss concealment methods.
Jerry D. Gibson
GLOBECOM1
2004 A multiple description speech coder based on AMR-WB for mobile ad hoc networks
abstract
To address the challenging task of achieving effective voice communication over mobile ad hoc networks (MANETs), we introduce a new multiple description (MD) speech coder based on the AMR-WB (adaptive multirate wideband) standard. The MD coder splits the bitstream of the AMR-WB coder into two redundant sub-streams by directly selecting overlapping subsets of encoded data generated for each frame. The sub-streams are transmitted on different network paths. When both sub-streams arrive at the decoder, an output identical to that of AMR-WB is recovered. If only one substream arrives at the decoder, degraded, but still acceptable, speech quality is obtained. The performance of the multiple description coding system is tested for MANETS with a network simulator and an informal listening test. The results demonstrate that this approach makes effective use of the channel capacity and provides reliable end-to-end connections for voice communication in MANETs.
Allen Gersho, Jerry D. Gibson, Vladimir Cuperman
ICASSP (1)3
2004 Selective bit-error checking at the MAC layer for voice over mobile ad hoc networks with IEEE 802.11
abstract
Mobile ad hoc networks (MANET) have more severe operating conditions than traditional wireless networks. The MAC protocol of IEEE 802.11 mitigates collisions and ensures error-free packet transmissions at the cost of limiting capacity and increasing latency. For voice transmission over MANETs this cost should be minimized. We propose and examine selective error checking (SEC) at the MAC layer of 802.11 that takes advantage of the fact that many of the speech bits can tolerate errors while other bits must be protected for effective reconstruction of the speech. Simulation results demonstrate that the network performance and the speech quality are substantially improved by modifying the MAC layer with SEC to suit a particular GSM speech compression standard, the narrow-band adaptive multirate (NB-AMR) coder operating at a rate of 7.95 kbps.
I. D. Chakares, Allen Gersho, Elizabeth M. Belding, Jerry D. Gibson
WCNC5
2004 Application of NB/WB AMR speech codecs in the 30-kHz TDMA system
abstract
A new system enhancement method is proposed for the EIA/TIA-136 system offering both channel operational range extension and improved performance within the current operational range. The existing time-division multiple-access (TDMA) (136) speech codec, the IS-641 enhanced full rate vocoder, operates at a fixed bit rate and does not allow the reallocation of bits to channel error protection as channel conditions degrade. The research presented here investigates the application of the narrow-band adaptive multirate (NB-AMR) speech codec and the wide-band AMR (WB-AMR) codec, both originally designed for the 200 kHz GSM channel, in the TDMA (TIA/EIA-136) 30-kHz system. In particular, we investigate adaptively allocating bits between NB/WB speech coding and error control coding within the limited channel bandwidth. Four modes out of 17 have been carefully chosen for the new TDMA/AMR system. Switching between codec rates as channel conditions change produces range extension below a C/I of 15 dB while also improving performance in the existing operational range above 15 dB. We keep the time slot formats unchanged so that our method is completely compatible with existing 136 systems.
Jerry D. Gibson
IEEE Trans. Wirel. Commun.3
2003 Realization Adaptive Strategies and a Comparison of Fourier Bandwidth, Shannon Bandwidth, and Campbell Bandwidth
abstract
In 1960, Campbell derived a quantity that he defined as the coefficient rate of a random process that involves the process spectral entropy. However, no potential applications of the coefficient rate were identified. Two new derivations of Campbell's rate coefficient rate are presented. One derivation solidifies the interpretation of this quantity as a coefficient rate and allows definition of an effective bandwidth for the process. The second derivation implies a new approach for realization adaptive source compression. The coefficient rate can be used for realization adaptive coefficient selection in a sequence of source representations. Furthermore, the effective bandwidth is designated as Campbell bandwidth and contrasted with Fourier bandwidth and Shannon bandwidth. Several specific examples are presented that illustrate the differences among the three quantities.
Wenye Yang, Jerry D. Gibson
DCC3
2003 Channel estimation and data detection for MIMO-OFDM systems
abstract
The use of multiple antennas at both the transmitter and receiver can significantly increase the channel capacity. These systems are called the multiple-input multiple-output (MIMO) systems. By using orthogonal frequency division multiplexing (OFDM) transmission techniques, the MIMO-OFDM system can achieve high spectral efficiency, which makes it an attractive candidate for high-data-rate wireless applications. In this paper, we propose a convolutionally coded MIMO-OFDM system with EM-based channel estimation and a QRD-M data detection algorithm. In our systems, one training symbol is transmitted from each transmit antenna for the MIMO channel estimation at the receiver. With the channel estimates available, we apply the QRD-M algorithm on the estimated channel matrix for suboptimal data detection with reasonable computational cost. The bit error rate (BER) and packet error rate (PER) performance of the MIMO-OFDM systems are compared. In the simulations, the bit error rate performance of our systems is 9 (or 5) dB better than that of uncoded (or coded) BLAST systems.
Jiang Yue 0002, Kyeong Jin Kim, Jerry D. Gibson, Ronald A. Iltis
GLOBECOM3
2003 A new discrete spectral modeling method and an application to CELP coding
abstract
In this letter, a new discrete spectral modeling method is proposed based on a comparison of distance measure performance in an adaptive filtering context. It is shown that a new fast converging adaptive algorithm yields a more accurate estimate of the spectral envelope of the speech spectrum by minimizing COSH distance rather than the Itakura-Saito distance measure. We apply discrete spectral all-pole modeling to code-excited linear predictive coding by refining the short-term synthesis filter coefficients originally obtained by the linear prediction method. Simulation results show the enhanced harmonic and formant structure in the speech spectrum and that better speech quality is obtained.
Jerry D. Gibson
IEEE Signal Process. Lett.2
2002 Bandwidth scalable speech coding and rate switching
abstract
Bit rate scalability in speech coders can be used to achieve SNR scalable speech coding or bandwidth scalable coding of speech and audio. While SNR scalability has been studied theoretically and practically, the fundamental limits of bandwidth scalability have received substantially less attention. We utilize classical rate distortion theory subject to a weighted squared error fidelity criterion to analyze the performance of subband structured bandwidth scalable coders. We propose the application of bandwidth scalable speech coding for rate switching in wireless environments and develop a bandwidth scalable coder that builds upon narrowband adaptive multirate speech coders. The bandwidth scalable coders exhibit performance competitive with wideband speech coders at the same rates.
Jerry D. Gibson
ICASSP2
2002 Performance of OFDM systems with space-time coding
abstract
Space-time coding is a promising channel coding technique for wireless communications. Combined with antenna diversity, space-time coding can provide improved performance while maintaining the same data transmission rate, which makes it an attractive candidate for high-data-rate transmission systems such as orthogonal frequency division multiplexing (OFDM) systems. In this paper, we compare the performance of space-time coded OFDM systems with that of OFDM systems defined by IEEE 802.11a over AWGN channels and slow fading channels. Although space-time block coding does not have coding gain, it can be used as an inner coding method to obtain diversity gain. Space-time trellis coding obtains 1 dB coding gain and 9 dB diversity gain.
Jiang Yue 0002, Jerry D. Gibson
WCNC2
2001 Universal successive refinement of CELP speech coders
abstract
Many speech coding standards are based upon code-excited linear prediction (CELP), and it is desirable to develop layered coding methods that are compatible with this installed base of coders. We propose a layered speech coding structure that is universally compatible with all CELP-based coders. This structure encodes the reconstruction error signal from layer 1 using a low-delay, adaptive tree coder based upon the mean squared error (MSE) criterion. We note that rate distortion optimal successive refinement is achievable using two different distortion criteria and we derive expressions for the rate distortion function under autoregressive Gaussian assumptions on the source and the two different distortion measures. We demonstrate the universality of the approach by developing two-layer coders for a 3.65 kbps CELP coder, G.723.1, and G.729. We show that our layering method is favorably competitive with the MPEG-4 layering method at 8.7 kbps for both clean and noisy speech. Using tree coding and the MSE criterion in layer 2 improves speech naturalness when coding noisy speech.
Jerry D. Gibson
ICASSP2
2001 Frequency selectivity via the SpEnt methodology for wideband speech compression
abstract
In speech and audio coding, frequency selectivity of the basis functions is an important property of the codec. The more precise the frequency selectivity, the less chance there is for audible coding effects due to uncanceled aliasing. We use Campbell's (1960) coefficient rate and the spectral entropy (SpEnt) of the source random process as a guide to formulate adaptive nonuniform modulated lapped biorthogonal transforms (NMLBT). The use of the NMLBT allows for efficient implementation of a time-varying transform which possesses both good frequency and time resolution at all instances, without the need for transitional filters. By coupling the SpEnt methodology with modulated lapped biorthogonal transforms (MLBT), we develop band combining strategies to produce an adaptive NMLBT. Due to the nature of the SpEnt methodology, the new frequency selection process comprises a non-linear approximation method to determine the best n basis functions to represent the current speech frame. We implement a wideband speech compression scheme based on this strategy and verify its improved performance in coding speech and audio signals at 16 and 24 kbps.
Mark G. Kokes, Jerry D. Gibson
ICASSP2
2001 Parameter interpolation to enhance the frame erasure robustness of CELP coders in packet networks
abstract
Frame erasure (FE) robustness is an important quality measure for voice over IP networks (VoIP). The recovery of the erased frames from the received information is crucial to realize this robustness. We allow the lost frames to be recovered from both the "previous" and "next" good frames. We first give quantitative distortion comparisons between predictive and interpolative frame recovery. Then we add FE-robust LSF coding modes to the popular ITU G.723.1 and G.729 CELP coders. These FE-robust modes utilize intraframe LSF VQ and invoke no bit-rate increase for the G.723.1 coder and a small increase (0.4 kb/s) for G.729. Simulations show that FE robust coding with interpolation achieves average spectral distortions 0.7-1.8 dB smaller than that of the original coders. Significant quality improvement was achieved by combined implementation of FE robust coding, LSF and pitch interpolation, and a proposed fixed codebook excitation recovery method.
Jerry D. Gibson
ICASSP2
2001 Network model coding gain for smoothed descriptions coding
abstract
Numerous image coders have been proposed which improve the reconstructed image quality in the presence of transmission errors. Some of the more successful examples involve partitioning the image during the encoding process, and transmitting each 'description' of the image over a separate channel. However, analysis of the performance of such coders, given a network model, has been lacking until when the performance of generic coders and Gaussian sources was investigated. We build on this previous work by analyzing the performance of a specific partitioned coding algorithm called 'smoothed descriptions coding' coupled with a network model. Unlike other commonly considered coders, this coder has the pleasing feature of adding virtually no overhead to the encoded bit-stream. The results show that smoothed description coding can provide a gain over single-description coding. For test images, the largest gain was found to be approximately 7.4 dB at a moderate network loading of /spl rho/=0.4.
Justin Ridge, Jerry D. Gibson
ICC2
2001 Multiframe video coding for improved performance over wireless channels
abstract
We propose and evaluate a multi-frame extension to block motion compensation (BMC) coding of videoconferencing-type video signals for wireless channels. The multi-frame BMC (MF-BMC) coder makes use of the redundancy that exists across multiple frames in typical videoconferencing sequences to achieve additional compression over that obtained by using the single frame BMC (SF-BMC) approach, such as in the base-level H.263 codec. The MF-BMC approach also has an inherent ability of overcoming some transmission errors and is thus more robust when compared to the SF-BMC approach. We model the error propagation process in MF-BMC coding as a multiple Markov chain and use Markov chain analysis to infer that the use of multiple frames in motion compensation increases robustness. The Markov chain analysis is also used to devise a simple scheme which randomizes the selection of the frame (amongst the multiple previous frames) used in BMC to achieve additional robustness. The MF-BMC coders proposed are a multi-frame extension of the base level H.263 coder and are found to be more robust than the base level H.263 coder when subjected to simulated errors commonly encountered on wireless channels.
Madhukar Budagavi, Jerry D. Gibson
IEEE Trans. Image Process.2
2000 A Three-Layer, Two Description Image Coder
abstract
Summary form only given. We propose a new DCT-based image coder that provides a balance in three areas: compression, progressive coding, and robustness. We design a unique form of progressive coding. The scheme is described as a three-layer, two description image coder because it provides three levels of progressive coding where the last two layers refine the image independent of each other. This is the first image (source) coder of this type. Our coding technique increases robustness to lost information when compared to traditional progressive coding methods, and it maintains progressivity. The proposed coding technique implements two levels of coefficient scaling. The goal of the first level of coefficient scaling is to obtain a coarse representation of the original DCT coefficients by using large quantization step sizes. Next, the coarse representation of each DCT block is subtracted from the original, producing a quantization error or residual, which is quantized with smaller step sizes. At this point, we have two quantized representations taken from the original DCT coefficients. To create two layers of image refinement, the residual data is divided into two descriptions such that adding either description to the coarse representation provides a significant improvement in image quality. Thus, the coding scheme creates three layers of successive refinement of image quality, where the first layer provides a coarse or base level of quality and the two following layers provide independent, full image descriptions for refinement. We show that the proposed layered image coding scheme provides good performance for practical packet erasure rates experienced in the two refinement layers.
Fred W. Ware, Jerry D. Gibson
Data Compression Conference2
2000 Permuted smoothed descriptions and refinement coding for images
abstract
We consider the problem of transmitting compressed still images over lossy channels. In particular, we examine the situation where the data stream is partitioned into two independent channels, as is often considered in the multiple descriptions approach to image compression. We introduce a coder design called smoothed descriptions, which matches a data partitioning method utilized at the encoder to the error concealment technique employed at the decoder. This approach has the advantage of inserting minimal overhead into the transmitted data streams, so that system performance is undiminished when there are no packet losses over the channel. We show that, by using a combination of DC averaging and maximal smoothing to conceal errors, performance comparable to or better than multiple descriptions can be achieved for packet loss rates up to 5%. By adding a feedback loop that requests retransmission of important data, we also demonstrate the need to exploit latency whenever possible.
Justin Ridge, Fred W. Ware, Jerry D. Gibson
IEEE J. Sel. Areas Commun.3
1999 Image refinement for lossy channels with relaxed latency constraints
abstract
Numerous image coding approaches have been devised which result in high-quality reconstructions despite the presence of channel errors. This paper considers the case where a channel degrades to such a degree that decoder feedback must be applied in order for the reconstruction to be 'acceptable'. It considers several means of generating refinement data, and applies these methods to a multi-channel coding technique previously dubbed 'smoothed description coding'. Results are presented which demonstrate that satisfactory reconstruction quality can be attained in extremely lossy channels.
Justin Ridge, Fred W. Ware, Jerry D. Gibson
WCNC3
1999 Tree coding combined with TDHS for speech coding at 6.4 and 4.8 kbps
Insung Lee, Jerry D. Gibson
Speech Commun.2
1999 Efficient pitch filter encoding for variable rate speech processing
abstract
Analysis-by-synthesis techniques are used in a wide variety of speech coding standards and applications for rates below 16 kbps. The presence of a long-term predictor, commonly known as the adaptive codebook, is critical to coder performance at the lower rates. Unfortunately, the encoding rate and computational requirements for high-quality encoding of pitch filter parameters can be excessive. Several popular approaches explore the trade-off between predictor order, allocated bit rate, and computational requirements for long-term predictor optimization. We investigate the relative performance of several long-term predictor structures and present a new approach to vector quantization of the pitch filter coefficients having a subjective quality equivalent to other schemes, but at a lower coding rate and requiring significantly less closed-loop computation. The performance is evaluated in a variable-rate CELP coder at an average rate of 2 kbps and in Federal Standard 1016 CELP.
Stan A. McClellan, Jerry D. Gibson, B. Keith Rutherford
IEEE Trans. Speech Audio Process.2
1998 Speech coding in mobile radio communications
abstract
Speech coding, the efficient representation of speech in digital form, is one of the key technologies in current and evolving digital cellular and wireless voice communications offerings. The speech coders in existing standards exhibit a level of sophistication and performance unimaginable just 15 years ago. We outline the characteristics of the mobile communications problem with respect to speech coders and point out the principal issues in speech coder design for these applications. Speech coding methods in existing mobile communications standards are described and contrasted. The limitations imposed by the wireless channel and by background impairments are discussed, and approaches to addressing their resulting effects are presented. Suggestions for future research in speech coding for the mobile communications problem are outlined.
Madhukar Budagavi, Jerry D. Gibson
Proc. IEEE2
1998 Lossy Source Coding
abstract
Lossy coding of speech, high-quality audio, still images, and video is commonplace today. However, in 1948, few lossy compression systems were in service. Shannon introduced and developed the theory of source coding with a fidelity criterion, also called rate-distortion theory. For the first 25 years of its existence, rate-distortion theory had relatively little impact on the methods and systems actually used to compress real sources. Today, however, rate-distortion theoretic concepts are an important component of many lossy compression techniques and standards. We chronicle the development of rate-distortion theory and provide an overview of its influence on the practice of lossy source coding.
Toby Berger, Jerry D. Gibson
IEEE Trans. Inf. Theory2
1997 Time-correlation analysis of a class of nonstationary signals with an application to radar imaging
abstract
A new method of nonstationary signal analysis, called time-correlation analysis (TCA), is applied to a class of nonstationary random signals containing time distortion. The method reveals a relationship between the TCA summary statistics and the distortion and leads to two nonparametric estimators for the distortion function. An example is given that demonstrates the application of the TCA method to motion compensation problems in radar imaging.
Ta-Hsin Li, Jerry D. Gibson
ICASSP2
1997 Error Propagation in Motion Compensated Video Over Wireless Channels
abstract
The use of multiple frames in block motion compensation (BMC) provides us with a new framework for increasing the robustness of video coders. We study the error robustness properties of multiframe BMC (MF-BMC) coders by modeling the error propagation process in the MF-BMC approach as a multiple Markov chain. Using Markov chain analysis we find that the probability of error propagation can decrease (i.e. robustness can increase) by (i) the use of additional frames in BMC, and by (ii) decreasing the probability with which blocks in the immediate previous frame are chosen for motion compensation. We outline a simple scheme that modifies the probabilities of macroblock prediction to achieve additional robustness. The MF-BMC coders proposed in this paper are a multiframe extension of the base level H.263 coder and are found to be more robust than the base level H.263 coder when subjected to simulated errors commonly encountered on wireless channels.
Madhukar Budagavi, Jerry D. Gibson
ICIP (2)2
1997 Variable-rate CELP based on subband flatness
abstract
Code-excited linear prediction (CELP) is the predominant methodology for communications quality speech coding below 8 kbps, and several variable-rate CELP schemes have been discussed in the literature, including QCELP, the variable-rate wideband digital cellular mobile radio speech coding standard specified in IS-95. A key component of these speech coders is the detection and classification of speech activity, and several cues for rate variation have been studied, such as measuring the short-term speech energy, deciding whether the speech is voiced or unvoiced, or making more sophisticated phonetic classifications. We present a new method for rate variation based on a measure of subband spectral flatness, called spectral entropy. Spectral entropy is a normalized indicator of the texture of the input spectrum and is thus less dependent on speech and background noise energy variations. We present some results on the use of spectral entropy for voice activity detection across subbands and then evaluate using spectral entropy for deriving mode and rate allocation cues for a variable-rate CELP coder operating at an average rate of 2 kbps. To achieve communications quality speech at this rate, we develop a new split-band vector quantization (VQ) technique for representing the line spectral pairs and a multiple codebook approach for efficiently quantizing the coefficients of a three-tap pitch predictor, called lag-indexed VQ.
Stan A. McClellan, Jerry D. Gibson
IEEE Trans. Speech Audio Process.2
1996 Lag-indexed VQ for pitch filter coding
abstract
The presence of a pitch predictor is critical to low-rate performance of CELP coders. Unfortunately the rate required for high-quality encoding of pitch filter parameters is often a large fraction of the available bandwidth. Moreover, the application of analysis-by-synthesis techniques to pitch filter optimization can require excessive computation. We present a vector quantization (VQ) approach for coding pitch filter parameters which maintains a subjective quality equivalent to other coding schemes while requiring lower (variable) rate with less closed-loop computation than other VQ techniques.
Stan A. McClellan, Jerry D. Gibson
ICASSP2
1996 Packet video for heterogeneous networks using CU-SeeMe
abstract
Software based desktop videoconferencing tools are developed to demonstrate techniques necessary for video delivery in heterogeneous packet networks. Pyramidal compression, congestion avoidance, end-to-end delivery, and predictive rate control results are presented.
Tom Brown, Sharif Sazzad, Charles Schroeder, Pierce E. Cantrell, Jerry D. Gibson
ICIP (1)5
1996 Speech analysis and segmentation by parametric filtering
abstract
A new set of digital signal processing techniques for detecting changes in a speech signal is considered. The overall approach is called parametric filtering, and it yields several promising new diagnostics for speech analysis and segmentation including in particular, the demodulated lag-one autocorrelation /spl gamma//sub /spl theta//(/spl eta/), the time-correlation analysis plot, and the /spl gamma//sub /spl theta//(/spl eta/)-based distortion measures. Initial experiments described in this paper establish the potential significance of the parametric filtering method and these new diagnostics for speech analysis and segmentation.
Ta-Hsin Li, Jerry D. Gibson
IEEE Trans. Speech Audio Process.2
1995 Image coding with uniform and piecewise-uniform vector quantizers
abstract
New lattice vector quantizer design procedures for nonuniform sources that yield excellent performance while retaining the structure required for fast quantization are described. Analytical methods for truncating and scaling lattices to be used in vector quantization are given, and an analytical technique for piecewise-linear multidimensional companding is presented. The uniform and piecewise-uniform lattice vector quantizers are then used to quantize the discrete cosine transform coefficients of images, and their objective and subjective performance and complexity are contrasted with other lattice vector quantizers and with LBG training-mode designs.
Dae-Gwon Jeong, Jerry D. Gibson
IEEE Trans. Image Process.2
1994 Spectral entropy: an alternative indicator for rate allocation?
abstract
We introduce an approach to speech segment classification that differs from the usual energy, correlation, and zero-crossing criteria. Instead, we measure the gross shape of the short-term speech spectrum using spectral entropy to derive some indication of effective bandwidth. We propose to lower the required encoding rate by compensating for dynamic variations in signal bandwidth. We show that the spectral entropy can be used effectively to determine regions of voicing activity even in extreme background noise.>
Stan A. McClellan, Jerry D. Gibson
ICASSP (1)2
1994 A constrained joint source/channel coder design
abstract
The design of joint source/channel coders in situations where there is residual redundancy at the output of the source coder is examined. It has previously been shown that this residual redundancy can be used to provide error protection without a channel coder. In this paper, this approach is extended to conventional source coder/convolutional coder combinations. A family of nonbinary encoders is developed which more efficiently use the residual redundancy in the source coder output. It is shown through simulation results that the proposed systems outperform conventional source-channel coder pairs with gains of greater than 9 dB in the reconstruction SNR at high probability of error.>
Khalid Sayood, Fuling Liu, Jerry D. Gibson
IEEE J. Sel. Areas Commun.3
1994 Analysis of the smoothed residual-driven algorithm for speech coders
abstract
The smoothed residual-driven (SRD) algorithm was devised recently to adapt the short-term predictor coefficients in low-delay speech coders. This algorithm uses a smoothed version of the prediction residual for coefficient adaptation to provide good-quality speech while maintaining robustness to channel errors. The SRD algorithm is further investigated here to understand the effect of the smoothing approximation on the algorithm tracking capability for tone, AR, and speech inputs in the presence of channel errors. The transversal structure SRD algorithm is studied here. The relationship between the minimum MSE values with the SRD and signal-driven (SD) algorithms is derived as a function of the predictor coefficients. Tracking analysis shows that the SRD algorithm can track the input if the step size /spl mu/ is sufficiently small. The LMS-SRD algorithm is stable and succeeds in tracking all inputs. The LMS-SRD algorithm is compared with other existing algorithms such as the SD, residual-driven (RD), and LMS-CCITT algorithms.>
Seung H. Nam, Jerry D. Gibson
IEEE Trans. Speech Audio Process.2
1994 Low delay tree coding of speech at 8 kbit/s
abstract
A low delay tree coder for coding speech at 8 kbit/s that incorporates a new backward coefficient adaptation structure is presented. The coder has an algorithmic delay of 2.5 ms, produces good quality speech over ideal channels, and maintains robustness to bit errors for independent bit error rates up to and including 10/sup -2/. The tree coder consists of a randomly populated fractional rate tree, a weighted squared error distortion measure, the (M,l) tree search algorithm, an incremental path map symbol release rule, a long-term predictor, an all-pole short-term predictor, and the newly-obtained parameter adaptation algorithms; The sensitivity to channel errors is reduced by using the receiver excitation sequence for adaptation of the short-term predictor coefficients, but the ability to track rapid changes in the speech is retained by shaping the excitation sequence with an all-zero filter. The output speech of the tree coder is evaluated using segmental signal-to-noise ratios (SNRSEG), narrow band spectrograms, a paired comparison with pulse code modulation (PCM), and an equivalent noise paired comparison using modulated noise reference unit (MNRU).>
Hong Chae Woo, Jerry D. Gibson
IEEE Trans. Speech Audio Process.2
1994 A bound on the rate of a system for encoding an unknown Gaussian autoregressive source
abstract
To obtain the rate distortion function of an autoregressive (AR) source when the source parameters are fixed but unknown, it is common simply to insert the estimated parameters into the known-spectrum rate distortion expressions. This approach, although asymptotically correct, ignores any error in estimating the parameters, ignores distortion incurred by the need to encode the parameters for transmission to the receiver, and does not include a component for the side information. We derive a lower bound to the rate distortion performance of a relatively general system for encoding an unknown Gaussian AR source with respect to the mean-squared error (MSE) fidelity criterion. Asymptotically in the estimator frame length n and in the encoder blocklength N, the lower bound equals the usual Gaussian rate distortion function for small distortions.>
Jerry D. Gibson
IEEE Trans. Inf. Theory1
1993 Tree coding combined with harmonic scaling of speech at 6.4 kbps
Insung Lee, Jerry D. Gibson
ICASSP (2)2
1993 Analysis of the smoothed residual driven algorithm for speech coders
Seung H. Nam, Jerry D. Gibson
ICASSP (2)2
1993 Uniform and piecewise uniform lattice vector quantization for memoryless Gaussian and Laplacian sources
abstract
Lattice vector quantizer design procedures for nonuniform sources are presented. The procedures yield lattice vector quantizers with excellent performance and retaining the structure required for fast quantization. Analytical methods for truncating and scaling lattices to be used in vector quantizations are given, and their utility is demonstrated for independent and identically distributed (i.i.d.) Gaussian and Laplacian sources. An analytical technique for piecewise linear multidimensional compandor designs is evaluated for i.i.d. Gaussian and Laplacian sources by comparing its performance to that of the other vector quantizers.>
Dae-Gwon Jeong, Jerry D. Gibson
IEEE Trans. Inf. Theory2
1991 Smoothed DPCM codes
abstract
Interpolative differential pulse code modulation (IDPCM) is known to outperform DPCM at rate 1 b/sample for several synthetic source models. The authors use minimum-mean-squared error (MMSE) fixed-lag smoothing with DPCM to develop a code generator using delayed decoding. This smoothed DPCM (SDPCM) code generator is compared to DPCM and IDPCM code generators at rates 1 and 2 b/sample for tree coding several synthetic sources and to a DPCM code generator at 2 b/sample for speech sources.>
Wen-Whei Chang, Jerry D. Gibson
IEEE Trans. Commun.2
1991 Fractional rate multitree speech coding
abstract
The authors present both forward and backward adaptive speech coders that operate at 9.6, 12, and 16 kb/s using integer and fractional rate trees, weighted squared error distortion measures, the (M,L) tree search algorithm, and incremental path map symbol release. They introduce the concept of multitree source codes and illustrate how the multitree structure allows scalar quantizer-based codes and scalar adaptation rules to be used for fractional rate tree coding. With a frequency weighted distortion measure, the forward and backward adaptive multitree coders produce near toll quality speech at 16 kb/s, while the backward adaptive 9.6 kb/s multitree coder substantially outperforms adaptive predictive coding and has an encoding delay of less than 2 ms. Performance results are present in terms of unweighted and weighted signal-to-noise ratio and segmental signal-to-noise ratio, sound spectrograms, and subjective listening tests.>
Jerry D. Gibson, Wen-Whei Chang
IEEE Trans. Commun.1
1990 A comparison of backward adaptive prediction algorithms in low delay speech coders
abstract
Simulation results are presented comparing the performances of four backward adaptive lattice algorithms for updating the short-term predictors in adaptive predictive coding (APC) and differential pulse code modulation (DPCM) code generators for low-delay tree coding of speech at 16 and 9.6 kb/s. The algorithms studied are the least-squares lattice, the exponential window lattice, the signal-driven lattice, and the residual-driven lattice. Ideal channels are emphasized, and comparisons are based primarily upon frequency-weighted signal-to-noise ratio and subjective listening tests.>
Jerry D. Gibson, Yoon-Chae Cheong, Hong Chae Woo, Wen-Whei Chang
ICASSP1
1990 Path map symbol release rules and the exponential metric tree
abstract
To design a tree coder for source coding with a fidelity criterion, one must choose a suitable code generator, an efficient tree search algorithm, an appropriate distortion measure, and a path map symbol release rule. The performance of several path map symbol release rules when used with exhaustive searching of the exponential metric tree is investigated. The average single-letter distortion of fixed-length symbol release rules and two variable-length symbol release rules are derived for shallow search depths and compared to simulation results. The incremental or single-symbol release rule is shown to yield the best performance.>
Wen-Whei Chang, Jerry D. Gibson
IEEE Trans. Inf. Theory2
1989 Lattice vector quantization for image coding
abstract
The authors investigate the application of several lattice vector quantizers, namely D/sub N/ for N>
Dae-Gwon Jeong, Jerry D. Gibson
ICASSP2
1989 Filtering of colored noise for speech enhancement and coding
abstract
A report is presented on experiments using a colored-noise assumption Kalman filter to enhance speech additively contaminated by colored noise, such as helicopter noise and jeep noise, with a particular application to linear predictive coding (LPC) of noisy speech. The results indicate that the colored-noise Kalman filter provides a significant gain in SNR, a clear improvement in the sound spectrogram, and an audible improvement in output speech quality. The authors demonstrate that such gains are unavailable with white noise assumption Kalman and Wiener filters. The colored-noise prefilter greatly enhances the quality and intelligibility of LPC output speech for noisy inputs.>
Boneung Koo, Jerry D. Gibson, Steven D. Gray
ICASSP2
1989 Multipulse-based codebooks for CELP coding at 7 kbps
abstract
A report is presented on the design of CELP codebooks matched in some ways to multipulse sequences that produce toll-quality speech. The authors designed a multipulse LPC system that achieves toll quality at 15 kb/s. They then studied the properties, especially the statistical properties, of these successful multipulse excitations for a variety of speakers and speech segments. CELP codebooks were developed which match the properties of the multipulse sequences to yield a final CELP coder for operation at 7 kb/s.>
Hong Chae Woo, Jerry D. Gibson
ICASSP2
1988 Estimation and vector quantization of noisy speech
abstract
The block and alphabet-constrained formulations are compared for the problem of vector quantization of noisy speech. In the optimum estimator/source-coder structures, a training mode vector quantizer due to Linde et al. (1980) is used as the source coder for the estimator outputs in all cases, and three block estimators and five alphabet-constrained estimators are examined. Objective and subjective performance results are obtained for all eight estimators used in conjunction with training mode vector quantizers at a rate of 1 bit/dimension, for dimensions 1,2,. . . ,8, and 2 bits/dimension, for dimensions 1,2,3, and 4, on five sentences, The results show the superiority of the alphabet-constrained approach, using the frame-adaptive Kalman filter, with improvements in output signal-to-noise ratio over the block approach of 20%.>
Jerry D. Gibson, Thomas R. Fischer, Boneung Koo
ICASSP1
1988 Backward adaptive tree coding of speech at 16 kbps
abstract
The authors investigate the performance of eight fully backward adaptive DPCM-based code generators with exhaustive searching to shallow depths. In particular, they compare the performance of fixed (matched and unmatched) autoregressive (AR), fixed (matched and unmatched) moving average (MA), undamped and damped Kalman AR, and undamped and damped gradient MA code generators in DPCM searched to a maximum depth of 5 with incremental single symbol release. SNR, subjective listening test, and spectrogram performance comparisons are employed. The performance gain due to adaptive prediction over unmatched fixed prediction is greater than the gain provided by multipath searching over single path searching.>
Jerry D. Gibson, Greg B. Haschke
ICASSP1
1988 Predictive trellis coded quantization of speech
abstract
Trellis coded quantization is incorporated into a predictive coding structure for encoding sampled speech. Systems are developed using fixed prediction/fixed residual encoding, fixed prediction/adaptive residual encoding, and adaptive prediction/adaptive residual encoding. For a fully adaptive 16 kbps speech coding system, segmental signal-to-noise ratios in the range of 17.5 to 20.2 dB are obtained for a variety of speakers and test sentences. Reconstructed speech obtained from this system can be described as being of excellent communications quality.>
Michael W. Marcellin, Thomas R. Fischer, Jerry D. Gibson
ICASSP3
1987 Experiments on video teleconferencing algorithms at 56 kilobits/sec
abstract
The design of a 56 kilobits/sec. (kbps) coder to transmit monochrome video signals should include several bit rate reduction techniques. In this work, some already existing techniques are simplified and combined with a new edge detection algorithm to produce a simple but efficient coder. Both signal-to-noise ratios and perceptual comparisons indicate that, without loss in performance, significant reduction in transmission data rate can be accomplished by properly recognizing the information bearing parts of an image.
Michael Maragoudakis, Jerry D. Gibson
ICASSP2
1987 Digital coding of waveforms: Principles and applications to speech and video
Jerry D. Gibson
Proc. IEEE1
1986 Estimation and Optimum Source Coding of Noisy Sources
Thomas R. Fischer, Jerry D. Gibson
ICC2
1986 Experimental Comparison of All-Pole, All-Zero, and Pole-Zero Predictors for ADPCM Speech Coding
abstract
Simulation results are presented which compare the performance of all-pole, all-zero, and pole-zero predictors in ADPCM at data rates of 16 and 32 kbits/s over both ideal and noisy channels. Separate backward adaptive gradient algorithms are used to adapt the poles and the zeros independently. The performance indicators used are signal-to-quantization noise ratio (SNR), signal-to-prediction error ratio (SPER), segmental SNR (SNRSEG), and subjective listening tests. For speech sources, the all-zero and pole-zero predictors produce SNR and SNRSEG values that are approximately 1-3 dB higher than those generated by the all-pole predictor. Subjective listening tests reveal that an eighth-order all-zero predictor performs as well or better than an allpole predictor for all conditions studied.
Boneung Koo, Jerry D. Gibson
IEEE Trans. Commun.2
1986 Preposterior analysis for differential encoder design
abstract
A modified differential encoding structure is proposed and optimized based upon the concept of preposterior analysis from the theory of alphabet-constrained data compression. Using preposterior analysis, the quantizer input sequence is chosen to minimize the expected distortion over a fixed but arbitrary interval, say,Nsamples long. By computing the expected distortion over future inputs, preposterior analysis allows future behavior to be modeled, but without an encoding delay as in tree coding. The optimized quantizer input sequence is not simply the prediction error as in classical differential pulse code modulation, but it is a weighted combination of the current prediction error and past encoding errors. The optimization is accomplished using a backward dynamic programming argument.
Jerry D. Gibson, Thomas R. Fischer, Boneung Koo
IEEE Trans. Inf. Theory1
1985 Backward Adaptive Lattice and Transversal Predictors in ADPCM
abstract
Four different backward adaptive predictors and a fixed predictor are compared for use in an adaptive differential pulse code modulation (ADPCM) system for coding speech at 16 kilobits/second (kbits/s). For noise-free channels, the four adaptive predictors, a least squares lattice, a least mean square lattice, a Kalman transversal form, and a gradient transversal form, all exceed the fixed predictor performance as well as the performance of a continuously variable slope delta (CVSD) modulation system. For bit error rates (BER's) of 10-3or greater, the transversal predictor performance falls below that of the fixed predictor and CVSD; however, the lattice structures maintain their performance advantage. The least squares lattice predictor has the best objective and subjective performance for both noiseless and noisy channels. All systems perform poorly for a BER of 10-2. To extend the performance of ADPCM with a least squares lattice predictor down to a BER of 10-2, the sampling rate is reduced and a selective coding scheme is devised. The resulting ADPCM system maintains excellent performance through a BER of 10-2and outperforms CVSD for noise-free and noisy channels. The dynamic range, tandeming performance, and behavior for noisy inputs for the ADPCM system and CVSD are investigated.
Randall C. Reininger, Jerry D. Gibson
IEEE Trans. Commun.2
1984 Backward adaptive lattice and transversl predictors for ADPCM
abstract
Adaptive predictors are important components of high performance differential pulse code modulation (DPCM) systems. This work compares the noise-free and noisy channel performance of DPCM at 16 kilobits/sec for four backward adaptive predictors: (1) an adaptive gradient transversal predictor, (2) a Kalman transversal predictor, (3) an adaptive gradient lattice predictor, and (4) a least squares lattice predictor. Subjective performance comparisons for five sentences rate the least squares lattice best, followed in order by the Kalman transversal, the gradient lattice, and the gradient transversal. All adaptive predictors exceed the performance of a fixed predictor. Monte Carlo simulations reveal that the lattice predictors outperform the transversal predictors for bit error rates of 10-3or greater. A fully backward adaptive DPCM system is investigated which uses the least squares lattice predictor and limited channel coding and is shown to outperform delta modulation for noise-free channels and for noisy channels with bit error rates as high as 10-2.
Randall C. Reininger, Jerry D. Gibson
ICASSP2
1984 Hamming Coding of DCT-Compressed Images Over Noisy Channels
abstract
Theoretical and simulation results of using Hamming codes with the two-dimensional discrete cosine transform (2D-DCT) at a transmitted data rate of 1 bit/pixel over a binary symmetric channel (BSC) are presented. The design bit error rate (BER) of interest is 10-2. The (7, 4), (15, 11), and (31, 26) Hamming codes are used to protect the most important bits in each 16 by 16 transformed block, where the most important bits are determined by calculating the mean squared reconstruction error (MSE) contributed by a channel error in each individual bit. A theoretical expression is given which allows the number of protected bits to achieve minimum MSE for each code rate to be computed. By comparing these minima, the best code and bit allocation can be found. Objective and subjective performance results indicate that using the (7, 4) Hamming code to protect the most important 2D-DCT coefficients can substantially improve reconstructed image quality at a BER of 10-2. Furthermore, the allocation of 33 out of the 256 bits per block to channel coding does not noticeably degrade reconstructed image quality in the absence of channel errors.
David R. Comstock, Jerry D. Gibson
IEEE Trans. Commun.2
1984 Self-Orthogonal Convolutional Coding for the DPCM-AQB Speech Encoder
abstract
A fixed-tap differential pulse code modulation (DPCM) system with a robust backward-adaptive Jayant quantizer is investigated for speech encoding at 16-40 kbits/s using binary phase shift keying over an additive white Gaussian noise channel. The performance of this system becomes unacceptable as the channel bit error rate(P_{b})approaches 10-2. Using high-rate, long constraint length, self-orthogonal convolutional codes, the DPCM system performance is much-improved for10^{-4} < P_{b} < 10^{-2}depending on the transmitted data rate. The use of high-rate(n - 1)/n, n = 2,3,4,, and 5 codes minimizes the number of bits allocated to channel coding, and decoding complexity is reduced by employing self-orthogonal codes which admit threshold decoding. Subjectively, while there is additional quantization noise with channel coding, the irritating popping and squeaking sounds due to channel errors are eliminated.
Charles C. Moore, Jerry D. Gibson
IEEE Trans. Commun.2
1984 An algorithm for uniform vector quantizer design
abstract
A vector quantizer maps ak-dimensional vector into one of a finite set of output vectors or "points". Although certain lattices have been shown to have desirable properties for vector quantization applications, there are as yet no algorithms available in the quantization literature for building quantizers based on these lattices. An algorithm for designing vector quantizers based on the root latticesA_{n}, D_{n}, andE_{n}and their duals is presented. Also, a coding scheme that has general applicability to all vector quantizers is presented. A four-dimensional uniform vector quantizer is used to encode Laplacian and gamma-distributed sources at entropy rates of one and two bits/sample and is demonstrated to achieve performance that compares favorably with the rate distortion bound and other scalar and vector quantizers. Finally, an application using uniform four- and eight-dimensional vector quantizers for encoding the discrete cosine transform coefficients of an image at0.5bit/pel is presented, which visibly illustrates the performance advantage of vector quantization over scalar quantization.
Khalid Sayood, Jerry D. Gibson, Martin C. Rost
IEEE Trans. Inf. Theory2
1983 Soft Decision Demodulation and Transform Coding of Images
abstract
This paper describes a transform image coding system that uses soft decision demodulation to control channel errors. In soft decision demodulation, if certain received bits of a codeword representing a coefficient are unreliable, then the codeword is rejected and the corresponding coefficient is replaced with an estimate. By monitoring the three highest energy DCT coefficients, the reconstructed image quality can be improved for a channel with a bit error probability of 10-2.
Randall C. Reininger, Jerry D. Gibson
IEEE Trans. Commun.2
1983 Distributions of the Two-Dimensional DCT Coefficients for Images
abstract
For a two-dimensional discrete cosine transform (DCT) image coding system, there have been different assumptions concerning the distributions of the transform coefficients. This paper presents results of distribution tests that indicate that for many images the statistics of the coefficients are best approximated by a Gaussian distribution for the DC coefficient and a Laplacian distribution for the other coefficients. Furthermore, from a simulation of the DCT coding System it is shown that the assumption that the coefficients are Laplacian yields a higher actual output signal-to-noise ratio and a much better agreement between theory and simulation than the Gaussian assumption.
Randall C. Reininger, Jerry D. Gibson
IEEE Trans. Commun.2
1982 Alphabet-constrained data compression
abstract
The optimal data compression problem is posed in terms of an alphabet constraint rather than an entropy constraint. Solving the optimal alphabet-constrained data compression problem yields explicit source encoder/decoder designs, which is in sharp contrast to other approaches. The alphabet-constrained approach is shown to have the additional advantages that (1) classical waveform encoding schemes, such as pulse code modulation (PCM), differential pulse code modulation (DPCM), and delta modulation (DM), as well as rate distortion theory motivated tree/trellis coders fit within this theory; (2) the concept of preposterior analysis in data compression is introduced, yielding a rich. new class of coders: and (3) it provides a conceptual framework for the design of joint source/channel coders for noisy channel applications. Examples are presented of single-path differential encoding, delayed (or tree) encoding, preposterior analysis, and source coding over noisy channels.
Jerry D. Gibson, Thomas R. Fischer
IEEE Trans. Inf. Theory1
1981 Bounds on performance and dynamic bit allocation for sub-band coders
abstract
Optimal sub-band coder design is considered based on Shannon's bounds for the rate of a continuous source with a given spectral density. Minimization of these bounds with respect to band partitioning and rate per band is considered. For fixed sub-band bandwidths, the optimum rate allocation for each band is derived. The resulting expressions can be solved in real-time to yield a dynamic rate allocution scheme.
Jerry D. Gibson
ICASSP1
1981 Incremental tree coding of speech
abstract
Tree coding of speech has been investigated by several workers. Virtually all of these investigations have involved incremental tree coding in that no matter how deep the tree is searched, only a single path map symbol is released at a time. As noted by Gray, even if a good long-term fit is found, the first step in the fit may be a poor one, thus yielding large sample distortions. Hence, it is important to stay on a path long enough to achieve the promised long-term distortion value. The relative frequency of path switching for the single symbol release rule is investigated for the(M,L)and truncated Viterbi tree search algorithms, various search depths, and different code generators. In addition, two multiple symbol release rules are investigated. One rule releases a fixed number of path symbols at a time, while the other rule releases a variable number of path symbols, the exact number depending on how many symbols are required for the average sample distortion to be less than or equal to theL-depth path average distortion. Speech sources are considered exclusively.
Andrew C. Goris, Jerry D. Gibson
IEEE Trans. Inf. Theory2
1980 Experimental comparison of forward and backward adaptive prediction in DPCM
abstract
Results of experimental comparisons of forward- and backward-adaptive prediction in differential pulse code modulation (DPCM) of speech are presented. Two different types of comparisons were conducted. In one comparison, both predictors were used with the same three/five-level pitch compensating quantizer (PCQ). For this comparison, forward prediction clearly outperforms backward prediction, but with the penalty of a 10% increase in data rate due to the need to transmit coefficients. In the second comparison, the forward-prediction DPCM system and the backward-prediction DPCM system are constrained to have the same data rate of 16 kbits/sec. The backward-adaptive predictor outperforms forward prediction for this latter comparison. The speech data base for the simulations is one sentence spoken by a male speaker in four different languages, English, French, German, and Arabic. The performance comparisons are based on signal-to-quantization noise ratio, signal-to-prediction error ratio, sound spectrograms, and formal subjective listening tests.
Jerry D. Gibson, Louis C. Sauter
ICASSP1
1980 Kalman Backward Adaptive Predictor Coefficient Identification in ADPCM with PCQ
abstract
Kalman backward adaptive predictor coefficient identification is combined with a modified pitch-compensating quantizer (MPCQ) to produce a high-performance adaptive differential pulse code modulation (ADPCM) system for operation at data rates of 12-16 kbits/s. The Kalman/MPCQ system is compared to an ADPCM system using a Kalman algorithm and robust Jayant qnantization and to a system with a fixed-tap predictor and MPCQ. The performance indicators are signal-to-quantization noise ratio (SNR), sound spectrogram analyses, and formal subjective listening tests. The SNR comparisons indicate that the Kalman/ MPCQ system has the highest SNR, followed by the fixed-tap/MPCQ system, and then the Kalman/robust Jayant system. Subjective listening test results show that the Kalman/MPCQ system is preferred over the fixed-tap/MPCQ system 100 percent of the time and over the Kalman/ robust Jayant system 80 percent of the time. Kalman adaptation thus provides an important perceptual effect not evident in the SNR's. The previously catastrophic effects of transmission errors on backward adaptive prediction are eliminated by simple ADPCM system modifications that do not affect the SNR or subjective quality of the output in the absence of errors for the five sentences studied. The problem of tandeming with a linear predictive coder (LPC) is investigated by using LPC processed speech as input to the three ADPCM systems and by using the output of the three ADPCM systems as input to an LPC analysis algorithm. For the LPC to ADPCM connection, the two systems with the MPCQ produce good quality output speech, while the system with robust Jayant quantization exhibits a fading phenomenon. For the ADPCM into LPC analysis, all three systems produce speech of approximately the same quality, with the fixedtap system being slightly, noisier. Using a distance measure proposed by Itakura, the predictor coefficients computed from the three ADPCM system outputs are compared with the predictor coefficients calculated from the uncontaminated speech. According to this distance measure, the coefficients computed from the Kalman/MPCQ system output are much closer to the desired coefficients than are those computed by the other two systems.
Jerry D. Gibson, Victor P. Berglund, Louis C. Sauter
IEEE Trans. Commun.1
1979 Optimal estimation and speech analysis
abstract
Optimal estimation algorithms for the identification of the predictor coefficients for both undistorted and distorted speech are developed by first specifying appropriate state space message and observation models. Those portions of the models and the algorithms which impede their utility are discussed. A technique for identifying the predictor coefficient vector single-stage state transition matrix is derived, and experimental results using speech data are presented.
Jerry D. Gibson, Andrew C. Goris
ICASSP1
1978 Sequentially Adaptive Backward Prediction in ADPCM Speech Coders
abstract
Several questions concerning the performance in ADPCM systems of sequentially adaptive backward predictors based on the adaptive gradient and Kalman-type algorithms are addressed. Using a Jayant-type adaptive quantizer, it is shown that for bit rates less than 16 kbits/s with second order predictors and for bit rates less 18.4 kbits/s with fourth order predictors, backward-adaptive predictors have a definite performance advantage over fixed-tap predictors, since the latter may cause system divergence. For higher bit rates, the adaptive gradient predictor offers no advantage over a second order fixed-tap predictor; however, the Kalman predictor produces a substantial performance increment over the fixed-tap predictor. It is also shown that the Kalman predictor maintains a significant advantage over the adaptive gradient predictor for all bit rates from 12.8 to 32 kbits/s. Finally, it is noted that the ADPCM system divergence that occurs for fixed, multiple-tap predictors and a Jayant quantizer is caused by predictor mismatch with the input signal coupled with the infinite quantizer memory. This problem can be corrected by a modification to the quantizer adaptation logic.
Jerry D. Gibson
IEEE Trans. Commun.1
1978 Fixed-Tap ADPCM System Divergence and a Bound on the Robust Quantizer Overload Point
abstract
The divergence of ADPCM systems with fixed, multipletap predictors and a Jayant quantizer is investigated. It is shown that system divergence occurs due to excessive quantization noise in the feedback loop coupled with the infinite quantizer memory. Further, divergence may result for even finer quantization if the predictor is poorly matched with the system input. New insight into quantizer/ predictor interaction is provided by a demonstration that for all average speech data available in the literature and more than one feedback tap, the system that describes the quantization noise evolution is unstable whenever the predictor is stable. It is noted that robust quantizer designs originally proposed for transmission error suppression are also effective in preventing the ADPCM system divergence problem discussed here, and a bound on the robust quantizer overload point is derived which illustrates the effect of the finite quantizer memory. Simulation results which validate the bound are presented.
Jerry D. Gibson, Edward A. Cross
IEEE Trans. Commun.1
1974 Sequentially Adaptive Prediction and Coding of Speech Signals
abstract
A new method of speech digitization called residual encoding is introduced, and its application to the speech digitization problem is studied. The residual encoding system is a form of differential pulse code modulation which utilizes both an adaptive quantizer and an adaptive predictor. The residual encoder differs from previous systems in two ways. First, a sequential estimation method is used to continuously update the predictor coefficients, and second, the predictor coefficients are not transmitted, but are extracted from the estimate of the speech signal at both the transmitter and receiver. No form of pitch extraction is employed. The residual encoding system with a Kalman filter or a stochastic approximation algorithm for identifying the predictor coefficients has produced good quality speech at a data rate of 16 kbit/s.
Jerry D. Gibson, Stephen K. Jones, James L. Melsa
IEEE Trans. Commun.1