Paul Mermelstein

dblp:98/2246 · DBLP profile ↗
← Back
68ranked-venue papers
9as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 6 first-authorComputer networks · 24Artificial intelligence and machine learning · 7 · 1 first-authorTheory of computation · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
8 papers
Physical-layer communications · 79% Cellular and mobile networks · 10% Network optimization and economics · 4%
Computer graphics and multimedia
3 papers
Audio and music processing · 100%
Artificial intelligence
4 papers
Speech recognition and synthesis · 68% Planning, search and constraint satisfaction · 31% Image recognition and object detection · 1%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Physical-layer communications
code-division multiple access
0.142002
Interference subspace rejection: a framework for multiuser detection in wideband CDMA · IEEE J. Sel. Areas Commun. 2002
A new receiver structure for asynchronous CDMA: STAR-the spatio-temporal array-receiver · IEEE J. Sel. Areas Commun. 1998
Adaptive Traffic Admission for Integrated Services in CDMA Wireless-Access Networks · IEEE J. Sel. Areas Commun. 1996
Physical-layer communications › signal detection
multiuser detection
0.122002
Interference subspace rejection: a framework for multiuser detection in wideband CDMA · IEEE J. Sel. Areas Commun. 2002
Impact of synchronization on performance of enhanced array-receivers in wideband CDMA networks · IEEE J. Sel. Areas Commun. 2001
Physical-layer communications
interference cancellation
0.012002
Interference subspace rejection: a framework for multiuser detection in wideband CDMA · IEEE J. Sel. Areas Commun. 2002
Physical-layer communications › synchronization
code synchronization
0.012001
Impact of synchronization on performance of enhanced array-receivers in wideband CDMA networks · IEEE J. Sel. Areas Commun. 2001
Physical-layer communications
spread spectrum and CDMA
0.012001
Impact of synchronization on performance of enhanced array-receivers in wideband CDMA networks · IEEE J. Sel. Areas Commun. 2001
Physical-layer communications
synchronization
0.012001
Impact of synchronization on performance of enhanced array-receivers in wideband CDMA networks · IEEE J. Sel. Areas Commun. 2001
Physical-layer communications › multiple access
CDMA systems
0.021996
Common packet data channel (CPDC) for integrated wireless DS-CDMA networks · IEEE J. Sel. Areas Commun. 1996
Effects of diversity power control, and bandwidth an the capacity of microcellular CDMA systems · IEEE J. Sel. Areas Commun. 1994
Physical-layer communications › channel estimation
blind identification
0.011998
A new receiver structure for asynchronous CDMA: STAR-the spatio-temporal array-receiver · IEEE J. Sel. Areas Commun. 1998
Physical-layer communications
channel estimation
0.011998
A new receiver structure for asynchronous CDMA: STAR-the spatio-temporal array-receiver · IEEE J. Sel. Areas Commun. 1998
Physical-layer communications
receiver design
0.011998
A new receiver structure for asynchronous CDMA: STAR-the spatio-temporal array-receiver · IEEE J. Sel. Areas Commun. 1998
Audio and music processing
speech coding
0.031994
Statistical recovery of wideband speech from narrowband speech · IEEE Trans. Speech Audio Process. 1994
Ensuring predictor tracking in ADPCM speech coders under noisy transmission conditions · IEEE J. Sel. Areas Commun. 1988
Subjective Evaluation of a 4.8 kbit/s Residual-Excited Linear Prediction Coder · IEEE Trans. Commun. 1981
Network optimization and economics
admission control
0.011996
Adaptive Traffic Admission for Integrated Services in CDMA Wireless-Access Networks · IEEE J. Sel. Areas Commun. 1996
Cellular and mobile networks › interference management
interference estimation
0.011996
Adaptive Traffic Admission for Integrated Services in CDMA Wireless-Access Networks · IEEE J. Sel. Areas Commun. 1996
Cellular and mobile networks
radio resource management
0.011996
Adaptive Traffic Admission for Integrated Services in CDMA Wireless-Access Networks · IEEE J. Sel. Areas Commun. 1996
Wireless networking
random access
0.011996
Common packet data channel (CPDC) for integrated wireless DS-CDMA networks · IEEE J. Sel. Areas Commun. 1996
Vehicular, aerial and satellite networks › satellite multiple access
spread ALOHA
0.011996
Common packet data channel (CPDC) for integrated wireless DS-CDMA networks · IEEE J. Sel. Areas Commun. 1996
Physical-layer communications
spread spectrum
0.011996
Rapid Acquisition Algorithms for Synchronization of Bursty Transmission in CDMA Microcellular and Personal Wireless Systems · IEEE J. Sel. Areas Commun. 1996
Audio and music processing › acoustic signal processing › audio signal reconstruction
audio super-resolution
0.011994
Statistical recovery of wideband speech from narrowband speech · IEEE Trans. Speech Audio Process. 1994
Physical-layer communications
antenna arrays
0.012002
Interference subspace rejection: a framework for multiuser detection in wideband CDMA · IEEE J. Sel. Areas Commun. 2002
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search › best-first search
a* search
0.011993
A*-admissible heuristics for rapid lexical access · IEEE Trans. Speech Audio Process. 1993
Natural language and speech › Speech recognition and synthesis
lexical access
0.011993
A*-admissible heuristics for rapid lexical access · IEEE Trans. Speech Audio Process. 1993
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
speaker adaptation
0.011990
Speaker Adaptation in a Large-Vocabulary Gaussian HMM Recognizer · IEEE Trans. Pattern Anal. Mach. Intell. 1990
Audio and music processing › speech coding
adaptive differential pulse code modulation
0.011988
Ensuring predictor tracking in ADPCM speech coders under noisy transmission conditions · IEEE J. Sel. Areas Commun. 1988
Coding theory › source coding
predictive coding
0.011988
Ensuring predictor tracking in ADPCM speech coders under noisy transmission conditions · IEEE J. Sel. Areas Commun. 1988
Network performance modeling
queueing analysis
0.011996
Common packet data channel (CPDC) for integrated wireless DS-CDMA networks · IEEE J. Sel. Areas Commun. 1996
Network optimization and economics
resource allocation
0.011996
Adaptive Traffic Admission for Integrated Services in CDMA Wireless-Access Networks · IEEE J. Sel. Areas Commun. 1996
Audio and music processing › audio representation
spectral modeling
0.011994
Statistical recovery of wideband speech from narrowband speech · IEEE Trans. Speech Audio Process. 1994
Cellular and mobile networks › power control
closed-loop power control
0.011994
Effects of diversity power control, and bandwidth an the capacity of microcellular CDMA systems · IEEE J. Sel. Areas Commun. 1994
Physical-layer communications
diversity
0.011994
Effects of diversity power control, and bandwidth an the capacity of microcellular CDMA systems · IEEE J. Sel. Areas Commun. 1994
Physical-layer communications › diversity
multipath diversity
0.011994
Effects of diversity power control, and bandwidth an the capacity of microcellular CDMA systems · IEEE J. Sel. Areas Commun. 1994

Methods — techniques the papers use, named apart from their topics

simulation · 0.0timing estimation · 0.0linear receivers · 0.0channel estimation · 0.0space-time processing · 0.0blind equalization · 0.0array processing · 0.0performance analysis · 0.0kalman filter · 0.0congestion control · 0.0statistical recovery function · 0.0spectral prediction · 0.0residual-signal-driven adaptation · 0.0lattice prediction · 0.0a* admissible heuristics · 0.0spectral mapping · 0.0covariance matrix estimation · 0.0subjective listening test · 0.0
YearPublicationVenuePosition
2007 A Spectrum-Efficient Multicarrier CDMA Array-Receiver with Diversity-Based Enhanced Time and Frequency Synchronization
abstract
This paper proposes a spectrum-efficient spatio-temporal array-receiver for multi-carrier CDMA systems named MC-STAR. First, we derive a new post-correlation model for MC-CDMA that characterizes the structure of the channel in space, time and frequency. Based on this model, we introduce a new multi-carrier array-receiver with rapid and accurate joint synchronization in time and frequency. There, we exploit jointly the spatial, temporal and frequency diversities as well as the intrinsic inter-carrier correlation (termed hereafter frequency gain) to improve the channel identification and the synchronization operations. In addition, based on a new link/system-level performance analysis, with a band-limited chip waveform assumption, we provide a comparative performance study of MC-STAR over two multi-carrier CDMA air-interface configurations, namely MT-CDMA and MC-DS-CDMA, in the most realistic operating conditions. Link/system-level results confirm the advantages of MT-CDMA in increasing throughput and bandwidth efficiency. The current trend is to design radio air-interfaces with flat fading subcarriers. In contrast, with MC-STAR we show that the positive effects of multipath diversity and frequency gain over large strongly-overlapping subcarriers is more significant than the negative effects of multipath and multi-carrier interference.
Besma Smida, Sofiène Affes, Paul Mermelstein
IEEE Trans. Wirel. Commun.4
2006 Performance Evaluation of an Enhanced Wideband CDMA Receiver Using Channel Measurements
abstract
The spatio-temporal array-receiver (STAR) decomposes generic wideband CDMA channel responses across various parameter dimensions (e.g., time-delays, multipath components, etc...) and extracts the associated time-varying parameters (i.e., analysis) before reconstructing the channel (i.e., synthesis) with increased accuracy. This work verifies the performance of STAR by comparing the results achieved with generic and measured channels for an average multipath power profile of [0, -4, -8] dB and a vehicular speed below 30 Km/h. The results suggest that losses due to operations with real 5 MHz channels are only 1 dB in SNR and 20-30% in capacity with DBPSK and single transmit and receive antennas. The corresponding SNR threshold for operation with real channels is about 5 dB.
Karim Cheikhrouhou, Sofiène Affes, Ahmed Elderini, Besma Smida, Paul Mermelstein, Belhassen Sultana, Venkatesh Sampath
VTC Fall5
2005 Multicarrier-CDMA STAR with time and frequency synchronization
abstract
This paper proposes a spectrum-efficient spatio-temporal array-receiver (STAR) for multi-carrier CDMA systems named MC-STAR. First, we derive a new post-correlation model for MC-CDMA that supports both the MT-CDMA and MC-DS-CDMA air-interfaces. Based on this model, we introduce a new multi-carrier receiver with rapid and accurate joint synchronization in time and frequency. We also exploit the intrinsic subcarrier correlation to improve the channel identification and the synchronization operations. We analyze the performance of MC-STAR in an unknown time-varying Rayleigh channel with multipath, carrier offset and cross-correlation between subcarrier channels. Simulation results confirm the accuracy of the joint time/frequency synchronization. They also confirm that for each MC-STAR configuration there exists an optimum number of subcarriers which results in maximum throughput. A higher number of subcarriers increases the inter-carrier interference while a lower number of subcarriers reduces the frequency gain. With four receiving antennas and five MT-CDMA subcarriers in 5 MHz bandwidth, MC-STAR provides about 1.2 bps/Hz at low mobility for DBPSK, i.e., an increase of 30% in spectrum efficiency over DS-CDMA.
Besma Smida, Sofiène Affes, Paul Mermelstein
ICC4
2004 Call Admission on the Uplink and Downlink of a CDMA System Based on Total Received and Transmitted Powers
abstract
We consider the problem of call admission and resource management in a code-division multiple-access (CDMA) wireless network supporting several types of services over a range of transmission rates and offering possibly different grades of service. Resource requirements are considered separately for the uplink and downlink. The high-level objective is to design a simple admission scheme that ensures adequate signal-to-interference ratios for both the incoming call (if accepted), as well as previously admitted calls. Our approach is based on two key ideas: 1) an integrated measure of resource utilization that is agnostic to the details of the traffic mix and 2) an estimate of the additional resources required to accommodate the new call seeking admission. The current work considers estimation of the total received power distribution on the uplink and the total transmitted power distribution on the downlink, and prediction of their displacements as a result of admitting a new call. The total received/transmitted power distributions are estimated based on data obtained from the power control module. The displacement of total received/transmitted power is predicted based on the characteristics of the incoming call and the current resource utilization. Dynamic call capacities are compared with static capacities to indicate the effectiveness of the proposed algorithm in achieving high network utilization with low probability of overload.
Sonia Aïssa, Joy Kuri, Paul Mermelstein
IEEE Trans. Wirel. Commun.3
2003 Downlink MIMO multiuser detection with interference subspace rejection
abstract
We proposed recently a new technique for multiuser detection in CDMA networks, denoted interference subspace rejection (ISR), and evaluated its performance on the uplink. This paper extends its application to the downlink (DL). On the DL the information about interference is sparse, e.g., spreading factor (SF) and modulation of interferers may not be known, which makes the task much more challenging. We present three new ISR variants, which require no prior knowledge of the interfering users. The new solutions are applicable to MIMO systems and can accommodate any modulation, coding, spreading factor, and connection type. A new dynamic power-assisted channelization code allocation (DACCA) technique significantly reduces implementation complexity at the receiving mobile. Simulations under practically reasonable conditions suggest that increased user capacities and data-rates are attainable with downlink interference subspace rejection (DLISR) and system capacity increases linearly with the number of antennas. Capacity gains are at least 3 dB over the single-user detector and increase to 8 dB for high data-rates with 16-QAM.
Henrik Hansen, Sofiène Affes, Paul Mermelstein
GLOBECOM3
2003 Enhanced interference suppression for spectrum-efficient high data-rate transmissions over wideband CDMA networks
abstract
Recently, we developed an efficient multiuser upgrade of the single-user spatio-temporal array-receiver (STAR), referred to as interference subspace rejection (ISR) (Affes, S. et al., IEEE J. Selec. Areas Comm., vol.20, no.2, p.287-302, 2002). The resulting STAR-ISR receiver offers a number of implementation modes covering a large range in performance and complexity for wideband CDMA networks. Here we apply STAR-ISR to high data-rate (HDR) transmissions by extending operation to delay spreads larger than the symbol duration. Low processing gain situations render symbol timing extremely difficult, especially with RAKE-type receivers. For HDR links of 512 Kbps in 5 MHz bandwidth, simulations indicate that the simplest mode of STAR-ISR outperforms the 2D-RAKE-PIC by a factor of 4 to 5 in spectrum efficiency while requiring the same order of complexity. This gain increases by up to a factor of 7.5 in high Doppler.
Sofiène Affes, Karim Cheikhrouhou, Paul Mermelstein
ICASSP (4)3
2003 Joint optimization of short-term and long-term predictors in CELP speech coders
abstract
The objective of this work is to investigate whether joint optimization of short-term and long-term predictors manifests significant advantages over the sequential optimization in speech coding. We propose a new joint optimization method based on Wiener filtering. The proposed analysis model resolves the pitch-bias problem of classical LPC analysis by considering the contribution of the long-term predictor while optimizing the short-term predictor. Our approach to joint optimization is based on analysis-by-synthesis and guarantees the synthesis filter stability. By applying our proposed joint optimization approach to CELP coding we obtain superior objective and subjective performance relative to CELP coding with sequential optimization. To provide voice quality equivalent to that of sequentially optimized CELP, the jointly optimized coder needs fewer FCB pulses and requires a reduced bit budget for LPC quantization. Our listening tests suggest that the JCELP coder at 4.25 kbps is equivalent in quality to the G.729 at 8 kbps.
Houman Zarrinkoub, Paul Mermelstein
ICASSP (2)2
2003 Efficient use of pilot signals in wideband CDMA array-receivers
abstract
We extend application a new scheme for efficient use of pilot signals in wideband CDMA array-receivers from the pilot-channel to the pilot-symbol case. The new scheme exploits the pilot signals for the simple resolution of the sign ambiguity arising in BPSK-decision-directed blind channel identification and achieves significant spectrum efficiency gains and power or overhead savings over the same array-receiver versions, which use pilots for conventional channel identification only. Both analysis and simulations suggest that pilot channel and pilot-symbol array-receiver versions, either with conventional or new pilot use, have similar performance at weak Doppler. They also indicate increasing performance gains with increasing Doppler due to the improved use of the pilot information. For a data rate of 144 Kbps with 60 Kmph speed, simulations indicate efficiently gains due to new pilot use of about 25 and 70% in the pilot-channel and pilot-symbol cases, respectively.
Sofiène Affes, Nahi Kandil, Paul Mermelstein
ICC3
2003 Joint optimization of short-term and long-term predictors in CELP speech coders
abstract
The objective of this work is to investigate whether joint optimization of short-term and long-term predictors manifests significant advantages over the sequential optimization in speech coding. We propose a new joint optimization method based on Wiener filtering. The proposed analysis model resolves the pitch-bias problem of classical LPC analysis by considering the contribution of the long-term predictor while optimizing the short-term predictor. Our approach to joint optimization is based on analysis-by-synthesis and guarantees the synthesis filter stability. By applying our proposed joint optimization approach to CELP coding we obtain superior objective and subjective performance relative to CELP coding with sequential optimization. To provide voice quality equivalent to that of sequentially optimized CELP, the jointly optimized coder needs fewer FCB pulses and requires a reduced bit budget for LPC quantization. Our listening tests suggest that the JCELP coder at 4.25 kbps is equivalent in quality to the G.729 at 8 kbps.
Houman Zarrinkoub, Paul Mermelstein
ICME2
2003 Capacity gain of zone division for a position-based resource allocation algorithm in WCDMA uplink data transmission
abstract
We present a position-based resource allocation algorithm and evaluate the gain of zone division by a simulator that realistically replicates the imperfect operations of a real system. I he positions are important in the sense of how they affect the level of interference in adjacent cells (or sectors), and is a function of the distance to adjacent base stations (BSs) and the shadowing effects of the propagation environment. The queued packet algorithm (QPA) introduced here from [E.C. Haddad, 2002] is based on the occupancy of virtual queues for users grouped in zones where the resource requirements may be considered as similar. We evaluate the gain of a 4-zone division versus a 2-zone division for an assignment region at capacity for various propagation conditions and mobile distributions in the system. We consider stable operation of the system with long-term fairness achieved between users in different zones of the same cell. We show that the capacity gain of 4-zone division ranges between 8% and 15% compared to the 2-zone division.
Elias Chafic Haddad, Charles L. Despins, Paul Mermelstein
PIMRC3
2003 Forward-link soft-handoff in CDMA with multiple-antenna selection and fast joint power control
abstract
We consider forward-link soft-handoff with multiple antenna selection and fast joint power control at high data rates in a cellular code-division multiple-access network, where signals are directed to a mobile station (MS) from antennas located at the same or different base stations. The total power transmitted to any mobile is divided among the active antennas selected according to the momentary channel conditions so as to maximize the signal-interference ratio at each MS. Multiple-antenna selection is used to mitigate the effects of both short- and long-term fading, and achieve the best soft-handoff with respect to system capacity and complexity. To achieve capacity gains with soft-handoff, we derive optimum handoff thresholds corresponding to the optimum handoff region in different cell environments. Numerical results demonstrate that under high Doppler spread and large handoff-delay conditions, the proposed soft-handoff employing two transmit antennas and the optimum handoff threshold achieves a significant gain in microcell environments, but not in macrocell environments.
Sofiène Affes, Paul Mermelstein
IEEE Trans. Wirel. Commun.3
2003 Uplink packet scheduling in the presence of interference cancellation in multi-rate wireless CDMA networks
abstract
Abstract We consider packet scheduling and rate assignment on the uplink of a packet data wireless CDMA network in the presence of imperfect interference cancellation (IC) and limited user transmission rates, and subject to in‐cell and out‐of‐cell resource limitations. The objective is to propose and implement a system level position‐based flow control algorithm that accounts for a limited IC capability provided by power control for multi‐user detection. The proposed algorithm assigns packets to be transmitted to separate queues, one for each spatial zone within which packets generate roughly the same in‐cell interference and impose equal interference on a neighboring base station. Given the cell partitioning into zones, the algorithm dynamically adapts to the resource constraints and efficiently uses IC to provide for fairness in serving the various queues without giving up the objective of maximizing data throughput. Throughput and fairness are the two conflicting objectives to be optimized. We show that the joint use of IC and location‐based scheduling is able to achieve complete fairness with negligible loss in throughput even under stringent resource limitations. The IC technique implemented is based on the interference subspace rejection (ISR) technique. We investigate both successive and group cancellation modes of ISR. Through the zone‐based grouping of users, the flow control algorithm provides a high flexibility in taking advantage of IC and is general enough to adapt to situations with constraints on the transmission rates. Results provided show how group‐based scheduling with group‐cancellation can provide for high fairness even under stringent out‐of‐cell resource limitations. Copyright © 2003 John Wiley & Sons, Ltd.
Sonia Aïssa, Amine Maaref, Paul Mermelstein
Wirel. Commun. Mob. Comput.3
2002 Carrier frequency offset recovery for CDMA array-receivers in selective Rayleigh-fading channels
abstract
We propose a carrier frequency offset recovery (CFOR) module for CDMA to operate with STAR, the spatio-temporal array-receiver. The new CFOR module implements simple linear regression (LR), similar in approach to that implemented for time-delay synchronization in STAR. Simulations in selective Rayleigh-fading channels at various mobile speeds show that an increasing carrier frequency offset (CFO) severely degrades the capacity of a CDMA system, the relative loss being more significant at low mobility and transmission rate. CFOR with STAR reduces the effect of CFO and compensates almost completely the capacity loss. Relative capacity gains due to CFOR increase with the CFO. For nomadic voice and data-rates, the capacity gain is in the range of 270 and 150% respectively, at a CFO of about 1 ppm.
Sofiène Affes, Paul Mermelstein
VTC Spring3
2002 Enhanced array-receiver operation with turbo codes for increasing the capacity of wideband CDMA networks
abstract
STAR, the spatio-temporal array-receiver was shown to outperform RAKE-type array-receivers and to increase the capacity of wideband CDMA networks. Turbo codecs were also identified as offering significant performance improvements. We demonstrate the gain achieved by turbo codecs over conventional convolutional codecs in a STAR-based CDMA system. For high data-rates of 153.6 Kbps with low mobility, link-level simulations on the uplink indicate that turbo codecs gain 2 to 3 dB in required SNR at a BER<10/sup -6/. System-level simulations confirm that SNR gains almost double capacity, from 7 to 13 mobiles/cell for blind STAR (i.e., without pilot) and from 10 to 22 mobiles/cell for pilot-aided STAR.
Sofiène Affes, Paul Mermelstein
VTC Spring3
2002 Interference subspace rejection: a framework for multiuser detection in wideband CDMA
abstract
We present a unifying framework for a new class of receivers that employ linearly-constrained interference cancellation (IC). The associated multiuser detectors operate in various modes and options ranging in performance from that of IC detectors to that of linear receivers, yet provide more attractive performance/complexity tradeoffs. They exploit both space and time diversities as well as the array-processing capabilities of multiple antennas and carry out simultaneous channel and timing estimation, signal combining and interference rejection. Additionally, they can operate on both links and in multiple mixed-rate traffic scenarios. The improved performance can be translated to increased utilization of wideband code division multiple access networks, particularly at high data rates.
Sofiène Affes, Henrik Hansen, Paul Mermelstein
IEEE J. Sel. Areas Commun.3
2002 A blind coherent spatiotemporal processor of orthogonal Walsh-modulated CDMA signals
abstract
Abstract A more efficient detection of orthogonal Walsh‐modulated Code Division Multiple Access (CDMA) signals is required not only for better exploitation of current IS‐95 CDMA systems but also for future cdma2000 3G networks. Integration of adaptive antennas at the base station has been recognized as one key lever to increasing capacity and spectrum efficiency. Prospective array‐receiver solutions such as the 2‐D‐RAKE improve performance; however, in the absence of a pilot signal, they have to implement noncoherent detection. In this work, we propose a space‐time processor that achieves coherent detection of orthogonal Walsh‐modulated CDMA signals without a pilot. We assess its performance in spatially correlated Rayleigh‐fading. Simulation results for voice links of 9.6 Kbps indicate that up to an antenna‐correlation factor of 0.8, the proposed receiver outperforms the conventional 2‐D‐RAKE's capacity by 90% in nonselective fading. This gain shrinks fast at higher correlation factors. In selective fading, however, it maintains about a 130% gain in capacity over the entire correlation range. For data links of 153.6 Kbps, this performance advantage increases up to 180–200%. With high‐speed mobiles, however, it vanishes quickly at correlation factors beyond 0.5. Overall, the capacity gains of the proposed spatiotemporal receiver structure increase with reduced relative Doppler (Doppler‐frequency/symbol‐rate ratio). This may arise from increased spatiotemporal diversity, higher transmission rates, and/or slower mobility. Copyright © 2002 John Wiley & Sons, Ltd.
Sofiène Affes, Paul Mermelstein
Wirel. Commun. Mob. Comput.2
2001 Interference subspace rejection in wideband CDMA: Modes for high data-rate operation
abstract
This paper extends our study on a multi-user receiver structure for base-station receivers with antenna arrays in multicellular systems. The receiver employs a beamforming structure with constraints that nulls the signal component in appropriate interference subspaces. Here we introduce a new mode, ISR-D (diversities), which finds and suppresses the subspace of the identified paths of all known interferers. A frame extension technique is proposed which can be applied to all available ISR modes to increase the dimensionality of the observation space and thereby avoid noise amplification as a result of subspace suppression, as well as allow asynchronous transmission. Performance differences arise between the modes due to different sensitivities to channel identification and data detection errors. For homogeneous high data-rate situations ISR-DX manifests the best performance. However, due to its reduced complexity, ISR-TRX appears to offer the best complexity-performance tradeoffs.
Henrik Hansen, Sofiène Affes, Paul Mermelstein
GLOBECOM3
2001 Interference subspace rejection in wideband CDMA: modes for mixed-power operation
abstract
Multiuser CDMA detectors suppress interference to provide improved noise immunity, increased capacity, higher data rates and reduced precision requirements for power control. A new multiuser detector structure is formulated which offers a number of implementation modes ranging in performance from that of interference cancellation (IC) detectors to that of linear receivers, yet provides more attractive performance/complexity tradeoffs. It exploits both space and time diversities as well as the array-processing capabilities enabled by multiple antennas and carries out simultaneous channel and timing estimation, signal combining and interference rejection. The improved performance enables increased utilization of wideband CDMA networks, especially at high data rates.
Sofiène Affes, Henrik Hansen, Paul Mermelstein
ICC3
2001 Impact of synchronization on performance of enhanced array-receivers in wideband CDMA networks
abstract
The synchronization performance of the receivers may severely impact the capacity or spectrum efficiency of wideband code division multiple access (CDMA) networks. Evaluations of the uplink capacity for improved spatio-temporal array receivers (STAR) and the two-dimensional RAKE (2-D-RAKE) with both perfect and active synchronization indicate that synchronization with the early-late gate component of a RAKE-type receiver may constitute a bottleneck to performance improvement. Enhancement of synchronization remains a key issue. Results also suggest that STAR offers a promising alternative to the 2-D-RAKE with early-late gate, offering an average increase in spectrum efficiency up to 100% in the presence of synchronization errors at both data rates of 9.6 and 128 kb/s. This gain further increases in high Doppler and fast multipath delay drifts. Data oversampling above the chip-rate favors STAR even more. Significant simplifications are introduced in the formulation of the STAR receiver which result in a complexity comparable to the 2-D-RAKE.
Karim Cheikhrouhou, Sofiène Affes, Paul Mermelstein
IEEE J. Sel. Areas Commun.3
2000 A high capacity CDMA array-receiver requiring reduced pilot power
abstract
The use of coherent modulation for the uplink of wireless CDMA systems requires pilot signals which contribute to the interference seen by other users. In previous work we showed that blind array-receivers outperform pilot-channel assisted array-receivers in capacity for various operating conditions. These array-receivers avoid additional interference due to pilot signals and achieve better channel identification using relatively stronger data signals. However, for BPSK signals they identify the channel within a sign ambiguity and require differential modulation and decoding of coherently detected bits. In this contribution we implement coherent modulation and detection and further increase the advantage of this blind channel identification scheme by introducing a new pilot called "pilot-sign". This pilot simply allows resolution of the channel sign ambiguity after its estimation by long-term averaging of the pilot-channel combiner output. Analysis indicates that the resulting pilot-sign assisted array-receiver requires very weak pilot power ratios, in the range of a fraction of a percent, and allows very significant pilot power savings and large capacity gains compared to pilot-channel array-receivers.
Sofiène Affes, Abdelrhani Louzi, Nahi Kandil, Paul Mermelstein
GLOBECOM4
2000 Pitch-synchronous linear-prediction analysis by synthesis with reduced pulse densities
abstract
An important step toward achieving a high-quality 4 kb/s speech codec is reducing the coding-rate of the stochastic codebook component to near 2 kb/s. The increased reconstruction error in the residual that such low-rate quantization implies motivates the search for techniques that reduce the perceptibility of the errors in the reconstructed signal. Pitch-synchronous estimation of the linear-prediction filter and pitch-synchronous updating of the adaptive codebook reduce the coefficient-estimation error and increase the relative contribution of the adaptive codebook component to the synthesized signal, thereby reducing audible noise. However, pitch synchronous analysis normally results in a variable-rate coder. To obtain a fixed-rate representation, we introduce an efficient representation of the stochastic codebook component using a pulse density of one pulse per 2 ms and signed magnitudes specified by 2 bits per pulse-pair. The resulting reconstructions are evaluated for CELP coders corresponding to classical and generalized-pitch-predictor designs. In both cases speech quality comparable to 8 kb/s G.729 is achieved.
Driss Guerchi, Yasheng Qian, Paul Mermelstein
ICASSP3
1999 Analysis by synthesis speech coding with generalized pitch prediction
abstract
A new analysis-by-synthesis speech coding structure is presented for high-quality speech coding in the 4 to 8 kb/s range. CELP with generalized pitch prediction (GPP-CELP) differs from classical code-excited linear prediction (CELP) in that for voiced segments it is the speech signal that is decomposed into a component predictable with the aid of the adaptive codebook (ACB) and a nonpredictable aperiodic component, not the LPC residual. The spectrum of the aperiodic component is estimated by linear-prediction analysis. An approximation to the aperiodic component is synthesized from a stochastic codebook of sparse pulse sequences and its spectrum is shaped by the LPC synthesis filter. The ACB contains samples of the past reconstructed signal, low-passed to increase the pitch prediction gain. For voiced segments the new structure yields higher pitch prediction gain and lower linear-prediction gain than classical CELP. Subjective and objective comparisons reveal significant advantages for GPP-CELP over classical CELP.
Paul Mermelstein, Yasheng Qian
ICASSP1
1999 A beamformer for CDMA with enhanced near-far resistance
abstract
The spatio-temporal array-receiver (STAR) achieves good performance in CDMA with multiple receiving antennas where the interference can be characterized as AWGN uncorrelated with the signal. To enhance its near-far resistance in correlated noise environments, we introduce optimal combining of the spatio-temporal components. Nearly as good performance can be obtained with a low complexity adaptive beamformer combination of the antenna and multipath diversity branches. Simulation results indicate that the modified STAR manifests significant gain in near-far resistance over its original version, more so with more branches. They also suggest that exploiting additional temporal correlation fingers as interference references can further improve near-far resistance in poor diversity situations.
Henrik Hansen, Sofiène Affes, Paul Mermelstein
ICC3
1999 Call admission on the uplink of a CDMA system based on total received power
abstract
We consider the problem of admitting calls to a CDMA wireless network supporting stream and packet services over a range of transmission rates and offering possibly different grades of service. Resource requirements are considered separately for the uplink and downlink. The high-level objective is to design a simple admission scheme for the uplink that ensures adequate signal-to-interference ratios for both the incoming call (if accepted), as well as previously admitted calls. Our approach is based on two key ideas: (a) a measure of resource utilization that is independent of how resources are allocated among the ongoing calls, and (b) an estimate of the additional resources required because of the new call seeking admission. The current work considers measurement of the received interference power distribution and estimation of its displacement as a result of admitting a new call. The interference distribution is obtained from the power control module. The interference displacement is estimated based on the characteristics of the incoming call and the current resource utilization. Dynamic call capacities are compared with static capacities to indicate the effectiveness of the algorithm in achieving high network utilization with low probability of overload.
Joy Kuri, Paul Mermelstein
ICC2
1998 Signal processing improvements for smart antenna signals in IS-95 CDMA
abstract
We propose an efficient signal processing technique to exploit smart antennas in IS-95 CDMA. It integrates decision feedback identification (DFI) and new decision variables into the adaptive antenna beamforming scheme. These two features permit implementation of a coherent detection on the uplink without a pilot and spatio-temporal maximum ratio combining (MRC). With two antennas and nonselective fading, the additional gain in SNR achieved by either feature is as significant as the gain obtained by simple antenna beamforming over diversity combining. Up to a 2 dB performance advantage is obtained in selective fading as compared to simple beamforming.
Sofiène Affes, Paul Mermelstein
PIMRC2
1998 A new receiver structure for asynchronous CDMA: STAR-the spatio-temporal array-receiver
abstract
We propose a spatio-temporal array-receiver (STAR) for asynchronous code division multiple access (CDMA), using a new space/time structural approach. First, STAR performs blind identification and equalization of the propagation channel from each mobile transmitter. Second, it provides fast and accurate estimates of the number, relative magnitude, and delay of the multipath components. From this space/time separation, STAR reconstructs the identified channel with respect to a partially revealed space/time structure and reduces identification errors by the order of the ratio of the processing gain and the number of paths. Therefore, STAR offers a high potential for increasing capacity, with relatively low computational complexity. Simulations confirm the good multipath acquisition and tracking properties of STAR in the presence of strong interference and fast Doppler.
Sofiène Affes, Paul Mermelstein
IEEE J. Sel. Areas Commun.2
1997 Spatio-Temporal Array-Receiver for Multipath Tracking in Cellular CDMA
abstract
Efficient multipath tracking is a crucial requirement for high performance in asynchronous cellular CDMA. Despite the recent increased focus on use of array receivers at the base stations, the problem of time-delay estimation for individual paths has not been specifically addressed. In this contribution we propose simple and fast multipath tracking procedures in a nonstationary CDMA environment, using the spatio-temporal array-receiver (STAR). Both the time-delays and their number are tracked in time in a partially blind scheme with a low order of computational complexity. We show by simulations the tracking capability of STAR at a high interference level.
Sofiène Affes, Paul Mermelstein
ICC (3)2
1996 Switched prediction and quantization of LSP frequencies
abstract
We present new results on switched prediction and quantization of the line spectral frequencies applicable to high-quality coding of speech at rates below 8 kb/s. The predictor-quantizer exploits both the temporal inter-frame correlations among the line-spectral frequency vectors of successive frames as well as the spatial within-frame correlations. Best results are obtained by separately training two split vector quantizers, one on predictable frames and one on non-predictable frames. Line-frequency differences are found to yield the lowest quantization error for both the high and low order quantizers for the non-predictable condition, as well as the low order quantizer of the predictable condition. With 10 ms frame separation, 18 bit split VQs, one for each mode, with 1-bit mode indication yield a mean spectral distortion near 1 dB. With 20 ms frame separation, 20 bit quantizers are required.
Houman Zarrinkoub, Paul Mermelstein
ICASSP2
1996 Adaptive Traffic Admission for Integrated Services in CDMA Wireless-Access Networks
abstract
Code-division multiple-access (CDMA) is a serious candidate for personal communication systems at 1.9 GHz in North America. We consider the issue of bandwidth management in a CDMA integrated wireless-access network with heterogeneous services. A framework for adaptive connection admission in the up-link direction is proposed. It is based on estimation of the interference at the base station receivers. The estimation algorithm employs a linear Kalman filter which is driven by a measurement of the interference and by predicted traffic parameters of the admitted connections. We derived several generic variants of the control architecture for the up-link direction to assess the main characteristics of the framework and to determine the trade-offs between complexity and performance. They vary from a fixed strategy with fixed power control to an adaptive strategy with full information about network state and adaptive power control. A numerical study of the proposed framework shows that the estimated value of the average interference adapts well to the real value under nonstationary and nonuniform environment. This feature results in high network utilization for arbitrary traffic conditions.
Zbigniew Dziong, Ming Jia, Paul Mermelstein
IEEE J. Sel. Areas Commun.3
1996 Common packet data channel (CPDC) for integrated wireless DS-CDMA networks
abstract
A common packet data channel (CPDC) architecture is proposed to support bursty, packet-based services in direct sequence code-division multiple-access (DS-CDMA) integrated wireless access networks. The architecture employs an asynchronous transfer mode (ATM) strategy in the forward CPDC link and a spread ALOHA-type random access strategy in the reverse CPDC link. A congestion control algorithm using base station broadcast and portable terminal random delay call reattempt is described. A performance analysis of the CPDC architecture and algorithms is carried out, and formulas for the bit error rate, blocking probability, system delay time, transmission time, and waiting time for packet data calls are derived. The interference caused by a CPDC to stream services in the network is determined, and the capacity of a CPDC is evaluated in terms of the number of packet data subscribers that can be served with a specified grade of service (GOS).
Salvatore D. Morgera, Paul Mermelstein
IEEE J. Sel. Areas Commun.3
1996 Rapid Acquisition Algorithms for Synchronization of Bursty Transmission in CDMA Microcellular and Personal Wireless Systems
abstract
Two rapid synchronization acquisition algorithms applicable to spread spectrum links of code division multiple access (CDMA) personal communication systems are proposed and evaluated. The algorithms operate within a self-referencing matched filter synchronizer structure, and are particularly useful in reducing synchronization overhead on links designed to carry packet-type services. The main distinguishing characteristic between the two schemes is that one uses hard-decision while the other uses soft-decision detection. The proposed schemes are especially applicable to reverse link transmissions in quasisynchronous CDMA systems in which timing at portable terminals is established via pilot and synchronization signals received on respective code-division channels from the home base. If discontinuous (bursty) transmission is used on reverse links, the acquisition process is required for each transmission burst because of the propagation time uncertainty. Analysis of the algorithms on additive white Gaussian noise (AWGN) and Rayleigh fading channels reveals that their performance depends significantly on the choice of synchronizer parameters and the average despread signal-to-interference ratio (SIR). When this choice is proper, acquisition over a single preamble of relatively short duration can be achieved with high probability. The soft-decision scheme introduces a performance advantage of between 4-9 dB depending on the length of the synchronizing preamble.
Witold A. Krzymien, Ahmad Jalali, Paul Mermelstein
IEEE J. Sel. Areas Commun.3
1996 Low bit-rate video transmission over fading channels for wireless microcellular systems
abstract
We consider the transmission of QCIF resolution (176/spl times/144 pixels) video signals over wireless channels at transmission rates of 64 kb/s and below. The bursty nature of the errors on the wireless channel requires careful control of transmission performance without unduly increasing the overhead for error protection. A dual-rate source coder is presented that adaptively selects a coding rate according to the current channel conditions. An automatic repeat request (ARQ) error control technique is employed to retransmit erroneous data-frames. The source coding rate is selected based on the occupancy level of the ARQ transmission buffer. Error detection followed by retransmission results in less overhead than forward error correction for the same quality. Simulation results are provided for the statistics of the frame-error bursts of the proposed system over code division multiple access (CDMA) channels with average bit error rates of 10/sup -3/ to 10/sup -4/.
Masoud R. K. Khansari, Ahmad Jalali, Eric Dubois 0002, Paul Mermelstein
IEEE Trans. Circuits Syst. Video Technol.4
1995 Common packet data channel (CPDC) architecture for CDMA integrated wireless access networks
abstract
A common packet data channel (CPDC) architecture is proposed to support bursty, packet-based services in DS-CDMA integrated wireless access networks. The architecture employs a TDMA strategy in the CPDC downlink and a spread ALOHA-type random access strategy in the CPDC uplink. A congestion control algorithm using base station broadcast and portable terminal random delay call reattempt is described. A performance analysis of the CPDC architecture and algorithms is carried out, and formulae for the bit error rate, blocking probability, system delay time, and waiting time for packet data calls are derived. The capacity of a CPDC is evaluated in terms of the number of packet data subscribers that can be served with a specified grade of service (GOS).
Salvatore D. Morgera, Paul Mermelstein
PIMRC3
1995 Capacity estimates for mixed-rate traffic on the integrated wireless access network
abstract
The integrated wireless access network provides for simultaneous transmission of speech, data, and video signals on a shared spectrum basis. All services employ the same chip rate and rate differences manifest themselves in variations on the resulting processing gain and/or transmitted power. The capacity of the system to carry stream traffic is estimated as a function of system bandwidth, source data rate and path loss exponent. Both stream and packet traffic capacities are of interest, however, the focus is on the stream traffic, both low rate and high rate sources, some constant rate and some variable rate. The bandwidth efficiency, defined as the number of simultaneous calls per cell or sector per MHz of bandwidth, is found to increase due to improved multiplexing gains as well as due to better averaging of the multiuser interference. On the downlink there is an additional source of multiplexing due to the addition of the powers transmitted to different portables, subject to a maximum total transmitted power constraint. The results suggest that aggregating contiguous blocks of available spectrum increases the capacity, bandwidth efficiency and peak service rate at the cost of the added complexity of transmission at higher chip rates.
Paul Mermelstein, Srinivas Kandala
PIMRC1
1994 Multi-band residual coding of CELP codecs at 8 kb/s
abstract
We explore the benefits for CELP coding of speech at 8 kb/s of dividing the residual signal after pitch filtering into three band-passed components and using separate codebooks to represent each component. Minimization of the perceptually weighted error between the input signal and the reconstructed signal is divided into several band-limited minimization operations where the lowest frequency match dominates the quality of the result. For equal total numbers of bits allocated to code the residual in a 5 ms frame, spectral division of the coding operation results on the average in a better match than temporal division into subframes. These results permit the design of a high quality speech codec at 8 kb/s with modest delay and low complexity.>
Paul Mermelstein, M. Saikaly
ICASSP (2)1
1994 Approaches to Layered Coding for Dual-Rate Wireless Video Transmission
abstract
Visual communications over wireless networks require the efficient and robust coding of video signals for transmission over wireless links having time-varying channel capacity. The authors compare several schemes for encoding video data into two priority streams, thereby enabling the transmission of video data over wireless links to be switched between two bit rates. An H.261 (p/spl times/64) algorithm is modified to implement each candidate scheme. The algorithms are evaluated for a microcellular wireless environment and a clear-channel bit rate of 65 kb/s. The results show that by combining layering with automatic-repeat request (wireless-)link control, almost-wireline visual quality can be achieved. >
Masoud R. K. Khansari, Awais Zakauddin, Wai-Yip Chan, Eric Dubois 0002, Paul Mermelstein
ICIP (1)5
1994 An improved acquisition algorithm for the synchronization of CDMA personal wireless systems
abstract
An improved synchronization acquisition algorithm applicable to spread-spectrum links of code division multiple access (CDMA) personal communication systems is proposed. The algorithm operates within a self-referencing matched filter synchronizer structure, and is particularly useful to reduce synchronization overhead on links designed to carry packet-type services. The proposed scheme involves a soft-decision detection algorithm which improves the acquisition performance by between 4 and 9 dB (depending on the choice of system parameters) in comparison to the one described previously by the authors (1994).
Witold A. Krzymien, Ahmad Jalali, Paul Mermelstein
PIMRC3
1994 On channel models for microcellular CDMA systems
abstract
Measured channel impulse response estimates are used to predict the parameters of the channel model needed to evaluate the performance of Rake receivers in CDMA systems. A channel model consisting of a number of resolvable multipath components is assumed. The parameters of interest are the relative average powers of the multipath components, the probability distribution of multipath fading on each path, and the cross-correlation coefficient of multipath fading among the different multipath components. The paper presents statistics on these parameters for measurements made in outdoor urban and indoor microcellular environments at 910 MHz.>
Geng Wu, Ahmad Jalali, Paul Mermelstein
VTC3
1994 Effects of diversity power control, and bandwidth an the capacity of microcellular CDMA systems
abstract
We evaluate the capacity and bandwidth efficiency of microcellular CDMA systems. Power control, multipath diversity system bandwidth, and path loss exponent are seen to have a major impact on the capacity. The CDMA system considered uses convolutional codes, orthogonal signalling, multipath/antenna diversity with noncoherent combining, and fast closed-loop power control on the uplink (portable-to-base) direction. On the downlink (base-to-portable), convolutional codes, BPSK modulation with pilot-signal-assisted coherent reception, and multipath diversity are employed. Both fast and slow power control are considered for the downlink. The capacity of the CDMA system is evaluated in a multicell environment taking into account shadow fading, path loss, fast fading, and closed-loop power control. Fast power control on the downlink increases the capacity significantly. Capacity is also significantly impacted by the path loss exponent. Narrowband CDMA (system bandwidth of 1.25 MHz) requires artificial multipath generation on the downlink to achieve adequate capacity. For smaller path loss exponents, which are more likely in microcellular environments, artificial multipath diversity of an order of as high as 4 may be needed. Wideband CDMA systems (10 MHz bandwidth) achieve greater efficiencies in terms of capacity per MHz.< >
Ahmad Jalali, Paul Mermelstein
IEEE J. Sel. Areas Commun.2
1994 Books on tape as training data for continuous speech recognition
Gilles Boulianne, Patrick Kenny, Matthew Lennig, Douglas D. O'Shaughnessy, Paul Mermelstein
Speech Commun.5
1994 Statistical recovery of wideband speech from narrowband speech
abstract
We present an algorithm to generate wideband speech from a narrow band version of the same. The main body of the algorithm is a statistical recovery function (SRF), which predicts the highband spectrum based solely on the narrowband spectrum. The performance of the algorithm has been measured both in terms of spectral distortion and spectral signal-to-noise ratio (SNR). We obtained a 3 dB gain in SNR for the reconstructed wideband speech as compared to the narrowband speech.>
Yan Ming Cheng, Douglas D. O'Shaughnessy, Paul Mermelstein
IEEE Trans. Speech Audio Process.3
1993 A*-admissible heuristics for rapid lexical access
abstract
A new class of A* algorithms for Viterbi phonetic decoding subject to lexical constraints is presented. This type of algorithm can be made to run substantially faster than the Viterbi algorithm in an isolated word recognizer having a vocabulary of 1600 words. In addition, multiple recognition hypotheses can be generated on demand and the search can be constrained in respect conditions on phone durations in such a way that computational requirements are substantially reduced. Results are presented on a 60000-word recognition task.>
Patrick Kenny, Rene Hollan, Vishwa Gupta, Matthew Lennig, Paul Mermelstein, Douglas D. O'Shaughnessy
IEEE Trans. Speech Audio Process.5
1992 Hybrid segmental-LVQ/HMM for large vocabulary speech recognition
abstract
The authors have assessed the possibility of modeling phone trajectories to accomplish speech recognition. This approach has been considered as one of the ways to model context-dependency in speech recognition based on the acoustic variability of phones in the current database. A hybrid segmental learning vector quantization/hidden Markov model (SLVQ/HMM) system has been developed and evaluated on a telephone speech database. The authors have obtained 85.27% correct phrase recognition with SLVQ alone. By combining the likelihoods issued by SLVQ and by HMM, the authors have obtained 94.5% correct phrase recognition, a small improvement over that obtained with HMM alone.>
Yan Ming Cheng, Douglas D. O'Shaughnessy, Vishwa Gupta, Patrick Kenny, Matthew Lennig, Paul Mermelstein, Sarangarajan Parthasarathy
ICASSP6
1992 Improving the speech quality of cellular mobile systems under heavy fading
abstract
Digital cellular systems experience heavy fading for short intervals during which the encoded bit stream may be severely corrupted. Requirements exist to detect the corrupted data slots and replace the corresponding information in the decoded speech stream by predicted speech based on previously corrected received data. The paper reports on error detection and speech regeneration techniques which achieve considerable quality improvements over those proposed for North American cellular use in IS-54, yet are compatible with it. Error detection is enhanced by setting a maximum likelihood threshold for Viterbi decoding and rejecting frames where the decoded sequence exceeds the threshold. Rejected frames are regenerated using the most recent correctly received speech coding parameters.>
Huan-yu Su, Paul Mermelstein
ICASSP2
1992 HMM training on unconstrained speech for large vocabulary, continuous speech recognition
Gilles Boulianne, Patrick Kenny, Matthew Lennig, Douglas D. O'Shaughnessy, Paul Mermelstein
ICSLP5
1992 Statistical recovery of wideband speech from narrowband speech
abstract
We present an algorithm to generate wideband speech from a narrow band version of the same. The main body of the algorithm is a statistical recovery function (SRF), which predicts the highband spectrum based solely on the narrowband spectrum. The performance of the algorithm has been measured both in terms of spectral distortion and spectral signal-to-noise ratio (SNR). We obtained a 3 dB gain in SNR for the reconstructed wideband speech as compared to the narrowband speech. >
Yan Ming Cheng, Douglas D. O'Shaughnessy, Paul Mermelstein
ICSLP3
1991 Using phoneme duration and energy contour information to improve large vocabulary isolated-word recognition
abstract
Minimum duration constraints and energy thresholds for phonemes were used to increase the recognition accuracy of an 86000-word speaker-trained isolated word recognizer. Minimum duration constraints force the phoneme models to map to acoustic segments longer than the duration minima for the phonemes. Such constraints result in significant lowering of likelihoods of many incorrect word choices, improving the accuracy of acoustic recognition and recognition with the language model. The phoneme models were also improved by correcting the segmentation of the phonemes in the training set. During training, the boundaries between phonemes are not marked accurately. Energy is used to correct these boundaries. Application of an energy threshold improves the segment boundaries between stops and sonorants (vowels, liquids, and glides), between fricatives and sonorants, between affricates and sonorants. and between breath noise and sonorants. On two speakers, the overall reduction in errors using minimum durations and energy thresholds is from 27.3% to 23.1% for acoustic recognition and from 14.3% to 8.8% with the language model.>
Vishwa Gupta, Matthew Lennig, Paul Mermelstein, Patrick Kenny, Franz Seitz, Douglas D. O'Shaughnessy
ICASSP3
1991 A*-admissible heuristics for rapid lexical access
abstract
The authors present a new class of A* algorithms for Viterbi phonetic decoding subject to lexical constraints. They show that this type of algorithm can be made to run substantially faster than the Viterbi algorithm in an isolated word recognizer having a vocabulary of 1600 words and that it runs very quickly on a 60000-word recognition task. In addition, multiple recognition hypotheses can be generated on demand and the search can be constrained to respect conditions on phone durations in such a way that computational requirements are substantially reduced.>
Patrick Kenny, Rene Hollan, Vishwa Gupta, Matthew Lennig, Paul Mermelstein, Douglas D. O'Shaughnessy
ICASSP5
1991 Energy, duration and Markov models
abstract
We present a new stochastic model for the energy and duration of phone segments which takes account of the speech rate, the loudness of the signal and the effects of stress and pre-pausal lengthening and we show how the block Viterbi decoding algorithm can be used to integrate it with phone-based HMM speech recognizers. The model has been implemented on an isolated-word data-base and a preliminary experiment gives a modest improvement in word recognition accuracy.
Patrick Kenny, Sarangarajan Parthasarathy, Vishwa Gupta, Matthew Lennig, Paul Mermelstein, Douglas D. O'Shaughnessy
EUROSPEECH5
1990 Acoustic recognition component of an 86000-word speech recognizer
abstract
Recent results obtained with a hidden Markov model (HMM)-based acoustic recognizer using a virtually unlimited vocabulary (86000 words) to perform speaker-dependent isolated-word recognition are described. The task domain of this recognizer is quite general, consisting of paragraphs read from various newspapers, books, and magazines. The results of a comparative acoustic recognition study using various types of HMMs and various amounts of training data (from 700 to about 4000 words) are presented. The models explored include context-dependent allophonic HMMs (including generalized diphone and triphone models with unimodal Gaussian output densities) and context-independent phonemic HMMs (using either unimodal or mixture densities). Experimental results indicate that phonemic HMMs with many components in the mixture output densities provide the highest acoustic recognition accuracy. The acoustic recognition accuracy for a total of about 7000 test words spoken by four male and five female speakers is 82%. Recognition accuracy after application of the language model increases to 92%.>
Vishwa Gupta, Matthew Lennig, Patrick Kenny, Paul Mermelstein
ICASSP5
1990 Speaker Adaptation in a Large-Vocabulary Gaussian HMM Recognizer
abstract
The problem of using a small amount of speech data to adapt a set of Gaussian HMMs (hidden Markov models) that have been trained on one speaker to recognize the speech of another is considered. The authors experimented with a phoneme-dependent spectral mapping for adapting the mean vectors of the multivariate Gaussian distributions (a method analogous to the confusion matrix method that has been used to adapt discrete HMMs), and a heuristic for estimating covariance matrices from small amounts of data. The best results were obtained by training the mean vectors individually from the adaptation data and using the heuristic to estimate distinct covariance matrices for each phoneme.>
Patrick Kenny, Matthew Lennig, Paul Mermelstein
IEEE Trans. Pattern Anal. Mach. Intell.3
1989 A locus model of coarticulation in an HMM speech recognizer
abstract
A novel type of hidden Markov model (HMM) has been developed to account explicitly for the context-dependent vowel acoustic transitions in consonant-vowel and vowel consonant phonetic environments. The major difference between this type of HMM and the standard Gaussian HMM is that the Gaussian mean vectors associated with the vowel HMM states, which are intended to model the vowel acoustic transitions, are set to be linearly interpolated values between those of the vowel steady state and those of the assumed locus for the adjacent consonant. The locus vectors, one for each consonant except for /h/, are trained together with all other HMM parameters using the Baum-Welch algorithm. The training procedure is fully automatic and converges to a local maximum. The model is incorporated in a phonetically based 75000-word vocabulary speech recognizer and provides a modest improvement in recognition rate over the standard approach.>
Patrick Kenny, Matthew Lennig, Vishwa Gupta, Paul Mermelstein
ICASSP5
1988 Modeling acoustic-phonetic detail in an HMM-based large vocabulary speech recognizer
abstract
The acoustic recognizer of the INRS-Telecommunications 60000-word-vocabulary isolated-word recognition system is discussed. The task of the acoustic recognizer is to generate a list of word hypotheses and their likelihoods based on the acoustic data for each input word. Two sets of experiments are reported in which such knowledge is incorporated into the hidden Markov models (HMMs) used during recognition. In the first set, vowel duration properties are used in the HMMs. In the second set, word-initial and word-final stop consonants are modeled as a sequence of context-dependent subphonemes. The performance of the recognizer is significantly improved by appropriate utilization of the context-dependent vowel-duration information and the context-dependent microsegmental properties of stop consonants. >
Matthew Lennig, Vishwa Gupta, Paul Mermelstein
ICASSP4
1988 Three probabilistic language models for a large-vocabulary speech recognizer
abstract
Relative performance is compared for three different language models applied to the linguistic decoding part of a 75000-word speech recognizer. These models are the trigram model, the tri-POS model (POS stands for parts of speech), and a smoothed trigram model with tied distributions for words three or more syllables long. The full trigram model gives the best performance but is most expensive in terms of data and storage requirements. The smoothed trigram and tri-POS models yield equivalent performance. For general text entry tasks, use of the tri-POS model is suggested since it is less sensitive to variations in the discourse domains. For applications specific to individual discourse domains, trigram models trained on data obtained from the target domain are recommended.>
Pierre Dumouchel, Vishwa Gupta, Matthew Lennig, Paul Mermelstein
ICASSP4
1988 Ensuring predictor tracking in ADPCM speech coders under noisy transmission conditions
abstract
The problem of predictor mistracking for narrowband signals in backward adaptive ADPCM (adaptive digital pulse code modulation) speech coders is shown to arise as a result of feedback from the signal reconstruction filter to the predictor adaptation process. A class of residual-signal-driven lattice predictors (PR) is defined that guarantees tracking for all signals without regard to the order of prediction. The LR predictor enhances speech and DTMF (dual-tone multifrequency) signal transmission performance in the presence of transmission errors. Under error-free transmission conditions, a segmental SNR (signal-to-noise ratio) drop for speech of nearly 2 dB may be encountered for the LR predictor relative to the classical signal-drive lattice predictor. In most practical telecommunication applications, however, this degradation is outweighed by the improved robustness of the predictor.>
Paul Yatrou, Paul Mermelstein
IEEE J. Sel. Areas Commun.2
1987 Integration of acoustic information in a large vocabulary word recognizer
abstract
This paper proposes a new way of using vector quantization for improving recognition performance for a 60,000 word vocabulary speaker-trained isolated word recognizer using a phonemic Markov model approach to speech recognition. We show that we can effectively increase the codebook size by dividing the feature vector into two vectors of lower dimensionality, and then quantizing and training each vector separately. For a small codebook size, integration of the results of the two parameter vectors provides significant improvement in recognition performance as compared to the quantizing and training of the entire feature set together. Even for a codebook size as small as 64, the results obtained when using the new quantization procedure are quite close to those obtained when using Gaussian distribution of the parameter vectors.
Vishwa Gupta, Matthew Lennig, Paul Mermelstein
ICASSP3
1986 Reducing signal delay in multipulse coding at 16kb/s
abstract
Block-band coders typically impose a signal delay of about 60 ms. This paper reports on multipulse-LPC coding structures which reduce signal delay to 2 ms. The proposed structures are based on I) a high update rate for LPC side-information, and II) the use of very short duration intervals for the placement of pulses. The low delay structure imposes a relatively more uniform distribution of pulses in time. The addition of pitch prediction, over very short blocks, is observed to overcome the degradation due to the uniform pulse distribution. An efficient pitch prediction approach and the use of a forward-adaptive lattice filter contribute to maintaining computational requirements within the reach of modem DSP chips.
Michael G. Berouti, J. Jachner, D. Sloan, Paul Mermelstein
ICASSP4
1984 Efficient computation and encoding of the multipulse excitation for LPC
abstract
This paper discusses the analysis techniques used to derive the excitation waveform for multipulse coding of speech. A computationally efficient formulation is derived for both covariance and correlation type analyses. These methods differ in the way block edges are treated. Several methods for pulse amplitude and position determination are given, ranging from a purely sequential one to one which reoptimizes pulse amplitudes at each step. It is shown that the reoptimization scheme has a nested structure that allows a reduction in the computations. An efficient method for pulse position coding is given. This method can essentially achieve the entropy limit for randomly placed pulses. Experimental results are given for typical configurations including computational requirements and speech quality assessments.
Michael G. Berouti, H. Garten, Peter Kabal, Paul Mermelstein
ICASSP4
1984 Decision rules for speaker-independent isolated word recognition
abstract
This study compares the recognition rates attainable with the aid of two different methods of generating reference templates from training words and two different decision rules. The test enviroment consists of isolated words from a small vocabulary spoken by a large number of speakers over the public telephone system. Experiments performed show that the use of individual word templates for references together with the k-nearest neighbor decision procedure substantially improves the performance in isolated word recognition. We attempted to minimize the computations involved in the k-nearest neighbor decision procedure by assuming that the dynamic time-warp distance was a metric, which would allow use of a 1-nearest neighbor decision rule with appropriately relabeled reference data. Results indicate that this step leads to an error rate exceeding that obtainable with the 1-nearest neighbor rule on the original nonrelabeled data.
Vishwa Gupta, Matthew Lennig, Paul Mermelstein
ICASSP3
1984 Prevention of Predictor Mistracking in ADPCM Coders
D. J. Millar, Paul Mermelstein
ICC (3)2
1982 Adaptive predictive coding of speech and voiceband data signals
abstract
We consider ADPCM coding techniques for speech and voiceband data signals that do not require previous identification of the type of signal to be encoded. After a review of the frequently conflicting requirements for prediction and quantization of these two types of signals, coders with backward adaptive prediction and backward (AQB) or forward (AQF) quantizer adaptation are discussed. We find that ADPCM-AQF possesses significant performance advantages relative to ADPCM-AQB for both speech and voiceband data signals at the cost of additional signal delay. 4800 b/s voiceband data can be transmitted satisfactorily at 32 kb/s even with multiple stages of tandeming and infrequent transmission errors, but 40 kb/s are required for reliable 9600 b/s voiceband data transmission with intermediate analog conversion.
Paul Mermelstein, D. J. Millar
ICASSP1
1982 Review of 'Spoken Language Generation and Understanding' (Simon, J.C., Ed.; 1980)
Paul Mermelstein
IEEE Trans. Inf. Theory1
1981 Subjective Evaluation of a 4.8 kbit/s Residual-Excited Linear Prediction Coder
abstract
A 4.8 kbit/s residual-excited linear prediction coder (RELP) with two subband coded basebands was systematically evaluated in terms of intelligibility and overall quality. Intelligibility degradation due to RELP-coding is found to be 6 percent without transmission errors, an additional 6.4 percent with 1 percent bit errors, and 9.8 percent in 10 dB SNR acoustic background noise. Quality of the RELP coded speech is midway between those of 3 and 4 bit log-PCM and is significantly higher than that of the pitch-excited linear prediction coder.
Mamoru Nakatsui, Dale C. Stevenson, Paul Mermelstein
IEEE Trans. Commun.3
1980 Experiments in syllable-based recognition of continuous speech
abstract
An exploratory implementation of a syllable-based recognizer is described. Continuous speech is first divided into syllabic units, and the units are then matched against syllable templates using a dynamic programming algorithm. A hierarchical transition network is used to limit the syllable search to possible continuations of the current partial sentence hypotheses. Competing hypotheses are pruned by a 'beam search'. Experiments are reported on automatic recognition of English sentences with a 70-word vocabulary and restricted syntax produced by one male speaker. 85% of the sentences were correctly recognized. Comparable results were obtained for a similar task in French using a female speaker. The method is computationally efficient: real-time performance on dedicated hardware should be obtainable without difficulty. A method of scaling the distance measures used in the syllable matching is described. This scaling takes into account variability in syllable production, both as a function of position within the syllable and as a function of the various spectral parameters being used.
Melvyn J. Hunt, Matthew Lennig, Paul Mermelstein
ICASSP3
1978 Recognition of monosyllabic words in continuous sentences using composite word templates
abstract
As a preliminary step in syllable recognition in continuous speech, differences in acoustic forms were studied when these are induced by variation in the syntactic role of the word in the sentences. A modified dynamic programming algorithm is presented that allows building up of reference information from a speaker's productions in the face of such variations. The limited data base studied included 169 productions occurring in 57 sentences and composed of 52 different monosyllabic CVC words spoken by two male speakers. When each speaker's productions were tested against reference data built up from different tokens of his productions, correct first choice recognition was attained for 83% and 90% of the words, respectively. A similarity metric that takes into account the detailed acoustic variation that each speech sound may undergo without loss of identifiability can be expected to improve on these results.
Paul Mermelstein
ICASSP1
1976 The syntax of acoustic segments
abstract
The grouping of acoustic segments into syllables and words reflects a structural organization similar to that found at many higher levels of linguistics. Evidence from speech perception studies indicates that the human interpretation of an acoustic segment is a function of its position in the stream of speech segments. Speech synthesis programs use segment-combination rules that depend on the manner of production classes of the constituent segments. This paper reviews considerations for use of the syntactic approach in automatic segmental analysis of speech for speech recognition applications. We find that the segmentation and labeling processes are sufficiently strongly connected to make context-independent segmentation unlikely to prove successful in practice. Voicing and manner of production are the features under strongest syntactic constraint while segments differing in place of production can be more freely substituted.
Paul Mermelstein
ICASSP1
1969 Computer Simulation of Articulatory Activity in Speech Production
Paul Mermelstein
IJCAI1
1964 Experiments on Computer Recognition of Connected Handwritten Words
Paul Mermelstein, Murray Eden
Inf. Control.1