VLDB 2026 Research / reviewers in the wild / expert
Costas S. Xydeas
dblp:39/4248 · also Costas Xydeas
· DBLP profile ↗
40ranked-venue papers
13as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 8 first-authorArtificial intelligence and machine learning · 10 · 2 first-authorComputer networks · 6 · 4 first-authorDatabases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
8 papers |
Image and video processing · 59% Image and video coding · 21% Audio and music processing · 20% | |
| Network and information security
1 paper |
Digital forensics and information hiding · 100% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing
image fusion |
0.1 | 3 | 2005 | Objective Image Fusion Performance Characterisation · ICCV 2005 Gradient-based multiresolution image fusion · IEEE Trans. Image Process. 2004 Evaluation of Image Fusion Performance with Visible Differences · ECCV (3) 2004 |
Image and video coding › image quality assessment
image fusion quality assessment |
0.0 | 1 | 2004 | Evaluation of Image Fusion Performance with Visible Differences · ECCV (3) 2004 |
Image and video processing › image fusion
multi-scale image fusion |
0.0 | 1 | 2004 | Gradient-based multiresolution image fusion · IEEE Trans. Image Process. 2004 |
Audio and music processing
speech coding |
0.0 | 4 | 1999 | Split matrix quantization of LPC parameters · IEEE Trans. Speech Audio Process. 1999 Frequency Compression of 7.6 kHz Speech Into 3.3 kHz Bandwidth · IEEE Trans. Commun. 1983 Sequential Adaptive Predictors for ADPCM Speech Encoders · IEEE Trans. Commun. 1982 |
Audio and music processing › speech coding
linear predictive coding |
0.0 | 1 | 1999 | Split matrix quantization of LPC parameters · IEEE Trans. Speech Audio Process. 1999 |
Audio and music processing › speech coding
line spectral frequency quantization |
0.0 | 1 | 1999 | Split matrix quantization of LPC parameters · IEEE Trans. Speech Audio Process. 1999 |
Image and video coding › quantization
vector quantization |
0.0 | 1 | 1999 | Split matrix quantization of LPC parameters · IEEE Trans. Speech Audio Process. 1999 |
Image and video processing › image fusion
multisensor image fusion |
0.0 | 1 | 2005 | Objective Image Fusion Performance Characterisation · ICCV 2005 |
Image and video processing › image representation
image pyramid |
0.0 | 1 | 2004 | Gradient-based multiresolution image fusion · IEEE Trans. Image Process. 2004 |
Image and video processing › image representation
multiresolution representation |
0.0 | 1 | 2004 | Gradient-based multiresolution image fusion · IEEE Trans. Image Process. 2004 |
Image and video coding › quality assessment
visible difference prediction |
0.0 | 1 | 2004 | Evaluation of Image Fusion Performance with Visible Differences · ECCV (3) 2004 |
Digital forensics and information hiding
steganography |
0.0 | 1 | 1984 | Embedding Data Into Pictures by Modulo Masking · IEEE Trans. Commun. 1984 |
Audio and music processing › speech coding
adaptive differential pulse code modulation |
0.0 | 1 | 1982 | Sequential Adaptive Predictors for ADPCM Speech Encoders · IEEE Trans. Commun. 1982 |
Audio and music processing › speech coding
adaptive predictive coding |
0.0 | 1 | 1982 | Sequential Adaptive Predictors for ADPCM Speech Encoders · IEEE Trans. Commun. 1982 |
Image and video coding › quantization
adaptive quantization |
0.0 | 1 | 1980 | Envelope Dynamic Ratio Quantizer · IEEE Trans. Commun. 1980 |
Audio and music processing › speech quality assessment
speech intelligibility |
0.0 | 1 | 1983 | Frequency Compression of 7.6 kHz Speech Into 3.3 kHz Bandwidth · IEEE Trans. Commun. 1983 |
Image and video coding › predictive coding
differential pulse code modulation |
0.0 | 1 | 1982 | Sequential Adaptive Predictors for ADPCM Speech Encoders · IEEE Trans. Commun. 1982 |
Methods — techniques the papers use, named apart from their topics
gradient information representation · 0.1visible difference metric · 0.0quadrature mirror filters · 0.0discrete wavelet transform · 0.0split vector quantization · 0.0modulo masking · 0.0luminance scrambling · 0.0adaptive frequency mapping · 0.0stochastic approximation · 0.0sequential gradient estimation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Probabilistic Multimodal Classification with dynamic feature selection
Asmar A. Khan, Costas S. Xydeas, Hassan Ahmed |
FUSION | 2 |
| 2013 | Multi-Component/Multi-Model AAM framework for face image modelingabstractAn image face modeling framework is proposed that aims to enhance the face modeling capability of the well known Active Appearance Model (AAM). AAM has been used successfully in person-specific related applications but it poses significant limitations when employed in generic face modeling. Thus this work is focused on the development of new face models which are generic in nature and which accurately fit unseen image faces, both in terms of shape and texture. For this purpose, images are decomposed into face related components which are subsequently clustered on the basis of shape similarities. Experimental results show that models generated through this novel framework can be significantly more effective than conventional AAM, in terms of both shape and texture. Muhammad A. Khan 0002, Costas S. Xydeas, Hassan Ahmed |
ICASSP | 2 |
| 2013 | Hierarchical Classification Fusion frameworkabstractThis paper presents a novel hierarchical Classification Fusion (CF) framework which operates on Abstract and Measurement levels simultaneously and thus exploits information patterns resulting from the output labels and posterior beliefs of individual classifiers. Furthermore the proposed classification fusion methodology allows for the decomposition of the input data, which is used to design individual classifiers, into subsets. This in turn permits individual classifiers to be re-designed per subset and in a manner that increases overall system classification performance. Experimental results are presented which demonstrate the potential of the proposed methodology in the case of multi-modal, multi-feature binary data classification problems. In addition the proposed CF design framework can be applied to multi class problems and is independent of the type of classifiers employed in the system. Asmar A. Khan, Costas S. Xydeas, Hassan Ahmed |
ICASSP | 2 |
| 2013 | Using macroscopic information in image segmentationabstractPost‐processing ‘macroscopically’ output‐segmented images obtained from conventional image segmentation (IS) techniques, leads into the concept of micro–macro IS (MMIS). MMIS pays extra attention to information extracted from relatively large image regions and as a result, overall system segmentation performance improves both subjectively and objectively. The proposed post‐processing scheme is generic, in the sense that can be used together with any other existing segmentation approach. Thus given an input‐segmented image, MMIS has the ability to automatically select an appropriate number of regions and classes in a way that helps object‐oriented visual information to become more apparent in the final segmented output image. Computer simulation results clearly indicate that significant IS performance benefits can be obtained by augmenting conventional IS schemes within an MMIS framework, with or without input images being corrupted by additive Gaussian noise Asmar Azar Khan, Costas S. Xydeas, Hassan Ahmed |
IET Image Process. | 2 |
| 2006 | Fuzzy systems design: direct and indirect approaches
Plamen Angelov 0001, Costas S. Xydeas |
Soft Comput. | 2 |
| 2005 | Objective Image Fusion Performance CharacterisationabstractImage fusion as a way of combining multiple image signals into a single fused image has in recent years been extensively researched for a variety of multisensor applications. Choosing an optimal fusion approach for each application from the plethora of algorithms available however, remains a largely open issue. A small number of metrics proposed so far provide only a rough, numerical estimate of fusion performance with limited understanding of the relative merits of different fusion schemes. This paper proposes a method for comprehensive, objective, image fusion performance characterisation using a fusion evaluation framework based on gradient information representation. The method provides an in-depth analysis of fusion performance by quantifying: information contributions by each sensor, fusion gain, fusion information loss and fusion artifacts (artificial information created). It is demonstrated on the evaluation of an extensive dataset of multisensor images fused with a wide range of established image fusion algorithms. The results demonstrate and quantify a number of well known issues concerning the performance of these schemes and provide a useful insight into a number of more subtle yet important fusion performance effects not immediately accessible to an observer. Vladimir S. Petrovic, Costas S. Xydeas |
ICCV | 2 |
| 2004 | Evaluation of Image Fusion Performance with Visible Differences
Vladimir S. Petrovic, Costas S. Xydeas |
ECCV (3) | 2 |
| 2004 | On-line identification of MIMO evolving Takagi- Sugeno fuzzy modelsabstractEvolving Takagi-Sugeno (eTS) fuzzy models and the method for their on-line identification has been recently introduced as an effective tool for design of flexible system models with minimum a priori information. Their structure develops on-line during the process of model identification itself. In this paper, this approach has been extended for the case of multi-input multi-output (MIMO) system model. Both parts of the identification algorithm, namely the unsupervised fuzzy rule-base antecedents learning by a recursive, noniterative clustering, and the supervised linear sub-model parameters learning by Kalman-filtering-based procedure, are extended for the MIMO case. The radius of influence of each fuzzy rule is considered a vector instead of a scalar as in the original eTS approach, allowing different areas of the data space to be covered by each input variable. As in the eTS, in MIMO eTS, the rule-base and parameters of the fuzzy model continually evolve by adding new rules with more summarization power and by modifying existing rules and parameters. Simulation results using a well-known benchmark are considered in this paper. Further investigation concern the application of MIMO eTS to predictive modeling of the speech spectrum magnitude, classification of multi-channel source modulation etc. Plamen Angelov 0001, Costas S. Xydeas, Dimitar P. Filev |
FUZZ-IEEE | 2 |
| 2004 | Behavior Modeling Using a Hierarchical HMM ApproachabstractWe introduce a new methodology for the hierarchical modeling of the behavior-with-time of players operating and interacting within a certain application domain. Behavior modelling and characterization are performed online, given that a number of observations are made or sensed at regular time intervals with respect to each player. A key element of this hierarchical behavior modeling system architecture is a new formulation of multiple hidden Markov models (HMM) with discrete densities operating in parallel, with each HMM accepting a single feature-related observation sequence. However the proposed classification approach recognizes the existence of possible dependencies between the observation sequences of the features obtained for a given player. This property is effectively exploited in a new dependent-multiHMM with discrete densities (DM-HMM-D) classification approach. The proposed methodology is applied in modeling the behavior of aircrafts operating in relatively simple 3D "air-patrol" situations. Computer simulation results demonstrate the significant gains that can be obtained in system classification and modeling performance when compared to those obtained while using conventional independent-multidiscrete hidden Markov model (IM-HMM-D) schemes. Shih-Yang Chiao, Costas S. Xydeas |
HIS | 2 |
| 2004 | Gradient-based multiresolution image fusionabstractA novel approach to multiresolution signal-level image fusion is presented for accurately transferring visual information from any number of input image signals, into a single fused image without loss of information or the introduction of distortion. The proposed system uses a "fuse-then-decompose" technique realized through a novel, fusion/decomposition system architecture. In particular, information fusion is performed on a multiresolution gradient map representation domain of image signal information. At each resolution, input images are represented as gradient maps and combined to produce new, fused gradient maps. Fused gradient map signals are processed, using gradient filters derived from high-pass quadrature mirror filters to yield a fused multiresolution pyramid representation. The fused output image is obtained by applying, on the fused pyramid, a reconstruction process that is analogous to that of conventional discrete wavelet transform. This new gradient fusion significantly reduces the amount of distortion artefacts and the loss of contrast information usually observed in fused images obtained from conventional multiresolution fusion schemes. This is because fusion in the gradient map domain significantly improves the reliability of the feature selection and information fusion processes. Fusion performance is evaluated through informal visual inspection and subjective psychometric preference tests, as well as objective fusion performance measurements. Results clearly demonstrate the superiority of this new approach when compared to conventional fusion systems. Vladimir S. Petrovic, Costas S. Xydeas |
IEEE Trans. Image Process. | 2 |
| 2003 | Total Quality Management for Electronic Educational MarketsabstractThe concept of an electronic educational market, which is an "open" system for the exchange and brokerage of electronic learning resources between institutions of higher education is introduced and its value and feasibility are demonstrated within the paradigm of the EducaNext portal. The brokerage system can deal with highly heterogeneous learning resources, ranging from asynchronous educational to educational activities like computer-mediated lectures and courses. The main aim of such an endeavour is to develop and validate a scalable exchange model, which embraces offers, enquiries, booking and controlled delivery of learning resources. The key innovation is to create and manage an open electronic educational market with a standardized way of describing the pedagogical, administrative and technical characteristics of learning resources. Electronic educational markets enable institutions to enrich their curricula with remotely sourced material. The emphasis here is placed on the quality management of electronic educational markets. Lampros K. Stergioulas, Hassan Ahmed, Costas S. Xydeas, Bernd Simon |
ICALT | 3 |
| 2003 | Model-based packet loss concealment for AMR codersabstractA general packet loss correction/concealment signal recovery framework is proposed for parametric speech coders. Both redundancy-based forward error correction (FEC) and receiver only (RO) techniques are considered in conjunction with the adaptive multi-rate (AMR) coder. The robust, low bit rate, high communications speech quality Manchester pitch synchronous (MPS) coder is employed in the proposed systems as a secondary coding process. Thus the performance of AMR/MPS-FEC/RO packetised speech transmission systems is considered. Subjective and objective computer simulation results clearly indicate the superiority of the proposed schemes over conventional AMR based systems, particularly at relatively high (>5%) packet loss rates. Costas S. Xydeas, Fotis Zafeiropoulos |
ICASSP (1) | 1 |
| 2003 | Classification of Decision-Behavior Patterns in Multivariate Computer Log Data Using Independent Component Analysis
Serafeim Fragos, Lampros K. Stergioulas, Costas S. Xydeas |
KES | 3 |
| 1999 | Segmental prototype interpolation codingabstractCurrent parametric speech coding schemes can achieve high communications quality speech at bit rates in the range of 2.4 to 1.5 kbits/sec. Most schemes sample and quantise, at regular intervals, the "tracks in time" generated by the parameters of the speech production model. As a result, reconstructed "parameter tracks" do not evolve "smoothly" with time. Furthermore, no advantage is taken of the "linguistic event" nature of speech. In this paper, model parameter "time tracks" are split into non-overlapping speech "event" related segments. These segment based evolutions of model parameters are then vector quantised to provide at the receiver a smooth and subjectively meaningful reconstruction. Thus the paper presents an application of this generic segmental speech model quantisation approach to a 1.5 kbits/sec prototype interpolation coding (PIC) system. Results indicate that the proposed methodology can almost halve the bit rate of this PIC system while preserving overall recovered speech quality. Costas S. Xydeas, Thomas M. Chapman |
ICASSP | 1 |
| 1999 | Secondary codebook storage quantisationabstractThis paper presents an intonation generation system for use in a text-to-speech synthesis system. The intonation generation system uses classification trees to predict intonation event location and regression trees to predict parameters relating to the F0 shape for the predicted events. The decision trees model intonation within the Tilt intonation model, which provides a parameterized description of fundmaental frequency and an intuitive labelling scheme. The event location trees predict an event class (e.g. accent, boundary, none) for each syllable in an utterance based on local and global context (e.g. stress, phrasing, part of speech). The parameter prediction trees then provide the parameterized description of each intonation event based on similar context features. Informal results of the full system are presented together with results for the individual components. 1. INTRODUCTION Most of the currently available speech synthesizers have some sort of intonation generation modul... Thomas M. Chapman, Costas S. Xydeas |
EUROSPEECH | 2 |
| 1999 | Split matrix quantization of LPC parametersabstractThis paper examines in detail the design issues and performance characteristics of linear predictive coding (LPC) split matrix quantization (SMQ). This efficient LPC quantization method which was proposed by Xydeas and Papanastasiou (1995) can be viewed as an extension of the conventional split vector quantization (SVQ) process. SMQ removes existing interframe/intraframe line spectral frequency (LSF) redundancy by applying VQ principles on trajectories of smoothly evolving, with time, LSF coefficients. Using a 20 ms LPC analysis frame size, "transparent" quantization is achieved at 900 b/s, whereas "high quality" LSF quantization is easily obtained at 650 b/s, Furthermore, the SMQ methodology offers valuable flexibility in the way quantization of LPC coefficients is performed and leads into several schemes of varying computational complexity/storage characteristics. Costas S. Xydeas, Charalampos Papanastasiou |
IEEE Trans. Speech Audio Process. | 1 |
| 1998 | Multicodebook vector quantization of LPC parametersabstractThis paper presents a novel and efficient variable bit rate LPC quantization approach. The proposed MCVQ framework allows a dynamic programming based minimum quantization distortion partitioning and quantization process to be performed on input LSP vector tracks in time. Variable duration segments of LSP vector tracks are classified into one of a finite number of language related events. Specific codebooks, designed optimally for each event type, are then employed to vector quantize the individual LSP vectors of a given segment. "high quality" LSP quantization can be easily achieved at an average of 700 bits/sec while "transparent" performance is obtained at an average rate of 800 bits/sec. Costas S. Xydeas, Thomas M. Chapman |
ICASSP | 1 |
| 1997 | Efficient mixed excitation models in LPC based prototype interpolation speech codersabstractThis paper presents a new and efficient method for modeling voiced, mixed excitation spectra in sinusoidal (SC) and prototype interpolation coding (PIC) systems. Speech harmonics are classified as "weak-voiced" or "strong-voiced" by simply examining the short-term residual magnitude spectrum. This information is encoded effectively in terms of fixed width frequency bands and is used to control sets of periodic and random sine wave oscillators which model the short-term mixed excitation nature of speech. In this way the model allows the mixing of periodic and random signal energy on a harmonic basis. The proposed methodology has been used in a 2.4 Kbits/sec speech coder, whose recovered speech quality is better than that of the 4.8 Kbits/sec DoD standard. Charalampos Papanastasiou, Costas S. Xydeas |
ICASSP | 2 |
| 1997 | A novel 1.7/2.4 kb/s DCT based prototype interpolation speech coding system
Costas S. Xydeas, Gokhan H. Ilk |
EUROSPEECH | 1 |
| 1996 | Source driven variable bit rate prototype interpolation codingabstractCurrent developments in high quality, low bit rate speech coders are focused at bit rates in the region of 2.4 kbps and below. This paper presents a source driven variable bit rate coding scheme which operates in conjunction with a prototype interpolation coder (PIC). The proposed SD-VBRC-PIC coder explores effectively the residual redundancy that can be found in the evolution with time trajectories of the PIC model parameters, using a variable synthesis frame size technique. Furthermore, this approach is generic and VBRC can be applied to other prototype interpolation coding systems. Computer simulation experiments show that on average up to a maximum of 30% bit rate savings can be achieved for the same recovered speech quality, as compared to the fixed bit rate of the PIC system under consideration. Costas S. Xydeas, Binshi Cao |
ICASSP | 1 |
| 1996 | Pitch synchronous multi-band (PSMB) coding of speech signals
Haiyun Yang, Soo-Ngee Koh, Pratab Sivaprakasapillai, Costas S. Xydeas |
Speech Commun. | 4 |
| 1996 | A comparison of system architectures for intelligent document understanding
Gary S. D. Farrow, Costas S. Xydeas, John P. Oakley, A. Khorabi, Nuria González-Prelcic |
Signal Process. Image Commun. | 2 |
| 1995 | Efficient coding of LSP parameters using split matrix quantisationabstractThis paper presents a new and efficient LPC quantisation scheme called split matrix quantisation (SMQ). The proposed method can be viewed as an extension of the conventional split vector quantisation process. It operates over N consecutive LPC frames and effectively divides a p/spl times/N LSP matrix into K submatrices which are then vector quantised independently. SMQ exploits the interframe redundancy that exists between consecutive sets of LSP coefficients and achieves "transparent" quantisation at 900 bits/sec. "High quality" LSP quantisation can be easily obtained at 750 bits/sec. These bit rates are based in a 20 msec LPC analysis frame size. Furthermore, SMQ is characterised by relatively low complexity and low storage requirements. Costas S. Xydeas, Charalampos Papanastasiou |
ICASSP | 1 |
| 1995 | Model matching in intelligent document understandingabstractIntelligent Document Understanding (IDU) is the process of converting scanned document pages into an electronic, processable form. We have previously presented a IDU system architecture suitable for this task which uses a hybrid bottom-up/top-down control strategy. In this paper we focus on a specific subproblem that arises within the chosen framework, concerned with selecting an appropriate page layout structure. A detailed analysis of the problem using an error propagation model, allows computationally simple search strategies to be developed. A multistage layout formation algorithm is proposed and its performance is critically assessed when implemented using two different Layout Object selection criterion. The first selection criterion is based on a maximal column area coverage; the second is based on a probabilistic Layout Object selection. Both techniques have been incorporated into the hybrid IDU system and the results presented indicate its superiority over previously reported systems. Gary S. D. Farrow, Costas S. Xydeas, John P. Oakley |
ICDAR | 2 |
| 1994 | Conversion of scanned documents to the open document architectureabstractThe paper presents a system for the conversion of scanned documents into the open document architecture. Unlike previous work in this field the authors use a combination of evidence sources to achieve greater robustness to document defects and noise introduced in the scanning process. Furthermore, they use optical character recognition in conjunction with other forms of image analysis as a means of detecting document structure. This enables enhanced document feature extraction and improved performance. They demonstrate the performance of the system on a specific class of input document.> Gary S. D. Farrow, Costas S. Xydeas, John P. Oakley |
ICASSP (5) | 2 |
| 1994 | Detecting the skew angle in document images
Gary S. D. Farrow, Mark Ireton, Costas S. Xydeas |
Signal Process. Image Commun. | 3 |
| 1993 | A long history quantization approach to scalar and vector quantization of LSP coefficients
Costas S. Xydeas, K. K. M. So |
ICASSP (2) | 1 |
| 1992 | The Manchester Multimedia Information System
Carole A. Goble, Michael O'Docherty, Peter Crowther, Mark Ireton, John P. Oakley, Costas S. Xydeas |
EDBT | 6 |
| 1991 | A novel graph-theoretic texture segmentation algorithmabstractA new texture segmentation algorithm is described which is invariant under image spatial rotations and linear gray level transformations. The algorithm exploits the properties of the shortest spanning tree and involves both local and global information. The spanning tree is formed by taking into account the relationship not only between neighboring pixels but also between pixels in the surrounding regions. As a result pixels which are not nearest neighbors can interact during the segmentation process, thus enabling an overall reduction in chaining effects and an improvement in the noise immunity characteristics of the system. Texture segmentation is achieved by optimal partitioning of the spanning tree in a hierarchical way so as to form a spanning forest which conforms to a homogeneity requirement for the regions.> H. Ahmed, C. N. Daskalakis, Costas S. Xydeas |
ICASSP | 3 |
| 1989 | On improving vector excitation coders through the use of spherical lattice codebooks (SLCs)abstractTwo novel techniques for use in VXC (vector excitation coding) speech coders are presented. The first enables massive excitation codebooks (>or=20 b) to be used at realizable complexities by using a novel spherical lattice codebook for the excitation codebook. The second technique is a generalization of the gain-optimized error measure which allows any number of gains to be calculated for each excitation vector. This multiple-gain VXC can be thought of as a hybrid between multipulse and VXC.> Mark Ireton, Costas S. Xydeas |
ICASSP | 2 |
| 1986 | Coding Scheme for Scanned Images of Handwriting and Graphics
Costas S. Xydeas, Murray J. J. Holt |
ICC | 1 |
| 1984 | Split-Band Coding of Speech Signals Using A Transform Technique
F. S. Yeoh, Costas S. Xydeas |
ICC (3) | 2 |
| 1984 | Embedding Data Into Pictures by Modulo MaskingabstractThree systems are proposed for embedding data into industrial quality monochrome analog pictures. The video signal on each scan line is sampled, and a data bit is inserted into a block of three or five pels by modulo masking scrambling the luminance level of only one pel in the block. Prior to transmission, the combined data and video sequence is converted into a continuous signal with a bandwidth that is no greater than that of the original video signal. Using six images each containing 65 536 pels, Systems 1 and 2 embedded an average of 17 430 and 8713 bits per image, while System 3 accommodated data at a constant rate of 21 760 bits/image. The data embedding procedures of Systems 1, 2, and 3 operated with average picture SNR's of 41, 44, and 30 dB, respectively, when the transmission channel was ideal. When the transmission was over a channel composed of a second-order Butterworth filter plus additive noise that yield a channel SNR of 40 dB, no bit errors occurred but System 3 offered the greater safety margin to bit errors than Systems 1 and 2. Costas S. Xydeas, Branko Kostic, Raymond Steele |
IEEE Trans. Commun. | 1 |
| 1983 | Frequency compression of 7.6 kHz speed into 3.3 kHz bandwidthabstractTelephone channels restrict the bandwidth of speech signals to approximately 0.3 to 3.3 kHz, with the consequence that the intelligibility of unvoiced sounds may be significantly impaired. To prevent this bandlimitation of unvoiced sounds while still confining the speech to the telephonic bandwidth, we propose a scheme which on recognizing the presence of unvoiced sounds extending to 7.6 kHz, frequency maps them into the band 0.3 to 3.3 kHz. Four mapping laws are considered and the unvoiced speech is compressed using each law. Frequency de-mapping is employed, and the law that has the best spectral match to the speech spectrum is selected. Voiced speech is bandlimited from 0.3 to 3.3 kHz. Results indicate that the adaptive frequency mapping algorithm significantly enhances the recovered speech compared to telephonic speech. P. J. Patrick, Raymond Steele, Costas S. Xydeas |
ICASSP | 3 |
| 1983 | Frequency Compression of 7.6 kHz Speech Into 3.3 kHz BandwidthabstractTelephone channels restrict the bandwidth of speech signals to approximately 0.3-3.3 kHz, with the consequence that the intelligibility of unvoiced sounds may be significantly impaired. To prevent this band limitation of unvoiced sounds while still confining the speech to the telephonic bandwidth, we propose a scheme which, on recognizing the presence of unvoiced sounds extending to 7.6 kHz, frequency maps them into the band 0.3-3.3 kHz. Four mapping laws are considered and the unvoiced speech is compressed using each law. Frequency demapping is employed, and the law that has the best spectral match to the speech spectrum is selected. Voiced speech is band limited from 0.3 to 3.3 kHz. Results measured over 16 ms, a phoneme, and word durations indicate that the adaptive frequency mapping algorithm significantly enhances the recovered speech compared to telephonic speech. Informal listening experiences support these findings. P. J. Patrick, Raymond Steele, Costas S. Xydeas |
IEEE Trans. Commun. | 3 |
| 1982 | A comparative study of DPCM-AQF speech coders for bit rates of 16-32 kb/sabstractA comparative study of DPCM-AQF speech digitizers employing various second order prediction algorithms is presented. A new prediction scheme called Correlation Switched Prediction is also discussed. Computer simulation results for the range of 16 to 32 kbits/sec, indicate an improved segmented signal-to-noise ratio (SNRSEG) performance when using Forward Block Adaptive (FBA) instead of Backward Sequentially Adaptive Prediction. The introduction, in DPCM-AQF, of the relatively simple Correlation Switched Predictor produces SNRSEG values comparable to those obtained from DPCM-AQF-FBA. A more complex prediction scheme which combines Correlation Switched and Backward Sequential Adaptive Prediction (CSP-SGEP) is also considered. DPCM-AQF using the latter prediction is shown to provide the best overall SNRSEG performance. Costas S. Xydeas, Cumhur Cengiz Evci |
ICASSP | 1 |
| 1982 | Sequential Adaptive Predictors for ADPCM Speech EncodersabstractThe sequential gradient estimation predictor is compared in detail to the stochastic approximation predictor, and both are evaluated in an ADPCM codec. A switched predictor having two coefficients is then described for use in a DPCM-AQF codec. This predictor divides the range of the correlation coefficient of the speech signal into zones, and as the correlation coefficient changes zones, the predictor coefficients undergo a substantial modification. By this method the adaptation rate of the predictor is improved, particularly during transitions between unvoiced and voiced sounds. Costas S. Xydeas, Cumhur Cengiz Evci, Raymond Steele |
IEEE Trans. Commun. | 1 |
| 1981 | Wideband quality speech encoders with bit rates of 16-32 kbits/sabstractTelephone channels restrict the bandwidth of speech signals to within the range of 0.3 to 3.4 kHz and in doing so cause significant attenuation and distortion in unvoiced sounds with resulting loss of intelligibility and quality. A bandwidth compression system is proposed in this paper, whereby 0.3 to 7.6 kHz band limited speech is frequency mapped into the 0.3 to 3.4 kHz telephonic bandwidth range and is subsequently digitized prior to transmission. The received signal is decoded and spectrally expanded using a corresponding frequency de-mapping process. Consideration is given to the digital encoding of the compressed signals for bit rates between 16 and 32 kbits/s. We find that the quality of the recovered speech is improved compared to band limited telephonic speech. P. J. Patrick, Costas S. Xydeas, Raymond Steele, Wai-kuen Cham |
ICASSP | 2 |
| 1980 | Envelope Dynamic Ratio QuantizerabstractThe envelope dynamic ratio quantizer (envelope DRQ) is an adaptive quantizer for speech signals. By utilizing the envelope of the speech, nonlinear elements, a fixed quantizer and a simple predictor, a closed-loop adaptive quantizer emerges having a high constant SNR over a wide dynamic range. The theory of the quantizer is presented, together with computer simulation results which show an improvement.compared to the one word memory APCM system. Finally, the simplicity of implementing the envelope DRQ is described. Costas S. Xydeas, M. N. Faruqui, Raymond Steele |
IEEE Trans. Commun. | 1 |
| 1979 | Sequential gradient estimation predictor for speech signalsabstractA sequential gradient estimation predictor [SGEP] for speech signals is presented. In a given sampling interval, each prediction coefficient in turn is increased and decreased in value by a prescribed amount, while the other coefficients are kept constant. Two predictions are then made and the better enables the coefficient to be modified in the correct direction, but by an amount determined by a number of factors. The superiority of SGEP over the stochastic approximation predictor is illustrated by waveforms and snr (signal to noise ratio) performance curves. Cumhur Cengiz Evci, Raymond Steele, Costas S. Xydeas |
ICASSP | 3 |