VLDB 2026 Research / reviewers in the wild / expert
Changxue Ma
dblp:09/5271
· DBLP profile ↗
17ranked-venue papers
11as first author
0since 2021 · last 2012
0009-0001-8527-6131ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 first-authorArtificial intelligence and machine learning · 9 · 5 first-authorComputer networks · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Audio and music processing · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › speaker diarization
speaker clustering |
0.1 | 1 | 2012 | Statistical Utterance Comparison for Speaker Clustering Using Factor Analysis · IEEE Trans. Speech Audio Process. 2012 |
Audio and music processing
speech processing |
0.0 | 1 | 1994 | A Frobenius norm approach to glottal closure detection from the speech signal · IEEE Trans. Speech Audio Process. 1994 |
Methods — techniques the papers use, named apart from their topics
factor analysis · 0.1eigenvoice model · 0.1eigenchannel model · 0.1total linear least squares · 0.0singular value decomposition · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Statistical Utterance Comparison for Speaker Clustering Using Factor AnalysisabstractWe propose a novel method of measuring the similarity between two or more speech utterances for speaker clustering, based on probability theory and factor analysis. The similarity function is formulated as the probability that the utterances originated from the same speaker, and uses statistical eigenvoice and eigenchannel models to incorporate physical knowledge of interspeaker and intraspeaker variabilities, allowing the similarity function to be trainable and robust. The comparison function can be efficiently computed using a compact set of sufficient statistics for each speech utterance, allowing the acoustic features to be discarded. We begin using only eigenvoices, and then show how the eigenchannels can be incorporated into the equation to result in an identical form but with a different set of sufficient statistics. We test the proposed model in a speaker clustering task using the CALLHOME telephone conversation corpus and show that it performs better than two other well-known similarity measures: the Cross-Likelihood Ratio (CLR) and Generalized Likelihood Ratio (GLR). Woojay Jeon, Changxue Ma, Dusan Macho |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Efficient search of music pitch contours using wavelet transforms and segmented dynamic time warpingabstractWe propose a method of music melody matching based on their "continuous" (or "time-frame-based") pitch contours. Most previous methods using frame-based contours either made limiting assumptions on the locations, musical scale, tempo, and/or rhythm of the queries in relations to the targets, or involved exhaustive dynamic time-warping procedures that were too computation-intensive for use in real scenarios. In the proposed method, variable-scale windowing and wavelet transformations are performed at an initial coarse search stage to efficiently match queries to targets. At the following fine search stage, we apply a novel segmented dynamic time warping (DTW) method for melody contours, computing a more accurate distance between the query and each of the candidate targets with less computation than traditional DTW. The method searches arbitrary target locations and explicitly adjusts for differences in tempo and musical scale between queries and targets as well as rhythmic inconsistencies within queries. At the same time, retrieval times are fast enough for use in real-world scenarios. Woojay Jeon, Changxue Ma |
ICASSP | 2 |
| 2011 | An utterance comparison model for speaker clustering using factor analysisabstractWe propose a novel utterance comparison model based on probability theory and factor analysis that computes the likelihood of two speech utterances originating from die same speaker. The model depends only on a set of statistics extracted from each utterance and can efficiently compare utterances using these statistics without requiring die indefinite storage of speech features. We apply the model as a distance metric for speaker clustering in die CALLHOME telephone conversation corpus to achieve competitive results compared to three other known similarity measures: the Generalized Likelihood Ratio, Cross-Likelihood Ratio, and eigenvoice distance. Woojay Jeon, Changxue Ma, Dusan Macho |
ICASSP | 2 |
| 2009 | Efficient speech indexing and search for embedded devices using unitermsabstractIn this paper, we present an efficient method of speech indexing and search using phoneme sequences called uniterms. In the indexing stage, a collection of uniterms and uniterm sequences is extracted from the target speech database by applying statistical scoring to each data item's phoneme lattice. In the search stage, each speech query's phoneme lattice is used to select candidate uniterms from the collection. These uniterms are applied in a speech recognition engine to convert the speech query into a uniterm lattice, from which we obtain a set of candidate uniterm sequences, each of which can be mapped to a search result item. Not only is this method a significant improvement over previous phoneme-based methods, it is shown that explicit sequential comparison of uniterms in query and target data can be avoided using the proposed method without loss of search performance. Avoiding sequential comparison allows better handling of transposition of words, and for the case where queries have word orders different from their intended targets, the proposed method can potentially bring about significant improvement. Changxue Ma, Woojay Jeon |
ICASSP | 1 |
| 2008 | Uniterm Voice Indexing and Search for Mobile DevicesabstractIn this paper we present two novel approaches for voice indexing and search. The first approach is a Uniterm based voice indexing and search scheme that can be used for the fast retrieval of voice tagged multimedia contents on mobile devices. Uniterms, a string of phonemes with high scores, are extracted, in the indexing stage, from the phoneme lattice. For retrieval, the Uniterms are scored against the latent lattice model from the query voice. The candidate Uniterm list is selected to generate the audio segments of which the best phoneme paths are retrieved. These best paths of the candidate audio segments are compared against the best path of query voice and the search results are generated among the best matches. The second approach is also presented for comparison where the indices are extracted from the lattice in the form of unigram and bigram feature vectors. Each audio segment is represented by a tf-idf modulated feature vector. The search process involves two stages: the coarse search looks up the index and quickly returns a set of candidates; the fine search then compares the best paths of the query voice to the phone lattices of the candidates by using dynamic programming. Experimental results show that the Uniterm approach is significantly better and it is feasible for such voice search approaches on voice tagged multimedia contents on mobile devices. Finally we introduce a real-time news broadcasting search system. Changxue Ma |
ICCCN | 1 |
| 2006 | Cross-language evaluation of voice-to-phoneme conversions for voice-tag application in embedded platforms
Yan Ming Cheng, Changxue Ma, Lynette Melnar |
INTERSPEECH | 2 |
| 2004 | Automatic phonetic base form generation based on maximum context treeabstractTo improve the performance and the usability of the speech recognition devices, it is necessary for most applications to allow users to enter new words or personalize words in the system vocabulary. The voicetagging technique is a simple example of using speaker dependent spoken samples to generate baseform transcriptions of the spoken words. More sophisticated techniques can use both spoken samples and text versions of the new words to generate baseform transcriptions. In this paper, we propose a maximum context tree (MCT) based approach to the problem. Comparison is made to the common decision tree based method and Pronunciation by Analogy (PbA) approach. The new approach gives exact baseform transcription for in-vocabulary words and it shows better performance than decision tree. It performs significantly better than PbA approach with less memory usages. MCT uses the word segment probability rather than frequency count used in PbA. MCT uses the full context for the focus letter to overcome the some deficiencies in the PbA approach. Changxue Ma |
INTERSPEECH | 1 |
| 2003 | Novel robust feature extraction based on spectrally masked channel energy ratio (SMaChER) for speech recognitionabstractComparing speech recognition performance between human beings and computer, the latter's performance degrades dramatically in a noisy mobile environment. Based on the perceptual study of speech sounds in terms of auditory masking and speech perception study, it has been suggested that the contribution of each auditory channel is rather independent as used in articulation index model. From the point of view of the signal process perspective, we can diminish the noise caused variability by using spectral masking so that the noise signal between spectrum gaps is masked and limit the error propagation between frequency (auditory) channels. Based on this observation, we propose a novel robust feature extraction based spectrally masked channel energy ratio (SMaChER). We show the significant error reduction rate on noisy data has been achieved. Changxue Ma |
ICASSP (2) | 1 |
| 2003 | An approach to multilingual acoustic modeling for portable devices
Yan Ming Cheng, Yuanjun Wei, Lynette Melnar, Changxue Ma |
INTERSPEECH | 5 |
| 2001 | A support vector machines-based rejection technique for speech recognitionabstractSupport vector machines represent a new approach to pattern classification developed from the theory of structural risk minimization. In this paper, we present an investigation into the application of support vector machines to the confidence measurement problem in speech recognition. Specifically, based on the results from an initial decoding of an utterance during speech recognition, we derive a feature vector consisting of parameters such as word score density, N-best word score density differences, relative word score and relative word duration as input to the confidence measurement process in which hypothetically correct utterances are accepted and utterances determined to be incorrect are rejected. We propose a new approach to training support vector machines. In this paper, we train and test a support vector machines classifier and compare the results with other statistical classification methods. Changxue Ma, Mark A. Randolph, Joe Drish |
ICASSP | 1 |
| 2001 | An approach to automatic phonetic baseform generation based on Bayesian networksabstractTo improve the performance and the usability of the speech recognition devices, It is necessary for most applications to allow users to enter new words or personalize words to the system vocabulary. Voice-tagging technique is a simple example that use speaker dependent spoken sample to generate baseform transcriptions of the spoken words. More sophisticated techniques can use both spoken samples and texts of the new words to generate baseform transcriptions. In this paper, we propose a new approach to the problem. We use Bayesian networks to model the letter-to-sound rule probabilities. Compared to the common decision tree based method, This new approach shows a definite advantage. Changxue Ma, Mark A. Randolph |
INTERSPEECH | 1 |
| 1994 | The masking of narrowband noise by broadband harmonic complex sounds and implications for the processing of speech sounds
Changxue Ma, Douglas D. O'Shaughnessy |
Speech Commun. | 1 |
| 1994 | A Frobenius norm approach to glottal closure detection from the speech signalabstractThe detection of glottal closure instants has been a necessary step in several applications of speech processing, such as voice source analysis, speech prosody manipulation and speech synthesis. The paper presents a new algorithm for glottal closure detection that compares favorably with other methods available in terms of robustness and computational efficiency. The authors propose to use the singular value decomposition (SVD) approach to detect the instants of glottal closure from the speech signal. The proposed SVD method amounts to calculating the Frobenius norms of signal matrices and therefore is computationally efficient. Moreover, it produces well-defined and reliable peaks that indicate the instants of glottal closure. Finally, with the introduction of the total linear least squares technique, two other proposed methods are reinvestigated and unified into the SVD framework.> Changxue Ma, Yves Kamp, Lei F. Willems |
IEEE Trans. Speech Audio Process. | 1 |
| 1993 | Connection between weighted LPC and higher-order statistics for AR model estimationabstractThis paper establishes the relationship between a weighted linear prediction method used for robust analysis of voiced speech and the autoregressive modelling based on higher-order statistics, known as cumulants Yves Kamp, Changxue Ma |
EUROSPEECH | 2 |
| 1993 | The influence of temporal processes on spectral masking patterns of harmonic complex tones and vowelsabstractThe detection of narrowband noise targets in maskers of equal-amplitude harmonics has been reported in an earlier contribution [4]. The results revealed that the masking patterns are strongly dependent on the fundamental frequency of the masker. In the present paper we report masking patterns of broadband harmonic complexes as a function of their spectral tilt and level, and the masking pattern of a vowel sound. We will focus on how the spectral masking patterns are influenced by temporal processes in the auditory system Changxue Ma, Armin Kohlrausch |
EUROSPEECH | 1 |
| 1993 | A psychophysical study of fourier phase and amplitude coding of speech
Changxue Ma, Douglas D. O'Shaughnessy |
EUROSPEECH | 1 |
| 1993 | Robust signal selection for linear prediction analysis of voiced speech
Changxue Ma, Yves Kamp, Lei F. Willems |
Speech Commun. | 1 |