Jorge F. Silva

dblp:13/10857 · DBLP profile ↗
← Back
40ranked-venue papers
24as first author
6since 2021 · last 2026
0000-0002-0256-282XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 9 first-authorTheory of computation · 6 · 6 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Enabling online maximum driving range prognostics in electric vehicles via uncertain event likelihood functions
abstract
The increasing adoption of electric vehicles (EVs) demands accurate methodologies for predicting their maximum driving range (MDR) under dynamic and uncertain operating conditions. Monte Carlo simulations (MCs) serve as a fundamental tool for MDR prognostics. However, their enormous computational cost makes real-time implementation impractical. In order to address this challenge, we propose a novel framework that replaces MC-based methods with Uncertain Event Likelihood Functions (UELFs) for real-time driving range prognostics, significantly reducing computational overhead while maintaining predictive accuracy. The UELF framework leverages probabilistic modeling and efficient numerical solutions to predict the likelihood of end-of-power availability events, complemented by machine learning (ML) models such as a stochastic dropout-based Gated Recurrent Unit for vehicle speed and a Light Gradient Boosting Machine for energy consumption prediction. These models are trained on data from various geographic and environmental conditions, ensuring generalization and robustness. Test results confirm that the UELF-based approach achieves accuracy comparable to MC simulations while reducing computational time by over 99 %, enabling seamless integration into online applications. Combining state-of-the-art ML methodologies with advanced uncertainty quantification strategies successfully bridges the gap between abstract mathematical models and practical engineering solutions, offering a scalable and efficient tool for advancing electromobility and supporting real-time decision-making algorithms in EV energy management and route planning.
Jorge E. García Bustos, Benjamín Brito Schiele, Bruno Masserano, Ricardo Salas-Espiñeira, Diego Troncoso-Kurtovic, David Acuña-Ureta, Francisco Jaramillo, Marcos E. Orchard, Jorge F. Silva, Aramis Pérez
Eng. Appl. Artif. Intell.9
2025 Understanding encoder-decoder structures in machine learning using information measures
abstract
We present a theory of representation learning to model and understand the role of encoder–decoder design in machine learning (ML) from an information-theoretic angle. We use two main information concepts, information sufficiency (IS) and mutual information loss to represent predictive structures in machine learning. Our first main result provides a functional expression that characterizes the class of probabilistic models consistent with an IS encoder–decoder latent predictive structure. This result formally justifies the encoder–decoder forward stages many modern ML architectures adopt to learn latent (compressed) representations for classification. To illustrate IS as a realistic and relevant model assumption, we revisit some known ML concepts and present some interesting new examples: invariant, robust, sparse, and digital models. Furthermore, our IS characterization allows us to tackle the fundamental question of how much performance could be lost, using the cross entropy risk, when a given encoder–decoder architecture is adopted in a learning setting. Here, our second main result shows that a mutual information loss quantifies the lack of expressiveness attributed to the choice of a (biased) encoder–decoder ML design. Finally, we address the problem of universal cross-entropy learning with an encoder–decoder design where necessary and sufficiency conditions are established to meet this requirement. In all these results, Shannon’s information measures offer new interpretations and explanations for representation learning.
Jorge F. Silva, Victor Faraggi, Camilo Ramírez, Alvaro Egaña, Eduardo Pavez
Signal Process.1
2024 Studying the Interplay between Information Loss and Operation Loss in Representations for Classification
abstract
Information-theoretic measures have been widely adopted for machine learning (ML) feature design. Inspired by this, we look at the relationship between information loss in the Shannon sense and the operation loss in the minimum probability of error (MPE) sense when considering a family of lossy representations. Our first result offers a lower bound on a weak form of information loss as a function of its respective operation loss when adopting a discrete encoder. When considering a general family of lossy continuous representations, we show that a form of vanishing information loss (a weak informational sufficiency (WIS)) implies a vanishing MPE loss. Our findings support the observation that selecting/designing representations that capture informational sufficiency is appropriate for learning. However, this selection is rather conservative if the intended goal is achieving MPE in classification. Supporting this, we show that it is possible to adopt an alternative notion of informational sufficiency (strictly weaker than pure sufficiency in the mutual information sense) to achieve operational sufficiency in learning. Furthermore, our new WIS condition is used to demonstrate the expressive power of digital encoders and the capacity of two existing compression-based algorithms to achieve lossless prediction in ML.
Jorge F. Silva, Felipe A. Tobar, Mario Vicuña, Felipe Cordova
J. Mach. Learn. Res.1
2022 A Data-Driven Quantization Design for Distributed Testing Against Independence with Communication Constraints
abstract
This paper studies the problem of designing a quantizer (encoder) for the task of distributed detection of independence subject to one-side communication (limited bits) constraints. By exploiting the asymptotic performance limits as an objective to train a quantization scheme, we propose an algorithm that addresses an info-max problem for this lossy compression task. Tools from machine learning are incorporated to facilitate our data-driven optimization. Experiments on synthetic data support our design principle and approximations, expressing that the devised solutions are effective in compressing data while preserving the relevant information for the underlying task of testing against independence.
Sebastian Espinosa, Jorge F. Silva, Pablo Piantanida
ICASSP2
2022 On Universal D-Semifaithful Coding for Memoryless Sources With Infinite Alphabets
abstract
The problem of variable length and fixed-distortion universal source coding (or D-semifaithful source coding) for stationary and memoryless sources on countably infinite alphabets ($\infty $-alphabets) is addressed in this paper. The main results of this work offer a set of sufficient conditions (from weaker to stronger) to obtain weak minimax universality, strong minimax universality, and corresponding achievable rates of convergences for the worst-case redundancy for the family of stationary memoryless sources whose densities are dominated by an envelope function (or the envelope family) on$\infty $-alphabets. An important implication of these results is that universal D-semifaithful source coding is not feasible for the complete family of stationary and memoryless sources on$\infty $-alphabets. To demonstrate this infeasibility, a sufficient condition for the impossibility is presented for the envelope family. Interestingly, it matches the well-known impossibility condition in the context of lossless (variable-length) universal source coding. More generally, this work offers a simple description of what is needed to achieve universal D-semifaithful coding for a family of distributions$\Lambda $. This reduces to finding a collection of quantizations of the product space at different block-lengths — reflecting the fixed distortion restriction — that satisfy two asymptotic requirements: the first is a universal quantization condition with respect to$\Lambda $, and the second is a vanishing information radius (I-radius) condition for$\Lambda $reminiscent of the condition known for lossless universal source coding.
Jorge F. Silva, Pablo Piantanida
IEEE Trans. Inf. Theory1
2021 Finite-Length Bounds on Hypothesis Testing Subject to Vanishing Type I Error Restrictions
abstract
A central problem in Binary Hypothesis Testing (BHT) is to determine the optimal tradeoff between the Type I error (referred to as false alarm) and Type II (referred to as miss) error. In this context, the exponential rate of convergence of the optimal miss error probability - as the sample size tends to infinity - given some (positive) restrictions on the false alarm probabilities is a fundamental question to address in theory. Considering the more realistic context of a BHT with a finite number of observations, this letter presents a new non-asymptotic result for the scenario with monotonic (sub-exponential decreasing) restriction on the Type I error probability, which extends the result presented by Strassen in 2009. Building on the use of concentration inequalities, we offer new upper and lower bounds to the optimal Type II error probability for the case of finite observations. Finally, the derived bounds are evaluated and interpreted numerically (as a function of the number samples) for some vanishing Type I error restrictions.
Sebastian Espinosa, Jorge F. Silva, Pablo Piantanida
IEEE Signal Process. Lett.2
2020 Data-Driven Representations for Testing Independence: A Connection with Mutual Information Estimation
abstract
From the design of a data-driven partition, this paper addresses the problem of testing independence between two multidimensional random variables from i.i.d. samples. The empirical log-likelihood statistics is adopted with the objective of approximating the sufficient statistics of a test against independence that knows the two distributions (the oracle test). It is shown that approximating the sufficient statistics of the oracle test (asymptotically) offers a connection with the problem of estimating mutual information. Applying these ideas in the context of a data-dependent tree-structured partition (TSP), we derive concrete sufficient conditions on the parameters of the TSP scheme to obtain a strongly consistent test of independence distribution-free over the family of joint probabilities equipped with densities.
Mauricio E. Gonzalez, Jorge F. Silva
ISIT2
2020 Universal Weak Variable-Length Source Coding on Countably Infinite Alphabets
abstract
Motivated from the fact that universal source coding on countably infinite alphabets (∞-alphabets) is not feasible, this work introduces the notion of “almost lossless source coding”. Analog to the weak variable-length source coding problem studied by Han (IEEE Trans. Inf. Theory, vol. 46, no. 4, pp. 1217-1226, Jul. 2000), almost lossless source coding aims at relaxing the lossless block-wise assumption to allow an average per-letter distortion that vanishes asymptotically as the block-length tends to infinity. In this setup, we show on one hand that Shannon entropy characterizes the minimum achievable rate (similarly to the case of finite alphabet sources) while on the other that almost lossless universal source coding becomes feasible for the family of finite-entropy stationary memoryless sources with ∞-alphabets. Furthermore, we study a stronger notion of almost lossless universality that demands uniform convergence of the average per-letter distortion to zero, where we establish a necessary and sufficient condition for the so-called family of “envelope distributions” to achieve it. Remarkably, this condition is the same necessary and sufficient condition needed for the existence of a strongly minimax (lossless) universal source code for the family of envelope distributions. Finally, we show that an almost lossless coding scheme offers faster rate of convergence for the (minimax) redundancy compared to the well-known information radius developed for the lossless case at the expense of tolerating a non-zero distortion that vanishes to zero as the block-length grows. This shows that even when lossless universality is feasible, an almost lossless scheme can offer different regimes on the rates of convergence of the (worst case) redundancy versus the (worst case) distortion.
Jorge F. Silva, Pablo Piantanida
IEEE Trans. Inf. Theory1
2019 Universal D-Semifaithfull Coding for Countably Infinite Alphabets
abstract
The problem of fixed-distortion universal source coding (or universal D-semifaithfull source coding) for stationary and memoryless sources on countably infinite alphabets is investigated. The main result of this paper offers sufficient conditions to achieve strong and weak minimax universality for D-semifaithfull source coding of any memoryless source defined by an envelope function on infinite alphabets. It is also shown that universal D-semifaithfull source coding is not feasible for the complete family of memoryless sources. Furthermore, a sufficient condition for impossibility is presented for the family of envelope distributions.
Jorge F. Silva, Pablo Piantanida
ISIT1
2018 Improving battery voltage prediction in an electric bicycle using altitude measurements and kernel adaptive filters
Felipe A. Tobar, Iván Castro, Jorge F. Silva, Marcos E. Orchard
Pattern Recognit. Lett.3
2017 The redundancy gains of almost lossless universal source coding over envelope families
abstract
The problem of almost lossless universal coding is revisited in this work. We study uniform rate of convergence for distortion and redundancy over a family of envelope distributions. In particular, we show that an almost lossless coding scheme offers faster rate of convergence for the (minimax) redundancy compared with the well-known information radius developed for the lossless case at the expense of tolerating a non-zero distortion that vanishes to zero as the block-length grows. Our results show that even when lossless universality is feasible, an almost lossless scheme can still offer different regimes on the rates of convergence of the redundancy versus the distortion.The problem of almost lossless universal coding is revisited in this work. We study uniform rate of convergence for distortion and redundancy over a family of envelope distributions. In particular, we show that an almost lossless coding scheme offers faster rate of convergence for the (minimax) redundancy compared with the well-known information radius developed for the lossless case at the expense of tolerating a non-zero distortion that vanishes to zero as the block-length grows. Our results show that even when lossless universality is feasible, an almost lossless scheme can still offer different regimes on the rates of convergence of the redundancy versus the distortion.
Jorge F. Silva, Pablo Piantanida
ISIT1
2016 Almost lossless variable-length source coding on countably infinite alphabets
abstract
Motivated from the fact that universal source coding on countably infinite alphabets is not feasible, the notion of almost lossless source coding is introduced. This idea -analog to the weak variable-length source coding problem proposed by Han [1]- aims at relaxing the lossless block-wise assumption to allow a distortion that vanishes asymptotically as the block-length goes to infinity1. In this setup, both feasibility and optimality results are derived for the case of memoryless sources defined on countably infinite alphabets. Our results show on one hand that Shannon entropy characterizes the minimum achievable rate (known statistics) while on the other that almost lossless universal source coding becomes feasible for the family of finite entropy stationary and memoryless sources with countably infinite alphabets.
Jorge F. Silva, Pablo Piantanida
ISIT1
2015 Information-Theoretic Measures and Sequential Monte Carlo Methods for Detection of Regeneration Phenomena in the Degradation of Lithium-Ion Battery Cells
abstract
This paper analyses and compares the performance of a number of approaches implemented for the detection of capacity regeneration phenomena (measured in ampere-hours) in the degradation trend of energy storage devices, particularly Lithium-Ion battery cells. All implemented approaches are based on a combination of information-theoretic measures and sequential Monte Carlo methods for state estimation in nonlinear, non-Gaussian dynamic systems. Properties of information measures are conveniently used to quantify the impact of process measurements on the posterior probability density function of the state, assuming that sub-optimal Bayesian estimation algorithms (such as classic or risk-sensitive particle filters) are to be used to obtain an empirical representation of the system uncertainty. The proposed anomaly detection strategies are tested and evaluated both in terms of (i) detection time (early detection) and (ii) false alarm rates. Verification of detection schemes is performed using simulated data for battery State-Of-Health accelerated degradation tests, to ensure absolute knowledge on the time instant where a regeneration phenomenon occurs.
Marcos E. Orchard, Matias S. Lacalle, Benjamín E. Olivares, Jorge F. Silva, Rodrigo Palma-Behnke, Pablo A. Estévez, Bernardo Severino, Williams Calderon-Munoz, Marcelo Cortes-Carmona
IEEE Trans. Reliab.4
2015 Particle-Filtering-Based Discharge Time Prognosis for Lithium-Ion Batteries With a Statistical Characterization of Use Profiles
abstract
We present the implementation of a particle-filtering-based prognostic framework that utilizes statistical characterization of use profiles to (i) estimate the state-of-charge (SOC), and (ii) predict the discharge time of energy storage devices (lithium-ion batteries). The proposed approach uses a novel empirical state-space model, inspired by battery phenomenology, and particle-filtering algorithms to estimate SOC and other unknown model parameters in real-time. The adaptation mechanism used during the filtering stage improves the convergence of the state estimate, and provides adequate initial conditions for the prognosis stage. SOC prognosis is implemented using a particle-filtering-based framework that considers a statistical characterization of uncertainty for future discharge profiles based on maximum likelihood estimates of transition probabilities for a two-state Markov chain. All algorithms have been trained and validated using experimental data acquired from one Li-Ion 26650 and two Li-Ion 18650 cells, and considering different operating conditions.
Daniel A. Pola, Hugo F. Navarrete, Marcos E. Orchard, Ricardo S. Rabie, Matías A. Cerda Munoz, Benjamín E. Olivares, Jorge F. Silva, Pablo A. Espinoza, Aramis Pérez
IEEE Trans. Reliab.7
2014 Precise best k-term approximation error analysis of ergodic processes
abstract
The characterization of ℓp-compressible random sequences is revisited and extended to the case of stationary and ergodic processes. The main result of this work offers a simple-to-check necessary and sufficient condition for a stationary and ergodic sequence to be ℓp-compressible in the sense proposed by Amini, Unser and Marvasti [1, Def. 6]. Furthermore, for non ℓp-compressible random sequences, we provide a closed-form expression for the best k-term relative approximation error given a rate of coefficients as the block-length tends to infinity.
Jorge F. Silva, Milan S. Derpich
ISIT1
2013 Shannon entropy estimation from convergence results in the countable alphabet case
abstract
In this paper new results for the Shannon entropy estimation and estimation of distributions, consistently in information divergence, are presented in the countable alphabet case. Sufficient conditions for the entropy convergence are adopted, including scenarios with both finitely and infinitely supported distributions. From this approach, new estimates, strong consistency results and rate of convergences are derived for various plug-in histogram-based schemes.
Jorge F. Silva, Patricio Parada
ITW1
2012 Shannon entropy convergence results in the countable infinite case
abstract
The convergence of the Shannon entropy in the countable infinity case is revisited and extended in this work. New results are presented that provide necessary and sufficient conditions for the convergence of the entropy in different settings, including scenarios with both finitely and infinitely supported measures. These results show some connections between the Shannon entropy convergence and the convergence in information divergence.
Jorge F. Silva, Patricio Parada
ISIT1
2012 On signal representations within the Bayes decision framework
Jorge F. Silva, Shri Narayanan
Pattern Recognit.1
2012 Analysis and design of Wavelet-Packet Cepstral coefficients for automatic speech recognition
Eduardo Pavez, Jorge F. Silva
Speech Commun.2
2012 Complexity-Regularized Tree-Structured Partition for Mutual Information Estimation
abstract
A new histogram-based mutual information estimator using data-driven tree-structured partitions (TSP) is presented in this paper. The derived TSP is a solution to a complexity regularized empirical information maximization, with the objective of finding a good tradeoff between the known estimation and approximation errors. A distribution-free concentration in equality for this tree-structured learning problem as well as finite sample performance bounds for the proposed histogram-based solution is derived. It is shown that this solution is density-free strongly consistent and that it provides, with an arbitrary high probability, an optimal balance between the mentioned estimation and approximation errors. Finally, for the emblematic scenario of independence, I(X; Y) = 0, it is shown that the TSP estimate converges to zero with o(e-n1/3+ log log n).
Jorge F. Silva, Shri Narayanan
IEEE Trans. Inf. Theory1
2011 Necessary and sufficient conditions for zero-rate density estimation
abstract
This work addresses the problem of universal density estimation under an operational data-rate constraint. We present a coding theorem that stipulates necessary and sufficient conditions to learn and transmit a memoryless source distribution with arbitrary precision (in total variations), under an asymptotic zero-rate regime, in bits per sample. In the process, we propose a concrete coding scheme to achieve this learning objective, adopting the Skeleton estimate developed by Y. Yatracos [1], [2].
Jorge F. Silva, Milan S. Derpich
ITW1
2011 Sufficient conditions for the convergence of the Shannon differential entropy
abstract
This work revisits and extends results concerning the convergence of the Shannon differential entropy. Concrete connections with the convergence of probability measures in the sense of total variations and (direct and reverse) information divergence are established. In particular, under uniform bounded conditions on the sequence of probability measures, the results stipulate that the convergence in information divergence is sufficient to guarantee the convergence of the differential entropy functional.
Jorge F. Silva, Patricio Parada
ITW1
2010 A near-optimal (minimax) tree-structured partition for mutual information estimation
abstract
A novel histogram-based mutual information estimator using data-driven tree-structured partitions (TSP) is presented in this work. The TSP is the solution of a complexity regularized empirical information maximization (EIM) criterion, with the objective to find a good tradeoff between the known estimation and approximation errors. We show that this solution is density-free strongly consistent and, furthermore, it provides a near-optimal balance between the mentioned variance-bias errors.
Jorge F. Silva, Shri Narayanan
ISIT1
2010 On data-driven histogram-based estimation for mutual information
abstract
The problem of mutual information (MI) estimation based on data-dependent partition is addressed in this work. Sufficient conditions are stipulated on a histogram-based construction to guarantee a strong consistent estimate for the MI. The practical implications of this result are in the specification of a range of design parameters for two data-driven histogram-based approaches - statistically equivalent blocks and tree-structure vector quantizations - to yield density-free strongly consistent estimates for the MI.
Jorge F. Silva, Shri Narayanan
ISIT1
2009 Histogram-based estimation for the divergence revisited
abstract
This work revisits and extends the problem of consistent divergence estimation using data-dependent partitions. For distributions defined on (¿d, B(¿d)), the main result characterizes sufficient conditions on a data-dependent partition scheme to get a strongly consistent histogram-based estimate of the divergence.
Shri Narayanan, Jorge F. Silva
ISIT2
2007 Optimal Wavelet Packets Decomposition Based on a Rate-Distortion Optimality Criterion
abstract
We address the problem of optimal decomposition of wavelet packets (WPs) for pattern recognition based on the minimum probability of error signal representation (MPE-SR) principle. The problem is formulated as a complexity regularized optimization, where the tree-indexed structure of the WP family is used to reduce it to a type of minimum cost tree pruning problem used in regression and classification trees (CART). MPE-SR solutions are obtained for a frame level phone recognition task showing promising performance results.
Jorge F. Silva, Shri Narayanan
ICASSP (3)1
2007 Information Theoretic Analysis of Direct Articulatory Measurements for Phonetic Discrimination
abstract
This paper focuses on the analysis of speech production signals (physical measurements from electromagnetic articulograph) from the perspective of phone discrimination. We explore two different signal representation schemes for the articulatory signals, one based on time-domain analysis and the other based on frequency domain. We quantify the amount of discrimination information offered by the speech production signals in identifying the phone labels through mutual information. Mutual information analyses establish that substantial discrimination information is present in the articulatory stream. Furthermore, phonological classification results with articulatory signals indicate higher accuracy compared to the acoustic signal.
Jorge F. Silva, Vivek Kumar Rangarajan Sridhar, Viktor Rozgic, Shri Narayanan
ICASSP (4)1
2007 Universal Consistency of Data-Driven Partitions for Divergence Estimation
abstract
This paper presents a general histogram based divergence estimator based on data-dependent partition. Sufficient conditions for the universal strong consistency of the data-driven divergence estimator, using Lugosi and Nobel's combinatorial notions for partition families, are presented. As a corollary this result is particularized for the emblematic case oflm-spacing quantization scheme.
Jorge F. Silva, Shri Narayanan
ISIT1
2007 Soft indexing of speech content for search in spoken documents
Ciprian Chelba, Jorge F. Silva, Alex Acero
Comput. Speech Lang.2
2006 Pruning Analysis for the Position Specific Posterior Lattices for Spoken Document Search
abstract
The paper presents the position specific posterior lattice (PSPL), a novel lossy representation of automatic speech recognition lattices that naturally lends itself to efficient indexing and subsequent relevance ranking of spoken documents. Two pruning techniques for generating word lattices are explored in this framework, where experiments performed on a collection of lecture recordings - MIT iCampus database - show that the spoken document ranking accuracy was improved by 20% - in the mean average precision sense - relative over the commonly used baseline of indexing the 1-best output from an automatic speech recognizer (ASR)
Jorge F. Silva, Ciprian Chelba, Alex Acero
ICASSP (1)1
2006 Automatic detection of voice onset time contrasts for use in pronunciation assessment
abstract
This study examines methods for recognizing different classes of phones from accented speech based on voice on-set time (VOT). These methods are tested on data from the Tball corpus of Los Angeles-area elementary school chil-dren [1]. The methods proposed and tested are: 1) to train models based on standard English VOT contrasts and then extract the VOT characteristics of the phones by measuring the duration of phone-level and sub-phone-level alignments, 2) to train phone models with explicit aspiration, and 3) to train different models for different phoneme classes of VOT times. Error rates of 23-53 % for different phone classes are reported for the rst method, 5-57 % for the second method, and 0-36 % for the third. The results show that different methods work better on different phone classes. We inter-pret these results in relation to past research on VOT, explain possible uses for these ndings, and propose directions for future research. Index Terms: recognition of speech variation, pronuncia-tion variation, voice onset time.
Abe Kazemzadeh, Joseph Tepperman, Jorge F. Silva, Hong You, Sungbok Lee, Abeer Alwan, Shri Narayanan
INTERSPEECH3
2006 Pronunciation verification of children²s speech for automatic literacy assessment
abstract
Arguably the most important part of automatically assessing a new reader's literacy is in verifying his pronunciation of read-aloud target words.But the pronunciation evaluation task is especially difficult in children, non-native speakers, and pre-literates.Traditional likelihood ratio thresholding methods do not generalize easily, and even expert human evaluators do not always agree on what constitutes an acceptable pronunciation.We propose new recognition-and alignment-based features in a decision tree classification framework, along with the use of prior linguistic information and human perceptual evaluations.Our classification methods demonstrate a 91% agreement with the voted results of 20 human evaluators who agree among themselves 85% of the time.
Joseph Tepperman, Jorge F. Silva, Abe Kazemzadeh, Hong You, Sungbok Lee, Abeer Alwan, Shri Narayanan
INTERSPEECH2
2006 Upper Bound Kullback-Leibler Divergence for Hidden Markov Models with Application as Discrimination Measure for Speech Recognition
abstract
This paper presents a criterion for defining an upper bound Kullback-Leibler divergence (UB-KLD) for Gaussian mixtures models (GMMs). An information theoretic interpretation of this indicator and an algorithm for calculating it based on similarity alignment between mixture components of the models are proposed. This bound is used to characterize an upper bound closed-form expression for the Kullback-Leibler divergence (KLD) for left-to-right transient hidden Markov models (HMMs), where experiments based on real speech data show that this indicator precisely follows the discrimination tendency of the actual KLD
Jorge F. Silva, Shri Narayanan
ISIT1
2006 Integration of Metadata in spoken Document Search Using Position Specific Posterior latices
abstract
This paper addresses the problem of integrating speech and text content sources for the document search problem, as well as its usefulness from an ad-hoc retrieval -keyword search - point of view. Position specific posterior latices (PSPL) is naturally extended to deal with both speech and text content, where a new relevance ranking framework is proposed for integrating the different sources of information available. Experimental results on the MIT iCampus corpus show a relative improvement of 302% in mean average precision (MAP) when using speech content and metadata as opposed to just metadata (which constitutes about 1% of the amount of words in the transcription of the speech content).
Jorge F. Silva, Ciprian Chelba, Alex Acero
SLT1
2006 Average divergence distance as a statistical discrimination measure for hidden Markov models
abstract
This paper proposes and evaluates a new statistical discrimination measure for hidden Markov models (HMMs) extending the notion of divergence, a measure of average discrimination information originally defined for two probability density functions. Similar distance measures have been proposed for the case of HMMs, but those have focused primarily on the stationary behavior of the models. However, in speech recognition applications, the transient aspects of the models have a principal role in the discrimination process and, consequently, capturing this information is crucial in the formulation of any discrimination indicator. This paper proposes the notion of average divergence distance (ADD) as a statistical discrimination measure between two HMMs, considering the transient behavior of these models. This paper provides an analytical formulation of the proposed discrimination measure, a justification of its definition based on the Viterbi decoding approach, and a formal proof that this quantity is well defined for a left-to-right HMM topology with a final nonemitting state, a standard model for basic acoustic units in automatic speech recognition (ASR) systems. Using experiments based on this discrimination measure, it is shown that ADD provides a coherent way to evaluate the discrimination dissimilarity between acoustic models.
Jorge F. Silva, Shri Narayanan
IEEE Trans. Speech Audio Process.1
2006 Modeling, estimating, and compensating low-bit rate coding distortion in speech recognition
abstract
A solution to the problem of speech recognition with signals distorted by low-bit rate coders is presented in this paper. A model for the coding-decoding distortion, a HMM compensation method to include this model, and an EM-based adaptation algorithm to estimate this distortion are proposed here. Medium vocabulary continuous-speech speaker-independent recognition experiments with 8 kbps G.729(CS-CELP), 13 kbps RPE-LTP (GSM), 5.3 kbps G723.1, 4.8 kbps FS-1016 and 32 kbps G.726(ADPCM) coders show that the approach described in this paper is able to dramatically reduce the effect of the coding distortion and, in some cases, gives a word accuracy higher than the baseline system with uncoded speech. Finally, the EM estimation algorithm requires only one adapting utterance and the approach described is certainly suitable for dialogue systems where just a few adapting utterances are available.
Néstor Becerra Yoma, Jorge F. Silva, Carlos Busso
IEEE Trans. Speech Audio Process.3
2004 A statistical discrimination measure for hidden Markov models based on divergence
Jorge F. Silva, Shri Narayanan
INTERSPEECH1
2003 Language model accuracy and uncertainty in noise cancelling in the stochastic weighted viterbi algorithm
Néstor Becerra Yoma, Ivan Brito, Jorge F. Silva
INTERSPEECH3
2002 MAP speaker adaptation of state duration distributions for speech recognition
abstract
This paper presents a framework for maximum a posteriori (MAP) speaker adaptation of state duration distributions in hidden Markov models (HMM). Four key issues of MAP estimation, namely analysis and modeling of state duration distributions, the choice of prior distribution, the specification of the parameters of the prior density and the evaluation of the MAP estimates, are tackled. Moreover, a comparison with an adaptation procedure based on maximum likelihood (ML) estimation is presented, and the problem of truncation of the state duration distribution is addressed from the statistical point of view. The results shown in this paper suggest that the speaker adaptation of temporal restrictions substantially improves the accuracy of speaker-independent (SI) HMM with clean and noisy speech. The method requires a low computational load and a small number of adapting utterances, and can be useful to follow the dynamics of the speaking rate in speech recognition.
Néstor Becerra Yoma, Jorge F. Silva
IEEE Trans. Speech Audio Process.2
2001 Speaker adaptation of output probabilities and state duration distributions for speech recognition
Néstor Becerra Yoma, Jorge F. Silva
INTERSPEECH2