VLDB 2026 Research / reviewers in the wild / expert
Pierre L. Dognin
dblp:68/8053
· DBLP profile ↗
33ranked-venue papers
13as first author
10since 2021 · last 2025
0000-0001-5688-6005ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 9 first-author · 4 since 2021Artificial intelligence and machine learning · 22 · 9 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Language models and text generation · 36% Trustworthy machine learning · 23% Knowledge representation and reasoning · 12% | |
| Theoretical computer science
2 papers |
Information theory · 82% Mathematical optimization · 18% |
Topics — the 29 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering |
0.9 | 1 | 2025 | Programming Refusal with Conditional Activation Steering · ICLR 2025 |
Natural language and speech › Language models and text generation
controllable text generation |
0.9 | 1 | 2025 | Programming Refusal with Conditional Activation Steering · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model safety
refusal behavior |
0.9 | 1 | 2025 | Programming Refusal with Conditional Activation Steering · ICLR 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | Programming Refusal with Conditional Activation Steering · ICLR 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
normative reasoning |
0.8 | 1 | 2024 | ComVas: Contextual Moral Values Alignment System · IJCAI 2024 |
Machine learning › Trustworthy machine learning › fairness
bias mitigation |
0.6 | 1 | 2022 | Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without Refitting · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › fairness › fairness criteria
demographic parity |
0.6 | 1 | 2022 | Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without Refitting · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › fairness › fairness criteria
equal opportunity |
0.6 | 1 | 2022 | Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without Refitting · NeurIPS 2022 |
Machine learning › Trustworthy machine learning
fairness |
0.6 | 1 | 2022 | Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without Refitting · NeurIPS 2022 |
Natural language and speech › Language models and text generation › text generation › data-to-text generation
graph-to-text generation |
0.5 | 1 | 2021 | ReGen: Reinforcement Learning for Text and Knowledge Base Generation using Pretrained Language Models · EMNLP (1) 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base construction |
0.5 | 1 | 2021 | ReGen: Reinforcement Learning for Text and Knowledge Base Generation using Pretrained Language Models · EMNLP (1) 2021 |
Machine learning › Representation and self-supervised learning
mutual information maximization |
0.5 | 1 | 2021 | Improved Mutual Information Estimation · AAAI 2021 |
Information theory › information measures › mutual information
mutual information estimation |
0.5 | 1 | 2021 | Improved Mutual Information Estimation · AAAI 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph reasoning
knowledge base completion |
0.4 | 1 | 2020 | DualTKB: A Dual Learning Bridge between Text and Knowledge Base · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › text generation › data-to-text generation
knowledge graph-to-text generation |
0.4 | 1 | 2020 | DualTKB: A Dual Learning Bridge between Text and Knowledge Base · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation
text generation |
0.4 | 1 | 2020 | Learning Implicit Text Generation via Feature Matching · ACL 2020 |
Machine learning › Generative modeling › generative adversarial network
conditional GAN |
0.4 | 1 | 2019 | Adversarial Semantic Alignment for Improved Image Captions · CVPR 2019 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.4 | 1 | 2019 | Wasserstein Barycenter Model Ensembling · ICLR (Poster) 2019 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2019 | Learning Implicit Generative Models by Matching Perceptual Features · ICCV 2019 |
Computer vision › Vision and language
image captioning |
0.4 | 1 | 2019 | Adversarial Semantic Alignment for Improved Image Captions · CVPR 2019 |
Machine learning › Generative modeling
implicit generative model |
0.4 | 1 | 2019 | Learning Implicit Generative Models by Matching Perceptual Features · ICCV 2019 |
Machine learning › Probabilistic and Bayesian machine learning
moment matching |
0.4 | 1 | 2019 | Learning Implicit Generative Models by Matching Perceptual Features · ICCV 2019 |
Computer vision › Image recognition and object detection › object detection
contextual reasoning |
0.2 | 1 | 2024 | ComVas: Contextual Moral Values Alignment System · IJCAI 2024 |
Machine learning › Deep learning architectures and training
neural network estimator |
0.1 | 1 | 2021 | Improved Mutual Information Estimation · AAAI 2021 |
Natural language and speech › Speech recognition and synthesis
acoustic modeling |
0.1 | 1 | 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced Languages · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
hidden markov model acoustic modeling |
0.1 | 1 | 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced Languages · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
low-resource speech recognition |
0.1 | 1 | 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced Languages · IEEE Trans. Speech Audio Process. 2012 |
Mathematical optimization › optimal transport
wasserstein barycenter |
0.1 | 1 | 2019 | Wasserstein Barycenter Model Ensembling · ICLR (Poster) 2019 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › continuous speech recognition
large vocabulary continuous speech recognition |
0.0 | 1 | 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced Languages · IEEE Trans. Speech Audio Process. 2012 |
Methods — techniques the papers use, named apart from their topics
reproducing kernel hilbert space · 1.0likelihood ratio estimation · 1.0donsker-varadhan bound · 1.0self-critical sequence training · 0.9activation steering · 0.9influence functions · 0.6infinitesimal jackknife · 0.6reinforcement learning · 0.5graph linearization · 0.5feature matching · 0.4wasserstein barycenter · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Contextual Value AlignmentabstractDeveloping value-aligned agents is a complex undertaking and an ongoing challenge in the field of AI. Indeed, designing Large Language Models (LLMs) that can balance multiple possibly conflicting moral values based on the context is a problem of paramount importance. In this paper, we propose a system that performs contextual value alignment based on contextual aggregation of possible responses. This aggregation is achieved by integrating a subset of possible LLM responses that are best suited to a user's input while taking into account features extracted about the user's moral preferences. The proposed system trained using the Moral Integrity Corpus displays better alignment to human values than state-of-the-art baselines. Pierre L. Dognin, Jesus Rios, Ronny Luss, Prasanna Sattigeri, Miao Liu 0001, Inkit Padhi, Matthew Riemer, Manish Nagireddy, Kush R. Varshney, Djallel Bouneffouf 0001 |
ICASSP | 1 |
| 2025 | Programming Refusal with Conditional Activation SteeringabstractLLMs have shown remarkable capabilities, but precisely controlling their response behavior remains challenging.
Existing activation steering methods alter LLM behavior indiscriminately, limiting their practical applicability in settings where selective responses are essential, such as content moderation or domain-specific assistants.
In this paper, we propose Conditional Activation Steering (CAST), which analyzes LLM activation patterns during inference to selectively apply or withhold activation steering based on the input context.
Our method is based on the observation that different categories of prompts activate distinct patterns in the model's hidden states.
Using CAST, one can systematically control LLM behavior with rules like "if input is about hate speech or adult content, then refuse" or "if input is not about legal advice, then refuse."
This allows for selective modification of responses to specific content while maintaining normal responses to other content, all without requiring weight optimization.
We release an open-source implementation of our framework. Bruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Erik Miehling, Pierre L. Dognin, Manish Nagireddy, Amit Dhurandhar |
ICLR | 5 |
| 2025 | Evaluating the Prompt Steerability of Large Language ModelsabstractErik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy, Elizabeth M. Daly, Kush R. Varshney, Eitan Farchi, Pierre Dognin, Jesus Rios, Djallel Bouneffouf, Miao Liu, Prasanna Sattigeri. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Erik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy, Elizabeth Daly, Kush R. Varshney, Eitan Farchi, Pierre L. Dognin, Jesus Rios, Djallel Bouneffouf 0001, Miao Liu 0001, Prasanna Sattigeri |
NAACL (Long Papers) | 7 |
| 2024 | ComVas: Contextual Moral Values Alignment System
Inkit Padhi, Pierre L. Dognin, Jesus Rios, Ronny Luss, Swapnaja Achintalwar, Matthew Riemer, Miao Liu 0001, Prasanna Sattigeri, Manish Nagireddy, Kush R. Varshney, Djallel Bouneffouf 0001 |
IJCAI | 2 |
| 2022 | Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without RefittingabstractIn consequential decision-making applications, mitigating unwanted biases in machine learning models that yield systematic disadvantage to members of groups delineated by sensitive attributes such as race and gender is one key intervention to strive for equity. Focusing on demographic parity and equality of opportunity, in this paper we propose an algorithm that improves the fairness of a pre-trained classifier by simply dropping carefully selected training data points. We select instances based on their influence on the fairness metric of interest, computed using an infinitesimal jackknife-based approach. The dropping of training points is done in principle, but in practice does not require the model to be refit. Crucially, we find that such an intervention does not substantially reduce the predictive performance of the model but drastically improves the fairness metric. Through careful experiments, we evaluate the effectiveness of the proposed approach on diverse tasks and find that it consistently improves upon existing alternatives. Prasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre L. Dognin, Kush R. Varshney |
NeurIPS | 4 |
| 2022 | Cloud-Based Real-Time Molecular Screening Platform with MolFormer
Brian M. Belgodere, Vijil Chenthamarakshan, Pierre L. Dognin, Toby Kurien, Igor Melnyk, Youssef Mroueh, Inkit Padhi, Mattia Rigotti, Jarret Ross, Yair Schiff, Richard A. Young |
ECML/PKDD (6) | 4 |
| 2022 | Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 ChallengeabstractImage captioning has recently demonstrated impressive progress largely owing to the introduction of neural network algorithms trained on curated dataset like MS-COCO. Often work in this field is motivated by the promise of deployment of captioning systems in practical applications. However, the scarcity of data and contexts in many competition datasets renders the utility of systems trained on these datasets limited as an assistive technology in real-world settings, such as helping visually impaired people navigate and accomplish everyday tasks. This gap motivated the introduction of the novel VizWiz dataset, which consists of images taken by the visually impaired and captions that have useful, task-oriented information. In an attempt to help the machine learning computer vision field realize its promise of producing technologies that have positive social impact, the curators of the VizWiz dataset host several competitions, including one for image captioning. This work details the theory and engineering from our winning submission to the 2020 captioning competition. Our work provides a step towards improved assistive image captioning systems. This article appears in the special track on AI & Society. Pierre L. Dognin, Igor Melnyk, Youssef Mroueh, Inkit Padhi, Mattia Rigotti, Jarret Ross, Yair Schiff, Richard A. Young, Brian M. Belgodere |
J. Artif. Intell. Res. | 1 |
| 2021 | Improved Mutual Information EstimationabstractWe propose to estimate the KL divergence using a relaxed likelihood ratio estimation in a Reproducing Kernel Hilbert space. We show that the dual of our ratio estimator for KL in the particular case of Mutual Information estimation corresponds to a lower bound on the MI that is related to the so called Donsker Varadhan lower bound. In this dual form, MI is estimated via learning a witness function discriminating between the joint density and the product of marginal, as well as an auxiliary scalar variable that enforces a normalization constraint on the likelihood ratio. By extending the function space to neural networks, we propose an efficient neural MI estimator, and validate its performance on synthetic examples, showing advantage over the existing baselines. We demonstrate its strength in large-scale self-supervised representation learning through MI maximization. Youssef Mroueh, Igor Melnyk, Pierre L. Dognin, Jarret Ross, Tom Sercu |
AAAI | 3 |
| 2021 | ReGen: Reinforcement Learning for Text and Knowledge Base Generation using Pretrained Language ModelsabstractAutomatic construction of relevant Knowledge Bases (KBs) from text, and generation of semantically meaningful text from KBs are both long-standing goals in Machine Learning.In this paper, we present ReGen, a bidirectional generation of text and graph leveraging Reinforcement Learning (RL) to improve performance.Graph linearization enables us to re-frame both tasks as a sequence to sequence generation problem regardless of the generative direction, which in turn allows the use of Reinforcement Learning for sequence training where the model itself is employed as its own critic leading to Self-Critical Sequence Training (SCST).We present an extensive investigation demonstrating that the use of RL via SCST benefits graph and text generation on WebNLG+ 2020 and TEKGEN datasets.Our system provides state-of-the-art results on WebNLG+ 2020 by significantly improving upon published results from the WebNLG 2020+ Challenge for both text-to-graph and graph-to-text generation tasks.More details in https Pierre L. Dognin, Inkit Padhi, Igor Melnyk |
EMNLP (1) | 1 |
| 2021 | Tabular Transformers for Modeling Multivariate Time SeriesabstractTabular datasets are ubiquitous in data science applications. Given their importance, it seems natural to apply state-of-the-art deep learning algorithms in order to fully unlock their potential. Here we propose neural network models that represent tabular time series that can optionally leverage their hierarchical structure. This results in two architectures for tabular time series: one for learning representations that is analogous to BERT and can be pre-trained end-to-end and used in downstream tasks, and one that is akin to GPT and can be used for generation of realistic synthetic tabular sequences. We demonstrate our models on two datasets: a synthetic credit card transaction dataset, where the learned representations are used for fraud detection and synthetic data generation, and on a real pollution dataset, where the learned encodings are used to predict atmospheric pollutant concentrations. Code and data are available at https://github.com/IBM/TabFormer. Inkit Padhi, Yair Schiff, Igor Melnyk, Mattia Rigotti, Youssef Mroueh, Pierre L. Dognin, Jerret Ross, Ravi Nair, Erik R. Altman |
ICASSP | 6 |
| 2020 | Learning Implicit Text Generation via Feature MatchingabstractInkit Padhi, Pierre Dognin, Ke Bai, Cícero Nogueira dos Santos, Vijil Chenthamarakshan, Youssef Mroueh, Payel Das. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Inkit Padhi, Pierre L. Dognin, Cícero Nogueira dos Santos, Vijil Chenthamarakshan, Youssef Mroueh |
ACL | 2 |
| 2020 | DualTKB: A Dual Learning Bridge between Text and Knowledge BaseabstractIn this work, we present a dual learning approach for unsupervised text to path and path to text transfers in Commonsense Knowledge Bases (KBs).We investigate the impact of weak supervision by creating a weakly supervised dataset and show that even a slight amount of supervision can significantly improve the model performance and enable better-quality transfers.We examine different model architectures, and evaluation metrics, proposing a novel Commonsense KB completion metric tailored for generative models.Extensive experimental results show that the proposed method compares very favorably to the existing baselines.This approach is a viable step towards a more advanced system for automatic KB construction/expansion and the reverse operation of KB conversion to coherent textual descriptions. Pierre L. Dognin, Igor Melnyk, Inkit Padhi, Cícero Nogueira dos Santos |
EMNLP (1) | 1 |
| 2019 | Adversarial Semantic Alignment for Improved Image CaptionsabstractIn this paper, we study image captioning as a conditional GAN training, proposing both a context-aware LSTM captioner and co-attentive discriminator, which enforces semantic alignment between images and captions. We empirically focus on the viability of two training methods: Self-critical Sequence Training (SCST) and Gumbel Straight-Through (ST) and demonstrate that SCST shows more stable gradient behavior and improved results over Gumbel ST, even without accessing discriminator gradients directly. We also address the problem of automatic evaluation for captioning models and introduce a new semantic score, and show its correlation to human judgement. As an evaluation paradigm, we argue that an important criterion for a captioner is the ability to generalize to compositions of objects that do not usually co-occur together. To this end, we introduce a small captioned Out of Context (OOC) test set. The OOC set, combined with our semantic score, are the proposed new diagnosis tools for the captioning community. When evaluated on OOC and MS-COCO benchmarks, we show that SCST-based training has a strong performance in both semantic score and human evaluation, promising to be a valuable new approach for efficient discrete GAN training. Pierre L. Dognin, Igor Melnyk, Youssef Mroueh, Jerret Ross, Tom Sercu |
CVPR | 1 |
| 2019 | Learning Implicit Generative Models by Matching Perceptual FeaturesabstractPerceptual features (PFs) have been used with great success in tasks such as transfer learning, style transfer, and super-resolution. However, the efficacy of PFs as key source of information for learning generative models is not well studied. We investigate here the use of PFs in the context of learning implicit generative models through moment matching (MM). More specifically, we propose a new effective MM approach that learns implicit generative models by performing mean and covariance matching of features extracted from pretrained ConvNets. Our proposed approach improves upon existing MM methods by: (1) breaking away from the problematic min/max game of adversarial learning; (2) avoiding online learning of kernel functions; and (3) being efficient with respect to both number of used moments and required minibatch size. Our experimental results demonstrate that, due to the expressiveness of PFs from pretrained deep ConvNets, our method achieves state-of-the-art results for challenging benchmarks. Cícero Nogueira dos Santos, Youssef Mroueh, Inkit Padhi, Pierre L. Dognin |
ICCV | 4 |
| 2019 | Wasserstein Barycenter Model Ensembling
Pierre L. Dognin, Igor Melnyk, Youssef Mroueh, Jerret Ross, Cícero Nogueira dos Santos, Tom Sercu |
ICLR (Poster) | 1 |
| 2015 | Evaluating Deep Scattering Spectra with deep neural networks on large scale spontaneous speech taskabstractDeep Scattering Network features introduced for image processing have recently proved useful in speech recognition as an alternative to log-mel features for Deep Neural Network (DNN) acoustic models. Scattering features use wavelet decomposition directly producing log-frequency spectrograms which are robust to local time warping and provide additional information within higher order coefficients. This paper extends previous works by showing how scattering features perform on a state-of-the-art spontaneous speech recognition utilizing DNN acoustic model. We revisit feature normalization and compression topics in an extensive study, putting emphasis on comparing models of the same size. We observe that scattering features outperform baseline log-mel in all conditions, with additional gains from multi-resolution processing. Petr Fousek, Pierre L. Dognin, Vaibhava Goel |
ICASSP | 2 |
| 2015 | Annealed dropout trained maxout networks for improved LVCSRabstractA significant barrier to progress in automatic speech recognition (ASR) capability is the empirical reality that techniques rarely “scale”-the yield of many apparently fruitful techniques rapidly diminishes to zero as the training criterion or decoder is strengthened, or the size of the training set is increased. Recently we showed that annealed dropout-a regularization procedure which gradually reduces the percentage of neurons that are randomly zeroed out during DNN training-leads to substantial word error rate reductions in the case of small to moderate training data amounts, and acoustic models trained based on the cross-entropy (CE) criterion [1]. In this paper we show that deep Maxout networks trained using annealed dropout can substantially improve the quality of commercial-grade LVCSR systems even when the acoustic model is trained with sequence-level training criterion, and on large amounts of data. Steven J. Rennie, Pierre L. Dognin, Vaibhava Goel |
ICASSP | 2 |
| 2013 | Combining stochastic average gradient and Hessian-free optimization for sequence training of deep neural networksabstractMinimum phone error (MPE) training of deep neural networks (DNN) is an effective technique for reducing word error rate of automatic speech recognition tasks. This training is often carried out using a Hessian-free (HF) quasi-Newton approach, although other methods such as stochastic gradient descent have also been applied successfully. In this paper we present a novel stochastic approach to HF sequence training inspired by recently proposed stochastic average gradient (SAG) method. SAG reuses gradient information from past updates, and consequently simulates the presence of more training data than is really observed for each model update. We extend SAG by dynamically weighting the contribution of previous gradients, and by combining it to a stochastic HF optimization. We term the resulting procedure DSAG-HF. Experimental results for training DNNs on 1500h of audio data show that compared to baseline HF training, DSAG-HF leads to better held-out MPE loss after each model parameter update, and converges to an overall better loss value. Furthermore, since each update in DSAG-HF takes place over smaller amount of data, this procedure converges in about half the time as baseline HF sequence training. Pierre L. Dognin, Vaibhava Goel |
ASRU | 1 |
| 2013 | Direct product based deep belief networks for automatic speech recognitionabstractIn this paper, we present new methods for parameterizing the connections of neural networks using sums of direct products. We show that low rank parameterizations of weight matrices are a subset of this set, and explore the theoretical and practical benefits of representing weight matrices using sums of Kronecker products. ASR results on a 50 hr subset of the English Broadcast News corpus indicate that the approach is promising. In particular, we show that a factorial network with more than 150 times less parameters in its bottom layer than its standard unconstrained counterpart suffers minimal WER degradation, and that by using sums of Kronecker products, we can close the gap in WER performance while maintaining very significant parameter savings. In addition, direct product DBNs consistently outperform standard DBNs with the same number of parameters. These results have important implications for research on deep belief networks (DBNs). They imply that we should be able to train neural networks with thousands of neurons and minimal restrictions much more rapidly than is currently possible, and that by using sums of direct products, it will be possible to train neural networks with literally millions of neurons tractably-an exciting prospect. Petr Fousek, Steven J. Rennie, Pierre L. Dognin, Vaibhava Goel |
ICASSP | 3 |
| 2013 | Adaptive stereo-based stochastic mapping
Shay Maymon, Pierre L. Dognin, Vaibhava Goel |
INTERSPEECH | 2 |
| 2012 | Factorial Hidden Restricted Boltzmann Machines for noise robust speech recognitionabstractWe present the Factorial Hidden Restricted Boltzmann Machine (FHRBM) for robust speech recognition. Speech and noise are modeled as independent RBMs, and the interaction between them is explicitly modeled to capture how speech and noise combine to generate observed noisy speech features. In contrast with RBMs, where the bottom layer of random variables is observed, inference in the FHRBM is intractable, scaling exponentially with the number of hidden units. We introduce variational algorithms for efficient approximate inference that scale linearly with the number of hidden units. Compared to traditional factorial models of noisy speech, which are based on GMMs, the FHRBM has the advantage that the representations of both speech and noise are highly distributed, allowing the model to learn a parts-based representation of noisy speech data that can generalize better to previously unseen noise compositions. Preliminary results suggest that the approach is promising. Steven J. Rennie, Petr Fousek, Pierre L. Dognin |
ICASSP | 3 |
| 2012 | Hidden Markov Acoustic Modeling With Bootstrap and Restructuring for Low-Resourced LanguagesabstractThis paper proposes an acoustic modeling approach based on bootstrap and restructuring to dealing with data sparsity for low-resourced languages. The goal of the approach is to improve the statistical reliability of acoustic modeling for automatic speech recognition (ASR) in the context of speed, memory and response latency requirements for real-world applications. In this approach, randomized hidden Markov models (HMMs) estimated from the bootstrapped training data are aggregated for reliable sequence prediction. The aggregation leads to an HMM with superior prediction capability at cost of a substantially larger size. For practical usage the aggregated HMM is restructured by Gaussian clustering followed by model refinement. The restructuring aims at reducing the aggregated HMM to a desirable model size while maintaining its performance close to the original aggregated HMM. To that end, various Gaussian clustering criteria and model refinement algorithms have been investigated in the full covariance model space before the conversion to the diagonal covariance model space in the last stage of the restructuring. Large vocabulary continuous speech recognition (LVCSR) experiments on Pashto and Dari have shown that acoustic models obtained by the proposed approach can yield superior performance over the conventional training procedure with almost the same run-time memory consumption and decoding speed. Peder A. Olsen, Pierre L. Dognin, Upendra V. Chaudhari, John R. Hershey, Bowen Zhou 0006 |
IEEE Trans. Speech Audio Process. | 5 |
| 2011 | Matched-condition robust Dynamic Noise AdaptationabstractIn this paper we describe how the model-based noise robustness algorithm for previously unseen noise conditions, Dynamic Noise Adaptation (DNA), can be made robust to matched data, without the need to do any system re-training. The approach is to do online model selection and averaging between two DNA models of noise: one that is tracking the evolving state of the background noise, and one clamped to the null mis-match hypothesis. The approach, which we call DNA with (matched) condition detection (DNA-CD), improves the performance of a commerical-grade speech recognizer that utilizes feature-space Maximum Mutual Information (fMMI), boosted MMI (bMMI), and feature-space Maximum Likelihood Linear Regression (fMLLR) compensation by 15% relative at signal-to-noise ratios (SNRs) below 10 dB, and over 8% relative overall. Steven J. Rennie, Pierre L. Dognin, Petr Fousek |
ASRU | 2 |
| 2011 | Robust speech recognition using dynamic noise adaptationabstractDynamic noise adaptation (DNA) is a model-based technique for improving automatic speech recognition (ASR) performance in noise. DNA has shown promise on artificially mixed data such as the Aurora II and DNA+Aurora II tasks - significantly outperforming well-known techniques like the ETSI AFE and fMLLR - but has never been tried on real data. In this paper, we present new results generated by commercial-grade ASR systems trained on large amounts of data. We show that DNA improves upon the performance of the spectral subtraction (SS) and stochastic fMLLR algorithms of our embedded recognizers, particularly in unseen noise conditions, and describe how DNA has been evolved to become suitable for deployment in low-latency ASR systems. DNA improves our best embedded system, which utilizes SS, fMLLR, and fMPE by over 22% relative at SNRs below 6 dB, reducing the word error rate in these adverse conditions from 4.24% to 3.29%. Steven J. Rennie, Pierre L. Dognin, Petr Fousek |
ICASSP | 2 |
| 2010 | Acoustic modeling with bootstrap and restructuring for low-resourced languages
Pierre L. Dognin, Upendra V. Chaudhari, Bowen Zhou 0006 |
INTERSPEECH | 3 |
| 2010 | Restructuring exponential family mixture modelsabstractVariational KL (varKL) divergence minimization was previously applied to restructuring acoustic models (AMs) using Gaussian mixture models by reducing their size while preserving their accuracy. In this paper, we derive a related varKL for exponential family mixture models (EMMs) and test its accuracy using the weighted local maximum likelihood agglomerative clustering technique. Minimizing varKL between a reference and a restructured AM led previously to the variational expectation maximization (varEM) algorithm; which we extend to EMMs. We present results on a clustering task using AMs trained on 50 hrs of Broadcast News (BN). EMMs are trained on fMMI-PLP features combined with frame level phone posterior probabilities given by the recently introduced sparse representation phone identification process. As we reduce model size, we test the word error rate using the standard BN test set and compare with baseline models of the same size, trained directly from data. Index Terms: KL divergence, variational approximation, variational expectation-maximization, exponential family distributions, acoustic model clustering. Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen |
INTERSPEECH | 1 |
| 2009 | A fast, accurate approximation to log likelihood of Gaussian mixture modelsabstractIt has been a common practice in speech recognition and elsewhere to approximate the log likelihood of a Gaussian mixture model (GMM) with the maximum component log likelihood. While often a computational necessity, the max approximation comes at a price of inferior modeling when the Gaussian components significantly overlap. This paper shows how the approximation error can be reduced by changing component priors. In our experiments the loss in word error rate due to max approximation, albeit small, is reduced by 50-100% at no cost in computational efficiency. Furthermore, we expect acoustic models will become larger with time and increase component overlap and word error rate loss. This makes reducing the approximation error more relevant. The techniques considered do not use the original data and can easily be applied as a post-processing step to any GMM. Pierre L. Dognin, Vaibhava Goel, John R. Hershey, Peder A. Olsen |
ICASSP | 1 |
| 2009 | Refactoring acoustic models using variational density approximationabstractIn model-based pattern recognition it is often useful to change the structure, or refactor, a model. For example, we may wish to find a Gaussian mixture model (GMM) with fewer components that best approximates a reference model. One application for this arises in speech recognition, where a variety of model size requirements exists for different platforms. Since the target size may not be known a priori, one strategy is to train a complex model and subsequently derive models of lower complexity. We present methods for reducing model size without training data, following two strategies: GMM-approximation and Gaussian clustering based on divergences. A variational expectation-maximization algorithm is derived that unifies these two approaches. The resulting algorithms reduce the model size by 50% with less than 4% increase in error rate relative to the same-sized model trained on data. In fact, for up to 35% reduction in size, the algorithms can improve accuracy relative to baseline. Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen |
ICASSP | 1 |
| 2009 | Refactoring acoustic models using variational expectation-maximization
Pierre L. Dognin, John R. Hershey, Vaibhava Goel, Peder A. Olsen |
INTERSPEECH | 1 |
| 2008 | Beyond linear transforms: efficient non-linear dynamic adaptation for noise robust speech recognition
Steven J. Rennie, Pierre L. Dognin |
INTERSPEECH | 2 |
| 2003 | A new spectral transformation for speaker normalization
Pierre L. Dognin, Amro El-Jaroudi |
INTERSPEECH | 1 |
| 2002 | The 2001 BYBLOS English large vocabulary conversational speech recognition systemabstractThis paper describes the BYBLOS system that BBN used to participate in the 2001 NIST Hub-5 evaluation benchmark. We outline the procedure used for training and decoding, and present the algorithmic improvements made to the system, along with experimental results. These improvements include a Gaussian splitting initialization procedure, the use of Linear Discriminant Analysis, and processing of additional acoustic training data. We also discuss our system combination and confidence-based thresholding methods. Experiments on an internal validation test set show that all these system improvements provide a 8.1 % relative reduction in word error rate compared to our 2000 LVCSR system. Spyridon Matsoukas, Thomas Colthurst, Owen Kimball, Alex Solomonoff, Fred Richardson, Carl Quillen, Herbert Gish, Pierre L. Dognin |
ICASSP | 8 |
| 2000 | Parameter optimization for vocal tract length normalizationabstractThis paper focuses on the optimization of model parameters for vocal tract length normalization (VTLN). For maximum likelihood (ML) based normalization techniques, the complexity of the VTL-models is a source of variation in system performance. An optimal complexity for the VTL-model that ensures best global word error rate is proposed. The choice of frequency warping factor also depends on the signal processing step of VTLN. A best set of parameters for the VTLN signal processing stage is proposed with extensive results for an optimal frequency range. Pierre L. Dognin, Amro El-Jaroudi, Jayadev Billa |
ICASSP | 1 |