VLDB 2026 Research / reviewers in the wild / expert
Hossein Hadian
dblp:151/8527
· DBLP profile ↗
13ranked-venue papers
6as first author
1since 2021 · last 2022
0000-0002-1083-6026ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 8 · 4 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › speech recognition
acoustic modeling |
0.3 | 1 | 2018 | Flat-Start Single-Stage Discriminatively Trained HMM-Based Models for ASR · IEEE ACM Trans. Audio Speech Lang. Process. 2018 |
Audio and music processing › speech recognition
discriminative training |
0.3 | 1 | 2018 | Flat-Start Single-Stage Discriminatively Trained HMM-Based Models for ASR · IEEE ACM Trans. Audio Speech Lang. Process. 2018 |
Audio and music processing
speech recognition |
0.3 | 1 | 2018 | Flat-Start Single-Stage Discriminatively Trained HMM-Based Models for ASR · IEEE ACM Trans. Audio Speech Lang. Process. 2018 |
Methods — techniques the papers use, named apart from their topics
neural network training · 0.3flat-start training · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Continual Learning Using Lattice-Free MMI for Speech RecognitionabstractContinual learning (CL), or domain expansion, recently became a popular topic for automatic speech recognition (ASR) acoustic modeling because practical systems have to be updated frequently in order to work robustly on types of speech not observed during initial training. While sequential adaptation allows tuning a system to a new domain, it may result in performance degradation on the old domains due to catastrophic forgetting. In this work we explore regularization-based CL for neural network acoustic models trained with the lattice-free maximum mutual information (LF-MMI) criterion. We simulate domain expansion by incrementally adapting the acoustic model on different public datasets that include several accents and speaking styles. We investigate two well-known CL techniques, elastic weight consolidation (EWC) and learning without forgetting (LWF), which aim to reduce forgetting by preserving model weights or network outputs. We additionally introduce a sequence-level LWF regularization, which exploits posteriors from the denominator graph of LF-MMI to further reduce forgetting. Empirical results show that the proposed sequence-level LWF can improve the best average word error rate across all domains by up to 9.4% relative compared with using regular LWF. Hossein Hadian, Arseniy Gorin |
ICASSP | 1 |
| 2020 | An Alternative to MFCCs for ASR
Pegah Ghahremani, Hossein Hadian, Daniel Povey, Hynek Hermansky, Sanjeev Khudanpur |
INTERSPEECH | 2 |
| 2019 | Using ASR Methods for OCRabstractHybrid deep neural network hidden Markov models (DNN-HMM) have achieved impressive results on large vocabulary continuous speech recognition (LVCSR) tasks. However, the recent approaches using DNN-HMM models are not explored much for text recognition. Inspired by the current work in automatic speech recognition (ASR) and machine translation, we present an open vocabulary sub-word text recognition system. The sub-word lexicon and sub-word language model (LM) helps in overcoming the challenge of recognizing out of vocabulary (OOV) words, and a time delay neural network (TDNN) and convolution neural network (CNN) based DNN-HMM optical model (OM) efficiently models the sequence dependency in the line image. We present results on 12 datasets with training data varying from 6k lines to 600k lines. The system is built for 8 languages, i.e., English, French, Arabic, Chinese, Farsi, Tamil, Russian, and Korean. We report competitive results on several commonly used handwritten and printed text datasets. Ashish Arora, L. Paola García-Perera, Shinji Watanabe 0001, Vimal Manohar, Yiwen Shao, Sanjeev Khudanpur, Chun-Chieh Chang, Babak Rekabdar, Bagher BabaAli, Daniel Povey, David Etter, Desh Raj, Hossein Hadian, Jan Trmal |
ICDAR | 13 |
| 2018 | Semi-Supervised Training of Acoustic Models Using Lattice-Free MMIabstractThe lattice-free MMI objective (LF-MMI) has been used in supervised training of state-of-the-art neural network acoustic models for automatic speech recognition (ASR). With large amounts of unsupervised data available, extending this approach to the semi-supervised scenario is of significance. Finite-state transducer (FST) based supervision used with LF-MMI provides a natural way to incorporate uncertainties when dealing with unsupervised data. In this paper, we describe various extensions to standard LF-MMI training to allow the use as supervision of lattices obtained via decoding of unsupervised data. The lattices are rescored with a strong LM. We investigate different methods for splitting the lattices and incorporating frame tolerances into the supervision FST. We report results on different subsets of Fisher English, where we achieve WER recovery of 59-64% using lattice supervision, which is significantly better than using just the best path transcription. Vimal Manohar, Hossein Hadian, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 2 |
| 2018 | A Time-Restricted Self-Attention Layer for ASRabstractSelf-attention - an attention mechanism where the input and output sequence lengths are the same - has recently been successfully applied to machine translation, caption generation, and phoneme recognition. In this paper we apply a restricted self-attention mechanism (with multiple heads) to speech recognition. By “restricted” we mean that the mechanism at a particular frame only sees input from a limited number of frames to the left and right. Restricting the context makes it easier to encode the position of the input - we use a I-hot encoding of the frame offset. We try introducing attention layers into TDNN architectures, and replacing LSTM layers with attention layers in TDNN+LSTM architectures. We show experiments on a number of ASR setups. We observe improvements compared to the TDNN and TDNN+LSTM baselines. Attention layers are also faster than LSTM layers in test time, since they lack recurrence. Daniel Povey, Hossein Hadian, Pegah Ghahremani, Ke Li 0018, Sanjeev Khudanpur |
ICASSP | 2 |
| 2018 | Acoustic Modeling from Frequency Domain Representations of Speech
Pegah Ghahremani, Hossein Hadian, Hang Lv 0001, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 2 |
| 2018 | End-to-end Speech Recognition Using Lattice-free MMI
Hossein Hadian, Hossein Sameti, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 1 |
| 2018 | Improving LF-MMI Using Unconstrained Supervisions for ASRabstractWe present our work on improving the numerator graph for discriminative training using the lattice-free maximum mutual information (MMI) criterion. Specifically, we propose a scheme for creating unconstrained numerator graphs by removing time constraints from the baseline numerator graphs. This leads to much smaller graphs and therefore faster preparation of training supervisions. By testing the proposed un-constrained supervisions using factorized time-delay neural network (TDNN) models, we observe 0.5% to 2.6% relative improvement over the state-of-the-art word error rates on various large-vocabulary speech recognition databases. Hossein Hadian, Daniel Povey, Hossein Sameti, Jan Trmal, Sanjeev Khudanpur |
SLT | 1 |
| 2018 | Flat-Start Single-Stage Discriminatively Trained HMM-Based Models for ASRabstractIn recent years, end-to-end approaches to automatic speech recognition have received considerable attention as they are much faster in terms of preparing resources. However, conventional multistage approaches, which rely on a pipeline of training hidden Markov models (HMM)-GMM models and tree-building steps still give the state-of-the-art results on most databases. In this study, we investigate flat-start one-stage training of neural networks using lattice-free maximum mutual information (LF-MMI) objective function with HMM for large vocabulary continuous speech recognition. We thoroughly look into different issues that arise in such a setup and propose a standalone system, which achieves word error rates (WER) comparable with that of the state-of-the-art multi-stage systems while being much faster to prepare. We propose to use full biphones to enable flat-start context-dependent (CD) modeling and show through experiments that our CD modeling approach can be almost as effective as regular tree-based CD modeling. We show that our flat-start LF-MMI setup together with this tree-free CD modeling technique achieves 10 to 25 % relative WER reduction compared to other end-to-end methods on well-known databases. The improvements are larger for smaller databases. Hossein Hadian, Hossein Sameti, Daniel Povey, Sanjeev Khudanpur |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Investigation of transfer learning for ASR using LF-MMI trained neural networksabstractIt is common in applications of ASR to have a large amount of data out-of-domain to the test data and a smaller amount of in-domain data similar to the test data. In this paper, we investigate different ways to utilize this out-of-domain data to improve ASR models based on Lattice-free MMI (LF-MMI). In particular, we experiment with multi-task training using a network with shared hidden layers; and we try various ways of adapting previously trained models to a new domain. Both types of methods are effective in reducing the WER versus in-domain models, with the jointly trained models generally giving more improvement. Pegah Ghahremani, Vimal Manohar, Hossein Hadian, Daniel Povey, Sanjeev Khudanpur |
ASRU | 3 |
| 2017 | Phone Duration Modeling for LVCSR Using Neural Networks
Hossein Hadian, Daniel Povey, Hossein Sameti, Sanjeev Khudanpur |
INTERSPEECH | 1 |
| 2015 | Telephony text-prompted speaker verification using i-vector representationabstractI-vectors have proved to be the most effective features for text-independent speaker verification in recent researches. In this article a new scheme is proposed to utilize i-vectors in text-prompted speaker verification in a simple while effective manner. In order to examine this scheme empirically, a telephony dataset of Persian month names is introduced. Experiments show that the proposed scheme reduces the EER by 31% compared to the state-of-the-art State-GMM-MAP method. Furthermore it is shown that using HMM instead of GMM for universal background modeling leads to 15% reduction in EER. Hossein Zeinali, Elaheh Kalantari, Hossein Sameti, Hossein Hadian |
ICASSP | 4 |
| 2014 | Active Learning in Noisy Conditions for Spoken Language Understanding
Hossein Hadian, Hossein Sameti |
COLING | 1 |