Abul Azad

dblp:258/6842 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
0since 2021 · last 2020
0000-0001-9700-5964ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
speech coding
0.412020
Robust Speech Filter and Voice Encoder Parameter Estimation Using the Phase-Phase Correlator · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Audio and music processing
speech enhancement
0.112020
Robust Speech Filter and Voice Encoder Parameter Estimation Using the Phase-Phase Correlator · IEEE ACM Trans. Audio Speech Lang. Process. 2020

Methods — techniques the papers use, named apart from their topics

statistical outlier test · 0.4robust mahalanobis distance · 0.4phase-phase correlator · 0.4
YearPublicationVenuePosition
2020 Robust Speech Filter and Voice Encoder Parameter Estimation Using the Phase-Phase Correlator
abstract
In recent years, linear prediction voice encoders have become very efficient in terms of computing execution time and channel bandwidth usage while providing, in the absence of impulsive noise, natural sounding synthetic speech signals. This good performance has been achieved via the use of a maximum likelihood parameter estimation of an auto-regressive model of order ten that best fits the speech signal under the assumption that the signal and the noise are Gaussian stochastic processes. However, this method breaks down in the presence of impulse noise, which is common in practice, resulting in harsh or non-intelligible audio signals. In this paper, we propose a robust estimator of correlation, the Phase-Phase correlator that is able to cope with impulsive noise. Utilizing this correlator, we develop a Robust Mixed Excitation Linear Prediction encoder that provides improved audio quality for voiced, unvoiced, and transition speech segments. This is achieved by applying a statistical test to robust Mahalanobis distances for identifying the outliers in the corrupted speech signal, which are then replaced with filtered signals. Simulation results reveal that the proposed estimator of correlator outperforms in variance, bias, and breakdown point compared to three other robust approaches based on the arcsin law, the polarity coincidence correlator, and the median-of-ratio estimator without sacrificing the encoder bandwidth efficiency and the compression gain while remaining compatible with real-time applications. Furthermore, in the presence of impulsive noise, the proposed speech encoder speech subjective quality outperforms the state-of-the-art in terms of mean opinion score.
Abul Azad, Lamine Mili
IEEE ACM Trans. Audio Speech Lang. Process.1