Badr Abdullah

dblp:186/2763 · also Badr M. Abdullah · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0002-1281-148XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3
YearPublicationVenuePosition
2025 It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems
abstract
Iuliia Zaitova, Badr M. Abdullah, Wei Xue, Dietrich Klakow, Bernd Möbius, Tania Avgustinova. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Iuliia Zaitova, Badr Abdullah, Dietrich Klakow, Bernd Möbius, Tania Avgustinova
ACL (1)2
2025 Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
Badr Abdullah, Matthew Baas, Bernd Möbius, Dietrich Klakow
INTERSPEECH1
2024 Self-Supervised Adaptive Pre-Training of Multilingual Speech Models for Language and Dialect Identification
abstract
Transformer-based, pre-trained speech models have shown striking performance when fine-tuned on various downstream tasks such as automatic speech recognition and spoken language identification (SLID). However, the problem of domain mismatch remains a challenge in this area, where the domain of the pre-training data might differ from that of the downstream labeled data used for fine-tuning. In multilingual tasks such as SLID, the pre-trained speech model may not support all the languages in the downstream task. To address this challenge, we propose self-supervised adaptive pre-training (SAPT) to adapt the pre-trained model to the target domain and languages of the downstream task. We apply SAPT to the XLSR-128 model and investigate the effectiveness of this approach for the SLID task. First, we demonstrate that SAPT improves XLSR performance on the FLEURS benchmark with substantial gains up to 40.1% for under-represented languages. Second, we apply SAPT on four different datasets in a few-shot learning setting, showing that our approach improves the sample efficiency of XLSR during fine-tuning. Our experiments provide strong empirical evidence that continual adaptation via self-supervision improves downstream performance for multilingual speech models.
Mohammed Maqsood Shaik, Dietrich Klakow, Badr Abdullah
ICASSP3
2024 Wave to Interlingua: Analyzing Representations of Multilingual Speech Transformers for Spoken Language Translation
Badr Abdullah, Mohammed Maqsood Shaik, Dietrich Klakow
INTERSPEECH1
2024 On the Encoding of Gender in Transformer-based ASR Representations
Aravind Krishnan, Badr Abdullah, Dietrich Klakow
INTERSPEECH2
2024 Field studies of the Artificial Intelligence model for defining indoor thermal comfort to acknowledge the adaptive aspect
Kanisius Karyono, Badr Abdullah, Alison J. Cotgrave, Ana Armada Bras, Jeff Cullen
Eng. Appl. Artif. Intell.2
2023 Ending the Blind Flight: Analyzing the Impact of Acoustic and Lexical Factors on WAV2VEC 2.0 in Air-Traffic Control
abstract
Transformer neural networks have shown remarkable success on standard automatic speech recognition (ASR) benchmarks. However, they are known to be less robust against domain mismatch, particularly with air traffic control (ATC) speech data. In the ATC domain, transformer-based ASR systems do usually not transfer across different datasets. The reasons for poor transferability across ATC datasets remain unclear. Our study investigates the influence of acoustic variability and lexical differences on the ASR performance across various ATC datasets. By fine-tuning and evaluating wav2vec 2.0 on synthetic ATC datasets, we examine the effect of acoustic variability on the model performance. Furthermore, we assess the effect of lexical differences by correlating language model perplexity with performance. Our findings reveal that a combination of acoustic and lexical mismatch causes the bad inter-dataset transferability and give insights on how to improve future ASR models for ATC.
Alexander Blatt, Badr Abdullah, Dietrich Klakow
ASRU2
2023 An Information-Theoretic Analysis of Self-supervised Discrete Representations of Speech
Badr Abdullah, Mohammed Maqsood Shaik, Bernd Möbius, Dietrich Klakow
INTERSPEECH1
2022 Integrating Form and Meaning: A Multi-Task Learning Model for Acoustic Word Embeddings
Badr Abdullah, Bernd Möbius, Dietrich Klakow
INTERSPEECH1
2021 Do Acoustic Word Embeddings Capture Phonological Similarity? An Empirical Study
abstract
Several variants of deep neural networks have been successfully employed for building parametric models that project variable-duration spoken word segments onto fixed-size vector representations, or acoustic word embeddings (AWEs). However, it remains unclear to what degree we can rely on the distance in the emerging AWE space as an estimate of word-form similarity. In this paper, we ask: does the distance in the acoustic embedding space correlate with phonological dissimilarity? To answer this question, we empirically investigate the performance of supervised approaches for AWEs with different neural architectures and learning objectives. We train AWE models in controlled settings for two languages (German and Czech) and evaluate the embeddings on two tasks: word discrimination and phonological similarity. Our experiments show that (1) the distance in the embedding space in the best cases only moderately correlates with phonological distance, and (2) improving the performance on the word discrimination task does not necessarily yield models that better reflect word phonological similarity. Our findings highlight the necessity to rethink the current intrinsic evaluations for AWEs.
Badr Abdullah, Marius Mosbach, Iuliia Zaitova, Bernd Möbius, Dietrich Klakow
Interspeech1
2020 A Closer Look at Linguistic Knowledge in Masked Language Models: The Case of Relative Clauses in American English
abstract
Transformer-based language models achieve high performance on various tasks, but we still lack understanding of the kind of linguistic knowledge they learn and rely on.We evaluate three models (BERT, RoBERTa, and ALBERT), testing their grammatical and semantic knowledge by sentence-level probing, diagnostic cases, and masked prediction tasks.We focus on relative clauses (in American English) as a complex phenomenon needing contextual information and antecedent identification to be resolved.Based on a naturalistic dataset, probing shows that all three models indeed capture linguistic knowledge about grammaticality, achieving high performance.Evaluation on diagnostic cases and masked prediction tasks considering fine-grained linguistic knowledge, however, shows pronounced model-specific weaknesses especially on semantic knowledge, strongly impacting models' performance.Our results highlight the importance of (a) model comparison in evaluation task and (b) building up claims of model performance and the linguistic knowledge they capture beyond purely probing-based evaluations.
Marius Mosbach, Stefania Degaetano-Ortlieb, Marie-Pauline Krielke, Badr Abdullah, Dietrich Klakow
COLING4
2020 Cross-Domain Adaptation of Spoken Language Identification for Related Languages: The Curious Case of Slavic Languages
abstract
State-of-the-art spoken language identification (LID) systems, which are based on end-to-end deep neural networks, have shown remarkable success not only in discriminating between distant languages but also between closely-related languages or even different spoken varieties of the same language. However, it is still unclear to what extent neural LID models generalize to speech samples with different acoustic conditions due to domain shift. In this paper, we present a set of experiments to investigate the impact of domain mismatch on the performance of neural LID systems for a subset of six Slavic languages across two domains (read speech and radio broadcast) and examine two low-level signal descriptors (spectral and cepstral features) for this task. Our experiments show that (1) out-of-domain speech samples severely hinder the performance of neural LID models, and (2) while both spectral and cepstral features show comparable performance within-domain, spectral features show more robustness under domain mismatch. Moreover, we apply unsupervised domain adaptation to minimize the discrepancy between the two domains in our study. We achieve relative accuracy improvements that range from 9% to 77% depending on the diversity of acoustic conditions in the source domain.
Badr Abdullah, Tania Avgustinova, Bernd Möbius, Dietrich Klakow
INTERSPEECH1
2019 A Smart Adaptive Lighting System for a Multifunctional Room
abstract
Young professionals and millennials who live alone or are living in small groups and seek practicality, trigger the trend of smaller, modular and micro houses and apartments which are faster and cheaper to build. Multifunctional or flexible room is one of the important parts of the home. This particular room needs well-designed lighting for comfort. It should give an adequate illuminance for every activity and even pattern of light. This paper presents the factors for developing the smart adaptive lighting system which can provide lighting comfort for the occupants. The simulation is being done in 5 scenarios in the LJMU BRE 2010 house model using DIALux Software with the dimmable type of LED independent luminaire. The proposed system structure uses a wireless sensor network (WSN) and big data processing as the main components. The design employs an Artificial Intelligence (AI) sub-system which has the capability to predict and adaptively regulate the illumination level based on the occupant needs or routine. The simulation shows that this system is able to give even lighting pattern for luminance values 200, 250, 300, 500, and 750 lux which are needed by the occupants. With the possibility of user-defined lighting values, this system can be developed to accommodate the needs of special groups of occupants such as the elder or disabled groups.
Kanisius Karyono, Badr Abdullah, Alison J. Cotgrave, Ana Armada Bras
DeSE2
2019 Industry 4.0 LabVIEW Based Industrial Condition Monitoring System for Industrial IoT System
abstract
As a result of a substantial shift in focus towards a more digital industry, multiple sectors of industry are now realising the potential of Industry 4.0 and Internet of Things (IoT) technology. The manufacturing industry in particular is subject to unexpected machine downtime from component wear over an extended period. With Industrial IoT (IIoT) technology implemented, there is the potential for gathering large quantities of data, which can be used for preventative maintenance. This research article addresses some of the technological requirements for developing an IoT industrial condition monitoring network, whose composition makes use of wireless devices along with conventional wired methods to enable a series of data capture and control operations in amongst a network of nodes. To provide a platform to host these operations, the industry standard fieldbus protocol Modbus TCP was used in conjunction with the LabVIEW development environment, where a bespoke graphical user interface was developed to provide control and a visual representation of the data collected. In addition, one of the nodes acted as the output for hardware displays, which in turn correlated the alarm status of the user interface. By using industry standard communication protocols, it was also possible to enable connectivity between real industry hardware, further extending the capabilities of the system.
Harvey William Picot, Muhammad Ateeq, Badr Abdullah, Jeff Cullen
DeSE3
2019 A Roadmap Towards the Smart Factory
abstract
Industry 4.0 is the transformation of industrial manufacturing through digitisation and the use of different emerging technological advancement, when coupled together forms the smart factory. However, the roadmap of adoption is a journey rather than an absolute solution. The objectives of this paper are to give general insights and a roadmap towards the smart factory. A six-gear roadmap concept is proposed and discussed together with different challenges and practical ways of overcoming them. The significance of this paper can serve as a steppingstone for a detailed strategic roadmap for a successful implementation and transformation into a smart factory.
Amr T. Sufian, Badr Abdullah, Muhammad Ateeq, Roderick Wah, David Clements
DeSE2
2018 Dynamic Extension of ASR Lexicon Using Wikipedia Data
abstract
Despite recent progress in developing Large Vocabulary Continuous Speech Recognition Systems (LVCSR), these systems suffer from-Of-Vocabulary words (OOV). In many cases, the OOV words are Proper Nouns (PNs). The correct recognition of PNs is essential for broadcast news, audio indexing, etc. In this article, we address the problem of OOV PN retrieval in the framework of broadcast news LVCSR. We focused on dynamic (document dependent) extension of LVCSR lexicon. To retrieve relevant OOV PNs, we propose to use a very large multipurpose text corpus: Wikipedia. This corpus contains a huge number of PNs. These PNs are grouped in semantically similar classes using word embedding. We use a two-step approach: first, we select OOV PN pertinent classes with a multi-class Deep Neural Network (DNN). Secondly, we rank the OOVs of the selected classes. The experiments on French broadcast news show that the Bi-GRU model outperforms other studied models. Speech recognition experiments demonstrate the effectiveness of the proposed methodology.
Badr Abdullah, Irina Illina, Dominique Fohr
SLT1