Shammur Absar Chowdhury

dblp:140/2718 · also S. A. Chowdhury 0001, Shammur A. Chowdhury · DBLP profile ↗
← Back
41ranked-venue papers
15as first author
23since 2021 · last 2025
0000-0002-1331-2543ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 12 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 10 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
abstract
Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects alongside Modern Standard Arabic (MSA), created using Machine Translation (MT) combined with human post-editing. We present AraDiCE, a benchmark for Arabic Dialect and Cultural Evaluation. We evaluate LLMs on dialect comprehension and generation, focusing specifically on low-resource Arabic dialects. Additionally, we introduce the first-ever fine-grained benchmark designed to evaluate cultural awareness across the Gulf, Egypt, and Levant regions, providing a novel dimension to LLM evaluation. Our findings demonstrate that while Arabic-specific models like Jais and AceGPT outperform multilingual models on dialectal tasks, significant challenges persist in dialect identification, generation, and translation. This work contributes ≈45K post-edited samples, a cultural benchmark, and highlights the importance of tailored training to improve LLM performance in capturing the nuances of diverse Arabic dialects and cultural contexts. We have released the dialectal translation models and benchmarks developed in this study (https://huggingface.co/datasets/QCRI/AraDiCE)
Basel Mousi, Nadir Durrani, Fatema Ahmad, Md. Arid Hasan, Maram Hasanain, Tameem Kabbani, Fahim Dalvi, Shammur Absar Chowdhury, Firoj Alam
COLING8
2025 SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
Firoj Alam, Md. Arid Hasan, Shammur Absar Chowdhury
INTERSPEECH3
2025 From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models
Asim Ersoy, Basel Mousi, Shammur Absar Chowdhury, Firoj Alam, Fahim Dalvi, Nadir Durrani
INTERSPEECH3
2025 SawtArabi: A Benchmark Corpus for Arabic TTS. Standard, Dialectal and Code-Switching
Vasista Sai Lodagala, Lamya Alkanhal, Daniel Izham, Shivam Mehta, Shammur Absar Chowdhury, Aqeelah Makki, Hamdy S. Hussein, Gustav Eje Henter, Ahmed Ali 0002
INTERSPEECH5
2024 Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in Arabic
abstract
This paper presents a novel Dialectal Sound and Vowelization Recovery framework, designed to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages, that extends beyond its standard orthographic sound sets.The proposed framework utilized quantized sequence of input with(out) continuous pretrained selfsupervised representation.We show the efficacy of the pipeline using limited data for Arabic, a dialect-rich language containing more than 22 major dialects.Phonetically correct transcribed speech resources for dialectal Arabic is scare.Therefore, we introduce Arab-Voice15 1 , a first of its kind, curated test set featuring 5 hours of dialectal speech across 15 Arab countries, with phonetically accurate transcriptions, including borrowed and dialectspecific sounds.We described in detail the annotation guideline along with the analysis of the dialectal confusion pairs.Our extensive evaluation includes both subjective -human perception tests and objective measures.Our empirical results, reported with three test sets, show that with only one and half hours of training data, our model improve character error rate by ≈ 7% in ArabVoice15 compared to the baseline.
Yassine El Kheir, Hamdy Mubarak, Ahmed Ali 0002, Shammur Absar Chowdhury
ACL (1)4
2024 LAraBench: Benchmarking Arabic AI with Large Language Models
abstract
Ahmed Abdelali, Hamdy Mubarak, Shammur Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Youssef Elshahawy, Ahmed Ali, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Ahmed Abdelali, Hamdy Mubarak, Shammur Absar Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Yousseif Elshahawy, Ahmed Ali 0002, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam
EACL (1)3
2024 Speech Collage: Code-Switched Audio Generation by Collaging Monolingual Corpora
abstract
Designing effective automatic speech recognition (ASR) systems for Code-Switching (CS) often depends on the availability of the transcribed CS resources. To address data scarcity, this paper introduces Speech Collage, a method that synthesizes CS data from monolingual corpora by splicing audio segments. We further improve the smoothness quality of audio generation using an overlap-add approach. We investigate the impact of generated data on speech recognition in two scenarios: using in-domain CS text and a zero-shot approach with synthesized CS text. Empirical results highlight up to 34.4% and 16.2% relative reductions in Mixed-Error Rate and Word-Error Rate for in-domain and zero-shot scenarios, respectively. Lastly, we demonstrate that CS augmentation bolsters the model’s code-switching inclination and reduces its monolingual bias.
Amir Hussein, Dorsa Zeinali, Ondrej Klejch, Matthew Wiesner, Brian Yan, Shammur Absar Chowdhury, Ahmed Ali 0002, Shinji Watanabe 0001, Sanjeev Khudanpur
ICASSP6
2024 L1-Aware Multilingual Mispronunciation Detection Framework
abstract
The phonological discrepancies between a speaker’s native (L1) and the non-native language (L2) serves as a major factor for mispronunciation. This paper introduces a novel multilingual Mispronunciation Detection and Diagnosis (MDD) architecture, L1-MultiMDD, enriched with L1-aware speech representation. An end-to-end speech encoder is trained on the input signal and its corresponding reference phoneme sequence. First, an attention mechanism is deployed to align the input audio with the reference phoneme sequence. Afterwards, the L1-L2-speech embedding are extracted from an auxiliary model, pretrained in a multi-task setup identifying L1 and L2 language, and are infused with the primary network. Finally, the L1-MultiMDD is then optimized for a unified multilingual phoneme recognition task using connectionist temporal classification (CTC) loss for the target languages: English, Arabic, and Mandarin. Our experiments demonstrate the effectiveness of the proposed L1-MultiMDD framework on both seen – L2-ARTIC, LATIC, and AraVoiceL2v2; and unseen – EpaDB and Speechocean762 datasets. The consistent gains in PER, and false rejection rate (FRR) across all target languages confirm our approach’s robustness, efficacy, and generalizability.
Yassine El Kheir, Shammur Absar Chowdhury, Ahmed Ali 0002
ICASSP2
2024 Children's Speech Recognition through Discrete Token Enhancement
Vrunda N. Sukhadia, Shammur Absar Chowdhury
INTERSPEECH2
2024 What do end-to-end speech models learn about speaker, language and channel information? A layer-wise and neuron-level analysis
abstract
Deep neural networks are inherently opaque and challenging to interpret. Unlike hand-crafted feature-based models, we struggle to comprehend the concepts learned and how they interact within these models. This understanding is crucial not only for debugging purposes but also for ensuring fairness in ethical decision-making. In our study, we conduct a post-hoc functional interpretability analysis of pretrained speech models using the probing framework (Hupkes et al., 2018). Specifically, we analyze utterance-level representations of speech models trained for various tasks such as speaker recognition and dialect identification. We conduct layer and neuron-wise analyses, probing for speaker, language, and channel properties. Our study aims to answer the following questions: (i) what information is captured within the representations? (ii) how is it represented and distributed? and (iii) can we identify a minimal subset of the network that possesses this information? Our results reveal several novel findings, including: (i) channel and gender information are distributed across the network, (ii) the information is redundantly available in neurons with respect to a task, (iii) complex properties such as dialectal information are encoded only in the task-oriented pretrained network, (iv) and is localised in the upper layers, (v) we can extract a minimal subset of neurons encoding the pre-defined property, (vi) salient neurons are sometimes shared between properties, (vii) our analysis highlights the presence of biases (for example gender) in the network. Our cross-architectural comparison indicates that: (i) the pretrained models capture speaker-invariant information, and (ii) CNN models are competitive with Transformer models in encoding various understudied properties.
Shammur Absar Chowdhury, Nadir Durrani, Ahmed Ali 0002
Comput. Speech Lang.1
2023 Multilingual Word Error Rate Estimation: E-Wer3
abstract
The success of the multilingual automatic speech recognition systems empowered many voice-driven applications. However, measuring the performance of such systems remains a major challenge, due to its dependency on manually transcribed speech data in both mono- and multilingual scenarios. In this paper, we propose a novel multilingual framework – eWER3 – jointly trained on acoustic and lexical representation to estimate word error rate. We demonstrate the effectiveness of eWER3 to (i) predict WER without using any internal states from the ASR and (ii) use the multilingual shared latent space to push the performance of the close-related languages. We show our proposed multilingual model outperforms the previous monolingual word error rate estimation method (eWER2) by an absolute 9% increase in Pearson correlation coefficient (PCC), with better overall estimation between the predicted and reference WER.
Shammur Absar Chowdhury, Ahmed Ali 0002
ICASSP1
2023 MyVoice: Arabic Speech Resource Collaboration Platform
Yousseif Elshahawy, Yassine El Kheir, Shammur Absar Chowdhury, Ahmed Ali 0002
INTERSPEECH3
2023 QVoice: Arabic Speech Pronunciation Learning Application
Yassine El Kheir, Fouad Khnaisser, Shammur Absar Chowdhury, Hamdy Mubarak, Shazia Afzal, Ahmed Ali 0002
INTERSPEECH3
2023 Emojis as anchors to detect Arabic offensive language and hate speech
abstract
Abstract We introduce a generic, language-independent method to collect a large percentage of offensive and hate tweets regardless of their topics or genres. We harness the extralinguistic information embedded in the emojis to collect a large number of offensive tweets. We apply the proposed method on Arabic tweets and compare it with English tweets—analyzing key cultural differences. We observed a constant usage of these emojis to represent offensiveness throughout different timespans on Twitter. We manually annotate and publicly release the largest Arabic dataset for offensive, fine-grained hate speech, vulgar, and violence content. Furthermore, we benchmark the dataset for detecting offensiveness and hate speech using different transformer architectures and perform in-depth linguistic analysis. We evaluate our models on external datasets—a Twitter dataset collected using a completely different method, and a multi-platform dataset containing comments from Twitter, YouTube, and Facebook, for assessing generalization capability. Competitive results on these datasets suggest that the data collected using our method capture universal characteristics of offensive language. Our findings also highlight the common words used in offensive communications, common targets for hate speech, specific patterns in violence tweets, and pinpoint common classification errors that can be attributed to limitations of NLP models. We observe that even state-of-the-art transformer models may fail to take into account culture, background, and context or understand nuances present in real-world data such as sarcasm.
Hamdy Mubarak, Sabit Hassan, Shammur Absar Chowdhury
Nat. Lang. Eng.3
2022 What can Speech and Language Tell us About the Working Alliance in Psychotherapy
abstract
We are interested in the problem of conversational analysis and its application to the health domain. Cognitive Behavioral Therapy is a structured approach in psychotherapy, allowing the therapist to help the patient to identify and modify the malicious thoughts, behavior, or actions. This cooperative effort can be evaluated using the Working Alliance Inventory Observer-rated Shortened – a 12 items inventory covering task, goal, and relationship – which has a relevant influence on therapeutic outcomes. In this work, we investigate the relation between this alliance inventory and the spoken conversations (sessions) between the patient and the psychotherapist. We have delivered eight weeks of e-therapy, collected their audio and video call sessions, and manually transcribed them. The spoken conversations have been annotated and evaluated with WAI ratings by professional therapists. We have investigated speech and language features and their association with WAI items. The feature types include turn dynamics, lexical entrainment, and conversational descriptors extracted from the speech and language signals. Our findings provide strong evidence that a subset of these features are strong indicators of working alliance. To the best of our knowledge, this is the first and a novel study to exploit speech and language for characterising working alliance
Sebastian P. Bayerl, Gabriel Roccabruna, Shammur Absar Chowdhury, Tommaso Ciulli, Morena Danieli, Korbinian Riedhammer, Giuseppe Riccardi
INTERSPEECH3
2022 ArCovidVac: Analyzing Arabic Tweets About COVID-19 Vaccination
abstract
The emergence of the COVID-19 pandemic and the first global infodemic have changed our lives in many different ways. We relied on social media to get the latest information about COVID-19 pandemic and at the same time to disseminate information. The content in social media consisted not only health related advice, plans, and informative news from policymakers, but also contains conspiracies and rumors. It became important to identify such information as soon as they are posted to make an actionable decision (e.g., debunking rumors, or taking certain measures for traveling). To address this challenge, we develop and publicly release the first largest manually annotated Arabic tweet dataset, ArCovidVac, for COVID-19 vaccination campaign, covering many countries in the Arab region. The dataset is enriched with different layers of annotation, including, (i) Informativeness more vs. less importance of the tweets); (ii) fine-grained tweet content types (e.g., advice, rumors, restriction, authenticate news/information); and (iii) stance towards vaccination (pro-vaccination, neutral, anti-vaccination). Further, we performed in-depth analysis of the data, exploring the popularity of different vaccines, trending hashtags, topics, and presence of offensiveness in the tweets. We studied the data for individual types of tweets and temporal changes in stance towards vaccine. We benchmarked the ArCovidVac dataset using transformer architectures for informativeness, content types, and stance detection.
Hamdy Mubarak, Sabit Hassan, Shammur Absar Chowdhury, Firoj Alam
LREC3
2022 Benchmarking Evaluation Metrics for Code-Switching Automatic Speech Recognition
abstract
Code-switching poses a number of challenges and opportunities for multilingual automatic speech recognition. In this paper, we focus on the question of robust and fair evaluation metrics. To that end, we develop a reference benchmark data set of code-switching speech recognition hypotheses with human judgments. We define clear guidelines for minimal editing of automatic hypotheses. We validate the guidelines using 4-way inter-annotator agreement. We evaluate a large number of metrics in terms of correlation with human judgments. The metrics we consider vary in terms of representation (orthographic, phonological, semantic), directness (intrinsic vs extrinsic), granularity (e.g. word, character), and similarity computation method. The highest correlation to human judgment is achieved using transliteration followed by text normalization. We release the first corpus for human acceptance of code-switching speech recognition results in dialectal Arabic/English conversation speech.
Injy Hamed, Amir Hussein, Oumnia Chellah, Shammur Absar Chowdhury, Hamdy Mubarak, Sunayana Sitaram, Nizar Habash, Ahmed Ali 0002
SLT4
2022 Textual Data Augmentation for Arabic-English Code-Switching Speech Recognition
abstract
The pervasiveness of intra-utterance code-switching (CS) in spoken content requires that speech recognition (ASR) systems handle mixed language. Designing a CS-ASR system has many challenges, mainly due to data scarcity, grammatical structure complexity, and domain mismatch. The most common method for addressing CS is to train an ASR system with the available transcribed CS speech, along with monolingual data. In this work, we propose a zero-shot learning methodology for CS-ASR by augmenting the monolingual data with artificially generating CS text. We based our approach on random lexical replacements and Equivalence Constraint (EC) while exploiting aligned translation pairs to generate random and grammatically valid CS content. Our empirical results show a 65.5% relative reduction in language model perplexity, and 7.7% in ASR WER on two ecologically valid CS test sets. The human evaluation of the generated text using EC suggests that more than 80% is of adequate quality.
Amir Hussein, Shammur Absar Chowdhury, Ahmed Abdelali, Najim Dehak, Ahmed Ali 0002, Sanjeev Khudanpur
SLT2
2021 QASR: QCRI Aljazeera Speech Resource A Large Scale Annotated Arabic Speech Corpus
abstract
Hamdy Mubarak, Amir Hussein, Shammur Absar Chowdhury, Ahmed Ali. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Hamdy Mubarak, Amir Hussein, Shammur Absar Chowdhury, Ahmed Ali 0002
ACL/IJCNLP (1)3
2021 Arabic Code-Switching Speech Recognition Using Monolingual Data
abstract
Code-switching in automatic speech recognition (ASR) is an important challenge due to globalization. Recent research in multilingual ASR shows potential improvement over monolingual systems. We study key issues related to multilingual modeling for ASR through a series of large-scale ASR experiments. Our innovative framework deploys a multi-graph approach in the weighted finite state transducers (WFST) framework. We compare our WFST decoding strategies with a transformer sequence to sequence system trained on the same data. Given a code-switching scenario between Arabic and English languages, our results show that the WFST decoding approaches were more suitable for the intersentential code-switching datasets. In addition, the transformer system performed better for intrasentential code-switching task. With this study, we release an artificially generated development and test sets, along with ecological code-switching test set, to benchmark the ASR performance.
Ahmed Ali 0002, Shammur Absar Chowdhury, Amir Hussein, Yasser Hifny
Interspeech2
2021 Towards One Model to Rule All: Multilingual Strategy for Dialectal Code-Switching Arabic ASR
abstract
With the advent of globalization, there is an increasing demand for multilingual automatic speech recognition (ASR), handling language and dialectal variation of spoken content. Recent studies show its efficacy over monolingual systems. In this study, we design a large multilingual end-to-end ASR using self-attention based conformer architecture. We trained the system using Arabic (Ar), English (En) and French (Fr) languages. We evaluate the system performance handling: (i) monolingual (Ar, En and Fr); (ii) multi-dialectal (Modern Standard Arabic, along with dialectal variation such as Egyptian and Moroccan); (iii) code-switching -- cross-lingual (Ar-En/Fr) and dialectal (MSA-Egyptian dialect) test cases, and compare with current state-of-the-art systems. Furthermore, we investigate the influence of different embedding/character representations including character vs word-piece; shared vs distinct input symbol per language. Our findings demonstrate the strength of such a model by outperforming state-of-the-art monolingual dialectal Arabic and code-switching Arabic ASR.
Shammur Absar Chowdhury, Amir Hussein, Ahmed Abdelali, Ahmed Ali 0002
Interspeech1
2021 Persona analytics: Analyzing the stability of online segments and content interests over time using non-negative matrix factorization
Jim Jansen, Soon-Gyo Jung, Shammur Absar Chowdhury, Joni Salminen
Expert Syst. Appl.3
2021 The ability of personas: An empirical evaluation of altering incorrect preconceptions about users
Joni Salminen, Soon-Gyo Jung, Shammur Absar Chowdhury, Dianne Ramirez Robillos, Jim Jansen
Int. J. Hum. Comput. Stud.3
2020 A Literature Review of Quantitative Persona Creation
abstract
Quantitative persona creation (QPC) has tremendous potential, as HCI researchers and practitioners can leverage user data from online analytics and digital media platforms to better understand their users and customers. However, there is a lack of a systematic overview of the QPC methods and progress made, with no standard methodology or known best practices. To address this gap, we review 49 QPC research articles from 2005 to 2019. Results indicate three stages of QPC research: Emergence, Diversification, and Sophistication. Sharing resources, such as datasets, code, and algorithms, is crucial to achieving the next stage (Maturity). For practitioners, we provide guiding questions for assessing QPC readiness in organizations.
Joni Salminen, Kathleen W. Guan, Soon-Gyo Jung, Shammur Absar Chowdhury, Jim Jansen
CHI4
2020 Personas and Analytics: A Comparative User Study of Efficiency and Effectiveness for a User Identification Task
abstract
Personas are a well-known technique in human computer interaction. However, there is a lack of rigorous empirical research evaluating personas relative to other methods. In this 34-participant experiment, we compare a persona system and an analytics system, both using identical user data, for efficiency and effectiveness for a user identification task. Results show that personas afford faster task completion than the analytics system, as well as outperforming analytics with significantly higher user identification accuracy. Qualitative analysis of think-aloud transcripts shows that personas have other benefits regarding learnability and consistency. However, the analytics system affords insights and capabilities that personas cannot due to inherent design differences. Findings support the use of personas to learn about users, empirically confirming some of the stated benefits in the literature, while also highlighting the limitations of personas that may necessitate the use of accompanying methods.
Joni Salminen, Soon-Gyo Jung, Shammur Absar Chowdhury, Sercan Sengün, Jim Jansen
CHI3
2020 What Does an End-to-End Dialect Identification Model Learn About Non-Dialectal Information?
Shammur Absar Chowdhury, Ahmed Ali 0002, Suwon Shon, James R. Glass
INTERSPEECH1
2020 Effects of Dialectal Code-Switching on Speech Modules: A Study Using Egyptian Arabic Broadcast Speech
Shammur Absar Chowdhury, Younes Samih, Mohamed Eldesouki, Ahmed Ali 0002
INTERSPEECH1
2020 A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection
abstract
Access to social media often enables users to engage in conversation with limited accountability. This allows a user to share their opinions and ideology, especially regarding public content, occasionally adopting offensive language. This may encourage hate crimes or cause mental harm to targeted individuals or groups. Hence, it is important to detect offensive comments in social media platforms. Typically, most studies focus on offensive commenting in one platform only, even though the problem of offensive language is observed across multiple platforms. Therefore, in this paper, we introduce and make publicly available a new dialectal Arabic news comment dataset, collected from multiple social media platforms, including Twitter, Facebook, and YouTube. We follow two-step crowd-annotator selection criteria for low-representative language annotation task in a crowdsourcing platform. Furthermore, we analyze the distinctive lexical content along with the use of emojis in offensive comments. We train and evaluate the classifiers using the annotated multi-platform dataset along with other publicly available data. Our results highlight the importance of multiple platform dataset for (a) cross-platform, (b) cross-domain, and (c) cross-dialect generalization of classifier performance.
Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Abdelali, Soon-Gyo Jung, Jim Jansen, Joni Salminen
LREC1
2019 Automatic classification of speech overlaps: Feature representation and algorithms
abstract
Overlapping speech is a natural and frequently occurring phenomenon in human–human conversations with an underlying purpose. Speech overlap events may be categorized as competitive and non-competitive. While the former is an attempt to grab the floor, the latter is an attempt to assist the speaker to continue the turn. The presence and distribution of these categories are indicative of the speakers’ states during the conversation. Therefore, understanding these manifestations is crucial for conversational analysis and for modeling human–machine dialogs. The goal of this study is to design computational models to classify overlapping speech segments of dyadic conversations into competitive vs. non-competitive acts using lexical and acoustic cues, as well as their surrounding context. The designed overlap representations are evaluated in both linear – Support Vector Machines (SVM) – and non-linear – feed-forward (FFNN), convolutional (CNN) and long short-term memory (LSTM) neural network – models. We experiment with lexical and acoustic representations and their combinations from both speaker channels in feature and hidden space. We observe that lexical word-embedding features significantly increase the overall F 1 -measure compared to both acoustic and bag-of-ngrams lexical representations , suggesting that lexical information can be utilized as a powerful cue for overlap classification. Our comparative study shows that the best computational architecture is an FFNN along with a combination of word embeddings and acoustic features.
Shammur Absar Chowdhury, Evgeny A. Stepanov, Morena Danieli, Giuseppe Riccardi
Comput. Speech Lang.1
2018 RNN Simulations of Grammaticality Judgments on Long-distance Dependencies
abstract
The paper explores the ability of LSTM networks trained on a language modeling task to detect linguistic structures which are ungrammatical due to extraction violations (extra arguments and subject-relative clause island violations), and considers its implications for the debate on language innatism. The results show that the current RNN model can correctly classify (un)grammatical sentences, in certain conditions, but it is sensitive to linguistic processing factors and probably ultimately unable to induce a more abstract notion of grammaticality, at least in the domain we tested.
Shammur Absar Chowdhury, Roberto Zamparelli
COLING1
2018 Depression Severity Estimation from Multiple Modalities
abstract
Depression is a major debilitating disorder which can affect people from all ages. With a continuous increase in the number of annual cases of depression, there is a need to develop automatic techniques for the detection of the presence and its severity. We explore different modalities (speech, behavioral characteristics, language and visual features extracted from face) to design and develop automatic methods for the detection of depression. In psychology literature, the eight-item Patient Health Questionnaire depression scale (PHQ-8) is well established as a tool for measuring the severity of depression. In this paper we aim to automatically predict the total sum of PHQ-8 scores from features extracted from the different modalities. We demonstrate that among the considered modalities, behavioral characteristic features extracted from speech yield the lowest MAE, outperforming the best system at the Audio/Visual Emotion Challenge (AVEC) 2017 depression sub-challenge.
Evgeny A. Stepanov, Stéphane Lathuilière, Shammur Absar Chowdhury, Arindam Ghosh 0004, Radu L. Vieriu, Nicu Sebe, Giuseppe Riccardi
HealthCom3
2017 A Deep Learning approach to modeling competitiveness in spoken conversations
abstract
The motivation behind the research on overlapping speech has always been dominated by the need to model human-machine interaction for dialog systems and conversation analysis. To have more complex insights of the interlocutors' intentions behind the interaction, we need to understand the type of overlaps. Overlapping speech signals the interlocutor's intention to grab the floor. This act could be a competitive or non-competitive act, which either signals a problem or indicates assistance in communication. In this paper, we present a Deep Learning approach to modeling competitiveness in overlapping speech using acoustic and lexical features and their combination. We compare a fully-connected feed-forward neural network to the Support Vector Machine (SVM) models on real call center human-human conversations. We have observed that feature combination with DNN (significantly) outperforms SVM models, both the individual feature baselines and the feature combination model by 4% and 2% respectively.
Shammur Absar Chowdhury, Giuseppe Riccardi
ICASSP1
2016 How Interlocutors Coordinate with each other within Emotional Segments?
abstract
In this paper, we aim to investigate the coordination of interlocutors behavior in different emotional segments. Conversational coordination between the interlocutors is the tendency of speakers to predict and adjust each other accordingly on an ongoing conversation. In order to find such a coordination, we investigated 1) lexical similarities between the speakers in each emotional segments, 2) correlation between the interlocutors using psycholinguistic features, such as linguistic styles, psychological process, personal concerns among others, and 3) relation of interlocutors turn-taking behaviors such as competitiveness. To study the degree of coordination in different emotional segments, we conducted our experiments using real dyadic conversations collected from call centers in which agent’s emotional state include empathy and customer’s emotional states include anger and frustration. Our findings suggest that the most coordination occurs between the interlocutors inside anger segments, where as, a little coordination was observed when the agent was empathic, even though an increase in the amount of non-competitive overlaps was observed. We found no significant difference between anger and frustration segment in terms of turn-taking behaviors. However, the length of pause significantly decreases in the preceding segment of anger where as it increases in the preceding segment of frustration.
Firoj Alam, Shammur Absar Chowdhury, Morena Danieli, Giuseppe Riccardi
COLING2
2016 Discourse connective detection in spoken conversations
abstract
Discourse parsing is an important task in Language Understanding with applications to human-human and human-machine communication modeling. However, most of the research has focused on written text, and parsers heavily rely on syntactic parsers that themselves have low performance on dialog data. In our work, we address the problem of analyzing the semantic relations between discourse units in human-human spoken conversations. In particular, in this paper we focus on the detection of discourse connectives which are the predicate of such relations. The discourse relations are drawn from the Penn Discourse Treebank annotation model and adapted to a domain-specific Italian human-human spoken conversations. We study the relevance of lexical and acoustic context in predicting discourse connectives. We observe that both lexical and acoustic context have mixed effect on the prediction of specific connectives. While the oracle of using lexical and acoustic contextual feature combinations is F1 = 68.53, the lexical context alone significantly outperforms the baseline by more than 10 points with F1= 64.93.
Giuseppe Riccardi, Evgeny A. Stepanov, Shammur Absar Chowdhury
ICASSP3
2016 Predicting User Satisfaction from Turn-Taking in Spoken Conversations
Shammur Absar Chowdhury, Evgeny A. Stepanov, Giuseppe Riccardi
INTERSPEECH1
2016 Transfer of Corpus-Specific Dialogue Act Annotation to ISO Standard: Is it worth it?
Shammur Absar Chowdhury, Evgeny A. Stepanov, Giuseppe Riccardi
LREC1
2015 Annotating and categorizing competition in overlap speech
abstract
Overlapping speech is a common and relevant phenomenon in human conversations, reflecting many aspects of discourse dynamics. In this paper, we focus on the pragmatic role of overlaps in turn-in-progress, where it can be categorized as competitive or non-competitive. Previous studies on these two categories have mostly relied on controlled scenarios and small datasets. In our study, we focus on call center data, with customers and operators engaged in problem-solving tasks. We propose and evaluate an annotation scheme for these two overlap categories in the context of spontaneous and in-vivo human conversations. We analyze the distinctive predictive characteristics of a very large set of high-dimensional acoustic feature. We obtained a significant improvement in classification results as well as significant reduction in the feature set size.
Shammur Absar Chowdhury, Morena Danieli, Giuseppe Riccardi
ICASSP1
2015 Selection and aggregation techniques for crowdsourced semantic annotation task
Shammur Absar Chowdhury, Marcos Calvo, Arindam Ghosh 0004, Evgeny A. Stepanov, Ali Orkan Bayer, Giuseppe Riccardi, Fernando García 0001, Emilio Sanchis Arnal
INTERSPEECH1
2015 The role of speakers and context in classifying competition in overlapping speech
abstract
Overlapping speech is one of the most frequently occurring events in the course of human-human conversations.Understanding the dynamics of overlapping speech is crucial for conversational analysis and for modeling human-machine dialog.Overlapping speech may signal the speaker's intention to grab the floor with a competitive vs non-competitive act.In this paper, we study the role of speakers, whether they initiate (overlapper) or not (overlappee) the overlap, and the context of the event.The speech overlap may be explained and predicted by the dialog context, the linguistic or acoustic descriptors.Our goal is to understand whether the competitiveness of the overlap is best predicted by the overlapper, the overlappee, the context or by their combinations.For each overlap and its context we have extracted acoustic, linguistic, and psycholinguistic features and combined decisions from the best classification models.The evaluation of the classifier has been carried out over call center human-human conversations.The results show that the complete knowledge of speakers' role and context highly contribute to the classification results when using acoustic and psycholinguistic features.Our findings also suggest that the lexical selections of the overlapper are good indicators of speaker's competitive or non-competitive intentions.
Shammur Absar Chowdhury, Morena Danieli, Giuseppe Riccardi
INTERSPEECH1
2014 Cross-language transfer of semantic annotation via targeted crowdsourcing
Shammur Absar Chowdhury, Arindam Ghosh 0004, Evgeny A. Stepanov, Ali Orkan Bayer, Giuseppe Riccardi, Ioannis Klasinas
INTERSPEECH1
2013 Motivational feedback in crowdsourcing: a case study in speech transcription
Giuseppe Riccardi, Arindam Ghosh 0004, Shammur Absar Chowdhury, Ali Orkan Bayer
INTERSPEECH3