VLDB 2026 Research / reviewers in the wild / expert
Anna Esposito
dblp:66/396
· DBLP profile ↗
57ranked-venue papers
17as first author
22since 2021 · last 2025
0000-0002-7268-1795ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 15 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 8 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Punctual or Continuous? Analyzing Depression Traces in Language and Paralanguage with Multiple Instance LearningabstractThe key-question addressed in this article is whether the traces of depression in speech are punctual (the pathology manifests itself at specific points in time) or continuous (the pathology manifests itself at every moment in time), where the expression “speech” refers to both speech signals and their transcriptions. For this reason, this work compares the performances of different approaches (both unimodal or multimodal) based either on the assumption that the traces are punctual or on the other one. In this way, it is possible to test which of the alternative assumptions is more realistic. The experiments were performed over a publicly available dataset (the Androids Corpus) and the results include F1 Scores up to 93.1%, among the best reported in literature for the Corpus. Furthermore, the results suggest that depression traces are punctual, but they appear so frequently that approaches based on the assumption of continuous traces still perform well. The conclusions discuss the implications of such observation. Rawan Alsarrani, Anna Esposito, Alessandro Vinciarelli |
ICMI | 2 |
| 2025 | Frontal Alpha Asymmetry as an Index of Willingness to Interact with Virtual Agents in Users with Depressive SymptomsabstractSocially engaging interactive systems, like virtual agents, can be employed as "companions" or "therapists", promoting wellbeing and mental health. To favor the actual usage of this kind of technologies, it is crucial to identify the features influencing users’ perception and acceptance toward them. Traditionally, user acceptance is assessed through self-report questionnaires which, however, do not shed light on the affective and motivational state of users during the interaction. These important aspects could be explored through EEG signal analysis, which would provide an objective measure of users’ preferences and usage intentions. This work investigates the relationship between users’ willingness to interact with virtual agents and frontal alpha activity, trying to provide useful insights on the affective and motivational states of users, with and without depressive symptoms, when interacting with happy, neutral and sad virtual agents. Rosa Milo, Terry Amorese, Marialucia Cuciniello, Antonio Perna, Gennaro Cordasco, Anna Esposito |
IJCNN | 6 |
| 2025 | Using AI explainable models and handwriting/drawing tasks for psychological well-beingabstractThis study addresses the increasing threat to Psychological Well-Being (PWB) posed by Depression, Anxiety, and Stress conditions. Machine learning methods have shown promising results for several psychological conditions. However, the lack of transparency in existing models impedes practical application. The study aims to develop explainable machine learning models for depression, anxiety and stress prediction, focusing on features extracted from tasks involving handwriting and drawing. Two hundred patients completed the Depression, Anxiety, and Stress Scale (DASS-21) and performed seven tasks related to handwriting and drawing. Extracted features, encompassing pressure, stroke pattern, time, space, and pen inclination, were used to train the explainable-by-design Entropy-based Logic Explained Network (e-LEN) model, employing first-order logic rules for explanation. Performance comparison was performed with XGBoost, enhanced by the SHAP explanation method. The trained models achieved notable accuracy in predicting depression (0.749 ±0.089), anxiety (0.721 ±0.088), and stress (0.761 ±0.086) through 10-fold cross-validation (repeated 20 times). The e-LEN model’s logic rules facilitated clinical validation, uncovering correlations with existing clinical literature. While performance remained consistent for depression and anxiety on an independent test dataset, a slight degradation was observed for stress prediction in the test task. Francesco Prinzi, Pietro Barbiero, Claudia Greco, Terry Amorese, Gennaro Cordasco, Pietro Liò, Salvatore Vitabile, Anna Esposito |
Inf. Syst. | 8 |
| 2025 | Exploring Emotion Expression Recognition in Older Adults Interacting With a Virtual CoachabstractThe EMPATHIC project aimed to design an emotionally expressive virtual coach capable of engaging healthy seniors to improve well-being and promote independent aging. In particular, the system's human sensing capabilities allow for the perception of emotional states to provide a personalized experience. This paper outlines the development of the emotion expression recognition module of the virtual coach, encompassing data collection, annotation design, and a first methodological approach, all tailored to the project requirements. With the latter, we investigate the role of various modalities, individually and combined, for discrete emotion expression recognition in this context: speech from audio, and facial expressions, gaze, and head dynamics from video. The collected corpus includes users from Spain, France, and Norway, and was annotated separately for the audio and video channels with distinct emotional labels, allowing for a performance comparison across cultures and label types. Results confirm the informative power of the modalities studied for the emotional categories considered, with multimodal methods generally outperforming others (around 68% accuracy with audio labels and 72-74% with video labels). The findings are expected to contribute to the limited literature on emotion recognition applied to older adults in conversational human-machine interaction, and guide the development of future systems. Cristina Palmero, Mikel de Velasco-Vázquez, Mohamed Amine Hmani, Aymen Mtibaa, Leila Ben Letaifa, Pau Buch-Cardona, Raquel Justo, Terry Amorese, Eduardo Gonzalez-Fraile, Begoña Fernández-Ruanova, Jofre Tenorio-Laranga, Anna Torp Johansen, Micaela Rodrigues da Silva, L. J. Martinussen, Maria Stylianou Korsnes, Gennaro Cordasco, Anna Esposito, Mounim A. El-Yacoubi, Dijana Petrovska-Delacrétaz, M. Inés Torres, Sergio Escalera |
IEEE Trans. Affect. Comput. | 17 |
| 2024 | Cross-Data Multilevel Attention for Depression Detection: Analyzing the Interplay Between Read and Spontaneous SpeechabstractThis work proposes a novel Cross-Data Multilevel Attention (CDMA) approach for multi-type speech-based depression detection, encompassing both read and spontaneous speech. The main novelty lies in analyzing the unique and common representations of the two types of speech and integrating them into a unified end-to-end framework with novel Intra-Type Multi-Local Attention (IT-MLA) and Cross-Type Global Attention (CT-GA) mechanisms. In particular, IT-MLA highlights depression-relevant information unique in either read or spontaneous speech via intra-modal attention-aware interactions. Furthermore, CT-GA further emphasises the depression-relevant common information in both read and spontaneous speech, with each type being guided by the other. These multiple enhanced representations are aggregated to produce the final predictions. Experiments conducted on a publicly available corpus of 104 speakers (including 52 diagnosed with depression by professional psychiatrists) demonstrate that the proposed CDMA achieves an F1 score of up to 92.5%, the highest performance recorded on this dataset. Fuxiang Tao, Xuri Ge, Anna Esposito, Alessandro Vinciarelli |
BIBM | 4 |
| 2024 | Assessing Privacy Risks of Attribute Inference Attacks Against Speech-Based Depression Detection SystemabstractMany AI applications now attempt to infer users’ mental health conditions, such as depression, from their speech data. In addition to the spoken words, the speech audio contains information about speaker’s identity and demographic attributes, exposing users to serious privacy risks. Previous efforts have primarily focused on developing deep models that preserve privacy; however, there have been few attempts to systematically assess and quantify privacy risks in such systems. We present the first framework for systematically assessing privacy risks in a multimodal (audio-lexical) depression detection system particularly looking at attribute inference attacks. Unlike past works that considered only white-box gender inference attacks against unimodal systems, our framework designs novel white-box and black-box attacks across multiple modalities against three protected speaker attributes: gender, age and education level. We present extensive results on a large, clinically validated dataset, demonstrating critical vulnerability of depression detection systems, where an adversary can infer speaker attributes with 59% - 68% accuracy even for inputs as short as 10 seconds of speech. Our results offer insights and guidelines to inform the development and benchmarking of privacy-preserving models for speech-based depression detection systems. Our code and data are available at: https://github.com/apr-aia/privacy_risks Basmah Alsenani, Anna Esposito, Alessandro Vinciarelli, Tanaya Guha |
ECAI | 2 |
| 2024 | Cultural Differences in the Assessment of Synthetic VoicesabstractThis research involved 88 young adults aged between 20 years and 35 years from two different countries, Spain and Italy. This work aims to explore preferences of the two groups toward synthetic voices, created for the experiment with variations in gender and quality for each language. The Spanish group was asked to evaluate the two high-quality voices of Elena and Pablo and the two low-quality voices of Maria and Juan while the Italian group was asked to assess the high-quality voices of Giulia and Antonio and the low-quality voices of Clara and Edoardo. The shortened and digitized version of the Virtual Agent Voice Acceptance Questionnaire (VAVAQ) was administered, respectively, in the Spanish or Italian version on the basis of the referring group to collect participants’ preferences. Due to the pandemic situation, participants were mainly contacted via email. Each participant was provided with a specific link. Outcomes revealed that Spanish and Italian young adults showed a greater appreciation toward the high-quality female voice compared to the other proposed voices. Regarding participants’ cross-cultural differences, Italian participants seem to judge the voices as more emotionally engaging than the Spanish participants whereas Spanish participants consider the audited voices as more natural and expressive than the Italian participants. Marialucia Cuciniello, Terry Amorese, Claudia Greco, Zoraida Callejas Carrión, Carl Vogel, Gennaro Cordasco, Anna Esposito |
Int. J. Neural Syst. | 7 |
| 2024 | Discriminative Power of Handwriting and Drawing Features in DepressionabstractThis study contributes knowledge on the detection of depression through handwriting/drawing features, to identify quantitative and noninvasive indicators of the disorder for implementing algorithms for its automatic detection. For this purpose, an original online approach was adopted to provide a dynamic evaluation of handwriting/drawing performance of healthy participants with no history of any psychiatric disorders ([Formula: see text]), and patients with a clinical diagnosis of depression ([Formula: see text]). Both groups were asked to complete seven tasks requiring either the writing or drawing on a paper while five handwriting/drawing features’ categories (i.e. pressure on the paper, time, ductus, space among characters, and pen inclination) were recorded by using a digitalized tablet. The collected records were statistically analyzed. Results showed that, except for pressure, all the considered features, successfully discriminate between depressed and nondepressed subjects. In addition, it was observed that depression affects different writing/drawing functionalities. These findings suggest the adoption of writing/drawing tasks in the clinical practice as tools to support the current depression detection methods. This would have important repercussions on reducing the diagnostic times and treatment formulation. Claudia Greco, Gennaro Raimo, Terry Amorese, Marialucia Cuciniello, Gavin McConvey, Gennaro Cordasco, Marcos Faúndez-Zanuy, Alessandro Vinciarelli, Zoraida Callejas Carrión, Anna Esposito |
Int. J. Neural Syst. | 10 |
| 2024 | HUM-CARD: A human crowded annotated real datasetabstractThe growth of data-driven approaches typical of Machine Learning leads to an ever-increasing need for large quantities of labeled data. Unfortunately, these attributions are often made automatically and/or crudely, thus destroying the very concept of “ground truth” they are supposed to represent. To address this problem, we introduce HUM-CARD, a dataset of human trajectories in crowded contexts manually annotated by nine experts in engineering and psychology, totaling approximately 5000 hours. Our multidisciplinary labeling process has enabled the creation of a well-structured ontology, accounting for both individual and contextual factors influencing human movement dynamics in shared environments. Preliminary and descriptive analyzes are presented, highlighting the potential benefits of this dataset and its methodology in various research challenges. Giovanni Di Gennaro, Claudia Greco, Amedeo Buonanno, Marialucia Cuciniello, Terry Amorese, Maria Santina Ler, Gennaro Cordasco, Francesco Palmieri 0001, Anna Esposito |
Inf. Syst. | 9 |
| 2024 | On the effects of obfuscating speaker attributes in privacy-aware depression detection
Nujud Aloshban, Anna Esposito, Alessandro Vinciarelli, Tanaya Guha |
Pattern Recognit. Lett. | 2 |
| 2023 | Multi-Local Attention for Speech-Based Depression DetectionabstractThis article shows that an attention mechanism, the Multi-Local Attention, can improve a depression detection approach based on Long Short-Term Memory Networks. Besides leading to higher performance metrics (e.g., Accuracy and F1 Score), Multi-Local Attention improves two other aspects of the approach, both important from an application point of view. The first is the effectiveness of a confidence score associated to the detection outcome at identifying speakers more likely to be classified correctly. The second is the amount of speaking time needed to classify a speaker as depressed or non-depressed. The experiments were performed over read speech and involved 109 participants (including 55 diagnosed with depression by professional psychiatrists). The results show accuracies up to 88.0% (F1 Score 88.0%). Fuxiang Tao, Xuri Ge, Anna Esposito, Alessandro Vinciarelli |
ICASSP | 4 |
| 2023 | The Androids Corpus: A New Publicly Available Benchmark for Speech Based Depression Detection
Fuxiang Tao, Anna Esposito, Alessandro Vinciarelli |
INTERSPEECH | 2 |
| 2022 | Thin Slices of Depression: Improving Depression Detection Performance Through Data SegmentationabstractThe computing community is making major efforts towards automatic detection of depression, a serious pathology that affects roughly 4.4% of the world’s population. One of the main difficulties is the collection of data aimed at training models capable to learn differences between depressed and non-depressed people. In fact, data collection in the depression domain requires the respect of rigorous ethical constraints that, inevitably, limit the size of the corpora that can be collected. This article proposes to address the problem by using the thin slices theory, i.e., the possibility to detect the inner state of an individual (depression in this case) through very short samples of behavior. In particular, the article shows that the performance of data-driven models can be improved by segmenting the data at disposition into thin slices and then training data-driven models over them. This increases the amount of samples at disposition and allows a relative F1 Score improvement by up to 16.2%. Rawan Alsarrani, Anna Esposito, Alessandro Vinciarelli |
ICASSP | 2 |
| 2022 | Android Robots vs Virtual Agents: which system differently aged users prefer?abstractThe growing presence of robots in our daily life brings out the need to develop systems that are ever more user-friendly, considering users' needs and preferences. This is necessary in particular when robots are developed to be introduced into welfare settings. For this reason, a study is proposed with the aim to investigate differently aged (young, middle-aged, and seniors) potential users' assessment of male android robots as opposed to male virtual agents, in order to compare interactive systems characterized by different levels of embodiment. 180 participants joined the experiment, which consisted of watching video clips depicting android robots and virtual agents, and subsequently fulfilling the RAQ (Robot Acceptance Questionnaire) and the VAAQ (Virtual Agent Acceptance Questionnaire). Results highlighted substantial differences in robots and agents' assessment, differences which seem to be affected by participants' age, as well. Claudia Greco, Terry Amorese, Marialucia Cuciniello, Gennaro Cordasco, Anna Esposito |
RO-MAN | 5 |
| 2022 | Age and gender effects on the human's ability to decode posed and naturalistic emotional faces
Anna Esposito, Terry Amorese, Marialucia Cuciniello, Maria Teresa Riviello, Gennaro Cordasco |
Pattern Anal. Appl. | 1 |
| 2022 | Guest Editorial: Special issue on computer vision and machine learning for healthcare applications
Cristina Palmero, M. Inés Torres, Anna Esposito, Sergio Escalera |
Pattern Anal. Appl. | 3 |
| 2022 | Assessing Facial Symmetry and Attractiveness using Augmented RealityabstractAbstract Facial symmetry is a key component in quantifying the perception of beauty. In this paper, we propose a set of facial features computed from facial landmarks which can be extracted at a low computational cost. We quantitatively evaluated the proposed features for predicting perceived attractiveness from human portraits on four benchmark datasets (SCUT-FBP, SCUT-FBP5500, FACES and Chicago Face Database). Experimental results showed that the performance of the proposed features is comparable to those extracted from a set with much denser facial landmarks. The computation of facial features was also implemented as an augmented reality (AR) app developed on Android OS. The app overlays four types of measurements and guidelines over a live video stream, while the facial measurements are computed from the tracked facial landmarks at run time. The developed app can be used to assist plastic surgeons in assessing facial symmetry when planning reconstructive facial surgeries. Wei Wei 0006, Edmond S. L. Ho, Kevin D. McCay, Robertas Damasevicius, Rytis Maskeliunas, Anna Esposito |
Pattern Anal. Appl. | 6 |
| 2022 | What an "Ehm" Leaks About You: Mapping Fillers into Personality Traits with Quantum Evolutionary Feature Selection AlgorithmsabstractThis work shows that fillers - short utterances like “ehm” and “uhm” - allow one to predict whether someone is above median along the Big-Five personality traits. The experiments have been performed over a corpus of 2,988 fillers uttered by 120 different speakers in spontaneous conversations. The results show that the prediction accuracies range between 74 and 82 percent depending on the particular trait. The proposed approach includes a feature selection step - based on Quantum Evolutionary Algorithms - that has been used to detect the personality markers, i.e., the subset of the features that better account for the prediction outcomes and, indirectly, for the personality of the speakers. The results show that only a relatively few features tend to be consistently selected, thus acting as reliable personality markers. Mohammad Tayarani, Anna Esposito, Alessandro Vinciarelli |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Synthetic vs Human Emotional Faces: What Changes in Humans' Decoding AccuracyabstractConsidered the increasing use of assistive technologies in the shape of virtual agents, it is necessary to investigate those factors which characterize and affect the interaction between the user and the agent, among these emerges the way in which people interpret and decode synthetic emotions, i.e., emotional expressions conveyed by virtual agents. For these reasons, an article is proposed, which involved 278 participants split in differently aged groups (young, middle-aged, and elders). Within each age group, some participants were administered a “naturalistic decoding task,” a recognition task of human emotional faces, while others were administered a “synthetic decoding task” namely emotional expressions conveyed by virtual agents. Participants were required to label pictures of female and male humans or virtual agents of different ages (young, middle-aged, and old) displaying static expressions of disgust, anger, sadness, fear, happiness, surprise, and neutrality. Results showed that young participants showed better recognition performances (compared to older groups) of anger, sadness, and neutrality, while female participants showed better recognition performances (compared to males) of sadness, fear, and neutrality; sadness and fear were better recognized when conveyed by real human faces, while happiness, surprise, and neutrality were better recognized when represented by virtual agents. Young faces were better decoded when expressing anger and surprise, middle-aged faces were better decoded when expressing sadness, fear, and happiness, while old faces were better decoded in the case of disgust; on average, female faces where better decoded compared to male ones. Terry Amorese, Marialucia Cuciniello, Alessandro Vinciarelli, Gennaro Cordasco, Anna Esposito |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2021 | The EMPATHIC Virtual Coach: a demoabstractThe main objective of the EMPATHIC project has been the design and development of a virtual coach to engage the healthy-senior user and to enhance well-being through awareness of personal status. The EMPATHIC approach addresses this objective through multimodal interactions supported by the GROW coaching model. The paper summarizes the main components of the EMPATHIC Virtual Coach (EMPATHIC-VC) and introduces a demonstration of the coaching sessions in selected scenarios. Javier Mikel Olaso, Alain Vázquez, Leila Ben Letaifa, Mikel de Velasco-Vázquez, Aymen Mtibaa, Mohamed Amine Hmani, Dijana Petrovska-Delacrétaz, Gérard Chollet, César Montenegro, Asier López-Zorrilla, Raquel Justo, Roberto Santana 0001, Jofre Tenorio-Laranga, Eduardo Gonzalez-Fraile, Begoña Fernández-Ruanova, Gennaro Cordasco, Anna Esposito, Kristin Beck Gjellesvik, Anna Torp Johansen, Maria Stylianou Korsnes, Colin Pickard, Cornelius Glackin, Gary Cahalane, Pau Buch-Cardona, Cristina Palmero, Sergio Escalera, Olga Gordeeva, Olivier Deroo, Anaïs Fernández, Daria Kyslitska, José Antonio Lozano 0001, M. Inés Torres, Stephan Schlögl |
ICMI | 17 |
| 2021 | A Lightweight Machine Learning Approach to Detect Depression from Speech AnalysisabstractThe growing number of people suffering from depression makes it increasingly necessary to find new approaches able to support medical experts in its diagnosis. The early detection of depressive symptoms is crucial in limiting the co-occurrence of associated behavioural disorders such as psycho-motor retardation symptoms and social withdrawal. Therefore, automatic detection systems represent promising solutions not only for supporting the early diagnosis of the disease but also for monitoring patient’s health status, thus improving both the quality of the care process and life quality of patients. At the light of these considerations, this paper proposes an automatic system exploiting a machine learning algorithm, to distinguish among depressed and healthy subjects through the analysis of selected acoustic features extracted from spontaneous speech narratives produced by healthy and depressed subjects. The proposed system achieves a classification accuracy of about 85%, proving to be a promising solution for supporting the diagnosis of depression in real-time in a reliable, fast, inexpensive and non-intrusive ways. Laura Verde, Gennaro Raimo, Federica Vitale, Bruno Carbonaro, Gennaro Cordasco, Stefano Marrone 0001, Anna Esposito |
ICTAI | 7 |
| 2021 | Language or Paralanguage, This is the Problem: Comparing Depressed and Non-Depressed Speakers Through the Analysis of Gated Multimodal Units
Nujud Aloshban, Anna Esposito, Alessandro Vinciarelli |
Interspeech | 2 |
| 2020 | Seniors' ability to decode differently aged facial emotional expressionsabstractThe present investigation aims at assessing elders' ability to decode facial emotional expressions conveyed by differently aged people in order to confirm (or disconfirm) the appropriateness of the “own age bias” theory, as well as investigate effects of different ages and different emotional categories. The study, involves 44 healthy elders (23 females), aged 65+ (mean age=75.09; SD=±7.9) which were requested to label 76 pictures depicting elders, middle-aged and young women and men displaying the six facial emotional expressions of disgust, anger, fear, sadness, happiness and neutrality. Results show a complex pattern of influences that calls for more deep investigations on the features to be accounted by providing socially and emotionally believable interfaces of effective and efficient algorithms to detect and decode their users' emotional facial expressions. Anna Esposito, Terry Amorese, Mauro N. Maldonato, Alessandro Vinciarelli, M. Inés Torres, Sergio Escalera, Gennaro Cordasco |
FG | 1 |
| 2020 | Impairments in decoding facial and vocal emotional expressions in high functioning autistic adults and adolescentsabstractThe present investigation shows that gender of stimuli, age, and emotional categories affects the ability of adults and adolescent with Autistic Spectrum Conditions (ASC) to decode facial and vocal emotional expressions. A total of 60 subjects participated to the research: 15 ASC and 15 control adolescents aged between 10-14 years; and 15 ASC and 15 control young adults aged between 20-24 years. Their tasks consisted in decoding: a) 24 adults and 24 children contemporary facial emotional expressions of happiness, sadness, anger, fear, surprise, and disgust; and b) 20 adult's vocal emotional expressions of the same abovementioned emotions (except disgust). Significant differences were observed between ASC and typically developed peers. The data suggest that gender, type (voices or faces) of stimuli, and participants' age affect the emotion recognition process making difficult the definition of a common and shared pattern of emotional expression's recognition compliance among autistic and control groups. These results suggest that efficient and effective e-health technologies need to be able to learn and adapt to user individual traits and subjective needs to offer personalized assistance and support. Anna Esposito, Italia Cirillo, Antonietta Maria Esposito, Leopoldina Fortunati, Gian Luca Foresti, Sergio Escalera, Nikolaos G. Bourbakis |
FG | 1 |
| 2020 | Detecting Depression in Less Than 10 Seconds: Impact of Speaking Time on Depression Detection SensitivityabstractThis article investigates whether it is possible to detect depression using less than 10 seconds of speech. The experiments have involved 59 participants (including 29 that have been diagnosed with depression by a professional psychiatrist) and are based on a multimodal approach that jointly models linguistic (what people say) and acoustic (how people say it) aspects of speech using four different strategies for the fusion of multiple data streams. On average, every interview has lasted for 242.2 seconds, but the results show that 10 seconds or less are sufficient to achieve the same level of recall (roughly 70%) observed after using the entire inteview of every participant. In other words, it is possible to maintain the same level of sensitivity (the name of recall in clinical settings) while reducing by 95%, on average, the amount of time requireed to collect the necessary data. Nujud Aloshban, Anna Esposito, Alessandro Vinciarelli |
ICMI | 2 |
| 2020 | First Workshop on Multimodal e-CoachesabstractT e-Coaches are promising intelligent systems that aims at supporting human everyday life, dispatching advices through different interfaces, such as apps, conversational interfaces and augmented reality interfaces. This workshop aims at exploring how e-coaches might benefit from spatially and time-multiplexed interfaces and from different communication modalities (e.g., text, visual, audio, etc.) according to the context of the interaction. Leonardo Angelini, Mira El Kamali, Elena Mugellini, Omar Abou Khaled, Yordan Dimitrov, Vera Veleva, Zlatka Gospodinova, Nadejda Miteva, Richard Wheeler, Zoraida Callejas Carrión, David Griol, Kawtar Benghazi Akhlaki, Manuel Noguera, Panagiotis D. Bamidis, Evdokimos I. Konstantinidis, Despoina Petsani, Andoni Beristain, Dimitrios I. Fotiadis, Gérard Chollet, M. Inés Torres, Anna Esposito, Hannes Schlieter |
ICMI | 21 |
| 2020 | Spotting the Traces of Depression in Read Speech: An Approach Based on Computational Paralinguistics and Social Signal Processing
Fuxiang Tao, Anna Esposito, Alessandro Vinciarelli |
INTERSPEECH | 2 |
| 2020 | Ethical issues in assistive ambient living technologies for ageing wellabstractAbstract Assistive Ambient Living (AAL) in ageing refers to any device used to support ageing related psychological and physical changes aimed at improving seniors’ quality of life and reducing caregivers’ burdens. The diffusion of these devices opens the ethical issues related to their use in the human personal space. This is particularly relevant when AAL technologies are devoted to the ageing population that exhibits special bio-psycho-social aspects and needs. In spite of this, relatively little research has focused on ethical issues that emerge from AAL technologies. The present article addresses ethical issues emerging when AAL technologies are implemented for assisting the elderly population and is aimed at raising awareness of these aspects among healthcare providers. The overall conclusion encourages a person-oriented approach when designing healthcare facilities. This process must be fulfilled in compliance with the general principles of ethics and individual nature of the person devoted to. This perspective will develop new research paradigms, paving the way for fulfilling essential ethical principles in the development of future generations of personalized AAL devices to support ageing people living independently at their home. Francesco Panico, Gennaro Cordasco, Carl Vogel, Luigi Trojano, Anna Esposito |
Multim. Tools Appl. | 5 |
| 2019 | The Dependability of Voice on Elders' Acceptance of Humanoid Agents
Anna Esposito, Terry Amorese, Marialucia Cuciniello, Maria Teresa Riviello, Antonietta Maria Esposito, Alda Troncone, Gennaro Cordasco |
INTERSPEECH | 1 |
| 2018 | Depression Speaks: Automatic Discrimination between Depressed and Non-Depressed Speakers Based on Nonverbal Speech FeaturesabstractThis article proposes an automatic approach - based on nonverbal speech features - aimed at the automatic discrimination between depressed and non-depressed speakers. The experiments have been performed over one of the largest corpora collected for such a task in the literature (62 patients diagnosed with depression and 54 healthy control subjects), especially when it comes to data where the depressed speakers have been diagnosed as such by professional psychiatrists. The results show that the discrimination can be performed with an accuracy of over 75% and the error analysis shows that the chances of correct classification do not change according to gender, depression-related pathology diagnosed by the psychiatrists or length of the pharmacological treatment (if any). Furthermore, for every depressed subject, the corpus includes a control subject that matches age, education level and gender. This ensures that the approach actually discriminates between depressed and non depressed speakers and does not simply capture differences resulting from other factors. Filomena Scibelli, Giorgio Roffo, Mohammad Tayarani, Luca Bartoli, Gaetano De Mattia, Anna Esposito, Alessandro Vinciarelli |
ICASSP | 6 |
| 2018 | Seniors' Sensing of Agents' Personality from Facial ExpressionsabstractThe presented study investigated the preferences of seniors towards artificial avatars showing personality both from a pragmatic and a hedonic point of view. Also, preferences for technological devices were considered. The involved participants were 45 adults (20 female) aged 65+ years in good health. They were asked to watch video clips of 4 agents (two males and two females) showing different personality traits (i.e. angry, depressed, joyful, and practical), and subsequently had to complete a questionnaire. Subjects were not informed about an avatar’s personality and not openly interviewed regarding this subject. Rather, the administered questionnaire was devoted to test their perception of agents and whether such complies with the intended characteristics. Results show that subjects prefer female agents with a positive personality (joyful and practical) on both pragmatic and hedonic dimensions of the interactive system. Anna Esposito, Stephan Schlögl, Terry Amorese, Antonietta Maria Esposito, M. Inés Torres, Francesco Masucci, Gennaro Cordasco |
ICCHP (2) | 1 |
| 2018 | Power Poses Affect Risk Tolerance and Skin Conductance LevelsabstractHumans are used to express their feelings of selfconfidence/ powerfulness or their distress/sadness through either expansive postures that occupy as much space as possible or closing postures occupying as less space as possible to avoid contact. This conduct suggests that feelings of selfconfidence/ powerfulness or distress/sadness change our body expressions/postures. It can be interesting to assess whether the reverse is also true, i.e. the way we arrange our body at a given moment would affect our feelings. The present research reports an investigation on such argument. To this aim, 50 subjects (25 females) aged between 23 and 31 years were requested to adopt either an expansive (high-powered) or contracted (low-powered) posture for as long as 3 minutes and then asked to bet money in a dice game. The results show that assuming high-power poses favors risk tolerant behaviors and rises feelings of powerfulness. This is not true in the case of low-power postures, which engender a sense of stress, sustained by a significant increase of skin conductance levels. Considerations are made on how to exploit these results for psychotherapy and rehabilitation purposes, as well as, for the implementation of artificial intelligent systems operating as tools for well-being and coaching. Davide Saggese, Gennaro Cordasco, Mauro N. Maldonato, Nikolaos G. Bourbakis, Alessandro Vinciarelli, Anna Esposito |
ICTAI | 6 |
| 2017 | How Traders' Appearances and Moral Descriptions Influence Receivers' Choices in the Ultimatum GameabstractThis work reports on a series of experiments involving 960 participants (aged between 20-30 years and equally balanced by gender), asked to play the receiver role in a modified version of the Ultimatum Game, where together with information on the offer's fairness (e.g. 40 (fair) vs 10 (unfair) of 100 euros), a photo depicted the trader's appearance (trustworthy vs. untrustworthy) and a text provided his moral description (honest vs. dishonest). Receivers were asked to motivate their decision in connection with the appearance, moral judgment, and fairness of the offer, and report on how these variables affected their emotional feelings. Data analysis shows that, in all conditions containing a fair offer, the trader's appearance plays a significant role in the receivers' decisions in terms of acceptance rate. Moral descriptions play a significant role only in conditions containing an unfair offer. However, when asked to motivate their choices, subjects do not feel the interference of the social appearance, rather they provide more or less equal number of motivations with reference to the amount of offers and moral judgments. As for the emotions driving their decisions, non-converging feelings are observed both at intra and inter group level. Anna Esposito, Antonietta Maria Esposito, Marilena Esposito, Filomena Scibelli, Gennaro Cordasco, Carl Vogel, Nikolaos G. Bourbakis |
ICTAI | 1 |
| 2017 | EMOTHAW: A Novel Database for Emotional State Recognition From Handwriting and DrawingabstractThe detection of negative emotions through daily activities such as writing and drawing is useful for promoting wellbeing. The spread of human-machine interfaces such as tablets makes the collection of handwriting and drawing samples easier. In this context, we present a first publicly available database which relates emotional states to handwriting and drawing, that we call EMOTHAW (EMOTion recognition from HAndWriting and draWing). This database includes samples of 129 participants whose emotional states, namely anxiety, depression, and stress, are assessed by the Depression-Anxiety-Stress Scales (DASS) questionnaire. Seven tasks are recorded through a digitizing tablet: pentagons and house drawing, words copied in handprint, circles and clock drawing, and one sentence copied in cursive writing. Records consist in pen positions, on-paper and in-air, time stamp, pressure, pen azimuth, and altitude. We report our analysis on this database. From collected data, we first compute measurements related to timing and ductus. We compute separate measurements according to the position of the writing device: on paper or in-air. We analyze and classify this set of measurements (referred to as features) using a random forest approach. This latter is a machine learning method, based on an ensemble of decision trees, which includes a feature ranking process. We use this ranking process to identify the features which best reveal a targeted emotional state. We then build random forest classifiers associated with each emotional state. We provide accuracy, sensitivity, and specificity evaluation measures obtained from cross-validation experiments. Our results show that anxiety and stress recognition perform better than depression recognition. Laurence Likforman-Sulem, Anna Esposito, Marcos Faúndez-Zanuy, Stéphan Clémençon, Gennaro Cordasco |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2016 | A Human-Like SPN Methodology for Deep Understanding of Technical DocumentsabstractThis paper deals with the Automatic Deep Understanding (ADU) of technical documents. Here we present a synergistic collaboration between two different modalities, a natural language text understanding (NLU) method and a diagram-image extraction & modeling (DIM) one for the deep understanding of technical documents. In particular, the NLU extracts the text from the document and determines the associations among the nouns and their interactions, by creating their stochastic Petri-net (SPN) graph model. The DIM extracts the diagrams from the document and produces their graph models. Then we combine (associate) these two models in a synergistic way, which leads to the deeper understanding of the technical document. Nikolaos G. Bourbakis, Adamantia Psarologou, Giorgia Rematska, Anna Esposito |
ICTAI | 4 |
| 2016 | Effects of Emotional Visual Scenes on the Ability to Decode Emotional MelodiesabstractAn effective change in Human Computer Interaction requires to account of how communication practices are transformed in different contexts, how users sense the interaction with a machine, and an efficient machine sensitivity in interpreting users' communicative signals, and activities. To this aims, the present paper investigates on whether and how positive and negative visual scenes may alter listeners' ability to decode emotional melodies. Emotional tunes were played alone and with, either positive, or negative, or neutral emotional scenes. Afterword, subjects (8 groups, each of 38 subjects, equally balanced by gender) were asked to decode the emotional feeling aroused by melodies ascribing them either emotional valences (positive, negative, I don't know) or emotional labels (happy, sad, fear, anger, another emotion, I don't know). It was found that dimensional emotional features rather than emotional labels strongly affect cognitive judgements of emotional melodies. Musical emotional information is most effectively retained when the task is to assign labels rather than valence values to melodies. In addition, significant misperception effects are observed when happy or positively judged melodies are concurrently played with negative scenes. Anna Esposito, Antonietta Maria Esposito, Marilena Esposito, Maria Teresa Riviello, Alessandro Vinciarelli, Nikolaos G. Bourbakis |
ICTAI | 1 |
| 2015 | On the Amount of Semantic Information Conveyed by GesturesabstractThis paper aims at investigating on whether and how semantic information conveyed by gestures supports communication effectiveness. The research hypothesis was operationalized as a word retrieval task. To this aim, 140 subjects (73 males, 67 females) aged between 18 and 35 years were recruited at the University of Salerno (Italy). They underwent a memory task after being requested to watch video-clips explicating 15 every-day words through 4 different presentation modes: a) only audio, b) only gestures, c) audio and articulatory (complex auditory) information, and d) audio, gestures, and articulatory (multimodal) information. It was found that semantic information is most effectively retained when conveyed through the multimodal mode, and that the only gestures outperforms the only audio and complex auditory mode. It was also found that females have higher significant recall ability than males, no matter the experimental condition. Anna Esposito, Jessica Vassallo, Antonietta Maria Esposito, Nikolaos G. Bourbakis |
ICTAI | 1 |
| 2015 | A Synthesis of Stochastic Petri Net (SPN) Graphs for Natural Language Understanding (NLU) Event/Action AssociationabstractThis paper focuses on the combination of StochasticPetri Net (SPN) graphs for event association in the context ofNatural Languages Understanding (NLU). Our general goal isto develop a new NLU methodology. In this paper we presentsome of its components which are: the use of AnaphoraResolution (AR), the extraction of kernel(s) based on thestructure of their parse trees, and the synthesis of SPN graphs.Each extracted kernel is represented with SPN graphs. Herewe define when, why and how to synthesize the produced SPNgraphs for event/action association, while preserving theirtiming and information flow. Finally, we provide examples ofcombined and uncombined SPN graphs of different NL textscreated using our proposed methodology. Adamantia Psarologou, Anna Esposito, Nikolaos G. Bourbakis |
ICTAI | 2 |
| 2015 | Needs and challenges in human computer interaction for processing social emotional information
Anna Esposito, Antonietta Maria Esposito, Carl Vogel |
Pattern Recognit. Lett. | 1 |
| 2014 | UM3I 2014: International Workshop on Understanding and Modeling Multiparty, Multimodal InteractionsabstractIn this paper, we present a brief summary of the international workshop on Modeling Multiparty, Multimodal Interactions. The UM3I 2014 workshop is held in conjunction with the ICMI 2014 conference. The workshop will highlight recent developments and adopted methodologies in the analysis and modeling of multiparty and multimodal interactions, the design and implementation principles of related human-machine interfaces, as well as the identification of potential limitations and ways of overcoming them. Samer Al Moubayed, Dan Bohus, Anna Esposito, Dirk Heylen, Maria Koutsombogera, Harris Papageorgiou, Gabriel Skantze |
ICMI | 3 |
| 2009 | Empty Speech Pause Detection in Spontaneous SpeechabstractThis work describes two new pause detection algorithms and compare their performance with four standard Voice Activity Detection (VAD) methods represented by the adaptive Long Term Spectral Divergence (LTSD) algorithm, the Likelihood Ratio Test (LRT) algorithm, the Neural Network thresholding and G.729. The proposed algorithms exploit the concept of adaptation in order to handle adverse conditions and spontaneous speech properties. The test data are recordings of spontaneous speech made in noisy environments. The experimental results show that the performance of proposed algorithms on noisy and even artificially cleaned speech are superior than that achieved by standard methods reported in literature . Vojtech Stejskal, Nikolaos G. Bourbakis, Anna Esposito |
ICTAI | 3 |
| 2008 | A Speaker Independent Approach to the Classification of Emotional Vocal ExpressionsabstractThe paper proposes a speaker independent procedure for classifying vocal expressions of emotion. The procedure is based on the splitting up of the emotion recognition process into two steps. In the first step, a combination of selected acoustic features is used to classify six emotions through a Bayesian Gaussian Mixture Model classifier (GMM). The two emotions that obtain the highest likelihood scores are selected for further processing in order to discriminate between them. For this purpose, a unique set of high-level acoustic features was identified using the Sequential Floating Forward Selection (SFFS) algorithm, and a GMM was used to separate between each couple of emotion. The mean classification rate is 81% with an improvement of 5% with respect to the most recent results obtained on the same database (75%). Hicham Atassi, Anna Esposito |
ICTAI (2) | 2 |
| 2008 | A Multimodal Interaction Scheme between a Blind User and the Tyflos Assistive PrototypeabstractThis paper presents the multimodal interaction scheme (visual and audio) used by the Tyflos prototype. Tyflos is a wearable prototype that provides reading and navigating assistance for visually impaired users. In particular, the Tyflos prototype integrates a wireless portable computer, cameras, range and GPS sensors, microphones, natural language processor, text-to-speech device, an ear speaker, a speech synthesizer, a 2D vibration vest and a digital audio recorder. Data collected by the Tyflos sensors is processed by appropriate modules, each of which is specialized in one or more tasks. In this paper we also present a Stochastic Petri-net model of the multimodal interaction scheme for both of the Tyflos capabilities, reading and navigation. Simple illustrative examples from reading and navigation cases are also presented to demonstrate the multimodal interaction. Nikolaos G. Bourbakis, Robert Keefer, Dimitrios Dakopoulos, Anna Esposito |
ICTAI (2) | 4 |
| 2008 | Cognitive Role of Speech Pauses and Algorithmic Considerations for their ProcessingabstractThis study investigates pausing strategies, focusing the attention on empty speech pauses. A cross-modal analysis (video and audio) of spontaneous narratives produced by male and female children and adults showed that a remarkable amount of empty speech pauses was used to signal new concepts in the speech flow and to segment discourse units such as clauses and paragraphs. Based on these results, an adaptive mathematical model for pause distribution was suggested, that exploits, as pause features, the absence of signal and/or the changes of energy over different acoustic dimensions strongly related to the auditory perception. These considerations inspired the formulation and the implementation of two pause detection procedures that proved to be more effective than the Likelihood Ratio Test (LRT) and Long-Term Spectral Divergence (LTSD) algorithms recently proposed in literature and applied for Voice Activity Detection (VAD). Anna Esposito, Vojtech Stejskal, Zdenek Smékal |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2007 | Multi-modal Interfaces for Interaction-Communication between Hearing and Visually Impaired Individuals: Problems and IssuesabstractOne important and challenging problem in human interaction is the communication between blind and deaf individuals. The challenge here involves several cases: (i) first case is a deaf person usually does not speak in order a blind person to hear him/her; (ii) second case is when a blind person speaks a deaf person cannot hear; (iii) third case is when a deaf person makes sign language signs a blind person cannot see them. Thus, this paper presents a study on multi-modal interfaces, issues and problems for establishing communication and interaction between blind and deaf persons. A system-prototype Tyflos-Koufos is proposed in an effort for offering solutions to these challenges. Nikolaos G. Bourbakis, Anna Esposito, Despina Kavraki |
ICTAI (2) | 2 |
| 2006 | Analysis of Invariant Meta-features for Learning and Understanding Disable People's Emotional Behavior Related to Their Health Conditions: A Case Studyabstract"There are million individuals with disabilities with traumatic emotional experiences due to their health issues leading frustration and depression, where a major factor for it is their emotional behavior". Emotion is a topic that has received much attention during the last few years, both in the context of speech synthesis, image understanding as well in automatic speech recognition, interactive dialogues systems and wearable computing. There are few promising studies on the emotional behavior of people with disabilities. These studies are partial due to the lack Information technology and engineering (ITE) techniques that make available a deeper and large scale non-invasive analysis and evaluation of the disabled people emotional behavior in order to provide tools and support for helping them to overcome social and health barriers. A quantitative and qualitative study of emotional invariant meta-features to support the development of emotionally rich man-machine interfaces (interactive dialogue systems and intelligent avatars) for people with disabilities is the subject of this paper Nikolaos G. Bourbakis, Anna Esposito, Despina Kavraki |
BIBE | 2 |
| 2006 | The Role of Timing in Speech Perception and Speech Production Processes and its Effects on Language Impaired IndividualsabstractPhoneme recognition strongly depends on the intrinsic duration of speech segments, phoneme spectral change's durations, and the relative timing of two overlapping events. Excerpts of fluent speech are not very well perceived below a given duration threshold, and "phonemic clauses" are signalled either by a speech pause, or the lengthening of the final word or syllable in the clause. If a sentence should be perceived as temporally fluent, changes made in the duration of one segment should be compensated by durational changes of adjacent segments. These data lead to the conclusion that temporal aspects have a primary role in vocal communication and their perception is basic for a correct exchange of information and a correct identification of the phonetic and phonologic characteristics of the vocal message. In the light of these considerations, the role that the perception of temporal phenomena plays in learning or reading language is investigated considering two theories proposed to explain reading and specific language impairments. The first theory assumes that poor readers are affected by general auditory deficits in temporal processing, whereas the second assumes that the above impairments arose from specific deficits in learning speech. Experimental data in favour of the first or the second theory are discussed, and at the light of the reported results, new experimental paradigms are suggested Anna Esposito, Nikolaos G. Bourbakis |
BIBE | 1 |
| 2006 | A comparison of acoustic coding models for speech-driven facial animation
Praveen K. Kakumanu, Anna Esposito, Oscar N. Garcia, Ricardo Gutierrez-Osuna |
Speech Commun. | 2 |
| 2005 | Audio/visual mapping with cross-modal hidden Markov modelsabstractThe audio/visual mapping problem of speech-driven facial animation has intrigued researchers for years. Recent research efforts have demonstrated that hidden Markov model (HMM) techniques, which have been applied successfully to the problem of speech recognition, could achieve a similar level of success in audio/visual mapping problems. A number of HMM-based methods have been proposed and shown to be effective by the respective designers, but it is yet unclear how these techniques compare to each other on a common test bed. In this paper, we quantitatively compare three recently proposed cross-modal HMM methods, namely the remapping HMM (R-HMM), the least-mean-squared HMM (LMS-HMM), and HMM inversion (HMMI). The objective of our comparison is not only to highlight the merits and demerits of different mapping designs, but also to study the optimality of the acoustic representation and HMM structure for the purpose of speech-driven facial animation. This paper presents a brief overview of these models, followed by an analysis of their mapping capabilities on a synthetic dataset. An empirical comparison on an experimental audio-visual dataset consisting of 75 TIMIT sentences is finally presented. Our results show that HMMI provides the best performance, both on synthetic and experimental audio-visual data. Shengli Fu, Ricardo Gutierrez-Osuna, Anna Esposito, Praveen K. Kakumanu, Oscar N. Garcia |
IEEE Trans. Multim. | 3 |
| 2005 | Speech-driven facial animation with realistic dynamicsabstractThis work presents an integral system capable of generating animations with realistic dynamics, including the individualized nuances, of three-dimensional (3-D) human faces driven by speech acoustics. The system is capable of capturing short phenomena in the orofacial dynamics of a given speaker by tracking the 3-D location of various MPEG-4 facial points through stereovision. A perceptual transformation of the speech spectral envelope and prosodic cues are combined into an acoustic feature vector to predict 3-D orofacial dynamics by means of a nearest-neighbor algorithm. The Karhunen-Loe/spl acute/ve transformation is used to identify the principal components of orofacial motion, decoupling perceptually natural components from experimental noise. We also present a highly optimized MPEG-4 compliant player capable of generating audio-synchronized animations at 60 frames/s. The player is based on a pseudo-muscle model augmented with a nonpenetrable ellipsoidal structure to approximate the skull and the jaw. This structure adds a sense of volume that provides more realistic dynamics than existing simplified pseudo-muscle-based approaches, yet it is simple enough to work at the desired frame rate. Experimental results on an audiovisual database of compact TIMIT sentences are presented to illustrate the performance of the complete system. Ricardo Gutierrez-Osuna, Praveen K. Kakumanu, Anna Esposito, Oscar N. Garcia, Adriana Bojórquez, José Luis Castillo, Isaac Rudomín |
IEEE Trans. Multim. | 3 |
| 2004 | A general framework for learning rules from dataabstractWith the aim of getting understandable symbolic rules to explain a given phenomenon, we split the task of learning these rules from sensory data in two phases: a multilayer perceptron maps features into propositional variables and a set of subsequent layers operated by a PAC-like algorithm learns Boolean expressions on these variables. The special features of this procedure are that: i) the neural network is trained to produce a Boolean output having the principal task of discriminating between classes of inputs; ii) the symbolic part is directed to compute rules within a family that is not known a priori; iii) the welding point between the two learning systems is represented by a feedback based on a suitability evaluation of the computed rules. The procedure we propose is based on a computational learning paradigm set up recently in some papers in the fields of theoretical computer science, artificial intelligence and cognitive systems. The present article focuses on information management aspects of the procedure. We deal with the lack of prior information about the rules through learning strategies that affect both the meaning of the variables and the description length of the rules into which they combine. The paper uses the task of learning to formally discriminate among several emotional states as both a working example and a test bench for a comparison with previous symbolic and subsymbolic methods in the field. Bruno Apolloni, Anna Esposito, Dario Malchiodi, Christos Orovas, Giorgio Palmas, John G. Taylor |
IEEE Trans. Neural Networks | 2 |
| 2002 | Holds as gestural correlates to empty and filled speech pauses
Anna Esposito, Susan Duncan, Francis K. H. Quek |
INTERSPEECH | 1 |
| 2000 | Approximation of continuous and discontinuous mappings by a growing neural RBF-based algorithm
Anna Esposito, Maria Marinaro, Domenico Oricchio, Silvia Scarpetta |
Neural Networks | 1 |
| 1998 | An incremental local radial basis function network
Anna Esposito, Maria Marinaro, Silvia Scarpetta |
ESANN | 1 |
| 1997 | The amplitudes of the peaks in the spectrum: data from /a/ contextabstractThis work is devoted to the study of the properties of the sound spectrum at the release of Italian stop consonants in vocalic contexts. The aim is to check if the amplitudes of the peaks in the spectrum can be used as acoustic attributes of the place of articulation of the consonants. This information is useful for defining an automatic algorithm which can discriminate among different place of articulation using simple data such as the values, in dB, of the maximum peaks in different frequency ranges. Moreover, different measurements have been performed (the spectra are computed at the release, averaged over 10 msec after the release, and using a smoothed spectrum) in order to define which measure retains more information about peak amplitudes. Materials and procedures The recording and measurements were made at the Research Laboratory of Electronics, Speech Communication Group, MIT, Cambridge, USA. The materials consisted in VCVC utterances produced by seven adult Italian speakers (three females and four males) in a sound-treated room and recorded on a high-quality magnetic tape recording system. The utterances were embedded in a carrier phrase. The measurements were made for the intervocalic consonant. Data were collected for all Italian vowels embedded in stop contexts. However, the results reported in the present paper are derived from the analysis of the stop consonants in the [a] context. The spectral representations used Supported by IIASS, CNR, and INFM Salerno University. Acknoledgements goes to M. Grabriella Di Benedetto for her useful comments and suggestions include a DFT spectrum, a smoothed DFT, a spectral averaging. The analysis window (Hamming window) was set to 3.1 msec. The spectrum at the consonant release, the averaged spectrum over the first 4 msec (for [b, d, g]) and over 10 msec (for [p, t, k]) after the release and, the k-averaged2 spectrum were computed using a software program developed by Klatt (1984). All spectra were preemphasized, and the spectral amplitudes were enhanced by modifying an overall spectral gain control parameter. The amplitudes of the maximum peaks in different frequency ranges were measured by visual examination. The amplitude attributes The peaks amplitudes measured in the different frequency ranges described above were compared in order to identify properties that can be useful to discriminate the place of articulation of each consonant. Initially averages of the maximum peak amplitudes in different frequency ranges were computed. However, even though some of these averages differ significantly from one consonant to another, the standard deviations were high and they overlapped. This effect is mostly due to the variability of the peak amplitudes among the speakers. For this reason we decided to exclude these measures and we start to look to the amplitudes of the maximum peaks in specified frequency ranges compared to the amplitudes of the maximum peaks in other The k-averaged spectrum was computed by measuring the VOT length of the voiceless consonant. The cursor was then placed on the waveform at the temporal sampling point corresponding to half the VOT length, and the spectrum was averaged over 5 msec to the left and 5 msec to the right of this sam- Anna Esposito |
EUROSPEECH | 1 |
| 1996 | Preprocessing and neural classification of English stop consonants [b, d, g, p, t, k]
Anna Esposito, Eugène C. Ezin, M. Ceccarelli |
ICSLP | 1 |
| 1994 | A neural network for error correcting decoding of binary linear codes
Anna Esposito, Salvatore Rampone, Roberto Tagliaferri |
Neural Networks | 1 |