Alessandro Vinciarelli

dblp:09/6044 · DBLP profile ↗
← Back
116ranked-venue papers
24as first author
30since 2021 · last 2026
0000-0002-9048-0524ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 12 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 8 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 33 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Predicting Fibromyalgia Pain From Self-reported Health Data Using Time-Series and Deep Learning Models
Delnia Alipour, Olga Perepelkina, Alessandro Vinciarelli, Tahir Janmohamed, Binh P. Nguyen, Simone Stumpf
AIME (1)3
2025 "SAM" - School Attachment Monitor for the Assessment of Attachment in Children aged 4-8
abstract
Attachment is a crucial part of children's psychological development and can impact social relationships and overall mental health and well-being.However, measures to assess attachment are traditionally costly in terms of time and use.Though previous interventions have tried to support the assessment of attachment in children, there have been few developments of a low-cost, fast, and efficient tool.Therefore, we present the demo for a School Attachment Monitor (SAM).The SAM is a digital tool that utilises artificial intelligence to assess attachment in children with 80% accuracy [21].Our demo showcases the SAM at a key point of development, exploring alternative methods of interaction to suit young children today and its deployability in schools.
Michael John Saiger, Yara Aleid, Helen Minnis, Alessandro Vinciarelli, Stephen A. Brewster
IDC4
2025 Punctual or Continuous? Analyzing Depression Traces in Language and Paralanguage with Multiple Instance Learning
abstract
The key-question addressed in this article is whether the traces of depression in speech are punctual (the pathology manifests itself at specific points in time) or continuous (the pathology manifests itself at every moment in time), where the expression “speech” refers to both speech signals and their transcriptions. For this reason, this work compares the performances of different approaches (both unimodal or multimodal) based either on the assumption that the traces are punctual or on the other one. In this way, it is possible to test which of the alternative assumptions is more realistic. The experiments were performed over a publicly available dataset (the Androids Corpus) and the results include F1 Scores up to 93.1%, among the best reported in literature for the Corpus. Furthermore, the results suggest that depression traces are punctual, but they appear so frequently that approaches based on the assumption of continuous traces still perform well. The conclusions discuss the implications of such observation.
Rawan Alsarrani, Anna Esposito, Alessandro Vinciarelli
ICMI3
2025 Multimodal Analysis of Disagreement in Dyadic Conversations: An Approach Based on Emotion Recognition
abstract
This article proposes a multimodal approach for the detection of disagreement in dyadic conversations, where disagreement means that people express different opinions about a topic under discussion. The key-assumption underlying the work is that people tend to manifest different emotions depending on whether they are disagreeing or not. Therefore, emotions can provide evidence that disagreement is taking place. The experiments were performed over a corpus of 684 clips involving 60 dyads (120 persons and roughly 8 hours of speech). Each clip revolves around a decision-making task and it is annotated in terms of the percentage of time people spend in disagreement. For the sake of reproducibility, the Glasgow Disagreement Corpus, the data used in the experiments, has been made accessible through a link available in the paper. The results show that a multimodal approach based on language and paralanguage can predict such a percentage with Mean Absolute Error 9.7 and correlation 0.52 between actual and predicted percentage of time spent in disagreement.
Areej Buker, Emily Smith, Olga Perepelkina, Alessandro Vinciarelli
ICMI4
2025 From Speech and PPG to EDA: Stress Detection Based on Cross-Modal Fine-Tuning of Foundation Models
abstract
Foundation Models trained to perform a certain task can be fine-tuned to other tasks with limited data and computational resources. The advantage of such practice is that it makes it possible to benefit, at least indirectly, from the large amounts of data and the major computational infrastructure necessary for training a Foundation Model. However, there is a limitation too, namely that the few organizations that have the major resources necessary to develop and train Foundation Models do it only for the modalities that are of interest to them. For this reason, this article proposes to fine-tune Foundation Models trained on speech and photoplethysmography signals to perform stress detection based on Electro-Dermal Activity, a modality for which no Foundation Model exists. To the best of our knowledge, this is one of the first works proposing experiments of this type and the results show state-of-the-art stress detection performances over a publicly available benchmark, even if speech and photoplethysmograhpy data differ significantly from Electro-Dermal Activity signals.
Alia Ahmed Al Dossary, Mathieu Chollet, Alessandro Vinciarelli
ICMI3
2025 Birds of a Feather Augment Together: Exploring Sonic Links Between Real and Virtual Worlds in Audio Augmented Reality
abstract
Augmented reality (AR) applications present virtual elements that are connected to the real world. While these may be aware of a user's geographical or visual context, the real sounds in a user's environment are rarely used in the AR experience. We investigate audio augmented reality (AAR) with sound as the primary output. We present the first evaluation of sonic linking, where an AAR application uses real sounds (bird calls and car engines) in a user's surroundings to drive the interaction. We developed two AAR applications to investigate how to design and use such links: a game where entities are spawned based on real-world sounds and a music player with sound-reactive filtering and volume adjustment. Design variations are compared to cover different AAR scenarios, types of sonic link, and existing unlinked equivalents. The results show that sonic linking can create a more augmented, engaging AAR experience, and may alter a user's relationship with their real-world surroundings, enabling new types of augmented reality applications.
Jacob Bhattacharyya, Alessandro Vinciarelli, Stephen A. Brewster
ISMAR2
2025 Sonomancer: Exploring Sonic Control Schemes for Audio Augmented Reality Games
abstract
Audio augmented reality (AAR) games allow for audio-only game experiences which blend real and virtual worlds. Despite being sound-based in their output, these games utilise a small number of control schemes, and rarely deploy control schemes that are also sonic in nature. We present the first evaluation of sonic control schemes for AAR gameplay. We present a systematic literature review that collates the control schemes and game scenarios deployed in existing AAR games, and a user study comparing these traditional control schemes with sonic controls -- speech, music, and sonic gesture -- in minigames corresponding to the most common game scenarios. Our results show that while there are some key differences between these control schemes, sonic controls show promise for use in AAR gameplay, and could be deployed in scenarios where AAR developers currently reach for established control schemes.
Jacob Bhattacharyya, Alessandro Vinciarelli, Stephen A. Brewster
Proc. ACM Hum. Comput. Interact.2
2025 Learning Classifier Performance in an Ensemble of Classifiers for Personality Prediction Using Laughter
abstract
This paper conducts a study on the acoustic features of laughter signals and performs experiments to identify the most relevant features. Different acoustic features are extracted from laughter signals, and studies are performed to identify the features that are more representative of the personality traits. An ensemble learning algorithm is proposed in this paper that gets the laughter signals from the speakers as input and predicts their personality traits. A Wagging algorithm is presented in this work that generates a set of diverse learning algorithms. A pruning algorithm is then proposed that finds the optimal subset of learning algorithms for the ensemble. In the proposed ensemble learning algorithm, a weighted averaging scheme is proposed to aggregate the output of the learning algorithms. In this scheme, during the training phase, a mechanism is adopted that measures the performance of each classifier in different areas in the feature space. These data are then used to generate a model of the performance of the basic classifiers at any given point in the feature space. In the classification phase, for any given test data record, this model is used to estimate the accuracy of the base learners. Then, a vote is performed among the classifiers, and this estimate is used to adjust the weight of the votes. The experiments are performed over a corpus of 1157 laughter bouts produced by 120 individuals to study whether laughter could be used to predict whether a person is above or below the median with respect to personality traits.
Mohammad-Hassan Tayarani-Najaran, Alessandro Vinciarelli
IEEE Trans. Affect. Comput.2
2024 Cross-Data Multilevel Attention for Depression Detection: Analyzing the Interplay Between Read and Spontaneous Speech
abstract
This work proposes a novel Cross-Data Multilevel Attention (CDMA) approach for multi-type speech-based depression detection, encompassing both read and spontaneous speech. The main novelty lies in analyzing the unique and common representations of the two types of speech and integrating them into a unified end-to-end framework with novel Intra-Type Multi-Local Attention (IT-MLA) and Cross-Type Global Attention (CT-GA) mechanisms. In particular, IT-MLA highlights depression-relevant information unique in either read or spontaneous speech via intra-modal attention-aware interactions. Furthermore, CT-GA further emphasises the depression-relevant common information in both read and spontaneous speech, with each type being guided by the other. These multiple enhanced representations are aggregated to produce the final predictions. Experiments conducted on a publicly available corpus of 104 speakers (including 52 diagnosed with depression by professional psychiatrists) demonstrate that the proposed CDMA achieves an F1 score of up to 92.5%, the highest performance recorded on this dataset.
Fuxiang Tao, Xuri Ge, Anna Esposito, Alessandro Vinciarelli
BIBM5
2024 Assessing Privacy Risks of Attribute Inference Attacks Against Speech-Based Depression Detection System
abstract
Many AI applications now attempt to infer users’ mental health conditions, such as depression, from their speech data. In addition to the spoken words, the speech audio contains information about speaker’s identity and demographic attributes, exposing users to serious privacy risks. Previous efforts have primarily focused on developing deep models that preserve privacy; however, there have been few attempts to systematically assess and quantify privacy risks in such systems. We present the first framework for systematically assessing privacy risks in a multimodal (audio-lexical) depression detection system particularly looking at attribute inference attacks. Unlike past works that considered only white-box gender inference attacks against unimodal systems, our framework designs novel white-box and black-box attacks across multiple modalities against three protected speaker attributes: gender, age and education level. We present extensive results on a large, clinically validated dataset, demonstrating critical vulnerability of depression detection systems, where an adversary can infer speaker attributes with 59% - 68% accuracy even for inputs as short as 10 seconds of speech. Our results offer insights and guidelines to inform the development and benchmarking of privacy-preserving models for speech-based depression detection systems. Our code and data are available at: https://github.com/apr-aia/privacy_risks
Basmah Alsenani, Anna Esposito, Alessandro Vinciarelli, Tanaya Guha
ECAI3
2024 Emotion Recognition for Multimodal Recognition of Attachment in School-Age Children
abstract
Attachment is a psychological construct describing the relationship between children and their caregivers. Attachment issues lead to difficulties in the relationships with others and, as a consequence, to higher chances to have negative experiences in adult life (e.g., antisocial behaviour, mental health problems, etc.). However, early detection of attachment issues can help to attenuate the risks. For this reason, this article addresses the problem of attachment recognition in school-age children. The main novelty of the proposed approach is that it is based on emotion recognition. The motivation behind such choice is that the way people regulate their emotions is a marker of their attachment condition. In the experiments, pre-trained models based on Neural Networks were used to extract features fed to attachment classifiers capable to identify children with attachment issues. The best result (F1 Score 73.0% and Accuracy 77.9% over a corpus including 104 children) was obtained with a multimodal approach outperforming, to a statistically significant extent, unimodal methodologies based on language or paralanguage.
Areej Buker, Alessandro Vinciarelli
ICMI2
2024 Is Distance a Modality? Multi-Label Learning for Speech-Based Joint Prediction of Attributed Traits and Perceived Distances in 3D Audio Immersive Environments
abstract
To the best of our knowledge, this article presents the first experiments on speech-based Automatic Personality Perception performed in a virtual immersive audio environment. The key-difference compared to all previous works in the literature is that, in a virtual immersive environment, people perceive not only the voice of the speakers, but also their position and distance in space. Therefore, it is possible to investigate for the first time whether people tend to attribute different traits to people speaking at different distances and, if so, whether this makes a difference in terms of Automatic Personality Perception. The experiments were performed over 360 recordings rendered at different distances (120 speakers including 60 female and 60 male). The results show that there are correlations between perceived distance and personality judgments. Furthermore, the experiments show that the performance in Automatic Personality Perception improves when taking perceived distance into account. These results are important because immersive environments are likely to become one of the main technological interfaces through which people interact with one another and with machines.
Evangelia Fringi, Nesreen Alshubaily, Lorenzo Picinali, Stephen A. Brewster, Tanaya Guha, Alessandro Vinciarelli
ICMI6
2024 Exploring 3D Human Pose Estimation and Forecasting from the Robot's Perspective: The HARPER Dataset
abstract
We introduce HARPER, a novel dataset for 3D body pose estimation and forecasting in dyadic interactions between users and Spot, the quadruped robot manufactured by Boston Dynamics. The key-novelty of HARPER is its focus on the robot’s perspective, i.e., on the data captured by the robot’s sensors. This makes 3D body pose analysis challenging, as being close to the ground results in only partial captures of humans. The scenario underlying HARPER includes 15 actions, of which 10 involve physical contact between the robot and users. The corpus contains recordings not only from Spot’s built-in stereo cameras but also from a 6-camera OptiTrack system, with all recordings synchronized. This setup leads to ground-truth skeletal representations with a precision of less than a millimeter. Additionally, the corpus includes reproducible benchmarks for 3D Human Pose Estimation, Human Pose Forecasting, and Collision Prediction, all based on publicly available baseline approaches. This enables future HARPER users to rigorously compare their results with those provided in this work.
Andrea Avogaro, Andrea Toaiari, Federico Cunico, Xiangmin Xu 0003, Haralambos Dafas, Alessandro Vinciarelli, Liying Li 0001, Marco Cristani
IROS6
2024 Discriminative Power of Handwriting and Drawing Features in Depression
abstract
This study contributes knowledge on the detection of depression through handwriting/drawing features, to identify quantitative and noninvasive indicators of the disorder for implementing algorithms for its automatic detection. For this purpose, an original online approach was adopted to provide a dynamic evaluation of handwriting/drawing performance of healthy participants with no history of any psychiatric disorders ([Formula: see text]), and patients with a clinical diagnosis of depression ([Formula: see text]). Both groups were asked to complete seven tasks requiring either the writing or drawing on a paper while five handwriting/drawing features’ categories (i.e. pressure on the paper, time, ductus, space among characters, and pen inclination) were recorded by using a digitalized tablet. The collected records were statistically analyzed. Results showed that, except for pressure, all the considered features, successfully discriminate between depressed and nondepressed subjects. In addition, it was observed that depression affects different writing/drawing functionalities. These findings suggest the adoption of writing/drawing tasks in the clinical practice as tools to support the current depression detection methods. This would have important repercussions on reducing the diagnostic times and treatment formulation.
Claudia Greco, Gennaro Raimo, Terry Amorese, Marialucia Cuciniello, Gavin McConvey, Gennaro Cordasco, Marcos Faúndez-Zanuy, Alessandro Vinciarelli, Zoraida Callejas Carrión, Anna Esposito
Int. J. Neural Syst.8
2024 On the effects of obfuscating speaker attributes in privacy-aware depression detection
Nujud Aloshban, Anna Esposito, Alessandro Vinciarelli, Tanaya Guha
Pattern Recognit. Lett.3
2023 Multi-Local Attention for Speech-Based Depression Detection
abstract
This article shows that an attention mechanism, the Multi-Local Attention, can improve a depression detection approach based on Long Short-Term Memory Networks. Besides leading to higher performance metrics (e.g., Accuracy and F1 Score), Multi-Local Attention improves two other aspects of the approach, both important from an application point of view. The first is the effectiveness of a confidence score associated to the detection outcome at identifying speakers more likely to be classified correctly. The second is the amount of speaking time needed to classify a speaker as depressed or non-depressed. The experiments were performed over read speech and involved 109 participants (including 55 diagnosed with depression by professional psychiatrists). The results show accuracies up to 88.0% (F1 Score 88.0%).
Fuxiang Tao, Xuri Ge, Anna Esposito, Alessandro Vinciarelli
ICASSP5
2023 Privacy Risks in Speech Emotion Recognition: A Systematic Study on Gender Inference Attack
abstract
Increasingly more applications now use deep networks to analyse speaker's affective states. An undesirable side effect is that models trained to perform one task (e.g, emotion from speech) can be attacked to infer other, possibly privacy-sensitive attributes (e.g., gender) of the speaker. The amount of information an attacker can infer through such attacks is called leakage, and this article presents the first systematic study of the interplay between gender leakage and the main characteristics of the attacker model (family, architecture and training condition). To this end, we define various attack scenarios, and perform extensive experiments to analyse privacy risks in Speech Emotion Recognition (SER). Results show that SER models can leak a speaker's gender with an accuracy of 51% to 95% (upper bound) depending on the attack condition. Furthermore, our results provide fresh insights on how to limit the effectiveness of possible attacks and, thereby, to ensure privacy preservation.
Basmah Alsenani, Tanaya Guha, Alessandro Vinciarelli
INTERSPEECH3
2023 Multiple Instance Learning for Inference of Child Attachment From Paralinguistic Aspects of Speech
abstract
Attachment is a psychological construct that accounts for the way children perceive their relationship with their caregivers. Depending on the attachment condition, a child can either be secure or insecure. Identifying as many insecure children as possible is important to mitigate the negative consequences of insecure attachment in adult life. For this reason, this article proposes an attachment recognition approach that, compared to other approaches, increases the Recall, the percentage of insecure children identified as such. The approach is based on Multiple Instance Learning, a body of methodologies dealing with data represented as "bags" of feature vectors. This is suitable for speech recordings because these are typically represented as vector sequences. The experiments involved 104 participants of age 5 to 9. The results show that insecure children can be identified with Recall up to 63.3% (accuracy up to 75%), an improvement with respect to most existing models.
Abeer A. N. Buker, Huda Alsofyani, Alessandro Vinciarelli
INTERSPEECH3
2023 The Androids Corpus: A New Publicly Available Benchmark for Speech Based Depression Detection
Fuxiang Tao, Anna Esposito, Alessandro Vinciarelli
INTERSPEECH3
2022 Thin Slices of Depression: Improving Depression Detection Performance Through Data Segmentation
abstract
The computing community is making major efforts towards automatic detection of depression, a serious pathology that affects roughly 4.4% of the world’s population. One of the main difficulties is the collection of data aimed at training models capable to learn differences between depressed and non-depressed people. In fact, data collection in the depression domain requires the respect of rigorous ethical constraints that, inevitably, limit the size of the corpora that can be collected. This article proposes to address the problem by using the thin slices theory, i.e., the possibility to detect the inner state of an individual (depression in this case) through very short samples of behavior. In particular, the article shows that the performance of data-driven models can be improved by segmenting the data at disposition into thin slices and then training data-driven models over them. This increases the amount of samples at disposition and allows a relative F1 Score improvement by up to 16.2%.
Rawan Alsarrani, Anna Esposito, Alessandro Vinciarelli
ICASSP3
2022 Attachment Recognition in School-Age Children: A Multimodal Approach Based on Language and Paralanguage Analysis
abstract
Attachment is the psychological construct accounting for whether parents address effectively physical and emotional needs of their children or not. The approach proposed in this work recognizes whether a child is secure or insecure, the two major attachment conditions an individual can belong to. The approach is based on the combination of language and paralanguage, what children say and how they say it. The experiments involved 104 children of age between 5 and 9 that were recorded while undergoing the Manchester Attachment Story Task, one of the main psychometric instruments child psychiatrists use to assess the attachment condition of children. The results show that it is possible to achieve an accuracy of up 74.6% (F1 Score 66.7%), meaning that the approach correctly identifies the attachment condition of a child three times out of four, on average.
Huda Alsofyani, Alessandro Vinciarelli
ICASSP2
2022 Automatic Detection of Reactive Attachment Disorder Through Turn-Taking Analysis in Clinical Child-Caregiver Sessions
Andrei Bîrladeanu, Helen Minnis, Alessandro Vinciarelli
INTERSPEECH3
2022 Which Model is Best: Comparing Methods and Metrics for Automatic Laughter Detection in a Naturalistic Conversational Dataset
Gordon Rennie, Olga Perepelkina, Alessandro Vinciarelli
INTERSPEECH3
2022 What an "Ehm" Leaks About You: Mapping Fillers into Personality Traits with Quantum Evolutionary Feature Selection Algorithms
abstract
This work shows that fillers - short utterances like “ehm” and “uhm” - allow one to predict whether someone is above median along the Big-Five personality traits. The experiments have been performed over a corpus of 2,988 fillers uttered by 120 different speakers in spontaneous conversations. The results show that the prediction accuracies range between 74 and 82 percent depending on the particular trait. The proposed approach includes a feature selection step - based on Quantum Evolutionary Algorithms - that has been used to detect the personality markers, i.e., the subset of the features that better account for the prediction outcomes and, indirectly, for the personality of the speakers. The results show that only a relatively few features tend to be consistently selected, thus acting as reliable personality markers.
Mohammad Tayarani, Anna Esposito, Alessandro Vinciarelli
IEEE Trans. Affect. Comput.3
2022 Synthetic vs Human Emotional Faces: What Changes in Humans' Decoding Accuracy
abstract
Considered the increasing use of assistive technologies in the shape of virtual agents, it is necessary to investigate those factors which characterize and affect the interaction between the user and the agent, among these emerges the way in which people interpret and decode synthetic emotions, i.e., emotional expressions conveyed by virtual agents. For these reasons, an article is proposed, which involved 278 participants split in differently aged groups (young, middle-aged, and elders). Within each age group, some participants were administered a “naturalistic decoding task,” a recognition task of human emotional faces, while others were administered a “synthetic decoding task” namely emotional expressions conveyed by virtual agents. Participants were required to label pictures of female and male humans or virtual agents of different ages (young, middle-aged, and old) displaying static expressions of disgust, anger, sadness, fear, happiness, surprise, and neutrality. Results showed that young participants showed better recognition performances (compared to older groups) of anger, sadness, and neutrality, while female participants showed better recognition performances (compared to males) of sadness, fear, and neutrality; sadness and fear were better recognized when conveyed by real human faces, while happiness, surprise, and neutrality were better recognized when represented by virtual agents. Young faces were better decoded when expressing anger and surprise, middle-aged faces were better decoded when expressing sadness, fear, and happiness, while old faces were better decoded in the case of disgust; on average, female faces where better decoded compared to male ones.
Terry Amorese, Marialucia Cuciniello, Alessandro Vinciarelli, Gennaro Cordasco, Anna Esposito
IEEE Trans. Hum. Mach. Syst.3
2021 I Feel it in Your Fingers: Inference of Self-Assessed Personality Traits from Keystroke Dynamics in Dyadic Interactive Chats
abstract
The question at the core of this work is whether it is possible to infer self-assessed personality traits from keystroke dynamics (the way people type on a keyboard). The experiments were performed over a corpus of 30 dyadic chats, 60 participants in total, collected through a text-based chat interface similar to those available in popular products (e.g., Skype). The results show that keystroke dynamics (typing speed, frequency of deletions, etc.) allow one to infer whether someone is below median or not along the Big Five personality traits. In particular, it was possible to achieve F1 Scores up to 72% depending on the trait. To the best of our knowledge, this is the first work aimed at recognizing personality traits through analysis of keystroke dynamics.
Abeer A. N. Buker, Alessandro Vinciarelli
ACII2
2021 Attachment Recognition in School Age Children Based on Automatic Analysis of Facial Expressions and Nonverbal Vocal Behaviour
abstract
Attachment is a psychological construct that accounts for whether children are secure (the parents meet their physical and emotional needs) or insecure (the parents do not meet their physical and emotional needs). Unless identified and supported early enough, insecure children develop higher chances of experiencing issues such as antisocial behaviour or suicidal tendencies. For this reason, this article proposes a multimodal approach for attachment recognition in school age children (5-9 years old). In particular, the approach infers the attachment condition of a child from facial expressions and nonverbal vocal behaviour. The experiments involved 104 children that were recorded while undergoing the Manchester Child Attachment Story Test, an instrument that child psychiatrists use often to identify insecure children. The results show that attachment can be recognized with accuracy up to 71.2% (F1 score 62.4%).
Huda Alsofyani, Alessandro Vinciarelli
ICMI2
2021 Language or Paralanguage, This is the Problem: Comparing Depressed and Non-Depressed Speakers Through the Analysis of Gated Multimodal Units
Nujud Aloshban, Anna Esposito, Alessandro Vinciarelli
Interspeech3
2021 Stacked Recurrent Neural Networks for Speech-Based Inference of Attachment Condition in School Age Children
Huda Alsofyani, Alessandro Vinciarelli
Interspeech2
2021 Infinite Feature Selection: A Graph-based Feature Filtering Approach
abstract
We propose a filtering feature selection framework that considers subsets of features as paths in a graph, where a node is a feature and an edge indicates pairwise (customizable) relations among features, dealing with relevance and redundancy principles. By two different interpretations (exploiting properties of power series of matrices and relying on Markov chains fundamentals) we can evaluate the values of paths (i.e., feature subsets) of arbitrary lengths, eventually go to infinite, from which we dub our framework Infinite Feature Selection (Inf-FS). Going to infinite allows to constrain the computational complexity of the selection process, and to rank the features in an elegant way, that is, considering the value of any path (subset) containing a particular feature. We also propose a simple unsupervised strategy to cut the ranking, so providing the subset of features to keep. In the experiments, we analyze diverse settings with heterogeneous features, for a total of 11 benchmarks, comparing against 18 widely-known comparative approaches. The results show that Inf-FS behaves better in almost any situation, that is, when the number of features to keep are fixed a priori, or when the decision of the subset cardinality is part of the process.
Giorgio Roffo, Simone Melzi, Umberto Castellani, Alessandro Vinciarelli, Marco Cristani
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Seniors' ability to decode differently aged facial emotional expressions
abstract
The present investigation aims at assessing elders' ability to decode facial emotional expressions conveyed by differently aged people in order to confirm (or disconfirm) the appropriateness of the “own age bias” theory, as well as investigate effects of different ages and different emotional categories. The study, involves 44 healthy elders (23 females), aged 65+ (mean age=75.09; SD=±7.9) which were requested to label 76 pictures depicting elders, middle-aged and young women and men displaying the six facial emotional expressions of disgust, anger, fear, sadness, happiness and neutrality. Results show a complex pattern of influences that calls for more deep investigations on the features to be accounted by providing socially and emotionally believable interfaces of effective and efficient algorithms to detect and decode their users' emotional facial expressions.
Anna Esposito, Terry Amorese, Mauro N. Maldonato, Alessandro Vinciarelli, M. Inés Torres, Sergio Escalera, Gennaro Cordasco
FG4
2020 Detecting Depression in Less Than 10 Seconds: Impact of Speaking Time on Depression Detection Sensitivity
abstract
This article investigates whether it is possible to detect depression using less than 10 seconds of speech. The experiments have involved 59 participants (including 29 that have been diagnosed with depression by a professional psychiatrist) and are based on a multimodal approach that jointly models linguistic (what people say) and acoustic (how people say it) aspects of speech using four different strategies for the fusion of multiple data streams. On average, every interview has lasted for 242.2 seconds, but the results show that 10 seconds or less are sufficient to achieve the same level of recall (roughly 70%) observed after using the entire inteview of every participant. In other words, it is possible to maintain the same level of sensitivity (the name of recall in clinical settings) while reducing by 95%, on average, the amount of time requireed to collect the necessary data.
Nujud Aloshban, Anna Esposito, Alessandro Vinciarelli
ICMI3
2020 Did the Children Behave?: Investigating the Relationship Between Attachment Condition and Child Computer Interaction
abstract
This work investigates the interplay between Child-Computer Interaction and attachment, a psychological construct that accounts for how children perceive their parents to be. In particular, the article makes use of a multimodal approach to test whether children with different attachment conditions tend to use differently the same interactive system. The experiments show that the accuracy in predicting usage behaviour changes, to a statistically significant extent, according to the attachment conditions of the 52 experiment participants (age-range 5 to 9). Such a result suggests that attachment-relevant processes are actually at work when people interact with technology, at least when it comes to children.
Dong-Bach Vo, Stephen A. Brewster, Alessandro Vinciarelli
ICMI3
2020 Spotting the Traces of Depression in Read Speech: An Approach Based on Computational Paralinguistics and Social Signal Processing
Fuxiang Tao, Anna Esposito, Alessandro Vinciarelli
INTERSPEECH3
2020 Speech Synthesis for the Generation of Artificial Personality
abstract
A synthetic voice personifies the system using it. In this work we examine the impact text content, voice quality and synthesis system have on the perceived personality of two synthetic voices. Subjects rated synthetic utterances based on the Big-Five personality traits and naturalness. The naturalness rating of synthesis output did not correlate significantly with any Big-Five characteristic except for a marginal correlation with openness. Although text content is dominant in personality judgments, results showed that voice quality change implemented using a unit selection synthesis system significantly affected the perception of the Big-Five, for example tense voice being associated with being disagreeable and lax voice with lower conscientiousness. In addition a comparison between a parametric implementation and unit selection implementation of the same voices showed that parametric voices were rated as significantly less neurotic than both the text alone and the unit selection system, while the unit selection was rated as more open than both the text alone and the parametric system. The results have implications for synthesis voice and system type selection for applications such as personal assistants and embodied conversational agents where developing an emotional relationship with the user, or developing a branding experience is important.
Matthew P. Aylett, Alessandro Vinciarelli, Mirjam Wester
IEEE Trans. Affect. Comput.2
2019 Automating the Administration and Analysis of Psychiatric Tests: The Case of Attachment in School Age Children
abstract
This article presents the School Attachment Monitor, a novel interactive system that can reliably administer the Manchester Child Attachment Story Task (a standard psychiatric test for the assessment of attachment in children) without the supervision of trained professionals. Attachment problems in children cause significant mental health issues and costs to society which technology has the potential to reduce. SAM collects, through instrumented doll-play games, enough information to allow a human assessor to manually identify the attachment status of children. Experiments show that the system successfully does this in 87.5% of cases. In addition, the experiments show that an automatic approach based on deep neural networks can map the information collected into the attachment condition of the children. The outcome SAM matches the judgment of expert human assessors in 82.8% of cases. This is the first time an automated tool has been successful in measuring attachment. This work has significant implications for psychiatry as it allows professionals to assess many more children cost effectively and to direct healthcare resources more accurately and efficiently to improve mental health.
Giorgio Roffo, Dong-Bach Vo, Mohammad Tayarani, Maki Rooksby, Alessandra Sorrentino, Simona Di Folco, Helen Minnis, Stephen A. Brewster, Alessandro Vinciarelli
CHI9
2019 Affective and behavioural computing: Lessons learnt from the First Computational Paralinguistics Challenge
Björn W. Schuller, Felix Weninger, Yue Zhang 0014, Fabien Ringeval, Anton Batliner, Stefan Steidl, Florian Eyben, Erik Marchi, Alessandro Vinciarelli, Klaus R. Scherer, Mohamed Chetouani, Marcello Mortillaro
Comput. Speech Lang.9
2018 Shaping Robot Gestures to Shape Users' Perception: The Effect of Amplitude and Speed on Godspeed Ratings
abstract
This work analyses the relationship between the way robots gesture and the way those gestures are perceived by human users. In particular, this work shows how modifying the amplitude and speed of a gesture affect the Godspeed scores given to those gestures, by means of an experiment involving 45 stimuli and 30 observers. The results suggest that shaping gestures aimed at manifesting the inner state of the robot (e.g., cheering or showing disappointment) tends to change the perception of Animacy (the dimension that accounts for how driven by endogenous factors the robot is perceived to be), while shaping gestures aimed at achieving an interaction effect (e.g., engaging and disengaging) tends to change the perception of Anthropomorphism, Likeability and Perceived Safety (the dimensions that account for the social aspects of the perception).
Amol A. Deshmukh, Bart G. W. Craenen, Alessandro Vinciarelli, Mary Ellen Foster
HAI3
2018 Depression Speaks: Automatic Discrimination between Depressed and Non-Depressed Speakers Based on Nonverbal Speech Features
abstract
This article proposes an automatic approach - based on nonverbal speech features - aimed at the automatic discrimination between depressed and non-depressed speakers. The experiments have been performed over one of the largest corpora collected for such a task in the literature (62 patients diagnosed with depression and 54 healthy control subjects), especially when it comes to data where the depressed speakers have been diagnosed as such by professional psychiatrists. The results show that the discrimination can be performed with an accuracy of over 75% and the error analysis shows that the chances of correct classification do not change according to gender, depression-related pathology diagnosed by the psychiatrists or length of the pharmacological treatment (if any). Furthermore, for every depressed subject, the corpus includes a control subject that matches age, education level and gender. This ensures that the approach actually discriminates between depressed and non depressed speakers and does not simply capture differences resulting from other factors.
Filomena Scibelli, Giorgio Roffo, Mohammad Tayarani, Luca Bartoli, Gaetano De Mattia, Anna Esposito, Alessandro Vinciarelli
ICASSP7
2018 Power Poses Affect Risk Tolerance and Skin Conductance Levels
abstract
Humans are used to express their feelings of selfconfidence/ powerfulness or their distress/sadness through either expansive postures that occupy as much space as possible or closing postures occupying as less space as possible to avoid contact. This conduct suggests that feelings of selfconfidence/ powerfulness or distress/sadness change our body expressions/postures. It can be interesting to assess whether the reverse is also true, i.e. the way we arrange our body at a given moment would affect our feelings. The present research reports an investigation on such argument. To this aim, 50 subjects (25 females) aged between 23 and 31 years were requested to adopt either an expansive (high-powered) or contracted (low-powered) posture for as long as 3 minutes and then asked to bet money in a dice game. The results show that assuming high-power poses favors risk tolerant behaviors and rises feelings of powerfulness. This is not true in the case of low-power postures, which engender a sense of stress, sustained by a significant increase of skin conductance levels. Considerations are made on how to exploit these results for psychotherapy and rehabilitation purposes, as well as, for the implementation of artificial intelligent systems operating as tools for well-being and coaching.
Davide Saggese, Gennaro Cordasco, Mauro N. Maldonato, Nikolaos G. Bourbakis, Alessandro Vinciarelli, Anna Esposito
ICTAI5
2018 Do We Really Like Robots that Match our Personality? The Case of Big-Five Traits, Godspeed Scores and Robotic Gestures
abstract
This work investigates the role of the attraction paradigm - the tendency to associate similarity and attraction in interpersonal relations - in Human-Robot Interaction. The experiment presented here involved 30 human observers who watched and rated 45 robotic gestures in terms of Big-Five personality traits and Godspeed scores. The results show that, for 24 of the 30 observers, there was a statistically significant correlation between the Godspeed scores and the perceived similarity between the robot's personality and their own. However, the association was positive for 15 subjects - meaning that for these there is a similarity-attraction effect - and negative for the other 9 - meaning that for these there is a complementarity-attraction effect. Furthermore, the strength of the effect depends on the particular trait under examination.
Bart G. W. Craenen, Amol A. Deshmukh, Mary Ellen Foster, Alessandro Vinciarelli
RO-MAN4
2018 Shaping Gestures to Shape Personalities: The Relationship Between Gesture Parameters, Attributed Personality Traits and Godspeed Scores
abstract
This work explores the role of personality as a mediation variable between the observable behaviour of a robot - gestures of different energy and spatial extension in the experiments of this work - and the subjective experience of its users as measure by the Godspeed questionnaire. The results show that, at least for some traits, the Big Five personality traits that the users attribute to a robot are predictive of the Godspeed scores, i.e., of the quality of the interaction the users have with the robot. In other words, robots that are attributed different personality traits tend to be perceived differently in relation to the quality of the interaction.
Bart G. W. Craenen, Amol A. Deshmukh, Mary Ellen Foster, Alessandro Vinciarelli
RO-MAN4
2018 The More I Understand it, the Less I Like it: The Relationship Between Understandability and Godspeed Scores for Robotic Gestures
abstract
This work investigates the relationship between the perception that people develop about a robot and the understandability of the gestures the latter displays. The experiments have involved 30 human observers that have rated 45 robotic gestures in terms of the Godspeed dimensions. At the same time, the observers have assigned a score to 10 possible interpretations (the same interpretations for all gestures). The results show that there is a statistically significant correlation between the understandability of the gestures - measured through an information theoretic approach - and all Godspeed scores. However, the correlation is positive in some cases (Anthropomorphism, Animacy and Perceived Intelligence), but negative in others (Perceived Safety and Likeability). In other words, higher understandability is not necessarily associated with more positive perceptions.
Amol A. Deshmukh, Bart G. W. Craenen, Mary Ellen Foster, Alessandro Vinciarelli
RO-MAN4
2017 SAM: The School Attachment Monitor
abstract
Secure Attachment relationships have been shown to minimise social and behavioural problems in children and boosts resilience to risks later on such as antisocial behaviour, heart pathologies, and suicide. Attachment assessment is an expensive and time-consuming process that is not often performed. The School Attachment Monitor (SAM) automates Attachment assessment to support expert assessors. It uses doll-play activities with the dolls augmented with sensors and the child's play recorded with cameras to provide data for assessment. Social signal processing tools are then used to analyse the data and to automatically categorize Attachment patterns. This paper presents the current SAM interactive prototype.
Dong-Bach Vo, Maki Rooksby, Mohammad Tayarani, Rui Huan, Alessandro Vinciarelli, Helen Minnis, Stephen A. Brewster
IDC5
2017 Infinite Latent Feature Selection: A Probabilistic Latent Graph-Based Ranking Approach
abstract
Feature selection is playing an increasingly significant role with respect to many computer vision applications spanning from object recognition to visual object tracking. However, most of the recent solutions in feature selection are not robust across different and heterogeneous set of data. In this paper, we address this issue proposing a robust probabilistic latent graph-based feature selection algorithm that performs the ranking step while considering all the possible subsets of features, as paths on a graph, bypassing the combinatorial problem analytically. An appealing characteristic of the approach is that it aims to discover an abstraction behind low-level sensory data, that is, relevancy. Relevancy is modelled as a latent variable in a PLSA-inspired generative process that allows the investigation of the importance of a feature when injected into an arbitrary set of cues. The proposed method has been tested on ten diverse benchmarks, and compared against eleven state of the art feature selection methods. Results show that the proposed approach attains the highest performance levels across many different scenarios and difficulties, thereby confirming its strong robustness while setting a new state of the art in feature selection domain.
Giorgio Roffo, Simone Melzi, Umberto Castellani, Alessandro Vinciarelli
ICCV4
2017 Evaluating robot facial expressions
abstract
This paper outlines a demonstration of the work carried out in the SoCoRo project investigating how far a neuro-typical population recognises facial expressions on a non-naturalistic robot face that are designed to show approval and disapproval. RFID-tagged objects are presented to an Emys robot head (called Alyx) and Alyx reacts to each with a facial expression. Participants are asked to put the object in a box marked 'Like' or 'Dislike'. This study is being extended to include assessment of participants' Autism Quotient using a validated questionnaire as a step towards using a robot to help train high-functioning adults with an Autism Spectrum Disorder in social signal recognition.
Ruth Aylett, Frank Broz, Ayan Ghosh, Peter E. McKenna, Gnanathusharan Rajendran, Mary Ellen Foster, Giorgio Roffo, Alessandro Vinciarelli
ICMI8
2017 Modulating the non-verbal social signals of a humanoid robot
abstract
In this demonstration we present a repertoire of social signals generated by the humanoid robot Pepper in the context of the EU-funded project MuMMER. The aim of this research is to provide the robot with the expressive capabilities required to interact with people in real-world public spaces such as shopping malls-and being able to control the non-verbal behaviour of such a robot is key to engaging with humans in an effective way. We propose an approach to modulating the non-verbal social signals of the robot based on systematically varying the amplitude and speed of the joint motions and gathering user evaluations of the resulting gestures. We anticipate that the humans' perception of the robot behaviour will be influenced by these modulations
Amol A. Deshmukh, Bart G. W. Craenen, Alessandro Vinciarelli, Mary Ellen Foster
ICMI3
2017 SAM: the school attachment monitor
abstract
Secure Attachment relationships have been shown to minimise social and behavioural problems in children and boosts resilience to risks such as antisocial behaviour, heart pathologies, and suicide later in life. Attachment assessment is an expensive and time-consuming process that is not often performed. The School Attachment Monitor (SAM) automates Attachment assessment to support expert assessors. It uses doll-play activities with the dolls augmented with sensors and the child's play recorded with cameras to provide data for assessment. Social signal processing tools are then used to analyse the data and to automatically categorize Attachment patterns. This paper presents the current SAM interactive prototype.
Dong-Bach Vo, Mohammad Tayarani, Maki Rooksby, Rui Huan, Alessandro Vinciarelli, Helen Minnis, Stephen A. Brewster
ICMI5
2017 Affective Reasoning for Big Social Data Analysis
abstract
This special section focuses on the introduction, presentation, and discussion of novel techniques that further develop and apply affective reasoning tools and techniques for big social data analysis. A key motivation for this special section, in particular, is to explore the adoption of novel affective reasoning frameworks and cognitive learning systems to go beyond a mere word-level analysis of natural language text and provide novel concept-level tools and techniques that allow a more efficient passage from (unstructured) natural language to (structured) machine-processable affective data, in potentially any domain. The selected papers aim to address the wide spectrum of issues related to affective computing research and, hence, better grasp the current limitations and opportunities related to this fast-evolving branch of artificial intelligence. Out of the 29 submissions received, 5 were accepted to appear in the special section. One of the accepted papers underwent 3 rounds of revisions, the rest were revised twice. The papers appearing in this issue are briefly summarized.
Erik Cambria, Amir Hussain 0001, Alessandro Vinciarelli
IEEE Trans. Affect. Comput.3
2017 The Pictures We Like Are Our Image: Continuous Mapping of Favorite Pictures into Self-Assessed and Attributed Personality Traits
abstract
Flickr allows its users to tag the pictures they like as “favorite”. As a result, many users of the popular photo-sharing platform produce galleries of favorite pictures. This article proposes new approaches, based on Computational Aesthetics, capable to infer the personality traits of Flickr users from the galleries above. In particular, the approaches map low-level features extracted from the pictures into numerical scores corresponding to the Big-Five Traits, both self-assessed and attributed. The experiments were performed over 60,000 pictures tagged as favorite by 300 users (the PsychoFlickr Corpus). The results show that it is possible to predict beyond chance both self-assessed and attributed traits. In line with the state-of-the-art of Personality Computing, these latter are predicted with higher effectiveness (correlation up to 0.68 between actual and predicted traits).
Cristina Segalin, Alessandro Perina, Marco Cristani, Alessandro Vinciarelli
IEEE Trans. Affect. Comput.4
2016 Effects of Emotional Visual Scenes on the Ability to Decode Emotional Melodies
abstract
An effective change in Human Computer Interaction requires to account of how communication practices are transformed in different contexts, how users sense the interaction with a machine, and an efficient machine sensitivity in interpreting users' communicative signals, and activities. To this aims, the present paper investigates on whether and how positive and negative visual scenes may alter listeners' ability to decode emotional melodies. Emotional tunes were played alone and with, either positive, or negative, or neutral emotional scenes. Afterword, subjects (8 groups, each of 38 subjects, equally balanced by gender) were asked to decode the emotional feeling aroused by melodies ascribing them either emotional valences (positive, negative, I don't know) or emotional labels (happy, sad, fear, anger, another emotion, I don't know). It was found that dimensional emotional features rather than emotional labels strongly affect cognitive judgements of emotional melodies. Musical emotional information is most effectively retained when the task is to assign labels rather than valence values to melodies. In addition, significant misperception effects are observed when happy or positively judged melodies are concurrently played with negative scenes.
Anna Esposito, Antonietta Maria Esposito, Marilena Esposito, Maria Teresa Riviello, Alessandro Vinciarelli, Nikolaos G. Bourbakis
ICTAI5
2016 Looking Good With Flickr Faves: Gaussian Processes for Finding Difference Makers in Personality Impressions
abstract
Flickr allows its users to generate galleries of "faves", i.e., pictures that they have tagged as favourite. According to recent studies, the faves are predictive of the personality traits that people attribute to Flickr users. This article investigates the phenomenon and shows that faves allow one to predict whether a Flickr user is perceived to be above median or not with respect to each of the Big-Five Traits (accuracy up to 79\% depending on the trait). The classifier - based on Gaussian Processes with a new kernel designed for this work - allows one to identify the visual characteristics of faves that better account for the prediction outcome.
Xiaoyu Xiong, Maurizio Filippone, Alessandro Vinciarelli
ACM Multimedia3
2015 Automatic personality perception: Prediction of trait attribution based on prosodic features extended abstract
abstract
This paper proposes a prosody based approach for Automatic Personality Perception. Social psychology has shown that whenever we listen to a voice for the first time, we spontaneously and unconsciously attribute personality traits to the speaker. The attribution process is not necessarily accurate, but it is important because it shapes our behavior towards others. The experiments of this work are performed over a corpus of 640 speech samples (322 individuals in total) assessed in terms of speaker's personality traits by 11 judges. The results show that it is possible to predict some of the personality traits with accuracy higher than 70%. The effect of different prosodic features has also been analyzed and compared with findings in the psychological literature.
Gelareh Mohammadi, Alessandro Vinciarelli
ACII2
2015 Measuring mimicry in task-oriented conversations: degree of mimicry is related to task difficulty
Vijay Solanki, Alessandro Vinciarelli, Jane Stuart-Smith, Rachel Smith
INTERSPEECH2
2015 A Survey on perceived speaker traits: Personality, likability, pathology, and the first challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001
Comput. Speech Lang.5
2015 Introduction
Björn W. Schuller, Stefan Steidl, Anton Batliner, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son
Comput. Speech Lang.4
2014 An Outline of Opportunities for Multimodal Research
abstract
This paper summarizes the contributions to the Workshop "Roadmapping the Future of Multimodal Interaction Research including Business Opportunities and Challenges". We present major challenges and ideas for making progress in the field of social signal processing and related fields as presented by the contributors of the workshop.
Dirk Heylen, Alessandro Vinciarelli
ICMI2
2014 The SSPNet-Mobile Corpus: Social Signal Processing Over Mobile Phones
Anna Polychroniou, Hugues Salamin, Alessandro Vinciarelli
LREC3
2014 Face-Based Automatic Personality Perception
abstract
Automatic Personality Perception is the task of automatically predicting the personality traits people attribute to others. This work presents experiments where such a task is performed by mapping facial appearance into the Big-Five personality traits, namely Openness, Conscientiousness, Extraversion, Agreeableness and Neuroticism. The experiments are performed over the pictures of the FERET corpus, originally collected for biometrics purposes, for a total of 829 individuals. The results show that it is possible to automatically predict whether a person is perceived to be above or below median with an accuracy close to 70 percent (depending on the trait).
Noura Al Moubayed, Yolanda Vazquez-Alvarez, Alex McKay, Alessandro Vinciarelli
ACM Multimedia4
2014 Predicting Continuous Conflict Perceptionwith Bayesian Gaussian Processes
abstract
Conflict is one of the most important phenomena of social life, but it is still largely neglected by the computing community. This work proposes an approach that detects common conversational social signals (loudness, overlapping speech, etc.) and predicts the conflict level perceived by human observers in continuous, non-categorical terms. The proposed regression approach is fully Bayesian and it adopts automatic relevance determination to identify the social signals that influence most the outcome of the prediction. The experiments are performed over the SSPNet Conflict Corpus, a publicly available collection of 1,430 clips extracted from televised political debates (roughly 12 hours of material for 138 subjects in total). The results show that it is possible to achieve a correlation close to 0.8 between actual and predicted conflict perception.
Samuel Kim, Fabio Valente, Maurizio Filippone, Alessandro Vinciarelli
IEEE Trans. Affect. Comput.4
2014 A Survey of Personality Computing
abstract
Personality is a psychological construct aimed at explaining the wide variety of human behaviors in terms of a few, stable and measurable individual characteristics. In this respect, any technology involving understanding, prediction and synthesis of human behavior is likely to benefit from Personality Computing approaches, i.e. from technologies capable of dealing with human personality. This paper is a survey of such technologies and it aims at providing not only a solid knowledge base about the state-of-the-art, but also a conceptual model underlying the three main problems addressed in the literature, namely Automatic Personality Recognition (inference of the true personality of an individual from behavioral evidence), Automatic Personality Perception (inference of personality others attribute to an individual based on her observable behavior) and Automatic Personality Synthesis (generation of artificial personalities via embodied agents). Furthermore, the article highlights the issues still open in the field and identifies potential application areas.
Alessandro Vinciarelli, Gelareh Mohammadi
IEEE Trans. Affect. Comput.1
2014 More Personality in Personality Computing
abstract
By explicitly describing what has been done in the past, surveys implicitly outline what can (and sometimes should) be done in the future. The insightful commentary by Wright contributes significantly to this latter aspect, especially when it comes to aligning Personality Computing with the latest developments in Personality Science. This response article tries to progress in such a direction by discussing on Wright's suggestions from a computing science point of view.
Alessandro Vinciarelli, Gelareh Mohammadi
IEEE Trans. Affect. Comput.1
2013 Reading between the turns: Statistical modeling for identity recognition and verification in chats
abstract
Identity safekeeping has recently become an important problem for the social web: as a case study, we focus here on instant messaging platforms, proposing novel soft-biometric cues for user recognition and verification. Specifically, we design a set of features encoding effectively how a person converses: since chats are crossbreeds of written text and face-to-face verbal communication, the features inherit equally from textual authorship attribution and conversational analysis of speech. Importantly, our cues ignore completely the semantics of the chat, relying solely on non-verbal aspects, taking care of possible privacy and ethical issues. We apply our approach on a novel dataset of 94 different individuals, whose chat conversations have been recorded for an average period of five months; recognition rate, intended as normalized AUC on CMC curve, is 95.73%, while verification rate amounts to 95.66%, as normalized AUC on ROC curve.
Giorgio Roffo, Cristina Segalin, Alessandro Vinciarelli, Vittorio Murino, Marco Cristani
AVSS3
2013 Who is persuasive?: the role of perceived personality and communication modality in social multimedia
abstract
Persuasive communication is part of everyone's daily life. With the emergence of social websites like YouTube, Facebook and Twitter, persuasive communication is now seen online on a daily basis. This paper explores the effect of multi-modality and perceived personality on persuasiveness of social multimedia content. The experiments are performed over a large corpus of movie review clips from Youtube which is presented to online annotators in three different modalities: only text, only audio and video. The annotators evaluated the persuasiveness of each review across different modalities and judged the personality of the speaker. Our detailed analysis confirmed several research hypotheses designed to study the relationships between persuasion, perceived personality and communicative channel, namely modality. Three hypotheses are designed: the first hypothesis studies the effect of communication modality on persuasion, the second hypothesis examines the correlation between persuasion and personality perception and finally the third hypothesis, derived from the first two hypotheses explores how communication modality influence the personality perception.
Gelareh Mohammadi, Sunghyun Park 0001, Kenji Sagae, Alessandro Vinciarelli, Louis-Philippe Morency
ICMI4
2013 Investigating fine temporal dynamics of prosodic and lexical accommodation
abstract
Conversational interaction is a dynamic activity in which participants engage in the construction of meaning and in establishing and maintaining social relationships. Lexical and prosodic accommodation have been observed in many studies as contributing importantly to these dimensions of social interaction. However, while previous works have considered accommodation mechanisms at global levels (for whole conversations, halves and thirds of conversations), this work investigates their evolution through repeated analysis at time intervals of increasing granularity to analyze the dynamics of alignment in a spoken language corpus. Results show that the levels of both prosodic and lexical accommodation fluctuate several times over the course of a conversation.
Francesca Bonin, Céline De Looze, Sucheta Ghosh, Emer Gilmartin, Carl Vogel, Anna Polychroniou, Hugues Salamin, Alessandro Vinciarelli, Nick Campbell 0001
INTERSPEECH8
2013 Annotation and detection of conflict escalation in Political debates
abstract
Conflict escalation in multi-party conversations refers to an increase \nin the intensity of conflict during conversations. Here we study annotation \nand detection of conflict escalation in broadcast political \ndebates towards a machine-mediated conflict management system. \nIn this regard, we label conflict escalation using crowd-sourced annotations \nand predict it with automatically extracted conversational \nand prosodic features. In particular, to annotate the conflict escalation \nwe deploy two different strategies, i.e., indirect inference and \ndirect assessment; the direct assessment method refers to a way that \nannotators watch and compare two consecutive clips during the annotation \nprocess, while the indirect inference method indicates that \neach clip is independently annotated with respect to the level of \nconflict then the level conflict escalation is inferred by comparing \nannotations of two consecutive clips. Empirical results with 792 \npairs of consecutive clips in classifying three types of conflict escalation, \ni.e., escalation, de-escalation, and constant, show that labels \nfrom direct assessment yield higher classification performance \n(45.3% unweighted accuracy (UA)) than the one from indirect inference \n(39.7% UA), although the annotations from both methods are \nhighly correlated (ρ = 0.74 in continuous values and 63% agreement \nin ternary classes).
Samuel Kim, Fabio Valente, Alessandro Vinciarelli
INTERSPEECH3
2013 The INTERSPEECH 2013 computational paralinguistics challenge: social signals, conflict, emotion, autism
abstract
International audience
Björn W. Schuller, Stefan Steidl, Anton Batliner, Alessandro Vinciarelli, Klaus R. Scherer, Fabien Ringeval, Mohamed Chetouani, Felix Weninger, Florian Eyben, Erik Marchi, Marcello Mortillaro, Hugues Salamin, Anna Polychroniou, Fabio Valente, Samuel Kim
INTERSPEECH4
2013 Unveiling the multimedia unconscious: implicit cognitive processes and multimedia content analysis
abstract
One of the main findings of cognitive sciences is that automatic processes of which we are unaware shape, to a significant extent, our perception of the environment. The phenomenon applies not only to the real world, but also to multimedia data we consume every day. Whenever we look at pictures, watch a video or listen to audio recordings, our conscious attention efforts focus on the observable content, but our cognition spontaneously perceives intentions, beliefs, values, attitudes and other constructs that, while being outside of our conscious awareness, still shape our reactions and behavior. So far, multimedia technologies have neglected such a phenomenon to a large extent. This paper argues that taking into account cognitive effects is possible and it can also improve multimedia approaches. As a supporting proof-of-concept, the paper shows not only that there are visual patterns correlated with the personality traits of 300 Flickr users to a statistically significant extent, but also that the personality traits (both self-assessed and attributed by others) of those users can be inferred from the images these latter post as "favourite".
Marco Cristani, Alessandro Vinciarelli, Cristina Segalin, Alessandro Perina
ACM Multimedia2
2013 Automatic Detection of Laughter and Fillers in Spontaneous Mobile Phone Conversations
abstract
This article presents experiments on automatic detection of laughter and fillers, two of the most important nonverbal behavioral cues observed in spoken conversations. The proposed approach is fully automatic and segments audio recordings captured with mobile phones into four types of interval: laughter, filler, speech and silence. The segmentation methods rely not only on probabilistic sequential models (in particular Hidden Markov Models), but also on Statistical Language Models aimed at estimating the a-priori probability of observing a given sequence of the four classes above. The experiments are speaker independent and performed over a total of 8 hours and 25 minutes of data (120 people in total). The results show that F1scores up to 0.64 for laughter and 0.58 for fillers can be achieved.
Hugues Salamin, Anna Polychroniou, Alessandro Vinciarelli
SMC3
2012 Automatic detection of conflicts in spoken conversations: Ratings and analysis of broadcast political debates
abstract
Automatic analysis of spoken conversations has recently searched for phenomena like agreement/disagreement in collaborative and non-conflictual discussions (e.g., meetings). This work adds a novel dimension investigating conflicts in spontaneous conversations. The study makes use of broadcasted political debates where conflicts naturally arise between participants. In the first part, an annotation scheme to rate the degree of conflict in conversations is described and applied to 12 hours of recordings. In the second part, the correlation between various prosodic/conversational features and the degree of conflict is investigated. In the third part, we perform automatic detection of the level of conflict based on those features showing an F-measure of 71.6% in three-level classification tasks.
Samuel Kim, Fabio Valente, Alessandro Vinciarelli
ICASSP3
2012 On Speaker-Independent Personality Perception and Prediction from Speech
abstract
In this paper, we present ongoing experiments and insights regarding automatic assessment of perceived personality. While within the INTERSPEECH Speaker Trait Challenge participants will train systems in order to recognize binary targets along the Big 5 personality trait, we will analyze and discuss properties of the data, the labeling scheme and the predictive quality. Conducting factor analyses, estimating reliability, and building regression models capturing dimensions of personality we compare all results to our former and current work and introduce a new extension of our personality database. Eventually, this paper contributes in methodology and understanding on how to asses the perceived personality from an unknown speaker by humans and machines.
Tim Polzehl, Katrin Schoenenberg, Sebastian Möller 0001, Florian Metze, Gelareh Mohammadi, Alessandro Vinciarelli
INTERSPEECH6
2012 The INTERSPEECH 2012 Speaker Trait Challenge
abstract
LIDIAP
Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001
INTERSPEECH5
2012 Conversationally-inspired stylometric features for authorship attribution in instant messaging
abstract
Authorship attribution (AA) aims at recognizing automatically the author of a given text sample. Traditionally applied to literary texts, AA faces now the new challenge of recognizing the identity of people involved in chat conversations. These share many aspects with spoken conversations, but AA approaches did not take it into account so far. Hence, this paper tries to fill the gap and proposes two novelties that improve the effectiveness of traditional AA approaches for this type of data: the first is to adopt features inspired by Conversation Analysis (in particular for turn-taking), the second is to extract the features from individual turns rather than from entire conversations. The experiments have been performed over a corpus of dyadic chat conversations (77 individuals in total). The performance in identifying the persons involved in each exchange, measured in terms of area under the Cumulative Match Characteristic curve, is 89.5%.
Marco Cristani, Giorgio Roffo, Cristina Segalin, Loris Bazzani, Alessandro Vinciarelli, Vittorio Murino
ACM Multimedia5
2012 Predicting the conflict level in television political debates: an approach based on crowdsourcing, nonverbal communication and gaussian processes
abstract
One of the most recent trends in multimedia indexing is to represent data in terms of the social and psychological phenomena that users perceive. In such a perspective this article proposes an approach for the automatic detection of conflict level in television political debates. The proposed approach includes the use of crowdsourcing techniques for modeling the perception of data consumers, the extraction of (language independent) nonverbal behavioral cues and the application of regression techniques based on Gaussian Processes. The experiments have been performed over 1430 clips of 30 seconds extracted from 45 political debates (roughly 12 hours of material). The results show that a correlation up to 0.8 can be achieved between the actual and predicted conflict level.
Samuel Kim, Maurizio Filippone, Fabio Valente, Alessandro Vinciarelli
ACM Multimedia4
2012 From speech to personality: mapping voice quality and intonation into personality differences
abstract
From a cognitive point of view, personality perception corresponds to capturing individual differences and can be thought of as positioning the people around us in an ideal personality space. The more similar the personality of two individuals, the closer their position in the space. This work shows that the mutual position of two individuals in the personality space can be inferred from prosodic features. The experiments, based on ordinal regression techniques, have been performed over a corpus of 640 speech samples comprising 322 individuals assessed in terms of personality traits by 11 human judges, which is the largest database of this type in the literature. The results show that the mutual position of two individuals can be predicted with up to 80% accuracy.
Gelareh Mohammadi, Antonio Origlia, Maurizio Filippone, Alessandro Vinciarelli
ACM Multimedia4
2012 Automatic Personality Perception: Prediction of Trait Attribution Based on Prosodic Features
abstract
Whenever we listen to a voice for the first time, we attribute personality traits to the speaker. The process takes place in a few seconds and it is spontaneous and unaware. While the process is not necessarily accurate (attributed traits do not necessarily correspond to the actual traits of the speaker), still it significantly influences our behavior toward others, especially when it comes to social interaction. This paper proposes an approach for the automatic prediction of the traits the listeners attribute to a speaker they never heard before. The experiments are performed over a corpus of 640 speech clips (322 identities in total) annotated in terms of personality traits by 11 assessors. The results show that it is possible to predict with high accuracy (more than 70 percent depending on the particular trait) whether a person is perceived to be in the upper or lower part of the scales corresponding to each of the Big -Five, the personality dimensions known to capture most of the individual differences.
Gelareh Mohammadi, Alessandro Vinciarelli
IEEE Trans. Affect. Comput.2
2012 Bridging the Gap between Social Animal and Unsocial Machine: A Survey of Social Signal Processing
abstract
Social Signal Processing is the research domain aimed at bridging the social intelligence gap between humans and machines. This paper is the first survey of the domain that jointly considers its three major aspects, namely, modeling, analysis, and synthesis of social behavior. Modeling investigates laws and principles underlying social interaction, analysis explores approaches for automatic understanding of social exchanges recorded with different sensors, and synthesis studies techniques for the generation of social behavior via various forms of embodiment. For each of the above aspects, the paper includes an extensive survey of the literature, points to the most important publicly available resources, and outlines the most fundamental challenges ahead.
Alessandro Vinciarelli, Maja Pantic, Dirk Heylen, Catherine Pelachaud, Isabella Poggi, Francesca D'Errico, Marc Schröder 0001
IEEE Trans. Affect. Comput.1
2012 Automatic Role Recognition in Multiparty Conversations: An Approach Based on Turn Organization, Prosody, and Conditional Random Fields
abstract
Roles are a key aspect of social interactions, as they contribute to the overall predictability of social behavior (a necessary requirement to deal effectively with the people around us), and they result in stable, possibly machine-detectable behavioral patterns (a key condition for the application of machine intelligence technologies). This paper proposes an approach for the automatic recognition of roles in conversational broadcast data, in particular, news and talk shows. The approach makes use of behavioral evidence extracted from speaker turns and applies conditional random fields to infer the roles played by different individuals. The experiments are performed over a large amount of broadcast material (around 50 h), and the results show an accuracy higher than 85%.
Hugues Salamin, Alessandro Vinciarelli
IEEE Trans. Multim.2
2011 Language-Independent Socio-Emotional Role Recognition in the AMI Meetings Corpus
abstract
Social roles are a coding scheme that characterizes the relationships between group members during a discussion and their roles “oriented toward the functioning of the group as a group”. This work presents an investigation on language-independent automatic social role recognition in AMI meetings based on turns statistics and prosodic features. At first, turn-taking statistics and prosodic features are integrated into a single generative conversation model which achieves a role recognition accuracy of 59%. This model is then extended to explicitly account for dependencies (or influence) between speakers achieving an accuracy of 65%. The last contribution consists in investigating the statistical dependencies between the formal and the social role that participants have; integrating the information related to the formal role in the model, the recognition achieves an accuracy of 68%.
Fabio Valente, Alessandro Vinciarelli
INTERSPEECH2
2011 Joint ACM workshop on human gesture and behavior understanding: (J-HGBU'11)
abstract
The ability to understand social signals of a person we are communicating with is the core of social intelligence. Social Intelligence is a facet of human intelligence that has been argued to be indispensable and perhaps the most important for success in life. At the same time, human-centric multimedia applications for humans and about humans are becoming increasingly important. 3D modeled human-objects, like bodies, heads and faces are exploited for animation, security, and human computer interaction, while three dimensional motion of arms, legs and local body features is used for more complete human gesture, activity and behavior analysis. The Joint Human Gesture and Behavior Understanding (J-HGBU) workshop event consists of two parts focusing on these complementary challenges: the Workshop on Multimedia Access to 3D Human Objects (MA3HO'11) and the Workshop on Social Signal Processing (SSPW'11).
Maja Pantic, Alex Pentland, Alessandro Vinciarelli, Rita Cucchiara, Mohamed Daoudi, Alberto Del Bimbo
ACM Multimedia3
2011 Humans as feature extractors: Combining prosody and personality perception for improved speaking style recognition
abstract
This paper presents experiments where natural and spontaneous cognitive processes, in particular those who lead to the attribution of personality traits to unacquainted people, are used as a natural form of feature extraction. In particular, personality assessments provided by human judges are used as features to distinguish between professional and non-professional speakers. The same task is performed with prosodic features extracted with a fully automatic process for comparison purposes. Furthermore both prosodic features and personality assessments are combined. The results show that the discrimination between professional and non-professional speaking styles can be performed with an accuracy of 87.2% when using prosodic features, of 75.5% when using personality assessments, and of 90.0% when using the combination of the two.
Gelareh Mohammadi, Alessandro Vinciarelli
SMC2
2011 Recent developments in social signal processing
abstract
Social signal processing has the ambitious goal of bridging the social intelligence gap between computers and humans. Nowadays, computers are not only the new interaction partners of humans, but also a privileged interaction medium for social exchange between humans. Consequently, enhancing machine abilities to interpret and reproduce social signals is a crucial requirement for improving computer-mediated communication and interaction. Furthermore, automated analysis of such signals creates a host of new applications and improvements to existing applications. The study of social signals benefits a wide range of domains, including human-computer interaction, interaction design, entertainment technology, ambient intelligence, health-care, and psychology. This paper briefly introduces the field and surveys its latest developments.
Albert Ali Salah, Maja Pantic, Alessandro Vinciarelli
SMC3
2011 Understanding social signals in multi-party conversations: Automatic recognition of socio-emotional roles in the AMI meeting corpus
abstract
Any social interaction is characterized by roles, patterns of behavior recognized as such by the interacting participants and corresponding to shared expectations that people hold about their own behavior as well as the behavior of others. In this respect, social roles are a key aspect of social interaction because they are the basis for making reasonable guesses about human behavior. Recognizing roles is a crucial need towards understanding (possibly in an automatic way) any social exchange, whether this means to identify dominant individuals, detect conflict, assess engagement or spot conversation highlights. This work presents an investigation on language-independent automatic social role recognition in AMI meetings, spontaneous multi-party conversations, based solely on turn organization and prosodic features. At first turn-taking statistics and prosodic features are integrated into a single generative conversation model which achieves an accuracy of 59%. This model is then extended to explicitly account for dependencies (or influence) between speakers achieving an accuracy of 65%. The last contribution consists in investigating the statistical dependency between the formal and the social role that participants have; integrating the information related to the formal role in the recognition model achieves an accuracy of 68%. The paper is concluded highlighting some future directions.
Alessandro Vinciarelli, Fabio Valente, Sree Harsha Yella, Ashtosh Sapru
SMC1
2010 Mobile social signal processing: vision and research issues
abstract
This paper introduces the First International Workshop on Mobile Social Signal Processing (SSP).The Workshop aims at bringing together the Mobile HCI and Social Signal Processing research communities.The former investigates approaches for effective interaction with mobile and wearable devices, while the latter focuses on modeling, analysis and synthesis of nonverbal behavior in human-human and humanmachine interactions.While dealing with similar problems, the two domains have different goals and methodologies.However, mutual exchange of expertise is likely to raise new research questions as well as to improve approaches in both domains.After providing a brief survey of Mobile HCI and SSP, the paper introduces general aspects of the workshop (including topics, keynote speakers and dissemination means).
Alessandro Vinciarelli, Roderick Murray-Smith, Hervé Bourlard
Mobile HCI1
2010 MM'10 workshop summary for SSPW: ACM workshop on social signal processing 2010
abstract
The Workshop on Social Signal Processing (SSPW) is the yearly event of the Social Signal Processing Network (EU-FP7 SSPNet project). This year's workshop programme consists of 4 premium Key Note Talks by Jeff Cohn, Alex Pentland. Justine Cassell, and Toyoaki Nishida, an oral session with 4 presentations, a poster session with 7 posters, and a panel session where the panelists will be the Key Note Speakers and the workshop organizers.
Maja Pantic, Alessandro Vinciarelli, Alex Pentland
ACM Multimedia2
2010 Automatic role recognition based on conversational and prosodic behaviour
abstract
This paper proposes an approach for the automatic recognition of roles in settings like news and talk-shows, where roles correspond to specific functions like Anchorman, Guest or Interview Participant. The approach is based on purely nonverbal vocal behavioral cues, including who talks when and how much (turn-taking behavior), and statistical properties of pitch, formants, energy and speaking rate (prosodic behavior). The experiments have been performed over a corpus of around 50 hours of broadcast material and the accuracy, percentage of time correctly labeled in terms of role, is up to 89%. Both turn-taking and prosodic behavior lead to satisfactory results. Furthermore, on one database, their combination leads to a statistically significant improvement.
Hugues Salamin, Alessandro Vinciarelli, Khiet P. Truong, Gelareh Mohammadi
ACM Multimedia2
2009 Implicit Human-Centered Tagging
abstract
This paper provides a general introduction to the concept of implicit human-centered tagging (IHCT) - the automatic extraction of tags from nonverbal behavioral feedback of media users. The main idea behind IHCT is that nonverbal behaviors displayed when interacting with multimedia data (e.g., facial expressions, head nods, etc.) provide information useful for improving the tag sets associated with the data. As such behaviors are displayed naturally and spontaneously, no effort is required from the users, and this is why the resulting tagging process is said to be "implicit". Tags obtained through IHCT are expected to be more robust than tags associated with the data explicitly, at least in terms of: generality (they make sense to everybody) and statistical reliability (all tags will be sufficiently represented). The paper discusses these issues in detail and provides an overview of pioneering efforts in the field.
Alessandro Vinciarelli, N. Suditu, Maja Pantic
ICME1
2009 Automatic role recognition in multiparty recordings using social networks and probabilistic sequential models
abstract
The automatic analysis of social interactions is attracting significant interest in the multimedia community. This work addresses one of the most important aspects of the problem, namely the recognition of roles in social exchanges. The proposed approach is based on Social Network Analysis, for the representation of individuals in terms of their interactions with others, and probabilistic sequential models, for the recognition of role sequences underlying the sequence of speakers in conversations. The experiments are performed over different kinds of data (around 90 hours of broadcast data and meetings), and show that the performance depends on how formal the roles are, i.e. on how much they constrain people behavior.
Sarah Favre, Alfred Dielmann, Alessandro Vinciarelli
ACM Multimedia3
2009 Social Computers for the Social Animal: State-of-the-Art and Future Perspectives of Social Signal Processing
Alessandro Vinciarelli
UMAP1
2009 Social signal processing: Survey of an emerging domain
Alessandro Vinciarelli, Maja Pantic, Hervé Bourlard
Image Vis. Comput.1
2009 Automatic Role Recognition in Multiparty Recordings: Using Social Affiliation Networks for Feature Extraction
abstract
Automatic analysis of social interactions attracts increasing attention in the multimedia community. This letter considers one of the most important aspects of the problem, namely the roles played by individuals interacting in different settings. In particular, this work proposes an automatic approach for the recognition of roles in both production environment contexts (e.g., news and talk-shows) and spontaneous situations (e.g., meetings). The experiments are performed over roughly 90 h of material (one of the largest databases used for role recognition in the literature) and show that the recognition effectiveness depends on how much the roles influence the behavior of people. Furthermore, this work proposes the first approach for modeling mutual dependences between roles and assesses its effect on role recognition performance.
Hugues Salamin, Sarah Favre, Alessandro Vinciarelli
IEEE Trans. Multim.3
2008 Role recognition in multiparty recordings using social affiliation networks and discrete distributions
abstract
This paper presents an approach for the recognition of roles in multiparty recordings. The approach includes two major stages: extraction of Social Affiliation Networks (speaker diarization and representation of people in terms of their social interactions), and role recognition (application of discrete probability distributions to map people into roles). The experiments are performed over several corpora, including broadcast data and meeting recordings, for a total of roughly 90 hours of material. The results are satisfactory for the broadcast data (around 80 percent of the data time correctly labeled in terms of role), while they still must be improved in the case of the meeting recordings (around 45 percent of the data time correctly labeled). In both cases, the approach outperforms significantly chance.
Sarah Favre, Hugues Salamin, John Dines, Alessandro Vinciarelli
ICMI4
2008 Social signals, their function, and automatic analysis: a survey
abstract
Social Signal Processing (SSP) aims at the analysis of social behaviour in both Human-Human and Human-Computer interactions. SSP revolves around automatic sensing and interpretation of social signals, complex aggregates of nonverbal behaviours through which individuals express their attitudes towards other human (and virtual) participants in the current social context. As such, SSP integrates both engineering (speech analysis, computer vision, etc.) and human sciences (social psychology, anthropology, etc.) as it requires multimodal and multidisciplinary approaches. As of today, SSP is still in its early infancy, but the domain is quickly developing, and a growing number of works is appearing in the literature. This paper provides an introduction to nonverbal behaviour involved in social signals and a survey of the main results obtained so far in SSP. It also outlines possibilities and challenges that SSP is expected to face in the next years if it is to reach its full maturity.
Alessandro Vinciarelli, Maja Pantic, Hervé Bourlard, Alex Pentland
ICMI1
2008 Role recognition for meeting participants: an approach based on lexical information and social network analysis
abstract
This paper presents experiments on the automatic recognition of roles in meetings. The proposed approach combines two sources of information: the lexical choices made by people playing different roles on one hand, and the Social Networks describing the interactions between the meeting participants on the other hand. Both sources lead to role recognition results significantly higher than chance when used separately, but the best results are obtained with their combination. Preliminary experiments obtained over a corpus of 138 meeting recordings (over 45 hours of material) show that around 70% of the time is labeled correctly in terms of role.
Neha P. Garg, Sarah Favre, Hugues Salamin, Dilek Hakkani-Tür, Alessandro Vinciarelli
ACM Multimedia5
2008 Social signal processing: state-of-the-art and future perspectives of an emerging domain
abstract
The ability to understand and manage social signals of a person we are communicating with is the core of social intelligence. Social intelligence is a facet of human intelligence that has been argued to be indispensable and perhaps the most important for success in life. This paper argues that next-generation computing needs to include the essence of social intelligence - the ability to recognize human social signals and social behaviours like politeness, and disagreement - in order to become more effective and more efficient. Although each one of us understands the importance of social signals in everyday life situations, and in spite of recent advances in machine analysis of relevant behavioural cues like blinks, smiles, crossed arms, laughter, and similar, design and development of automated systems for Social Signal Processing (SSP) are rather difficult. This paper surveys the past efforts in solving these problems by a computer, it summarizes the relevant findings in social psychology, and it proposes a set of recommendations for enabling the development of the next generation of socially-aware computing.
Alessandro Vinciarelli, Maja Pantic, Hervé Bourlard, Alex Pentland
ACM Multimedia1
2007 Role Recognition in Broadcast News using Bernoulli Distributions
abstract
This work presents an approach for the recognition of the roles played by speakers participating in radio broadcast news (e.g. anchorman or guest). The approach includes two main stages: the first is the split of the news recordings into single speaker segments using an unsupervised approach. The second is the application of Bernoulli Distributions for role modeling and recognition. The experiments are performed over a collection of 96 news bulletins (around 19 hours of material) and show that around 80 percent of the data time is labeled correctly in terms of role.
Alessandro Vinciarelli
ICME1
2007 Semantic Segmentation of Radio Programs using Social Network Analysis and Duration Distribution Modeling
abstract
This work presents and compare two approaches for the semantic segmentation of broadcast news: the first is based on social network analysis, the second is based on Poisson stochastic processes. The experiments are performed over 27 hours of material: preliminary results are obtained by addressing the problem of splitting different episodes of the same program into two parts corresponding to a news bulletin and a talk-show respectively. The results show that the transition point between the two parts can be detected with an average error of around three minutes, i.e. roughly 5 percent of each episode duration.
Alessandro Vinciarelli, Sarah Favre
ICME1
2007 Broadcast news story segmentation using social network analysis and hidden markov models
abstract
This paper presents an approach for the segmentation of broadcast news into stories. The main novelty of this work is that the segmentation process does not take into account the content of the news, i.e. what is said, but rather the structure of the social relationships between the persons that in the news are involved. The main rationale behind such an approach is that people interacting with each other are likely to talk about the same topics, thus social relationships are likely to be correlated to stories. The approach is based on Social Network Analysis (for the representation of social relationships) and Hidden Markov Models (for the mapping of social relationships into stories). The experiments are performed over 26 hours of radio news and the results show that a fully automatic process achieves a purity higher than 0.75.
Alessandro Vinciarelli, Sarah Favre
ACM Multimedia1
2007 Speakers Role Recognition in Multiparty Audio Recordings Using Social Network Analysis and Duration Distribution Modeling
abstract
This paper presents two approaches for speaker role recognition in multiparty audio recordings. The experiments are performed over a corpus of 96 radio bulletins corresponding to roughly 19 h of material. Each recording involves, on average, 11 speakers playing one among six roles belonging to a predefined set. Both proposed approaches start by segmenting automatically the recordings into single speaker segments, but perform role recognition using different techniques. The first approach is based on Social Network Analysis, the second relies on the intervention duration distribution across different speakers. The two approaches are used separately and combined and the results show that around 85% of the recording time can be labeled correctly in terms of role.
Alessandro Vinciarelli
IEEE Trans. Multim.1
2006 Sociometry based Multiparty Audio Recordings Segmentation
abstract
This paper shows how social network analysis, the sociological domain studying the interaction between people in specific social environments, can be used to assign roles to different speakers in multiparty recordings. The experiments presented in this work focus on radio news recordings involving around 11 speakers on average. Each of them is assigned automatically a role (e.g. anchorman or guest) without using any information related to their identity or the amount of time they talk. The results (obtained over 96 recordings for a total of around 19 hours) show that more than 85% of the recording time is correctly labeled in terms of role
Alessandro Vinciarelli
ICME1
2006 Application of Information Retrieval Technologies to Presentation Slides
abstract
Presentations are becoming an increasingly more common means of communication in working environments, and slides are often the necessary supporting material on which the presentations rely. In this paper, we describe a slide indexing and retrieval system in which the slides are captured as images (through a framegrabber) at the moment they are displayed during a presentation and then transcribed with an optical character recognition (OCR) system. In this context, we show that such an approach presents several advantages over the use of commercial software (API based) to obtain the slide transcriptions. We report a set of retrieval experiments conducted on a database of 26 real presentations (570 slides) collected at a workshop. The experiments show that the overall retrieval performance is close to that obtained using either a manual transcription of the slides or the API software. Moreover, the experiments show that the OCR-based approach outperforms significantly the API in extracting the text embedded in images and figures
Alessandro Vinciarelli, Jean-Marc Odobez
IEEE Trans. Multim.1
2005 OCR Based Slide Retrieval
abstract
This paper addresses the problem of acquiring, indexing and retrieving slides in the context of automatic oral presentation processing. Since the most suitable acquisition technique, in such a context, is the use of a framegrabber (a device capturing as images the slides displayed on a screen), the slides must be transcribed with an optical character recognition system. Retrieval experiments performed on a corpus of 570 slides (26 presentations) gathered at a workshop show that performance obtained with the OCR transcriptions are close to those obtained by extracting the text from the electronic version (pdf or ppt) of the slides (through apposite APIs).
N. Daddaoua, Jean-Marc Odobez, Alessandro Vinciarelli
ICDAR3
2005 Effect of segmentation method on video retrieval performance
abstract
This paper presents experiments that evaluate the effect of different video segmentation methods on text-based video retrieval. Segmentations relying on modalities like speech, video and text or their combination are compared with a baseline sliding window segmentation. The results suggest that even with the sliding window segmentation, acceptable performance can be obtained on a broadcast news retrieval task. Moreover, in the case where manually segmented data are available for training, the approach combining the different modalities can lead to IR results close to those obtained with a manual segmentation.
David Grangier, Alessandro Vinciarelli
ICME2
2005 Noisy Text Categorization
abstract
This work presents categorization experiments performed over noisy texts. By noisy, we mean any text obtained through an extraction process (affected by errors) from media other than digital texts (e.g., transcriptions of speech recordings extracted with a recognition system). The performance of a categorization system over the clean and noisy (Word Error Rate between approximately 10 and approximately 50 percent) versions of the same documents is compared. The noisy texts are obtained through handwriting recognition and simulation of optical character recognition. The results show that the performance loss is acceptable for Recall values up to 60-70 percent depending on the noise sources. New measures of the extraction process performance, allowing a better explanation of the categorization results, are proposed.
Alessandro Vinciarelli
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Application of information retrieval techniques to single writer documents
Alessandro Vinciarelli
Pattern Recognit. Lett.1
2004 Offline Recognition of Unconstrained Handwritten Texts Using HMMs and Statistical Language Models
abstract
This paper presents a system for the offline recognition of large vocabulary unconstrained handwritten texts. The only assumption made about the data is that it is written in English. This allows the application of Statistical Language Models in order to improve the performance of our system. Several experiments have been performed using both single and multiple writer data. Lexica of variable size (from 10,000 to 50,000 words) have been used. The use of language models is shown to improve the accuracy of the system (when the lexicon contains 50,000 words, the error rate is reduced by approximately 50 percent for single writer data and by approximately 25 percent for multiple writer data). Our approach is described in detail and compared with other methods presented in the literature to deal with the same problem. An experimental setup to correctly deal with unconstrained text recognition is proposed.
Alessandro Vinciarelli, Samy Bengio, Horst Bunke
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Markov Model Document Retrieval
abstract
This paper presents a new probabilistic approachto document retrieval based on the assumption thata Markov process can explain the process by whichhumans rank the relevance of documents to queries.The model ranks documents for retrieval based on theirprobability of relevance. Two training methods are presented. The model is compared with Latent Semantic Analysis (LSA) on two publicly available databases.The results show that the new algorithm achieves Precision/Recall performance equivalent to or better than LSA.
Michael Perrone, Alessandro Vinciarelli
ICDAR2
2003 Offline Recognition of Large Vocabulary Cursive Handwritten Text
abstract
This paper presents a system for the offline recognition of cursive handwritten lines of text. The system is based on continuous density HMMs and Statistical Language Models. The system recognizes data produced by a single writer. No a-priori knowledge is used about the content of the text to be recognized. Changes in the experimental setup with respect to the recognition of single words are highlighted. The results show a recognition rate of #85% with a lexicon containing 50'000 words. The experiments were performed over a publicly available database.
Alessandro Vinciarelli, Samy Bengio, Horst Bunke
ICDAR1
2003 Combining Online and Offline Handwriting Recognition
abstract
This work describes an online handwriting recognition system working in combination with an offline recognizer. The online input data is first transcribed by theonline recognizer, then converted into an offline bitmapand recognized by the offline system. The outputs ofthe two recognizers are then combined probabilisticallyresulting in a classifier out-performing both individualsystems. Experiments were performed over a databaseof single digits. The error rate of the online recognizeris reduced by 43% when the combination with the offlinesystem is applied.
Alessandro Vinciarelli, Michael Perrone
ICDAR1
2003 Combining neural gas and learning vector quantization for cursive character recognition
Francesco Camastra, Alessandro Vinciarelli
Neurocomputing2
2002 Estimating the Intrinsic Dimension of Data with a Fractal-Based Method
abstract
In this paper, the problem of estimating the intrinsic dimension of a data set is investigated. A fractal-based approach using the Grassberger-Procaccia algorithm is proposed. Since the Grassberger-Procaccia algorithm (1983) performs badly on sets of high dimensionality, an empirical procedure that improves the original algorithm has been developed. The procedure has been tested on data sets of known dimensionality and on time series of Santa Fe competition.
Francesco Camastra, Alessandro Vinciarelli
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 A survey on off-line Cursive Word Recognition
Alessandro Vinciarelli
Pattern Recognit.1
2002 Writer adaptation techniques in HMM based Off-Line Cursive Script Recognition
Alessandro Vinciarelli, Samy Bengio
Pattern Recognit. Lett.1
2001 Intrinsic Dimension Estimation of Data: An Approach Based on Grassberger-Procaccia's Algorithm
Francesco Camastra, Alessandro Vinciarelli
Neural Process. Lett.2
2001 Cursive character recognition by learning vector quantization
Francesco Camastra, Alessandro Vinciarelli
Pattern Recognit. Lett.2
2001 A new normalization technique for cursive handwritten words
Alessandro Vinciarelli, Jürgen Lüttin
Pattern Recognit. Lett.1