EDBT 2026 Demo / reviewers in the wild / expert
Daniel Gatica-Perez
dblp:55/1328 · also Daniel Gática-Pérez
· DBLP profile ↗
179ranked-venue papers
22as first author
23since 2021 · last 2026
0000-0001-5488-2182ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 78 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 77 · 14 first-author · 5 since 2021Artificial intelligence and machine learning · 32 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 11 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Computer networks · 5 · 1 first-authorSecurity and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building A Civic Tool for Community-Police Engagement to Adapt Neighborhood PolicingabstractData-driven policing often prioritizes incident records over residents’ lived experiences. In the Baltic city of Riga, with a history of distrust and limited community-police engagement, this can further alienate the public. To bridge this gap, we propose a Research through Design (RtD) inquiry into the development of Par drošu Rīgu, a civic tool for community-data-integrated policing. With municipal police, NGOs, and city staff, we ask how RtD enables stakeholder negotiation and which interaction qualities support trust and the use of combined community and incident data. The co-design process included workshops that surfaced divergent notions of safety; material probes designed as boundary objects to negotiate among stakeholders; and a pilot deployment showing how combining quantitative and qualitative data reshapes engagement and trust. Mixed-methods evaluation suggests increased officer-citizen interaction, but frictions in sustaining stakeholder collaboration. We contribute (i) an empirical RtD inquiry with public institutions, (ii) an artifact combining physical and dashboard interactions, and (iii) reflections on interaction design as a boundary-spanning practice for trust and infrastructuring. Ravinithesh Annapureddy, Stanislavs Seiko, Natalie Higham-James, William Droz, Alessandro Fornaroli, Sarah Vollmer, Britta Elena Hecking, Daniel Gatica-Perez |
DIS | 8 |
| 2026 | Framing Migration News with LLMs: Structured CoT as a Support for Human InterpretationabstractFrame analysis of migration news is a socially consequential task: media scholars and researchers who study how migration is narrated need tools that are not only accurate, but transparent, auditable, and accessible within the resource constraints typical of academic research groups. Existing LLM-based approaches rely on proprietary APIs and large models that raise concerns about data privacy, reproducibility and equitable access among media researchers. This work studies how a locally deployable open-source LLM can support interpretable frame analysis as an assistive tool. We introduce a Structured Chain-of-Thought (SCoT) prompting approach using Llama3-8B, enabling step-by-step justifications grounded in predefined framing categories. This structured design allows users to audit model outputs and examine alternative interpretations in a task that is inherently subjective. We evaluate our approach on a dataset of migration-related news and show that SCoT improves classification performance over zero-shot and few-shot baselines while remaining feasible on a single GPU. Then, we conduct a human-centered evaluation in which annotators assess the coherence and influence of “the model’s reasoning”. Results indicate that SCoT explanations are generally perceived as logical (mean score 4.1/5, though with notable variation across texts) and can prompt reflection on initial interpretations, even when disagreement persists. Our findings highlight both the potential and risks of LLM-assisted frame analysis. While structured reasoning can increase the traceability of model outputs and support critical interpretation, it can also influence human judgment in subtle ways. By enabling local deployment and emphasizing human-in-the-loop interaction, this work contributes to discussions on responsible and accessible computational tools for the study of socially impactful media narratives. David Alonso del Barrio, Daniel Gatica-Perez |
COMPASS | 3 |
| 2026 | Migrant Voices, Local News: Insights on Bridging Community Needs with Media ContentabstractResearch shows news consumption differs across demographics, yet little is known about non-mainstream audiences, especially in relation to local media. Our study addresses this gap by examining how French-speaking migrants in a mid-size European city engage with local news, and whether their needs are reflected in coverage. Eight community members participated in focus groups, whose insights guided the selection of natural language processing methods (topic modeling, information retrieval, sentiment analysis, and readability) applied to over 2000 hyper-local news articles. Results showed that while articles frequently covered local events, gaps remained in topics important to participants. Sentiment analysis revealed a generally positive tone, and readability measures indicated an intermediate-advanced French level, raising questions about accessibility for integration. Our work contributes to bridging the gap between local news platforms' content and diverse readers' needs, and could inform local media organizations about opportunities to expand their current news story coverage to appeal to more diverse audiences. David Alonso del Barrio, Paula Dolores Rescala, Victor Bros, Daniel Gatica-Perez |
IMX | 4 |
| 2025 | The Suisse Romande Local News DatasetabstractThis paper introduces a comprehensive dataset of news articles sourced from ESH Médias, a prominent local press agency in Romandy, the French-speaking region of Switzerland. The dataset encompasses all articles published on their digital platforms from January 2015 through June 2022. With over 130,000 articles written in French, this dataset offers a rich insight into local news from the French-speaking cantons of Switzerland. The articles cover a diverse range of topics and provide valuable material for Natural Language Processing and media studies. To respect privacy and legal considerations, journalists' names have been anonymized, and the dataset is made available for research purposes under a specific agreement with ESH Médias. The dataset adheres to the FAIR principles, and a detailed datasheet is provided to facilitate its use. The dataset is accessible via a DOI link. Victor Bros, Daniel Gatica-Perez |
ICWSM | 2 |
| 2025 | Inferring Mood-While-Eating with Smartphone Sensing and Community-Based Model PersonalizationabstractThe interplay between mood and eating episodes has been extensively researched within the fields of nutrition, psychology, and behavioral science, revealing a connection between the two. Previous studies have relied on questionnaires and mobile phone self-reports to investigate the relationship between mood and eating. In more recent work, phone sensor data has been utilized to characterize both eating behavior and mood independently, particularly in the context of mobile food diaries and mobile health applications. However, current literature exhibits several limitations: a lack of investigation into the generalization of mood inference models trained with data from various everyday life situations to specific contexts like eating; an absence of studies using sensor data to explore the intersection of mood and eating; and inadequate examination of model personalization techniques within limited label settings, a common challenge in mood inference (i.e., far fewer negative mood reports compared to positive or neutral reports). In this study, we examined the everyday eating and mood using two separate datasets from two different studies: (i) Mexico (N \({}_{MEX}\) = 84, 1,843 mood-while-eating reports with a label distribution of positive: 51.7%, neutral: 38.6%, and negative: 9.8%) in 2019, and (ii) eight countries (N \({}_{MUL}\) = 678, 329K mood reports, including 24K mood-while-eating reports with a label distribution of positive: 83%, neutral: 14.9%, and negative: 2.2%) in 2020, which contain both passive smartphone sensing and self-report data. Our results indicate that generic mood inference models experience a decline in performance in specific contexts, such as during eating, highlighting the issue of sub-context shifts in mobile sensing. Moreover, we discovered that population-level (non-personalized) and hybrid (partially personalized) modeling techniques fall short in the commonly used three-class mood inference task (positive, neutral, negative). Additionally, we found that user-level modeling posed challenges for the majority of participants due to insufficient labels and data in the negative class. To overcome these limitations, we implemented a novel community-based personalization approach, building models with data from a set of users similar to the target user. Our findings demonstrate that mood-while-eating can be inferred with accuracies 63.8% (with F1 score of 62.5) for the MEX dataset and 88.3% (with F1 score of 85.7) with the MUL dataset using community-based models, surpassing those achieved with traditional methods. Wageesha Bangamuarachchi, Anju Chamantha, Lakmal Meegahapola, Haeeun Kim, Salvador Ruiz-Correa, Indika Perera, Daniel Gatica-Perez |
ACM Trans. Comput. Heal. | 7 |
| 2024 | Learning About Social Context From Smartphone Data: Generalization Across Countries and Daily Life MomentsabstractUnderstanding how social situations unfold in people’s daily lives is relevant to designing mobile systems that can support users in their personal goals, well-being, and activities. As an alternative to questionnaires, some studies have used passively collected smartphone sensor data to infer social context (i.e., being alone or not) with machine learning models. However, the few existing studies have focused on specific daily life occasions and limited geographic cohorts in one or two countries. This limits the understanding of how inference models work in terms of generalization to everyday life occasions and multiple countries. In this paper, we used a novel, large-scale, and multimodal smartphone sensing dataset with over 216K self-reports collected from 581 young adults in five countries (Mongolia, Italy, Denmark, UK, Paraguay), first to understand whether social context inference is feasible with sensor data, and then, to know how behavioral and country-level diversity affects inferences. We found that several sensors are informative of social context, that partially personalized multi-country models (trained and tested with data from all countries) and country-specific models (trained and tested within countries) can achieve similar performance above 90% AUC, and that models do not generalize well to unseen countries regardless of geographic proximity. These findings confirm the importance of the diversity of mobile data, to better understand social context inference models in different countries. Aurel Ruben Mäder, Lakmal Meegahapola, Daniel Gatica-Perez |
CHI | 3 |
| 2024 | Human Interest or Conflict? Leveraging LLMs for Automated Framing Analysis in TV ShowsabstractIn the current media landscape, understanding the framing of information is crucial for critical consumption and informed decision making. Framing analysis is a valuable tool for identifying the underlying perspectives used to present information, and has been applied to a variety of media formats, including television programs. However, manual analysis of framing can be time-consuming and labor-intensive. This is where large language models (LLMs) can play a key role. In this paper, we propose a novel approach to use prompt-engineering to identify the framing of spoken content in television programs. Our findings indicate that prompt-engineering LLMs can be used as a support tool to identify frames, with agreement rates between human and machine reaching up to 43%. As LLMs are still under development, we believe that our approach has the potential to be refined and further improved. The potential of this technology for interactive media applications is vast, including the development of support tools for journalists, educational resources for students of journalism learning about framing and related concepts, and interactive media experiences for audiences. David Alonso del Barrio, Max Tiel, Daniel Gatica-Perez |
IMX | 3 |
| 2024 | ProGAP: Progressive Graph Neural Networks with Differential Privacy GuaranteesabstractGraph Neural Networks (GNNs) have become a popular tool for learning on graphs, but their widespread use raises privacy concerns as graph data can contain personal or sensitive information. Differentially private GNN models have been recently proposed to preserve privacy while still allowing for effective learning over graph-structured datasets. However, achieving an ideal balance between accuracy and privacy in GNNs remains challenging due to the intrinsic structural connectivity of graphs. In this paper, we propose a new differentially private GNN called ProGAP that uses a progressive training scheme to improve such accuracy-privacy trade-offs. Combined with the aggregation perturbation technique to ensure differential privacy, ProGAP splits a GNN into a sequence of overlapping submodels that are trained progressively, expanding from the first submodel to the complete model. Specifically, each submodel is trained over the privately aggregated node embeddings learned and cached by the previous submodels, leading to an increased expressive power compared to previous approaches while limiting the incurred privacy costs. We formally prove that ProGAP ensures edge-level and node-level privacy guarantees for both training and inference stages, and evaluate its performance on benchmark graph datasets. Experimental results demonstrate that ProGAP can achieve up to 5-10% higher accuracy than existing state-of-the-art differentially private GNNs. Our code is available at https://github.com/sisaman/ProGAP. Sina Sajadmanesh, Daniel Gatica-Perez |
WSDM | 2 |
| 2023 | Keep Sensors in Check: Disentangling Country-Level Generalization Issues in Mobile Sensor-Based Models with Diversity ScoresabstractMachine learning models trained with passive sensor data from mobile devices can be used to perform various inferences pertaining to activity recognition, context awareness, and health and well-being. Prior work has improved inference performance through the use of multimodal sensors (inertial, GPS, proximity, app usage, etc.) or improved machine learning. In this context, a few studies shed light on critical issues relating to the poor cross-country generalization of models due to distributional shifts across countries. However, these studies have largely relied on inference performance as a means of studying generalization issues, failing to investigate whether the root cause of the problem is linked to specific sensor modalities (independent variables) or the target attribute (dependent variable). In this paper, we study this issue in complex activities of daily living (ADL) inference task, involving 12 classes, by using a multimodal, multi-country dataset collected from 689 participants across eight countries. We first show that the ‘country of origin’ of data is captured by sensors and can be inferred from each modality separately, with an average accuracy of 65%. We then propose two diversity scores (DS) that measure how a country differentiates from others w.r.t. sensor modalities or activities. Using these diversity scores, we observed that both individual sensor modalities and activities have the ability to differentiate countries. However, while many activities capture country differences, only the ‘App usage’ and ‘Location’ sensors can do so. By dissecting country-level diversity across dependent and independent variables, we provide a framework to better understand model generalization issues across countries and country-level diversity of sensing modalities. Alexandre Nanchen, Lakmal Meegahapola, William Droz, Daniel Gatica-Perez |
AIES | 4 |
| 2023 | Complex Daily Activities, Country-Level Diversity, and Smartphone Sensing: A Study in Denmark, Italy, Mongolia, Paraguay, and UKabstractSmartphones enable understanding human behavior with activity recognition to support people’s daily lives. Prior studies focused on using inertial sensors to detect simple activities (sitting, walking, running, etc.) and were mostly conducted in homogeneous populations within a country. However, people are more sedentary in the post-pandemic world with the prevalence of remote/hybrid work/study settings, making detecting simple activities less meaningful for context-aware applications. Hence, the understanding of (i) how multimodal smartphone sensors and machine learning models could be used to detect complex daily activities that can better inform about people’s daily lives, and (ii) how models generalize to unseen countries, is limited. We analyzed in-the-wild smartphone data and ∼ 216K self-reports from 637 college students in five countries (Italy, Mongolia, UK, Denmark, Paraguay). Then, we defined a 12-class complex daily activity recognition task and evaluated the performance with different approaches. We found that even though the generic multi-country approach provided an AUROC of 0.70, the country-specific approach performed better with AUROC scores in [0.79-0.89]. We believe that research along the lines of diversity awareness is fundamental for advancing human behavior understanding through smartphones and machine learning, for more real-world utility across countries. Karim Assi, Lakmal Meegahapola, William Droz, Peter Kun, Amalia de Götzen, Miriam Bidoglia, Sally Stares, George Gaskell, Altangerel Chagnaa, Amarsanaa Ganbold, Tsolmon Zundui, Carlo Caprini, Daniele Miorandi, José Luis Zarza, Alethia Hume, Luca Cernuzzi, Ivano Bison, Marcelo Dario Rodas Britez, Matteo Busso, Ronald Chenu, Fausto Giunchiglia, Daniel Gatica-Perez |
CHI | 22 |
| 2023 | Framing the News: From Human Perception to Large Language Model InferencesabstractIdentifying the frames of news is important to understand the articles’ vision, intention, message to be conveyed, and which aspects of the news are emphasized. Framing is a widely studied concept in journalism, and has emerged as a new topic in computing, with the potential to automate processes and facilitate the work of journalism professionals. In this paper, we study this issue with articles related to the Covid-19 anti-vaccine movement. First, to understand the perspectives used to treat this theme, we developed a protocol for human labeling of frames for 1786 headlines of No-Vax movement articles of European newspapers from 5 countries. Headlines are key units in the written press, and worth of analysis as many people only read headlines (or use them to guide their decision for further reading.) Second, considering advances in Natural Language Processing (NLP) with large language models, we investigated two approaches for frame inference of news headlines: first with a GPT-3.5 fine-tuning approach, and second with GPT-3.5 prompt-engineering. Our work contributes to the study and analysis of the performance that these models have to facilitate journalistic tasks like classification of frames, while understanding whether the models are able to replicate human perception in the identification of these frames. David Alonso del Barrio, Daniel Gatica-Perez |
ICMR | 2 |
| 2023 | Referencing in YouTube Knowledge Communication VideosabstractIn recent years, there has been widespread concern about misinformation and hateful content on social media that are damaging societies. Being one of the most influential social media that practically serves as a new search engine, YouTube has accepted criticisms of being a major conduit of misinformation. However, it is often neglected that there exist communities on YouTube that aim to produce credible and informative content - usually falling under the educational category. One way to characterize this valuable content is to find references entailed to each video. While such citation practices function as a voluntary gatekeeping culture within the community, how they are actually done varies and remains unquestioned. Our study aims to investigate common citation practices in major knowledge communication channels on YouTube. After investigating 44 videos manually sampled from YouTube, we characterized two common referencing methods, namely bibliographies and in-video citations. We then selected 129 referenced resources, assessed and categorized their availability as being immediate, conditional, and absent. After relating the observed referencing methods to the characteristics of the knowledge communication community, we show that the usability of references could vary depending on viewers’ user profiles. Furthermore, we witnessed the use of rich-text technologies that can enrich the usability of online video resources. Finally, we discuss design implications for the platform to have a standardized referencing convention that can promote information credibility and improve user experience, especially valuable for the young audiences who tend to watch this content. Haeeun Kim, Daniel Gatica-Perez |
IMX | 2 |
| 2023 | GAP: Differentially Private Graph Neural Networks with Aggregation Perturbation
Sina Sajadmanesh, Ali Shahin Shamsabadi, Aurélien Bellet, Daniel Gatica-Perez |
USENIX Security Symposium | 4 |
| 2022 | Focus on People: Five Questions from Human-Centered ComputingabstractA substantial body of research in multimodal interaction has studied how people naturally interact –face-to-face and through machines– and developed technology to analyze, support, and extend such forms of interaction. The talk will share personal experiences and views on how audio-visual and ubiquitous research on social interaction has evolved over the past two decades. Five recurrent questions, then and now, include how to study interaction in everyday life; how to learn from and collaborate with the humanities and social sciences; how to think about data; how to address the challenges brought by automation; and how to engage and empower individuals and communities to take part in research projects. Today, the limitations of technology-centric solutions are more evident than ever. Future research with a people-first focus will continue to call for reflection, commitment, and action for a long-term alignment with societal needs and nature’s limits. Daniel Gatica-Perez |
ICMI | 1 |
| 2022 | Health Talk: Understanding Practices of Popular Professional YouTubersabstractPractices related to health are circulated widely on YouTube. With a health psychology perspective, we present a study to understand health and wellbeing-related practices of a group of popular, professional YouTubers from the audio-visual content they produce. We first identify, via polytextual thematic analysis, six thematic health-related categories, and use them to label a set of 2500 YouTube videos. Agreement among three independent annotators was acceptable for these health-related categories. We then present an analysis of speech transcriptions and visual content, demonstrating that distinctive patterns exist for these health-related categories. These include linguistic markers and specific scene types and objects. Finally, with an interpretability focus, we study the feasibility of classifying health-related video categories in a binary setting, and compare performance across features, finding best accuracy for linguistic features (74-87%), and various patterns of linguistic and visual relevance used for the classification of health categories. The results shows promise to support mixed-methods research in health psychology, combining manual analysis and data-driven methods. More generally, our work contributes to the understanding of current health practices shared and promoted on social video. Thanh-Trung Phan, Chloé Michoud, Lucia Volpato, María del Río Carral, Daniel Gatica-Perez |
MUM | 5 |
| 2022 | A Sensor-Driven Visit Detection System in Older Adults' Homes: Towards Digital Late-Life Depression Marker ExtractionabstractModern sensor technology is increasingly used in older adults to not only provide additional safety but also to monitor health status, often by means of sensor derived digital measures or biomarkers. Social isolation is a known risk factor for late-life depression, and a potential component of social-isolation is the lack of home visits. Therefore, home visits may serve as a digital measure for social isolation and late-life depression. Late-life depression is a common mental and emotional disorder in the growing population of older adults. The disorder, if untreated, can significantly decrease quality of life and, amongst other effects, leads to increased mortality. Late-life depression often goes undiagnosed due to associated stigma and the incorrect assumption that it is a normal part of ageing. In this work, we propose a visit detection system that generalizes well to previously unseen apartments - which may differ largely in layout, sensor placement, and size from apartments found in the semi-annotated training dataset. We find that by using a self-training-based domain adaptation strategy, a robust system to extract home visit information can be built (ROC AUC = 0.773). We further show that the resulting visit information correlates well with the common geriatric depression scale screening tool ( ρ = -0.87, p = 0.001), providing further support for the idea of utilizing the extracted information as a potential digital measure or even as a digital biomarker to monitor the risk of late-life depression. Narayan Schütz, Angela Botros, Sami Ben Hassen, Hugo Saner, Philipp Buluschek, Prabitha Urwyler, Bruno Pais, Valérie Santschi, Daniel Gatica-Perez, René Müri, Tobias Nef |
IEEE J. Biomed. Health Informatics | 9 |
| 2021 | The Theory, Practice, and Ethical Challenges of Designing a Diversity-Aware Platform for Social RelationsabstractDiversity-aware platform design is a paradigm that responds to the ethical challenges of existing social media platforms. Available platforms have been criticized for minimizing users' autonomy, marginalizing minorities, and exploiting users' data for profit maximization. This paper presents a design solution that centers the well-being of users. It presents the theory and practice of designing a diversity-aware platform for social relations. In this approach, the diversity of users is leveraged in a way that allows like-minded individuals to pursue similar interests or diverse individuals to complement each other in a complex activity. The end users of the envisioned platform are students, who participate in the design process. Diversity-aware platform design involves numerous steps, of which two are highlighted in this paper: 1) defining a framework and operationalizing the "diversity" of students, 2) collecting "diversity" data to build diversity-aware algorithms. The paper further reflects on the ethical challenges encountered during the design of a diversity-aware platform. Laura Schelenz, Ivano Bison, Matteo Busso, Amalia de Götzen, Daniel Gatica-Perez, Fausto Giunchiglia, Lakmal Meegahapola, Salvador Ruiz-Correa |
AIES | 5 |
| 2021 | Locally Private Graph Neural NetworksabstractGraph Neural Networks (GNNs) have demonstrated superior performance in learning node representations for various graph inference tasks. However, learning over graph data can raise privacy concerns when nodes represent people or human-related variables that involve sensitive or personal information. In this paper, we study the problem of node data privacy, where graph nodes (e.g., social network users) have potentially sensitive data that is kept private, but they could be beneficial for a central server for training a GNN over the graph. To address this problem, we propose a privacy-preserving, architecture-agnostic GNN learning framework with formal privacy guarantees based on Local Differential Privacy (LDP). Specifically, we develop a locally private mechanism to perturb and compress node features, which the server can efficiently collect to approximate the GNN's neighborhood aggregation step. Furthermore, to improve the accuracy of the estimation, we prepend to the GNN a denoising layer, called KProp, which is based on the multi-hop aggregation of node features. Finally, we propose a robust algorithm for learning with privatized noisy labels, where we again benefit from KProp's denoising capability to increase the accuracy of label inference for node classification. Extensive experiments conducted over real-world datasets demonstrate that our method can maintain a satisfying level of accuracy with low privacy loss. Sina Sajadmanesh, Daniel Gatica-Perez |
CCS | 2 |
| 2021 | The impact of personality in using technology to ask and offer help: The experience of the Chatbot "UC - Paraguay"abstractIn the context of the project “WeNet: Internet of us” we are studying the role of diversity in relation to Internetmediated social interactions. In this paper, in particular, we analyze a possible relationship between personality aspects and social interaction mediated by digital platforms. More specifically, we rely on the five personality traits (Extraversion, Agreeableness, Conscientiousness, Emotional Stability and Openness to Experience), commonly referred to as “Big-five”, and associate them to automatically extracted behavioral characteristics derived from the experience of using a Chatbot for a closed community of students at the Universidad Católica “Nuestra Señora de la Asunción” (UC). The personality data comes from a self-report made by the users through questionnaires. The main results show very positive appraisals about the use of the Chatbot in terms of user experience and its main functionalities. As for the role of personality in relation to the main use of the Chatbot, although further experience is required to confirm trends, the results suggest that there are some correlations between some of the five personality traits and the length of questions and answers as well as the selection of best answers. José Luis Zarza, Alethia Hume, Luca Cernuzzi, Daniel Gatica-Perez, Ivano Bison |
CLEI | 4 |
| 2021 | Approximating the Mental Lexicon from Clinical Interviews as a Support Tool for Depression DetectionabstractDepression disorder is one of the major causes of disability in the world that can lead to tragic outcomes. In this paper, we propose a method for using an approximation to a mental lexicon to model the communication process of depressed and non-depressed participants in spontaneous North American English clinical interviews. Our approach, inspired by the Lexical Availability theory, identifies the most relevant vocabulary of the interviewed participant, and use it as features in a classification process. We performed an in-depth evaluation on the DAIC-WOZ [20] and the E-DAIC [11] clinical datasets. Obtained results indicate that our approach can compete against recent contextual embeddings when modeling and identifying depression. We show the generalization capabilities of our algorithm using outside data, reaching a macro F1 = 0.83 and F1 = 0.80 in the DAIC-WOZ and E-DAIC datasets respectively. An analysis of our method’s interpretability allows understanding how the classifier is making its decisions. During this process, we observed strong connections between our obtained results and previous research from the psychological field. Esaú Villatoro-Tello, Gabriela Ramírez-de-la-Rosa, Daniel Gatica-Perez, Mathew Magimai-Doss, Héctor Jiménez-Salazar |
ICMI | 3 |
| 2021 | Trust Indicators and Explainable AI: A Study on User Perceptions
Delphine Ribes Lemay, Nicolas Henchoz, Hélène Portier, Lara Défayes, Thanh-Trung Phan, Daniel Gatica-Perez, Andreas Sonderegger |
INTERACT (2) | 6 |
| 2021 | UrbanMM'21: 1st International Workshop on Multimedia Computing for Urban DataabstractUnderstanding complex processes that give cities their form traditionally relied primarily on the analysis of various open data statistics in relation to e.g. neighbourhood demographics, economy and mobility. However, recent years have seen an unprecedented increase in the availability and use of city-related sensors, participatory data and social multimedia. As the valuable information about urban challenges is usually encoded across multiple modalities, such as visual (e.g. panoramic, satellite and user-contributed images), text (e.g. social media and participatory data) and open data statistics, extracting this information requires effective multimedia analysis tools. This Workshop will showcase the power of multimedia computing in addressing various urban challenges, ranging from event detection and analysis, location recommendation and crowdedness estimation to more efficient handling of citizen reports and modelling and improving city liveability. In addition, it will serve as an impulse for the multimedia community to intensify research on these interesting, challenging and truly multimodal problems. Stevan Rudinac, Alessandro Bozzon, Tat-Seng Chua, Suzanne Little, Daniel Gatica-Perez, Kiyoharu Aizawa |
ACM Multimedia | 5 |
| 2021 | Declarative Variables in Online Dating: A Mixed-Method Analysis of a Mimetic-Distinctive MechanismabstractDeclarative variables of self-description have a long-standing tradition in matchmaking media. With the advent of online dating platforms and their brand positioning, the volume and semantics of variables vary greatly across apps. However, a variable landscape across multiple platforms, providing an in-depth understanding of the dating structure offered to users, has hitherto been absent in the literature. In this study, more than 300 declarative variables from 22 Anglophone and Francophone dating apps are examined. A mixed-method research design is used, combining hierarchical classification with an interview analysis of nine founders and developers in the industry. We present a new typology of variables in nine categories and a classification of dating apps, which highlights a double mimetic-distinctive mechanism in the variable definition and reflects the dating market. From the interviews, we extract three main factors concerning the economic and sociotechnical framework of coding practices, the actors' personal experience, and the development methodologies including user traces that influence this mechanism. This work, which to our knowledge is the most extensive thus far on dating app declarative variables, provides a new perspective on the analysis of the intersection between developers and users of online dating, which is mediated through variables, among other components. Jessica Pidoux, Pascale Kuntz, Daniel Gatica-Perez |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2020 | Understanding Applicants' Reactions to Asynchronous Video Interviews Through Self-reports and Nonverbal CuesabstractAsynchronous video interviews (AVIs) are increasingly used by organizations in their hiring process. In this mode of interviewing, the applicants are asked to record their responses to predefined interview questions using a webcam via an online platform. AVIs have increased usage due to employers' perceived benefits in terms of costs and scale. However, little research has been conducted regarding applicants' reactions to these new interview methods. In this work, we investigate applicants' reactions to an AVI platform using self-reported measures previously validated in psychology literature. We also investigate the connections of these measures with nonverbal behavior displayed during the interviews. We find that participants who found the platform creepy and had concerns about privacy reported lower interview performance compared to participants who did not have such concerns. We also observe weak correlations between nonverbal cues displayed and these self-reported measures. Finally, inference experiments achieve overall low-performance w.r.t. to explaining applicants' reactions. Overall, our results reveal that participants who are not at ease with AVIs (i.e., high creepy ambiguity score) might be unfairly penalized. This has implications for improved hiring practices using AVIs. Skanda Muralidhar, Emmanuelle Patricia Kleinlogel, Eric Mayor, Adrian Bangerter, Marianne Schmid Mast, Daniel Gatica-Perez |
ICMI | 6 |
| 2020 | Alone or With Others? Understanding Eating Episodes of College Students with Mobile SensingabstractUnderstanding food consumption patterns and contexts using mobile sensing is fundamental to build mobile health applications that require minimal user interaction to generate mobile food diaries. Many available mobile food diaries, both commercial and in research, heavily rely on self-reports, and this dependency limits the long term adoption of these apps by people. The social context of eating (alone, with friends, with family, with a partner, etc.) is an important self-reported feature that influences aspects such as food type, psychological state while eating, and the amount of food, according to prior research in nutrition and behavioral sciences. In this work, we use two datasets regarding the everyday eating behavior of college students in two countries, namely Switzerland (Nch=122) and Mexico (Nmx=84), to examine the relation between the social context of eating and passive sensing data from wearables and smartphones. Moreover, we design a classification task, namely inferring eating-alone vs. eating-with-others episodes using passive sensing data and time of eating, obtaining accuracies between 77% and 81%. We believe that this is a first step towards understanding more complex social contexts related to food consumption using mobile sensing. Lakmal Meegahapola, Salvador Ruiz-Correa, Daniel Gatica-Perez |
MUM | 3 |
| 2020 | Protecting Mobile Food Diaries from Getting too PersonalabstractSmartphone applications that use passive sensing to support human health and well-being primarily rely on: (a) generating low-dimensional representations from high-dimensional data streams; (b) making inferences regarding user behavior; and (c) using those inferences to benefit application users. Meanwhile, sometimes these datasets are shared with third parties as well. Human-centered ubiquitous systems need to ensure that sensitive attributes of users are protected when applications provide utility to people based on such behavioral inferences. In this paper, we demonstrate that inferences of sensitive attributes of users (gender, body mass index category) are possible using low-dimensional and sparse data coming from mobile food diaries (a combination of sensor data and self-reports). After exposing this potential risk, we demonstrate how deep learning techniques can be used for feature transformation to preserve sensitive user information while achieving high accuracies for application-related inferences (e.g. inferring the type of consumed food). Our work is based on two datasets of daily eating behavior of 160 young adults from Switzerland (NCH=122) and Mexico (NMX=38). Results show that using the proposed approach, accuracies in the order of 75%-90% can be achieved for application related inferences, while reducing the sensitive inference to almost random performance. Lakmal Meegahapola, Salvador Ruiz-Correa, Daniel Gatica-Perez |
MUM | 3 |
| 2020 | Learning Urban Nightlife Routines from Mobile DataabstractThe use of smartphone sensing for public health studies is appealing to understand routines. We present an approach to learn nightlife routines in a smartphone sensing dataset volunteered by 184 young people (1586 weekend nights with location data captured between 8PM and 4AM.) Human activity is represented at two levels, namely as the types of places visited and as the areas of the city where those places are. Routines extracted with two topic models (Latent Dirichlet Allocation and Hierarchical Dirichlet Process) are semantically meaningful and represent different moments of the weekend night, depicting activities such as pub crawling. The inference capacity of the routine representation is demonstrated with two classification tasks of value for alcohol research (alcohol consumption throughout the night, and heavy alcohol consumption.) The results suggest that nightlife routine mining could be used as a complementary tool to traditional survey-based methods in public health studies, and also inform other institutional actors interested in understanding and supporting youth well-being. Ada Pozo, Thanh-Trung Phan, Daniel Gatica-Perez |
MUM | 3 |
| 2019 | My Own Private Nightlife: Understanding Youth Personal Spaces from Crowdsourced VideoabstractPrivate nightlife environments of young people are likely characterized by their physical attributes, particular ambiance, and activities, but relatively little is known about it from social media studies. For instance, recent work has documented ambiance and physical characteristics of homes using pictures from Airbnb, but questions remain on whether this kind of curated data reliably represents everyday life situations. To describe the physical and ambiance features of homes of youth using manual annotations and machine-extracted features, we used a unique dataset of 301 crowdsourced videos of home environments recorded in-situ by young people on weekend nights. Agreement among five independent annotators was high for most studied variables. Results of the annotation task revealed various patterns of youth home spaces, such as the type of room attended (e.g., living room and bedroom), the number and gender of friends present, and the type of ongoing activities (e.g., watching TV alone; or drinking, chatting and eating in the presence of others.) Then, object and scene visual features of places, extracted via deep learning, were found to correlate with ambiances, while sound features did not. Finally, the results of a regression task for inferring ambiances from those features showed that six of the ambiance categories can be inferred with R 2 in the [0.21, 0.69] range. Our work is novel with regard to the type of data (crowdsourced videos of real homes of young people) and the analytical design (combined use of manual annotation and deep learning to identify relevant cues), and contributes to the understanding of home environments represented through digital media. Thanh-Trung Phan, Florian Labhart, Daniel Gatica-Perez |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2019 | Modeling Dyadic and Group Impressions with Intermodal and Interperson FeaturesabstractThis article proposes a novel feature-extraction framework for inferring impression personality traits, emergent leadership skills, communicative competence, and hiring decisions. The proposed framework extracts multimodal features, describing each participant’s nonverbal activities. It captures intermodal and interperson relationships in interactions and captures how the target interactor generates nonverbal behavior when other interactors also generate nonverbal behavior. The intermodal and interperson patterns are identified as frequent co-occurring events based on clustering from multimodal sequences. The proposed framework is applied to the SONVB corpus, which is an audiovisual dataset collected from dyadic job interviews, and the ELEA audiovisual data corpus, which is a dataset collected from group meetings. We evaluate the framework on a binary classification task involving 15 impression variables from the two data corpora. The experimental results show that the model trained with co-occurrence features is more accurate than previous models for 14 out of 15 traits. Shogo Okada, Laurent Son Nguyen, Oya Aran, Daniel Gatica-Perez |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2018 | Vlogging Over Time: Longitudinal Impressions and Behavior in YouTubeabstractYouTube vlogging, as a popular genre of ubiquitous social video, engages people in entertainment, civic, and social activities. Although several aspects of vlogging have been studied in media studies and multimedia analysis, the longitudinal angle of vlogging regarding recognition of personal state and trait impressions from behavior has not been yet analyzed. We present a study using behavioral data of vloggers who posted vlogs on YouTube for a period between three and six years. We use online crowdsourcing to collect a rich set of 21 impression variables for each video, including perceived personality, mood, skills, and expertise. Acoustic and motion features are extracted to characterize basic nonverbal behavior. The analysis shows that only a couple of perceived variables, including perceived expertise and perceived quality of audio and video, display weak temporal patterns. Furthermore, we show that the use of longitudinal data helps to improve the automatic inference of impressions for several of the impression variables. Daniel Gatica-Perez, Dairazalia Sanchez-Cortes, Trinh Minh Tri Do, Dinesh Babu Jayagopi, Kazuhiro Otsuka |
MUM | 1 |
| 2018 | Facing Employers and Customers: What Do Gaze and Expressions Tell About Soft Skills?abstractEye gaze and facial expressions are central to face-to-face social interactions. These behavioral cues and their connections to first impressions have been widely studied in psychology and computing literature, but limited to a single situation. Utilizing ubiquitous multimodal sensors coupled with advances in computer vision and machine learning, we investigate the connections between these behavioral cues and perceived soft skills in two diverse workplace situations (job interviews and reception desk). Pearson's correlation analysis shows a moderate connection between certain facial expressions, eye gaze cues and perceived soft skills in job interviews (r ϵ [-30,30]) and desk (r ϵ [20,36]) situations. Results of our computational framework to infer perceived soft skills indicates a low predictive power of eye gaze, facial expressions, and their combination in both interviews (R2 ϵ [0.02,0.21]) and desk (R2 ϵ [0.05,0.15]) situations. Our work has important implications for employee training and behavioral feedback systems. Skanda Muralidhar, Rémy Siegfried, Jean-Marc Odobez, Daniel Gatica-Perez |
MUM | 4 |
| 2018 | DrinkSense: Characterizing Youth Drinking Behavior Using SmartphonesabstractAlcohol consumption is the number one risk factor for morbidity and mortality among young people. In late adolescence and early adulthood, excessive drinking and intoxication are more common than in any other life period, increasing the risk of adverse physical and psychological health consequences. In this paper, we examine the feasibility of using smartphone sensor data and machine learning to automatically characterize and classify drinking behavior of young adults in an urban, ecologically valid nightlife setting. Our work has two contributions. First, we use previously unexplored data from a large-scale mobile crowdsensing study involving 241 young participants in two urban areas in a European country, which includes phone data (location. accelerometer, Wifi, Bluetooth, battery, screen, and app usage) along with self-reported, fine-grain data on individual alcoholic drinks consumed on Friday and Saturday nights over a three-month period. Second,we build a machine learning methodology to infer whether an individual consumed alcohol on a given weekend night, based on her/his smartphone data contributed between 8 PM and 4 AM. We found that accelerometer data is the most informative single cue, and that a combination of features results in an overall accuracy of 76.6 percent. Darshan Santani, Trinh Minh Tri Do, Florian Labhart, Sara Landolt, Emmanuel Kuntsche, Daniel Gatica-Perez |
IEEE Trans. Mob. Comput. | 6 |
| 2018 | Maya Codical Glyph Segmentation: A Crowdsourcing ApproachabstractThis paper focuses on the crowd-annotation of an ancient Maya glyph dataset derived from the three ancient codices that survived up to date. More precisely, nonexpert annotators are asked to segment glyph-blocks into their constituent glyph entities. As a means of supervision, available glyph variants are provided to the annotators during the crowdsourcing task. Compared to object recognition in natural images or handwriting transcription tasks, designing an engaging task and dealing with crowd behavior is challenging in our case. This challenge originates from the inherent complexity of Maya writing and an incomplete understanding of the signs and semantics in the existing catalogs. We elaborate on the evolution of the crowdsourcing task design, and discuss the choices for providing supervision during the task. We analyze the distributions of similarity and task difficulty scores, and the segmentation performance of the crowd. A unique dataset of over 9000 Maya glyphs from 291 categories individually segmented from the three codices was created and will be made publicly available thanks to this process. This dataset lends itself to automatic glyph classification tasks. We provide baseline methods for glyph classification using traditional shape descriptors and convolutional neural networks. Gulcan Can, Jean-Marc Odobez, Daniel Gatica-Perez |
IEEE Trans. Multim. | 3 |
| 2018 | Check Out This Place: Inferring Ambiance From Airbnb PhotosabstractAirbnb is changing the landscape of the hospitality industry, and to this day, little is known about the inferences that guests make about Airbnb listings. Our work constitutes a first attempt at understanding how potential Airbnb guests form first impressions from images, one of the main modalities featured on the platform. We contribute to the multimedia community by proposing the novel task of automatically predicting human impressions of ambiance from pictures of listings on Airbnb. We collected Airbnb images, focusing on the countries Switzerland and Mexico as case studies, and used crowdsourcing mechanisms to gather annotations on physical and ambiance attributes, finding that agreement among raters was high for most of the attributes. Our cluster analysis showed that both physical and psychological attributes could be grouped into three clusters. We then extracted state-of-the-art features from the images to automatically infer the annotated variables in a regression task. Results show the feasibility of predicting ambiance impressions of homes on Airbnb, with up to 42% of the variance explained by our model, and best results were obtained using activation layers of deep convolutional neural networks trained on the Places dataset, a collection of scene-centric images. Laurent Son Nguyen, Salvador Ruiz-Correa, Marianne Schmid Mast, Daniel Gatica-Perez |
IEEE Trans. Multim. | 4 |
| 2017 | How may I help you? behavior and impressions in hospitality service encountersabstractIn the service industry, customers often assess quality of service based on the behavior, perceived personality, and other attributes of the front line service employees they interact with. Interpersonal communication during these interactions is key to determine customer satisfaction and perceived service quality. We present a computational framework to automatically infer perceived performance and skill variables of employees interacting with customers in a hotel reception desk setting using nonverbal behavior, studying a dataset of 169 dyadic interactions involving students from a hospitality management school. We also study the connections between impressions of Big-5 personality traits, attractiveness, and performance of receptionists. In regression tasks, our automatic framework achieves R2=0.30 for performance impressions using audio-visual nonverbal cues, compared to 0.35 using personality impressions, while attractiveness impressions had low predictive power. We also study the integration of nonverbal behavior and Big-5 personality impressions towards increasing regression performance (R2 = 0.37). Skanda Muralidhar, Marianne Schmid Mast, Daniel Gatica-Perez |
ICMI | 3 |
| 2017 | Insiders and Outsiders: Comparing Urban Impressions between Population GroupsabstractThere is a growing interest in social and urban computing to employ crowdsourcing as means to gather impressions of urban perception for indoor and outdoor environments. Previous studies have established that reliable estimates of urban perception can be obtained using online crowdsourcing systems, but implicitly assumed that the judgments provided by the crowd are not dependent on the background knowledge of the observer. In this paper, we investigate how the impressions of outdoor urban spaces judged by online crowd annotators, compare with the impressions elicited by the local inhabitants, along six physical and psychological labels. We focus our study in a developing city where understanding and characterization of these socio-urban perceptions is of societal importance. We found statistically significant differences between the two population groups. Locals perceived places to be more dangerous and dirty, when compared with online crowd workers; while online annotators judged places to be more interesting in comparison to locals. Our results highlight the importance of the degree of familiarity with urban spaces and background knowledge while rating urban perceptions, which is lacking in some of the existing work in urban computing. Darshan Santani, Salvador Ruiz-Correa, Daniel Gatica-Perez |
ICMR | 3 |
| 2017 | Venues in Social Media: Examining Ambiance Perception Through Scene SemanticsabstractWe address the question of what visual cues, including scene objects and demographic attributes, contribute to the automatic inference of perceived ambiance in social media venues. We first use a state-of-art, deep scene semantic parsing method and a face attribute extractor to understand how different cues present in a scene relate to human perception of ambiance on Foursquare images of social venues. We then analyze correlational links between visual cues and thirteen ambiance variables, as well as the ability of the semantic attributes to automatically infer place ambiance. We study the effect of the type and amount of image data used for learning, and compare regression results to previous work, showing that the proposed approach results in marginal-to-moderate performance increase for up to ten of the ambiance dimensions, depending on the corpus. Yassir Benkhedda, Darshan Santani, Daniel Gatica-Perez |
ACM Multimedia | 3 |
| 2017 | Examining linguistic content and skill impression structure for job interview analytics in hospitalityabstractFirst impressions are critical to professional interactions especially in the context of employment interviews. This work investigates connections between linguistic content and first impressions in job interviews and the structure of ten soft skills and overall impressions. Towards this, we transcribe 169 role-played job interviews conducted at a hospitality school and analyze the linguistic content using off-the-shelf software. To understand the structure of the soft skill impressions, we conduct a principal component analysis. We then develop methods to automatically infer impressions using verbal and nonverbal features and their combination. Results indicate low predictive power of verbal cues for overall impression (R2 = 0.11). Combined verbal and nonverbal cues explain up to 34% of variance, a marginal improvement over R2 = 0.32 using only nonverbal cues. The use of principal components reveals a major component associated to overall positive and negative impressions that when used as labels for supervised learning results in a regression performance of R2 = 0.41. Skanda Muralidhar, Daniel Gatica-Perez |
MUM | 2 |
| 2017 | Healthy #fondue #dinner: analysis and inference of food and drink consumption patterns on instagramabstractSocial media generate large-scale data to study food and drink consumption in everyday life. Using Instagram posts in Switzerland over five years, our goal is two-fold. First, we extract key food & drink consumption patterns, through the lenses of a data-driven dictionary of popular items extracted from hashtags, and of a food categorization system used by the Swiss Federal government for national statistics purposes. Patterns related to spatial and temporal distributions of food & drink consumption, demographics, and eating events are extracted and compared to official statistics. Second, using the insights from this analysis, we define two eating event classification tasks, including a two-class task (healthy vs. unhealthy) and a six-class task (the three main meals break-fast/lunch/dinner/plus brunch/coffee/tea). Both tasks use hash-tags as labels for supervised learning. We study how content (hashtags and food categories), context (time and location), and social features (likes) can discriminate these eating events. A random forest and a combination of content and context features can classify healthy vs. unhealthy eating posts with 85.8% accuracy, and the six daily eating occasions with 61.7% accuracy. Thanh-Trung Phan, Daniel Gatica-Perez |
MUM | 2 |
| 2017 | Rapport with Virtual Agents: What Do Human Social Cues and Personality Explain?abstractRapport has been recognized as an important aspect of relationship building. While rapport in the context of human-human interaction has been widely studied, how it can be established and maintained in human-agent interaction has been studied only recently. Our study investigates how social cues and personality of a human interacting with an agent can be used for automatic prediction of rapport in this context. We conduct experiments with two emotional virtual agents. Alongside the audio-visual data, we also collect human personality measures and two measures of rapport: self-reported rapport and rapport judged by observers. The social cues, such as turn-taking patterns and facial expressions are extracted from audio-visual data. Our results show that the most significant cues that infer the rapport judgments are the number of turn-taking cues and pauses. We also find that some of the significant social cues related to rapport are similar to those reported in previous psychology literature. We also confirm previous findings on how human personality plays an important role in perceiving the interaction with agents-people who score high in extraversion and agreeableness report higher rapport with both agents. Finally, the rapport prediction results suggest that automatic analysis of social phenomena in human-agent interaction could be a feasible method for agent evaluation. Aleksandra Cerekovic, Oya Aran, Daniel Gatica-Perez |
IEEE Trans. Affect. Comput. | 3 |
| 2016 | The night is young: urban crowdsourcing of nightlife patternsabstractWe present a mobile crowdsourcing study to capture and examine the nightlife patterns of two youth populations in Switzerland. Our contributions are three fold. First, we developed a smartphone application to capture data on places, social context and nightlife activities, and to record mobile videos capturing the ambiance of places. Second, we conducted an "in-the-wild" study with more than 200 participants over a period of three months in two Swiss cities, resulting in a total of 1,394 unique place visits and 843 videos that spread across place categories (including personal homes and public parks), social and ambiance variables. Finally, we investigated the use of automatic ambiance features to estimate the loudness and brightness of places at scale, and found that while features are reliable with respect to video content, videos do not always reflect the place ambiance reported by people in-situ. We believe that the developed methodology provides an opportunity to understand the physical mobility, activities, and social context of youth as they experience different aspects of nightlife. Darshan Santani, Joan-Isaac Biel, Florian Labhart, Jasmine Truong, Sara Landolt, Emmanuel Kuntsche, Daniel Gatica-Perez |
UbiComp | 7 |
| 2016 | Stressful first impressions in job interviewsabstractStress can impact many aspects of our lives, such as the way we interact and work with others, or the first impressions that we make. In the past, stress has been most commonly assessed through self-reported questionnaires; however, advancements in wearable technology have enabled the measurement of physiological symptoms of stress in an unobtrusive manner. Using a dataset of job interviews, we investigate whether first impressions of stress (from annotations) are equivalent to physiological measurements of the electrodermal activity (EDA). We examine the use of automatically extracted nonverbal cues stemming from both the visual and audio modalities, as well EDA stress measurements for the inference of stress impressions obtained from manual annotations. Stress impressions were found to be significantly negatively correlated with hireability ratings i.e individuals who were perceived to be more stressed were more likely to obtained lower hireability scores. The analysis revealed a significant relationship between audio and visual features but low predictability and no significant effects were found for the EDA features. While some nonverbal cues were more clearly related to stress, the physiological cues were less reliable and warrant further investigation into the use of wearable sensors for stress detection. Ailbhe Finnerty, Skanda Muralidhar, Laurent Son Nguyen, Fabio Pianesi, Daniel Gatica-Perez |
ICMI | 5 |
| 2016 | Training on the job: behavioral analysis of job interviews in hospitalityabstractFirst impressions play a critical role in the hospitality industry and have been shown to be closely linked to the behavior of the person being judged.In this work, we implemented a behavioral training framework for hospitality students with the goal of improving the impressions that other people make about them. We outline the challenges associated with designing such a framework and embedding it in the everyday practice of a real hospitality school. We collected a dataset of 169 laboratory sessions where two role-plays were conducted, job interviews and reception desk scenarios, for a total of 338 interactions. For job interviews, we evaluated the relationship between automatically extracted nonverbal cues and various perceived social variables in a correlation analysis. Furthermore, our system automatically predicted first impressions from job interviews in a regression task, and was able to explain up to 32% of the variance, thus extending the results in existing literature, and showing gender differences, corroborating previous findings in psychology. This work constitutes a step towards applying social sensing technologies to the real world by designing and implementing a living lab for students of an international hospitality management school. Skanda Muralidhar, Laurent Son Nguyen, Denise Frauendorfer, Jean-Marc Odobez, Marianne Schmid Mast, Daniel Gatica-Perez |
ICMI | 6 |
| 2016 | InnerView: Learning Place Ambiance from Social Media ImagesabstractIn the recent past, there has been interest in characterizing the physical and social ambiance of urban spaces to understand how people perceive and form impressions of these environments based on physical and psychological constructs. Building on our earlier work on characterizing ambiance of indoor places, we present a methodology to automatically infer impressions of place ambiance, using generic deep learning features extracted from images publicly shared on Foursquare. We base our methodology on a corpus of 45,000 images from 300 popular places in six cities on Foursquare. Our results indicate the feasibility to automatically infer place ambiance with a maximum R2 of 0.53 using features extracted from a pre-trained convolutional neural network. We found that features extracted from deep learning with convolutional nets consistently outperformed individual and combinations of several low-level image features (including Color, GIST, HOG and LBP) to infer all the studied 13 ambiance dimensions. Our work constitutes a first study to automatically infer ambiance impressions of indoor places from deep features learned from images shared on social media. Darshan Santani, Rui Hu 0010, Daniel Gatica-Perez |
ACM Multimedia | 3 |
| 2016 | Dites-moi: wearable feedback on conversational behaviorabstractInterpersonal communication skills are critical in certain industry sectors like sales and marketing. Recent advances in wearable technology are enabling the design of real-time behavioral feedback tools for apprentices in aforementioned industries. This paper describes the design and implementation of a conversational behavior awareness tool based on Google Glass. The goal of the system is to provide real-time feedback to young sales apprentices about the amount of time they talk in an interaction with a client. We evaluated our system with a pilot study involving 15 apprentices (ages 16--20). Overall, participants found the system fun, little distracting and useful. Furthermore, manual coding of the recorded videos, showed that wearable sensing and real-time feedback did not negatively influence the dyadic social interaction. Skanda Muralidhar, Jean Marcel dos Reis Costa, Laurent Son Nguyen, Daniel Gatica-Perez |
MUM | 4 |
| 2016 | Hirability in the Wild: Analysis of Online Conversational Video ResumesabstractOnline social media is changing the personnel recruitment process. Until now, resumes were among the most widely used tools for the screening of job applicants. The advent of inexpensive sensors combined with the success of online video platforms has enabled the introduction of a new type of resume. Video resumes are short video messages where job applicants present themselves to potential employers. Online video resumes represent an opportunity to study the formation of first impressions in an employment context at a scale never attempted before, and to our knowledge they have not been studied from a behavioral standpoint. We collected a dataset of 939 conversational English-speaking video resumes from YouTube. Annotations of demographics, skills, and first impressions were collected using the Amazon Mechanical Turk crowdsourcing platform. Basic demographics were then analyzed to understand the population who uses video resumes to find a job, and results showed that applicants mainly consisted of young people looking for internship and junior positions. We developed a computational framework for the prediction of organizational first impressions, where the inference and nonverbal cue extraction steps were fully automated. Results demonstrate automatic prediction of first impressions of up to 27% of the variance explained for extraversion, and up to 20% for social and communication skills. Laurent Son Nguyen, Daniel Gatica-Perez |
IEEE Trans. Multim. | 2 |
| 2015 | I Would Hire You in a Minute: Thin Slices of Nonverbal Behavior in Job InterviewsabstractIn everyday life, judgments people make about others are based on brief excerpts of interactions, known as thin slices. Inferences stemming from such minimal information can be quite accurate, and nonverbal behavior plays an important role in the impression formation. Because protagonists are strangers, employment interviews are a case where both nonverbal behavior and thin slices can be predictive of outcomes. In this work, we analyze the predictive validity of thin slices of real job interviews, where slices are defined by the sequence of questions in a structured interview format. We approach this problem from an audio-visual, dyadic, and nonverbal perspective, where sensing, cue extraction, and inference are automated. Our study shows that although nonverbal behavioral cues extracted from thin slices were not as predictive as when extracted from the full interaction, they were still predictive of hirability impressions with $R^2$ values up to $0.34$, which was comparable to the predictive validity of human observers on thin slices. Applicant audio cues were found to yield the most accurate results. Laurent Son Nguyen, Daniel Gatica-Perez |
ICMI | 2 |
| 2015 | Personality Trait Classification via Co-Occurrent Multiparty Multimodal Event DiscoveryabstractThis paper proposes a novel feature extraction framework from mutli-party multimodal conversation for inference of personality traits and emergent leadership. The proposed framework represents multi modal features as the combination of each participant's nonverbal activity and group activity. This feature representation enables to compare the nonverbal patterns extracted from the participants of different groups in a metric space. It captures how the target member outputs nonverbal behavior observed in a group (e.g. the member speaks while all members move their body), and can be available for any kind of multiparty conversation task. Frequent co-occurrent events are discovered using graph clustering from multimodal sequences. The proposed framework is applied for the ELEA corpus which is an audio visual dataset collected from group meetings. We evaluate the framework for binary classification task of 10 personality traits. Experimental results show that the model trained with co-occurrence features obtained higher accuracy than previously related work in 8 out of 10 traits. In addition, the co-occurrence features improve the accuracy from 2 % up to 17 %. Shogo Okada, Oya Aran, Daniel Gatica-Perez |
ICMI | 3 |
| 2015 | CommuniSense: Crowdsourcing Road Hazards in NairobiabstractNairobi is one of the fastest growing metropolitan cities and a major business and technology powerhouse in Africa. However, Nairobi currently lacks monitoring technologies to obtain reliable data on traffic and road infrastructure conditions. In this paper, we investigate the use of mobile crowdsourcing as means to gather and document Nairobi's road quality information. We first present the key findings of a city-wide road quality survey about the perception of existing road quality conditions in Nairobi. Based on the survey's findings, we then developed a mobile crowdsourcing application, called CommuniSense, to collect road quality data. The application serves as a tool for users to locate, describe, and photograph road hazards. We tested our application through a two-week field study amongst 30 participants to document various forms of road hazards from different areas in Nairobi. To verify the authenticity of user-contributed reports from our field study, we proposed to use online crowdsourcing using Amazon's Mechanical Turk (MTurk) to verify whether submitted reports indeed depict road hazards. We found 92% of user-submitted reports to match the MTurkers judgements. While our prototype was designed and tested on a specific city, our methodology is applicable to other developing cities Darshan Santani, Jidraph Njuguna, Tierra Bills, Aisha Walcott-Bryant, Reginald E. Bryant, Jonathan Ledgard, Daniel Gatica-Perez |
MobileHCI | 7 |
| 2015 | Loud and Trendy: Crowdsourcing Impressions of Social Ambiance in Popular Indoor Urban PlacesabstractNew research cutting across architecture, urban studies, and psychology is contextualizing the understanding of urban spaces according to the perceptions of their inhabitants. One fundamental construct that relates place and experience is ambiance, which is defined as "the mood or feeling associated with a particular place". We posit that the systematic study of ambiance dimensions in cities is a new domain for which multimedia research can make pivotal contributions. We present a study to examine how images collected from social media can be used for the crowdsourced characterization of indoor ambiance impressions in popular urban places. We design a crowdsourcing framework to understand suitability of social images as data source to convey place ambiance, to examine what type of images are most suitable to describe ambiance, and to assess how people perceive places socially from the perspective of ambiance along 13 dimensions. Our study is based on 50,000 Foursquare images collected from 300 popular places across six cities worldwide. The results show that reliable estimates of ambiance can be obtained for several of the dimensions. Furthermore, we found that most aggregate impressions of ambiance are similar across popular places in all studied cities. We conclude by presenting a multidisciplinary research agenda for future research in this domain. Darshan Santani, Daniel Gatica-Perez |
ACM Multimedia | 2 |
| 2015 | Happy and agreeable?: multi-label classification of impressions in social videoabstractThe mobile and ubiquitous nature of conversational social video has placed video blogs among the most popular forms of online video. For this reason, there has been an increasing interest in conducting studies of human behavior from video blogs in affective and social computing. In this context, we consider the problem of mood and personality trait impression inference using verbal and nonverbal audio-visual features. Under a multi-label classification framework, we show that for both mood and personality trait binary label sets, not only the simultaneous inference of multiple labels is feasible, but also that classification accuracy increases moderately for several labels, compared to a single-label approach. The multi-label method we consider naturally exploits label correlations, which motivate our approach, and our results are consistent with models proposed in psychology to define human emotional states and personality. Our approach points to the automatic specification of co-occurring emotional states and personality, by inferring several labels at once, compared to single-label approaches. We also propose a new set of facial features, based on emotion valence from facial expressions, and analyze their suitability in the multi-label framework. Gilberto Chávez-Martínez, Salvador Ruiz-Correa, Daniel Gatica-Perez |
MUM | 3 |
| 2015 | A probabilistic kernel method for human mobility prediction with smartphones
Trinh Minh Tri Do, Olivier Dousse, Markus Miettinen, Daniel Gatica-Perez |
Pervasive Mob. Comput. | 4 |
| 2015 | Theme issue from ISWC 2013
Daniel Gatica-Perez, Daniel Roggen, Masaaki Fukumoto |
Pers. Ubiquitous Comput. | 1 |
| 2015 | What Your Face Vlogs About: Expressions of Emotion and Big-Five Traits Impressions in YouTubeabstractSocial video sites where people share their opinions and feelings are increasing in popularity. The face is known to reveal important aspects of human psychological traits, so the understanding of how facial expressions relate to personal constructs is a relevant problem in social media. We present a study of the connections between automatically extracted facial expressions of emotion and impressions of Big-Five personality traits in YouTube vlogs (i.e., video blogs). We use the Computer Expression Recognition Toolbox (CERT) system to characterize users of conversational vlogs. From CERT temporal signals corresponding to instantaneously recognized facial expression categories, we propose and derive four sets of behavioral cues that characterize face statistics and dynamics in a compact way. The cue sets are first used in a correlation analysis to assess the relevance of each facial expression of emotion with respect to Big-Five impressions obtained from crowd-observers watching vlogs, and also as features for automatic personality impression prediction. Using a dataset of 281 vloggers, the study shows that while multiple facial expression cues have significant correlation with several of the Big-Five traits, they are only able to significantly predict Extraversion impressions with moderate values of R2. Lucia Teijeiro-Mosquera, Joan-Isaac Biel, José Luis Alba-Castro, Daniel Gatica-Perez |
IEEE Trans. Affect. Comput. | 4 |
| 2015 | In the Mood for Vlog: Multimodal Inference in Conversational Social VideoabstractThe prevalent “share what's on your mind” paradigm of social media can be examined from the perspective of mood: short-term affective states revealed by the shared data. This view takes on new relevance given the emergence of conversational social video as a popular genre among viewers looking for entertainment and among video contributors as a channel for debate, expertise sharing, and artistic expression. From the perspective of human behavior understanding, in conversational social video both verbal and nonverbal information is conveyed by speakers and decoded by viewers. We present a systematic study of classification and ranking of mood impressions in social video, using vlogs from YouTube. Our approach considers eleven natural mood categories labeled through crowdsourcing by external observers on a diverse set of conversational vlogs. We extract a comprehensive number of nonverbal and verbal behavioral cues from the audio and video channels to characterize the mood of vloggers. Then we implement and validate vlog classification and vlog ranking tasks using supervised learning methods. Following a reliability and correlation analysis of the mood impression data, our study demonstrates that, while the problem is challenging, several mood categories can be inferred with promising performance. Furthermore, multimodal features perform consistently better than single-channel features. Finally, we show that addressing mood as a ranking problem is a promising practical direction for several of the mood categories studied. Dairazalia Sanchez-Cortes, Shiro Kumano, Kazuhiro Otsuka, Daniel Gatica-Perez |
ACM Trans. Interact. Intell. Syst. | 4 |
| 2015 | Let Your Body Speak: Communicative Cue Extraction on Natural Interaction Using RGBD DataabstractEmployment interviews are relevant scenarios for the study of social interaction. In this setting, social skills play an important role, even though the interactions between potential employers and candidates are often limited. One fundamental aspect of social interaction is the use of nonverbal communication , which affects how we are socially perceived. We present a method to automatically extract body communicative cues from one-on-one conversations recorded with Kinect devices. First, we find the three-dimensional position of hands and head of the subject, and, aided by training data, we infer the upper body pose. Then, we use the inferred poses to perform action recognition and build person-specific activity descriptors. We evaluate our system with both domain-specific and public, generic datasets, and show competitive performance. Alvaro Marcos-Ramiro, Daniel Pizarro-Perez, Marta Marrón Romera, Daniel Gatica-Perez |
IEEE Trans. Multim. | 4 |
| 2014 | Capturing Upper Body Motion in Conversation: An Appearance Quasi-Invariant ApproachabstractWe address the problem of body communication retrieval and measuring in seated conversations by means of markerless motion capture. In psychological studies, the use of automatic methods is key to reduce the subjectivity present in manual behavioral coding used to extract these cues. These studies usually involve hundreds of subjects with different clothing, non-acted poses, or different distances to the camera in uncalibrated, RGB-only video. However, range cameras are not yet common in psychology research, especially in existing recordings. Therefore, it becomes highly relevant to develop a fast method that is able to work in these conditions. Given the known relationship between depth and motion estimates, we propose to robustly integrate highly appearance-invariant image motion features in a machine learning approach, complemented with an effective tracking scheme. We evaluate the method's performance with existing databases and a database of upper body poses displayed in job interviews that we make public, showing that in our scenario it is comparable to that of Kinect without using a range camera, and state-of-the-art w.r.t. the HumanEva and ChaLearn 2011 evaluation datasets. Alvaro Marcos-Ramiro, Daniel Pizarro-Perez, Marta Marrón Romera, Daniel Gatica-Perez |
ICMI | 4 |
| 2014 | Automatic Blinking Detection towards Stress DiscoveryabstractWe present a robust method to automatically detect blinks in video sequences of conversations, aimed to discovering stress. Psychological studies have shown a relationship between blink frequency and dopamine levels, which in turn are affected by stress. Task performance correlates through an inverted U shape to both dopamine and stress levels. This shows the importance of automatic blink detection as a way of reducing human coding burden. We use an off-the-shelf face tracker in order to extract the eye region. Then, we perform per-pixel classification of the extracted eye images to later identify blinks through their dynamics. We evaluate the performance of our system with a job interview database with annotations of psychological variables, and show statistically significant correlation between perceived stress resistance and the automatically detected blink patterns. Alvaro Marcos-Ramiro, Daniel Pizarro-Perez, Marta Marrón Romera, Daniel Gatica-Perez |
ICMI | 4 |
| 2014 | The Workshop on Computational Personality Recognition 2014abstractThe Workshop on Computational Personality Recognition aims to define the state-of-the-art in the field and to provide tools for future standard evaluations in personality recognition tasks. In the WCPR14 we released two different datasets: one of Youtube Vlogs and one of Mobile Phone interactions. We structured the workshop in two tracks: an open shared task, where participants can do any kind of experiment, and a competition. We also distinguished two tasks: A) personality recognition from multimedia data, and B) personality recognition from text only. In this paper we discuss the results of the workshop. Fabio Celli, Bruno Lepri, Joan-Isaac Biel, Daniel Gatica-Perez, Giuseppe Riccardi, Fabio Pianesi |
ACM Multimedia | 4 |
| 2014 | Automatic Maya hieroglyph retrieval using shape and context informationabstractWe propose an automatic Maya hieroglyph retrieval method integrating shape and glyph context information. Two recent local shape descriptors, Gradient Field Histogram of Orientation Gradient (GF-HOG) and Histogram of Orientation Shape Context (HOOSC), are evaluated. To encode the context information, we propose to convert each Maya glyph block into a first-order Markov chain and apply the co-occurrence of neighbouring glyphs. The retrieval results obtained based on visual matching are therefore re-ranked. Experimental results show that our method can significantly improve the glyph retrieval accuracy even with a basic co-occurrence model. Furthermore, two unique glyph datasets are contributed which can be used as novel shape benchmarks in future research. Rui Hu 0010, Carlos Pallan, Guido Krempel, Jean-Marc Odobez, Daniel Gatica-Perez |
ACM Multimedia | 5 |
| 2014 | Where and what: Using smartphones to predict next locations and applications in daily life
Trinh Minh Tri Do, Daniel Gatica-Perez |
Pervasive Mob. Comput. | 2 |
| 2014 | A probabilistic approach to mining mobile phone data sequences
Katayoun Farrahi, Daniel Gatica-Perez |
Pers. Ubiquitous Comput. | 2 |
| 2014 | The Places of Our Lives: Visiting Patterns and Automatic Labeling from Longitudinal Smartphone DataabstractThe location tracking functionality of modern mobile devices provides unprecedented opportunity to the understanding of individual mobility in daily life. Instead of studying raw geographic coordinates, we are interested in understanding human mobility patterns based on sequences of place visits which encode, at a coarse resolution, most daily activities. This paper presents a study on place characterization in people's everyday life based on data recorded continuously by smartphones. First, we study human mobility from sequences of place visits, including visiting patterns on different place categories. Second, we address the problem of automatic place labeling from smartphone data without using any geo-location information. Our study on a large-scale data collected from 114 smartphone users over 18 months confirm many intuitions, and also reveals findings regarding both regularly and novelty trends in visiting patterns. Considering the problem of place labeling with 10 place categories, we show that frequently visited places can be recognized reliably (over 80 percent) while it is much more challenging to recognize infrequent places. Trinh Minh Tri Do, Daniel Gatica-Perez |
IEEE Trans. Mob. Comput. | 2 |
| 2014 | Broadcasting Oneself: Visual Discovery of Vlogging StylesabstractWe present a data-driven approach to discover different styles that people use to present themselves in online video blogging (vlogging). By vlogging style, we denote the combination of conscious and unconscious choices that the vlogger made during the production of the vlog, affecting the video quality, appearance, and structure. A compact set of vlogging styles is discovered using clustering methods based on a fast and robust spatio-temporal descriptor to characterize the visual activity in a vlog. On 2268 YouTube vlogs, our results show that the vlogging styles are differentiated with respect to the vloggers' level of editing and conversational activity in the video. Furthermore, we show that these automatically discovered styles relate to vloggers with different personality trait impressions and to vlogs that receive different levels of social attention. Oya Aran, Joan-Isaac Biel, Daniel Gatica-Perez |
IEEE Trans. Multim. | 3 |
| 2014 | Mining Crowdsourced First Impressions in Online Social VideoabstractWhile multimedia and social computing research have used crowdsourcing techniques to annotate objects, actions, and scenes in social video sites like YouTube, little work has addressed the crowdsourcing of personal and social traits in online social video or social media content in general. In this paper, we address the problems of (1) crowdsourcing the annotation of first impressions of video bloggers (vloggers) personal and social traits in conversational YouTube videos, and (2) mining the impressions with the goal of modeling the interplay of different vlogger facets. First, we design a human annotation task to crowdsource impressions of vloggers that extends a tradition of studies of personality impressions with the addition of attractiveness and mood impressions. Second, we propose a probabilistic framework using Topic Models to discover prototypical impressions that are data driven, and that combine multiple facets of vloggers. Finally, we address the task of automatically predicting topic impressions using nonverbal and verbal content extracted from videos and comments. Our study of 442 YouTube vlogs and 2,210 annotations collected in Mechanical Turk supports recent literature showing the feasibility to crowdsource interpersonal human impression with comparable quality to what is reported in social psychology research, and provides insights on the interplay among human first impressions. We also show that topic models are useful to discover meaningful prototypical impressions that can be validated by humans, and that different topics can be predicted using different sources of information from vloggers' nonverbal and verbal content, as well as comments from the audience. Joan-Isaac Biel, Daniel Gatica-Perez |
IEEE Trans. Multim. | 2 |
| 2014 | Hire me: Computational Inference of Hirability in Employment Interviews Based on Nonverbal BehaviorabstractUnderstanding the basis on which recruiters form hirability impressions for a job applicant is a key issue in organizational psychology and can be addressed as a social computing problem. We approach the problem from a face-to-face, nonverbal perspective where behavioral feature extraction and inference are automated. This paper presents a computational framework for the automatic prediction of hirability. To this end, we collected an audio-visual dataset of real job interviews where candidates were applying for a marketing job. We automatically extracted audio and visual behavioral cues related to both the applicant and the interviewer. We then evaluated several regression methods for the prediction of hirability scores and showed the feasibility of conducting such a task, with ridge regression explaining 36.2% of the variance. Feature groups were analyzed, and two main groups of behavioral cues were predictive of hirability: applicant audio features and interviewer visual cues, showing the predictive validity of cues related not only to the applicant, but also to the interviewer. As a last step, we analyzed the predictive validity of psychometric questionnaires often used in the personnel selection process, and found that these questionnaires were unable to predict hirability, suggesting that hirability impressions were formed based on the interaction during the interview rather than on questionnaire data. Laurent Son Nguyen, Denise Frauendorfer, Marianne Schmid Mast, Daniel Gatica-Perez |
IEEE Trans. Multim. | 4 |
| 2013 | The vernissage corpus: a conversational human-robot-interaction dataset
Dinesh Babu Jayagopi, Samira Sheikhi, David Klotz, Johannes Wienke, Jean-Marc Odobez, Sebastian Wrede 0001, Vasil Khalidov, Laurent Nyugen, Britta Wrede, Daniel Gatica-Perez |
HRI | 10 |
| 2013 | One of a kind: inferring personality impressions in meetingsabstractWe present an analysis on personality prediction in small groups based on trait attributes from external observers. We use a rich set of automatically extracted audio-visual nonverbal features, including speaking turn, prosodic, visual activity, and visual focus of attention features. We also investigate whether the thin sliced impressions of external observers generalize to the whole meeting in the personality prediction task. Using ridge regression, we have analyzed both the regression and classification performance of personality prediction. Our experiments show that the extraversion trait can be predicted with high accuracy in a binary classification task and visual activity features give higher accuracies than audio ones. The highest accuracy for the extraversion trait, is 75\%, obtained with a combination of audio-visual features. Openness to experience trait also has a significant accuracy, only when the whole meeting is used as the unit of processing. Oya Aran, Daniel Gatica-Perez |
ICMI | 2 |
| 2013 | Cross-domain personality prediction: from video blogs to small group meetingsabstractIn this study, we investigate the use of social media content as a domain to learn personality trait impressions, particularly extraversion. Our aim is to transfer the knowledge that can be extracted from conversational videos in video blogging sites to small group settings to predict the extraversion trait with nonverbal cues. We use YouTube data containing personality impression scores of 442 people as the source domain and a small-group meeting data from a total of 102 people as our target domain. Our results show that, for the extraversion trait, by using user-created video blogs, as part of the training data, and a small amount of adaptation data from the target domain, we are able to achieve higher prediction accuracies than using only the data recorded in small group settings. Oya Aran, Daniel Gatica-Perez |
ICMI | 2 |
| 2013 | Hi YouTube!: personality impressions and verbal content in social videoabstractDespite the evidence that social video conveys rich human personality information, research investigating the automatic prediction of personality impressions in vlogging has shown that, amongst the Big-Five traits, automatic nonverbal behavioral cues are useful to predict mainly the Extraversion trait. This finding, also reported in other conversational settings, indicates that personality information may be coded in other behavioral dimensions like the verbal channel, which has been less studied in multimodal interaction research. In this paper, we address the task of predicting personality impressions from vloggers based on what they say in their YouTube videos. First, we use manual transcripts of vlogs and verbal content analysis techniques to understand the ability of verbal content for the prediction of crowdsourced Big-Five personality impressions. Second, we explore the feasibility of a fully-automatic framework in which transcripts are obtained using automatic speech recognition (ASR). Our results show that the analysis of error-free verbal content is useful to predict four of the Big-Five traits, three of them better than using nonverbal cues, and that the errors caused by the ASR system decrease the performance significantly. Joan-Isaac Biel, Vagia Tsiminaki, John Dines, Daniel Gatica-Perez |
ICMI | 4 |
| 2013 | Inferring social activities with mobile sensor networksabstractWhile our daily activities usually involve interactions with others, the current methods on activity recognition do not often exploit the relationship between social interactions and human activity. This paper addresses the problem of interpreting social activity from human interactions captured by mobile sensing networks. Our first goal is to discover different social activities such as chatting with friends from interaction logs and then characterize them by the set of people involved, and the time and location of the occurring event. Our second goal is to perform automatic labeling of the discovered activities using predefined semantic labels such as coffee breaks, weekly meetings, or random discussions. Our analysis was conducted on a real-life interaction network sensed with Bluetooth and infrared sensors of about fifty subjects who carried sociometric badges over 6 weeks. We show that the proposed system reliably recognized coffee breaks with 99% accuracy, while weekly meetings were recognized with 88% accuracy. Trinh Minh Tri Do, Kyriaki Kalimeri, Bruno Lepri, Fabio Pianesi, Daniel Gatica-Perez |
ICMI | 5 |
| 2013 | A semi-automated system for accurate gaze coding in natural dyadic interactionsabstractIn this paper we propose a system capable of accurately coding gazing events in natural dyadic interactions. Contrary to previous works, our approach exploits the actual continuous gaze direction of a participant by leveraging on remote RGB-D sensors and a head pose-independent gaze estimation method. Our contributions are: i) we propose a system setup built from low-cost sensors and a technique to easily calibrate these sensors in a room with minimal assumptions; ii) we propose a method which, provided short manual annotations, can automatically detect gazing events in the rest of the sequence; iii) we demonstrate on substantially long, natural dyadic data that high accuracy can be obtained, showing the potential of our system. Our approach is non-invasive and does not require collaboration from the interactors. These characteristics are highly valuable in psychology and sociology research. Kenneth Alberto Funes Mora, Laurent Son Nguyen, Daniel Gatica-Perez, Jean-Marc Odobez |
ICMI | 3 |
| 2013 | Multimodal analysis of body communication cues in employment interviewsabstractHand gestures and body posture are intimately linked to speech as they are used to enrich the vocal content, and are therefore inherently multimodal. As an important part of nonverbal behavior, body communication carries relevant information that can reveal social constructs as diverse as personality, internal states, or job interview outcomes. In this work, we analyze body communication cues in real dyadic employment interviews, where the protagonists of the interaction are seated. We use a mixture of body communicative features based on manual annotations and automated extraction methods to successfully predict two key organizational constructs, namely personality and job interview ratings. Our work also confirms the multimodal nature of body communication and shows that the speaking status can be used to improve the prediction performance of personality and hirability. Laurent Son Nguyen, Alvaro Marcos-Ramiro, Marta Marrón Romera, Daniel Gatica-Perez |
ICMI | 4 |
| 2013 | From Foursquare to My Square: Learning Check-in Behavior from Multiple Sources
Eric Malmi, Trinh Minh Tri Do, Daniel Gatica-Perez |
ICWSM | 3 |
| 2013 | Speaking swiss: languages and venues in foursquareabstractDue to increasing globalization, urban societies are becoming more multicultural. The availability of large-scale digital mobility traces e.g. from tweets or checkins provides an opportunity to explore multiculturalism that until recently could only be addressed using survey-based methods. In this paper we examine a basic facet of multiculturalism through the lens of language use across multiple cities in Switzerland. Using data obtained from Foursquare over 330 days, we present a descriptive analysis of linguistic differences and similarities across five urban agglomerations in a multicultural, western European country. Darshan Santani, Daniel Gatica-Perez |
ACM Multimedia | 2 |
| 2013 | Inferring mood in ubiquitous conversational videoabstractConversational social video is becoming a worldwide trend. Video communication allows a more natural interaction, when aiming to share personal news, ideas, and opinions, by transmitting both verbal content and nonverbal behavior. However, the automatic analysis of natural mood is challenging, since it is displayed in parallel via voice, face, and body. This paper presents an automatic approach to infer 11 natural mood categories in conversational social video using single and multimodal nonverbal cues extracted from video blogs (vlogs) from YouTube. The mood labels used in our work were collected via crowdsourcing. Our approach is promising for several of the studied mood categories. Our study demonstrates that although multimodal features perform better than single channel features, not always all the available channels are needed to accurately discriminate mood in videos. Dairazalia Sanchez-Cortes, Joan-Isaac Biel, Shiro Kumano, Junji Yamato, Kazuhiro Otsuka, Daniel Gatica-Perez |
MUM | 6 |
| 2013 | Discovering places of interest in everyday life from smartphone data
Raúl Montoliu, Jan Blom, Daniel Gatica-Perez |
Multim. Tools Appl. | 3 |
| 2013 | Special Issue on the Mobile Data Challenge
Daniel Gatica-Perez, Juha K. Laurila, Jan Blom |
Pervasive Mob. Comput. | 1 |
| 2013 | From big smartphone data to worldwide research: The Mobile Data Challenge
Juha K. Laurila, Daniel Gatica-Perez, Imad Aad, Jan Blom, Olivier Bornet, Trinh Minh Tri Do, Olivier Dousse, Julien Eberle, Markus Miettinen |
Pervasive Mob. Comput. | 2 |
| 2013 | Mining large-scale smartphone data for personality studies
Gokul Chittaranjan, Jan Blom, Daniel Gatica-Perez |
Pers. Ubiquitous Comput. | 3 |
| 2013 | Human interaction discovery in smartphone proximity networks
Trinh Minh Tri Do, Daniel Gatica-Perez |
Pers. Ubiquitous Comput. | 2 |
| 2013 | Wordless Sounds: Robust Speaker Diarization Using Privacy-Preserving Audio RepresentationsabstractThis paper investigates robust privacy-sensitive audio features for speaker diarization in multiparty conversations: i.e., a set of audio features having low linguistic information for speaker diarization in a single and multiple distant microphone scenarios. We systematically investigate Linear Prediction (LP) residual. Issues such as prediction order and choice of representation of LP residual are studied. Additionally, we explore the combination of LP residual with subband information from 2.5 kHz to 3.5 kHz and spectral slope. Next, we propose a supervised framework using deep neural architecture for deriving privacy-sensitive audio features. We benchmark these approaches against the traditional Mel Frequency Cepstral Coefficients (MFCC) features for speaker diarization in both the microphone scenarios. Experiments on the RT07 evaluation dataset show that the proposed approaches yield diarization performance close to the MFCC features on the single distant microphone dataset. To objectively evaluate the notion of privacy in terms of linguistic information, we perform human and automatic speech recognition tests, showing that the proposed approaches to privacy-sensitive audio features yield much lower recognition accuracies compared to MFCC features. Sree Hari Krishnan Parthasarathi, Hervé Bourlard, Daniel Gatica-Perez |
IEEE Trans. Speech Audio Process. | 3 |
| 2013 | The YouTube Lens: Crowdsourced Personality Impressions and Audiovisual Analysis of VlogsabstractDespite an increasing interest in understanding human perception in social media through the automatic analysis of users' personality, existing attempts have explored user profiles and text blog data only. We approach the study of personality impressions in social media from the novel perspective of crowdsourced impressions, social attention, and audiovisual behavioral analysis on slices of conversational vlogs extracted from YouTube. Conversational vlogs are a unique case study to understand users in social media, as vloggers implicitly or explicitly share information about themselves that words, either written or spoken cannot convey. In addition, research in vlogs may become a fertile ground for the study of video interactions, as conversational video expands to innovative applications. In this work, we first investigate the feasibility of crowdsourcing personality impressions from vlogging as a way to obtain judgements from a variate audience that consumes social media video. Then, we explore how these personality impressions mediate the online video watching experience and relate to measures of attention in YouTube. Finally, we investigate on the use of automatic nonverbal cues as a suitable lens through which impressions are made, and we address the task of automatic prediction of vloggers' personality impressions using nonverbal cues and machine learning techniques. Our study, conducted on a dataset of 442 YouTube vlogs and 2210 annotations collected in Amazon's Mechanical Turk, provides new findings regarding the suitability of collecting personality impressions from crowdsourcing, the types of personality impressions that emerge through vlogging, their association with social attention, and the level of utilization of nonverbal cues in this particular setting. In addition, it constitutes a first attempt to address the task of automatic vlogger personality impression prediction using nonverbal cues, with promising results. Joan-Isaac Biel, Daniel Gatica-Perez |
IEEE Trans. Multim. | 2 |
| 2012 | Contextual conditional models for smartphone-based human mobility predictionabstractHuman behavior is often complex and context-dependent. This paper presents a general technique to exploit this "multidimensional" contextual variable for human mobility prediction. We use an ensemble method, in which we extract different mobility patterns with multiple models and then combine these models under a probabilistic framework. The key idea lies in the assumption that human mobility can be explained by several mobility patterns that depend on a sub-set of the contextual variables and these can be learned by a simple model. We showed how this idea can be applied to two specific online prediction tasks: what is the next place a user will visit? and how long will he stay in the current place?. Using smartphone data collected from 153 users during 17 months, we show the potential of our method in predicting human mobility in real life. Trinh Minh Tri Do, Daniel Gatica-Perez |
UbiComp | 2 |
| 2012 | StressSense: detecting stress in unconstrained acoustic environments using smartphonesabstractStress can have long term adverse effects on individuals' physical and mental well-being. Changes in the speech production process is one of many physiological changes that happen during stress. Microphones, embedded in mobile phones and carried ubiquitously by people, provide the opportunity to continuously and non-invasively monitor stress in real-life situations. We propose StressSense for unobtrusively recognizing stress from human voice using smartphones. We investigate methods for adapting a one-size-fits-all stress model to individual speakers and scenarios. We demonstrate that the StressSense classifier can robustly identify stress across multiple individuals in diverse acoustic environments: using model adaptation StressSense achieves 81% and 76% accuracy for indoor and outdoor environments, respectively. We show that StressSense can be implemented on commodity Android phones and run in real-time. To the best of our knowledge, StressSense represents the first system to consider voice based stress detection and model adaptation in diverse real-life conversational situations using smartphones. Hong Lu 0006, Denise Frauendorfer, Mashfiqui Rabbi, Marianne Schmid Mast, Gokul Chittaranjan, Andrew T. Campbell, Daniel Gatica-Perez, Tanzeem Choudhury |
UbiComp | 7 |
| 2012 | FaceTube: predicting personality from facial expressions of emotion in online conversational videoabstractThe advances in automatic facial expression recognition make possible to mine and characterize large amounts of data, opening a wide research domain on behavioral understanding. In this paper, we leverage the use of a state-of-the-art facial expression recognition technology to characterize users of a popular type of online social video, conversational vlogs. First, we propose the use of several activity cues to characterize vloggers based on frame-by-frame estimates of facial expressions of emotion. Then, we present results for the task of automatically predicting vloggers' personality impressions using facial expressions and the Big-Five traits. Our results are promising, specially for the case of the Extraversion impression, and in addition our work poses interesting questions regarding the representation of multiple natural facial expressions occurring in conversational video. Joan-Isaac Biel, Lucia Teijeiro-Mosquera, Daniel Gatica-Perez |
ICMI | 3 |
| 2012 | Linking speaking and looking behavior patterns with group composition, perception, and performanceabstractThis paper addresses the task of mining typical behavioral patterns from small group face-to-face interactions and linking them to social-psychological group variables. Towards this goal, we define group speaking and looking cues by aggregating automatically extracted cues at the individual and dyadic levels. Then, we define a bag of nonverbal patterns (Bag-of-NVPs) to discretize the group cues. The topics learnt using the Latent Dirichlet Allocation (LDA) topic model are then interpreted by studying the correlations with group variables such as group composition, group interpersonal perception, and group performance. Our results show that both group behavior cues and topics have significant correlations with (and predictive information for) all the above variables. For our study, we use interactions with unacquainted members i.e. newly formed groups. Dinesh Babu Jayagopi, Dairazalia Sanchez-Cortes, Kazuhiro Otsuka, Junji Yamato, Daniel Gatica-Perez |
ICMI | 5 |
| 2012 | Modeling dominance effects on nonverbal behaviors using granger causalityabstractIn this paper we modeled the effects that dominant people might induce on the nonverbal behavior (speech energy and body motion) of the other meeting participants using Granger causality technique. Our initial hypothesis that more dominant people have generalized higher influence was not validated when using the DOME-AMI corpus as data source. However, from the correlational analysis some interesting patterns emerged: contradicting our initial hypothesis dominant individuals are not accounting for the majority of the causal flow in a social interaction. Moreover, they seem to have more intense causal effects as their causal density was significantly higher. Finally dominant individuals tend to respond to the causal effects more often with complementarity than with mimicry. Kyriaki Kalimeri, Bruno Lepri, Oya Aran, Dinesh Babu Jayagopi, Daniel Gatica-Perez, Fabio Pianesi |
ICMI | 5 |
| 2012 | Using self-context for multimodal detection of head nods in face-to-face interactionsabstractHead nods occur in virtually every face-to-face discussion. As part of the backchannel domain, they are not only used to express a 'yes', but also to display interest or enhance communicative attention. Detecting head nods in natural interactions is a challenging task as head nods can be subtle, both in amplitude and duration. In this study, we make use of findings in psychology establishing that the dynamics of head gestures are conditioned on the person's speaking status. We develop a multimodal method using audio-based self-context to detect head nods in natural settings. We demonstrate that our multimodal approach using the speaking status of the person under analysis significantly improved the detection rate over a visual-only approach. Laurent Son Nguyen, Jean-Marc Odobez, Daniel Gatica-Perez |
ICMI | 3 |
| 2012 | The Good, the Bad, and the Angry: Analyzing Crowdsourced Impressions of Vloggers
Joan-Isaac Biel, Daniel Gatica-Perez |
ICWSM | 2 |
| 2012 | Checking in or checked in: comparing large-scale manual and automatic location disclosure patternsabstractStudies on human mobility are built on two fundamentally different data sources: manual check-in data that originates from location-based social networks and automatic check-in data that can be automatically collected through various smartphone sensors. In this paper, we analyze the differences and similarities of manual check-ins from Foursquare and automatic check-ins from Nokia's Mobile Data Challenge. Several new findings follow from our analysis: (1) While automatic checking-in overall results in more visits than manual checking-in, the check-in levels are comparable when visiting new places. (2) Daily and weekly check-in activity patterns are similar for both systems except for Saturdays -- when manual check-ins are relatively more probable. (3) A recently proposed rank distribution to describe human mobility, so far validated on manual check-in data, also holds for automatic check-in data given a slight modification to the definition of rank. (4) The patterns described by automatic check-ins are in general more predictable. We also address the question of whether it is possible to find matching places across the two check-in systems. Our analysis shows that while this is challenging in areas such as city centers, our method achieves an accuracy of 51% for places that are not homes of phone users. Eric Malmi, Trinh Minh Tri Do, Daniel Gatica-Perez |
MUM | 3 |
| 2012 | Assessing the impact of language style on emergent leadership perception from ubiquitous audioabstractLeaders stand out for what they say and how they say it. This work describes the impact of the language style of emergent leaders in small group discussions based on 7 hours of audio from English spoken discussions recorded with a ubiquitous platform. For the language style analysis, word categories are extracted from manual transcriptions of the discussions as well as from automatically detected keywords. The most relevant word categories are then used to predict the emergent leader in each group. Our findings reveal that non-privacy sensitive word categories like amount of words, conjunctions and assent are good predictors of emergent leadership. The emergent leader can be correctly inferred in a fully automatic approach with up to 82% accuracy using categories derived from keywords, and up to 86% using categories derived from full manual transcriptions. Dairazalia Sanchez-Cortes, Petr Motlícek, Daniel Gatica-Perez |
MUM | 3 |
| 2012 | Privacy-sensitive recognition of group conversational context with sociometers
Dinesh Babu Jayagopi, Taemie Jung Kim, Alex Pentland, Daniel Gatica-Perez |
Multim. Syst. | 4 |
| 2012 | Inferring competitive role patterns in reality TV show through nonverbal analysis
Bogdan Raducanu, Daniel Gatica-Perez |
Multim. Tools Appl. | 2 |
| 2012 | A Nonverbal Behavior Approach to Identify Emergent Leaders in Small GroupsabstractIdentifying emergent leaders in organizations is a key issue in organizational behavioral research, and a new problem in social computing. This paper presents an analysis on how an emergent leader is perceived in newly formed, small groups, and then tackles the task of automatically inferring emergent leaders, using a variety of communicative nonverbal cues extracted from audio and video channels. The inference task uses rule-based and collective classification approaches with the combination of acoustic and visual features extracted from a new small group corpus specifically collected to analyze the emergent leadership phenomenon. Our results show that the emergent leader is perceived by his/her peers as an active and dominant person; that visual information augments acoustic information; and that adding relational information to the nonverbal cues improves the inference of each participant's leadership rankings in the group. Dairazalia Sanchez-Cortes, Oya Aran, Marianne Schmid Mast, Daniel Gatica-Perez |
IEEE Trans. Multim. | 4 |
| 2012 | Learning Semantics From Multimedia Web Resources: An Introduction to the Special IssueabstractThe thirteen papers in this special issue focus on effective techniques for learning semantics from multimedia Web resources. Qi Tian 0001, Jinhui Tang 0001, Marcel Worring, Daniel Gatica-Perez |
IEEE Trans. Multim. | 4 |
| 2012 | Introduction to the special section of best papers of ACM multimedia 2011abstractNo abstract available. Daniel Gatica-Perez, Gang Hua 0001, Wei Tsang Ooi, Pål Halvorsen |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2011 | Exploiting observers' judgements for nonverbal group interaction analysisabstractIncorporating annotators' knowledge into a machine-learning framework for detecting psychological traits using multimodal data is an open issue in human communication and social computing. We present a model that is designed to exploit the subjective judgements of multiple annotators on a social trait labeling task. Our two-stage model first estimates a ground truth by modeling the annotators using both the annotations and annotators' self-reported confidences. In the second stage, we train a classifier using the estimated ground truth as labels. We also define ways to verify the consistency of our model and validate it using annotations and nonverbal cues for a dominance estimation task in a group interaction scenario on the publicly available DOME corpus, in addition to synthetically generated data. Our models give satisfactory results, outperforming the commonly used majority voting as well as other approaches in the literature. Gokul Chittaranjan, Oya Aran, Daniel Gatica-Perez |
FG | 3 |
| 2011 | Smartphone usage in the wild: a large-scale analysis of applications and contextabstractThis paper presents a large-scale analysis of contextualized smartphone usage in real life. We introduce two contextual variables that condition the use of smartphone applications, namely places and social context. Our study shows strong dependencies between phone usage and the two contextual cues, which are automatically extracted based on multiple built-in sensors available on the phone. By analyzing continuous data collected on a set of 77 participants from a European country over 9 months of actual usage, our framework automatically reveals key patterns of phone application usage that would traditionally be obtained through manual logging or questionnaire. Our findings contribute to the large-scale understanding of applications and context, bringing out design implications for interfaces on smartphones. Trinh Minh Tri Do, Jan Blom, Daniel Gatica-Perez |
ICMI | 3 |
| 2011 | You Are Known by How You Vlog: Personality Impressions and Nonverbal Behavior in YouTube
Joan-Isaac Biel, Oya Aran, Daniel Gatica-Perez |
ICWSM | 3 |
| 2011 | LP Residual Features for Robust, Privacy-Sensitive Speaker DiarizationabstractWe present a comprehensive study of linear prediction residual for speaker diarization on single and multiple distant microphone conditions in privacy-sensitive settings, a requirement to analyze a wide range of spontaneous conversations. Two representations of the residual are compared, namely real-cepstrum and MFCC, with the latter performing better. Experiments on RT06eval show that residual with subband information from 2.5 kHz to 3.5 kHz and spectral slope yields a performance close to traditional MFCC features. As a way to objectively evaluate privacy in terms of linguistic information, we perform phoneme recognition. Residual features yield low phoneme accuracies compared to traditional MFCC features. Sree Hari Krishnan Parthasarathi, Hervé Bourlard, Daniel Gatica-Perez |
INTERSPEECH | 3 |
| 2011 | Contextual Grouping: Discovering Real-Life Interaction Types from Longitudinal Bluetooth DataabstractBy exploiting built-in sensors, mobile smart phone have become attractive options for large-scale sensing of human behavior as well as social interaction. In this paper, we present a new probabilistic model to analyze longitudinal dynamic social networks created by the physical proximity of people sensed continuously by the phone Bluetooth sensors. A new probabilistic model is proposed in order to jointly infer emergent grouping modes of the community together with their temporal context. We present experimental results on a Bluetooth proximity network sensed with mobile smart-phones over 9 months of continuous real-life, and show the effectiveness of our method. Trinh Minh Tri Do, Daniel Gatica-Perez |
Mobile Data Management (1) | 2 |
| 2011 | People-centric mobile sensing with a pragmatic twist: from behavioral data points to active user involvementabstractMobile phones have recently been used to collect largescale continuous data about human behavior. This people centric sensing paradigm is useful not only from a scientific point of view: Contextual user data has pragmatic value, too. Individuals whose data is collected in such long-term people centric sensing projects can be engaged in user centric design activities aiming to generate data driven services that benefit the end user. This paper demonstrates the value of such user centric approach. In a two-stage approach, we analyse mobile phone data to extract mobile phone usage c ategories. We then go on to interview the participants concerning their perceptions toward context-aware services. The two sta ges, combined as we present here, offer a clear value in ter ms of provi ding complementary insights, both to researchers and users, about the feasibility of and the expectations about personalized mobile services. Jan Blom, Daniel Gatica-Perez, Niko Kiukkonen |
Mobile HCI | 2 |
| 2011 | Searching the past: an improved shape descriptor to retrieve maya hieroglyphsabstractArchaeologists often spend significant time looking at traditional printed catalogs to identify and classify historical images. Our collaborative efforts between archaeologists and multimedia researchers seek to develop a tool to retrieve two specific types of ancient Maya visual information: hieroglyphs and iconographic elements. Towards that goal we present two contributions in this paper. The first one is the introduction and analysis of a new dataset of 3400+ Maya hieroglyphs, whose compilation involved manual search, annotation and segmentation by experts. This dataset presents several challenges for visual description and automatic retrieval as it is rich in complex visual details. The second and main contribution is the in-depth analysis of the Histogram Of Orientation Shape Context (HOOSC), and more precisely, the development of 4 improvements that were designed to handle the visual complexity of Maya hieroglyphs: open contours, mixture of thick and thin lines, hatches, large instance variability, and a variety of internal details. Experiments demonstrate that the adequate combination of our improvements to retrieve Maya hieroglyphs, provides results with roughly 20% more precision compared to the original HOOSC descriptor. Complementary results with the MPEG-7 shape dataset validate (or not) the proposed improvements, showing that the design of appropriate descriptors depends on the nature of the shapes one deals with. Edgar Roman-Rangel, Carlos Pallan, Jean-Marc Odobez, Daniel Gatica-Perez |
ACM Multimedia | 4 |
| 2011 | Analyzing Ancient Maya Glyph Collections with Contextual Shape Descriptors
Edgar Roman-Rangel, Carlos Pallan, Jean-Marc Odobez, Daniel Gatica-Perez |
Int. J. Comput. Vis. | 4 |
| 2011 | Estimating Dominance in Multi-Party Meetings Using Speaker DiarizationabstractWith the increase in cheap commercially available sensors, recording meetings is becoming an increasingly practical option. With this trend comes the need to summarize the recorded data in semantically meaningful ways. Here, we investigate the task of automatically measuring dominance in small group meetings when only a single audio source is available. Past research has found that speaking length as a single feature, provides a very good estimate of dominance. For these tasks we use speaker segmentations generated by our automated faster than real-time speaker diarization algorithm, where the number of speakers is not known beforehand. From user-annotated data, we analyze how the inherent variability of the annotations affects the performance of our dominance estimation method. We primarily focus on examining of how the performance of the speaker diarization and our dominance tasks vary under different experimental conditions and computationally efficient strategies, and how this would impact on a practical implementation of such a system. Despite the use of a state-of-the-art speaker diarization algorithm, speaker segments can be noisy. On conducting experiments on almost 5 hours of audio-visual meeting data, our results show that the dominance estimation is robust to increasing diarization noise. Hayley Hung, Gerald Friedland, Daniel Gatica-Perez |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Privacy-Sensitive Audio Features for Speech/Nonspeech DetectionabstractThe goal of this paper is to investigate features for speech/nonspeech detection (SND) having low linguistic information from the speech signal. Towards this, we present a comprehensive study of privacy-sensitive features for SND in multiparty conversations. Our study investigates three different approaches to privacy-sensitive features. These approaches are based on: 1) simple, instantaneous feature extraction methods; 2) excitation source information based methods; and 3) feature obfuscation methods such as local (within 130 ms) temporal averaging and randomization applied on excitation source information. To evaluate these approaches for SND, we use multiparty conversational meeting data of nearly 450 hours. On this dataset, we evaluate these features and benchmark them against standard spectral shape based features such as Mel frequency perceptual linear prediction (MFPLP). Fusion strategies combining excitation source with simple features show that comparable performance can be obtained in both close-talking and far-field microphone scenarios. As one way to objectively evaluate the notion of privacy, we conduct phoneme recognition studies on TIMIT. While excitation source features yield phoneme recognition accuracies in between the simple features and the MFPLP features, obfuscation methods applied on the excitation features yield low phoneme accuracies in conjunction with SND performance comparable to that of MFPLP features. Sree Hari Krishnan Parthasarathi, Daniel Gatica-Perez, Hervé Bourlard, Mathew Magimai-Doss |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2011 | Discovering routines from large-scale human locations using probabilistic topic modelsabstractIn this work, we discover the daily location-driven routines that are contained in a massive real-life human dataset collected by mobile phones. Our goal is the discovery and analysis of human routines that characterize both individual and group behaviors in terms of location patterns. We develop an unsupervised methodology based on two differing probabilistic topic models and apply them to the daily life of 97 mobile phone users over a 16-month period to achieve these goals. Topic models are probabilistic generative models for documents that identify the latent structure that underlies a set of words. Routines dominating the entire group's activities, identified with a methodology based on the Latent Dirichlet Allocation topic model, include “going to work late”, “going home early”, “working nonstop” and “having no reception (phone off)” at different times over varying time-intervals. We also detect routines which are characteristic of users, with a methodology based on the Author-Topic model. With the routines discovered, and the two methods of characterizing days and users, we can then perform various tasks. We use the routines discovered to determine behavioral patterns of users and groups of users. For example, we can find individuals that display specific daily routines, such as “going to work early” or “turning off the mobile (or having no reception) in the evenings”. We are also able to characterize daily patterns by determining the topic structure of days in addition to determining whether certain routines occur dominantly on weekends or weekdays. Furthermore, the routines discovered can be used to rank users or find subgroups of users who display certain routines. We can also characterize users based on their entropy. We compare our method to one based on clustering using K-means. Finally, we analyze an individual's routines over time to determine regions with high variations, which may correspond to specific events. Katayoun Farrahi, Daniel Gatica-Perez |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2011 | VlogSense: Conversational behavior and social attention in YouTubeabstractWe introduce the automatic analysis of conversational vlogs (VlogSense, for short) as a new research domain in social media. Conversational vlogs are inherently multimodal, depict natural behavior, and are suitable for large-scale analysis. Given their diversity in terms of content, VlogSense requires the integration of robust methods for multimodal analysis and for social media understanding. We present an original study on the automatic characterization of vloggers' audiovisual nonverbal behavior, grounded in work from social psychology and behavioral computing. Our study on 2,269 vlogs from YouTube shows that several nonverbal cues are significantly correlated with the social attention received by videos. Joan-Isaac Biel, Daniel Gatica-Perez |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2010 | Evaluating the robustness of privacy-sensitive audio features for speech detection in personal audio log scenariosabstractPersonal audio logs are often recorded in multiple environments. This poses challenges for robust front-end processing, including speech/nonspeech detection (SND). Motivated by this, we investigate the robustness of four different privacy-sensitive features for SND, namely energy, zero crossing rate, spectral flatness, and kurtosis. We study early and late fusion of these features in conjunction with modeling temporal context. These combinations are evaluated in mismatched conditions on a dataset of nearly 450 hours. While both combinations yield improvements over individual features, generally feature combinations perform better. Comparisons with a state-of-the-art spectral based and a privacy-sensitive feature set are also provided. Sree Hari Krishnan Parthasarathi, Mathew Magimai-Doss, Hervé Bourlard, Daniel Gatica-Perez |
ICASSP | 4 |
| 2010 | Fusing Audio-Visual Nonverbal Cues to Detect Dominant People in Group ConversationsabstractThis paper addresses the multimodal nature of social dominance and presents multimodal fusion techniques to combine audio and visual nonverbal cues for dominance estimation in small group conversations. We combine the two modalities both at the feature extraction level and at the classifier level via score and rank level fusion. The classification is done by a simple rule-based estimator. We perform experiments on a new 10-hour dataset derived from the popular AMI meeting corpus. We objectively evaluate the performance of each modality and each cue alone and in combination. Our results show that the combination of audio and visual cues is necessary to achieve the best performance. Oya Aran, Daniel Gatica-Perez |
ICPR | 2 |
| 2010 | Voices of Vlogging
Joan-Isaac Biel, Daniel Gatica-Perez |
ICWSM | 2 |
| 2010 | All things mobile: the present and future of mobile phone computingabstractThis is the summary of the panel All Things Mobile: The Present and Future of Mobile Phone Computing. Daniel Gatica-Perez |
ACM Multimedia | 1 |
| 2010 | Modeling human behavior with mobile phonesabstractIn just a few years, mobile phones have emerged as the ultimate multimedia device. This is the summary of a proposed tutorial on Modeling Human Behavior with Mobile Phones, which aims to present the scientific and technological state-of-the-art in mobile phone-based modeling of large-scale human behavior from a coherent perspective, and hopes to motivate further work in this domain by the multimedia research community. Daniel Gatica-Perez |
ACM Multimedia | 1 |
| 2010 | Kodak moments and Flickr diamonds: how users shape large-scale mediaabstractIn today's age of digital multimedia deluge, a clear understanding of the dynamics of online communities is capital. Users have abandoned their role of passive consumers and are now the driving force behind large-scale media repositories, whose dynamics and shaping factors are not yet fully understood. In this paper we present a novel human-centered analysis of two major photo sharing websites, Flickr and Kodak Gallery. On a combined dataset of over 5 million tagged photos, we investigate fundamental differences and similarities at the level of tag usage and propose a joint probabilistic topic model to provide further insight into semantic differences between the two communities. Our results show that the effects of the users' motivations and needs can be strongly observed in this large-scale data, in the form of what we call Kodak Moments and Flickr Diamonds. They are an indication that system designers should carefully take into account the target audience and its needs. Radu Andrei Negoescu, Alexander C. Loui, Daniel Gatica-Perez |
ACM Multimedia | 3 |
| 2010 | By their apps you shall understand them: mining large-scale patterns of mobile phone usageabstractMobile phones are becoming more and more widely used nowadays, and people do not use the phone only for communication: there is a wide variety of phone applications allowing users to select those that fit their needs. Aggregated over time, application usage patterns exhibit not only what people are consistently interested in but also the way in which they use their phones, and can help improving phone design and personalized services. This work aims at mining automatically usage patterns from apps data recorded continuously with smartphones. A new probabilistic framework for mining usage patterns is proposed. Our methodology involves the design of a bag-of-apps model that robustly represents level of phone usage over specific times of the day, and the use of a probabilistic topic model that jointly discovers patterns of usage over multiple applications and describes users as mixtures of such patterns. Our framework is evaluated using 230 000+ hours of real-life app phone log data, demonstrates that relevant patterns of usage can be extracted, and is objectively validated on a user retrieval task with competitive performance. Trinh Minh Tri Do, Daniel Gatica-Perez |
MUM | 2 |
| 2010 | Recognizing conversational context in group interaction using privacy-sensitive mobile sensorsabstractThe availability of mobile sociometric sensors allows Computer-Supported Cooperative Work (CSCW) designers the possibility to enhance online meeting support through automatic recognition of conversational context. This paper addresses the task of discriminating one conversational context against another, specifically brainstorming from decision-making interactions using easily computable nonverbal behavioral cues. We hypothesize that the difference in the dynamics between brainstorming and decision-making discussions is significant and measurable using speech activity based nonverbal cues. We employ a set of nonverbal cues to characterize the entire group by the aggregation (both temporal and person-wise) of their nonverbal behavior. Our results on a dataset collected using privacy-sensitive sociometric badges show that the floor-occupation patterns in a brain-storming interaction are different from a decision-making interaction and we can obtain a discrimination accuracy as high as 87.5%. Dinesh Babu Jayagopi, Taemie Jung Kim, Alex Pentland, Daniel Gatica-Perez |
MUM | 4 |
| 2010 | Discovering human places of interest from multimodal mobile phone dataabstractIn this paper, a new framework to discover places-of-interest from multimodal mobile phone data is presented. Mobile phones have been used as sensors to obtain location information from users’ real lives. Two levels of clustering are used to obtain places of interest. First, user location points are grouped using a time-based clustering technique which discovers stay points while dealing with missing location data. The second level performs clustering on the stay points to obtain stay regions. A grid-based clustering algorithm has been used for this purpose. To obtain more user location points, a client-server system has been installed on the mobile phones, which is able to obtain location information by integrating GPS, Wifi, GSM and accelerometer sensors, among others. An extensive set of experiments have been performed to show the benefits of using the proposed framework, using data from the real life of 8 users over 5 continuous months of natural phone usage. Raúl Montoliu, Daniel Gatica-Perez |
MUM | 2 |
| 2010 | Estimating Cohesion in Small Groups Using Audio-Visual Nonverbal BehaviorabstractCohesiveness in teams is an essential part of ensuring the smooth running of task-oriented groups. Research in social psychology and management has shown that good cohesion in groups can be correlated with team effectiveness or productivity, so automatically estimating group cohesion for team training can be a useful tool. This paper addresses the problem of analyzing group behavior within the context of cohesion. Four hours of audio-visual group meeting data were used for collecting annotations on the cohesiveness of four-participant teams. We propose a series of audio and video features, which are inspired by findings in the social sciences literature. Our study is validated on a set of 61 2-min meeting segments which showed high agreement amongst human annotators when asked to identify meetings that have high or low cohesion. Hayley Hung, Daniel Gatica-Perez |
IEEE Trans. Multim. | 2 |
| 2010 | Mining Group Nonverbal Conversational Patterns Using Probabilistic Topic ModelsabstractThe automatic discovery of group conversational behavior is a relevant problem in social computing. In this paper, we present an approach to address this problem by defining a novel group descriptor called bag of group-nonverbal-patterns (NVPs) defined on brief observations of group interaction, and by using principled probabilistic topic models to discover topics. The proposed bag of group NVPs allows fusion of individual cues and facilitates the eventual comparison of groups of varying sizes. The use of topic models helps to cluster group interactions and to quantify how different they are from each other in a formal probabilistic sense. Results of behavioral topics discovered on the Augmented Multi-Party Interaction (AMI) meeting corpus are shown to be meaningful using human annotation with multiple observers. Our method facilitates “group behavior-based” retrieval of group conversational segments without the need of any previous labeling. Dinesh Babu Jayagopi, Daniel Gatica-Perez |
IEEE Trans. Multim. | 2 |
| 2010 | Modeling Flickr Communities Through Probabilistic Topic-Based AnalysisabstractWith the increased presence of digital imaging devices, there also came an explosion in the amount of multimedia content available online. Users have transformed from passive consumers of media into content creators and have started organizing themselves in and around online communities. Flickr has more than 30 million users and over 3 billion photos, and many of them are tagged and public. One very important aspect in Flickr is the ability of users to organize in self-managed communities called groups. This paper examines an unexplored problem, which is jointly analyzing Flickr groups and users. We show that although users and groups are conceptually different, in practice they can be represented in a similar way via a bag-of-tags derived from their photos, which is amenable for probabilistic topic modeling. We then propose a probabilistic topic model representation learned in an unsupervised manner that allows the discovery of similar users and groups beyond direct tag-based strategies, and we demonstrate that higher-level information such as topics of interest are a viable alternative. On a dataset containing users of 10 000 Flickr groups and over 1 milion photos, we show how this common topic-based representation allows for a novel analysis of the groups-users Flickr ecosystem, which results into new insights about the structure of the entities in this social media source. We demonstrate novel practical applications of our topic-based representation, such as similarity-based exploration of entities, or single and multi-topic tag-based search, which address current limitations in the ways Flickr is used today. Radu Andrei Negoescu, Daniel Gatica-Perez |
IEEE Trans. Multim. | 2 |
| 2009 | You are fired! Nonverbal role analysis in competitive meetingsabstractThis paper addresses the problem of social interaction analysis in competitive meetings, using nonverbal cues. For our study, we made use of ldquoThe Apprenticerdquo reality TV show, which features a competition for a real, highly paid corporate job. Our analysis is centered around two tasks regarding a person's role in a meeting: predicting the person with the highest status and predicting the fired candidates. The current study was carried out using nonverbal audio cues. Results obtained from the analysis of a full season of the show, representing around 90 minutes of audio data, are very promising (up to 85.7% of accuracy in the first case and up to 92.8% in the second case). Our approach is based only on the nonverbal interaction dynamics during the meeting without relying on the spoken words. Bogdan Raducanu, Jordi Vitrià, Daniel Gatica-Perez |
ICASSP | 3 |
| 2009 | Characterizing conversational group dynamics using nonverbal behaviourabstractThis paper addresses the novel problem of characterizing conversational group dynamics. It is well documented in social psychology that depending on the objectives a group, the dynamics are different. For example, a competitive meeting has a different objective from that of a collaborative meeting. We propose a method to characterize group dynamics based on the joint description of a group members' aggregated acoustical nonverbal behaviour to classify two meeting datasets (one being cooperative-type and the other being competitive-type). We use 4.5 hours of real behavioural multi-party data and show that our methodology can achieve a classification rate of upto 100%. Dinesh Babu Jayagopi, Bogdan Raducanu, Daniel Gatica-Perez |
ICME | 3 |
| 2009 | Learning and predicting multimodal daily life patterns from cell phonesabstractIn this paper, we investigate the multimodal nature of cell phone data in terms of discovering recurrent and rich patterns in people's lives. We present a method that can discover routines from multiple modalities (location and proximity) jointly modeled, and that uses these informative routines to predict unlabeled or missing data. Using a joint representation of location and proximity data over approximately 10 months of 97 individuals' lives, Latent Dirichlet Allocation is applied for the unsupervised learning of topics describing people's most common locations jointly with the most common types of interactions at these locations. We further successfully predict where and with how many other individuals users will be, for people with both highly and lowly varying lifestyles. Katayoun Farrahi, Daniel Gatica-Perez |
ICMI | 2 |
| 2009 | Discovering group nonverbal conversational patterns with topicsabstractThis paper addresses the problem of discovering conversational group dynamics from nonverbal cues extracted from thin-slices of interaction. We first propose and analyze a novel thin-slice interaction descriptor - a bag of group nonverbal patterns - which robustly captures the turn-taking behavior of the members of a group while integrating its leader's position. We then rely on probabilistic topic modeling of the interaction descriptors which, in a fully unsupervised way, is able to discover group interaction patterns that resemble prototypical leadership styles proposed in social psychology. Our method, validated on the Augmented Multi-Party Interaction (AMI) meeting corpus, facilitates the retrieval of group conversational segments where semantically meaningful group behaviours emerge, without the need of any previous labeling. Dinesh Babu Jayagopi, Daniel Gatica-Perez |
ICMI | 2 |
| 2009 | Speaker change detection with privacy-preserving audio cuesabstractIn this paper we investigate a set of privacy-sensitive audio features for speaker change detection (SCD) in multiparty conversations. These features are based on three different principles: characterizing the excitation source information using linear prediction residual, characterizing subband spectral information shown to contain speaker information, and characterizing the general shape of the spectrum. Experiments show that the performance of the privacy-sensitive features is comparable or better than that of the state-of-the-art full-band spectral-based features, namely, mel frequency cepstral coefficients, which suggests that socially acceptable ways of recording conversations in real-life is feasible. Sree Hari Krishnan Parthasarathi, Mathew Magimai-Doss, Daniel Gatica-Perez, Hervé Bourlard |
ICMI | 3 |
| 2009 | Investigating privacy-sensitive features for speech detection in multiparty conversationsabstractWe investigate four different privacy-sensitive features, namely energy, zero crossing rate, spectral flatness, and kurtosis, for speech detection in multiparty conversations. We liken this scenario to a meeting room and define our datasets and annotations accordingly. The temporal context of these features is modeled. With no temporal context, energy is the best performing single feature. But by modeling temporal context, kurtosis emerges as the most effective feature. Also, we combine the features. Besides yielding a gain in performance, certain combinations of features also reveal that a shorter temporal context is sufficient. We then benchmark other privacy-sensitive features utilized in previous studies. Our experiments show that the performance of all the privacy-sensitive features modeled with context is close to that of state-of-the-art spectral-based features, without extracting and using any features that can be used to reconstruct the speech signal. Sree Hari Krishnan Parthasarathi, Mathew Magimai-Doss, Hervé Bourlard, Daniel Gatica-Perez |
INTERSPEECH | 4 |
| 2009 | Wearing a YouTube hat: directors, comedians, gurus, and user aggregated behaviorabstractWhile existing studies on YouTube's massive user-generated video content have mostly focused on the analysis of videos, their characteristics, and network properties, little attention has been paid to the analysis of users' long-term behavior as it relates to the roles they self-define and (explicitly or not) play in the site. In this paper, we present a novel statistical analysis of aggregated user behavior in YouTube from the novel perspective of user categories, a feature that allows people to ascribe to popular roles and to potentially reach certain communities. Using a sample of 270,000 users, we found that a high level of interaction and participation is concentrated on a relatively small, yet significant, group of users, following recognizable patterns of personal and social involvement. Based on our analysis, we also show that by using simple behavioral features from user profiles, people can be automatically classified according to their category with accuracy rates of up to 73%. Joan-Isaac Biel, Daniel Gatica-Perez |
ACM Multimedia | 2 |
| 2009 | Flickr hypergroupsabstractThe amount of multimedia content available online constantly increases, and this leads to problems for users who search for content or similar communities. Users in Flickr often self-organize in user communities through Flickr Groups. These groups are particularly interesting as they are a natural instantiation of the content~+~relations social media paradigm. We propose a novel approach to group searching through hypergroup discovery. Starting from roughly 11,000 Flickr groups' content and membership information, we create three different bag-of-word representations for groups, on which we learn probabilistic topic models. Finally, we cast the hypergroup discovery as a clustering problem that is solved via probabilistic affinity propagation. We show that hypergroups so found are generally consistent and can be described through topic-based and similarity-based measures. Our proposed solution could be relatively easily implemented as an application to enrich Flickr's traditional group search. Radu Andrei Negoescu, Brett Adams, Dinh Q. Phung, Svetha Venkatesh, Daniel Gatica-Perez |
ACM Multimedia | 5 |
| 2009 | Automatic nonverbal analysis of social interaction in small groups: A review
Daniel Gatica-Perez |
Image Vis. Comput. | 1 |
| 2009 | Modeling Dominance in Group Conversations Using Nonverbal Activity CuesabstractDominance - a behavioral expression of power - is a fundamental mechanism of social interaction, expressed and perceived in conversations through spoken words and audiovisual nonverbal cues. The automatic modeling of dominance patterns from sensor data represents a relevant problem in social computing. In this paper, we present a systematic study on dominance modeling in group meetings from fully automatic nonverbal activity cues, in a multi-camera, multi-microphone setting. We investigate efficient audio and visual activity cues for the characterization of dominant behavior, analyzing single and joint modalities. Unsupervised and supervised approaches for dominance modeling are also investigated. Activity cues and models are objectively evaluated on a set of dominance-related classification tasks, derived from an analysis of the variability of human judgment of perceived dominance in group discussions. Our investigation highlights the power of relatively simple yet efficient approaches and the challenges of audiovisual integration. This constitutes the most detailed study on automatic dominance modeling in meetings to date. Dinesh Babu Jayagopi, Hayley Hung, Chuohao Yeo, Daniel Gatica-Perez |
IEEE Trans. Speech Audio Process. | 4 |
| 2008 | Identifying dominant people in meetings from audio-visual sensorsabstractThis paper provides an overview of the area of automated dominance estimation in group meetings. We describe research in social psychology and use this to explain the motivations behind suggested automated systems. With the growth in availability of conversational data captured in meeting rooms, it is possible to investigate how multi-sensor data allows us to characterize non-verbal behaviors that contribute towards dominance. We use an overview of our own work to address the challenges and opportunities in this area of research. Hayley Hung, Daniel Gatica-Perez |
FG | 2 |
| 2008 | Estimating the dominant person in multi-party conversations using speaker diarization strategiesabstractIn this paper, we apply speaker diarization strategies from a single source to the task of estimating the dominant person in a group meeting. Previous work has shown that speaking length is strongly correlated with perceived dominance. Here we investigate this in more depth by considering two dominance tasks where there is full and majority agreement amongst ground-truth annotators. In addition, we investigate how 24 different speed-up and algorithmic strategies, and source types lead to interesting outcomes when applied to dominance estimation. We obtained the best performance of 77% using our slowest scheme and a single distant microphone (SDM). Within the top 3 out of 24 performing experiments in both dominance tasks, we show that we can use the furthest SDM, with no prior knowledge of the number of speakers and the fastest diarization scheme, which performs 1.3 times faster than real-time. Hayley Hung, Gerald Friedland, Daniel Gatica-Perez |
ICASSP | 4 |
| 2008 | Investigating automatic dominance estimation in groups from visual attention and speaking activityabstractWe study the automation of the visual dominance ratio (VDR); a classic measure of displayed dominance in social psychology literature, which combines both gaze and speaking activity cues. The VDR is modified to estimate dominance in multi-party group discussions where natural verbal exchanges are possible and other visual targets such as a table and slide screen are present. Our findings suggest that fully automated versions of these measures can estimate effectively the most dominant person in a meeting and can match the dominance estimation performance when manual labels of visual attention are used. Hayley Hung, Dinesh Babu Jayagopi, Sileye O. Ba, Jean-Marc Odobez, Daniel Gatica-Perez |
ICMI | 5 |
| 2008 | Predicting two facets of social verticality in meetings from five-minute time slices and nonverbal cuesabstractThis paper addresses the automatic estimation of two aspects of social verticality (status and dominance) in small-group meetings using nonverbal cues. The correlation of nonverbal behavior with these social constructs have been extensively documented in social psychology, but their value for computational models is, in many cases, still unknown. We present a systematic study of automatically extracted cues - including vocalic, visual activity, and visual attention cues - and investigate their relative effectiveness to predict both the most-dominant person and the high-status project manager from relative short observations. We use five hours of task-oriented meeting data with natural behavior for our experiments. Our work suggests that, although dominance and role-based status are related concepts, they are not equivalent and are thus not equally explained by the same nonverbal cues. Furthermore, the best cues can correctly predict the person with highest dominance or role-based status with an accuracy of 70% approximately. Dinesh Babu Jayagopi, Sileye O. Ba, Jean-Marc Odobez, Daniel Gatica-Perez |
ICMI | 4 |
| 2008 | What did you do today?: discovering daily routines from large-scale mobile dataabstractWe present a framework built from two Hierarchical Bayesian topic models to discover human location-driven routines from mobile phones. The framework uses location-driven bag representations of people's daily activities obtained from celltower connections. Using 68 000+ hours of real-life human data from the Reality Mining dataset, we successfully discover various types of routines. The first studied model, Latent Dirichlet Allocation (LDA), automatically discovers characteristic routines for all individuals in the study, including "going to work at 10am", "leaving work at night", or "staying home for the entire evening". In contrast, the second methodology with the Author Topic model (ATM) finds routines characteristic of a selected groups of users, such as "being at home in the mornings and evenings while being out in the afternoon", and ranks users by their probability of conforming to certain daily routines. Katayoun Farrahi, Daniel Gatica-Perez |
ACM Multimedia | 2 |
| 2008 | Predicting the dominant clique in meetings through fusion of nonverbal cuesabstractThis paper addresses the problem of automatically predicting the dominant clique (i.e., the set of K-dominant people) in face-to-face small group meetings recorded by multiple audio and video sensors. For this goal, we present a framework that integrates automatically extracted nonverbal cues and dominance prediction models. Easily computable audio and visual activity cues are automatically extracted from cameras and microphones. Such nonverbal cues, correlated to human display and perception of dominance, are well documented in the social psychology literature. The effectiveness of the cues were systematically investigated as single cues as well as in unimodal and multimodal combinations using unsupervised and supervised learning approaches for dominant clique estimation. Our framework was evaluated on a five-hour public corpus of teamwork meetings with third-party manual annotation of perceived dominance. Our best approaches can exactly predict the dominant clique with 80.8% accuracy in four-person meetings in which multiple human annotators agree on their judgments of perceived dominance. Dinesh Babu Jayagopi, Hayley Hung, Chuohao Yeo, Daniel Gatica-Perez |
ACM Multimedia | 4 |
| 2008 | Topickr: flickr groups and users reloadedabstractWith the increased presence of digital imaging devices there also came an explosion in the amount of multimedia content available online. Users have transformed from passive consumers of media into content creators. Flickr.com is such an example of an online community, with over 2 billion photos (and more recently, videos as well), most of which are publicly available. The user interaction with the system also provides a plethora of metadata associated with this content, and in particular tags. One very important aspect in Flickr is the ability of users to organize in self-managed communities called groups. Although users and groups are conceptually different, in practice they can be represented in the same way: a bag-of-tags, which is amenable for probabilistic topic modeling. We present a topic-based approach to represent Flickr users and groups and demonstrate it with a web application, Topickr, that allows similarity based exploration of Flickr entities using their topic-based representation, learned in an unsupervised manner. Radu Andrei Negoescu, Daniel Gatica-Perez |
ACM Multimedia | 2 |
| 2008 | Tracking the Visual Focus of Attention for a Varying Number of Wandering PeopleabstractWe define and address the problem of finding the visual focus of attention for a varying number of wandering people (VFOA-W), determining where the people's movement is unconstrained. VFOA-W estimation is a new and important problem with mplications for behavior understanding and cognitive science, as well as real-world applications. One such application, which we present in this article, monitors the attention passers-by pay to an outdoor advertisement. Our approach to the VFOA-W problem proposes a multi-person tracking solution based on a dynamic Bayesian network that simultaneously infers the (variable) number of people in a scene, their body locations, their head locations, and their head pose. For efficient inference in the resulting large variable-dimensional state-space we propose a Reversible Jump Markov Chain Monte Carlo (RJMCMC) sampling scheme, as well as a novel global observation model which determines the number of people in the scene and localizes them. We propose a Gaussian Mixture Model (GMM) and Hidden Markov Model (HMM)-based VFOA-W model which use head pose and location information to determine people's focus state. Our models are evaluated for tracking performance and ability to recognize people looking at an outdoor advertisement, with results indicating good performance on sequences where a moderate number of people pass in front of an advertisement. Kevin Smith 0001, Sileye O. Ba, Jean-Marc Odobez, Daniel Gatica-Perez |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2007 | Unsupervised Speech/Non-Speech Detection for Automatic Speech Recognition in Meeting RoomsabstractThe goal of this work is to provide robust and accurate speech detection for automatic speech recognition (ASR) in meeting room settings. The solution is based on computing long-term modulation spectrum, and examining specific frequency range for dominant speech components to classify speech and non-speech signals for a given audio signal. Manually segmented speech segments, short-term energy, short-term energy and zero-crossing based segmentation techniques, and a recently proposed multi layer perceptron (MLP) classifier system are tested for comparison purposes. Speech recognition evaluations of the segmentation methods are performed on a standard database and tested in conditions where the signal-to-noise ratio (SNR) varies considerably, as in the cases of close-talking headset, lapel, distant microphone array output, and distant microphone. The results reveal that the proposed method is more reliable and less sensitive to mode of signal acquisition and unforeseen conditions. Hari Krishna Maganti, Petr Motlícek, Daniel Gatica-Perez |
ICASSP (4) | 3 |
| 2007 | Using audio and video features to classify the most dominant person in a group meetingabstractThe automated extraction of semantically meaningful information from multi-modal data is becoming increasingly necessary due to the escalation of captured data for archival. A novel area of multi-modal data labelling, which has received relatively little attention, is the automatic estimation of the most dominant person in a group meeting. In this paper, we provide a framework for detecting dominance in group meetings using different audio and video cues. We show that by using a simple model for dominance estimation we can obtain promising results. Hayley Hung, Dinesh Babu Jayagopi, Chuohao Yeo, Gerald Friedland, Sileye O. Ba, Jean-Marc Odobez, Kannan Ramchandran, Nikki Mirghafori, Daniel Gatica-Perez |
ACM Multimedia | 9 |
| 2007 | Modeling Semantic Aspects for Cross-Media Image IndexingabstractTo go beyond the query-by-example paradigm in image retrieval, there is a need for semantic indexing of large image collections for intuitive text-based image search. Different models have been proposed to learn the dependencies between the visual content of an image set and the associated text captions, then allowing for the automatic creation of semantic indices for unannotated images. The task, however, remains unsolved. In this paper, we present three alternatives to learn a Probabilistic Latent Semantic Analysis model (PLSA) for annotated images, and evaluate their respective performance for automatic image indexing. Under the PLSA assumptions, an image is modeled as a mixture of latent aspects that generates both image features and text captions, and we investigate three ways to learn the mixture of aspects. We also propose a more discriminative image representation than the traditional Blob histogram, concatenating quantized local color information and quantized local texture descriptors. The first learning procedure of a PLSA model for annotated images is a standard EM algorithm, which implicitly assumes that the visual and the textual modalities can be treated equivalently. The other two models are based on an asymmetric PLSA learning, allowing to constrain the definition of the latent space on the visual or on the textual modality. We demonstrate that the textual modality is more appropriate to learn a semantically meaningful latent space, which translates into improved annotation performance. A comparison of our learning algorithms with respect to recent methods on a standard dataset is presented, and a detailed evaluation of the performance shows the validity of our framework. Florent Monay, Daniel Gatica-Perez |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | A Thousand Words in a SceneabstractThis paper presents a novel approach for visual scene modeling and classification, investigating the combined use of text modeling methods and local invariant features. Our work attempts to elucidate (1) whether a text-like bag-of-visterms representation (histogram of quantized local visual features) is suitable for scene (rather than object) classification, (2) whether some analogies between discrete scene representations and text documents exist, and (3) whether unsupervised, latent space models can be used both as feature extractors for the classification task and to discover patterns of visual co-occurrence. Using several data sets, we validate our approach, presenting and discussing experiments on each of these issues. We first show, with extensive experiments on binary and multi-class scene classification tasks using a 9,500-image data set, that the bag-of-visterms representation consistently outperforms classical scene classification approaches. In other data sets we show that our approach competes with or outperforms other recent, more complex, methods. We also show that Probabilistic Latent Semantic Analysis (PLSA) generates a compact scene representation, discriminative for accurate classification, and more robust than the bag-of-visterms representation when less labeled training data is available. Finally, through aspect-based image ranking experiments, we show the ability of PLSA to automatically extract visually meaningful scene patterns, making such representation useful for browsing image collections. Pedro Quelhas, Florent Monay, Jean-Marc Odobez, Daniel Gatica-Perez, Tinne Tuytelaars |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2007 | Audiovisual Probabilistic Tracking of Multiple Speakers in MeetingsabstractTracking speakers in multiparty conversations constitutes a fundamental task for automatic meeting analysis. In this paper, we present a novel probabilistic approach to jointly track the location and speaking activity of multiple speakers in a multisensor meeting room, equipped with a small microphone array and multiple uncalibrated cameras. Our framework is based on a mixed-state dynamic graphical model defined on a multiperson state-space, which includes the explicit definition of a proximity-based interaction model. The model integrates audiovisual (AV) data through a novel observation model. Audio observations are derived from a source localization algorithm. Visual observations are based on models of the shape and spatial structure of human heads. Approximate inference in our model, needed given its complexity, is performed with a Markov Chain Monte Carlo particle filter (MCMC-PF), which results in high sampling efficiency. We present results-based on an objective evaluation procedure-that show that our framework 1) is capable of locating and tracking the position and speaking activity of multiple meeting participants engaged in real conversations with good accuracy, 2) can deal with cases of visual clutter and occlusion, and 3) significantly outperforms a traditional sampling-based approach Daniel Gatica-Perez, Guillaume Lathoud, Jean-Marc Odobez, Iain McCowan |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Speech Enhancement and Recognition in Meetings With an Audio-Visual Sensor ArrayabstractThis paper addresses the problem of distant speech acquisition in multiparty meetings, using multiple microphones and cameras. Microphone array beamforming techniques present a potential alternative to close-talking microphones by providing speech enhancement through spatial filtering. Beamforming techniques, however, rely on knowledge of the speaker location. In this paper, we present an integrated approach, in which an audio– visual multiperson tracker is used to track active speakers with high accuracy. Speech enhancement is then achieved using microphone array beamforming followed by a novel postfiltering stage. Finally, speech recognition is performed to evaluate the quality of the enhanced speech signal. The approach is evaluated on data recorded in a real meeting room for stationary speaker, moving speaker, and overlapping speech scenarios. The results show that the speech enhancement and recognition performance achieved using our approach are significantly better than a single table-top microphone and are comparable to a lapel microphone for some of the scenarios. The results also indicate that the audio–visual-based system performs significantly better than audio-only system, both in terms of enhancement and recognition. This reveals that the accurate speaker tracking provided by the audio–visual sensor array proved beneficial to improve the recognition performance in a microphone array-based speech recognition system. Hari Krishna Maganti, Daniel Gatica-Perez, Iain McCowan |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Modeling Interactions from Email CommunicationabstractE-mail plays an important role as a medium for the spread of information, ideas, and influence among its users. We present a framework to learn topic-based interactions between pairs of E-mail users, i.e., the extent to which the E-mail topic dynamics of one user are likely to be affected by the others. The proposed framework is built on the influence model and the probabilistic latent semantic analysis (PLSA) language model. This paper makes two contributions. First, we model interactions between E-mail users using the semantic content of E-mail body, instead of E-mail header. Second, our framework models not only E-mail topic dynamics of individual E-mail users, but also the interactions within a group of individuals. Experiments on the Enron E-mail corpus show some interesting results that are potentially useful to discover the hierarchy of the Enron organization Dong Zhang 0001, Daniel Gatica-Perez, Deb Roy, Samy Bengio |
ICME | 2 |
| 2006 | Speaker localization for microphone array-based ASR: the effects of accuracy on overlapping speechabstractAccurate speaker location is essential for optimal performance of distant speech acquisition systems using microphone array techniques. However, to the best of our knowledge, no comprehensive studies on the degradation of automatic speech recognition (ASR) as a function of speaker location accuracy in a multi-party scenario exist. In this paper, we describe a framework for evaluation of the effects of speaker location errors on a microphone array-based ASR system, in the context of meetings in multi-sensor rooms comprising multiple cameras and microphones. Speakers are manually annotated in videos in different camera views, and triangulation is used to determine an accurate speaker location. Errors in the speaker location are then induced in a systematic manner to observe their influence on speech recognition performance. The system is evaluated on real overlapping speech data collected with simultaneous speakers in a meeting room. The results are compared with those obtained from close-talking headset microphones, lapel microphones, and speaker location based on audio-only and audio-visual information approaches. Hari Krishna Maganti, Daniel Gatica-Perez |
ICMI | 2 |
| 2006 | Detection and application of influence rankings in small group meetingsabstractWe address the problem of automatically detecting participant's influence levels in meetings. The impact and social psychological background are discussed. The more influential a participant is, the more he or she influences the outcome of a meeting. Experiments on 40 meetings show that application of statistical (both dynamic and static) models while using simply obtainable features results in a best prediction performance of 70.59% when using a static model, a balanced training set, and three discrete classes: high, normal and low. Application of the detected levels are shown in various ways i.e. in a virtual meeting environment as well as in a meeting browser system. Copyright 2006 ACM. Rutger Rienks, Dong Zhang 0001, Daniel Gatica-Perez, Wilfried M. Post |
ICMI | 3 |
| 2006 | Tracking the multi person wandering visual focus of attentionabstractEstimating the wandering visual focus of attention (WVFOA) for multiple people is an important problem with many applications in human behavior understanding. One such application, addressed in this paper, monitors the attention of passers-by to outdoor advertisements. This paper investigates the problem of tracking the wandering visual focus-of-attention (VFOA) of multiple people, an important problem with many applications in human behavior understanding. We address the specific problem of monitoring attention to outdoor advertisements. To solve the WVFOA problem, we propose a multi-person tracking approach based on a hybrid Dynamic Bayesian Network that simultaneously infers the number of people in the scene, their body and head locations, and their head pose, in a joint state-space formulation that is amenable for person interaction modeling. The model exploits both global measurements and individual observations for the VFOA. For inference in the resulting high-dimensional state-space, we propose a trans-dimensional Markov Chain Monte Carlo (MCMC) sampling scheme, which not only handles a varying number of people, but also efficiently searches the state-space by allowing person-part state updates. Our model was rigorously evaluated for tracking and its ability to recognize when people look at an outdoor advertisement using a realistic data set. Kevin Smith 0001, Sileye O. Ba, Daniel Gatica-Perez, Jean-Marc Odobez |
ICMI | 3 |
| 2006 | Human-centered computing: a multimedia perspectiveabstractHuman-Centered Computing (HCC) is a set of methodologies that apply to any field that uses computers, in any form, in applications in which humans directly interact with devices or systems that use computer technologies. In this paper, we give an overview of HCC from a Multimedia perspective. We describe what we consider to be the three main areas of Human-Centered Multimedia (HCM): media production, analysis, and interaction. In addition, we identify the core characteristics of HCM, describe example applications, and propose a research agenda for HCM. Copyright 2006 ACM. Alejandro Jaimes, Nicu Sebe, Daniel Gatica-Perez |
ACM Multimedia | 3 |
| 2006 | Embedding Motion in Model-Based Stochastic TrackingabstractParticle filtering is now established as one of the most popular methods for visual tracking. Within this framework, there are two important considerations. The first one refers to the generic assumption that the observations are temporally independent given the sequence of object states. The second consideration, often made in the literature, uses the transition prior as the proposal distribution. Thus, the current observations are not taken into account, requiring the noise process of this prior to be large enough to handle abrupt trajectory changes. As a result, many particles are either wasted in low likelihood regions of the state space, resulting in low sampling efficiency, or more importantly, propagated to distractor regions of the image, resulting in tracking failures. In this paper, we propose to handle both considerations using motion. We first argue that, in general, observations are conditionally correlated, and propose a new model to account for this correlation, allowing for the natural introduction of implicit and/or explicit motion measurements in the likelihood term. Second, explicit motion measurements are used to drive the sampling process towards the most likely regions of the state space. Overall, the proposed model handles abrupt motion changes and filters out visual distractors, when tracking objects with generic models based on shape or color distribution. Results were obtained on head tracking experiments using several sequences with moving camera involving large dynamics. When compared against the Condensation Algorithm, they have demonstrated the superior tracking performance of our approach. Jean-Marc Odobez, Daniel Gatica-Perez, Sileye O. Ba |
IEEE Trans. Image Process. | 2 |
| 2006 | Modeling individual and group actions in meetings with layered HMMsabstractWe address the problem of recognizing sequences of human interaction patterns in meetings, with the goal of structuring them in semantic terms. The investigated patterns are inherently group-based (defined by the individual activities of meeting participants, and their interplay), and multimodal (as captured by cameras and microphones). By defining a proper set of individual actions, group actions can be modeled as a two-layer process, one that models basic individual activities from low-level audio-visual (AV) features,and another one that models the interactions. We propose a two-layer hidden Markov model (HMM) framework that implements such concept in a principled manner, and that has advantages over previous works. First, by decomposing the problem hierarchically, learning is performed on low-dimensional observation spaces, which results in simpler models. Second, our framework is easier to interpret, as both individual and group actions have a clear meaning, and thus easier to improve. Third, different HMMs can be used in each layer, to better reflect the nature of each subproblem. Our framework is general and extensible, and we illustrate it with a set of eight group actions, using a public 5-hour meeting corpus. Experiments and comparison with a single-layer HMM baseline system show its validity. Dong Zhang 0001, Daniel Gatica-Perez, Samy Bengio, Iain McCowan |
IEEE Trans. Multim. | 2 |
| 2005 | Using Particles to Track Varying Numbers of Interacting PeopleabstractIn this paper, we present a Bayesian framework for the fully automatic tracking of a variable number of interacting targets using a fixed camera. This framework uses a joint multi-object state-space formulation and a trans-dimensional Markov Chain Monte Carlo (MCMC) particle filter to recursively estimates the multi-object configuration and efficiently search the state-space. We also define a global observation model comprised of color and binary measurements capable of discriminating between different numbers of objects in the scene. We present results which show that our method is capable of tracking varying numbers of people through several challenging real-world tracking situations such as full/partial occlusion and entering/leaving the scene. Kevin Smith 0001, Daniel Gatica-Perez, Jean-Marc Odobez |
CVPR (1) | 2 |
| 2005 | Semi-Supervised Adapted HMMs for Unusual Event DetectionabstractWe address the problem of temporal unusual event detection. Unusual events are characterized by a number of features (rarity, unexpectedness, and relevance) that limit the application of traditional supervised model-based approaches. We propose a semi-supervised adapted hidden Markov model (HMM) framework, in which usual event models are first learned from a large amount of (commonly available) training data, while unusual event models are learned by Bayesian adaptation in an unsupervised manner. The proposed framework has an iterative structure, which adapts a new unusual event model at each iteration. We show that such a framework can address problems due to the scarcity of training data and the difficulty in pre-defining unusual events. Experiments on audio, visual, and audiovisual data streams illustrate its effectiveness, compared with both supervised and unsupervised baseline methods. Dong Zhang 0001, Daniel Gatica-Perez, Samy Bengio, Iain McCowan |
CVPR (1) | 2 |
| 2005 | Detecting Group Interest-Level in MeetingsabstractFinding relevant segments in meeting recordings is important for summarization, browsing, and retrieval purposes. In this paper, we define relevance as the interest-level that meeting participants manifest as a group during the course of their interaction (as perceived by an external observer), and investigate the automatic detection of segments of high-interest from audio-visual cues. This is motivated by the assumption that there is a relationship between segments of interest to participants, and those of interest to the end user, e.g. of a meeting browser. We first address the problem of human annotation of group interest-level. On a 50-meeting corpus, recorded in a room equipped with multiple cameras and microphones, we found that the annotations generated by multiple people exhibit a good degree of consistency, providing a stable ground-truth for automatic methods. For the automatic detection of high-interest segments, we investigate a methodology based on hidden Markov models (HMM) and a number of audio and visual features. Single- and multi-stream approaches were studied. Using precision and recall as performance measures, the results suggest that the automatic detection of group interest-level is promising, and that while audio in general constitutes the predominant modality in meetings, the use of a multi-modal approach is beneficial. Daniel Gatica-Perez, Iain McCowan, Dong Zhang 0001, Samy Bengio |
ICASSP (1) | 1 |
| 2005 | Modeling Scenes with Local Descriptors and Latent AspectsabstractWe present a new approach to model visual scenes in image collections, based on local invariant features and probabilistic latent space models. Our formulation provides answers to three open questions:(l) whether the invariant local features are suitable for scene (rather than object) classification; (2) whether unsupennsed latent space models can be used for feature extraction in the classification task; and (3) whether the latent space formulation can discover visual co-occurrence patterns, motivating novel approaches for image organization and segmentation. Using a 9500-image dataset, our approach is validated on each of these issues. First, we show with extensive experiments on binary and multi-class scene classification tasks, that a bag-of-visterm representation, derived from local invariant descriptors, consistently outperforms state-of-the-art approaches. Second, we show that probabilistic latent semantic analysis (PLSA) generates a compact scene representation, discriminative for accurate classification, and significantly more robust when less training data are available. Third, we have exploited the ability of PLSA to automatically extract visually meaningful aspects, to propose new algorithms for aspect-based image ranking and context-sensitive image segmentation. Pedro Quelhas, Florent Monay, Jean-Marc Odobez, Daniel Gatica-Perez, Tinne Tuytelaars, Luc Van Gool |
ICCV | 4 |
| 2005 | Speech Acquisition in Meetings with an Audio-Visual Sensor ArrayabstractClose-talk headset microphones have been traditionally used for speech acquisition in a number of applications, as they naturally provide a higher signal-to-noise ratio - needed for recognition tasks than single distant microphones. However, in multi-party conversational settings like meetings, microphone arrays represent an important alternative to close-talking microphones, as they allow for localisation and tracking of speakers and signal-independent enhancement, while providing a non-intrusive, hands-free operation mode. In this article, we investigate the use of an audio-visual sensor array, composed of a small table-top microphone array and a set of cameras, for speaker tracking and speech enhancement in meetings. Our methodology first fuses audio and video for person tracking, and then integrates the output of the tracker with a beamformer for speech enhancement. We compare and discuss the features of the resulting speech signal with respect to that obtained from single close-talking and table-top microphones. Iain McCowan, Maganto Hari Krishna, Daniel Gatica-Perez, Darren Moore, Sileye O. Ba |
ICME | 3 |
| 2005 | Semi-supervised meeting event recognition with adapted HMMsabstractThis paper investigates the use of unlabeled data to help labeled data for audio-visual event recognition in meetings. To deal with situations in which it is difficult to collect enough labeled data to capture event characteristics, but collecting a large amount of unlabeled data is easy, we present a semi-supervised framework using HMM adaptation techniques. Instead of directly training one model for each event, we first train a well-estimated general event model for all events using both labeled and unlabeled data, and then adapt the general model to each specific event model using its own labeled data. We illustrate the proposed approach with a set of eight audio-visual events defined in meetings. Experiments and comparison with the fully-supervised baseline method show the validity of the proposed semi-supervised approach. Dong Zhang 0001, Daniel Gatica-Perez, Samy Bengio |
ICME | 2 |
| 2005 | Multimodal multispeaker probabilistic tracking in meetingsabstractTracking speakers in multiparty conversations constitutes a fundamental task for automatic meeting analysis. In this paper, we present a probabilistic approach to jointly track the location and speaking activity of multiple speakers in a multisensor meeting room, equipped with a small microphone array and multiple uncalibrated cameras. Our framework is based on a mixed-state dynamic graphical model defined on a multiperson state-space, which includes the explicit definition of a proximity-based interaction model. The model integrates audio-visual (AV) data through a novel observation model. Audio observations are derived from a source localization algorithm. Visual observations are based on models of the shape and spatial structure of human heads. Approximate inference in our model, needed given its complexity, is performed with a Markov Chain Monte Carlo particle filter (MCMC-PF), which results in high sampling efficiency. We present results -based on an objective evaluation procedure-that show that our framework (1) is capable of locating and tracking the position and speaking activity of multiple meeting participants engaged in real conversations with good accuracy; (2) can deal with cases of visual clutter and partial occlusion; and (3) significantly outperforms a traditional sampling-based approach. Daniel Gatica-Perez, Guillaume Lathoud, Jean-Marc Odobez, Iain McCowan |
ICMI | 1 |
| 2005 | Learning Influence among Interacting Markov ChainsabstractWe present a model that learns the influence of interacting Markov chains within a team. The proposed model is a dynamic Bayesian network (DBN) with a two-level structure: individual-level and group-level. Individual level models actions of each player, and the group-level models actions of the team as a whole. Experiments on synthetic multi-player games and a multi-party meeting corpus show the effectiveness of the proposed model. Dong Zhang 0001, Daniel Gatica-Perez, Samy Bengio, Deb Roy |
NIPS | 2 |
| 2005 | Automatic Analysis of Multimodal Group Actions in MeetingsabstractThis paper investigates the recognition of group actions in meetings. A framework is employed in which group actions result from the interactions of the individual participants. The group actions are modeled using different HMM-based approaches, where the observations are provided by a set of audiovisual features monitoring the actions of individuals. Experiments demonstrate the importance of taking interactions into account in modeling the group actions. It is also shown that the visual modality contains useful information, even for predominantly audio-based events, motivating a multimodal approach to meeting analysis. Iain McCowan, Daniel Gatica-Perez, Samy Bengio, Guillaume Lathoud, Mark Barnard, Dong Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Order Matters: A Distributed Sampling Method for Multi-Object TrackingabstractMulti-Object tracking (MOT) is an important problem in a number of vision applications. For particle filter (PF) tracking, as the number of objects tracked increases, the search space for random sampling explodes in dimension. Partitioned sampling (PS) solves this problem by partitioning the search space, then searching each partition sequentially. However, sequential weighted resampling steps cause an impoverishment effect that increases with the number of objects. This effect depends on the specific order in which the partitions are explored, creating an erratic and undesirable performance. We propose a method to search the state space that fairly distributes these impoverishment effects between the objects by defining a set of mixture components and performing PS in each of these components using one of a small set of representative object orderings. Using synthetic and real data, we show that our method retains the overall performance and reduced computational cost of PS, while improving performance in scenes where the impoverishment effect is significant. 1 Kevin Smith 0001, Daniel Gatica-Perez |
BMVC | 2 |
| 2004 | PLSA-based image auto-annotation: constraining the latent spaceabstractWe address the problem of unsupervised image auto-annotation with probabilistic latent space models. Unlike most previous works, which build latent space representations assuming equal relevance for the text and visual modalities, we propose a new way of modeling multi-modal co-occurrences, constraining the definition of the latent space to ensure its consistency in semantic terms (words), while retaining the ability to jointly model visual information. The concept is implemented by a linked pair of Probabilistic Latent Semantic Analysis (PLSA) models. On a 16000-image collection, we show with extensive experiments that our approach significantly outperforms previous joint models. Florent Monay, Daniel Gatica-Perez |
ACM Multimedia | 2 |
| 2003 | An implicit motion likelihood for tracking with particle filtersabstractParticle filters are now established as the most popular method for visual tracking. Within this framework, it is generally assumed that the data are temporally independent given the sequence of object states. In this paper, we argue that in general the data are correlated, and that modeling such dependency should improve tracking robustness. To take data correlation into account, we propose a new model which can be interpreted as introducing a likelihood on implicit motion measurements. The proposed model allows to filter out visual distractors when tracking objects with generic models based on shape or color distribution representations, as shown by the reported experiments. Jean-Marc Odobez, Sileye O. Ba, Daniel Gatica-Perez |
BMVC | 3 |
| 2003 | Modeling human interaction in meetingsabstractThe paper investigates the recognition of group actions in meetings by modeling the joint behaviour of participants. Many meeting actions, such as presentations, discussions and consensus, are characterised by similar or complementary behaviour across participants. Recognising these meaningful actions is an important step towards the goal of providing effective browsing and summarisation of processed meetings. A corpus of meetings was collected in a room equipped with a number of microphones and cameras. The corpus was labeled in terms of a predefined set of meeting actions characterised by global behaviour. In experiments, audio and visual features for each participant are extracted from the raw data and the interaction of participants is modeled using HMM-based approaches. Initial results on the corpus demonstrate the ability of the system to recognise the set of meeting actions. Iain McCowan, Samy Bengio, Daniel Gatica-Perez, Guillaume Lathoud, Florent Monay, Darren Moore, Pierre Wellner, Hervé Bourlard |
ICASSP (4) | 3 |
| 2003 | On automatic annotation of meeting databasesabstractIn this paper, we discuss meetings as an application domain for multimedia content analysis. Meeting databases are a rich data source suitable for a variety of audio, visual and multi-modal tasks, including speech recognition, people and action recognition, and information retrieval. We specifically focus on the task of semantic annotation of audio-visual (AV) events, where annotation consists of assigning labels (event names) to the data. In order to develop an automatic annotation system in a principled manner, it is essential to have a well-defined task, a standard corpus and an objective performance measure. In this work we address each of these issues to automatically annotate events based on participant interactions. Mark Barnard, Samy Bengio, Hervé Bourlard, Daniel Gatica-Perez, Iain McCowan |
ICIP (3) | 4 |
| 2003 | Audio-visual speaker tracking with importance particle filtersabstractWe present a probabilistic method for audio-visual (AV) speaker tracking, using an uncalibrated wide-angle camera and a micro- phone array. The algorithm fuses 2-D object shape and audio information via importance particle filters (I-PFs), allowing for the asymmetrical integration of AV information in a way that efficiently exploits the complementary features of each modality. Audio localization information is used to generate an importance sampling (IS) function, which guides the random search process of a particle filter towards regions of the configuration space likely to contain the true configuration (a speaker). The measurement process integrates contour-based and audio observations, which results in reliable head tracking in realistic scenarios. We show that imperfect single modalities can be combined into an algorithm that automatically initializes and tracks a speaker, switches between multiple speakers, tolerates visual clutter, and recovers from total AV object occlusion, in the context of a multimodal meeting room. Daniel Gatica-Perez, Guillaume Lathoud, Iain McCowan, Jean-Marc Odobez, Darren Moore |
ICIP (3) | 1 |
| 2003 | A Hierarchical Keyframe User Interface for Browsing Video over the Internet
Maël Guillemot, Pierre Wellner, Daniel Gatica-Perez, Jean-Marc Odobez |
INTERACT | 3 |
| 2003 | On image auto-annotation with latent space modelsabstractImage auto-annotation, i.e., the association of words to whole images, has attracted considerable attention. In particular, unsupervised, probabilistic latent variable models of text and image features have shown encouraging results, but their performance with respect to other approaches remains unknown. In this paper, we apply and compare two simple latent space models commonly used in text analysis, namely Latent Semantic Analysis (LSA) and Probabilistic LSA (PLSA). Annotation strategies for each model are discussed. Remarkably, we found that, on a 8000-image dataset, a classic LSA model defined on keywords and a very basic image representation performed as well as much more complex, state-of-the-art methods. Furthermore, non-probabilistic methods (LSA and direct image matching) outperformed PLSA on the same dataset. Florent Monay, Daniel Gatica-Perez |
ACM Multimedia | 2 |
| 2003 | Finding structure in home videos by probabilistic hierarchical clusteringabstractAccessing, organizing, and manipulating home videos present technical challenges due to their unrestricted content and lack of storyline. We present a methodology to discover cluster structure in home videos, which uses video shots as the unit of organization, and is based on two concepts: (1) the development of statistical models of visual similarity, duration, and temporal adjacency of consumer video segments and (2) the reformulation of hierarchical clustering as a sequential binary Bayesian classification process. A Bayesian formulation allows for the incorporation of prior knowledge of the structure of home video and offers the advantages of a principled methodology. Gaussian mixture models are used to represent the class-conditional distributions of intra- and inter-segment visual and temporal features. The models are then used in the probabilistic clustering algorithm, where the merging order is a variation of highest confidence first, and the merging criterion is maximum a posteriori. The algorithm does not need any ad-hoc parameter determination. We present extensive results on a 10-h home-video database with ground truth which thoroughly validate the performance of our methodology with respect to cluster detection, individual shot-cluster labeling, and the effect of prior selection. Daniel Gatica-Perez, Alexander C. Loui, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Probabilistic home video structuring: feature selection and performance evaluationabstractWe previously proposed a method to find the cluster structure in home videos based on statistical models of visual and temporal features of video segments and sequential binary Bayesian classification. In this paper, we present analysis and improved results on two key issues: feature selection and performance evaluation, using a ten-hour database (30 video clips, 1,075,000 frames). From multiple features and similarity measures, visual features are selected in order to minimize the empirical probability of misclassification. Temporal features are chosen to reflect the patterns existing in both shot and cluster duration and adjacency. Finally, we describe a detailed performance evaluation procedure that includes cluster detection, individual shot-cluster labeling, and prior selection. Daniel Gatica-Perez, Alexander C. Loui, Ming-Ting Sun |
ICIP (1) | 1 |
| 2002 | Linking objects in videos by importance samplingabstractWe present an approach to create hyper-links between video segments that contain objects of interest, based on video structuring, object definition, and stochastic object localization in the video structure. Localization is formulated in the metric mixture model framework, which allows for the joint probabilistic modeling of a (user-defined) set of color appearance exemplars and their geometric transformations. Candidate object configurations are drawn from a prior distribution using importance sampling - which guides the search towards regions of the configuration space likely to contain the correct object configuration, thus avoiding exhaustive processing - and evaluated using Bayes' rule. Results of linking real objects (with changes of size and pose) in several home videos illustrate the performance of the method. Daniel Gatica-Perez, Ming-Ting Sun |
ICME (2) | 1 |
| 2001 | Consumer Video Structuring by Probabilistic Merging of Video SegmentsabstractAccessing, organizing, and manipulating home videos constitutes a technical challenge due to their unrestricted content and the lack of storyline. In this paper, we present a methodology for structuring consumer video, based on the development of statistical models of similarity and adjacency between video segments in a probabilistic formulation. Learned Gaussian mixture models of inter-segment visual similarity, temporal adjacency, and segment duration are used to represent the classconditional densities of observed features. Such models are then used in a sequential merging algorithm consisting of a binary Bayes classifier, where the merging order is determined by a variation of Highest Confidence First (HCF), and the merging criterion is Maximum a Posteriori (MAP). The merging algorithm can be efficiently implemented and does not need any empirical parameter determination. Finally, the representation of the merging sequence by a tree provides for hierarchical, nonlinear access to the video content. Results on an eight-hour home video database illustrate the validity of our approach. Daniel Gatica-Perez, Ming-Ting Sun, Alexander C. Loui |
ICME | 1 |
| 2001 | Semantic video object extraction using four-band watershed and partition lattice operatorsabstractWe conceive the problem of multiple semantic video object (SVO) extraction as an issue of designing extensive operators on a complete lattice of partitions. As a result, we propose a framework based on spatial partition generation and application of optimal operators on the generated partitions. Based on a statistical analysis of the watershed algorithm, we develop a multivalued morphological spatial segmentation method that incorporates an edge-driven marker extraction algorithm and a growing method which integrates both color and edge information. Having embedded the problem in the partition lattice framework, we propose a spatio-temporal regional maximum likelihood operator for extraction purposes. Some theoretical properties of the operator are established. Experimental results on several MPEG-4 test video sequences show that our scheme improves the precision of the extracted SVO boundaries compared to traditional watershed algorithms and provides accurate tracking of multiple SVOs in both static and moving camera scenarios. Furthermore, this scheme can be extended to deal with more general interactive video authoring systems. Daniel Gatica-Perez, Chuang Gu, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2001 | Multiview extensive partition operators for semantic video object extractionabstractOcclusion/disocclusion is one of the fundamental problems for semantic video object (SVO) extraction, where pixel-wise accuracy is required. This issue is critical because the degradation in tracking due to object occlusion/disocclusion significantly increases the amount of user interaction required in off-line video editing applications. We present an approach based on the application of an extensive operator on a lattice of partitions, which exploits information from various views of the scene, based on a probabilistic formulation. Our multiview operator builds on the regional application of the maximum a posteriori principle, by integrating a single-view region classification stage with a multiview stage that improves classification for those disoccluded regions labeled as uncertain. Results on several real sequences show that our approach improves the SVO tracking compared to the single-view case and that, as a result, increases the quality of the extracted SVOs and reduces the total amount of user interaction. Daniel Gatica-Perez, Ming-Ting Sun, Chuang Gu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2001 | Extensive partition operators, gray-level connected operators, and region merging/classification segmentation algorithms: theoretical linksabstractThe relation between morphological gray-level connected operators and segmentation algorithms based on region merging/classification strategies has been pointed out several times in the literature. However, to the best of our knowledge, the formal relation between them has not been established. This paper presents the link between the two domains based on the observation that both connected operators and segmentation algorithms share a key mechanism: they simultaneously operate on images and on partitions, and therefore they can be described as operations on a joint image-partition model. As a result, we analyze both segmentation algorithms and connected operators by defining operators on complete product lattices, that explicitly model gray-level and partition attributes. In the first place, starting with a complete lattice of partitions, we initially define the concept of the segmentation model as a mapping in a product lattice, whose elements are three-tuples consisting of a partition, an image that models the partition attributes, and an image that represents the gray-level model associated to the segmentation. Then, assuming a conditional ordering relation, we show that any region merging/classification segmentation algorithm can be defined as an extensive operator in such a complete product lattice, in the second place, we proposed a very similar lattice-based extended representation of gray-level functions in the context of connected operators, that highlights the mathematical analogy with segmentation algorithms, but in which the ordering relation is different. We use this framework to show that every region merging/classification segmentation algorithm indeed corresponds to a connected operator. While this result provides an explanation to previous work in the area, it also opens possibilities for further analysis in the two domains. From this perspective, we additionally study some theoretical properties of a general region merging segmentation algorithm. Daniel Gatica-Perez, Chuang Gu, Ming-Ting Sun, Salvador Ruiz-Correa |
IEEE Trans. Image Process. | 1 |
| 2000 | Generating Video Objects by Multiple-View Extensive Partition Lattice OperatorsabstractOcclusion/disocclusion is one of the fundamental problems for semantic video object (SVO) extraction because the tracking degradation from occlusion/disocclusion creates difficulty for the pixel-wise accuracy requirement and adds to the amount of user interaction required in off-line video editing applications. We present an approach based on the application of a new extensive operator in a lattice of partitions, that exploits information from various views of the scene to address the disocclusion problem by using a regional Bayesian formulation. Our multiview operator builds on the application of the maximum a posteriori (MAP) principle, by integrating a single view region classification stage and a multiview stage that solves for disoccluded regions labeled as uncertain. Results on several real sequences show that this approach improves the quality of the extracted SVOs and reduces the total amount of user interaction. Daniel Gatica-Perez, Ming-Ting Sun, Chuang Gu |
ICIP | 1 |
| 2000 | Semiautomatic video object generation using multivalued watershed and partition lattice operatorsabstractWe conceive the problem of multiple semantic wide object (SVO) extraction as an issue of designing extensive operators on the lattice of partitions. As a result, we propose a framework based on spatial partition generation and application of optimal operators on the generated partitions. The first stage is obtained with a multivalued morphological spatial segmentation method that incorporates an edge-driven marker extraction algorithm and a growing method which integrates both color and edge information. For the second stage, we propose a new spatio-temporal regional maximum likelihood partition operator for extraction purposes. Subjective and objective evaluation of the experimental results obtained with our approach on several MPEG-4 test video sequences show that it accurately tracks multiple SVOs in several scenarios, while improving the SVO extraction precision compared to traditional watershed techniques. Daniel Gatica-Perez, Ming-Ting Sun, Chuang Gu |
ISCAS | 1 |
| 1999 | Semantic Video Object Extraction Based on Backward Tracking of Multivalued WatershedabstractWe present a novel algorithm for semantic video object (SVO) extraction, based on a new multivalued morphological spatial segmentation that integrates color and edge information, and object tracking by backward region classification. Our proposed spatial segmentation incorporates a new marker extraction method based on intensity edge information that improves the definition of the real borders of the scene objects, and a new distance criterion based on color and edge information to guide the watershed algorithm. Experimental results on several MPEG-4 test video sequences show that our algorithm improves the precision of the extracted SVO boundaries compared to the traditional water-shed technique, and that it is capable of tracking multiple SVOs in static and moving camera scenarios. Daniel Gatica-Perez, Ming-Ting Sun, Chuang Gu |
ICIP (2) | 1 |