Saturnino Luz

dblp:92/1880 · also Saturnino F. Luz-Filho · DBLP profile ↗
← Back
108ranked-venue papers
30as first author
23since 2021 · last 2026
0000-0001-8430-7875ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 59 · 16 first-author · 12 since 2021Artificial intelligence and machine learning · 49 · 11 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 12 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorTheory of computation · 2 · 2 first-author
YearPublicationVenuePosition
2026 Challenges in Reviewing Research Utilising Maps Visualizations: A Case Study of Cholera
abstract
This paper outlines challenges encountered in conducting a systematic scoping review of existing literature on the use of map visualizations and other related visual representations in research publications based on a case study of reported cholera outbreaks and epidemics. Unlike traditional reviews, this type of visualisation-based literature review requires the analysis of visuals included, both during article screening and data extraction. Based on a review of the use of visuals in articles on cholera, we describe the methods, search parameters and criteria adopted in the review process and its resulting outcomes. The aim is to highlight the needs of the reviewers during this process, and present the implications of the challenges encountered in relation to the development of more effective visual tools for supporting article screening and map visualization analysis in this type of novel and much-needed literature reviews. While focusing on the case study of map visualizations for cholera research, we would argue that much of the work presented here also applies to literature reviews of any research involving other forms of visualizations and visual material in general.
Saturnino Luz, Shane Sheehan, Masood Masoodian
AVI1
2026 Advances in Map-based Interfaces and Interactions - MAPII
abstract
Tools, applications and services relying on map-based interfaces are so ubiquitous in our modern digital world that often the important roles which maps play in supporting users’ interactions with them go unnoticed. Yet, without further advances in the design and development of map-based interfaces and interactions it is highly unlikely that future such systems will meet the increasing needs, expectations and requirements of their users. Therefore, it is important to provide regular multidisciplinary venues for the sharing of research, practical and professional expertise, learnings, practices, and experiences related to design, development and evaluation of advanced map-based interfaces and interactions across an ever-increasing range of fields. Over the past few years, the MAPII workshop has been providing one such much-needed forum that has brought together experts in maps and mapping from various disciplines, spanning sciences, humanities, art and design.
Masood Masoodian, Saturnino Luz
AVI2
2025 Supervised Machine Learning and Active Learning for Surrogate Outcome Detection in Clinical Protocols
abstract
A surrogate outcome is a substitute measure for a patient-final outcome; a direct measure of how a person feels, functions and/or survives. Intermediate outcomes can be defined as standardised functionality measures. Surrogate outcomes are increasingly used in clinical trials to accelerate drug approvals, yet the reporting of their use as primary endpoints in study protocols remains inconsistent. This is an issue because the primary endpoint is used to determine treatment efficacy and thus is crucial for interpreting results and making regulatory decisions. We considered various supervised learning techniques for detecting surrogate outcome usage using a three-class text classification approach. We evaluated this approach on two nervous system trial protocol datasets (EUCT-NS and NS-HRA). This study serves as a proof of concept for the potential of machine learning to automate the detection of primary surrogate endpoint use in clinical trial protocols, thus addressing gaps in reporting.
Onyeka Obuaya, Saturnino Luz, Shane Sheehan, Rod Taylor, Christopher Weir
CBMS2
2025 Investigating the Relationship between Signs of Depression and Alzheimer's Dementia in Speech
abstract
Depression and Alzheimer's dementia frequently co-occur, with depression often occurring at early stages of Alzheimer's disease (AD). We investigate whether knowledge of a person's depression score gained through analysis of acoustic features of the person's speech can help detect AD. We analysed data from 239 participants, where$n=162$individuals had a clinical diagnosis of AD, to explore correlations between depression scores, cognitive test scores and AD status. First, we created multilayer perceptron models for depression assessment using speech features extracted through our active data representation method and wav2vec features. We then used outputs from the depression models as predictors in our Alzheimer's detection models. Depression scores were strongly correlated with Alzheimer's diagnosis in this data set (Spearman's correlation$r = -0.610$and$r = -0.536,\ p < 0.01$for correlations between depression scores and diagnosis and cognitive test scores, respectively). We found that depression scores could predict AD fairly accurately$(Acc=0.84)$, and that depression could be detected in speech with moderate accuracy$(Acc=0.79)$. Using these predictions and acoustic features, our dementia model reached$Acc=0.70$. We also investigated classification of AD from speech in the presence and absence of depression, obtaining the same level of accuracy for this imbalanced prediction task.
Fasih Haider, Stina Saunders, Craig Ritchie, Saturnino Luz
CCNC4
2025 Automatic recognition of rodent call types using deep supervectors
abstract
Rats are gregarious rodents who naturally live in diverse social groups and communicate in part through ultrasonic vocalisations (USVs). USVs encode significant information about affective state and play an important role in social behaviour. Monitoring USVs is a non-invasive method of adding richness to data in a variety of experimental paradigms. However, manual analysis of USVs requires a significant amount of human effort. We propose a new method for automatic classification of USVs which could help automate analysis and thus reduce human input. The proposed method introduces a novel approach to USV representation called deep supervectors (DSV), which combines diverse deep embeddings feature sets extracted through our active data representation (ADR) method. The DSV method is evaluated on a multiclass recognition task involving 14 different types of rodent calls. The performance of DSV is compared to that obtained by state-of-the-art deep embeddings (alexNet, googleNet, squeezeNet and resNet). The proposed method achieves an Unweighted Average Recall (UAR) of 32.80% and outperforms both customs (8.07%) and deep embeddings (29.94%). Deep supervectors outperform pre-trained DNN in 6 out of 8 cases and fusion of the top six DSV with the top two deep embeddings improves the UAR to 37.22% in this challenging classification task, showing an improvement of 29% over the majority guess of 8%. By combining related classes, the UAR will reach 51.00%. The context of this application is the automation of the process for life-sciences laboratory technicians and scientists.
Fasih Haider, Raven Hickson, Peter Kind, Saturnino Luz
ICASSP4
2025 Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS Challenge
abstract
Dementia is associated with various cognitive impairments and typically manifests only after significant progression, making intervention at this stage often ineffective. To address this issue, the Prediction and Recognition of Cognitive Decline through Spontaneous Speech (PROCESS) Signal Processing Grand Challenge invites participants to focus on early-stage dementia detection. We provide a new spontaneous speech corpus for this challenge. This corpus includes answers from three prompts designed by neurologists to better capture the cognition of speakers. Our baseline models achieved an F1-score of 55.0% on the classification task and an RMSE of 2.98 on the regression task.
Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar, Sophie Young, Hend Elghazaly, Fritz Peters, Caitlin H. Illingworth, Dorota Braun, Ronan O'Malley, Simon Bell, Daniel Blackburn, Fasih Haider, Saturnino Luz, Heidi Christensen
ICASSP14
2025 Affective and Physiological Responses to Immersive Intangible Cultural Heritage Experiences in Extended Reality
abstract
As part of the Intangible Cultural Heritage, Bridging the Past, Present and Future (INT-ACT) project, we investigate how Extended Reality (XR) technologies can meaningfully engage users with cultural content by monitoring their physiological and affective responses. While immersive XR systems offer new ways of exploring heritage, their impact on users’ internal states remains underexplored. In this study, we present a multimodal experimental setup using EmotiBit, a wearable biosensing platform, to monitor real-time physiological signals during cultural XR interaction. Participants were evaluated across three activities representing varying cognitive and sensory loads: immersive interaction with the INT-ACT XR Demonstrator, composing work emails (a low-stimulation control task), and passive movie watching. Our aim is to quantify how cultural XR experiences influence biosignals such as electrodermal activity and heart rate. The findings reveal distinct physiological patterns across conditions, suggesting that biosignal monitoring can inform the design of adaptive XR environments that are responsive to user states. This work contributes to INT-ACT’s broader objective of creating intelligent, inclusive, and emotionally resonant cultural heritage experiences.
Fasih Haider, Sofia de la Fuente, Alicia Núñez García, Saturnino Luz
ICMI4
2025 Effective Map-Based Interfaces and Interactions
Masood Masoodian, Saturnino Luz
INTERACT (4)2
2025 An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
abstract
Speech-based depression detection tools could aid early screening. Here, we propose an interpretable speech foundation model approach to enhance the clinical applicability of such tools. We introduce a speech-level Audio Spectrogram Transformer (AST) to detect depression using long-duration speech instead of short segments, along with a novel interpretation method that reveals prediction-relevant acoustic features for clinician interpretation. Our experiments show the proposed model outperforms a segment-level AST, highlighting the impact of segment-level labelling noise and the advantage of leveraging longer speech duration for more reliable depression detection. Through interpretation, we observe our model identifies reduced loudness and F0 as relevant depression signals, aligning with documented clinical findings. This interpretability supports a responsible AI approach for speech-based depression detection, rendering such tools more clinically applicable.
Qingkun Deng, Saturnino Luz, Sofia de la Fuente
INTERSPEECH2
2025 Zero-Shot Speech-Based Depression and Anxiety Assessment with LLMs
abstract
The use of Large Language Models (LLMs) for psychological state assessment from speech has gained significant interest, particularly in analysing and predicting mental health. In this paper, we explore the potential of eight instruct-tuned LLMs (Llama-3.1-8B, Ministral, Gemma-2-9B, Phi-4, Mistral, DeepSeek-Qwen, QwQ-Preview and Llama-3.3-70B) in a zero-shot setting to predict Hospital Anxiety and Depression Scale (HADS) depression and anxiety scores from one-to-two minute spontaneous speech recordings from the PsyVoiD database. We evaluate how transcript quality affects LLM responses by comparing performance using ground-truth transcriptions versus transcripts generated by Whisper models of different sizes. Spearman correlation coefficients and statistical analysis demonstrate significant and notable potential of the LLMs to predict psychological states in a zero-shot setting.
Erfan Loweimi, Sofia de la Fuente, Saturnino Luz
INTERSPEECH3
2025 Visualisations to guide enriched proteome analyses
abstract
Quantitative mass spectrometry based proteomics data is high dimensional and when enriched with biological information for every measured protein, it widens the scope of analysis for end users. In silico analyses where data enrichment aids analysis goals can be further enhanced with tailored visualisations to guide differential pattern activity in proteomic data. In this study, we employed the core analysis tool of a single software platform, Qiagen’s Ingenuity Pathway Analysis (IPA), which is widely used in omic analyses, to explore comparative profiling of multi-source proteomic data. The central aim was to construct added visualisations from the wealth of exportable features available in IPA to provide further visual guides for users to utilise the built in tool metrics and domain based understanding in relation to their data.
Somya Iqbal, Shane Sheehan, Saturnino Luz
IV3
2024 Managing Personal Health and Well-being by Integrating Data Streams using Interactive Visualizations
abstract
There has been in a rapid growth in the use of consumer health monitoring devices integrated with smartphones that help people to record and analyse data relating to different factors contributing to their health and well-being. However, these factors are often presented separately, making it hard for users to visualise potential relationships among different health data streams, and missing an opportunity to use other information sources available on their devices to contextualise these health indicators. Here we present a practical concept for integrating health and other contextual data with different analytic and visualization tools as the basis for an interactive health and well-being management dashboard.
Saturnino Luz, Masood Masoodian
AVI1
2024 Map-based Interfaces and Interactions (MAPII 2024)
abstract
Maps have been used for centuries as tools for exploring the real and the imagined, the physical and the metaphysical worlds. Today, in the world of technology, maps also play an important role as underlying representation tools, forming the basis of a wide range of digital devices, applications, and services. Despite this, there are hardly any venues for sharing of research and design expertise, learnings, practices, and experiences of the use of maps and map-like visualizations in the context of visual interfaces and interactions. This workshop aims to fill this gap by providing a much-needed interdisciplinary forum focusing on map-based interfaces and interactions.
Masood Masoodian, Saturnino Luz
AVI2
2024 Connected Speech-Based Cognitive Assessment in Chinese and English
Saturnino Luz, Sofia de la Fuente, Fasih Haider, Davida Fromm, Brian MacWhinney, Alyssa Lanzi, Ya-Ning Chang, Chia-Ju Chou, Yi-Chien Liu
INTERSPEECH1
2023 Multilingual Alzheimer's Dementia Recognition through Spontaneous Speech: A Signal Processing Grand Challenge
abstract
This Signal Processing Grand Challenge (SPGC) targets a difficult automatic prediction problem of societal and medical relevance, namely, the detection of Alzheimer’s Dementia (AD). Participants were invited to employ signal processing and machine learning methods to create predictive models based on spontaneous speech data. The Challenge has been designed to assess the extent to which predictive models built based on speech in one language (English) generalise to another language (Greek). To the best of our knowledge no work has investigated acoustic features of the speech signal in multilingual AD detection. Our baseline system used conventional machine learning algorithms with Active Data Representation of acoustic features, achieving accuracy of 73.91% on AD detection, and 4.95 root mean squared error on cognitive score prediction.
Saturnino Luz, Fasih Haider, Davida Fromm, Ioulietta Lazarou, Ioannis Kompatsiaris, Brian MacWhinney
ICASSP1
2023 Designing for Map-Based Interfaces and Interactions
Masood Masoodian, Saturnino Luz
INTERACT (4)2
2022 Map-based Interfaces and Interactions
abstract
Maps and map-like visualizations are increasingly being used as the basis for many types of interactive tools, services and applications. This workshop brings together researchers and practitioners whose work on such map-based interfaces and interactions is currently dispersed across different publication venues and disciplines. As such, this workshop aims to act as a common space for sharing map-related research expertise and knowledge coming from the fields of visualization, user interface design, interaction design, visual design and cartography.
Masood Masoodian, Saturnino Luz
AVI2
2022 Corpus Summarization and Exploration using Multi-Mosaics
abstract
In fields such as translation studies and computational linguistics, various tools are used to analyze the content of text corpora, and extract keywords and other entities for analysis. Concordancing – arranging passages of text corpus in alphabetical order of user-defined keywords – is one of most widely used forms of text analysis. This paper describes Multi-Mosaics, a tool for text analysis using multiple implicitly linked Concordance Mosaic visualisations. Multi-Mosaics supports examining linguistic relationships within the context windows surrounding multiple extracted keywords.
Shane Sheehan, Saturnino Luz, Masood Masoodian
AVI2
2022 An Automated Mood Diary for Older User's using Ambient Assisted Living Recorded Speech
Fasih Haider, Saturnino Luz
INTERSPEECH2
2022 Task-based Quantitative Evaluation of the Concordance Mosaic Visualization
abstract
Researchers working in areas such as lexicography, translation studies, and computational linguistics, use a combination of automated and semi-automated tools to analyze the content of text corpora. Concordancing - or the arranging of passages of a textual corpus in alphabetical order according to user-defined keywords - is one of the oldest and still most widely used forms of text analysis. Concordance Mosaic is an interactive concordance visualization which emphasises quantitative information such as word frequency. While Concordance Mosaic is in active use by humanities scholars, no quantitative evaluation of the technique exists. In this paper, the Concordance Mosaic is quantitatively evaluated in comparison to a typical concordance browser. The comparison is evaluated using speed and accuracy on identified corpus analysis actions.
Shane Sheehan, Masood Masoodian, Saturnino Luz
IV3
2021 Affect Recognition Through Scalogram and Multi-Resolution Cochleagram Features
Fasih Haider, Saturnino Luz
Interspeech2
2021 Detecting Cognitive Decline Using Speech Only: The ADReSSo Challenge
abstract
Building on the success of the ADReSS Challenge at Interspeech 2020, which attracted the participation of 34 teams from across the world, the ADReSSo Challenge targets three difficult automatic prediction problems of societal and medical relevance, namely: detection of Alzheimer's Dementia, inference of cognitive testing scores, and prediction of cognitive decline. This paper presents these prediction tasks in detail, describes the datasets used, and reports the results of the baseline classification and regression models we developed for each task. A combination of acoustic and linguistic features extracted directly from audio recordings, without human intervention, yielded a baseline accuracy of 78.87% for the AD classification task, an MMSE prediction root mean squared (RMSE) error of 5.28, and 68.75% accuracy for the cognitive decline prediction task.
Saturnino Luz, Fasih Haider, Sofia de la Fuente, Davida Fromm, Brian MacWhinney
Interspeech1
2021 Emotion recognition in low-resource settings: An evaluation of automatic feature selection methods
Fasih Haider, Senja Pollak, Pierre Albert, Saturnino Luz
Comput. Speech Lang.4
2020 TeMoCo-Doc: A visualization for supporting temporal and contextual analysis of dialogues and associated documents
abstract
A common task in a number of application areas is to create textual documents based on recorded audio data. Visualizations designed to support such tasks require linking temporal audio data with contextual data contained in the resulting documents. In this paper, we present a tool for the visualization of temporal and contextual links between recorded dialogues and their summary documents.
Shane Sheehan, Saturnino Luz, Pierre Albert, Masood Masoodian
AVI2
2020 Automatic Recognition of Low-Back Chronic Pain Level and Protective Movement Behaviour using Physical and Muscle Activity Information
abstract
Automatic recognition of low-back chronic pain and movement behaviour in humans could be a useful technology in health monitoring and providing effective rehabilitation advice. Physical and muscle activity information can be used in automating this process in combination with machine learning and feature engineering methods. This paper presents a method for automatic recognition of chronic pain and movement behaviour using our recently proposed `Active Data Representation' (ADR) method, and applies it to two tasks of the EmoPain 2020 Challenge using physical and muscle activity features. The ADR method is used for the transformation of the physical and muscle activity features for the classification tasks. Our results show that ADR outperforms the LSTM challenge baseline model in terms of Matthews correlation coefficient (0.43) and F score (61.21) for the recognition of chronic pain and movement behaviour respectively in hold-out validation settings. Although a decrease in performance is observed on the test dataset, ADR still outperforms the challenge baseline for the recognition of chronic pain and movement behaviour tasks.
Fasih Haider, Pierre Albert, Saturnino Luz
FG3
2020 Alzheimer's Dementia Recognition Through Spontaneous Speech: The ADReSS Challenge
abstract
The ADReSS Challenge at INTERSPEECH 2020 defines a shared task through which different approaches to the automated recognition of Alzheimer's dementia based on spontaneous speech can be compared. ADReSS provides researchers with a benchmark speech dataset which has been acoustically pre-processed and balanced in terms of age and gender, defining two cognitive assessment tasks, namely: the Alzheimer's speech classification task and the neuropsychological score regression task. In the Alzheimer's speech classification task, ADReSS challenge participants create models for classifying speech as dementia or healthy control speech. In the the neuropsychological score regression task, participants create models to predict mini-mental state examination scores. This paper describes the ADReSS Challenge in detail and presents a baseline for both tasks, including feature extraction procedures and results for classification and regression models. ADReSS aims to provide the speech and language Alzheimer's research community with a platform for comprehensive methodological comparisons. This will hopefully contribute to addressing the lack of standardisation that currently affects the field and shed light on avenues for future research and clinical applicability.
Saturnino Luz, Fasih Haider, Sofia de la Fuente, Davida Fromm, Brian MacWhinney
INTERSPEECH1
2019 TeMoCo: A Visualization Tool for Temporal Analysis of Multi-party Dialogues in Clinical Settings
abstract
We present a tool for visualization of transcripts of multi-party dialogues, with application to the analysis of communication in medical teamwork. The visualization is based on a "temporal mosaic" metaphor, which provides a temporal overview of dialogues and supports the tasks of transcript browsing and information access, by segmenting the dialogue and laying out the keywords of the different segments on interactive visual "tiles". The tool has been tested on a corpus of transcribed dialogues among the members of a (simulated) critical care team. An analytical evaluation is presented which demonstrates the potential uses of the tool in an educational setting and highlights areas for improvements.
Shane Sheehan, Pierre Albert, Saturnino Luz, Masood Masoodian
CBMS3
2019 Attitude Recognition Using Multi-resolution Cochleagram Features
abstract
Attitudes play an important role in human communication. Models and algorithms for automatic recognition of attitudes therefore may have applications in areas where successful communication and interaction are crucial, such as health-care, education and digital entertainment. This paper focuses on the task of categorizing speaker attitudes using speech features. Data extracted from video recordings are employed in training and testing of predictive models consisting of different sets of speech features. A novel attitude recognition approach using Multi-Resolution Cochleagram (MRCG) features is proposed. The results show that MRCG feature set outperforms the feature sets most commonly used in computational paralinguistic tasks, including emobase, eGeMAPS and ComParE, in terms of attitude recognition accuracy for decision tree, 1-nearest neighbour and random forest classifiers. Analysis of the results suggests that MRCG features contribute information not captured by these existing feature sets. Indeed, while the ComParE feature set provides slightly better results than MRCG features for support vector machine classifiers, the fusion of the existing feature sets with the new MRCG features improves on those results. Overall, with the addition of MRCG, the attitude recognition method proposed in this study achieves accuracy scores approximately 11 points higher than reported in previous studies.
Fasih Haider, Saturnino Luz
ICASSP2
2019 A Searching and Automatic Video Tagging Tool for Events of Interest during Volleyball Training Sessions
abstract
Quick and easy access to performance data during matches and training sessions is important for both players and coaches. While there are many video tagging systems available, these systems require manual effort. This paper proposes a system architecture that automatically supplements video recording by detecting events of interests in volleyball matches and training sessions to provide tailored and interactive multi-modal feedback.
Fahim A. Salim, Fasih Haider, Sena Busra Yengec Tasdemir, Vahid Naghashi, Izem Tengiz, Kübra Cengiz, Dees B. W. Postma, Robby van Delden, Dennis Reidsma, Saturnino Luz, Bert-Jan van Beijnum
ICMI10
2019 A System for Real-Time Privacy Preserving Data Collection for Ambient Assisted Living
Fasih Haider, Saturnino Luz
INTERSPEECH2
2019 Analysing patterns of right brain-hemisphere activity prior to speech articulation for identification of system-directed speech
Fasih Haider, Hayakawa Akira, Carl Vogel, Nick Campbell 0001, Saturnino Luz
Speech Commun.5
2018 COMFRE: a visualization for comparing word frequencies in linguistic tasks
abstract
Comparing frequency distributions is a basic task in statistics and in disciplines that rely on statistical analysis, such as corpus linguistics. However, support for comparing word frequencies between different corpora in corpus linguistics tasks such as lexical analysis and corpus-based translation studies, is often limited to fairly basic techniques like tabular word lists. While other visualizations such as word clouds do exist, they are not widely used in linguistic analysis tasks due to their lack of precision, unsuitability to dealing with the high frequencies of common words, and lack of effective mechanisms for direct manipulation. In this paper, we propose a visualization for comparing word frequencies across two corpora using a combination of slope charts and histogram contours. An interactive implementation of this visualization is also presented. The design of visualization, and the development of the prototype, have been guided through the involvement of expert linguist users.
Shane Sheehan, Masood Masoodian, Saturnino Luz
AVI3
2018 On-Talk and Off-Talk Detection: A Discrete Wavelet Transform Analysis of Electroencephalogram
abstract
Spoken interaction with a machine results in a behaviour that is not very common in face-to-face human communication:Off-Talk, which is defined as speech utterances that are not directed to an immediate interlocutor, the machine, but to another person or even oneself. It is our contention that a system which is able to detect theOff-Talkutterances can interact with a human in a more efficient manner by acknowledging that the utterances are not directed to the system and hence, not replying toOff-Talkutterances. In this paper, we demonstrate the discrimination power of a wide range of Electroencephalogram (EEG) frequency bands using wavelet transform analysis and propose models forOn-TalkandOff-Talkdetection using audio and EEG signals, and their fusion. Our study shows that the EEG signal can identify the occurrence ofOff-Talkutterances with promising accuracy and its fusion with audio features adds a slight improvement in these results.
Fasih Haider, Hayakawa Akira, Saturnino Luz, Carl Vogel, Nick Campbell 0001
ICASSP3
2018 SAAMEAT: Active Feature Transformation and Selection Methods for the Recognition of User Eating Conditions
abstract
Automatic recognition of eating conditions of humans could be a useful technology in health monitoring. The audio-visual information can be used in automating this process, and feature engineering approaches can reduce the dimensionality of audio-visual information. The reduced dimensionality of data (particularly feature subset selection) can assist in designing a system for eating conditions recognition with lower power, cost, memory and computation resources than a system which is designed using full dimensions of data. This paper presents Active Feature Transformation (AFT) and Active Feature Selection (AFS) methods, and applies them to all three tasks of the ICMI 2018 EAT Challenge for recognition of user eating conditions using audio and visual features. The AFT method is used for the transformation of the Mel-frequency Cepstral Coefficient and ComParE features for the classification task, while the AFS method helps in selecting a feature subset. Transformation by Principal Component Analysis (PCA) is also used for comparison. We find feature subsets of audio features using the AFS method (422 for Food Type, 104 for Likability and 68 for Difficulty out of 988 features) which provide better results than the full feature set. Our results show that AFS outperforms PCA and AFT in terms of accuracy for the recognition of user eating conditions using audio features. The AFT of visual features (facial landmarks) provides less accurate results than the AFS and AFT sets of audio features. However, the weighted score fusion of all the feature set improves the results.
Fasih Haider, Senja Pollak, Eleni Zarogianni, Saturnino Luz
ICMI4
2018 Improving Response Time of Active Speaker Detection Using Visual Prosody Information Prior to Articulation
Fasih Haider, Saturnino Luz, Carl Vogel, Nick Campbell 0001
INTERSPEECH2
2018 An Active Feature Transformation Method for Attitude Recognition of Video Bloggers
Fasih Haider, Fahim A. Salim, Owen Conlan, Saturnino Luz
INTERSPEECH4
2018 Speech Rate Calculations with Short Utterances: A Study from a Speech-to-Speech, Machine Translation Mediated Map Task
Hayakawa Akira, Carl Vogel, Saturnino Luz, Nick Campbell 0001
LREC3
2018 The Metalogue Debate Trainee Corpus: Data Collection and Annotations
Volha Petukhova, Andrei Malchanau, Youssef Oualil, Dietrich Klakow, Saturnino Luz, Fasih Haider, Nick Campbell 0001, Dimitris Koryzis, Dimitris Spiliotopoulos, Pierre Albert, Nicklas Linz, Jan Alexandersson
LREC5
2017 Trust, Ethics and Access: Challenges in Studying the Work of Multi-disciplinary Medical Teams
abstract
This paper highlights the challenges for researchers when undertaking research on multidisciplinary medical teams (MDTs) in real-world healthcare settings, and suggests ways in which these challenges may be addressed.
Bridget Kane, Saturnino Luz
CBMS2
2017 Longitudinal Monitoring and Detection of Alzheimer's Type Dementia from Spontaneous Speech Data
abstract
A method for detection of Alzheimers type dementia though analysis of vocalisation features that can be easily extracted from spontaneous speech is presented. Unlike existing approaches, this method does not rely on transcriptions of the patients speech. Tests of the proposed method on a data set of spontaneous speech recordings of Alzheimers patients (n=214) and elderly controls (n=184) show that accuracy of 68% can be achieved with a Bayesian classifier operating on features extracted through simple algorithms for voice activity detection and speech rate tracking.
Saturnino Luz
CBMS1
2017 Ecoepidemiological Simulation as a Serious Game Engine Module
abstract
The integration of an agent-based simulation model as a component of a game engine for serious games targeting prevention and health promotion in the context of infectious diseases is described. It is argued that a combination of agent-based modelling and serious games can help provide a more realistic picture of disease spread than conventional ecoepidemiological models, by facilitating the integration of more detailed multidisciplinary expert knowledge. In addition, agent-based simulations provide engaging game mechanics, thereby fostering citizen engagement in the collection of up-to-date real-world data which can be used to improve the model.
Masood Masoodian, Saturnino Luz
CBMS2
2017 Speech Rate Comparison When Talking to a System and Talking to a Human: A Study from a Speech-to-Speech, Machine Translation Mediated Map Task
Hayakawa Akira, Carl Vogel, Saturnino Luz, Nick Campbell 0001
INTERSPEECH3
2017 Visual, Laughter, Applause and Spoken Expression Features for Predicting Engagement Within TED Talks
Fasih Haider, Fahim A. Salim, Saturnino Luz, Carl Vogel, Owen Conlan, Nick Campbell 0001
INTERSPEECH3
2017 Temporal Visualization of Energy Consumption Loads Using Time-Tone
abstract
Feedback plays an important role in assisting users to better understand their energy consumption behaviour. This is particularly true when users want to change their behaviour in order to reduce their energy consumption, and to manage their usage more effectively so as to avoid putting unnecessary load on energy providers. This paper presents the time-tone visualization, which aims to assist users by displaying variations in energy consumption by different categories of household devices over time, and their respective contributions to the total energy usage load. A user study conducted to compare time-tone against area-charts shows that although the two visualizations are comparable, time-tone is more effective for cases where there are large variations in energy usage loads.
Masood Masoodian, Ida Buchwald, Saturnino Luz, Elisabeth André
IV3
2016 Time-load: Visualization of Energy Consumption Loads Over Time
abstract
It has been suggested that providing feedback allows people to better understand their energy consumption behaviour and take the necessary actions to reduce energy consumption. Here, we present the time-load visualization, which shows percentages of energy usage by different categories of devices that contribute to the total energy consumption load over time. Time-load aims to assist users with visualization of variations in energy consumption by different categories over time, and their respective, and often very unbalanced, contributions to the total energy usage load.
Masood Masoodian, Saturnino Luz
AVI2
2016 Presentation quality assessment using acoustic information and hand movements
abstract
This study focuses on prosodic and gestural features that contribute to the positive judgement of public oral presentations. The general hypothesis is that certain prosodic characteristics, such as high pitch variation and perceived loudness, together with the production of natural hand gestures, influence the audience's perception of the speaker as a good presenter. Being able to identify features that can give an indication of a good presenter is useful for applications in the field of skills training, where automatic feedback could be provided to trainees at the end of their presentation about the extent to which they have been able to use their voices and gestures to keep the audience engaged. For this reason, we also propose a method, based on prosodic and visual features, able to categorise presentation quality with high accuracy.
Fasih Haider, Loredana Cerrato, Nick Campbell 0001, Saturnino Luz
ICASSP4
2016 Talking to a System and Talking to a Human: A Study from a Speech-to-Speech, Machine Translation Mediated Map Task
Hayakawa Akira, Saturnino Luz, Nick Campbell 0001
INTERSPEECH2
2016 The ILMT-s2s Corpus ― A Multimodal Interlingual Map Task Corpus
Hayakawa Akira, Saturnino Luz, Loredana Cerrato, Nick Campbell 0001
LREC2
2015 Clinical Training and Teamwork: Learning and Feedback
abstract
MDTMs are now a feature of routine hospital work and provide a valuable learning opportunity for education and practice development. The popularity of the forum as a patient management mechanism has had a negative counter effect on the educational function of the forum. Behavioural interventions and technical supports are identified based on long term ethnographic studies to restore the educational benefits of the forum. The potential for re-developing the MDTM into a rich educational resource that will assist in clinical education, professional development, provide an evidence base for guideline development by integrating clinical outcome feedback into the meeting record is proposed.
Bridget Kane, Saturnino Luz
CBMS2
2015 A Serious Game for Improving Community-Based Prevention of Neglected Diseases
abstract
Community-based healthcare strategies are becoming increasingly important in developing sustainable practices for prevention of neglected and emerging diseases in remote regions. In this paper, we discuss the use of "serious games" as one of the strategies for improving local populations' knowledge of the causes, preventive measures, and treatment options for neglected tropical diseases. We illustrate the potential of such a strategy for fostering engagement between local communities and healthcare workers by presenting a serious game architecture we have developed in collaboration with medical researchers and practitioners working in Amazonia. Although this first game focuses on Leishmaniases, it can be extended to easily include other similar diseases. This prototype game has been presented to a group of experts on neglected tropical diseases, and their opinions are reported here.
Masood Masoodian, Saturnino Luz, Manuel Cesario, Raquel Rangel Cesario, Bill Rogers 0001, Diones A. Borges
CBMS2
2015 Analyzing Multimodality of Video for User Engagement Assessment
abstract
These days, several hours of new video content is uploaded to the internet every second. It is simply impossible for anyone to see every piece of video which could be engaging or even useful to them. Therefore it is desirable to identify videos that might be regarded as engaging automatically, for a variety of applications such as recommendation and personalized video segmentation etc. This paper explores how multimodal characteristics of video, such as prosodic, visual and paralinguistic features, can help in assessing user engagement with videos. The approach proposed in this paper achieved good accuracy (maximum F score of 96.93 %) through a novel combination of features extracted directly from video recordings, demonstrating the potential of this method in identifying engaging content.
Fahim A. Salim, Fasih Haider, Owen Conlan, Saturnino Luz, Nick Campbell 0001
ICMI4
2015 Detection of cognitive states and their correlation to speech recognition performance in speech-to-speech machine translation systems
Hayakawa Akira, Fasih Haider, Loredana Cerrato, Nick Campbell 0001, Saturnino Luz
INTERSPEECH5
2015 Medical teamwork, collaboration and patient-centred care
abstract
How can information communication technology be better applied to make our health-care systems more effective, dependable and safer for patients, and deliver economic efficiency? Economists challen...
Bridget Kane, Saturnino Luz
Behav. Inf. Technol.2
2015 Disease surveillance and patient care in remote regions: an exploratory study of collaboration among health-care professionals in Amazonia
abstract
The development and deployment of information technology, particularly mobile tools, to support collaboration between different groups of health-care professionals has been viewed as a promising way to improve disease surveillance and patient care in remote regions. The effects of global climate change combined with rapid changes to land cover and use in Amazonia are believed to be contributing to the spread of vector-borne emerging and neglected diseases. This makes empowering and providing support for local health-care providers all the more important. We investigate the use of information technology in this context to support professionals whose activities range from diagnosing diseases and monitoring their spread to developing policies to deal with outbreaks. An analysis of stakeholders, their roles and requirements, is presented which encompasses results of fieldwork and of a process of design and prototyping complemented by questionnaires and targeted interviews. Findings are analysed with respect to the tasks of diagnosis, training of local health-care professionals, and gathering, sharing and visualisation of data for purposes of epidemiological research and disease surveillance. Methodological issues regarding the elicitation of cooperation and collaboration requirements are discussed and implications are drawn with respect to the use of technology in tackling emerging and neglected diseases.
Saturnino Luz, Masood Masoodian, Manuel Cesario
Behav. Inf. Technol.1
2015 Wizard of Oz Experimentation for Language Technology Applications: Challenges and Tools
abstract
Wizard of OZ (WOZ) is a well-established method for simulating the functionality and user experience of future systems. Using a human wizard to mimic certain operations of a potential system is particularly useful in situations where extensive engineering effort would otherwise be needed to explore the design possibilities offered by such operations. The WOZ method has been widely used in connection with speech and language technologies, but advances in sensor technology and pattern recognition as well as new application areas such as human-robot interaction have made it increasingly relevant to the design of a wider range of interactive systems. In such cases achieving acceptable performance at the user interface level often hinges on resource intensive improvements such as domain tuning, which are better done once the overall design is relatively stable. While WOZ is recognised as a valuable prototyping technique, surprisingly little effort has been put into exploring it from a methodological point of view. Starting from a survey of the literature, this paper presents a systematic investigation and analysis of the design space for WOZ for language technology applications, and proposes a generic architecture for tool support that supports the integration of components for speech recognition and synthesis as well as for machine translation. This architecture is instantiated in WebWOZ - a new web-based open-source WOZ prototyping platform. The viability of generic support is explored empirically through a series of evaluations. Researchers from a variety of backgrounds were able to create experiments, independent of their previous experience with WOZ. The approach was further validated through a number of real experiments, which also helped to identify a number of possibilities for additional support, and flagged potential issues relating to consistency in Wizard performance.
Stephan Schlögl, Gavin Doherty, Saturnino Luz
Interact. Comput.3
2014 Designing a serious game for community-based disease prevention in the Amazon
abstract
Many developing regions around the world rely on community-based healthcare strategies and practices to deal with prevention and control of often neglected diseases, by educating the local population and healthcare professionals, on the mechanisms by which such diseases spread and how they can be controlled. In this paper we describe a multiplayer serious game designed to raise awareness, and foster adoption of preventive measures among local citizens and community-health professionals about Leishmaniosis. We also discuss how the underlying concept for this game and its mechanics have been iteratively designed and developed in collaboration with a group of people with relevant medical and research expertise as well as practical knowledge resulting from working with our target population.
Saturnino Luz, Masood Masoodian, Manuel Cesario, Raquel Rangel Cesario, Bill Rogers 0001
Advances in Computer Entertainment1
2014 Readability of a background map layer under a semi-transparent foreground layer
abstract
This study investigates the readability (interpretability) of information presented on a geographical map onto which a semi-transparent multivariate selection layer has been overlaid. The investigation is based on an information visualization prototype developed for a mobile platform (tablet devices) which aimed at supporting epidemiologists and medical staff in field data collection and epidemiological interpretation tasks. Different factors are analysed under varying transparency (alpha blending) levels, including: map interpretation task (covering "seeing map" and "reading map" tasks), legend symbol and map area type. Our results complement other studies that focused on the readability characteristics of items displayed on semi-transparent foreground layers developed in the context of "toolglass" interfaces. The implications of these results to the usability of transparency variable selection layers in geographical map applications are also discussed.
Saturnino Luz, Masood Masoodian
AVI1
2014 A graph based abstraction of textual concordances and two renderings for their interactive visualisation
abstract
Concordancing, or the arranging of passages of a textual corpus in alphabetical order according to user-defined keywords, is one of the oldest and still most widely used forms of text analysis. It finds applications in areas such as lexicography, computational linguistics, translation studies and computer-assisted machine translation. Yet, the basic form of visualisation employed in the analysis of textual concordances has remained essentially the same since the keyword-in-context technique was introduced, over fifty years ago. This paper presents a generalisation of this technique as an analytical abstraction of concordances represented as undirected graphs, and then characterises keywords in terms of graph eccentricity properties. We illustrate this proposal with two distinct visual renderings: a mosaic (space-filling) display and a bi-directional hierarchical display. These displays can be used in isolation or in conjunction with traditional keyword-in-context components in an overview-plus-detail pattern, or as synchronised views. We discuss scenarios of use for these arrangements in lexicographical corpus analysis, in translation studies and in text comparison tasks.
Saturnino Luz, Shane Sheehan
AVI1
2014 Fostering smart energy applications through advanced visual interfaces
abstract
There is an increasing need for technology that assist people with more effective monitoring and management of their energy generation and consumption. In recent years a considerable number of research activities have resulted in a multitude of new ICT-supported tools and services for both the private energy consumer market, as well as for energy related business and industries (e.g., utility and grid companies, facility management, etc.). This workshop focuses on advanced interaction, interface, and visualization techniques for energy-related applications, tools, and services. It brings together researchers and practitioners from a diverse range of background, including interaction design, human-computer interaction, visualization, computer games, and other fields concerned with the development of advanced visual interfaces for smart energy applications.
Masood Masoodian, Elisabeth André, Saturnino Luz, Thomas Rist
AVI3
2014 Expanding the HCI Agenda in Healthcare
abstract
Designing technology for use in healthcare, and its evaluation in the healthcare setting, deserves special attention because of the nature of the special context of use. Biological hazards and the risk of infection, issues of privacy and security, system response times, as well as human factors and patient safety are identified as areas deserving of special attention. We give examples and quote from clinician interviews for illustration, and we argue that increased focus from the HCI community on these areas will bring tangible benefits of health information systems to healthcare staff, and ultimately improve patient services.
Bridget Kane, Saturnino Luz
CBMS2
2014 MLA'14: Third Multimodal Learning Analytics Workshop and Grand Challenges
abstract
This paper summarizes the third Multimodal Learning Analytics Workshop and Grand Challenges (MLA'14). This subfield of Learning Analytics focuses on the interpretation of the multimodal interactions that occurs in learning environments, both digital and physical. This is a hybrid event that includes presentations about methods and techniques to analyze and merge the different signals captured from these environments (workshop session) and more concrete results from the application of Multimodal Learning Analytics techniques to predict the performance of students while solving math problems or presenting in the classroom (challenges sessions). A total of eight articles will be presented in this event. The main conclusion from this event is that Multimodal Learning Analytics is a desirable research endeavour that could produce results that can be currently applied to improve the learning process.
Xavier Ochoa 0001, Marcelo Worsley, Katherine Chiluiza, Saturnino Luz
ICMI4
2014 Interlingual map task corpus collection
Hayakawa Akira, Nick Campbell 0001, Saturnino Luz
INTERSPEECH3
2013 In-class use of the nu-case mobile telehealth system in a medical school
abstract
As telehealth systems are increasingly being adopted by healthcare providers, a growing number of medical practitioners need to learn to use such systems. Despite this, however, most medical schools do not currently include telehealth-related content in their curriculum. In this paper we demonstrate, through examples from our own experience, how the Problem Based Learning method can be used to incorporate telehealth systems in medical science education. We also present a survey of medical students and faculty, conducted at a university in Brazil, to gain the participants' opinions on the nu-case mobile telehealth system. The results of the survey show that nu-case provides support for a range of tasks that a telehealth system needs to cater for in the context of our application area.
Marcelo Ballaben Carloni, Manuel Cesario, Raquel Rangel Cesario, Saturnino Luz, Edson Margarido, Masood Masoodian, Daniel Facciolo Pires
CBMS4
2013 Developing a framework for evaluation of technology use at multidisciplinary meetings in healthcare
abstract
Identifying an appropriate method to evaluate the use of technology during patient case discussions at multidisciplinary medical team (MDT) meetings, is problematic. A number of approaches conducted over an extended period of study are described and the lessons learned are explained. A framework is proposed to serve as a basis for evaluation of technology use in these complex collaborative work settings that incorporates hospital, technology and people perspectives.
Bridget Kane, Saturnino Luz, Pieter J. Toussaint
CBMS2
2013 Shared decision making needs a communication record
abstract
Increasing dependability in collaboration work among health professionals will directly improve patient outcomes, and reduce healthcare costs. Our research examines the development of a shared visual display to facilitate data entry and validation of an electronic record during multidisciplinary team meeting discussion, where specialists discuss patient symptoms, test results, and image findings. The problem of generating an electronic record for patient files that will serve as a record of collaboration, communication and a guide for later tasks is addressed through use of the shared visual display. Shortcomings in user-informed designed, structured data-entry screens became evident when in actual use. Time constraints prompt the synopsis of discussion in acronyms, free text, abbreviations, and the use of inferences. We demonstrate how common ground, team cohesiveness and the use of a shared visual display can improve dependability, but these factors can also provide a false sense of security and increase vulnerability in the patient management system.
Bridget Kane, Pieter J. Toussaint, Saturnino Luz
CSCW3
2013 Automatic identification of experts and performance prediction in the multimodal math data corpus through analysis of speech interaction
abstract
An analysis of multiparty interaction in the problem solving sessions of the Multimodal Math Data Corpus is presented. The analysis focuses on non-verbal cues extracted from the audio tracks. Algorithms for expert identification and performance prediction (correctness of solution) are implemented based on patterns of speech activity among session participants. Both of these categorisation algorithms employ an underlying graph-based representation of dialogues for each individual problem solving activities. The proposed Bayesian approach to expert prediction proved quite effective, reaching accuracy levels of over 92\% with as few as 6 dialogues of training data. Performance prediction was not quite as effective. Although the simple graph-matching strategy employed for predicting incorrect solutions improved considerably over a Monte Carlo simulated baseline (F1 score increased by a factor of 2.3), there is still much room for improvement in this task.
Saturnino Luz
ICMI1
2013 WebWOZ: A Platform for Designing and Conducting Web-based Wizard of Oz Experiments
Stephan Schlögl, Saturnino Luz, Gavin Doherty
SIGDIAL Conference2
2012 Supporting collaboration among healthcare professionals and disease surveillance in remote areas
abstract
It is estimated that global climate change and regional land use and cover changes in the Amazon region will contribute to the spread of vector-borne diseases such as bartonellosis and leishmaniasis. The large geographical distances and the sparsity of human settlements in the region pose challenges to the collaboration among health professionals whose goals range from diagnosing diseases to monitoring their spread. This paper presents work in progress on a system to support the tasks of local healthcare professionals and enabling collection, compilation, sharing and visualisation of data for purposes of epidemiological research and disease surveillance in remote regions.
Saturnino Luz, Masood Masoodian, Manuel Cesario, Bill Rogers 0001
CBMS1
2012 Time-based Geographical Mapping of Communicable Diseases
abstract
Information visualisation methods can potentially be employed to assist the work of epidemiologists and other health care professionals in mapping the spread of communicable diseases in remote areas, where the task of disease surveillance encompasses temporal elements such as changes in climate, land use and population movements. This paper presents an investigation into the use of time-based visualisation techniques coupled with geographical maps and support for distributed mobile collection of patient data. This study has focused on the development of an information visualisation prototype designed for use by epidemiology researchers on mobile platforms (tablets and smart phones). The prototyping activity has involved the participation of prospective users working in the Amazon region. Initial results are presented and discussed.
Manuel Cesario, Matthew Jervis, Saturnino Luz, Masood Masoodian, Bill Rogers 0001
IV3
2012 Collaboration in Translation: The Impact of Increased Reach on Cross-organisational Work
Gavin Doherty, Nikiforos Karamanis, Saturnino Luz
Comput. Support. Cooperative Work.3
2012 Cross-cultural assessment of automatically generated multimodal referring expressions in a virtual world
Ielka van der Sluis, Saturnino Luz, Werner Breitfuss, Mitsuru Ishizuka, Helmut Prendinger
Int. J. Hum. Comput. Stud.2
2012 The nonverbal structure of patient case discussions in multidisciplinary medical team meetings
abstract
Meeting analysis has a long theoretical tradition in social psychology, with established practical ramifications in computer science, especially in computer supported cooperative work. More recently, a good deal of research has focused on the issues of indexing and browsing multimedia records of meetings. Most research in this area, however, is still based on data collected in laboratories, under somewhat artificial conditions. This article presents an analysis of the discourse structure and spontaneous interactions at real-life multidisciplinary medical team meetings held as part of the work routine in a major hospital. It is hypothesized that the conversational structure of these meetings, as indicated by sequencing and duration of vocalizations, enables segmentation into individual patient case discussions. The task of segmenting audio-visual records of multidisciplinary medical team meetings is described as a topic segmentation task, and a method for automatic segmentation is proposed. An empirical evaluation based on hand labelled data is presented, which determines the optimal length of vocalization sequences for segmentation, and establishes the competitiveness of the method with approaches based on more complex knowledge sources. The effectiveness of Bayesian classification as a segmentation method, and its applicability to meeting segmentation in other domains are discussed.
Saturnino Luz
ACM Trans. Inf. Syst.1
2011 On record keeping at multidisciplinary team meetings
abstract
This paper explores issues related to record keeping at multidisciplinary medical team (MDT) meetings. Based on questionnaire and interview data with MDT members of various specialities, roles and teams, the information priorities for inclusion in a MDT meeting are identified. The utility and need for records after the meeting is discussed, and methods for gathering the information considered. Concerns are expressed that real-time data gathering at the meeting takes more time and risks turning the meeting into a group form-filling exercise. The value of interactive discussion among multidisciplinary peers is restated. The difficulties identified are discussed in the context of design implications for record-keeping at meetings. The role of records as a co-ordinating mechanism for tasks conducted after the meeting is emphasised, The dichotomy of having a record of a i) detailed prescriptive treatment plan, or ii) detailed diagnostic information with little treatment plan articulated, is explained.
Bridget Kane, Saturnino Luz
CBMS2
2011 Comparing Static Gantt and Mosaic Charts for Visualization of Task Schedules
abstract
A mosaic chart has been proposed for representation of events on a timeline. While early studies demonstrated the effectiveness of mosaics in supporting visualization of multimedia records on a meeting browser, the usability of mosaics as a static timeline visualization has not been studied in more general settings. This paper investigates the use of the mosaic charts for visualization of project schedules. A user study was conducted to compare a building project schedule encoded alternatively as a mosaic or as a Gantt chart. Although the study focused on static graphs, for which the Gantt technique is usually very effective, results showed that the users were as fast and accurate at answering the questions using the mosaic representation as they were using Gantt charts. The analysis and experiment indicated algorithmic, space-filling and interpretation limitations of the mosaic technique. We suggest possible design improvements to overcome some of these limitations.
Saturnino Luz, Masood Masoodian
IV1
2011 Translation practice in the workplace: contextual analysis and implications for machine translation
Nikiforos Karamanis, Saturnino Luz, Gavin Doherty
Mach. Transl.2
2010 Assessing support requirements for multidisciplinary team meetings
abstract
This paper profiles multidisciplinary team activity (MDT) in a typical teaching hospital setting and reports on a survey conducted among the teams to establish the information needs and constraints which affect their interaction. Support for pre-meeting work and post-meeting responsibilities is considered important in enabling the interaction at team meetings to be effective. Acoustics in the meeting room have higher priority over visual displays, although both `hearing the discussion' and `seeing images and colleagues' is important at meetings. Issues of time include scheduling and timing and are difficult to manage, particularly when individual roles belong to several MDTs and span more than one hospital. The potential is identified for automatic processing of some decisions by simple algorithm, which potentially will allow for more time at meetings to discuss complex patient cases in more detail.
Bridget Kane, Ken O'Byrne, Saturnino Luz
CBMS3
2010 Translation practice in the workplace and Machine Translation
Nikiforos Karamanis, Saturnino Luz, Gavin Doherty
EAMT2
2010 The relevance of timing, pauses and overlaps in dialogues: detecting topic changes in scenario based meetings
Saturnino Luz, Jing Su 0001
INTERSPEECH1
2010 Supporting Collaborative Transcription of Recorded Speech with a 3D Game Interface
Saturnino Luz, Masood Masoodian, Bill Rogers 0001
KES (4)1
2010 Assessing the effectiveness of conversational features for dialogue segmentation in medical team meetings and in the AMI corpus
Saturnino Luz, Jing Su 0001
SIGDIAL Conference1
2010 Fieldwork for requirements: Frameworks for mobile healthcare applications
Gavin Doherty, Joseph McKnight, Saturnino Luz
Int. J. Hum. Comput. Stud.3
2009 Assimilating information and offering a medical opinion in remote and co-located meetings
abstract
Discussion on patient data, among hospital staff, plays an increasingly important role in inter-specialist communication. Effectiveness of a discussion depends, among other factors, on how well its participants perceive, assimilate and interpret information exchanged during a discussion. This paper reports a field study conducted to assess information assimilation among medical observer participants during PCDs in a hospital. Medically trained observer participants undertook a questionnaire at multi-disciplinary medical team meetings (MDTMs) in teleconference and co-located settings. Results show that participants are more likely to offer opinions in teleconference while their expectations on the long-term effects of treatment are more realistic in co-located PCDs than in teleconference PCDs. Surprisingly, the presentation of clinical findings, radiology and pathology is perceived to be clearer in teleconference, and respondents believe that they follow the discussion, know the patient management plan and understand the basis for decisions, better in teleconference than in co-located PCDs. While a higher educational value is attributed to teleconference PCDs, evidence suggests a trend to have more errors in teleconference, less critical evaluation and no expression of disagreement with patient management decisions made in teleconference.
Bridget Kane, Saturnino Luz
CBMS2
2009 Classification of patient case discussions through analysis of vocalisation graphs
abstract
This paper investigates the use of amount and structure of talk as a basis for automatic classification of patient case discussions in multidisciplinary medical team meetings recorded in a real-world setting. We model patient case discussions as vocalisation graphs, building on research from the fields of interaction analysis and social psychology. These graphs are "content free" in that they only encode patterns of vocalisation and silence. The fact that it does not rely on automatic transcription makes the technique presented in this paper an attractive complement to more sophisticated speech processing methods as a means of indexing medical team meetings. We show that despite the simplicity of the underlying representation mechanism, accurate classification performance (F-scores: F_1 = 0.98, for medical patient case discussions, and F_1 = 0.97, for surgical case discussions) can be achieved with a simple k-nearest neighbour classifier when vocalisations are represented at the level of individual speakers. Possible applications of the method in health informatics for storage and retrieval of multimedia medical meeting records are discussed.
Saturnino Luz, Bridget Kane
ICMI1
2009 Chronos: A Tool for Interactive Scheduling and Visualisation of Task Hierarchies
abstract
Visualisation and structuring of tasks in a schedule, from relatively simple activities such as meeting scheduling to more complex ones such as project planning, has been traditionally supported by timeline representations similar to Gantt charts. Despite their popularity, Gantt charts suffer from a number of shortcomings, including poor representation of detail and inefficient use of screen real-estate,particularly when a large number of parallel tasks, each of which requiring its own representation space, need to be displayed. We have devised an alternative visualisation,called temporal mosaic, which addresses some of these shortcomings while utilising space more efficiently. We have recently shown that as a static visualisation temporal mosaics outperform Gantt charts in terms of their ability to convey time-based scheduling information to users engaged in various temporal inference tasks. This paper extends that research by presenting interactive techniques which support creation, and dynamic visualisation of task hierarchies and relationships. These techniques are illustrated through a system for direct manipulation of schedules in both Gantt chart and temporal mosaic formats.
Saturnino Luz, Masood Masoodian, Daniel McKenzie, Wim Vanden Broeck
IV1
2009 Evaluating an Algorithm for the Generation of Multimodal Referring Expressions in a Virtual World: A Pilot Study
Werner Breitfuss, Ielka van der Sluis, Saturnino Luz, Helmut Prendinger, Mitsuru Ishizuka
IVA3
2009 Achieving Diagnosis by Consensus
Bridget Kane, Saturnino Luz
Comput. Support. Cooperative Work.2
2008 A system for dynamic 3D visualisation of speech recognition paths
abstract
This paper presents an interactive visualisation system that assists users of semi-automatic speech transcription systems to assess alternative recognition results in real time and provide feedback to the speech recognition back-end in an intuitive manner. This prototype uses the OpenGL libraries to implement an animated 3D visual representation of alternative recognition results generated by the Sphinx automatic speech recognition system. It is expected that displaying alternatives dynamically will facilitate early detection of recognition errors and encourage user interaction, which in turn can be used to improve future recognition performance.
Saturnino Luz, Masood Masoodian, Bill Rogers 0001
AVI1
2008 Taking Lessons from Teleconference to Improve Same Time, Same Place Interaction
abstract
Performance on an information gathering task is shown to be superior in teleconference. Analysis of errors in an exercise revealed the data sources used in co-located and teleconference scenarios. The use of a visual display for text data, in addition to the audio source, is demonstrated in both co-located and teleconference discussions. Audio was used as a source of information more frequently in teleconference which resulted in an overall improvement in task performance. The lesson learned from the higher performance in teleconference, can be applied to improve performance at co-located meetings. Providing appropriate visual data together within audio enhanced spaces can be expected to improve the communication event and reduce medical errors. Results support proposals for the incorporation of physical spaces to improve communication in everyday work in co-operative workplaces, such as hospitals.
Bridget Kane, Saturnino Luz
CBMS2
2007 Active Learning with History-Based Query Selection for Text Categorisation
Michael Davy, Saturnino Luz
ECIR2
2007 Dimensionality reduction for active learning with nearest neighbour classifier in text categorisation problems
abstract
Dimensionality reduction techniques are commonly used in text categorisation problems to improve training and classification efficiency as well as to avoid overfitting. The best performing dimensionality reduction techniques for text categorisation are supervised, hence utilise the label information of the training data. Active learning is used to reduce the number of labelled training examples for problems where obtaining label information is expensive. Since the vast majority of data supplied to active learning are unlabelled, supervised dimensionality reduction techniques cannot be readily employed. For this reason, active learning in text categorisation problems do not perform dimensionality reduction thereby restricting the choice of classifier. In this paper we investigate unsupervised dimensionality reduction techniques in active learning for text categorisation problems. Two unsupervised techniques are investigated, namely document frequency and principal components analysis. We empirically show increased performance of active learning, using a k-nearest neighbour classifier, when dimensionality reduction is applied using the unsupervised techniques.
Michael Davy, Saturnino Luz
ICMLA2
2007 Visualisation of Parallel Data Streams with Temporal Mosaics
abstract
Despite its popularity and widespread use, timeline visualisation suffers from shortcomings which limit its use for displaying multiple data streams when the number of streams increases to more than a handful. This paper presents the Temporal Mosaic technique for visualisation of parallel time-based streams which addresses some of these shortcomings. Temporal mosaics provide a compact way of representing parallel streams of events by allocating a fixed drawing area to time intervals and partitioning that area according to the number of concurrent events. A user study is presented which compares this technique to a standard timeline representation technique in which events are depicted as horizontal bars and multiple streams are drawn in parallel along a vertical axis. Results of this user study show that users of the temporal mosaic visualisation perform significantly better at detecting concurrency, interval overlaps and inactivity than users of standard timelines.
Saturnino Luz, Masood Masoodian
IV1
2007 Meeting browsing
Matt-Mouley Bouamrane, Saturnino Luz
Multim. Syst.2
2007 An analytical evaluation of search by content and interaction patterns on multimodal meeting records
Matt-Mouley Bouamrane, Saturnino Luz
Multim. Syst.2
2006 Probing the Use and Value of Video for Multi-Disciplinary Medical Teams in Teleconference
abstract
Face-to-face interactions, via teleconferencing link, are investigated for multidisciplinary medical team (MDT) meetings. Comparison is made between colocated MDT meetings and those held in teleconference. Attitudes towards video changed positively, over a period of 8 months, following participants' experience of teleconferencing. Analysis of display screen use reveals 60% of case discussion time was time spent in face-to-face view with remote sites, contrary to expressed views of its relative unimportance. The value of the video link in MDT meetings, held in teleconference, is found to have higher than expected value
Bridget Kane, Saturnino Luz
CBMS2
2006 Navigating Multimodal Meeting Recordings with the Meeting Miner
Matt-Mouley Bouamrane, Saturnino Luz
FQAS2
2006 Gathering a corpus of multimodal computer-mediated meetings
Saturnino Luz, Matt-Mouley Bouamrane, Masood Masoodian
LREC1
2006 History-based visual mining of semi-structured audio and text
abstract
Accessing specific or salient parts of multimedia recordings remains a challenge as there is no obvious way of structuring and representing a mix of space-based and time-based media. A number of approaches have been proposed which usually involve translating the continuous component of the multimedia recording into a space-based representation, such as text from audio through automatic speech recognition and images from video (keyframes). In this paper, we present a novel technique which defines retrieval units in terms of a log of actions performed on space-based artifacts, and exploits timing properties and extended concurrency to construct a visual presentation of text and speech data. This technique can be easily adapted to any mix of space-based artifacts and continuous media
Matt-Mouley Bouamrane, Saturnino Luz, Masood Masoodian
MMM2
2006 Multidisciplinary Medical Team Meetings: An Analysis of Collaborative Working with Special Attention to Timing and Teleconferencing
Bridget Kane, Saturnino Luz
Comput. Support. Cooperative Work.2
2005 Enabling Change in Healthcare Structures through Teleconferencing
abstract
Developments in teleconferencing capabilities have made changes in work practices possible and are facilitating working partnerships between institutions. This study of a teleconferencing initiative examines how the work of a multi-disciplinary team (MDT) is affected by extending the meeting to remote locations. Challenges are highlighted, particularly in relation to the management of information flows, development of norms (meeting protocols) and record keeping.
Bridget Kane, Saturnino Luz, Gerard Menezes, Donal P. Hollywood
CBMS2
2005 Automatic Hypertext Keyphrase Detection
Daniel Kelleher, Saturnino Luz
IJCAI2
2005 A Model for Meeting Content Storage and Retrieval
abstract
This paper presents a model for storage of remote Internet-based multimedia meetings and information retrieval from textual and time-based content. The model builds on a theory of content mapping that exploits temporal and contextual relationships between media streams. Two prototypes are presented which illustrate the application of the model to a virtual meeting environment, and to a system for visualisation of meeting records on mobile devices. Implications of the proposed content mapping model with respect to interface design and non-linear browsing of time-based media are also discussed.
Saturnino Luz, Masood Masoodian
MMM1
2004 A mobile system for non-linear access to time-based data
abstract
Conventional interfaces for visualisation of time-based media support access to sequential data in a linear fashion. We present two visualisation interfaces for a mobile application that supports non-linear, structured browsing of multimedia recordings by exploiting certain features of concurrent multimedia streams. The system is built on a content mapping framework which automatically creates links between text and audio data by establishing "temporal neighbourhoods". It illustrates how non-linear browsing may be particularly valuable for devices with limited screen real-estate.
Saturnino Luz, Masood Masoodian
AVI1
2003 Compact Visualisation of Multimedia Interaction Records
abstract
We present an approach to building compact and effective information visualisation interfaces to allow users to browse and access content stored in continuous media. We introduce a content mapping framework, which supports time-based indexing of recordings of computer supported collaborative activity. In order to illustrate the temporal aspect of content mapping, we present HANMER, a system designed to provide fast access to audio-textual meeting recordings on personal digital assistants (PDAs). HANMER embodies techniques that exploit the processing capabilities and functionality of modern PDAs for visualisation and retrieval of multimedia meeting data.
Saturnino Luz, Masood Masoodian
IV1
2000 A Software Toolkit for Sharing and Accessing Corpora Over the Internet
Saturnino Luz
LREC1
2000 Using Tableaux to Automate the Lambek and Other Categorial Calculi
Saturnino Luz
Inf. Comput.1
1999 SMALTO: advising interface designers on the use of speech in multimodal systems
abstract
Describes a tool which provides system designers with advice on the use of speech input and/or output modalities in combination with other modalities in the design of multimodal systems. The tool, which we call SMALTO (Speech Modality AuxiLiary TOol), implements a theory of speech functionality and incorporates structured data extracted from a large corpus of claims about speech functionality from the recent literature on speech and multimodality. We envisage that SMALTO can be used both as a stand-alone hypertext system and as part of a complete design environment.
Niels Ole Bernsen, Saturnino Luz
MMSP2
1999 Meeting browser: a system for visualising and accessing audio in multicast meetings
abstract
We describe a tool which provides interactive graphical user-support within an Internet-based virtual meeting place where users may be connected via heterogeneous systems (e.g. workstation, PDA, or mobile telephone) and where spoken language is the main modality of interaction. Communicative turns are integrated with non-acoustic data to form the meeting history. The tool is accessed through a graphical user interface (GUI) component based on a "musical score" metaphor.
Saturnino Luz, David M. Roy
MMSP1
1996 Grammar Specification in Categorial Logics and Theorem Proving
Saturnino Luz
CADE1