Matthew Purver

dblp:06/3180 · DBLP profile ↗
← Back
48ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0003-2297-1273ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploring Social Bias in Slovenia: The EEC-SL Dataset
Jaya Caporusso, Damar Hoogland, Boshko Koloski, Matthew Purver, Senja Pollak, Spela Vintar
LREC4
2026 FineDialFact: A Benchmark for Fine-Grained Dialogue Fact Verification
abstract
Large language models are known to produce hallucinations - factually incorrect or fabricated information - which poses significant challenges for many natural language processing applications, such as dialogue systems. As a result, detecting hallucinations has become a critical area of research. Current approaches to hallucination detection in dialogue systems primarily focus on verifying the factual consistency of generated responses. However, these responses often contain a mix of accurate, inaccurate or non-verifiable facts, making the use of a single factual label overly simplistic and coarse-grained. In this paper, we introduce a benchmark, FineDialFact, for fine-grained dialogue fact verification, which involves verifying atomic facts extracted from dialogue responses. To support this, we construct a dataset based on publicly available dialogue datasets and evaluate it using various baseline methods. Experimental results demonstrate that methods incorporating Chain-of-Thought reasoning can enhance performance in dialogue fact verification. Despite this, the best F1-score achieved on the HybriDialogue, an open-domain dialogue dataset, is only 0.74, indicating that the benchmark remains a challenging task for future research. We release our dataset and code at https://github.com/XiangyanChen/FineDialFact.
Xiangyan Chen, Yujian Gan, Arkaitz Zubiaga, Matthew Purver
LREC5
2026 Mono- and cross-lingual evaluation of representation language models on less-resourced languages
abstract
The current dominance of large language models in natural language processing is based on their contextual awareness. For text classification, text representation models, such as ELMo, BERT, and BERT derivatives, are typically fine-tuned for a specific problem. Most existing work focuses on English; in contrast, we present a large-scale multilingual empirical comparison of several monolingual and multilingual ELMo and BERT models using 14 classification tasks in nine languages. The results show, that the choice of best model largely depends on the task and language used, especially in a cross-lingual setting. In monolingual settings, monolingual BERT models tend to perform the best among BERT models. Among ELMo models, the ones trained on large corpora dominate. Cross-lingual knowledge transfer is feasible on most tasks already in a zero-shot setting without losing much performance.
Matej Ulcar, Ales Zagar, Carlos Santos Armendariz, Andraz Repar, Senja Pollak, Matthew Purver, Marko Robnik-Sikonja
Comput. Speech Lang.6
2025 Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities in comprehending and analyzing lengthy sequential inputs, owing to their extensive context windows that allow processing millions of tokens in a single forward pass. However, this paper uncovers a surprising limitation: LLMs fall short when handling long input sequences. We investigate this issue using three datasets and two tasks (sentiment analysis and news categorization) across various LLMs, including Claude 3, Gemini Pro, GPT 3.5 Turbo, Llama 3 Instruct, and Mistral Instruct models. To address this limitation, we propose and evaluate ad-hoc solutions that substantially enhance LLMs’ performance on long input sequences by up to 50%, while reducing API cost and latency by up to 93% and 50%, respectively.
Peyman Hosseini, Ignacio Castro, Iacopo Ghinassi, Matthew Purver
COLING4
2025 A Dataset for Expert Reviewer Recommendation with Large Language Models as Zero-shot Rankers
abstract
The task of reviewer recommendation is increasingly important, with main techniques utilizing general models of text relevance. However, state of the art (SotA) systems still have relatively high error rates. Two possible reasons for this are: a lack of large datasets and the fact that large language models (LLMs) have not yet been applied. To fill these gaps, we first create a substantial new dataset, in the domain of Internet specification documents; then we introduce the use of LLMs and evaluate their performance. We find that LLMs with prompting can improve on SotA in some cases, but that they are not a cure-all: this task provides a challenging setting for prompt-based methods
Vanja M. Karan, Stephen McQuistin, Ryo Yanagida, Colin Perkins, Gareth Tyson, Ignacio Castro, Patrick G. T. Healey, Matthew Purver
COLING8
2024 A Computational Analysis of the Dehumanisation of Migrants from Syria and Ukraine in Slovene News Media
abstract
Dehumanisation involves the perception and/or treatment of a social group’s members as less than human. This phenomenon is rarely addressed with computational linguistic techniques. We adapt a recently proposed approach for English, making it easier to transfer to other languages and to evaluate, introducing a new sentiment resource, the use of zero-shot cross-lingual valence and arousal detection, and a new method for statistical significance testing. We then apply it to study attitudes to migration expressed in Slovene newspapers, to examine changes in the Slovene discourse on migration between the 2015-16 migration crisis following the war in Syria and the 2022-23 period following the war in Ukraine. We find that while this discourse became more negative and more intense over time, it is less dehumanising when specifically addressing Ukrainian migrants compared to others.
Jaya Caporusso, Damar Hoogland, Mojca Brglez, Boshko Koloski, Matthew Purver, Senja Pollak
LREC/COLING5
2024 When Cohesion Lies in the Embedding Space: Embedding-Based Reference-Free Metrics for Topic Segmentation
abstract
In this paper we propose a new framework and new methods for the reference-free evaluation of topic segmentation systems directly in the embedding space. Specifically, we define a common framework for reference-free, embedding-based topic segmentation metrics, and show how this applies to an existing metric. We then define new metrics, based on a previously defined cohesion score, Average Relative Proximity. Using this approach, we show that Large Language Models (LLMs) yield features that, if used correctly, can strongly correlate with traditional topic segmentation metrics based on costly and rare human annotations, while outperforming existing reference-free metrics borrowed from clustering evaluation in most domains. We then show that smaller language models specifically fine-tuned for different sentence-level tasks can outperform LLMs several orders of magnitude larger. Via a thorough comparison of our metric’s performance across different datasets, we see that conversational data present the biggest challenge in this framework. Finally, we analyse the behaviour of our metrics in specific error cases, such as those of under-generation and moving of ground truth topic boundaries, and show that our metrics behave more consistently than other reference-free methods.
Iacopo Ghinassi, Lin Wang 0009, Chris Newell, Matthew Purver
LREC/COLING4
2024 Denoising Labeled Data for Comment Moderation Using Active Learning
abstract
Noisily labeled textual data is ample on internet platforms that allow user-created content. Training models, such as offensive language detection models for comment moderation, on such data may prove difficult as the noise in the labels prevents the model to converge. In this work, we propose to use active learning methods for the purposes of denoising training data for model training. The goal is to sample examples the most informative examples with noisy labels with active learning and send them to the oracle for reannotation thus reducing the overall cost of reannotation. In this setting we tested three existing active learning methods, namely DBAL, Variance of Gradients (VoG) and BADGE. The proposed approach to data denoising is tested on the problem of offensive language detection. We observe that active learning can be effectively used for the purposes of data denoising, however care should be taken when choosing the algorithm for this purpose.
Andraz Pelicon, Mladen Karan, Ravi Shekhar, Matthew Purver, Senja Pollak
LREC/COLING4
2024 Temporal Network Analysis of Email Communication Patterns in a Long Standing Hierarchy
abstract
An important concept in organisational behaviour is how hierarchy affects the voice of individuals, whereby members of a given organisation exhibit differing power relations based on their hierarchical position. Although there have been prior studies of the relationship between hierarchy and voice, they tend to focus on more qualitative small-scale methods and do not account for structural aspects of the organisation. This paper develops large-scale computational techniques utilising temporal network analysis to measure the effect that organisational hierarchy has on communication patterns throughout an organisation, focusing on the structure of pairwise interactions between individuals. To this end, we focus on one major organisation as a case study --- the Internet Engineering Task Force (IETF) --- a major technical standards development organisation for the Internet. A particularly useful feature of the IETF is a transparent hierarchy, where participants take on explicit roles (e.g., Area Directors, Working Group Chairs), and because its processes are open we have visibility into the communication of people at different hierarchy levels over a long time period. Exploiting this, we utilise a temporal network dataset of 989,911 email interactions among 23,741 participants to study how hierarchy impacts communication patterns. We show that the middle levels of the IETF are growing in terms of their dominance in communications. Higher levels consistently experience a higher proportion of incoming communication than lower levels, with higher levels initiating more communications too. We find that, overall, communication tends to flow "up" the hierarchy more than "down". Finally, we find that communication with higher-levels is associated with future communication more than for lower-levels, which we interpret as "facilitation". We conclude by discussing the implications this has on patterns within the wider IETF and the impact our analysis can have for other organisations.
Matthew Russell Barnes, Mladen Karan, Stephen McQuistin, Colin Perkins, Gareth Tyson, Matthew Purver, Ignacio Castro, Richard G. Clegg
ICWSM6
2024 Analyzing and Enhancing Clarification Strategies for Ambiguous References in Consumer Service Interactions
abstract
Changling Li, Yujian Gan, Zhenrong Yang, Youyang Chen, Xinxuan Qiu, Yanni Lin, Matthew Purver, Massimo Poesio. Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2024.
Changling Li, Yujian Gan, Zhenrong Yang, Youyang Chen, Xinxuan Qiu, Yanni Lin, Matthew Purver, Massimo Poesio
SIGDIAL7
2023 Reformulating NLP tasks to Capture Longitudinal Manifestation of Language Disorders in People with Dementia
abstract
Dementia is associated with language disorders which impede communication.Here, we automatically learn linguistic disorder patterns by making use of a moderately-sized pre-trained language model and forcing it to focus on reformulated natural language processing (NLP) tasks and associated linguistic patterns.Our experiments show that NLP tasks that encapsulate contextual information and enhance the gradient signal with linguistic patterns benefit performance.We then use the probability estimates from the best model to construct digital linguistic markers measuring the overall quality in communication and the intensity of a variety of language disorders.We investigate how the digital markers characterize dementia speech from a longitudinal perspective.We find that our proposed communication marker is able to robustly and reliably characterize the language of people with dementia, outperforming existing linguistic approaches; and shows external validity via significant correlation with clinical markers of behaviour.Finally, our proposed linguistic disorder markers provide useful insights into gradual language impairment associated with disease progression.
Dimitris Gkoumas, Matthew Purver, Maria Liakata
EMNLP2
2023 Exploring Pre-Trained Neural Audio Representations for Audio Topic Segmentation
abstract
Recent works have shown that audio embeddings can improve automatic topic segmentation of formats such as radio shows. In this work we expand the work in that direction by showing how and which publicly available, pre-trained neural audio embeddings can perform the task, without the need of any further fine-tuning of the audio encoders. The ranking of the encoders suggest that neural encoders pre-trained for speaker diarization and general purpose audio classification are the best suited to be used as features, beating non-neural baselines. We show that we can obtain perfect results on a newly created random dataset similar to the one used in previous work. We also show for the first time results on real-world data, proving that our method can be applied to actual radio shows with good results, but the choice of audio encoders is extremely important in order to achieve those. Finally, by releasing the datasets we used we make the contribution of providing the first (to our knowledge) publicly available, free of charge datasets for audio topic segmentation of media products.
Iacopo Ghinassi, Matthew Purver, Huy Phan, Chris Newell
ICME2
2023 Multimodal Topic Segmentation of Podcast Shows with Pre-trained Neural Encoders
abstract
We present two multimodal models for topic segmentation of podcasts built on pre-trained neural text and audio embeddings. We show that results can be improved by combining different modalities; but also by combining different encoders from the same modality, especially general-purpose sentence embeddings with specifically fine-tuned ones. We also show that audio embeddings can be substituted with two simple features related to sentence duration and inter-sentential pauses with comparable results. Finally, we publicly release our two datasets, the first in our knowledge publicly and freely available multimodal datasets for topic segmentation.
Iacopo Ghinassi, Lin Wang 0009, Chris Newell, Matthew Purver
ICMR4
2022 The Web We Weave: Untangling the Social Graph of the IETF
Prashant Khare, Mladen Karan, Stephen McQuistin, Colin Perkins, Gareth Tyson, Matthew Purver, Patrick G. T. Healey, Ignacio Castro
ICWSM6
2022 Misspelling Semantics in Thai
abstract
User-generated content is full of misspellings. Rather than being just random noise, we hypothesise that many misspellings contain hidden semantics that can be leveraged for language understanding tasks. This paper presents a fine-grained annotated corpus of misspelling in Thai, together with an analysis of misspelling intention and its possible semantics to get a better understanding of the misspelling patterns observed in the corpus. In addition, we introduce two approaches to incorporate the semantics of misspelling: Misspelling Average Embedding (MAE) and Misspelling Semantic Tokens (MST). Experiments on a sentiment analysis task confirm our overall hypothesis: additional semantics from misspelling can boost the micro F1 score up to 0.4-2%, while blindly normalising misspelling is harmful and suboptimal.
Pakawat Nakwijit, Matthew Purver
LREC2
2021 Towards Robustness of Text-to-SQL Models against Synonym Substitution
abstract
Yujian Gan, Xinyun Chen, Qiuping Huang, Matthew Purver, John R. Woodward, Jinxia Xie, Pengsheng Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yujian Gan, Qiuping Huang, Matthew Purver, John R. Woodward, Jinxia Xie, Pengsheng Huang
ACL/IJCNLP (1)4
2021 Exploring Underexplored Limitations of Cross-Domain Text-to-SQL Generalization
abstract
Recently, there has been significant progress in studying neural networks for translating text descriptions into SQL queries under the zeroshot cross-domain setting.Despite achieving good performance on some public benchmarks, we observe that existing text-to-SQL models do not generalize when facing domain knowledge that does not frequently appear in the training data, which may render the worse prediction performance for unseen domains.In this work, we investigate the robustness of text-to-SQL models when the questions require rarely observed domain knowledge.In particular, we define five types of domain knowledge and introduce Spider-DK (DK is the abbreviation of domain knowledge), a human-curated dataset based on the Spider benchmark for text-to-SQL translation.NL questions in Spider-DK are selected from Spider, and we modify some samples by adding domain knowledge that reflects real-world question paraphrases.We demonstrate that the prediction accuracy dramatically drops on samples that require such domain knowledge, even if the domain knowledge appears in the training set, and the model provides the correct predictions for related training samples. 1
Yujian Gan, Matthew Purver
EMNLP (1)3
2021 Evaluating Natural Language Descriptions Generated in a Workspace-Based Architecture
George A. Wright 0002, Matthew Purver
ICCC2
2021 Characterising the IETF through the lens of RFC deployment
abstract
Protocol standards, defined by the Internet Engineering Task Force (IETF), are crucial to the successful operation of the Internet. This paper presents a large-scale empirical study of IETF activities, with a focus on understanding collaborative activities, and how these underpin the publication of standards documents (RFCs). Using a unique dataset of 2.4 million emails, 8,711 RFCs and 4,512 authors, we examine the shifts and trends within the standards development process, showing how protocol complexity and time to produce standards has increased. With these observations in mind, we develop statistical models to understand the factors that lead to successful uptake and deployment of protocols, deriving insights to improve the standardisation process.
Stephen McQuistin, Mladen Karan, Prashant Khare, Colin Perkins, Gareth Tyson, Matthew Purver, Patrick G. T. Healey, Waleed Iqbal, Junaid Qadir 0001, Ignacio Castro
Internet Measurement Conference6
2021 Detecting Alzheimer's Disease Using Interactional and Acoustic Features from Spontaneous Speech
abstract
Alzheimer's Disease (AD) is a form of Dementia that manifests in cognitive decline including memory, language, and changes in behavior.Speech data has proven valuable for inferring cognitive status, used in many health assessment tasks, and can be easily elicited in natural settings.Much work focuses on analysis using linguistic features; here, we focus on non-linguistic features and their use in distinguishing AD patients from similar-age Non-AD patients with other health conditions in the Carolinas Conversation Collection (CCC) dataset.We used two types of features: patterns of interaction including pausing behaviour and floor control, and acoustic features including pitch, amplitude, energy, and cepstral coefficients.Fusion of the two kinds of features, combined with feature selection, obtains very promising classification results: classification accuracy of 90% using standard models such as support vector machines and logistic regression.We also obtain promising results using interactional features alone (87% accuracy), which can be easily extracted from natural conversations in daily life and thus have the potential for future implementation as a noninvasive method for AD diagnosis and monitoring.
Shamila Nasreen, Julian Hough, Matthew Purver
Interspeech3
2021 Alzheimer's Dementia Recognition Using Acoustic, Lexical, Disfluency and Speech Pause Features Robust to Noisy Inputs
abstract
We present two multimodal fusion-based deep learning models that consume ASR transcribed speech and acoustic data simultaneously to classify whether a speaker in a structured diagnostic task has Alzheimer's Disease and to what degree, evaluating the ADReSSo challenge 2021 data. Our best model, a BiLSTM with highway layers using words, word probabilities, disfluency features, pause information, and a variety of acoustic features, achieves an accuracy of 84% and RSME error prediction of 4.26 on MMSE cognitive scores. While predicting cognitive decline is more challenging, our models show improvement using the multimodal approach and word probabilities, disfluency and pause information over word-only models. We show considerable gains for AD classification using multimodal fusion and gating, which can effectively deal with noisy inputs from acoustic features and ASR hypotheses.
Morteza Rohanian, Julian Hough, Matthew Purver
Interspeech3
2021 Mitigating Topic Bias when Detecting Decisions in Dialogue
abstract
This work revisits the task of detecting decision-related utterances in multi-party dialogue.We explore performance of a traditional approach and a deep learning-based approach based on transformer language models, with the latter providing modest improvements.We then analyze topic bias in the models using topic information obtained by manual annotation.Our finding is that when detecting some types of decisions in our data, models rely more on topic specific words that decisions are about rather than on words that more generally indicate decision making.We further explore this by removing topic information from the train data.We show that this resolves the bias issues to an extent and, surprisingly, sometimes even boosts performance.
Mladen Karan, Prashant Khare, Patrick G. T. Healey, Matthew Purver
SIGDIAL4
2021 Rare-Class Dialogue Act Tagging for Alzheimer's Disease Diagnosis
abstract
Alzheimer's Disease (AD) is associated with many characteristic changes, not only in an individual's language, but also in the interactive patterns observed in dialogue.The most indicative changes of this latter kind tend to be associated with relatively rare dialogue acts (DAs), such as those involved in clarification exchanges and responses to particular kinds of questions.However, most existing work in DA tagging focuses on improving average performance, effectively prioritizing more frequent classes; it thus gives poor performance on these rarer classes and is not suited for application to AD analysis.In this paper, we investigate tagging specifically for rare class DAs, using a hierarchical BiLSTM model with various ways of incorporating information from previous utterances and DA tags in context.We show that this can give good performance for rare DA classes on both the general Switchboard corpus (SwDA) and an AD-specific conversational dataset, the Carolinas Conversation Collection (CCC); and that the tagger outputs then contribute useful information for distinguishing patients with and without AD.
Shamila Nasreen, Julian Hough, Matthew Purver
SIGDIAL3
2020 Creative Language Generation in a Society of Engagement and Reflection
George A. Wright 0002, Matthew Purver
ICCC2
2020 Multi-Modal Fusion with Gating Using Audio, Lexical and Disfluency Features for Alzheimer's Dementia Recognition from Spontaneous Speech
abstract
This paper is a submission to the Alzheimer's Dementia Recognition through Spontaneous Speech (ADReSS) challenge, which aims to develop methods that can assist in the automated prediction of severity of Alzheimer's Disease from speech data. We focus on acoustic and natural language features for cognitive impairment detection in spontaneous speech in the context of Alzheimer's Disease Diagnosis and the mini-mental state examination (MMSE) score prediction. We proposed a model that obtains unimodal decisions from different LSTMs, one for each modality of text and audio, and then combines them using a gating mechanism for the final prediction. We focused on sequential modelling of text and audio and investigated whether the disfluencies present in individuals' speech relate to the extent of their cognitive impairment. Our results show that the proposed classification and regression schemes obtain very promising results on both development and test sets. This suggests Alzheimer's Disease can be detected successfully with sequence modeling of the speech data of medical sessions.
Morteza Rohanian, Julian Hough, Matthew Purver
INTERSPEECH3
2020 CoSimLex: A Resource for Evaluating Graded Word Similarity in Context
abstract
State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Standard tasks and datasets for intrinsic evaluation of embeddings are based on judgements of similarity, but ignore context; standard tasks for word sense disambiguation take account of context but do not provide continuous measures of meaning similarity. This paper describes an effort to build a new dataset, CoSimLex, intended to fill this gap. Building on the standard pairwise similarity task of SimLex-999, it provides context-dependent similarity measures; covers not only discrete differences in word sense but more subtle, graded changes in meaning; and covers not only a well-resourced language (English) but a number of less-resourced languages. We define the task and evaluation metrics, outline the dataset collection methodology, and describe the status of the dataset so far.
Carlos Santos Armendariz, Matthew Purver, Matej Ulcar, Senja Pollak, Nikola Ljubesic, Mark Granroth-Wilding
LREC2
2020 How Furiously Can Colourless Green Ideas Sleep? Sentence Acceptability in Context
abstract
We study the influence of context on sentence acceptability. First we compare the acceptability ratings of sentences judged in isolation, with a relevant context, and with an irrelevant context. Our results show that context induces a cognitive load for humans, which compresses the distribution of ratings. Moreover, in relevant contexts we observe a discourse coherence effect that uniformly raises acceptability. Next, we test unidirectional and bidirectional language models in their ability to predict acceptability ratings. The bidirectional models show very promising results, with the best model achieving a new state-of-the-art for unsupervised acceptability prediction. The two sets of experiments provide insights into the cognitive aspects of sentence processing and central issues in the computational modeling of text and discourse.
Jey Han Lau, Carlos Santos Armendariz, Matthew Purver, Shalom Lappin
Trans. Assoc. Comput. Linguistics3
2019 Detecting Depression with Word-Level Multimodal Fusion
abstract
Copyright © 2019 ISCA Semi-structured clinical interviews are frequently used diagnostic tools for identifying depression during an assessment phase. In addition to the lexical content of a patient's responses, multimodal cues concurrent with the responses are indicators of their motor and cognitive state, including those derivable from their voice quality and gestural behaviour. In this paper, we use information from different modalities in order to train a classifier capable of detecting the binary state of a subject (clinically depressed or not), as well as the level of their depression. We propose a model that is able to perform modality fusion incrementally after each word in an utterance using a time-dependent recurrent approach in a deep learning set-up. To mitigate noisy modalities, we utilize fusion gates that control the degree to which the audio or visual modality contributes to the final prediction. Our results show the effectiveness of word-level multimodal fusion, achieving state-of-the-art results in depression detection and outperforming early feature-level and late fusion techniques.
Morteza Rohanian, Julian Hough, Matthew Purver
INTERSPEECH3
2017 Opening Up and Closing Down Discussion: Experimenting with Epistemic Status in Conversation
Shauna Concannon, Patrick G. T. Healey, Matthew Purver
CogSci3
2015 Conceptualizing Creativity: From Distributional Semantics to Conceptual Spaces
Kat Agres, Stephen McGregor, Matthew Purver, Geraint A. Wiggins
ICCC3
2014 Strongly Incremental Repair Detection
abstract
We present STIR (STrongly Incremental Repair detection), a system that detects speech repairs and edit terms on transcripts incrementally with minimal latency.STIR uses information-theoretic measures from n-gram models as its principal decision features in a pipeline of classifiers detecting the different stages of repairs.Results on the Switchboard disfluency tagged corpus show utterance-final accuracy on a par with state-of-the-art incremental repair detection methods, but with better incremental accuracy, faster time-to-detection and less computational overhead.We evaluate its performance using incremental metrics and propose new repair processing evaluation standards.
Julian Hough, Matthew Purver
EMNLP2
2014 Evaluating Neural Word Representations in Tensor-Based Compositional Settings
abstract
We provide a comparative study between neural word representations and traditional vector spaces based on cooccurrence counts, in a number of compositional tasks.We use three different semantic spaces and implement seven tensor-based compositional models, which we then test (together with simpler additive and multiplicative approaches) in tasks involving verb disambiguation and sentence similarity.To check their scalability, we additionally evaluate the spaces using simple compositional methods on larger-scale tasks with less constrained language: paraphrase detection and dialogue act tagging.In the more constrained tasks, co-occurrence vectors are competitive, although choice of compositional method is important; on the largerscale tasks, they are outperformed by neural word embeddings, which show robust, stable performance across the tasks.
Dmitrijs Milajevs, Dimitri Kartsaklis, Mehrnoosh Sadrzadeh, Matthew Purver
EMNLP4
2014 Computational Creativity: A Philosophical Approach, and an Approach to Philosophy
Stephen McGregor, Geraint A. Wiggins, Matthew Purver
ICCC3
2012 Whose turn is it anyway? Same- and cross-person compound contributions in dialogue
Christine Howes, Patrick G. T. Healey, Matthew Purver
CogSci3
2012 Finishing each other's ... Responding to incomplete contributions in dialogue
Christine Howes, Patrick G. T. Healey, Matthew Purver, Arash Eshghi
CogSci3
2012 Experimenting with Distant Supervision for Emotion Classification
Matthew Purver, Stuart Adam Battersby
EACL1
2012 Predicting Adherence to Treatment for Schizophrenia from Dialogue Transcripts
Christine Howes, Matthew Purver, Rosemarie McCabe, Patrick G. T. Healey, Mary Lavelle
SIGDIAL Conference2
2010 The CALO Meeting Assistant System
abstract
The CALO Meeting Assistant (MA) provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper presents the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, topic identification and segmentation, question-answer pair identification, action item recognition, decision extraction, and summarization.
Gökhan Tür, Andreas Stolcke, L. Lynn Voss, Stanley Peters, Dilek Hakkani-Tür, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri
IEEE Trans. Speech Audio Process.17
2009 Cascaded Lexicalised Classifiers for Second-Person Reference Resolution
Matthew Purver, Raquel Fernández, Matthew Frampton, Stanley Peters
SIGDIAL Conference1
2009 Split Utterances in Dialogue: a Corpus Study
Matthew Purver, Christine Howes, Eleni Gregoromichelaki, Patrick G. T. Healey
SIGDIAL Conference1
2008 Meeting adjourned: off-line learning interfaces for automatic meeting understanding
abstract
Upcoming technologies will automatically identify and extract certain types of general information from meetings, such as topics and the tasks people agree to do. We explore interfaces for presenting this information to users after a meeting is completed, using two post-meeting interfaces that display information from topics and action items respectively. These interfaces also provide an excellent forum for obtaining user feedback about the performance of classification algorithms, allowing the system to learn and improve with time. We describe how we manage the delicate balance of obtaining necessary feedback without overburdening users. We also evaluate the effectiveness of feedback from one interface on improvement of future action item detection.
Patrick Ehlen, Matthew Purver, John Niekrasz, Kari Lee, Stanley Peters
IUI2
2008 The CALO meeting speech recognition and understanding system
abstract
The CALO Meeting Assistant provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper summarizes the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, question-answer pair identification, action item recognition, decision extraction, and summarization.
Gökhan Tür, Andreas Stolcke, L. Lynn Voss, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Dilek Hakkani-Tür, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Stanley Peters, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri
SLT17
2008 A Probabilistic Model of Meetings That Combines Words and Discourse Features
abstract
In order to determine the points at which meeting discourse changes from one topic to another, probabilistic models were used to approximate the process through which meeting transcripts were produced. Gibbs sampling was used to estimate the values of random variables in the models, including the locations of topic boundaries. This paper shows how discourse features were integrated into the Bayesian model and reports empirical evaluations of the benefit obtained through the inclusion of each feature and of the suitability of alternative models of the placement of topic boundaries. It demonstrates how multiple cues to segmentation can be combined in a principled way, and empirical tests show a clear improvement over previous work.
Mike Dowman, Virginia Savova, Thomas L. Griffiths 0001, Konrad P. Kording, Josh Tenenbaum, Matthew Purver
IEEE Trans. Speech Audio Process.6
2007 Disambiguating Between Generic and Referential "You" in Dialog
Matthew Purver, Daniel Jurafsky
ACL2
2006 Unsupervised Topic Modelling for Multi-Party Spoken Discourse
abstract
We present a method for unsupervised topic modelling which adapts methods used in document classification (Blei et al., 2003; Griffiths and Steyvers, 2004) to unsegmented multi-party discourse transcripts. We show how Bayesian inference in this generative model can be used to simultaneously address the problems of topic segmentation and topic identification: automatically segmenting multi-party meetings into topically coherent segments with performance which compares well with previous unsupervised segmentation-only methods (Galley et al., 2003) while simultaneously extracting topics which rate highly when assessed for coherence by human judges. We also show that this method appears robust in the face of off-topic dialogue and speech recognition errors.
Matthew Purver, Konrad P. Kording, Thomas L. Griffiths 0001, Josh Tenenbaum
ACL1
2006 Robust interpretation in dialogue by combining confidence scores with contextual features
abstract
We present an approach to dialogue management and interpretation that evaluates and selects amongst candidate dialogue moves based on features at multiple levels. Multiple interpretation methods can be combined, multiple speech recognition and parsing hypotheses tested, and multiple candidate dialogue moves considered to choose the highest scoring hypothesis overall. We integrate hypotheses generated from shallow slot-filling methods and from relatively deep parsing, using pragmatic information. We show that this gives more robust performance than using either approach alone, allowing n-best list reordering to correct errors in speech recognition or parsing. Index Terms: dialogue management, robust interpretation 1.
Matthew Purver, Florin Ratiu, Lawrence Cavedon
INTERSPEECH1
2006 CHAT: a conversational helper for automotive tasks
abstract
Spoken dialogue interfaces, mostly command-and-control, become more visible in applications where attention needs to be shared with other tasks, such as driving a car. The deployment of the simple dialog systems, instead of more sophisticated ones, is partly because the computing platforms used for such tasks have been less powerful and partly because certain issues from these cognitively challenging tasks have not been well addressed even in the most advanced dialog systems. This paper reports the progress of our research effort in developing a robust, wide-coverage, and cognitive load-sensitive spoken dialog interface called CHAT: Conversational Helper for Automotive Tasks. Our research in the past few years has led to promising results, including high task completion rate, dialog efficiency, and improved user experience. Index Terms: dialog systems, cognitive load, robustness
Fuliang Weng, Sebastian Varges, Badri Raghunathan, Florin Ratiu, Heather Pon-Barry, Brian Lathrop, Harry Bratt, Tobias Scheideck, Matthew Purver, Annie Lien, Madhuri Raya, Stanley Peters, J. Russell, Lawrence Cavedon, Elizabeth Shriberg, Hauke Schmidt, R. Prieto
INTERSPEECH11
2004 Context-Based Incremental Generation for Dialogue
Matthew Purver, Ruth Kempson
INLG1