Yvette Graham

dblp:05/8150 · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0001-6741-4855ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 10 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Towards Emotional Intelligence in Conversational AI: How Well Can LLMs Recognise Emotion in Conversations?
abstract
The ability to express contextually appropriate emotions remains a defining challenge in conversational AI. Meeting this challenge requires accurately annotated corpora, a need that inevitably collides with practical constraints. Manual annotation, while reliable, is often prohibitively expensive and time-consuming. Consequently, automatic emotion recognition has become a critical component in the fine-tuning pipeline, with Large Language Models (LLMs) emerging as a promising alternative to human annotators. Despite their considerable potential for generating training and evaluation data, the use of LLMs for this purpose introduces a subtle but significant risk: the creation of circular, potentially biased evaluation protocols that frequently go unacknowledged in empirical reporting. In this paper, we present a systematic evaluation of LLMs for conversational emotion recognition, using human annotation as a gold standard to quantify the degree to which reliance on LLMs may introduce error into both fine-tuning and evaluation workflows. We conduct a comprehensive assessment of multiple LLMs for emotion labeling across conversational contexts, and additionally examine a second critical factor: the impact of conversational context on LLM performance. Our results indicate that although LLMs benefit from access to prior conversational context, their utilization of such context differs substantially from human conversational understanding, suggesting that LLMs may not rely on the temporal ordering of conversational context in the same way humans do.
Islam A. Hassan, Yvette Graham
SIGIR2
2024 REFINE-LM: Mitigating Language Model Stereotypes via Reinforcement Learning
abstract
With the introduction of (large) language models, there has been significant concern about the unintended bias such models may inherit from their training data. A number of studies have shown that such models propagate gender stereotypes, as well as geographical and racial bias, among other biases. While existing works tackle this issue by preprocessing data and debiasing embeddings, the proposed methods require a lot of computational resources and annotation effort while being limited to certain types of biases. To address these issues, we introduce REFINE-LM, a debiasing method that uses reinforcement learning to handle different types of biases without any fine-tuning. By training a simple model on top of the word probability distribution of a LM, our bias agnostic reinforcement learning method enables model debiasing without human annotations or significant computational resources. Experiments conducted on a wide range of models, including several LMs, show that our method (i) significantly reduces stereotypical biases while preserving LMs performance; (ii) is applicable to different types of biases, generalizing across contexts such as gender, ethnicity, religion, and nationality-based biases; and (iii) it is not expensive to train.
Rameez Qureshi, Naïm Es-Sebbani, Luis Galárraga, Yvette Graham, Miguel Couceiro, Zied Bouraoui
ECAI4
2024 ALoRA: Allocating Low-Rank Adaptation for Fine-tuning Large Language Models
abstract
Zequan Liu, Jiawen Lyn, Wei Zhu, Xing Tian, Yvette Graham. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zequan Liu, Jiawen Lyn, Xing Tian, Yvette Graham
NAACL-HLT5
2023 Graph-Based Video-Language Learning with Multi-Grained Audio-Visual Alignment
abstract
Video-language learning has attracted significant attention in the fields of multimedia, computer vision and natural language processing in recent years. One of the key challenges in this area is how to effectively integrate visual and linguistic information to enable machines to understand video content and query information. In this work, we leverage graph-based representations and multi-grained audio-visual alignment to address this challenge. First, our approach starts by transforming video and query inputs into visual-scene graphs and semantic role graphs using a visual-scene parser and semantic role labeler respectively. These graphs are then encoded using graph neural networks to obtain enriched representations and combined to obtain a video-query joint representation that enhances the semantic expressivity of the inputs. Second, to achieve accurate matching of relevant parts of audio and visual features, we propose a multi-grained alignment module that aligns the audio and visual features at multiple scales. This enables us to effectively fuse the audio and visual information in a way that is consistent with the semantic-level information captured by the graph-based representations. Experiments on five representative datasets collected for Video Retrieval and Video Question Answering tasks show that our approach outperforms the literature on several metrics. Our extensive ablation studies demonstrate the effectiveness of graph-based representation and multi-grained audio-visual alignment.
Chenyang Lyu, Wenxi Li, Tianbo Ji, Longyue Wang, Liting Zhou, Cathal Gurrin, Linyi Yang, Yi Yu 0001, Yvette Graham, Jennifer Foster
ACM Multimedia9
2023 Memento: a prototype search engine for LSC 2021
abstract
Abstract In this extended paper, we describe our lifelog retrieval system called Memento which participated in the 2021 Lifelog Search Challenge in detail. Memento leverages semantic representations of images and textual queries projected into a common latent space to facilitate effective retrieval, aiming to bridge the existing semantic gap between complex visual scenes/events and user information needs expressed as textual and faceted queries. Our system also has a minimalist user interface which includes functionalities such as visual data filtering and temporal search. Finally, we include a comparative analysis of Memento’s performance at LSC 2021 and suggest improvements for future iterations of the system.
Naushad Alam, Yvette Graham
Multim. Tools Appl.2
2022 Achieving Reliable Human Assessment of Open-Domain Dialogue Systems
abstract
Evaluation of open-domain dialogue systems is highly challenging and development of better techniques is highlighted time and again as desperately needed.Despite substantial efforts to carry out reliable live evaluation of systems in recent competitions, annotations have been abandoned and reported as too unreliable to yield sensible results.This is a serious problem since automatic metrics are not known to provide a good indication of what may or may not be a high-quality conversation.Answering the distress call of competitions that have emphasized the urgent need for better evaluation techniques in dialogue, we present the successful development of human evaluation that is highly reliable while still remaining feasible and low cost.Self-replication experiments reveal almost perfectly repeatable results with a correlation of r = 0.969.Furthermore, due to the lack of appropriate methods of statistical significance testing, the likelihood of potential improvements to systems occurring due to chance is rarely taken into account in dialogue evaluation, and the evaluation we propose facilitates application of standard tests.Since we have developed a highly reliable evaluation method, new insights into system performance can be revealed.We therefore include a comparison of state-of-the-art models (i) with and without personas, to measure the contribution of personas to conversation quality, as well as (ii) prescribed versus freely chosen topics.Interestingly with respect to personas, results indicate that personas do not positively contribute to conversation quality as expected.
Tianbo Ji, Yvette Graham, Gareth J. F. Jones, Chenyang Lyu, Qun Liu 0001
ACL (1)2
2022 An Exploration into the Benefits of the CLIP model for Lifelog Retrieval
abstract
In this paper, we attempt to fine-tune the CLIP (Contrastive Language-Image Pre-Training) model on the Lifelog Question Answering dataset (LLQA) to investigate retrieval performance of the fine-tuned model over the zero-shot baseline model. We train the model adopting a weight space ensembling approach using a modified loss function to take into account the differences in our dataset (LLQA) when compared with the dataset the CLIP model was originally pretrained on. We further evaluate our fine-tuned model using visual as well as multimodal queries on multiple retrieval tasks, demonstrating improved performance over the zero-shot baseline model.
Ly-Duyen Tran, Naushad Alam, Yvette Graham, Linh Khanh Vo, Nghiem Tuong Diep, Binh T. Nguyen 0001, Liting Zhou, Cathal Gurrin
CBMI3
2022 Evaluation of Automatically Generated Video Captions Using Vision and Language Models
abstract
Vision and language models are easily transferred to other tasks. In particular, they have been demonstrated to work well in the evaluation of automatic image captioning. This has made it possible to evaluate systems without the need for references or additional information apart from the image and the caption. However, these models do not provide a straightforward way of evaluating videos. In this paper, we propose using these models for video captioning evaluation. We explore the use of both single image-based evaluation and different methods to include data from multiple frames. Experiments demonstrate that using clustering methods to select a few frames to compute the final score gives an excellent correlation with human judgment. The bias in the human annotations can also influence the metric, so we propose filtering the human assessments to discard outliers and improve the evaluation process.
Luis Lebron Casas, Yvette Graham, Noel E. O'Connor, Kevin McGuinness
ICIP2
2022 BERTHA: Video Captioning Evaluation Via Transfer-Learned Human Assessment
abstract
Evaluating video captioning systems is a challenging task as there are multiple factors to consider; for instance: the fluency of the caption, multiple actions happening in a single scene, and the human bias of what is considered important. Most metrics try to measure how similar the system generated captions are to a single or a set of human-annotated captions. This paper presents a new method based on a deep learning model to evaluate these systems. The model is based on BERT, which is a language model that has been shown to work well in multiple NLP tasks. The aim is for the model to learn to perform an evaluation similar to that of a human. To do so, we use a dataset that contains human evaluations of system generated captions. The dataset consists of the human judgments of the captions produces by the system participating in various years of the TRECVid video to text task. BERTHA obtain favourable results, outperforming the commonly used metrics in some setups.
Luis Lebron Casas, Yvette Graham, Kevin McGuinness, Konstantinos Kouramas, Noel E. O'Connor
LREC2
2021 Improving Unsupervised Question Answering via Summarization-Informed Question Generation
abstract
Question Generation (QG) is the task of generating a plausible question for a given pair.Template-based QG uses linguistically-informed heuristics to transform declarative sentences into interrogatives, whereas supervised QG uses existing Question Answering (QA) datasets to train a system to generate a question given a passage and an answer.A disadvantage of the heuristic approach is that the generated questions are heavily tied to their declarative counterparts.A disadvantage of the supervised approach is that they are heavily tied to the domain/language of the QA dataset used as training data.In order to overcome these shortcomings, we propose an unsupervised QG method which uses questions generated heuristically from summaries as a source of training data for a QG system.We make use of freely available news summary data, transforming declarative summary sentences into appropriate questions using heuristics informed by dependency parsing, named entity recognition and semantic role labeling.The resulting questions are then combined with the original news articles to train an end-to-end neural QG model.We extrinsically evaluate our approach using unsupervised QA: our QG model is used to generate synthetic QA pairs for training a QA model.Experimental results show that, trained with only 20k English Wikipedia-based synthetic QA pairs, the QA model substantially outperforms previous unsupervised models on three in-domain datasets (SQuAD1.1,Natural Questions, TriviaQA) and three out-of-domain datasets (NewsQA, BioASQ, DuoRC), demonstrating the transferability of the approach.
Chenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster, Xin Jiang 0002, Qun Liu 0001
EMNLP (1)3
2020 Contrasting Human Opinion of Non-factoid Question Answering with Automatic Evaluation
abstract
Evaluation in non-factoid question answering tasks generally takes the form of computation of automatic metric scores for systems on a sample test set of questions against human-generated reference answers. Conclusions drawn from the scores produced by automatic metrics inevitably lead to important decisions about future directions. Metrics commonly applied include ROUGE, adopted from the related field of summarization, BLEU and Meteor, both of the latter originally developed for evaluation of machine translation. In this paper, we pose the important question, given that question answering is evaluated by application of automatic metrics originally designed for other tasks, to what degree do the conclusions drawn from such metrics correspond to human opinion about system-generated answers? We take the task of machine reading comprehension (MRC) as a case study and to address this question, provide a new method of human evaluation developed specifically for the task at hand.
Tianbo Ji, Yvette Graham, Gareth J. F. Jones
CHIIR2
2020 Improving Document-Level Sentiment Analysis with User and Product Context
abstract
Past work that improves document-level sentiment analysis by encoding user and product information has been limited to considering only the text of the current review.We investigate incorporating additional review text available at the time of sentiment prediction that may prove meaningful for guiding prediction.Firstly, we incorporate all available historical review text belonging to the author of the review in question.Secondly, we investigate the inclusion of historical reviews associated with the current product (written by other users).We achieve this by explicitly storing representations of reviews written by the same user and about the same product and force the model to memorize all reviews for one particular user and product.Additionally, we drop the hierarchical architecture used in previous work to enable words in the text to directly attend to each other.Experiment results on IMDB, Yelp 2013 and Yelp 2014 datasets show improvement to state-of-the-art of more than 2 percentage points in the best case.
Chenyang Lyu, Jennifer Foster, Yvette Graham
COLING3
2020 Statistical Power and Translationese in Machine Translation Evaluation
abstract
The term translationese has been used to describe features of translated text, and in this paper, we provide detailed analysis of potential adverse effects of translationese on machine translation evaluation.Our analysis shows differences in conclusions drawn from evaluations that include translationese in test data compared to experiments that tested only with text originally composed in that language.For this reason we recommend that reverse-created test data be omitted from future machine translation test sets.In addition, we provide a reevaluation of a past machine translation evaluation claiming human-parity of MT.One important issue not previously considered is statistical power of significance tests applied to comparison of human and machine translation.Since the very aim of past evaluations was the investigation of ties between human and MT systems, power analysis is of particular importance, to avoid, for example, claims of human parity simply corresponding to Type II error resulting from the application of a low powered test.We provide detailed analysis of tests used in such evaluations to provide an indication of a suitable minimum sample size for future studies.
Yvette Graham, Barry Haddow, Philipp Koehn
EMNLP (1)1
2019 Exploring the Impact of Training Data Bias on Automatic Generation of Video Captions
Alan F. Smeaton, Yvette Graham, Kevin McGuinness, Noel E. O'Connor, Seán Quinn, Eric Arazo Sanchez
MMM (1)2
2018 Translating Pro-Drop Languages With Reconstruction Models
abstract
Pronouns are frequently omitted in pro-drop languages, such as Chinese, generally leading to significant challenges with respect to the production of complete translations. To date, very little attention has been paid to the dropped pronoun (DP) problem within neural machine translation (NMT). In this work, we propose a novel reconstruction-based approach to alleviating DP translation problems for NMT models. Firstly, DPs within all source sentences are automatically annotated with parallel information extracted from the bilingual training corpus. Next, the annotated source sentence is reconstructed from hidden representations in the NMT model. With auxiliary training objectives, in the terms of reconstruction scores, the parameters associated with the NMT model are guided to produce enhanced hidden representations that are encouraged as much as possible to embed annotated DP information. Experimental results on both Chinese-English and Japanese-English dialogue translation tasks show that the proposed approach significantly and consistently improves translation performance over a strong NMT baseline, which is directly built on the training data annotated with DPs.
Longyue Wang, Zhaopeng Tu, Shuming Shi 0001, Tong Zhang 0001, Yvette Graham, Qun Liu 0001
AAAI5
2017 Further Investigation into Reference Bias in Monolingual Evaluation of Machine Translation
abstract
Monolingual evaluation of Machine Translation (MT) aims to simplify human assessment by requiring assessors to compare the meaning of the MT output with a reference translation, opening up the task to a much larger pool of genuinely qualified evaluators.Monolingual evaluation runs the risk, however, of bias in favour of MT systems that happen to produce translations superficially similar to the reference and, consistent with this intuition, previous investigations have concluded monolingual assessment to be strongly biased in this respect.On re-examination of past analyses, we identify a series of potential analytical errors that force some important questions to be raised about the reliability of past conclusions, however.We subsequently carry out further investigation into reference bias via direct human assessment of MT adequacy via quality controlled crowd-sourcing.Contrary to both intuition and past conclusions, results show no significant evidence of reference bias in monolingual evaluation of MT.
Qingsong Ma, Yvette Graham, Timothy Baldwin, Qun Liu 0001
EMNLP2
2017 Can machine translation systems be evaluated by the crowd alone
abstract
Abstract Crowd-sourced assessments of machine translation quality allow evaluations to be carried out cheaply and on a large scale. It is essential, however, that the crowd's work be filtered to avoid contamination of results through the inclusion of false assessments. One method is to filter via agreement with experts, but even amongst experts agreement levels may not be high. In this paper, we present a new methodology for crowd-sourcing human assessments of translation quality, which allows individual workers to develop their own individual assessment strategy. Agreement with experts is no longer required, and a worker is deemed reliable if they are consistent relative to their own previous work. Individual translations are assessed in isolation from all others in the form of direct estimates of translation quality. This allows more meaningful statistics to be computed for systems and enables significance to be determined on smaller sets of assessments. We demonstrate the methodology's feasibility in large-scale human evaluation through replication of the human evaluation component of Workshop on Statistical Machine Translation shared translation task for two language pairs, Spanish-to-English and English-to-Spanish. Results for measurement based solely on crowd-sourced assessments show system rankings in line with those of the original evaluation. Comparison of results produced by the relative preference approach and the direct estimate method described here demonstrate that the direct estimate method has a substantially increased ability to identify significant differences between translation systems.
Yvette Graham, Timothy Baldwin, Alistair Moffat, Justin Zobel
Nat. Lang. Eng.1
2016 Is all that Glitters in Machine Translation Quality Estimation really Gold?
abstract
Human-targeted metrics provide a compromise between human evaluation of machine translation, where high inter-annotator agreement is difficult to achieve, and fully automatic metrics, such as BLEU or TER, that lack the validity of human assessment. Human-targeted translation edit rate (HTER) is by far the most widely employed human-targeted metric in machine translation, commonly employed, for example, as a gold standard in evaluation of quality estimation. Original experiments justifying the design of HTER, as opposed to other possible formulations, were limited to a small sample of translations and a single language pair, however, and this motivates our re-evaluation of a range of human-targeted metrics on a substantially larger scale. Results show significantly stronger correlation with human judgment for HBLEU over HTER for two of the nine language pairs we include and no significant difference between correlations achieved by HTER and HBLEU for the remaining language pairs. Finally, we evaluate a range of quality estimation systems employing HTER and direct assessment (DA) of translation adequacy as gold labels, resulting in a divergence in system rankings, and propose employment of DA for future quality estimation evaluations.
Yvette Graham, Timothy Baldwin, Meghan Dowling, Maria Eskevich, Teresa Lynn, Lamia Tounsi
COLING1
2016 Achieving Accurate Conclusions in Evaluation of Automatic Machine Translation Metrics
abstract
Automatic Machine Translation metrics, such as BLEU, are widely used in empirical evaluation as a substitute for human assessment.Subsequently, the performance of a given metric is measured by its strength of correlation with human judgment.When a newly proposed metric achieves a stronger correlation over that of a baseline, it is important to take into account the uncertainty inherent in correlation point estimates prior to concluding improvements in metric performance.Confidence intervals for correlations with human judgment are rarely reported in metric evaluations, however, and when they have been reported, the most suitable methods have unfortunately not been applied.For example, incorrect assumptions about correlation sampling distributions made in past evaluations risk over-estimation of significant differences in metric performance.In this paper, we provide analysis of each of the issues that may lead to inaccuracies before providing detail of a method that overcomes previous challenges.Additionally, we propose a new method of translation sampling that in contrast achieves genuine high conclusivity in evaluation of the relative performance of metrics.
Yvette Graham, Qun Liu 0001
HLT-NAACL1
2015 Improving Evaluation of Machine Translation Quality Estimation
abstract
Yvette Graham. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Yvette Graham
ACL (1)1
2015 Re-evaluating Automatic Summarization with BLEU and 192 Shades of ROUGE
abstract
We provide an analysis of current evaluation methodologies applied to summarization metrics and identify the following areas of concern: (1) movement away from evaluation by correlation with human assessment; (2) omission of important components of human assessment from evaluations, in addition to large numbers of metric variants; (3) absence of methods of significance testing improvements over a baseline.We outline an evaluation methodology that overcomes all such challenges, providing the first method of significance testing suitable for evaluation of summarization metrics.Our evaluation reveals for the first time which metric variants significantly outperform others, optimal metric variants distinct from current recommended best variants, as well as machine translation metric BLEU to have performance on-par with ROUGE for the purpose of evaluation of summarization systems.We subsequently replicate a recent large-scale evaluation that relied on, what we now know to be, suboptimal ROUGE variants revealing distinct conclusions about the relative performance of state-of-the-art summarization systems.
Yvette Graham
EMNLP1
2015 Accurate Evaluation of Segment-level Machine Translation Metrics
abstract
Evaluation of segment-level machine translation metrics is currently hampered by: (1) low inter-annotator agreement levels in human assessments; (2) lack of an effective mechanism for evaluation of translations of equal quality; and (3) lack of methods of significance testing improvements over a baseline.In this paper, we provide solutions to each of these challenges and outline a new human evaluation methodology aimed specifically at assessment of segment-level metrics.We replicate the human evaluation component of WMT-13 and reveal that the current state-of-the-art performance of segment-level metrics is better than previously believed.Three segment-level metrics -METEOR, NLEPOR and SENTBLEU-MOSES -are found to correlate with human assessment at a level not significantly outperformed by any other metric in both the individual language pair assessment for Spanish-to-English and the aggregated set of 9 language pairs.
Yvette Graham, Timothy Baldwin, Nitika Mathur
HLT-NAACL1
2014 Is Machine Translation Getting Better over Time?
abstract
Recent human evaluation of machine translation has focused on relative pref-erence judgments of translation quality, making it difficult to track longitudinal im-provements over time. We carry out a large-scale crowd-sourcing experiment to estimate the degree to which state-of-the-art performance in machine translation has increased over the past five years. To fa-cilitate longitudinal evaluation, we move away from relative preference judgments and instead ask human judges to provide direct estimates of the quality of individ-ual translations in isolation from alternate outputs. For seven European language pairs, our evaluation estimates an aver-age 10-point improvement to state-of-the-art machine translation between 2007 and 2012, with Czech-to-English translation standing out as the language pair achiev-ing most substantial gains. Our method of human evaluation offers an economi-cally feasible and robust means of per-forming ongoing longitudinal evaluation of machine translation. 1
Yvette Graham, Timothy Baldwin, Alistair Moffat, Justin Zobel
EACL1
2014 Testing for Significance of Increased Correlation with Human Judgment
abstract
Automatic metrics are widely used in ma-chine translation as a substitute for hu-man assessment. With the introduction of any new metric comes the question of just how well that metric mimics human assessment of translation quality. This is often measured by correlation with hu-man judgment. Significance tests are gen-erally not used to establish whether im-provements over existing methods such as BLEU are statistically significant or have occurred simply by chance, however. In this paper, we introduce a significance test for comparing correlations of two metrics, along with an open-source implementation of the test. When applied to a range of metrics across seven language pairs, tests show that for a high proportion of metrics, there is insufficient evidence to conclude significant improvement over BLEU. 1
Yvette Graham, Timothy Baldwin
EMNLP1
2008 Packed rules for automatic transfer-rule induction
Yvette Graham, Josef van Genabith
EAMT1