Arash Eshghi

dblp:58/3222 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
9since 2021 · last 2024
0000-0003-4711-4091ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Clarifying Completions: Evaluating How LLMs Respond to Incomplete Questions
abstract
People understand and produce language incrementally on a word by word basis. This gives rise to many characteristic conversational phenomena including long mid-sentence pauses that are followed by incremental clarification requests (iCRs) intended to recover the rest of the truncated turn (see Fig. 1; (A), (B), (C)). The ability to generate iCRs is important in natural conversational AI systems, and crucial to their accessibility to users with memory impairment. In this paper, we collect, release and analyse SLUICE-CR: a large corpus of 3000 human produced iCRs. We then use this corpus to probe the incremental processing capability of a number of state of the art LLMs by evaluating the quality of the model’s generated iCRs in response to incomplete questions. Our evaluations show that the ability to generate contextually appropriate iCRs only emerges at larger LLM sizes, and only when prompted with example iCRs from our corpus. They also indicate that autoregressive LMs are, in principle, able to both understand and generate language incrementally.
Angus Addlesee, Oliver Lemon, Arash Eshghi
LREC/COLING3
2024 Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models
abstract
In dialogue, the addressee may initially misunderstand the speaker and respond erroneously, often prompting the speaker to correct the misunderstanding in the next turn with a Third Position Repair (TPR).The ability to process and respond appropriately to such repair sequences is thus crucial in conversational AI systems.In this paper, we first collect, analyse, and publicly release BLOCKWORLD-REPAIRS: a dataset of multi-modal TPR sequences in an instruction-following manipulation task that is, by design, rife with referential ambiguity.We employ this dataset to evaluate several state-ofthe-art Vision and Language Models (VLM) across multiple settings, focusing on their capability to process and accurately respond to TPRs and thus recover from miscommunication.We find that, compared to humans, all models significantly underperform in this task.We then show that VLMs can benefit from specialised losses targeting relevant tokens during fine-tuning, achieving better performance and generalising better to new scenarios.Our results suggest that these models are not yet ready to be deployed in multi-modal collaborative settings where repairs are common, and highlight the need to design training regimes and objectives that facilitate learning from interaction.Our code and data are available at www.github. com/JChiyah/blockworld-repairs
Francisco Javier Chiyah Garcia, Alessandro Suglia, Arash Eshghi
EMNLP3
2024 Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs
abstract
Evaluating Large Language Models (LLMs) on reasoning benchmarks demonstrates their ability to solve compositional questions.However, little is known of whether these models engage in genuine logical reasoning or simply rely on implicit cues to generate answers.In this paper, we investigate the transitive reasoning capabilities of two distinct LLM architectures, LLaMA 2 and Flan-T5, by manipulating facts within two compositional datasets: QASC and Bamboogle.We controlled for potential cues that might influence the models' performance, including (a) word/phrase overlaps across sections of test input; (b) models' inherent knowledge during pre-training or fine-tuning; and (c) Named Entities.Our findings reveal that while both models leverage (a), Flan-T5 shows more resilience to experiments (b and c), having less variance than LLaMA 2. This suggests that models may develop an understanding of transitivity through fine-tuning on knowingly relevant datasets, a hypothesis we leave to future work 1 .
Houman Mehrafarin, Arash Eshghi, Ioannis Konstas
EMNLP2
2024 Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
abstract
This study explores replacing Transformers in Visual Language Models (VLMs) with Mamba, a recent structured state space model (SSM) that demonstrates promising performance in sequence modeling.We test models up to 3B parameters under controlled conditions, showing that Mamba-based VLMs outperforms Transformers-based VLMs in captioning, question answering, and reading comprehension.However, we find that Transformers achieve greater performance in visual grounding and the performance gap widens with scale.We explore two hypotheses to explain this phenomenon: 1) the effect of task-agnostic visual encoding on the updates of the hidden states, and 2) the difficulty in performing visual grounding from the perspective of in-context multimodal retrieval.Our results indicate that a task-aware encoding yields minimal performance gains on grounding, however, Transformers significantly outperform Mamba at incontext multimodal retrieval.Overall, Mamba shows promising performance on tasks where the correct output relies on a summary of the image but struggles when retrieval of explicit information from the context is required 1 .
Georgios Pantazopoulos, Malvina Nikandrou, Alessandro Suglia, Oliver Lemon, Arash Eshghi
EMNLP5
2023 Multitask Multimodal Prompted Training for Interactive Embodied Task Completion
abstract
Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh 0001, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia
EMNLP5
2023 No that's not what I meant: Handling Third Position Repair in Conversational Question Answering
abstract
The ability to handle miscommunication is crucial to robust and faithful conversational AI.People usually deal with miscommunication immediately as they detect it, using highly systematic interactional mechanisms called repair.One important type of repair is Third Position Repair (TPR) whereby a speaker is initially misunderstood but then corrects the misunderstanding as it becomes apparent after the addressee's erroneous response (see Fig. 1).Here, we collect and publicly release REPAIR-QA 1 , the first large dataset of TPRs in a conversational question answering (QA) setting.The data is comprised of the TPR turns, corresponding dialogue contexts, and candidate repairs of the original turn for execution of TPRs.We demonstrate the usefulness of the data by training and evaluating strong baseline models for executing TPRs.For stand-alone TPR execution, we perform both automatic and human evaluations on a fine-tuned T5 model, as well as OpenAI's GPT-3 LLMs.Additionally, we extrinsically evaluate the LLMs' TPR processing capabilities in the downstream conversational QA task.The results indicate poor out-of-thebox performance on TPR's by the GPT-3 models, which then significantly improves when exposed to REPAIR-QA.
Vevake Balaraman, Arash Eshghi, Ioannis Konstas, Ioannis Papaioannou
SIGDIAL2
2023 'What are you referring to?' Evaluating the Ability of Multi-Modal Dialogue Models to Process Clarificational Exchanges
abstract
Referential ambiguities arise in dialogue when a referring expression does not uniquely identify the intended referent for the addressee.Addressees usually detect such ambiguities immediately and work with the speaker to repair it using meta-communicative, Clarificational Exchanges (CE 1 ): a Clarification Request (CR) and a response.Here, we argue that the ability to generate and respond to CRs imposes specific constraints on the architecture and objective functions of multi-modal, visually grounded dialogue models.We use the SIMMC 2.0 dataset to evaluate the ability of different state-of-the-art model architectures to process CEs, with a metric that probes the contextual updates that arise from them in the model.We find that language-based models are able to encode simple multi-modal semantic information and process some CEs, excelling with those related to the dialogue history, whilst multi-modal models can use additional learning objectives to obtain disentangled object representations, which become crucial to handle complex referential ambiguities across modalities overall 2 .
Francisco Javier Chiyah Garcia, Alessandro Suglia, Arash Eshghi, Helen Hastie
SIGDIAL3
2022 Demonstrating EMMA: Embodied MultiModal Agent for Language-guided Action Execution in 3D Simulated Environments
abstract
Alessandro Suglia, Bhathiya Hemanthage, Malvina Nikandrou, Georgios Pantazopoulos, Amit Parekh, Arash Eshghi, Claudio Greco, Ioannis Konstas, Oliver Lemon, Verena Rieser. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022.
Alessandro Suglia, Bhathiya Hemanthage, Malvina Nikandrou, Georgios Pantazopoulos, Amit Parekh 0001, Arash Eshghi, Claudio Greco 0002, Ioannis Konstas, Oliver Lemon, Verena Rieser
SIGDIAL6
2021 A Study of Automatic Metrics for the Evaluation of Natural Language Explanations
abstract
As transparency becomes key for robotics and AI, it will be necessary to evaluate the methods through which transparency is provided, including automatically generated natural language (NL) explanations.Here, we explore parallels between the generation of such explanations and the much-studied field of evaluation of Natural Language Generation (NLG).Specifically, we investigate which of the NLG evaluation measures map well to explanations.We present the ExBAN corpus: a crowd-sourced corpus of NL explanations for Bayesian Networks.We run correlations comparing human subjective ratings with NLG automatic measures.We find that embedding-based automatic NLG evaluation methods, such as BERTScore and BLEURT, have a higher correlation with human ratings, compared to word-overlap metrics, such as BLEU and ROUGE.This work has implications for Explainable AI and transparent robotic and autonomous systems.
Miruna-Adriana Clinciu, Arash Eshghi, Helen Hastie
EACL2
2020 A Comprehensive Evaluation of Incremental Speech Recognition and Diarization for Conversational AI
abstract
Automatic Speech Recognition (ASR) systems are increasingly powerful and more accurate, but also more numerous with several options existing currently as a service (e.g.Google, IBM, and Microsoft).Currently the most stringent standards for such systems are set within the context of their use in, and for, Conversational AI technology.These systems are expected to operate incrementally in real-time, be responsive, stable, and robust to the pervasive yet peculiar characteristics of conversational speech such as disfluencies and overlaps.In this paper we evaluate the most popular of such systems with metrics and experiments designed with these standards in mind.We also evaluate the speaker diarization (SD) capabilities of the same systems which will be particularly important for dialogue systems designed to handle multi-party interaction.We found that Microsoft has the leading incremental ASR system which preserves disfluent materials and IBM has the leading incremental SD system in addition to the ASR that is most robust to speech overlaps.Google strikes a balance between the two but none of these systems are yet suitable to reliably handle natural spontaneous conversations in real-time.
Angus Addlesee, Yanchao Yu, Arash Eshghi
COLING3
2019 Data-Efficient Goal-Oriented Conversation with Dialogue Knowledge Transfer Networks
abstract
Igor Shalyminov, Sungjin Lee, Arash Eshghi, Oliver Lemon. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Igor Shalyminov, Arash Eshghi, Oliver Lemon
EMNLP/IJCNLP (1)3
2019 Few-Shot Dialogue Generation Without Annotated Data: A Transfer Learning Approach
abstract
Learning with minimal data is one of the key challenges in the development of practical, production-ready goal-oriented dialogue systems.In a real-world enterprise setting where dialogue systems are developed rapidly and are expected to work robustly for an evergrowing variety of domains, products, and scenarios, efficient learning from a limited number of examples becomes indispensable.In this paper, we introduce a technique to achieve state-of-the-art dialogue generation performance in a few-shot setup, without using any annotated data.We do this by leveraging background knowledge from a larger, more highly represented dialogue sourcenamely, the MetaLWOz dataset.We evaluate our model on the Stanford Multi-Domain Dialogue Dataset, consisting of human-human goal-oriented dialogues in in-car navigation, appointment scheduling, and weather information domains.We show that our few-shot approach achieves state-of-the art results on that dataset by consistently outperforming the previous best model in terms of BLEU and Entity F1 scores, while being more data-efficient by not requiring any data annotation.
Igor Shalyminov, Arash Eshghi, Oliver Lemon
SIGdial3
2017 Bootstrapping incremental dialogue systems from minimal data: the generalisation power of dialogue grammars
abstract
We investigate an end-to-end method for automatically inducing task-based dialogue systems from small amounts of unannotated dialogue data.It combines an incremental semantic grammar -Dynamic Syntax and Type Theory with Records (DS-TTR) -with Reinforcement Learning (RL), where language generation and dialogue management are a joint decision problem.The systems thus produced are incremental: dialogues are processed word-by-word, shown previously to be essential in supporting natural, spontaneous dialogue.We hypothesised that the rich linguistic knowledge within the grammar should enable a combinatorially large number of dialogue variations to be processed, even when trained on very few dialogues.Our experiments show that our model can process 74% of the Facebook AI bAbI dataset even when trained on only 0.13% of the data (5 dialogues).It can in addition process 65% of bAbI+, a corpus 1 we created by systematically adding incremental dialogue phenomena such as restarts and self-corrections to bAbI.We compare our model with a state-of-theart retrieval model, memn2n (Bordes et al., 2017).We find that, in terms of semantic accuracy, memn2n shows very poor robustness to the bAbI+ transformations even when trained on the full bAbI dataset.
Arash Eshghi, Igor Shalyminov, Oliver Lemon
EMNLP1
2017 VOILA: An Optimised Dialogue System for Interactively Learning Visually-Grounded Word Meanings (Demonstration System)
abstract
We present VOILA: an optimised, multimodal dialogue agent for interactive learning of visually grounded word meanings from a human user.VOILA is: (1) able to learn new visual categories interactively from users from scratch; (2) trained on real human-human dialogues in the same domain, and so is able to conduct natural spontaneous dialogue; (3) optimised to find the most effective trade-off between the accuracy of the visual categories it learns and the cost it incurs to users.VOILA is deployed on Furhat 1 , a humanlike, multi-modal robot head with backprojection of the face, and a graphical virtual character.
Yanchao Yu, Arash Eshghi, Oliver Lemon
SIGDIAL Conference2
2016 Incremental Generation of Visually Grounded Language in Situated Dialogue (demonstration system)
abstract
We present a multi-modal dialogue system for interactive learning of perceptually grounded word meanings from a human tutor (Yu et al., ).The system integrates an incremental, semantic, and bidirectional grammar framework -Dynamic Syntax and Type Theory with Records (DS-TTR 1 , (Eshghi et al., 2012;Kempson et al., 2001)) -with a set of visual classifiers that are learned throughout the interaction and which ground the semantic/contextual representations that it produces (c.f.Kennington & Schlangen (2015) where words, rather than semantic atoms, are grounded in visual classifiers).Our approach extends Dobnik et al. ( 2012) in integrating perception (vision in this case) and language within a single formal system: Type Theory with Records (TTR (Cooper, 2005)).The combination of deep semantic representations in TTR with an incremental grammar (Dynamic Syntax) allows for complex multi-turn dialogues to be parsed and generated (Eshghi et al., 2015).These include clarification interaction, corrections, ellipsis and utterance continuations (see e.g. the dialogue in Fig. 1).
Yanchao Yu, Arash Eshghi, Oliver Lemon
INLG2
2016 Training an adaptive dialogue policy for interactive learning of visually grounded word meanings
abstract
We present a multi-modal dialogue system for interactive learning of perceptually grounded word meanings from a human tutor.The system integrates an incremental, semantic parsing/generation framework -Dynamic Syntax and Type Theory with Records (DS-TTR) -with a set of visual classifiers that are learned throughout the interaction and which ground the meaning representations that it produces.We use this system in interaction with a simulated human tutor to study the effects of different dialogue policies and capabilities on accuracy of learned meanings, learning rates, and efforts/costs to the tutor.We show that the overall performance of the learning agent is affected by (1) who takes initiative in the dialogues;(2) the ability to express/use their confidence level about visual attributes; and (3) the ability to process elliptical and incrementally constructed dialogue turns.Ultimately, we train an adaptive dialogue policy which optimises the trade-off between classifier accuracy and tutoring costs.
Yanchao Yu, Arash Eshghi, Oliver Lemon
SIGDIAL Conference2
2012 Finishing each other's ... Responding to incomplete contributions in dialogue
Christine Howes, Patrick G. T. Healey, Matthew Purver, Arash Eshghi
CogSci4
2008 Communication Spaces
Patrick G. T. Healey, Graham White 0001, Arash Eshghi, Ahmad J. Reeves, Ann Light
Comput. Support. Cooperative Work.3