Alessandro Suglia

dblp:184/4588 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0002-3177-5197ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Playpen: An Environment for Exploring Learning From Dialogue Game Feedback
abstract
Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia
EMNLP16
2025 CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts
abstract
Malvina Nikandrou, Georgios Pantazopoulos, Nikolas Vitsakis, Ioannis Konstas, Alessandro Suglia. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Malvina Nikandrou, Georgios Pantazopoulos, Nikolas Vitsakis, Ioannis Konstas, Alessandro Suglia
NAACL (Long Papers)5
2024 Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks
abstract
Evaluating the generalisation capabilities of multimodal models based solely on their performance on out-of-distribution data fails to capture their true robustness.This work introduces a comprehensive evaluation framework that systematically examines the role of instructions and inputs in the generalisation abilities of such models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity.The proposed framework uncovers the resilience of multimodal models to extreme instruction perturbations and their vulnerability to observational changes, raising concerns about overfitting to spurious correlations.By employing this evaluation framework on current Transformerbased multimodal models for robotic manipulation tasks, we uncover limitations and suggest future advancements should focus on architectural and training innovations that better integrate multimodal inputs, enhancing a model's generalisation prowess by prioritising sensitivity to input content over incidental correlations.1 L1 L2 L3 L4 (a) Trained on Original; Evaluated on Original Cross-Attn + Obj-Centric 79.3 78.8 72.3 48.6 Cross-Attn + Patches 63.0 62.0 44.9 13.9 Concatenate + Obj-Centric 79.2 78.8 77.1 49.2 Concatenate + Patches 68.0 66.3 52.9 23.4(b) Trained on Original; Evaluated on Paraphrases Cross-Attn + Obj-Centric 78.6 77.6 69.8 47.1 Cross-Attn + Patches 61.1 58.5 45.3 16.8 Concatenate + Obj-Centric 71.5 72.2 62.7 43.0 Concatenate + Patches 61.3 57.0 46.0 20.5 (c) Trained on Paraphrases; Evaluated on Original Cross-Attn + Obj-Centric 82.7 81.8 77.4 48.0 Cross-Attn + Patches 63.9 63.0 49.5 20.4 Concatenate + Obj-Centric 80.4 78.2 74.8 49.0 Concatenate + Patches 67.1 62.8 52.0 19.8(d) Trained on Paraphrases; Evaluated on Paraphrases Cross-Attn + Obj-Centric 77.4 77.5 70.8 48.6 Cross-Attn + Patches 62.2 61.0 45.7 16.1 Concatenate + Obj-Centric 68.8 67.2 59.6 46.0 Concatenate + Patches 67.2 67.8 60.5 46.
Amit Parekh 0001, Nikolas Vitsakis, Alessandro Suglia, Ioannis Konstas
EMNLP3
2024 Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models
abstract
In dialogue, the addressee may initially misunderstand the speaker and respond erroneously, often prompting the speaker to correct the misunderstanding in the next turn with a Third Position Repair (TPR).The ability to process and respond appropriately to such repair sequences is thus crucial in conversational AI systems.In this paper, we first collect, analyse, and publicly release BLOCKWORLD-REPAIRS: a dataset of multi-modal TPR sequences in an instruction-following manipulation task that is, by design, rife with referential ambiguity.We employ this dataset to evaluate several state-ofthe-art Vision and Language Models (VLM) across multiple settings, focusing on their capability to process and accurately respond to TPRs and thus recover from miscommunication.We find that, compared to humans, all models significantly underperform in this task.We then show that VLMs can benefit from specialised losses targeting relevant tokens during fine-tuning, achieving better performance and generalising better to new scenarios.Our results suggest that these models are not yet ready to be deployed in multi-modal collaborative settings where repairs are common, and highlight the need to design training regimes and objectives that facilitate learning from interaction.Our code and data are available at www.github. com/JChiyah/blockworld-repairs
Francisco Javier Chiyah Garcia, Alessandro Suglia, Arash Eshghi
EMNLP2
2024 Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
abstract
This study explores replacing Transformers in Visual Language Models (VLMs) with Mamba, a recent structured state space model (SSM) that demonstrates promising performance in sequence modeling.We test models up to 3B parameters under controlled conditions, showing that Mamba-based VLMs outperforms Transformers-based VLMs in captioning, question answering, and reading comprehension.However, we find that Transformers achieve greater performance in visual grounding and the performance gap widens with scale.We explore two hypotheses to explain this phenomenon: 1) the effect of task-agnostic visual encoding on the updates of the hidden states, and 2) the difficulty in performing visual grounding from the perspective of in-context multimodal retrieval.Our results indicate that a task-aware encoding yields minimal performance gains on grounding, however, Transformers significantly outperform Mamba at incontext multimodal retrieval.Overall, Mamba shows promising performance on tasks where the correct output relies on a summary of the image but struggles when retrieval of explicit information from the context is required 1 .
Georgios Pantazopoulos, Malvina Nikandrou, Alessandro Suglia, Oliver Lemon, Arash Eshghi
EMNLP3
2024 Visually Grounded Language Learning: A Review of Language Games, Datasets, Tasks, and Models
abstract
In recent years, several machine learning models have been proposed. They are trained with a language modelling objective on large-scale text-only data. With such pretraining, they can achieve impressive results on many Natural Language Understanding and Generation tasks. However, many facets of meaning cannot be learned by “listening to the radio” only. In the literature, many Vision+Language (V+L) tasks have been defined with the aim of creating models that can ground symbols in the visual modality. In this work, we provide a systematic literature review of several tasks and models proposed in the V+L field. We rely on Wittgenstein’s idea of ‘language games’ to categorise such tasks into 3 different families: 1) discriminative games, 2) generative games, and 3) interactive games. Our analysis of the literature provides evidence that future work should be focusing on interactive games where communication in Natural Language is important to resolve ambiguities about object referents and action plans and that physical embodiment is essential to understand the semantics of situations and events. Overall, these represent key requirements for developing grounded meanings in neural models.
Alessandro Suglia, Ioannis Konstas, Oliver Lemon
J. Artif. Intell. Res.1
2023 Multitask Multimodal Prompted Training for Interactive Embodied Task Completion
abstract
Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh 0001, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia
EMNLP9
2023 'What are you referring to?' Evaluating the Ability of Multi-Modal Dialogue Models to Process Clarificational Exchanges
abstract
Referential ambiguities arise in dialogue when a referring expression does not uniquely identify the intended referent for the addressee.Addressees usually detect such ambiguities immediately and work with the speaker to repair it using meta-communicative, Clarificational Exchanges (CE 1 ): a Clarification Request (CR) and a response.Here, we argue that the ability to generate and respond to CRs imposes specific constraints on the architecture and objective functions of multi-modal, visually grounded dialogue models.We use the SIMMC 2.0 dataset to evaluate the ability of different state-of-the-art model architectures to process CEs, with a metric that probes the contextual updates that arise from them in the model.We find that language-based models are able to encode simple multi-modal semantic information and process some CEs, excelling with those related to the dialogue history, whilst multi-modal models can use additional learning objectives to obtain disentangled object representations, which become crucial to handle complex referential ambiguities across modalities overall 2 .
Francisco Javier Chiyah Garcia, Alessandro Suglia, Arash Eshghi, Helen Hastie
SIGDIAL2
2022 ACT-Thor: A Controlled Benchmark for Embodied Action Understanding in Simulated Environments
abstract
Artificial agents are nowadays challenged to perform embodied AI tasks. To succeed, agents must understand the meaning of verbs and how their corresponding actions transform the surrounding world. In this work, we propose ACT-Thor, a novel controlled benchmark for embodied action understanding. We use the AI2-THOR simulated environment to produce a controlled setup in which an agent, given a before-image and an associated action command, has to determine what the correct after-image is among a set of possible candidates. First, we assess the feasibility of the task via a human evaluation that resulted in 81.4% accuracy, and very high inter-annotator agreement (84.9%). Second, we design both unimodal and multimodal baselines, using state-of-the-art visual feature extractors. Our evaluation and error analysis suggest that only models that have a very structured representation of the actions together with powerful visual features can perform well on the task. However, they still fall behind human performance in a zero-shot scenario where the model is exposed to unseen (action, object) pairs. This paves the way for a systematic way of evaluating embodied AI agents that understand grounded actions.
Michael Hanna 0001, Federico Pedeni, Alessandro Suglia, Alberto Testoni, Raffaella Bernardi
COLING3
2022 Demonstrating EMMA: Embodied MultiModal Agent for Language-guided Action Execution in 3D Simulated Environments
abstract
Alessandro Suglia, Bhathiya Hemanthage, Malvina Nikandrou, Georgios Pantazopoulos, Amit Parekh, Arash Eshghi, Claudio Greco, Ioannis Konstas, Oliver Lemon, Verena Rieser. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022.
Alessandro Suglia, Bhathiya Hemanthage, Malvina Nikandrou, Georgios Pantazopoulos, Amit Parekh 0001, Arash Eshghi, Claudio Greco 0002, Ioannis Konstas, Oliver Lemon, Verena Rieser
SIGDIAL1
2021 An Empirical Study on the Generalization Power of Neural Representations Learned via Visual Guessing Games
abstract
Alessandro Suglia, Yonatan Bisk, Ioannis Konstas, Antonio Vergari, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Alessandro Suglia, Yonatan Bisk, Ioannis Konstas, Antonio Vergari, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon
EACL1
2020 CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language Learning
abstract
Alessandro Suglia, Ioannis Konstas, Andrea Vanzo, Emanuele Bastianelli, Desmond Elliott, Stella Frank, Oliver Lemon. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Alessandro Suglia, Ioannis Konstas, Andrea Vanzo, Emanuele Bastianelli, Desmond Elliott, Stella Frank, Oliver Lemon
ACL1
2020 Imagining Grounded Conceptual Representations from Perceptual Information in Situated Guessing Games
abstract
Alessandro Suglia, Antonio Vergari, Ioannis Konstas, Yonatan Bisk, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon. Proceedings of the 28th International Conference on Computational Linguistics. 2020.
Alessandro Suglia, Antonio Vergari, Ioannis Konstas, Yonatan Bisk, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon
COLING1
2020 Predicting Perceptual Speed from Search Behaviour
abstract
Perceptual Speed (PS) is a cognitive ability that is known to affect multiple factors in Information Retrieval (IR) such as a user's search performance and subjective experience. However PS tests are difficult to administer which limits the design of user-adaptive systems that can automatically infer PS to appropriately accommodate low PS users. Consequently, this paper evaluated whether PS can be automatically classified from search behaviour using several machine learning models trained on features extracted from TREC Common Core search task logs. Our results are encouraging: given a user's interactions from one query, a Decision Tree was able to predict a user's PS as low or high with 86% accuracy. Additionally, we identified different behavioural components for specific PS tests, implying that each PS test measures different aspects of a person's cognitive ability. These findings motivate further work for how best to design search systems that can adapt to individual differences.
Olivia Foulds, Alessandro Suglia, Leif Azzopardi, Martin Halvey
SIGIR2
2019 Bridging the gap between linked open data-based recommender systems and distributed representations
Pierpaolo Basile, Claudio Greco 0002, Alessandro Suglia, Giovanni Semeraro
Inf. Syst.3
2017 A Deep Architecture for Content-based Recommendations Exploiting Recurrent Neural Networks
abstract
In this paper we investigate the effectiveness of Recurrent Neural Networks (RNNs) in a top-N content-based recommendation scenario. Specifically, we propose a deep architecture which adopts Long Short Term Memory (LSTM) networks to jointly learn two embeddings representing the items to be recommended as well as the preferences of the user. Next, given such a representation, a logistic regression layer calculates the relevance score of each item for a specific user and we returns the top-N items as recommendations.
Alessandro Suglia, Claudio Greco 0002, Cataldo Musto, Marco de Gemmis, Pasquale Lops, Giovanni Semeraro
UMAP1