EDBT 2026 Demo / reviewers in the wild / expert
Ignacio Iacobacci
dblp:166/1757
· DBLP profile ↗
14ranked-venue papers
3as first author
8since 2021 · last 2024
0000-0002-8913-8561ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Trustworthy machine learning · 19% Representation and self-supervised learning · 13% Transfer learning and domain adaptation · 13% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 67% Image and video processing · 33% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% | |
| Theoretical computer science
1 paper |
Automata and formal languages · 100% |
Topics — the 24 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
scene understanding |
0.8 | 1 | 2024 | MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation · CVPR 2024 |
Visual content generation and editing
image editing |
0.8 | 1 | 2024 | MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation · CVPR 2024 |
Image and video processing › image decomposition › image separation
layer separation |
0.8 | 1 | 2024 | MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation · CVPR 2024 |
Visual content generation and editing › image generation
text-to-image generation |
0.8 | 1 | 2024 | MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation · CVPR 2024 |
Program synthesis and code generation › DSL-based synthesis
regular expression synthesis |
0.8 | 1 | 2024 | Correct and Optimal: The Regular Expression Inference Challenge · IJCAI 2024 |
Automata and formal languages › grammatical inference
regular expression inference |
0.8 | 1 | 2024 | Correct and Optimal: The Regular Expression Inference Challenge · IJCAI 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.7 | 1 | 2023 | A Systematic Study of Performance Disparities in Multilingual Task-Oriented Dialogue Systems · EMNLP 2023 |
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
0.7 | 1 | 2023 | A Systematic Study of Performance Disparities in Multilingual Task-Oriented Dialogue Systems · EMNLP 2023 |
Machine learning › Representation and self-supervised learning › word representation › word embedding
word sense representation |
0.6 | 2 | 2019 | LSTMEmbed: Learning Word and Sense Representations from a Large Semantically Annotated Corpus with Long Short-Term Memories · ACL (1) 2019 SensEmbed: Learning Sense Embeddings for Word and Relational Similarity · ACL (1) 2015 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.6 | 1 | 2022 | Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLU · EMNLP 2022 |
Machine learning › Learning paradigms
curriculum learning |
0.6 | 1 | 2022 | Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLU · EMNLP 2022 |
Machine learning › Transfer learning and domain adaptation › cross-lingual transfer
zero-shot cross-lingual transfer |
0.6 | 1 | 2022 | Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLU · EMNLP 2022 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.5 | 1 | 2021 | Improving Commonsense Causal Reasoning by Adversarial Training and Data Augmentation · AAAI 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
0.5 | 1 | 2021 | Improving Commonsense Causal Reasoning by Adversarial Training and Data Augmentation · AAAI 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.5 | 1 | 2021 | Improving Commonsense Causal Reasoning by Adversarial Training and Data Augmentation · AAAI 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | Improving Commonsense Causal Reasoning by Adversarial Training and Data Augmentation · AAAI 2021 |
Natural language and speech › Information extraction and text analysis
word sense disambiguation |
0.5 | 2 | 2016 | Embeddings for Word Sense Disambiguation: An Evaluation Study · ACL (1) 2016 SensEmbed: Learning Sense Embeddings for Word and Relational Similarity · ACL (1) 2015 |
Natural language and speech › Language models and text generation › text representation
contextualized word embeddings |
0.4 | 1 | 2020 | Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQA · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering |
0.4 | 1 | 2020 | Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQA · EMNLP (1) 2020 |
Machine learning › Representation and self-supervised learning › word representation
sense embeddings |
0.4 | 1 | 2019 | LSTMEmbed: Learning Word and Sense Representations from a Large Semantically Annotated Corpus with Long Short-Term Memories · ACL (1) 2019 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.2 | 1 | 2024 | MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation · CVPR 2024 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.2 | 2 | 2019 | LSTMEmbed: Learning Word and Sense Representations from a Large Semantically Annotated Corpus with Long Short-Term Memories · ACL (1) 2019 Embeddings for Word Sense Disambiguation: An Evaluation Study · ACL (1) 2016 |
Natural language and speech › Language models and text generation
natural language understanding |
0.2 | 1 | 2022 | Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLU · EMNLP 2022 |
Machine learning › Deep learning architectures and training
data augmentation |
0.1 | 1 | 2021 | Improving Commonsense Causal Reasoning by Adversarial Training and Data Augmentation · AAAI 2021 |
Methods — techniques the papers use, named apart from their topics
instance completion · 1.5image re-assembly · 1.5image decomposition · 1.5multilingual evaluation · 0.7training dynamics · 0.6difficulty metrics · 0.6synonym substitution · 0.5generative language model · 0.5discourse parser · 0.5adversarial training · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image GenerationabstractText-to-image generation has achieved astonishing results, yet precise spatial controllability and prompt fidelity remain highly challenging. This limitation is typically addressed through cumbersome prompt engineering, scene layout conditioning, or image editing techniques which often require hand drawn masks. Nonetheless, pre-existing works struggle to take advantage of the natural instance-level compositionality of scenes due to the typically flat nature of rasterized RGB output images. Towards adressing this challenge, we introduce MuLAn: a novel dataset comprising over 44K MUlti-Layer ANnotations of RGB images as multi-layer, instance-wise RGBA decompositions, and over 100K instance images. To build MuLAn, we developed a training free pipeline which decomposes a monocular RGB image into a stack of RGBA layers comprising of background and isolated instances. We achieve this through the use of pre-trained general-purpose models, and by developing three modules: image decomposition for instance discovery and extraction, instance completion to reconstruct occluded areas, and image re-assembly. We use our pipeline to create MuLAn-COCO and MuLAn-LAION datasets, which contain a variety of image decompositions in terms of style, composition and complexity. With MuLAn, we provide the first photorealistic resource providing instance decompo-sition and occlusion information for high quality images, opening up new avenues for text-to-image generative AI re-search. With this, we aim to encourage the development of novel generation and editing technology, in particular layer-wise solutions. MuLAn data resources are available at https://MuLAn-dataset.github.io/. Petru-Daniel Tudosiu, Yongxin Yang, Steven McDonagh 0001, Gerasimos Lampouras, Ignacio Iacobacci, Sarah Parisot |
CVPR | 7 |
| 2024 | Correct and Optimal: The Regular Expression Inference Challenge
Mojtaba Valizadeh, Philip John Gorinski, Ignacio Iacobacci, Martin Berger 0001 |
IJCAI | 3 |
| 2024 | HumanRankEval: Automatic Evaluation of LMs as Conversational AssistantsabstractMilan Gritta, Gerasimos Lampouras, Ignacio Iacobacci. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Milan Gritta, Gerasimos Lampouras, Ignacio Iacobacci |
NAACL-HLT | 3 |
| 2023 | A Systematic Study of Performance Disparities in Multilingual Task-Oriented Dialogue SystemsabstractAchieving robust language technologies that can perform well across the world's many languages is a central goal of multilingual NLP.In this work, we take stock of and empirically analyse task performance disparities that exist between multilingual task-oriented dialogue (TOD) systems.We first define new quantitative measures of absolute and relative equivalence in system performance, capturing disparities across languages and within individual languages.Through a series of controlled experiments, we demonstrate that performance disparities depend on a number of factors: the nature of the TOD task at hand, the underlying pretrained language model, the target language, and the amount of TOD annotated data.We empirically prove the existence of the adaptation bias and intrinsic biases in current TOD systems: e.g., TOD systems trained for Arabic or Turkish using annotated TOD data fully parallel to English TOD data still exhibit diminished TOD task performance.Beyond providing a series of insights into the performance disparities of TOD systems in different languages, our analyses offer practical tips on how to approach TOD data collection and system development for new languages. Songbo Hu, Han Zhou 0010, Moy Yuan, Milan Gritta, Guchun Zhang, Ignacio Iacobacci, Anna Korhonen, Ivan Vulic |
EMNLP | 6 |
| 2023 | Multi 3 WOZ: A Multilingual, Multi-Domain, Multi-Parallel Dataset for Training and Evaluating Culturally Adapted Task-Oriented Dialog SystemsabstractAbstract Creating high-quality annotated data for task-oriented dialog (ToD) is known to be notoriously difficult, and the challenges are amplified when the goal is to create equitable, culturally adapted, and large-scale ToD datasets for multiple languages. Therefore, the current datasets are still very scarce and suffer from limitations such as translation-based non-native dialogs with translation artefacts, small scale, or lack of cultural adaptation, among others. In this work, we first take stock of the current landscape of multilingual ToD datasets, offering a systematic overview of their properties and limitations. Aiming to reduce all the detected limitations, we then introduce Multi3WOZ, a novel multilingual, multi-domain, multi-parallel ToD dataset. It is large-scale and offers culturally adapted dialogs in 4 languages to enable training and evaluation of multilingual and cross-lingual ToD systems. We describe a complex bottom–up data collection process that yielded the final dataset, and offer the first sets of baseline scores across different ToD-related tasks for future reference, also highlighting its challenging nature. Songbo Hu, Han Zhou 0010, Mete Hergul, Milan Gritta, Guchun Zhang, Ignacio Iacobacci, Ivan Vulic, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 6 |
| 2022 | Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLUabstractCurriculum Learning (CL) is a technique of training models via ranking examples in a typically increasing difficulty trend with the aim of accelerating convergence and improving generalisability.Current approaches for Natural Language Understanding (NLU) tasks use CL to improve in-distribution data performance often via heuristic-oriented or task-agnostic difficulties.In this work, instead, we employ CL for NLU by taking advantage of training dynamics as difficulty metrics, i.e. statistics that measure the behavior of the model at hand on specific task-data instances during training and propose modifications of existing CL schedulers based on these statistics.Differently from existing works, we focus on evaluating models on in-distribution (ID), out-of-distribution (OOD) as well as zero-shot (ZS) cross-lingual transfer datasets.We show across several NLU tasks that CL with training dynamics can result in better performance mostly on zero-shot cross-lingual transfer and OOD settings with improvements up by 8.5% in certain cases.Overall, experiments indicate that training dynamics can lead to better performing models with smoother training compared to other difficulty metrics while being 20% faster on average.In addition, through analysis we shed light on the correlations of task-specific versus task-agnostic metrics 1 . Fenia Christopoulou, Gerasimos Lampouras, Ignacio Iacobacci |
EMNLP | 3 |
| 2021 | Improving Commonsense Causal Reasoning by Adversarial Training and Data AugmentationabstractDetermining the plausibility of causal relations between clauses is a commonsense reasoning task that requires complex inference ability. The general approach to this task is to train a large pretrained language model on a specific dataset. However, the available training data for the task is often scarce, which leads to instability of model training or reliance on the shallow features of the dataset. This paper presents a number of techniques for making models more robust in the domain of causal reasoning. Firstly, we perform adversarial training by generating perturbed inputs through synonym substitution. Secondly, based on a linguistic theory of discourse connectives, we perform data augmentation using a discourse parser for detecting causally linked clauses in large text, and a generative language model for generating distractors. Both methods boost model performance on the Choice of Plausible Alternatives (COPA) dataset, as well as on a Balanced COPA dataset, which is a modified version of the original data that has been developed to avoid superficial cues, leading to a more challenging benchmark. We show a statistically significant improvement in performance and robustness on both datasets, even with only a small number of additionally generated data points. Ieva Staliunaite, Philip John Gorinski, Ignacio Iacobacci |
AAAI | 3 |
| 2021 | Conversation Graph: Data Augmentation, Training and Evaluation for Non-Deterministic Dialogue ManagementabstractTask-oriented dialogue systems typically rely on large amounts of high-quality training data or require complex handcrafted rules. However, existing datasets are often limited in size con- sidering the complexity of the dialogues. Additionally, conventional training signal in- ference is not suitable for non-deterministic agent behavior, namely, considering multiple actions as valid in identical dialogue states. We propose the Conversation Graph (ConvGraph), a graph-based representation of dialogues that can be exploited for data augmentation, multi- reference training and evaluation of non- deterministic agents. ConvGraph generates novel dialogue paths to augment data volume and diversity. Intrinsic and extrinsic evaluation across three datasets shows that data augmentation and/or multi-reference training with ConvGraph can improve dialogue success rates by up to 6.4%. Milan Gritta, Gerasimos Lampouras, Ignacio Iacobacci |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQAabstractMany NLP tasks have benefited from transferring knowledge from contextualized word embeddings, however the picture of what type of knowledge is transferred is incomplete.This paper studies the types of linguistic phenomena accounted for by language models in the context of a Conversational Question Answering (CoQA) task.We identify the problematic areas for the finetuned RoBERTa, BERT and DistilBERT models through systematic error analysis -basic arithmetic (counting phrases), compositional semantics (negation and Semantic Role Labeling), and lexical semantics (surprisal and antonymy).When enhanced with the relevant linguistic knowledge through multitask learning, the models improve in performance.Ensembles of the enhanced models yield a boost between 2.2 and 2.7 points in F1 score overall, and up to 42.1 points in F1 on the hardest question classes.The results show differences in ability to represent compositional and lexical information between RoBERTa, BERT and DistilBERT. Ieva Staliunaite, Ignacio Iacobacci |
EMNLP (1) | 2 |
| 2020 | Auxiliary Capsules for Natural Language UnderstandingabstractLately, joint training of Intent detection and Slot filling has become the best-performing approach in the field of Natural Language Understanding (NLU). In this work we extend the newly introduced application of Capsule Networks for NLU to a multi-task learning environment, using relevant auxiliary tasks. Specifically, our models perform joint Intent classification and Slot filling with the aid of Named Entity Recognition (NER) and Part of Speech (POS) tagging tasks. This allows us to exploit the hierarchical relationships between the Intents of the utterances and the different features of input text, not only Slots but also Named Entity mentions, Parts of Speech, quantity indications, etc. The models developed in this work are evaluated on standard benchmarks, achieving state-of-the-art results on the SNIPS dataset while outperforming the best commercial systems on several low-resource datasets. Ieva Staliunaite, Ignacio Iacobacci |
ICASSP | 2 |
| 2019 | LSTMEmbed: Learning Word and Sense Representations from a Large Semantically Annotated Corpus with Long Short-Term MemoriesabstractWhile word embeddings are now a de facto standard representation of words in most NLP tasks, recently the attention has been shifting towards vector representations which capture the different meanings, i.e., senses, of words.In this paper we explore the capabilities of a bidirectional LSTM model to learn representations of word senses from semantically annotated corpora.We show that the utilization of an architecture that is aware of word order, like an LSTM, enables us to create better representations.We assess our proposed model on various standard benchmarks for evaluating semantic representations, reaching state-of-the-art performance on the SemEval-2014 word-to-sense similarity task.We release the code and the resulting word and sense embeddings at http://lcl.uniroma1. it/LSTMEmbed. Ignacio Iacobacci, Roberto Navigli |
ACL (1) | 1 |
| 2017 | Embedding Words and Senses Together via Joint Knowledge-Enhanced TrainingabstractWord embeddings are widely used in Natural Language Processing, mainly due to their success in capturing semantic information from massive corpora.However, their creation process does not allow the different meanings of a word to be automatically separated, as it conflates them into a single vector.We address this issue by proposing a new model which learns word and sense embeddings jointly.Our model exploits large corpora and knowledge from semantic networks in order to produce a unified vector space of word and sense embeddings.We evaluate the main features of our approach both qualitatively and quantitatively in a variety of tasks, highlighting the advantages of the proposed method in comparison to stateof-the-art word-and sense-based models. Massimiliano Mancini, José Camacho-Collados, Ignacio Iacobacci, Roberto Navigli |
CoNLL | 3 |
| 2016 | Embeddings for Word Sense Disambiguation: An Evaluation StudyabstractRecent years have seen a dramatic growth in the popularity of word embeddings mainly owing to their ability to capture semantic information from massive amounts of textual content.As a result, many tasks in Natural Language Processing have tried to take advantage of the potential of these distributional models.In this work, we study how word embeddings can be used in Word Sense Disambiguation, one of the oldest tasks in Natural Language Processing and Artificial Intelligence.We propose different methods through which word embeddings can be leveraged in a state-of-the-art supervised WSD system architecture, and perform a deep analysis of how different parameters affect performance.We show how a WSD system that makes use of word embeddings alone, if designed properly, can provide significant performance improvement over a state-ofthe-art WSD system that incorporates several standard WSD features. Ignacio Iacobacci, Mohammad Taher Pilehvar, Roberto Navigli |
ACL (1) | 1 |
| 2015 | SensEmbed: Learning Sense Embeddings for Word and Relational SimilarityabstractIgnacio Iacobacci, Mohammad Taher Pilehvar, Roberto Navigli. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Ignacio Iacobacci, Mohammad Taher Pilehvar, Roberto Navigli |
ACL (1) | 1 |