EDBT 2026 Demo / reviewers in the wild / expert
Federico Bianchi 0001
dblp:122/8815-1
· DBLP profile ↗
24ranked-venue papers
10as first author
17since 2021 · last 2025
0000-0003-0776-361XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 8 first-author · 17 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | h4rm3l: A Language for Composable Jailbreak Attack SynthesisabstractDespite their demonstrated valuable capabilities, state-of-the-art (SOTA) widely deployed large language models (LLMs) still have the potential to cause harm to society due to the ineffectiveness of their safety filters, which can be bypassed by prompt transformations called jailbreak attacks. Current approaches to LLM safety assessment, which employ datasets of templated prompts and benchmarking pipelines, fail to cover sufficiently large and diverse sets of jailbreak attacks, leading to the widespread deployment of unsafe LLMs. Recent research showed that novel jailbreak attacks could be derived by composition; however, a formal composable representation for jailbreak attacks, which, among other benefits, could enable the exploration of a large compositional space of jailbreak attacks through program synthesis methods, has not been previously proposed. We introduce h4rm3l, a novel approach that addresses this gap with a human-readable domain-specific language (DSL). Our framework comprises: (1) The h4rm3l DSL, which formally expresses jailbreak attacks as compositions of parameterized string transformation primitives. (2) A synthesizer with bandit algorithms that efficiently generates jailbreak attacks optimized for a target black box LLM. (3) The h4rm3l red-teaming software toolkit that employs the previous two components and an automated harmful LLM behavior classifier that is strongly aligned with human judgment. We demonstrate h4rm3l's efficacy by synthesizing a dataset of 2656 successful novel jailbreak attacks targeting 6 SOTA open-source and proprietary LLMs (GPT-3.5, GPT-4o, Claude-3-Sonnet, Claude-3-Haiku, Llama-3-8B, and Llama-3-70B), and by benchmarking those models against a subset of these synthesized attacks. Our results show that h4rm3l's synthesized attacks are diverse and more successful than existing jailbreak attacks in literature, with success rates exceeding 90% on SOTA LLMs. Warning: This paper and related research artifacts contain offensive and potentially disturbing prompts and model-generated content. Moussa Doumbouya, Ananjan Nandi, Gabriel Poesia, Davide Ghilardi, Anna Goldie, Federico Bianchi 0001, Daniel Jurafsky, Christopher D. Manning |
ICLR | 6 |
| 2024 | Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsabstractTraining large language models to follow instructions makes them perform better on a wide range of tasks and generally become more helpful. However, a perfectly helpful model will follow even the most malicious instructions and readily generate harmful content.
In this paper, we raise concerns over the safety of models that only emphasize helpfulness, not harmlessness, in their instruction-tuning.
We show that several popular instruction-tuned models are highly unsafe. Moreover, we show that adding just 3\% safety examples (a few hundred demonstrations) when fine-tuning a model like LLaMA can substantially improve its safety. Our safety-tuning does not make models significantly less capable or helpful as measured by standard benchmarks. However, we do find exaggerated safety behaviours, where too much safety-tuning makes models refuse perfectly safe prompts if they superficially resemble unsafe ones. As a whole, our results illustrate trade-offs in training LLMs to be helpful and training them to be safe. Federico Bianchi 0001, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Daniel Jurafsky, Tatsunori B. Hashimoto, James Zou 0001 |
ICLR | 1 |
| 2024 | How Well Can LLMs Negotiate? NegotiationArena Platform and AnalysisabstractNegotiation is the basis of social interactions; humans negotiate everything from the price of cars to how to share common resources. With rapidly growing interest in using large language models (LLMs) to act as agents on behalf of human users, such LLM agents would also need to be able to negotiate. In this paper, we study how well LLMs can negotiate with each other. We develop NegotiationArena: a flexible framework for evaluating and probing the negotiation abilities of LLM agents. We implemented three types of scenarios in NegotiationArena to assess LLM's behaviors in allocating shared resources (ultimatum games), aggregate resources (trading games) and buy/sell goods (price negotiations). Each scenario allows for multiple turns of flexible dialogues between LLM agents to allow for more complex negotiations. Interestingly, LLM agents can significantly boost their negotiation outcomes by employing certain behavioral tactics. For example, by pretending to be desolate and desperate, LLMs can improve their payoffs by 20% when negotiating against the standard GPT-4. We also quantify irrational negotiation behaviors exhibited by the LLM agents, many of which also appear in humans. Together, NegotiationArena offers a new environment to investigate LLM interactions, enabling new insights into LLM's theory of mind, irrationality, and reasoning abilities Federico Bianchi 0001, Patrick John Chia, Mert Yüksekgönül, Jacopo Tagliabue, Daniel Jurafsky, James Zou 0001 |
ICML | 1 |
| 2024 | XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language ModelsabstractPaul Röttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, Dirk Hovy. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Paul Röttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi 0001, Dirk Hovy |
NAACL-HLT | 5 |
| 2023 | When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?
Mert Yüksekgönül, Federico Bianchi 0001, Pratyusha Kalluri, Daniel Jurafsky, James Zou 0001 |
ICLR | 2 |
| 2023 | EvalRS 2023: Well-Rounded Recommender Systems for Real-World DeploymentsabstractEvalRS aims to bring together practitioners from industry and academia to foster a debate on rounded evaluation of recommender systems, with a focus on real-world impact across a multitude of deployment scenarios. Recommender systems are often evaluated only through accuracy metrics, which fall short of fully characterizing their generalization capabilities and miss important aspects, such as fairness, bias, usefulness, informativeness. This workshop builds on the success of last year's workshop at CIKM, but with a broader scope and an interactive format. Federico Bianchi 0001, Patrick John Chia, Jacopo Tagliabue, Ciro Greco, Gabriel de Souza P. Moreira, Davide Eynard, Fahd Husain, Claudio Pomo |
KDD | 1 |
| 2023 | Beyond Digital "Echo Chambers": The Role of Viewpoint Diversity in Political DiscussionabstractIncreasingly taking place in online spaces, modern political conversations are typically perceived to be unproductively affirming---siloed in so called "echo chambers" of exclusively like-minded discussants. Yet, to date we lack sufficient means to measure viewpoint diversity in conversations. To this end, in this paper, we operationalize two viewpoint metrics proposed for recommender systems and adapt them to the context of social media conversations. This is the first study to apply these two metrics (Representation and Fragmentation) to real world data and to consider the implications for online conversations specifically. We apply these measures to two topics---daylight savings time (DST), which serves as a control, and the more politically polarized topic of immigration. We find that the diversity scores for both Fragmentation and Representation are lower for immigration than for DST. Further, we find that while pro-immigrant views receive consistent pushback on the platform, anti-immigrant views largely operate within echo chambers. We observe less severe yet similar patterns for DST. Taken together, Representation and Fragmentation paint a meaningful and important new picture of viewpoint diversity. Rishav Hada, Amir Ebrahimi Fard, Sarah Shugars, Federico Bianchi 0001, Patrícia G. C. Rossini, Dirk Hovy, Rebekah Tromble, Nava Tintarev |
WSDM | 4 |
| 2023 | Viewpoint: Artificial Intelligence Accidents Waiting to Happen?abstractArtificial Intelligence (AI) is at a crucial point in its development: stable enough to be used in production systems, and increasingly pervasive in our lives. What does that mean for its safety? In his book Normal Accidents, the sociologist Charles Perrow proposed a framework to analyze new technologies and the risks they entail. He showed that major accidents are nearly unavoidable in complex systems with tightly coupled components if they are run long enough. In this essay, we apply and extend Perrow’s framework to AI to assess its potential risks. Today’s AI systems are already highly complex, and their complexity is steadily increasing. As they become more ubiquitous, different algorithms will interact directly, leading to tightly coupled systems whose capacity to cause harm we will be unable to predict. We argue that under the current paradigm, Perrow’s normal accidents apply to AI systems and it is only a matter of time before one occurs. This article appears in the AI & Society track. Federico Bianchi 0001, Amanda Cercas Curry, Dirk Hovy |
J. Artif. Intell. Res. | 1 |
| 2022 | "It's Not Just Hate": A Multi-Dimensional Perspective on Detecting Harmful Speech OnlineabstractWell-annotated data is a prerequisite for good Natural Language Processing models.Too often, though, annotation decisions are governed by optimizing time or annotator agreement.We make a case for nuanced efforts in an interdisciplinary setting for annotating offensive online speech.Detecting offensive content is rapidly becoming one of the most important real-world NLP tasks.However, most datasets use a single binary label, e.g., for hate or incivility, even though each concept is multi-faceted.This modeling choice severely limits nuanced insights, but also performance.We show that a more fine-grained multi-label approach to predicting incivility and hateful or intolerant content addresses both conceptual and performance issues.We release a novel dataset of over 40,000 tweets about immigration from the US and UK, annotated with six labels for different aspects of incivility and intolerance.Our dataset not only allows for a more nuanced understanding of harmful speech online, models trained on it also outperform or match performance on benchmark datasets.Warning: This paper contains examples of hateful language some readers might find offensive. Federico Bianchi 0001, Stefanie Anja Hills, Patrícia G. C. Rossini, Dirk Hovy, Rebekah Tromble, Nava Tintarev |
EMNLP | 1 |
| 2022 | SocioProbe: What, When, and Where Language Models Learn about SociodemographicsabstractPre-trained language models (PLMs) have outperformed other NLP models on a wide range of tasks.Opting for a more thorough understanding of their capabilities and inner workings, researchers have established the extend to which they capture lower-level knowledge like grammaticality, and mid-level semantic knowledge like factual understanding.However, there is still little understanding of their knowledge of higher-level aspects of language.In particular, despite the importance of sociodemographic aspects in shaping our language, the questions of whether, where, and how PLMs encode these aspects, e.g., gender or age, is still unexplored.We address this research gap by probing the sociodemographic knowledge of different single-GPU PLMs on multiple English data sets via traditional classifier probing and information-theoretic minimum description length probing.Our results show that PLMs do encode these sociodemographics, and that this knowledge is sometimes spread across the layers of some of the tested PLMs.We further conduct a multilingual analysis and investigate the effect of supplementary training to further explore to what extent, where, and with what amount of pre-training data the knowledge is encoded.Our overall results indicate that sociodemographic knowledge is still a major challenge for NLP.PLMs require large amounts of pre-training data to acquire the knowledge and models that excel in general language understanding do not seem to own more knowledge about these aspects. Anne Lauscher, Federico Bianchi 0001, Samuel R. Bowman, Dirk Hovy |
EMNLP | 2 |
| 2022 | Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced LanguagesabstractHate speech is a global phenomenon, but most hate speech datasets so far focus on Englishlanguage content.This hinders the development of more effective hate speech detection models in hundreds of languages spoken by billions across the world.More data is needed, but annotating hateful content is expensive, timeconsuming and potentially harmful to annotators.To mitigate these issues, we explore dataefficient strategies for expanding hate speech detection into under-resourced languages.In a series of experiments with mono-and multilingual models across five non-English languages, we find that 1) a small amount of target-language fine-tuning data is needed to achieve strong performance, 2) the benefits of using more such data decrease exponentially, and 3) initial fine-tuning on readily-available English data can partially substitute targetlanguage data and improve model generalisability.Based on these findings, we formulate actionable recommendations for hate speech detection in low-resource language settings. Paul Röttger, Debora Nozza, Federico Bianchi 0001, Dirk Hovy |
EMNLP | 3 |
| 2021 | Cross-lingual Contextualized Topic Models with Zero-shot LearningabstractFederico Bianchi, Silvia Terragni, Dirk Hovy, Debora Nozza, Elisabetta Fersini. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Federico Bianchi 0001, Silvia Terragni, Dirk Hovy, Debora Nozza, Elisabetta Fersini |
EACL | 1 |
| 2021 | BERTective: Language Models and Contextual Information for Deception DetectionabstractSpotting a lie is challenging but has an enormous potential impact on security as well as private and public safety.Several NLP methods have been proposed to classify texts as truthful or deceptive.In most cases, however, the target texts' preceding context is not considered.This is a severe limitation, as any communication takes place in context, not in a vacuum, and context can help to detect deception.We study a corpus of Italian dialogues containing deceptive statements and implement deep neural models that incorporate various linguistic contexts.We establish a new state-of-theart identifying deception and find that not all context is equally useful to the task.Only the texts closest to the target, if from the same speaker (rather than questions by an interlocutor), boost performance.We also find that the semantic information in language models such as BERT contributes to the performance.However, BERT alone does not capture the implicit knowledge of deception cues: its contribution is conditional on the concurrent use of attention to learn cues from BERT's representations. Tommaso Fornaciari, Federico Bianchi 0001, Massimo Poesio, Dirk Hovy |
EACL | 2 |
| 2021 | SWEAT: Scoring Polarization of Topics across Different CorporaabstractUnderstanding differences of viewpoints across corpora is a fundamental task for computational social sciences.In this paper, we propose the Sliced Word Embedding Association Test (SWEAT), a novel statistical measure to compute the relative polarization of a topical wordset across two distributional representations.To this end, SWEAT uses two additional wordsets, deemed to have opposite valence, to represent two different poles.We validate our approach and illustrate a case study to show the usefulness of the introduced measure. Federico Bianchi 0001, Marco Marelli, Paolo Nicoli, Matteo Palmonari |
EMNLP (1) | 1 |
| 2021 | Language in a (Search) Box: Grounding Language Learning in Real-World Human-Machine InteractionabstractWe investigate grounded language learning through real-world data, by modelling a teacher-learner dynamics through the natural interactions occurring between users and search engines; in particular, we explore the emergence of semantic generalization from unsupervised dense representations outside of synthetic environments.A grounding domain, a denotation function and a composition function are learned from user data only.We show how the resulting semantics for noun phrases exhibits compositional properties while being fully learnable without any explicit labelling.We benchmark our grounded semantics on compositionality and zero-shot inference tasks, and we show that it provides better results and better generalizations than SOTA non-grounded models, such as word2vec and BERT. Federico Bianchi 0001, Ciro Greco, Jacopo Tagliabue |
NAACL-HLT | 1 |
| 2021 | HONEST: Measuring Hurtful Sentence Completion in Language ModelsabstractLanguage models have revolutionized the field of NLP.However, language models capture and proliferate hurtful stereotypes, especially in text generation.Our results show that 4.3% of the time, language models complete a sentence with a hurtful word.These cases are not random, but follow language and genderspecific patterns.We propose a score to measure hurtful sentence completions in language models (HONEST).It uses a systematic template-and lexicon-based bias evaluation methodology for six languages.Our findings suggest that these models replicate and amplify deep-seated societal stereotypes about gender roles.Sentence completions refer to sexual promiscuity when the target is female in 9% of the time, and in 4% to homosexuality when the target is male.The results raise questions about the use of these models in production settings. Debora Nozza, Federico Bianchi 0001, Dirk Hovy |
NAACL-HLT | 2 |
| 2021 | Towards bridging the neuro-symbolic gap: deep deductive reasoners
Monireh Ebrahimi, Aaron Eberhart, Federico Bianchi 0001, Pascal Hitzler |
Appl. Intell. | 3 |
| 2020 | "You Sound Just Like Your Father" Commercial Machine Translation Systems Include Stylistic BiasesabstractThe main goal of machine translation has been to convey the correct content. Stylistic considerations have been at best secondary. We show that as a consequence, the output of three commercial machine translation systems (Bing, DeepL, Google) make demographically diverse samples from five languages "sound" older and more male than the original. Our findings suggest that translation models reflect demographic bias in the training data. This opens up interesting new research avenues in machine translation to take stylistic considerations into account. Dirk Hovy, Federico Bianchi 0001, Tommaso Fornaciari |
ACL | 2 |
| 2020 | The Embeddings That Came in From the Cold: Improving Vectors for New and Rare Products with Content-Based InferenceabstractTraining product embeddings in a multi-tenant scenario involves solving the challenges of ever changing catalogs across dozens of deployments, without supervision. In this work, we detail how we deal with new and rare products when building neural representations at scale: we show how to inject product knowledge into behavior-based embeddings to provide the best accuracy with minimal engineering changes in existing infrastructure and without additional manual effort. Jacopo Tagliabue, Bingqing Yu, Federico Bianchi 0001 |
RecSys | 3 |
| 2020 | Tough Tables: Carefully Evaluating Entity Linking for Tabular DataabstractTable annotation is a key task to improve querying the Web and support the Knowledge Graph population from legacy sources (tables). Last year, the SemTab challenge was introduced to unify different efforts to evaluate table annotation algorithms by providing a common interface and several general-purpose datasets as a ground truth. The SemTab dataset is useful to have a general understanding of how these algorithms work, and the organizers of the challenge included some artificial noise to the data to make the annotation trickier. However, it is hard to analyze specific aspects in an automatic way. For example, the ambiguity of names at the entity-level can largely affect the quality of the annotation. In this paper, we propose a novel dataset to complement the datasets proposed by SemTab. The dataset consists of a set of high-quality manually-curated tables with non-obviously linkable cells, i.e., where values are ambiguous names, typos, and misspelled entity names not appearing in the current version of the SemTab dataset. These challenges are particularly relevant for the ingestion of structured legacy sources into existing knowledge graphs. Evaluations run on this dataset show that ambiguity is a key problem for entity linking algorithms and encourage a promising direction for future work in the field. Vincenzo Cutrona, Federico Bianchi 0001, Ernesto Jiménez-Ruiz, Matteo Palmonari |
ISWC (2) | 2 |
| 2019 | Training Temporal Word Embeddings with a CompassabstractTemporal word embeddings have been proposed to support the analysis of word meaning shifts during time and to study the evolution of languages. Different approaches have been proposed to generate vector representations of words that embed their meaning during a specific time interval. However, the training process used in these approaches is complex, may be inefficient or it may require large text corpora. As a consequence, these approaches may be difficult to apply in resource-scarce domains or by scientists with limited in-depth knowledge of embedding models. In this paper, we propose a new heuristic to train temporal word embeddings based on the Word2vec model. The heuristic consists in using atemporal vectors as a reference, i.e., as a compass, when training the representations specific to a given time interval. The use of the compass simplifies the training process and makes it more efficient. Experiments conducted using state-of-the-art datasets and methodologies suggest that our approach outperforms or equals comparable approaches while being more robust in terms of the required corpus size. Valerio Di Carlo, Federico Bianchi 0001, Matteo Palmonari |
AAAI | 2 |
| 2019 | On the composition and recommendation of multi-feature paths: a comprehensive approach
Vincenzo Cutrona, Federico Bianchi 0001, Michele Ciavotta, Andrea Maurino |
GeoInformatica | 2 |
| 2018 | Towards Encoding Time in Text-Based Entity Embeddings
Federico Bianchi 0001, Matteo Palmonari, Debora Nozza |
ISWC (1) | 1 |
| 2017 | Actively Learning to Rank Semantic Associations for Personalized Contextual Exploration of Knowledge Graphs
Federico Bianchi 0001, Matteo Palmonari, Marco Cremaschi, Elisabetta Fersini |
ESWC (1) | 1 |