Jacopo Tagliabue

dblp:78/11468 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-8634-6122ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Safe, Untrusted, "Proof-Carrying" AI Agents: Toward the Agentic Lakehouse
Jacopo Tagliabue, Ciro Greco
IEEE Big Data1
2024 FaaS and Furious: abstractions and differential caching for efficient data pre-processing
abstract
Data pre-processing pipelines are the bread and butter of any successful AI project. We introduce a novel programming model for pipelines in a data lakehouse, allowing users to interact declaratively with assets in object storage. Motivated by real-world industry usage patterns, we exploit these new abstractions with a columnar and differential cache to maximize iteration speed for data scientists, who spent most of their time in pre-processing – adding or removing features, restricting or relaxing time windows, wrangling current or older datasets. We show how the new cache works transparently across programming languages, schemas and time windows, and provide preliminary evidence on its efficiency on standard data workloads.
Jacopo Tagliabue, Ryan Curtin, Ciro Greco
IEEE Big Data1
2024 How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
abstract
Negotiation is the basis of social interactions; humans negotiate everything from the price of cars to how to share common resources. With rapidly growing interest in using large language models (LLMs) to act as agents on behalf of human users, such LLM agents would also need to be able to negotiate. In this paper, we study how well LLMs can negotiate with each other. We develop NegotiationArena: a flexible framework for evaluating and probing the negotiation abilities of LLM agents. We implemented three types of scenarios in NegotiationArena to assess LLM's behaviors in allocating shared resources (ultimatum games), aggregate resources (trading games) and buy/sell goods (price negotiations). Each scenario allows for multiple turns of flexible dialogues between LLM agents to allow for more complex negotiations. Interestingly, LLM agents can significantly boost their negotiation outcomes by employing certain behavioral tactics. For example, by pretending to be desolate and desperate, LLMs can improve their payoffs by 20% when negotiating against the standard GPT-4. We also quantify irrational negotiation behaviors exhibited by the LLM agents, many of which also appear in humans. Together, NegotiationArena offers a new environment to investigate LLM interactions, enabling new insights into LLM's theory of mind, irrationality, and reasoning abilities
Federico Bianchi 0001, Patrick John Chia, Mert Yüksekgönül, Jacopo Tagliabue, Daniel Jurafsky, James Zou 0001
ICML4
2023 EvalRS 2023: Well-Rounded Recommender Systems for Real-World Deployments
abstract
EvalRS aims to bring together practitioners from industry and academia to foster a debate on rounded evaluation of recommender systems, with a focus on real-world impact across a multitude of deployment scenarios. Recommender systems are often evaluated only through accuracy metrics, which fall short of fully characterizing their generalization capabilities and miss important aspects, such as fairness, bias, usefulness, informativeness. This workshop builds on the success of last year's workshop at CIKM, but with a broader scope and an interactive format.
Federico Bianchi 0001, Patrick John Chia, Jacopo Tagliabue, Ciro Greco, Gabriel de Souza P. Moreira, Davide Eynard, Fahd Husain, Claudio Pomo
KDD3
2023 eCom'23: The SIGIR 2023 Workshop on eCommerce
abstract
eCommerce Information Retrieval (IR) is receiving increasing attention in the academic literature and is an essential component of some of the largest web sites (e.g. Airbnb, Alibaba, Amazon, eBay, Facebook, Flipkart, Lowes's, Taobao, Target). SIGIR has for several years seen sponsorship from eCommerce organizations, reflecting the importance of IR research to them. The purpose of this workshop is (1) to bring together researchers and practitioners of eCommerce IR to discuss topics unique to it, (2) to determine how to use eCommerce's unique combination of free text, structured data, and customer behavior data to improve search relevance, and (3) to examine how to build datasets and evaluate algorithms in this domain.
Surya Kallumadi, Yubin Kim 0001, Tracy Holloway King, Shervin Malmasi, Maarten de Rijke, Jacopo Tagliabue
SIGIR6
2022 eCom'22: The SIGIR 2022 Workshop on eCommerce
abstract
eCommerce Information Retrieval (IR) is receiving increasing attention in the academic literature and is an essential component of some of the world's largest web sites (e.g. Airbnb, Alibaba, Amazon, eBay, Facebook, Flipkart, Lowe's, Taobao, and Target). SIGIR has for several years seen sponsorship from eCommerce organisations, reflecting the importance of IR research to them. The purpose of this workshop is (1) to bring together researchers and practitioners of eCommerce IR to discuss topics unique to it, (2) to determine how to use eCommerce's unique combination of free text, structured data, and customer behavioral data to improve search relevance, and (3) to examine how to build datasets and evaluate algorithms in this domain. Since eCommerce customers often do not know exactly what they want to buy (i.e. navigational and spearfishing queries are rare), recommendations are valuable for inspiration and serendipitous discovery as well as basket building.
Ajinkya Kale, Surya Kallumadi, Tracy Holloway King, Shervin Malmasi, Maarten de Rijke, Jacopo Tagliabue
SIGIR6
2022 On the plurality of graphs
abstract
Abstract We investigate David Lewis’ signalling games dynamics in the context of neural agents as embedded in realistic graph structures. We follow in the footstep of both the deep learning tradition—in moving away from pure rule-based models—, and the symbolic literature—in leveraging discrete, top-down structures to constrain the learning process. Through a series of experiments in which we systematically vary the social graphs connecting the players, we are able to show for the first time that the dynamics of language emergence within a population of neural agents is strongly influenced by the underlying graph topology constraining their interactions.
Nicole Fitzgerald, Jacopo Tagliabue
J. Log. Comput.2
2021 Language in a (Search) Box: Grounding Language Learning in Real-World Human-Machine Interaction
abstract
We investigate grounded language learning through real-world data, by modelling a teacher-learner dynamics through the natural interactions occurring between users and search engines; in particular, we explore the emergence of semantic generalization from unsupervised dense representations outside of synthetic environments.A grounding domain, a denotation function and a composition function are learned from user data only.We show how the resulting semantics for noun phrases exhibits compositional properties while being fully learnable without any explicit labelling.We benchmark our grounded semantics on compositionality and zero-shot inference tasks, and we show that it provides better results and better generalizations than SOTA non-grounded models, such as word2vec and BERT.
Federico Bianchi 0001, Ciro Greco, Jacopo Tagliabue
NAACL-HLT3
2021 You Do Not Need a Bigger Boat: Recommendations at Reasonable Scale in a (Mostly) Serverless and Open Stack
abstract
We argue that immature data pipelines are preventing a large portion of industry practitioners from leveraging the latest research on recommender systems. We propose our template data stack for machine learning at “reasonable scale”, and show how many challenges are solved by embracing a serverless paradigm. Leveraging our experience, we detail how modern open source tools can provide a pipeline processing terabytes of data with minimal infrastructure work.
Jacopo Tagliabue
RecSys1
2020 The Embeddings That Came in From the Cold: Improving Vectors for New and Rare Products with Content-Based Inference
abstract
Training product embeddings in a multi-tenant scenario involves solving the challenges of ever changing catalogs across dozens of deployments, without supervision. In this work, we detail how we deal with new and rare products when building neural representations at scale: we show how to inject product knowledge into behavior-based embeddings to provide the best accuracy with minimal engineering changes in existing infrastructure and without additional manual effort.
Jacopo Tagliabue, Bingqing Yu, Federico Bianchi 0001
RecSys1
2019 Less (Data) Is More: Why Small Data Holds the Key to the Future of Artificial Intelligence
abstract
The claims that big data holds the key to enterprise successes and that Artificial Intelligence is going to replace humanity have become increasingly more popular over the past few years, both in academia and in the industry. However, while these claims may indeed capture some truth, they have also been massively oversold, or so we contend here. The goal of this paper is two-fold. First, we provide a qualified defence of the value of less data within the context of AI. This is done by carefully reviewing two distinct problems for big data driven AI, namely a) the limited track record of Deep Learning in key areas such as Natural Language Processing, b) the regulatory and business significance of being able to learn from few data points. Second, we briefly sketch what we refer to as a case of AI with humans and for humans, namely an AI paradigm whereby the systems we build are privacy-oriented and focused on human-machine collaboration, not competition. Combining our claims above, we conclude that when seen through the lens of cognitively inspired AI, the bright future of the discipline is about less data, not more, and more humans, not fewer.
Ciro Greco, Andrea Polonioli, Jacopo Tagliabue
DATA3