Jacopo Tagliabue

dblp:78/11468 · DBLP profile ↗
← Back
8ranked-venue papers in the field
4as first author
6since 2021 · last 2025
0000-0001-8634-6122ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 4 (2 first)Big Data, Cloud & Distributed Data Systems · 2 (2 first)Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2025 Safe, Untrusted, "Proof-Carrying" AI Agents: Toward the Agentic Lakehouse
Jacopo Tagliabue, Ciro Greco
IEEE Big Data1
2024 FaaS and Furious: abstractions and differential caching for efficient data pre-processing
abstract
Data pre-processing pipelines are the bread and butter of any successful AI project. We introduce a novel programming model for pipelines in a data lakehouse, allowing users to interact declaratively with assets in object storage. Motivated by real-world industry usage patterns, we exploit these new abstractions with a columnar and differential cache to maximize iteration speed for data scientists, who spent most of their time in pre-processing – adding or removing features, restricting or relaxing time windows, wrangling current or older datasets. We show how the new cache works transparently across programming languages, schemas and time windows, and provide preliminary evidence on its efficiency on standard data workloads.
Jacopo Tagliabue, Ryan Curtin, Ciro Greco
IEEE Big Data1
2023 EvalRS 2023: Well-Rounded Recommender Systems for Real-World Deployments
abstract
EvalRS aims to bring together practitioners from industry and academia to foster a debate on rounded evaluation of recommender systems, with a focus on real-world impact across a multitude of deployment scenarios. Recommender systems are often evaluated only through accuracy metrics, which fall short of fully characterizing their generalization capabilities and miss important aspects, such as fairness, bias, usefulness, informativeness. This workshop builds on the success of last year's workshop at CIKM, but with a broader scope and an interactive format.
Federico Bianchi 0001, Patrick John Chia, Jacopo Tagliabue, Ciro Greco, Gabriel de Souza P. Moreira, Davide Eynard, Fahd Husain, Claudio Pomo
KDD3
2023 eCom'23: The SIGIR 2023 Workshop on eCommerce
abstract
eCommerce Information Retrieval (IR) is receiving increasing attention in the academic literature and is an essential component of some of the largest web sites (e.g. Airbnb, Alibaba, Amazon, eBay, Facebook, Flipkart, Lowes's, Taobao, Target). SIGIR has for several years seen sponsorship from eCommerce organizations, reflecting the importance of IR research to them. The purpose of this workshop is (1) to bring together researchers and practitioners of eCommerce IR to discuss topics unique to it, (2) to determine how to use eCommerce's unique combination of free text, structured data, and customer behavior data to improve search relevance, and (3) to examine how to build datasets and evaluate algorithms in this domain.
Surya Kallumadi, Yubin Kim 0001, Tracy Holloway King, Shervin Malmasi, Maarten de Rijke, Jacopo Tagliabue
SIGIR6
2022 eCom'22: The SIGIR 2022 Workshop on eCommerce
abstract
eCommerce Information Retrieval (IR) is receiving increasing attention in the academic literature and is an essential component of some of the world's largest web sites (e.g. Airbnb, Alibaba, Amazon, eBay, Facebook, Flipkart, Lowe's, Taobao, and Target). SIGIR has for several years seen sponsorship from eCommerce organisations, reflecting the importance of IR research to them. The purpose of this workshop is (1) to bring together researchers and practitioners of eCommerce IR to discuss topics unique to it, (2) to determine how to use eCommerce's unique combination of free text, structured data, and customer behavioral data to improve search relevance, and (3) to examine how to build datasets and evaluate algorithms in this domain. Since eCommerce customers often do not know exactly what they want to buy (i.e. navigational and spearfishing queries are rare), recommendations are valuable for inspiration and serendipitous discovery as well as basket building.
Ajinkya Kale, Surya Kallumadi, Tracy Holloway King, Shervin Malmasi, Maarten de Rijke, Jacopo Tagliabue
SIGIR6
2021 You Do Not Need a Bigger Boat: Recommendations at Reasonable Scale in a (Mostly) Serverless and Open Stack
abstract
We argue that immature data pipelines are preventing a large portion of industry practitioners from leveraging the latest research on recommender systems. We propose our template data stack for machine learning at “reasonable scale”, and show how many challenges are solved by embracing a serverless paradigm. Leveraging our experience, we detail how modern open source tools can provide a pipeline processing terabytes of data with minimal infrastructure work.
Jacopo Tagliabue
RecSys1
2020 The Embeddings That Came in From the Cold: Improving Vectors for New and Rare Products with Content-Based Inference
abstract
Training product embeddings in a multi-tenant scenario involves solving the challenges of ever changing catalogs across dozens of deployments, without supervision. In this work, we detail how we deal with new and rare products when building neural representations at scale: we show how to inject product knowledge into behavior-based embeddings to provide the best accuracy with minimal engineering changes in existing infrastructure and without additional manual effort.
Jacopo Tagliabue, Bingqing Yu, Federico Bianchi 0001
RecSys1
2019 Less (Data) Is More: Why Small Data Holds the Key to the Future of Artificial Intelligence
abstract
The claims that big data holds the key to enterprise successes and that Artificial Intelligence is going to replace humanity have become increasingly more popular over the past few years, both in academia and in the industry. However, while these claims may indeed capture some truth, they have also been massively oversold, or so we contend here. The goal of this paper is two-fold. First, we provide a qualified defence of the value of less data within the context of AI. This is done by carefully reviewing two distinct problems for big data driven AI, namely a) the limited track record of Deep Learning in key areas such as Natural Language Processing, b) the regulatory and business significance of being able to learn from few data points. Second, we briefly sketch what we refer to as a case of AI with humans and for humans, namely an AI paradigm whereby the systems we build are privacy-oriented and focused on human-machine collaboration, not competition. Combining our claims above, we conclude that when seen through the lens of cognitively inspired AI, the bright future of the discipline is about less data, not more, and more humans, not fewer.
Ciro Greco, Andrea Polonioli, Jacopo Tagliabue
DATA3