Ciro Greco

dblp:244/2252 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0007-0359-4130ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Safe, Untrusted, "Proof-Carrying" AI Agents: Toward the Agentic Lakehouse
Jacopo Tagliabue, Ciro Greco
IEEE Big Data2
2024 FaaS and Furious: abstractions and differential caching for efficient data pre-processing
abstract
Data pre-processing pipelines are the bread and butter of any successful AI project. We introduce a novel programming model for pipelines in a data lakehouse, allowing users to interact declaratively with assets in object storage. Motivated by real-world industry usage patterns, we exploit these new abstractions with a columnar and differential cache to maximize iteration speed for data scientists, who spent most of their time in pre-processing – adding or removing features, restricting or relaxing time windows, wrangling current or older datasets. We show how the new cache works transparently across programming languages, schemas and time windows, and provide preliminary evidence on its efficiency on standard data workloads.
Jacopo Tagliabue, Ryan Curtin, Ciro Greco
IEEE Big Data3
2023 EvalRS 2023: Well-Rounded Recommender Systems for Real-World Deployments
abstract
EvalRS aims to bring together practitioners from industry and academia to foster a debate on rounded evaluation of recommender systems, with a focus on real-world impact across a multitude of deployment scenarios. Recommender systems are often evaluated only through accuracy metrics, which fall short of fully characterizing their generalization capabilities and miss important aspects, such as fairness, bias, usefulness, informativeness. This workshop builds on the success of last year's workshop at CIKM, but with a broader scope and an interactive format.
Federico Bianchi 0001, Patrick John Chia, Jacopo Tagliabue, Ciro Greco, Gabriel de Souza P. Moreira, Davide Eynard, Fahd Husain, Claudio Pomo
KDD4
2021 Language in a (Search) Box: Grounding Language Learning in Real-World Human-Machine Interaction
abstract
We investigate grounded language learning through real-world data, by modelling a teacher-learner dynamics through the natural interactions occurring between users and search engines; in particular, we explore the emergence of semantic generalization from unsupervised dense representations outside of synthetic environments.A grounding domain, a denotation function and a composition function are learned from user data only.We show how the resulting semantics for noun phrases exhibits compositional properties while being fully learnable without any explicit labelling.We benchmark our grounded semantics on compositionality and zero-shot inference tasks, and we show that it provides better results and better generalizations than SOTA non-grounded models, such as word2vec and BERT.
Federico Bianchi 0001, Ciro Greco, Jacopo Tagliabue
NAACL-HLT2
2019 Less (Data) Is More: Why Small Data Holds the Key to the Future of Artificial Intelligence
abstract
The claims that big data holds the key to enterprise successes and that Artificial Intelligence is going to replace humanity have become increasingly more popular over the past few years, both in academia and in the industry. However, while these claims may indeed capture some truth, they have also been massively oversold, or so we contend here. The goal of this paper is two-fold. First, we provide a qualified defence of the value of less data within the context of AI. This is done by carefully reviewing two distinct problems for big data driven AI, namely a) the limited track record of Deep Learning in key areas such as Natural Language Processing, b) the regulatory and business significance of being able to learn from few data points. Second, we briefly sketch what we refer to as a case of AI with humans and for humans, namely an AI paradigm whereby the systems we build are privacy-oriented and focused on human-machine collaboration, not competition. Combining our claims above, we conclude that when seen through the lens of cognitively inspired AI, the bright future of the discipline is about less data, not more, and more humans, not fewer.
Ciro Greco, Andrea Polonioli, Jacopo Tagliabue
DATA1