EDBT 2026 Demo / reviewers in the wild / expert
Tiziano Fagni
dblp:72/4019
· DBLP profile ↗
12ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0003-1921-7456ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Query processing and optimization · 61% Information retrieval · 39% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
query result caching |
0.1 | 1 | 2006 | Boosting the performance of Web search engines: Caching and prefetching query results by exploiting historical usage data · ACM Trans. Inf. Syst. 2006 |
Query processing and optimization › runtime optimization › prefetching
query result prefetching |
0.1 | 1 | 2006 | Boosting the performance of Web search engines: Caching and prefetching query results by exploiting historical usage data · ACM Trans. Inf. Syst. 2006 |
Information retrieval
search engines |
0.1 | 1 | 2006 | Boosting the performance of Web search engines: Caching and prefetching query results by exploiting historical usage data · ACM Trans. Inf. Syst. 2006 |
Information retrieval
query log analysis |
0.0 | 1 | 2006 | Boosting the performance of Web search engines: Caching and prefetching query results by exploiting historical usage data · ACM Trans. Inf. Syst. 2006 |
Methods — techniques the papers use, named apart from their topics
static-dynamic cache partitioning · 0.1replacement policy · 0.1concurrent caching · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and TargetsabstractThe rise of online platforms exacerbated the spread of hate speech, demanding scalable and effective detection. However, the accuracy of hate speech detection systems heavily relies on human-labeled data, which is inherently susceptible to biases. While previous work has examined the issue, the interplay between the characteristics of the annotator and those of the target of the hate are still unexplored. We fill this gap by leveraging an extensive dataset with rich socio-demographic information of both annotators and targets, uncovering how human biases manifest in relation to the target's attributes. Our analysis surfaces the presence of widespread biases, which we quantitatively describe and characterize based on their intensity and prevalence, revealing marked differences. Furthermore, we compare human biases with those exhibited by persona-based LLMs. Our findings indicate that while persona-based LLMs do exhibit biases, these differ significantly from those of human annotators. Overall, our work offers new and nuanced results on human biases in hate speech annotations, as well as fresh insights into the design of AI-driven hate speech detection systems. Tommaso Giorgi, Lorenzo Cima, Tiziano Fagni, Marco Avvenuti, Stefano Cresci |
ICWSM | 3 |
| 2025 | The DSA Transparency Database: Auditing Self-reported Moderation Actions by Social MediaabstractSince September 2023, the Digital Services Act (DSA) obliges large online platforms to submit detailed data on each moderation action they take within the European Union (EU) to the DSA Transparency Database. From its inception, this centralized database has sparked scholarly interest as an unprecedented and potentially unique trove of data on real-world online moderation. Here, we thoroughly analyze all 353.12M records submitted by the eight largest social media platforms in the EU during the first 100 days of the database. Specifically, we conduct a platform-wise comparative study of their: volume of moderation actions, grounds for decision, types of applied restrictions, types of moderated content, timeliness in undertaking and submitting moderation actions, and use of automation. Furthermore, we systematically cross-check the contents of the database with the platforms' own transparency reports. Our analyses reveal that (i) the platforms adhered only in part to the philosophy and structure of the database, (ii) the structure of the database is partially inadequate for the platforms' reporting needs, (iii) the platforms exhibited substantial differences in their moderation actions, (iv) a remarkable fraction of the database data is inconsistent, (v) the platform X (formerly Twitter) presents the most inconsistencies. Our findings have far-reaching implications for policymakers and scholars across diverse disciplines. They offer guidance for future regulations that cater to the reporting needs of online platforms in general, but also highlight opportunities to improve and refine the database itself. Amaury Trujillo, Tiziano Fagni, Stefano Cresci |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2024 | Evaluating large language models for user stance detection on X (Twitter)abstractAbstract Current stance detection methods employ topic-aligned data, resulting in many unexplored topics due to insufficient training samples. Large Language Models (LLMs) pre-trained on a vast amount of web data offer a viable solution when training data is unavailable. This work introduces Tweets2Stance - T2S , an unsupervised stance detection framework based on zero-shot classification, i.e. leveraging an LLM pre-trained on Natural Language Inference tasks. T2S detects a five-valued user’s stance on social-political statements by analyzing their X (Twitter) timeline. The Ground Truth of a user’s stance is obtained from Voting Advice Applications (VAAs). Through comprehensive experiments, a T2S’s optimal setting was identified for each election. Linguistic limitations related to the language model are further addressed by integrating state-of-the-art LLMs like GPT-4 and Mixtral into the T2S framework. The T2S framework’s generalization potential is demonstrated by measuring its performance (F1 and MAE scores) across nine datasets. These datasets were built by collecting tweets from competing parties’ Twitter accounts in nine political elections held in different countries from 2019 to 2021. The results, in terms of F1 and MAE scores, outperformed all baselines and approached the best scores for each election. This showcases the ability of T2S, particularly when combined with state-of-the-art LLMs, to generalize across different cultural-political contexts. Margherita Gambini, Caterina Senette, Tiziano Fagni, Maurizio Tesconi |
Mach. Learn. | 3 |
| 2023 | From Tweets to Stance: An Unsupervised Framework for User Stance Detection on Twitter
Margherita Gambini, Caterina Senette, Tiziano Fagni, Maurizio Tesconi |
DS | 3 |
| 2022 | Fine-grained Prediction of Political Leaning on Social Media with Unsupervised Deep LearningabstractPredicting the political leaning of social media users is an increasingly popular task, given its usefulness for electoral forecasts, opinion dynamics models and for studying the political dimension of polarization and disinformation. Here, we propose a novel unsupervised technique for learning fine-grained political leaning from the textual content of social media posts. Our technique leverages a deep neural network for learning latent political ideologies in a representation learning task. Then, users are projected in a low-dimensional ideology space where they are subsequently clustered. The political leaning of a user is automatically derived from the cluster to which the user is assigned. We evaluated our technique in two challenging classification tasks and we compared it to baselines and other state-of-the-art approaches. Our technique obtains the best results among all unsupervised techniques, with micro F1 = 0.426 in the 8-class task and micro F1 = 0.772 in the 3-class task. Other than being interesting on their own, our results also pave the way for the development of new and better unsupervised approaches for the detection of fine-grained political leaning. Tiziano Fagni, Stefano Cresci |
J. Artif. Intell. Res. | 1 |
| 2018 | Picture it in your mind: generating high level visual representations from textual descriptions
Fabio Carrara, Andrea Esuli, Tiziano Fagni, Fabrizio Falchi, Alejandro Moreo |
Inf. Retr. J. | 3 |
| 2010 | Selecting negative examples for hierarchical text classification: An experimental comparisonabstractAbstract Hierarchical text classification (HTC) approaches have recently attracted a lot of interest on the part of researchers in human language technology and machine learning, since they have been shown to bring about equal, if not better, classification accuracy with respect to their “flat” counterparts while allowing exponential time savings at both learning and classification time. A typical component of HTC methods is a “local” policy for selecting negative examples: Given a category c, its negative training examples are by default identified with the training examples that are negative for c and positive for the categories which are siblings of c in the hierarchy. However, this policy has always been taken for granted and never been subjected to careful scrutiny since first proposed 15 years ago. This article proposes a thorough experimental comparison between this policy and three other policies for the selection of negative examples in HTC contexts, one of which (BEST LOCAL (k)) is being proposed for the first time in this article. We compare these policies on the hierarchical versions of three supervised learning algorithms (boosting, support vector machines, and naïve Bayes) by performing experiments on two standard TC datasets, REUTERS‐21578 and RCV1‐V2. Tiziano Fagni, Fabrizio Sebastiani 0001 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | Boosting multi-label hierarchical text categorization
Andrea Esuli, Tiziano Fagni, Fabrizio Sebastiani 0001 |
Inf. Retr. | 2 |
| 2006 | MP-Boost: A Multiple-Pivot Boosting Algorithm and Its Application to Text Categorization
Andrea Esuli, Tiziano Fagni, Fabrizio Sebastiani 0001 |
SPIRE | 2 |
| 2006 | TreeBoost.MH: A Boosting Algorithm for Multi-label Hierarchical Text Categorization
Andrea Esuli, Tiziano Fagni, Fabrizio Sebastiani 0001 |
SPIRE | 2 |
| 2006 | Boosting the performance of Web search engines: Caching and prefetching query results by exploiting historical usage dataabstractThis article discusses efficiency and effectiveness issues in caching the results of queries submitted to a Web search engine (WSE). We propose SDC (Static Dynamic Cache), a new caching strategy aimed to efficiently exploit the temporal and spatial locality present in the stream of processed queries. SDC extracts from historical usage data the results of the most frequently submitted queries and stores them in astatic,read-onlyportion of the cache. The remaining entries of the cache are dynamically managed according to a given replacement policy and are used for those queries that cannot be satisfied by the static portion. Moreover, we improve the hit ratio of SDC by using an adaptive prefetching strategy, which anticipates future requests by introducing a limited overhead over the back-end WSE. We experimentally demonstrate the superiority of SDC over purely static and dynamic policies by measuring the hit ratio achieved on three large query logs by varying the cache parameters and the replacement policy used for managing the dynamic part of the cache. Finally, we deploy and measure the throughput achieved by a concurrent version of our caching system. Our tests show how the SDC cache can be efficiently exploited by many threads that concurrently serve the queries of different users. Tiziano Fagni, Raffaele Perego 0001, Fabrizio Silvestri, Salvatore Orlando 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2004 | A Highly Scalable Parallel Caching System for Web Search Engine Results
Tiziano Fagni, Raffaele Perego 0001, Fabrizio Silvestri |
Euro-Par | 1 |