Lorenzo Vaiani

dblp:210/1584 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-3605-1577ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Lightweight Strategies to Mitigate Small VideoLLM's Challenges in Video Summarization
Lorenzo Vaiani, Luca Cagliero, Wiktoria Woronko
DEXA (1)1
2025 KIEPrompter: Leveraging Lightweight Models' Predictions for Cost-Effective Key Information Extraction using Vision LLMs
abstract
Key information extraction (KIE) from visually rich documents, such as receipts and forms, involves a deep understanding of textual, visual, and layout feature information. Transformers fine-tuned for KIE achieve state-of-the-art performance but lack generality and portability across different domains. In contrast, vision large language models (VLLMs) offer higher flexibility and zero-shot capability but fall short with domain-specific layout relations unless performing a resource-demanding supervised fine-tuning. To reach the best compromise solution between lightweight models and VLLMs, we propose KIEPrompter, a cost-effective LLM-based KIE approach that leverages the predictions of lightweight models as external knowledge injected into VLLM prompts. By incorporating these auxiliary predictions, VLLMs are guided to attend relevant multimodal content without ad hoc training. The accuracy results achieved by KIEPrompter in three benchmark document collections are superior to those of VLLMs in both zero-shot and layout-sensitive scenarios. We compare various strategies for incorporating lightweight model predictions, ranging from coarse-grained predictions without explicit confidence scores to fine-grained per-element network logits. We also demonstrate that our approach is robust to the absence of specific classes in trained lightweight models, as the VLLMs' pre-training compensates for the limited generality of lightweight models.
Lorenzo Vaiani, Yihao Ding, Luca Cagliero, Jean Lee, Paolo Garza, Josiah Poon, Soyeon Caren Han
CIKM1
2025 Cross-modal consistency types in multimodal social data
abstract
Social media content, such as internet memes or tweets, are nowadays largely or mainly multimodal. Machine learning models often need to jointly process images and text to solve complex tasks such as hate speech detection or sentiment analysis. For example, the misogyny of a meme cannot be accurately predicted while considering the visual and textual modalities separately. Similarly, sentiment annotations for tweets’ images and text can be discordant. Detecting the samples with inconsistent modality contributions is particularly relevant to analyze machine learning model performance and explain classification errors. In this paper, we formalize the types of cross-modal consistency by differentiating between consistent cases and not. Cross-modal consistency denotes whether all modalities agree on the label (i.e., full consistency) or not (i.e., inconsistency). When the visual and textual modalities are discordant, we distinguish the cases in which a joint analysis of multimodal features is sufficient to solve the issue from those requiring a human agreement (i.e., NOR consistency). We also propose a CLIP-based architecture to predict the cross-modal consistency types and identify the modalities causing the inconsistency. The results achieved on benchmark datasets show that cross-modal consistency annotation is cost-effective, i.e., it provides relevant insights into model predictions while requiring a limited extra human effort.
Lorenzo Vaiani, Luca Cagliero, Paolo Garza, Jason Ravagli
Knowl. Based Syst.1
2024 On Leveraging Multi-Page Element Relations in Visually-Rich Documents
abstract
Thanks to the rapid progress of the digitalization process, Visually-Rich Documents (VRDs) such as PDF files or scanned documents have become among the most widespread sources of knowledge. However, Question Answering on VRDs is challenged by the presence of multi-page relationships between document elements such as tables, figures, sections. This paper addresses a specific Visual Question Answering subtask from VDRs where answer generation leverages pairwise element relations in multi-page documents. We explore the performance of text-only and multimodal Transformer-based architectures as well as open-source Large Language Models. The results show that multimodal Transformers outperform the other tested methods, particularly when training samples contain explicit textual references to the elements in the document layout.
Davide Napolitano, Lorenzo Vaiani, Luca Cagliero
COMPSAC2
2024 Efficient Neural Network-Based Estimation of Interval Shapley Values
abstract
The use of Shapley Values (SVs) to explain machine learning model predictions is established. Recent research efforts have been devoted to generating efficient Neural Network-based SVs estimates. However, the variability of the generated estimates, which depend on the selected data sampling, model, and training parameters, brings the reliability of such estimates into question. By leveraging the concept of Interval SVs, we propose to incorporate SVs uncertainty directly into the learning process. Specifically, we explain ensemble models composed of multiple predictors, each one generating potentially different outcomes. Unlike all existing approaches, the explainer design is tailored to Interval SVs learning instead of SVs only. We present three new Network-based explainers relying on different ISV paradigms, i.e., a Multi-Task Learning network inspired by the Shapley value's weighted least squares characterization and two Interval Shapley-Like Value Neural estimators. The experiments thoroughly evaluate the new approaches on ten benchmark datasets, looking for the best compromise between intervals’ accuracy and explainers’ efficiency.
Davide Napolitano, Lorenzo Vaiani, Luca Cagliero
IEEE Trans. Knowl. Data Eng.2
2023 ITALIC: An Italian Intent Classification Dataset
abstract
Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects.We introduce ITALIC, the first largescale speech dataset designed for intent classification in Italian.The dataset comprises 16,521 crowdsourced audio samples recorded by 70 speakers from various Italian regions and annotated with intent labels and additional metadata.We explore the versatility of ITALIC by evaluating current state-of-the-art speech and text models.Results on intent classification suggest that increasing scale and running language adaptation yield better speech models, monolingual text models outscore multilingual ones, and that speech recognition on ITALIC is more challenging than on existing Italian benchmarks.We release both the dataset and the annotation scheme to streamline the development of new Italian SLU models and language-specific datasets.
Alkis Koudounas, Moreno La Quatra, Lorenzo Vaiani, Luca Colomba, Giuseppe Attanasio, Eliana Pastor, Luca Cagliero, Elena Baralis
INTERSPEECH3
2022 How Much Attention Should we Pay to Mosquitoes?
abstract
Mosquitoes are a major global health problem. They are responsible for the transmission of diseases and can have a large impact on local economies. Monitoring mosquitoes is therefore helpful in preventing the outbreak of mosquito-borne diseases. In this paper, we propose a novel data-driven approach that leverages Transformer-based models for the identification of mosquitoes in audio recordings. The task aims at detecting the time intervals corresponding to the acoustic mosquito events in an audio signal. We formulate the problem as a sequence tagging task and train a Transformer-based model using a real-world dataset collecting mosquito recordings. By leveraging the sequential nature of mosquito recordings, we formulate the training objective so that the input recordings do not require fine-grained annotations. We show that our approach is able to outperform baseline methods using standard evaluation metrics, albeit suffering from unexpectedly high false negatives detection rates. In view of the achieved results, we propose future directions for the design of more effective mosquito detection models.
Moreno La Quatra, Lorenzo Vaiani, Alkis Koudounas, Luca Cagliero, Paolo Garza, Elena Baralis
ACM Multimedia2