EDBT 2026 Demo / reviewers in the wild / expert
Alex Falcon
dblp:250/9383
· DBLP profile ↗
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-6325-9066ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Retrieving Relevant Metaverses Using Hierarchical FeaturesabstractMetaverse environments offer immersive, multimedia-rich experiences with growing relevance in education, entertainment, and cultural applications. The ability to grasp the contents of these environments, consisting of many rooms or subspaces, is a key building block for implementing Metaverse retrieval systems. However, current methods remain limited, as they are not designed to separate local (e.g., individual multimedia elements, and room-level details) from global, Metaverse-level semantics. Moreover, public datasets on similar topics do not capture the complexity of multi-room environments filled with multimedia contents. Our contributions are twofold. First, we introduce HiCALM, a hierarchical Metaverse retrieval framework built around the structural hierarchy of Metaverse environments. HiCALM models Metaverses in a bottom-up way, progressively grasping local and global semantics. To reduce the gap with textual data and achieve high-performance text-to-Metaverse retrieval, a novel cross-modal hierarchical loss supervises the process by teaching the model to associate the hierarchical visual features with textual information, also extracted hierarchically. Second, to overcome the absence of suitable datasets for the task, we present Museums3k, a large-scale dataset of 3,000 virtual museums annotated with detailed descriptions, each composed of multiple rooms populated with diverse multimedia elements; and GamingMV, a smaller dataset with data coming from 239 gaming-related real-world Metaverses. Through extensive quantitative and qualitative experiments, we show that HiCALM achieves considerable improvements in text-to-Metaverse retrieval, obtaining up to 95.0% R@1 and 60.0% \(\text{nDCG}_{3}@10\) on Museums3k (more than +40% R@1 and +11% nDCG, compared to existing solutions), and up to 62.0% R@1 and 62.6% \(\text{nDCG}_{3}@10\) on GamingMV (more than +23% R@1 and +5% nDCG). Ali Abdari, Alex Falcon, Giuseppe Serra 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Reproducibility Companion Paper: AdOCTeRA - Adaptive Optimization Constraints for Improved text-guided Retrieval of ApartmentsabstractThis reproducibility Companion paper supports the approach presented in our paper titled ''AdOCTeRA: Adaptive Optimization Constraints for Improved Text-Guided Retrieval of Apartments.'' In that work, we addressed the problem of apartment retrieval using textual descriptions. Specifically, we proposed a novel adaptive approach that leverages the similarity between apartment descriptions in the dataset to enforce different levels of distance-a small distance between similar elements, a slightly larger distance between less similar elements, and an even larger distance between dissimilar elements. Ali Abdari, Alex Falcon, Giuseppe Serra 0001, Qiushi Huang |
ICMR | 2 |
| 2025 | HM3: Hierarchical Modeling of Multimedia Metaverses on 10000 Thematic Museums via Theme-aware Contrastive Loss FunctionabstractThe Metaverse and its immersive environments are gaining significant attention due to their potential applications across various fields, from healthcare to art. As their numbers grow, it becomes difficult to effectively search through them and identify those of interest to the user. Recently, Metaverses were modeled as multimedia-rich 3D scenarios. However, existing works on retrieving them via text have several shortcomings, including the lack of experimentation with joint analysis of heterogeneous multimedia formats within the Metaverse, the use of small-scale datasets with randomly aggregated elements, and the consequent lack of thematic coherence in retrieval methods. To address these issues, we introduce SAVAGE, a novel synthetic dataset of 10,000 thematic exhibitions containing both real-world paintings and generated video artworks. Moreover, we propose HM3, a new hierarchical methodology for Metaverse Retrieval which captures all the contents of the room and integrates both images and videos, while its training is guided by a novel theme-aware loss function. Experiments on SAVAGE demonstrate the effectiveness of HM3 in modelling museums. The method also shows considerable improvements on an existing dataset of Metaverses, with ablation studies and qualitative analyses confirming the utility of the proposed theme-aware loss function. Gianluca Macrì, Lorenzo Bazzana, Alex Falcon, Giuseppe Serra 0001 |
ICMR | 3 |
| 2025 | HierArtEx: Hierarchical Representations and Art Experts Supporting the Retrieval of Museums in the Metaverse
Alex Falcon, Ali Abdari, Giuseppe Serra 0001 |
MMM (2) | 1 |
| 2025 | ALCER3D: Adaptive Learning Constraints for Enhanced Retrieval of Complex Indoor 3D ScenariosabstractThe Metaverse is growing rapidly, resulting in thousands of rich virtual universes. This results in a difficult search process for the user, making advanced search tools a necessity. Existing methods leverage contrastive learning to obtain a function mapping a 3D scene and its textual descriptions into similar representations. However, Metaverse scenarios are complex, multimedia-rich 3D scenes containing many elements, making cross-modal alignment difficult. For instance, a museum dedicated to Van Gogh is unrelated to Warhol, yet it shares similarities with Matisse or Monet. To make the mapping functions aware of these nuances, we propose a novel learning strategy to integrate Adaptive Optimization Constraints, computing data-dependent distances using a language-based method we design and enforcing them between the representations at training time. This novelty sets our approach apart from standard procedures enforcing the same distance. We validate the effectiveness of two datasets, one including 6000 apartments, and a novel dataset of 3000 museums that we collect. We observe consistent improvements compared to existing methods. Moreover, we obtain better generalization when with very complex scenarios, e.g. on the museums dataset it obtains an average R@1 of 5.2% compared to 1.2% obtained by existing methods. Finally, the source code is available athttps://github.com/aliabdari/ALCER3D. Alex Falcon, Ali Abdari, Giuseppe Serra 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Introduction to the Special Issue on Text-Multimedia Retrieval: Retrieving Multimedia Data by Means of Natural Language
Alex Falcon, Giuseppe Serra 0001, Sergio Escalera, Michael Wray |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | AdOCTeRA: Adaptive Optimization Constraints for improved Text-guided Retrieval of ApartmentsabstractNowadays, it is common for workers to relocate to new countries while seeking better job opportunities, or to live as digital nomads. While doing so, they face the problem of finding a new place to call home, requiring them to trust online advertisements or to physically visit the apartment. Recently, the research community investigated the possibility of performing the search on the Metaverse, hence reducing time and costs related to traveling and limiting carbon emissions. The methods available are based on state-of-the-art cross-modal retrieval techniques, which learn a joint embedding space by mapping apartment-descriptions pairs close. However, these methodologies push all the other pairs far away in the embedding space. In this paper, we identify this decision as a limitation, since different apartments are likely to share many aspects. To overcome it, we propose AdOCTeRA, which automatically separates the apartments into three classes -- very similar, slightly similar, and dissimilar -- and proposes adaptive optimization constraints for each of them. We validate our methodology on a large dataset of more than 6000 apartments, obtaining considerable relative improvements over the previous state-of-the-art (+3.8% R@5 and +7.3% R@10), and consistent improvements over the baseline across all the experiments. The source code is available at \hrefhttps://github.com/aliabdari/AdOCTeRA https://github.com/aliabdari/AdOCTeRA Ali Abdari, Alex Falcon, Giuseppe Serra 0001 |
ICMR | 2 |
| 2024 | A Language-Based Solution to Enable Metaverse Retrieval
Ali Abdari, Alex Falcon, Giuseppe Serra 0001 |
MMM (3) | 2 |
| 2024 | Improving semantic video retrieval models by training with a relevance-aware online mining strategyabstractTo retrieve a video via a multimedia search engine, a textual query is usually created by the user and then used to perform the search. Recent state-of-the-art cross-modal retrieval methods learn a joint text-video embedding space by using contrastive loss functions, which maximize the similarity of positive pairs while decreasing that of the negative pairs. Although the choice of these pairs is fundamental for the construction of the joint embedding space, the selection procedure is usually driven by the relationships found within the dataset: a positive pair is commonly formed by a video and its own caption, whereas unrelated video-caption pairs represent the negative ones. We hypothesize that this choice results in a retrieval system with limited semantics understanding, as the standard training procedure requires the system to discriminate between groundtruth and negative even though there is no difference in their semantics. Therefore, differently from the previous approaches, in this paper we propose a novel strategy for the selection of both positive and negative pairs which takes into account both the annotations and the semantic contents of the captions. By doing so, the selected negatives do not share semantic concepts with the positive pair anymore, and it is also possible to discover new positives within the dataset. Based on our hypothesis, we provide a novel design of two popular contrastive loss functions, and explore their effectiveness on three heterogeneous state-of-the-art approaches. The extensive experimental analysis conducted on two datasets, EPIC-Kitchens-100 and MSR-VTT, validates the effectiveness of the proposed strategy, observing, e.g., more than +20% nDCG on EPIC-Kitchens-100. Furthermore, these results are corroborated with qualitative evidence both supporting our hypothesis and explaining why the proposed strategy effectively overcomes it. Alex Falcon, Giuseppe Serra 0001, Oswald Lanz |
Comput. Vis. Image Underst. | 1 |
| 2023 | Video question answering supported by a multi-task learning objectiveabstractAbstract Video Question Answering (VideoQA) concerns the realization of models able to analyze a video, and produce a meaningful answer to visual content-related questions. To encode the given question, word embedding techniques are used to compute a representation of the tokens suitable for neural networks. Yet almost all the works in the literature use the same technique, although recent advancements in NLP brought better solutions. This lack of analysis is a major shortcoming. To address it, in this paper we present a twofold contribution about this inquiry and its relation with question encoding. First of all, we integrate four of the most popular word embedding techniques in three recent VideoQA architectures, and investigate how they influence the performance on two public datasets: EgoVQA and PororoQA. Thanks to the learning process, we show that embeddings carry question type-dependent characteristics. Secondly, to leverage this result, we propose a simple yet effective multi-task learning protocol which uses an auxiliary task defined on the question types. By using the proposed learning strategy, significant improvements are observed in most of the combinations of network architecture and embedding under analysis. Alex Falcon, Giuseppe Serra 0001, Oswald Lanz |
Multim. Tools Appl. | 1 |
| 2022 | Relevance-based Margin for Contrastively-trained Video Retrieval ModelsabstractVideo retrieval using natural language queries has attracted increasing interest due to its relevance in real-world applications, from intelligent access in private media galleries to web-scale video search. Learning the cross-similarity of video and text in a joint embedding space is the dominant approach. To do so, a contrastive loss is usually employed because it organizes the embedding space by putting similar items close and dissimilar items far. This framework leads to competitive recall rates, as they solely focus on the rank of the groundtruth items. Yet, assessing the quality of the ranking list is of utmost importance when considering intelligent retrieval systems, since multiple items may share similar semantics, hence a high relevance. Moreover, the aforementioned framework uses a fixed margin to separate similar and dissimilar items, treating all non-groundtruth items as equally irrelevant. In this paper we propose to use a variable margin: we argue that varying the margin used during training based on how much relevant an item is to a given query, i.e. a relevance-based margin, easily improves the quality of the ranking lists measured through nDCG and mAP. We demonstrate the advantages of our technique using different models on EPIC-Kitchens-100 and YouCook2. We show that even if we carefully tuned the fixed margin, our technique (which does not have the margin as a hyper-parameter) would still achieve better performance. Finally, extensive ablation studies and qualitative analysis support the robustness of our approach. Code will be released at \urlhttps://github.com/aranciokov/RelevanceMargin-ICMR22. Alex Falcon, Swathikiran Sudhakaran, Giuseppe Serra 0001, Sergio Escalera, Oswald Lanz |
ICMR | 1 |
| 2022 | A Feature-space Multimodal Data Augmentation Technique for Text-video RetrievalabstractEvery hour, huge amounts of visual contents are posted on social media and user-generated content platforms. To find relevant videos by means of a natural language query, text-video retrieval methods have received increased attention over the past few years. Data augmentation techniques were introduced to increase the performance on unseen test examples by creating new training samples with the application of semantics-preserving techniques, such as color space or geometric transformations on images. Yet, these techniques are usually applied on raw data, leading to more resource-demanding solutions and also requiring the shareability of the raw data, which may not always be true, e.g. copyright issues with clips from movies or TV series. To address this shortcoming, we propose a multimodal data augmentation technique which works in the feature space and creates new videos and captions by mixing semantically similar samples. We experiment our solution on a large scale public dataset, EPIC-Kitchens-100, and achieve considerable improvements over a baseline method, improved state-of-the-art performance, while at the same time performing multiple ablation studies. We release code and pretrained models on Github at https://github.com/aranciokov/FSMMDA\_VideoRetrieval. Alex Falcon, Giuseppe Serra 0001, Oswald Lanz |
ACM Multimedia | 1 |