Marco Fisichella

dblp:40/8342 · DBLP profile ↗
← Back
24ranked-venue papers in the field
5as first author
14since 2021 · last 2026
0000-0002-6894-1101ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 13 (2 first)Database Systems & Data Management · 6 (2 first)Data Mining & Knowledge Discovery · 3Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 End-to-End Deep Entity Resolution without Labelled Instances
Franziska Neuhof, Marco Fisichella, George Papadakis 0001
DOLAP2
2026 SMBench: No-code benchmarking of learning-based entity matching
abstract
Entity Resolution (ER) constitutes a challenging data integration task that is typically addressed through the Filtering-Verification framework. Filtering reduces the quadratic search space in an unsupervised manner that relies on heuristics, whereas verification performs matching, usually through a machine or a deep learning-based approach. Numerous solutions have been proposed for each step, but analyzing their combined performance constitutes a non-trivial task, due to technical and methodological challenges, while the literature typically examines them as orthogonal tasks. We facilitate the benchmarking of state-of-the-art verification algorithms under realistic settings, applying them to the candidate pairs generated by established filtering approaches from popular real-world datasets. To democratize this benchmarking, we developed an open-source, hands-off Web application, called SMBench, which allows users to perform a wealth of experiments through an intuitive user interface that requires no coding or ER expertise. SMBench is publicly available at https://smbench.kbs.uni-hannover.de , while its code is released through https://github.com/erbench/erbench . We delve into its frontend and backend, elaborating on the technologies used for their implementation as well as on the state-of-the-art ER methods they support. Using SMBench, we perform an extended experimental analysis that combines 3 filtering methods with 7 verification approaches, applying them to 9 datasets. The experimental results lead to interesting insights into the relative effectiveness, time and memory efficiency of the considered methods.
Oleh Astappiev, Franziska Neuhof, Marco Fisichella, George Papadakis 0001
Inf. Syst.3
2025 ITL-LIME: Instance-Based Transfer Learning for Enhancing Local Explanations in Low-Resource Data Settings
abstract
Explainable Artificial Intelligence (XAI) methods, such as Local Interpretable Model-Agnostic Explanations (LIME), have advanced the interpretability of black-box machine learning models by approximating their behavior locally using interpretable surrogate models. However, LIME's inherent randomness in perturbation and sampling can lead to locality and instability issues, especially in scenarios with limited training data. In such cases, data scarcity can result in the generation of unrealistic variations and samples that deviate from the true data manifold. Consequently, the surrogate model may fail to accurately approximate the complex decision boundary of the original model. To address these challenges, we propose a novel Instance-based Transfer Learning LIME framework (ITL-LIME) that enhances explanation fidelity and stability in data-constrained environments. ITL-LIME introduces instance transfer learning into the LIME framework by leveraging relevant real instances from a related source domain to aid the explanation process in the target domain. Specifically, we employ clustering to partition the source domain into clusters with representative prototypes. Instead of generating random perturbations, our method retrieves pertinent real source instances from the source cluster whose prototype is most similar to the target instance. These are then combined with the target instance's neighboring real instances. To define a compact locality, we further construct a contrastive learning-based encoder as a weighting mechanism to assign weights to the instances from the combined set based on their proximity to the target instance. Finally, these weighted source and target instances are used to train the surrogate model for explanation purposes. Experimental evaluation with real-world datasets demonstrates that ITL-LIME greatly improves the stability and fidelity of LIME explanations in scenarios with limited data. Our code is available at https://github.com/rehanrazaa/ITL-LIME.
Rehan Raza, Guanjin Wang, Kevin Kok Wai Wong, Hamid Laga, Marco Fisichella
CIKM5
2025 1st Workshop on Detecting Trust, Authority, Sense and Knowledge in Online News Media Production
Giovanni Fulantelli, Davide Taibi 0002, Sergio Splendore, Marco Fisichella
WSDM4
2025 Fed-FUEL: fairness and utility enhancing agnostic federated learning framework
abstract
Abstract Federated learning (FL) is an emerging communication-efficient and collaborative learning paradigm of machine learning with privacy guarantees. As these advancements unfold, adapting FL for fairness-aware learning becomes crucial. In this context, we propose a pre-processing fairness and utility (balanced accuracy) enhancing agnostic federated learning framework (Fed-FUEL) that mitigates discrimination embedded in the non-independent identically distributed data. We contribute a novel adaptive data manipulation method that mitigates discrimination embedded in the data at client side during optimization, resulting in an optimized and fair centralized server. This pre-processing approach abstracts the model architecture from the equation, offering a significant advantage in a federated environment. This abstraction not only facilitates a broader application across diverse model architectures without necessitating modifications but also sidesteps the potential complexities and inefficiencies associated with model-specific in-processing methods. Extensive experiments with a range of publicly available datasets demonstrate that our method outperforms the competing baselines in terms of both discrimination mitigation and predictive performance. Our model effectively adapts to both statistical and causal fairness notions, as shown through our experiments.
Maryam Badar, Raneen Younis, Sandipan Sikdar, Wolfgang Nejdl, Marco Fisichella
Data Min. Knowl. Discov.5
2025 Interpretable zero-shot stance detection with proactive content intervention
abstract
Zero-Shot Stance Detection (ZSSD) identifies an author’s stance towards unseen targets. Existing works have mainly focused on contrastive, meta, adversarial learning, or data augmentation but face issues like data scarcity, generalizability , and lack of coherence between text and targets. Moreover, stance detection must be interpretable to ensure transparency. Recent works with large language models (LLMs) aim to enhance unseen target knowledge or generate explanations but often rely excessively on explicit reasoning or provide coarse explanations, overlooking implicit cues and complicating interpretation. To address these challenges, we propose a novel interpretable multi-stage ZSSD framework. Stage 1 decodes explanations (rationales) justifying the stance while Stage 2 provides the final stance label, thus providing inherent interpretability in predicting stances. Extensive experiments prove that our approach outperforms other baselines with an average improvement in F1 scores of 27.99% with LLMs and 23.60% without LLMs for SemEval and 14.62% with LLMs and 25.24% without LLMs for VAST datasets for the ZSSD task, benefiting from the proposed pipeline architecture and interpretable design. Furthermore, to mitigate the harmful effects of offensive content and promote a more respectful online environment, we integrate an intervention module that leverages the contextual insights derived from our ZSSD framework with the ethics-based text generation power of LLMs to develop interventions. Automatic and human evaluation of LLM-generated interventions based on various proposed criteria provide insights into how LLMs perceive similar information from different perspectives, which can help foster morally sound and respectful online discourse.
Apoorva Upadhyaya, Wolfgang Nejdl, Marco Fisichella
Inf. Process. Manag.3
2024 Harnessing Empathy and Ethics for Relevance Detection and Information Categorization in Climate and COVID-19 Tweets
abstract
In this work, we aim to understand the general public perception of societal issues related to the current climate crisis and the COVID-19 pandemic on Twitter (X). Social media discussions on such matters often lead to misleading information, resulting in delays in initiatives proposed by governments or policymakers. Hence, we focus on extracting relevant information from the conversations on climate change and COVID that could be useful for authorities to curb the spread of potentially biased information by proposing the classification tasks of relevance detection (RD) and information categorization (IC). We first curate the datasets for the RD and IC tasks for the climate domain and extend the COVID-19 benchmark attention-worthy Twitter dataset for the IC task through manual annotation. We initially conduct experiments with LLMs and observe that LLMs can extract the relevant information in zero and few-shot settings based on multi-perspective reasoning in the form of cognitive empathy and ethical standards, but still perform worse than fine-tuned small language models. Based on the initial findings, we conclude that LLMs may not be the best extractor of relevant information, but induce cognitive empathy and ethical reasonings that can intuitively guide supervised models. To achieve this idea, we develop a cognitive empathy and ethical reasoning-based multi-tasking pipelined network for RD and IC tasks. Our proposed approach provides valuable insights that could be useful in real-world scenarios for governments, policymakers, and other researchers to decode the overall public outlook on societal issues.
Apoorva Upadhyaya, Wolfgang Nejdl, Marco Fisichella
CIKM3
2024 Open benchmark for filtering techniques in entity resolution
Franziska Neuhof, Marco Fisichella, George Papadakis 0001, Konstantinos Nikoletos, Nikolaus Augsten, Wolfgang Nejdl, Manolis Koubarakis
VLDB J.2
2023 Benchmarking Filtering Techniques for Entity Resolution
abstract
Entity Resolution is the task of identifying pairs of entity profiles that represent the same real-world object. To avoid checking a quadratic number of entity pairs, various filtering techniques have been proposed that fall into two main categories: (i) blocking workflows group together entity profiles with identical or similar signatures, and (ii) nearest-neighbor methods convert all entity profiles into vectors and identify the closest ones to every query entity. Unfortunately, the main techniques from these two categories have rarely been compared in the literature and, thus, their relative performance is unknown. We perform the first systematic experimental study that investigates the relative performance of the main representatives per category over numerous established datasets. Comparing techniques from different categories turns out to be a non-trivial task due to the various configuration parameters that are hard to fine-tune, but have a significant impact on performance. We consider a plethora of parameter configurations, optimizing each technique with respect to recall and precision targets. Both schema-agnostic and schema-based settings are evaluated. The experimental results provide novel insights into the effectiveness, the time efficiency and the scalability of the considered techniques.
George Papadakis 0001, Marco Fisichella, Franziska Schoger, Georgios M. Mandilaras, Nikolaus Augsten, Wolfgang Nejdl
ICDE2
2023 A Multi-Task Model for Sentiment Aided Stance Detection of Climate Change Tweets
abstract
Climate change has become one of the biggest challenges of our time. Social media platforms such as Twitter play an important role in raising public awareness and spreading knowledge about the dangers of the current climate crisis. With the increasing number of campaigns and communication about climate change through social media, the information could create more awareness and reach the general public and policy makers. However, these Twitter communications lead to polarization of beliefs, opinion-dominated ideologies, and often a split into two communities of climate change deniers and believers. In this paper, we propose a framework that helps identify denier statements on Twitter and thus classifies the stance of the tweet into one of the two attitudes towards climate change (denier/believer). The sentimental aspects of Twitter data on climate change are deeply rooted in general public attitudes toward climate change. Therefore, our work focuses on learning two closely related tasks: Stance Detection and Sentiment Analysis of climate change tweets. We propose a multi-task framework that performs stance detection (primary task) and sentiment analysis (auxiliary task) simultaneously. The proposed model incorporates the feature-specific and shared-specific attention frameworks to fuse multiple features and learn the generalized features for both tasks. The experimental results show that the proposed framework increases the performance of the primary task, i.e., stance detection by benefiting from the auxiliary task, i.e., sentiment analysis compared to its uni-modal and single-task variants.
Apoorva Upadhyaya, Marco Fisichella, Wolfgang Nejdl
ICWSM2
2023 FLAMES2Graph: An Interpretable Federated Multivariate Time Series Classification Framework
abstract
Increasing privacy concerns have led to decentralized and federated machine learning techniques that allow individual clients to consult and train models collaboratively without sharing private information. Some of these applications, such as medical and healthcare, require the final decisions to be interpretable. One common form of data in these applications is multivariate time series, where deep neural networks, especially convolutional neural networks based approaches, have established excellent performance in their classification tasks. However, promising results and performance of deep learning models are a black box, and their decisions cannot always be guaranteed and trusted. While several approaches address the interpretability of deep learning models for multivariate time series data in a centralized environment, less effort has been made in a federated setting. In this work, we introduce FLAMES2Graph, a new horizontal federated learning framework designed to interpret the deep learning decisions of each client. FLAMES2Graph extracts and visualizes those input subsequences that are highly activated by a convolutional neural network. Besides, an evolution graph is created to capture the temporal dependencies between the extracted distinct subsequences. The federated learning clients only share this temporal evolution graph with the centralized server instead of trained model weights to create a global evolution graph. Our extensive experiments on various datasets from well-known multivariate benchmarks indicate that the FLAMES2Graph framework significantly outperforms other state-of-the-art federated methods while keeping privacy and augmenting network decision interpretation.
Raneen Younis, Zahra Ahmadi, Abdul Hakmeh, Marco Fisichella
KDD4
2023 A Multi-task Model for Emotion and Offensive Aided Stance Detection of Climate Change Tweets
abstract
In this work, we address the United Nations Sustainable Development Goal 13: Climate Action by focusing on identifying public attitudes toward climate change on social media platforms such as Twitter. Climate change is threatening the health of the planet and humanity. Public engagement is critical to address climate change. However, climate change conversations on Twitter tend to polarize beliefs, leading to misinformation and fake news that influence public attitudes, often dividing them into climate change believers and deniers. Our paper proposes an approach to classify the attitude of climate change tweets (believe/deny/ambiguous) to identify denier statements on Twitter. Most existing approaches for detecting stances and classifying climate change tweets either overlook deniers’ tweets or do not have a suitable architecture. The relevant literature suggests that emotions and higher levels of toxicity are prevalent in climate change Twitter conversations, leading to a delay in appropriate climate action. Therefore, our work focuses on learning stance detection (main task) while exploiting the auxiliary tasks of recognizing emotions and offensive utterances. We propose a multimodal multitasking framework MEMOCLiC that captures the input data using different embedding techniques and attention frameworks, and then incorporates the learned emotional and offensive expressions to obtain an overall representation of the features relevant to the stance of the input tweet. Extensive experiments conducted on a novel curated climate change dataset and two benchmark stance detection datasets (SemEval-2016 and ClimateStance-2022) demonstrate the effectiveness of our approach.
Apoorva Upadhyaya, Marco Fisichella, Wolfgang Nejdl
WWW2
2023 Towards sentiment and Temporal Aided Stance Detection of climate change tweets
Apoorva Upadhyaya, Marco Fisichella, Wolfgang Nejdl
Inf. Process. Manag.2
2022 Partially-federated learning: A new approach to achieving privacy and effectiveness
Marco Fisichella, Gianluca Lax, Antonia Russo
Inf. Sci.1
2017 JustEvents: A Crowdsourced Corpus for Event Validation with Strict Temporal Constraints
Andrea Ceroni, Ujwal Gadiraju, Marco Fisichella
ECIR3
2016 Where the Event Lies: Predicting Event Occurrence in Textual Documents
abstract
Manually inspecting text in a document collection to assess whether an event occurs in it is a cumbersome task. Although a manual inspection can allow one to identify and discard false events, it becomes infeasible with increasing numbers of automatically detected events. In this paper, we present a system to automatize event validation, defined as the task of determining whether a given event occurs in a given document or corpus. In addition to supporting users seeking for information that corroborates a given event, event validation can also boost the precision of automatically detected event sets by discarding false events and preserving the true ones. The system allows to specify events, retrieves candidate web documents, and assesses whether events occur in them. The validation results are shown to the user, who can revise the decision of the system. The validation method relies on a supervised model to predict the occurrence of events in a non-annotated corpus. This system can also be used to build ground-truths for event corpora.
Andrea Ceroni, Ujwal Gadiraju, Jan Matschke, Simon Wingert, Marco Fisichella
SIGIR5
2015 Improving Event Detection by Automatically Assessing Validity of Event Occurrence in Text
abstract
Manually inspecting text to assess whether an event occurs in a document collection is an onerous and time consuming task. Although a manual inspection to discard the false events would increase the precision of automatically detected sets of events, it is not a scalable approach. In this paper, we automatize event validation, defined as the task of determining whether a given event occurs in a given document or corpus. The introduction of automatic event validation as a post-processing step of event detection can boost the precision of the detected event set, discarding false events and preserving the true ones. We propose a novel automatic method for event validation, which relies on a supervised model to predict the occurrence of events in a non-annotated corpus. The data for training the model is gathered by exploiting the crowdsourcing paradigm. Experiments on real-world events and documents show that our proposed method (i) outperforms the state-of-the-art event validation approach and (ii) increases the precision of event detection while preserving recall.
Andrea Ceroni, Ujwal Gadiraju, Marco Fisichella
CIKM3
2014 Predicting Pair Similarities for Near-Duplicate Detection in High Dimensional Spaces
Marco Fisichella, Andrea Ceroni, Fan Deng 0004, Wolfgang Nejdl
DEXA (2)1
2014 Towards an Entity-Based Automatic Event Validation
Andrea Ceroni, Marco Fisichella
ECIR2
2014 WikipEvent: Leveraging Wikipedia Edit History for Event Detection
Tuan Tran 0002, Andrea Ceroni, Mihai Georgescu, Kaweh Djafari Naini, Marco Fisichella
WISE (2)5
2012 Polynomial Asymptotic Complexity of Multiple-Objective OLAP Data Cube Compression
Alfredo Cuzzocrea, Marco Fisichella
IPMU (2)2
2011 Detecting Health Events on the Social Web to Enable Epidemic Intelligence
Marco Fisichella, Avare Stewart, Alfredo Cuzzocrea, Kerstin Denecke
SPIRE1
2010 Unsupervised public health event detection for epidemic intelligence
abstract
Recent pandemics such as Swine Flu have caused concern for public health officials. Given the ever increasing pace at which infectious diseases can spread globally, officials must be prepared to react sooner and with greater epidemic intelligence gathering capabilities. However, state-of-the-art systems for Epidemic Intelligence have not kept the pace with the growing need for more robust public health event detection. In this paper, we propose a game-changing approach where public health events are detected in an unsupervised manner. We address the problems associated with adapting an unsupervised learner to the medical domain and in doing so, propose an approach which combines aspects from different feature-based event detection methods. We evaluate our approach with a real world dataset with respect to the quality of article clusters. Our results show that we are able to achieve a precision of 66% and a recall of 81% when evaluated using manually annotated, real-world data. This shows promising results for the use of such techniques in this new problem setting.
Marco Fisichella, Avare Stewart, Kerstin Denecke, Wolfgang Nejdl
CIKM1
2010 Efficient Incremental Near Duplicate Detection Based on Locality Sensitive Hashing
Marco Fisichella, Fan Deng 0004, Wolfgang Nejdl
DEXA (1)1