EDBT 2026 Demo / reviewers in the wild / expert
Endri Kacupaj
dblp:266/0421
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2022
0000-0001-5012-0420ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Contrastive Representation Learning for Conversational Question Answering over Knowledge GraphsabstractThis paper addresses the task of conversational question answering (ConvQA) over knowledge graphs (KGs). The majority of existing ConvQA methods rely on full supervision signals with a strict assumption of the availability of gold logical forms of queries to extract answers from the KG. However, creating such a gold logical form is not viable for each potential question in a real-world scenario. Hence, in the case of missing gold logical forms, the existing information retrieval-based approaches use weak supervision via heuristics or reinforcement learning, formulating ConvQA as a KG path ranking problem. Despite missing gold logical forms, an abundance of conversational contexts, such as entire dialog history with fluent responses and domain information, can be incorporated to effectively reach the correct KG path. This work proposes a contrastive representation learning-based approach to rank KG paths effectively. Our approach solves two key challenges. Firstly, it allows weak supervision-based learning that omits the necessity of gold annotations. Second, it incorporates the conversational context (entire dialog history and domain information) to jointly learn its homogeneous representation with KG paths to improve contrastive representations for effective path ranking. We evaluate our approach on standard datasets for ConvQA, on which it significantly outperforms existing baselines on all domains and overall. Specifically, in some cases, the Mean Reciprocal Rank (MRR) and [email protected] ranking metrics improve by absolute 10 and 18 points, respectively, compared to the state-of-the-art performance. Endri Kacupaj, Kuldeep Singh 0001, Maria Maleshkova, Jens Lehmann 0001 |
CIKM | 1 |
| 2021 | Demographic Aware Probabilistic Medical Knowledge Graph Embeddings of Electronic Medical Records
Aynur Guluzade, Endri Kacupaj, Maria Maleshkova |
AIME | 2 |
| 2021 | Conversational Question Answering over Knowledge Graphs with Transformer and Graph Attention NetworksabstractEndri Kacupaj, Joan Plepi, Kuldeep Singh, Harsh Thakkar, Jens Lehmann, Maria Maleshkova. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Endri Kacupaj, Joan Plepi, Kuldeep Singh 0001, Harsh Thakkar, Jens Lehmann 0001, Maria Maleshkova |
EACL | 1 |
| 2021 | ParaQA: A Question Answering Dataset with Paraphrase Responses for Single-Turn Conversation
Endri Kacupaj, Barshana Banerjee, Kuldeep Singh 0001, Jens Lehmann 0001 |
ESWC | 1 |
| 2021 | Context Transformer with Stacked Pointer Networks for Conversational Question Answering over Knowledge Graphs
Joan Plepi, Endri Kacupaj, Kuldeep Singh 0001, Harsh Thakkar, Jens Lehmann 0001 |
ESWC | 2 |
| 2021 | VOGUE: Answer Verbalization Through Multi-Task Learning
Endri Kacupaj, Shyamnath Premnadh, Kuldeep Singh 0001, Jens Lehmann 0001, Maria Maleshkova |
ECML/PKDD (3) | 1 |
| 2021 | GeoWINE: Geolocation based Wiki, Image, News and Event RetrievalabstractIn the context of social media, geolocation inference on news or events has become a very important task. In this paper, we present the GeoWINE (Geolocation-based Wiki-Image-News-Event retrieval) demonstrator, an effective modular system for multimodal retrieval which expects only a single image as input. The GeoWINE system consists of five modules in order to retrieve related information from various sources. The first module is a state-of-the-art model for geolocation estimation of images. The second module performs a geospatial-based query for entity retrieval using the Wikidata knowledge graph. The third module exploits four different image embedding representations, which are used to retrieve most similar entities compared to the input image. The last two modules perform news and event retrieval from EventRegistry and the Open Event Knowledge Graph (OEKG). GeoWINE provides an intuitive interface for end-users and is insightful for experts for reconfiguration to individual setups. The GeoWINE achieves promising results in entity label prediction for images on Google Landmarks dataset. The demonstrator is publicly available at http://cleopatra.ijs.si/geowine/. Golsa Tahmasebzadeh, Endri Kacupaj, Eric Müller-Budack, Sherzod Hakimov, Jens Lehmann 0001, Ralph Ewerth |
SIGIR | 2 |
| 2020 | MLM: A Benchmark Dataset for Multitask Learning with Multiple Languages and ModalitiesabstractIn this paper, we introduce the MLM (Multiple Languages and Modalities) dataset - a new resource to train and evaluate multitask systems on samples in multiple modalities and three languages. The generation process and inclusion of semantic data provide a resource that further tests the ability for multitask systems to learn relationships between entities. The dataset is designed for researchers and developers who build applications that perform multiple tasks on data encountered on the web and in digital archives. A second version of MLM provides a geo-representative subset of the data with weighted samples for countries of the European Union. We demonstrate the value of the resource in developing novel applications in the digital humanities with a motivating use case and specify a benchmark set of tasks to retrieve modalities and locate entities in the dataset. Evaluation of baseline multitask and single task systems on the full and geo-representative versions of MLM demonstrate the challenges of generalising on diverse data. In addition to the digital humanities, we expect the resource to contribute to research in multimodal representation learning, location estimation, and scene understanding. Jason Armitage, Endri Kacupaj, Golsa Tahmasebzadeh, Maria Maleshkova, Ralph Ewerth, Jens Lehmann 0001 |
CIKM | 2 |
| 2020 | VQuAnDa: Verbalization QUestion ANswering DAtasetabstractQuestion Answering (QA) systems over Knowledge Graphs (KGs) aim to provide a concise answer to a given natural language question. Despite the significant evolution of QA methods over the past years, there are still some core lines of work, which are lagging behind. This is especially true for methods and datasets that support the verbalization of answers in natural language. Specifically, to the best of our knowledge, none of the existing Question Answering datasets provide any verbalization data for the question-query pairs. Hence, we aim to fill this gap by providing the first QA dataset VQuAnDa that includes the verbalization of each answer. We base VQuAnDa on a commonly used large-scale QA dataset – LC-QuAD, in order to support compatibility and continuity of previous work. We complement the dataset with baseline scores for measuring future training and evaluation work, by using a set of standard sequence to sequence models and sharing the results of the experiments. This resource empowers researchers to train and evaluate a variety of models to generate answer verbalizations. Endri Kacupaj, Hamid Zafar, Jens Lehmann 0001, Maria Maleshkova |
ESWC | 1 |