Simon Gottschalk 0001

dblp:183/0640 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-2576-4640ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Aligning Visual Contrastive learning models via Preference Optimization
abstract
Contrastive learning models have demonstrated impressive abilities to capture semantic similarities by aligning representations in the embedding space. However, their performance can be limited by the quality of the training data and its inherent biases. While Preference Optimization (PO) methods such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) have been applied to align generative models with human preferences, their use in contrastive learning has yet to be explored. This paper introduces a novel method for training contrastive learning models using different PO methods to break down complex concepts. Our method systematically aligns model behavior with desired preferences, enhancing performance on the targeted task. In particular, we focus on enhancing model robustness against typographic attacks and inductive biases, commonly seen in contrastive vision-language models like CLIP. Our experiments demonstrate that models trained using PO outperform standard contrastive learning techniques while retaining their ability to handle adversarial challenges and maintain accuracy on other downstream tasks. This makes our method well-suited for tasks requiring fairness, robustness, and alignment with specific preferences. We evaluate our method for tackling typographic attacks on images and explore its ability to disentangle gender concepts and mitigate gender bias, showcasing the versatility of our approach.
Amirabbas Afzali, Borna Khodabandeh, Ali Rasekh, Mahyar JafariNodeh, Sepehr Kazemi Ranjbar, Simon Gottschalk 0001
ICLR6
2025 Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
abstract
Despite significant advances in Multimodal Large Language Models (MLLMs), understanding complex temporal dynamics in videos remains a major challenge. Our experiments show that current Video Large Language Model (Video-LLM) architectures have critical limitations in temporal understanding, struggling with tasks that require detailed comprehension of action sequences and temporal progression. In this work, we propose a Video-LLM architecture that introduces stacked temporal attention modules directly within the vision encoder. This design incorporates a temporal attention in vision encoder, enabling the model to better capture the progression of actions and the relationships between frames before passing visual tokens to the LLM. Our results show that this approach significantly improves temporal reasoning and outperforms existing models in video question answering tasks, specifically in action recognition. We improve on benchmarks including VITATECS, MVBench, and Video-MME by up to +5.5%. By enhancing the vision encoder with temporal structure, we address a critical gap in video understanding for Video-LLMs. Project page and code are available at: https://alirasekh.github.io/STAVEQ2/
Ali Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk 0001, Mohsen Fayyaz
NeurIPS4
2024 Spatially Constrained Transformer with Efficient Global Relation Modelling for Spatio-Temporal Prediction
abstract
Accurate spatio-temporal prediction is crucial for the sustainable development of smart cities. However, current approaches often struggle to capture important spatio-temporal relationships, particularly overlooking global relations among distant city regions. Most existing techniques predominantly rely on Convolutional Neural Networks (CNNs) to capture global relations. However, CNNs exhibit neighbourhood bias, making them insufficient for capturing distant relations. To address this limitation, we propose ST-SampleNet, a novel transformer-based architecture that combines CNNs with self-attention mechanisms to capture both local and global relations effectively. Moreover, as the number of regions increases, the quadratic complexity of self-attention becomes a challenge. To tackle this issue, we introduce a lightweight region sampling strategy that prunes non-essential regions and enhances the efficiency of our approach. Furthermore, we introduce a spatially constrained position embedding that incorporates spatial neighbourhood information into the self-attention mechanism, aiding in semantic interpretation and improving the performance of ST-SampleNet. Our experimental evaluation on three real-world datasets demonstrates the effectiveness of ST-SampleNet. Additionally, our efficient variant achieves a 40% reduction in computational costs with only a marginal compromise in performance, approximately 1%.
Ashutosh Sao, Simon Gottschalk 0001
ECAI2
2024 Event-Specific Document Ranking Through Multi-stage Query Expansion Using an Event Knowledge Graph
Sara Abdollahi, Tin Kuculo, Simon Gottschalk 0001
ECIR (2)3
2024 Generating a Question Answering Dataset About Geographic Changes in a Knowledge Graph
Michalis Mitsios, Dharmen Punjani, Sara Abdollahi, Simon Gottschalk 0001, Eleni Tsalapati, Elena Demidova, Manolis Koubarakis
EKAW4
2024 Evaluating Entity Importance in a Cross-National Context using Crowdsourcing and Best-Worst Scaling
abstract
Understanding the significance of entities in the context of a particular search topic necessitates a special focus on linguistic, cultural, and national aspects. This paper presents an innovative study on obtaining entity importance scores within a cross-national context. While setting up an extensive crowdsourcing task pool, we asked the crowd-workers to rank given entities within a particular general-domain search topic following the best-worst scaling method. Considering two different platforms (Amazon Mechanical Turk and Toloka AI), we focus on crowdworkers from the USA and Russia, speaking English and Russian, respectively. As a result of this, we reveal a strong impact of the national factor of crowdworkers on the entity importance annotations. By highlighting differences and commonalities in the importance scores across nations, our work provides advanced insights and a dataset for future research in the field of cross-national entity ranking. Our findings and insights can be applied beyond the general domain search topics, in particular, when working with recommendations (e.g., for eCommerce items) the consideration of the cross-national factor plays a vital role on the quality.
Aleksandr Perevalov, Sara Abdollahi, Simon Gottschalk 0001, Andreas Both 0001
KES3
2023 MetaCitta: Deep Meta-Learning for Spatio-Temporal Prediction Across Cities and Tasks
abstract
Abstract Accurate spatio-temporal prediction is essential for capturing city dynamics and planning mobility services. State-of-the-art deep spatio-temporal predictive models depend on rich and representative training data for target regions and tasks. However, the availability of such data is typically limited. Furthermore, existing predictive models fail to utilize cross-correlations across tasks and cities. In this paper, we propose MetaCitta, a novel deep meta-learning approach that addresses the critical challenges of data scarcity and model generalization. MetaCitta adopts the data from different cities and tasks in a generalizable spatio-temporal deep neural network. We propose a novel meta-learning algorithm that minimizes the discrepancy between spatio-temporal representations across tasks and cities. Our experiments with real-world data demonstrate that the proposed MetaCitta approach outperforms state-of-the-art prediction methods for zero-shot learning and pre-training plus fine-tuning. Furthermore, MetaCitta is computationally more efficient than the existing meta-learning approaches.
Ashutosh Sao, Simon Gottschalk 0001, Nicolas Tempelmeier, Elena Demidova
PAKDD (4)2
2023 LaSER: Language-specific event recommendation
abstract
While societal events often impact people worldwide, a significant fraction of events has a local focus that primarily affects specific language communities. Examples include national elections, the development of the Coronavirus pandemic in different countries, and local film festivals such as the César Awards in France and the Moscow International Film Festival in Russia. However, existing entity recommendation approaches do not sufficiently address the language context of recommendation. This article introduces the novel task of language-specific event recommendation, which aims to recommend events relevant to the user query in the language-specific context. This task can support essential information retrieval activities, including web navigation and exploratory search, considering the language context of user information needs. We propose LaSER, a novel approach toward language-specific event recommendation. LaSER blends the language-specific latent representations (embeddings) of entities and events and spatio-temporal event features in a learning to rank model. This model is trained on publicly available Wikipedia Clickstream data. The results of our user study demonstrate that LaSER outperforms state-of-the-art recommendation baselines by up to 33 percentage points in [email protected] concerning the language-specific relevance of recommended events.
Sara Abdollahi, Simon Gottschalk 0001, Elena Demidova
J. Web Semant.2
2022 QuoteKG: A Multilingual Knowledge Graph of Quotes
abstract
Abstract Quotes of public figures can mark turning points in history. A quote can explain its originator’s actions, foreshadowing political or personal decisions and revealing character traits. Impactful quotes cross language barriers and influence the general population’s reaction to specific stances, always facing the risk of being misattributed or taken out of context. The provision of a cross-lingual knowledge graph of quotes that establishes the authenticity of quotes and their contexts is of great importance to allow the exploration of the lives of important people as well as topics from the perspective of what was actually said. In this paper, we present QuoteKG, the first multilingual knowledge graph of quotes. We propose the QuoteKG creation pipeline that extracts quotes from Wikiquote, a free and collaboratively created collection of quotes in many languages, and aligns different mentions of the same quote. QuoteKG includes nearly one million quotes in 55 languages, said by more than 69, 000 people of public interest across a wide range of topics. QuoteKG is publicly available and can be accessed via a SPARQL endpoint.
Tin Kuculo, Simon Gottschalk 0001, Elena Demidova
ESWC2
2021 WorldKG: A World-Scale Geographic Knowledge Graph
abstract
OpenStreetMap is a rich source of openly available geographic information. However, the representation of geographic entities, e.g., buildings, mountains, and cities, within OpenStreetMap is highly heterogeneous, diverse, and incomplete. As a result, this rich data source is hardly usable for real-world applications. This paper presents WorldKG - a new geographic knowledge graph aiming to provide a comprehensive semantic representation of geographic entities in OpenStreetMap. We describe the WorldKG knowledge graph, including its ontology that builds the semantic dataset backbone, the extraction procedure of the ontology and geographic entities from OpenStreetMap, and the methods to enhance entity annotation. We perform statistical and qualitative dataset assessment, demonstrating the large scale and high precision of the semantic geographic information in WorldKG.
Alishiba Dsouza, Nicolas Tempelmeier, Ran Yu 0001, Simon Gottschalk 0001, Elena Demidova
CIKM4
2021 GeoVectors: A Linked Open Corpus of OpenStreetMap Embeddings on World Scale
abstract
OpenStreetMap (OSM) is currently the richest publicly available information source on geographic entities (e.g., buildings and roads) worldwide. However, using OSM entities in machine learning models and other applications is challenging due to the large scale of OSM, the extreme heterogeneity of entity annotations, and a lack of a well-defined ontology to describe entity semantics and properties. This paper presents GeoVectors - a unique, comprehensive world-scale linked open corpus of OSM entity embeddings covering the entire OSM dataset and providing latent representations of over 980 million geographic entities in 180 countries. The GeoVectors corpus captures semantic and geographic dimensions of OSM entities and makes these entities directly accessible to machine learning algorithms and semantic applications. We create a semantic description of the GeoVectors corpus, including identity links to the Wikidata and DBpedia knowledge graphs to supply context information. Furthermore, we provide a SPARQL endpoint - a semantic interface that offers direct access to the semantic and latent representations of geographic entities in OSM.
Nicolas Tempelmeier, Simon Gottschalk 0001, Elena Demidova
CIKM2
2020 Event-QA: A Dataset for Event-Centric Question Answering over Knowledge Graphs
abstract
Semantic Question Answering (QA) is a crucial technology to facilitate intuitive user access to semantic information stored in knowledge graphs. Whereas most of the existing QA systems and datasets focus on entity-centric questions, very little is known about these systems' performance in the context of events. As new event-centric knowledge graphs emerge, datasets for such questions gain importance. In this paper, we present the Event-QA dataset for answering event-centric questions over knowledge graphs. Event-QA contains 1000 semantic queries and the corresponding English, German and Portuguese verbalizations for EventKG - an event-centric knowledge graph with more than 970 thousand events.
Tarcísio Souza Costa, Simon Gottschalk 0001, Elena Demidova
CIKM2
2019 HapPenIng: Happen, Predict, Infer - Event Series Completion in a Knowledge Graph
Simon Gottschalk 0001, Elena Demidova
ISWC (1)1
2018 Towards Better Understanding Researcher Strategies in Cross-Lingual Event Analytics
Simon Gottschalk 0001, Viola Bernacchi, Richard Rogers, Elena Demidova
TPDL1
2018 EventKG: A Multilingual Event-Centric Temporal Knowledge Graph
abstract
One of the key requirements to facilitate semantic analytics of information regarding contemporary and historical events on the Web, in the news and in social media is the availability of reference knowledge repositories containing comprehensive representations of events and temporal relations. Existing knowledge graphs, with popular examples including DBpedia, YAGO and Wikidata, focus mostly on entity-centric information and are insufficient in terms of their coverage and completeness with respect to events and temporal relations. EventKG presented in this paper is a multilingual event-centric temporal knowledge graph that addresses this gap. EventKG incorporates over 690 thousand contemporary and historical events and over 2.3 million temporal relations extracted from several large-scale knowledge graphs and semi-structured sources and makes them available through a canonical representation. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Simon Gottschalk 0001, Elena Demidova
ESWC1
2017 MultiWiki: Interlingual Text Passage Alignment in Wikipedia
abstract
In this article, we address the problem of text passage alignment across interlingual article pairs in Wikipedia. We develop methods that enable the identification and interlinking of text passages written in different languages and containing overlapping information. Interlingual text passage alignment can enable Wikipedia editors and readers to better understand language-specific context of entities, provide valuable insights in cultural differences, and build a basis for qualitative analysis of the articles. An important challenge in this context is the tradeoff between the granularity of the extracted text passages and the precision of the alignment. Whereas short text passages can result in more precise alignment, longer text passages can facilitate a better overview of the differences in an article pair. To better understand these aspects from the user perspective, we conduct a user study at the example of the German, Russian, and English Wikipedia and collect a user-annotated benchmark. Then we propose MultiWiki, a method that adopts an integrated approach to the text passage alignment using semantic similarity measures and greedy algorithms and achieves precise results with respect to the user-defined alignment. The MultiWiki demonstration is publicly available and currently supports four language pairs.
Simon Gottschalk 0001, Elena Demidova
ACM Trans. Web1
2016 Analysing Temporal Evolution of Interlingual Wikipedia Article Pairs
abstract
Wikipedia articles representing an entity or a topic in different language editions evolve independently within the scope of the language-specific user communities. This can lead to different points of views reflected in the articles, as well as complementary and inconsistent information. An analysis of how the information is propagated across the Wikipedia language editions can provide important insights in the article evolution along the temporal and cultural dimensions and support quality control. To facilitate such analysis, we present MultiWiki -- a novel web-based user interface that provides an overview of the similarities and differences across the article pairs originating from different language editions on a timeline. MultiWiki enables users to observe the changes in the interlingual article similarity over time and to perform a detailed visual comparison of the article snapshots at a particular time point.
Simon Gottschalk 0001, Elena Demidova
SIGIR1