Andreas Weiler

dblp:121/1122 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0002-8385-2537ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 Website Segmentation Beyond Structure: A Benchmark on Functional and Digital Maturity Classes
Jasmin S. Saxer, Jonathan Gerber, Andreas Weiler, Michael Grossniklaus
ECIR (2)3
2026 Benchmarking state of the art website embedding methods for effective processing and analysis in the public sector
abstract
The ability to understand and process websites is crucial across various domains. It lays the foundation for machine understanding of websites. Specifically, website embedding proves invaluable when monitoring local government websites within the context of digital transformation. In this paper, we present a comparison of different state-of-the-art website embedding methods and their capability of creating a reasonable website embedding for our specific task. The models consist of visual, mixed, and textual-based embedding methods. We compare the models with a baseline model which embeds the header section of a website. We measure the performance of the models using zero-shot and transfer learning. We evaluate the performance of the models on three different datasets. Additionally to the embedding scoring, we evaluate the classification performance on these datasets. From the zero-shot models Homepage2Vec with visual, a combination of visual and textual embedding, performs best in general over all datasets. When applying transfer learning, TF-IDF & FNN, a text based model, outperforms the others in both cluster scoring as well as precision and F1-score in the classification task. However, time is an important factor when it comes to processing large data quantities. Thus, when additionally considering the time needed, our baseline model is a good alternative, being 1.88 times faster with a maximum decrease of 10 % in the F1-score.
Jonathan Gerber, Jasmin S. Saxer, Bruno B. Kreiner, Andreas Weiler
J. Intell. Inf. Syst.4
2025 WebClasSeg-25: A Dual-Classified Webpage Segmentation Dataset - Integrating Functional and Maturity-Based Analysis
abstract
Webpage segmentation is a crucial task in web analysis, enabling improvements in information retrieval, user experience, and automated web understanding.However, existing segmentation datasets often lack both comprehensive visual and textual segmentation, as well as classification systems that capture the functional and qualitative aspects of webpages.In this paper, we introduce a novel webpage segmentation dataset that addresses these gaps by providing both visual and textual segmentations, alongside two classification frameworks.The first framework defines a nominal classification of segments based on their functional roles, such as main content, header, footer, and navigation bar.The second introduces an ordinal classification assessing the digital maturity of webpage segments, offering a structured evaluation of their design evolution and complexity.By integrating both classification schemes into a single dataset, our approach enables a more holistic analysis of webpage structures.Furthermore, given the rapid evolution of web design conventions, content structures, and technological trends, our dataset is designed to reflect contemporary webpage characteristics, ensuring its relevance for modern applications in web analysis and machine learning.Additionally, we provide first results on visual segmentation, demonstrating the effectiveness of our dataset in practical applications.
Jonathan Gerber, Jasmin S. Saxer, Kimia Rabishokr, Bruno B. Kreiner, Andreas Weiler
SIGIR5
2024 Towards Website X-Ray for Europe's Municipalities: Unveiling Digital Transformation with Multimodal Embeddings
Jonathan Gerber, Bruno B. Kreiner, Jasmin S. Saxer, Andreas Weiler
iiWAS (1)4
2024 Digilog: Enhancing Website Embedding on Local Governments - A Comparative Analysis
Jonathan Gerber, Bruno B. Kreiner, Jasmin S. Saxer, Andreas Weiler
ISMIS4
2017 Editorial: Survey and Experimental Analysis of Event Detection Techniques for Twitter
abstract
Twitter's popularity as a source of up-to-date news and information is constantly increasing. In response to this trend, numerous event detection techniques have been proposed to cope with the rate and volume of Twitter data streams. Although most of these works conduct some evaluation of the proposed technique, a comparative study is often omitted. In this paper, we present a survey and experimental analysis of state-of-the-art event detection techniques for Twitter data streams. In order to conduct this study, we define a series of measures to support the quantitative and qualitative comparison. We demonstrate the effectiveness of these measures by applying them to event detection techniques as well as to baseline approaches using real-world Twitter streaming data.
Andreas Weiler, Michael Grossniklaus, Marc H. Scholl
Comput. J.1
2016 Stability Evaluation of Event Detection Techniques for Twitter
Andreas Weiler, Jöran Beel, Bela Gipp, Michael Grossniklaus
IDA1
2016 Situation monitoring of urban areas using social media data streams
Andreas Weiler, Michael Grossniklaus, Marc H. Scholl
Inf. Syst.1
2016 An evaluation of the run-time and task-based performance of event detection techniques for Twitter
Andreas Weiler, Michael Grossniklaus, Marc H. Scholl
Inf. Syst.1
2015 Run-Time and Task-Based Performance of Event Detection Techniques for Twitter
Andreas Weiler, Michael Grossniklaus, Marc H. Scholl
CAiSE1
2014 Discovering OLAP dimensions in semi-structured data
Svetlana Mansmann, Nafees Ur Rehman, Andreas Weiler, Marc H. Scholl
Inf. Syst.3
2013 OLAPing social media: the case of Twitter
abstract
Social networks are platforms where millions of users interact frequently and share variety of digital content with each other. Users express their feelings and opinions on every topic of interest. These opinions carry import value for personal, academic and commercial applications, but the volume and the speed at which these are produced make it a challenging task for researchers and the underlying technologies to provide useful insights to such data. We attempt to extend the established OLAP(On-line Analytical Processing) technology to allow multidimensional analysis of social media data by integrating text and opinion mining methods into the data warehousing system and by exploiting various knowledge discovery techniques to deal with semi-structured and unstructured data from social media.
Nafees Ur Rehman, Andreas Weiler, Marc H. Scholl
ASONAM2
2012 Building a Data Warehouse for Twitter Stream Exploration
abstract
In the recent year Twitter has evolved into an extremely popular social network and has revolutionized the ways of interacting and exchanging information on the Internet. By making its public stream available through a set of APIs Twitter has triggered a wave of research initiatives aimed at analysis and knowledge discovery from the data about its users and their messaging activities. While most of the projects and tools are tailored towards solving specific tasks, we pursue a goal of providing an application in dependent and universal analytical platform for supporting any kind of analysis and knowledge discovery. We employ the well established data warehousing technology with its underlying multidimensional data model, ETL routine for loading and consolidating data from different sources, OLAP functionality for exploring the data and data mining tools for more sophisticated analysis. In this work we describe the process of transforming the original stream into a set of related multidimensional cubes and demonstrate how the resulting data warehouse can be used for solving a variety of analytical tasks. We expect our proposed approach to be applicable for analyzing the data of other social networks as well.
Nafees Ur Rehman, Svetlana Mansmann, Andreas Weiler, Marc H. Scholl
ASONAM3
2012 Discovering OLAP dimensions in semi-structured data
abstract
With the standard OLAP technology, cubes are constructed from the input data based on the available data fields and known relationships between them. Structuring the data into a set of numeric measures distributed along a set of uniformly structured dimensions may be unrealistic for applications dealing with semi-structured data. We propose to extend the capabilities of OLAP via content-driven discovery of measures and dimensional characteristics in the original dataset. New structural elements are discovered by means of data mining and other techniques and are therefore prone to changes as the underlying dataset evolves. In this work we focus on the challenge of generating, maintaining, and querying such discovered elements of the cube.
Svetlana Mansmann, Nafees Ur Rehman, Andreas Weiler, Marc H. Scholl
DOLAP3
2012 Discovering Dynamic Classification Hierarchies in OLAP Dimensions
Nafees Ur Rehman, Svetlana Mansmann, Andreas Weiler, Marc H. Scholl
ISMIS3