VLDB 2026 Research / reviewers in the wild / expert
Tobias Meisen
dblp:99/4726
· DBLP profile ↗
11ranked-venue papers in the field
0as first author
10since 2021 · last 2026
0000-0002-1969-559XORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 6Database Systems & Data Management · 3Data Mining & Knowledge Discovery · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EXCODER: EXplainable Classification Of DiscretE Time Series Representations
Yannik Hahn, Antonin Königsfeld, Hasan Tercan, Tobias Meisen |
PAKDD (1) | 4 |
| 2025 | Out of Distribution Detection for Efficient Continual Learning in Quality Prediction for Arc WeldingabstractModern manufacturing relies heavily on fusion welding processes, including gas metal arc welding (GMAW). Despite significant advances in machine learning-based quality prediction, current models exhibit critical limitations when confronted with the inherent distribution shifts that occur in dynamic manufacturing environments. In this work, we extend the VQ-VAE Transformer architecture-previously demonstrating state-of-the-art performance in weld quality prediction-by leveraging its autoregressive loss as a reliable out-of-distribution (OOD) detection mechanism. Our approach exhibits superior performance compared to conventional reconstruction methods, embedding error-based techniques, and other established baselines. By integrating OOD detection with continual learning strategies, we optimize model adaptation, triggering updates only when necessary and thereby minimizing costly labeling requirements. We introduce a novel quantitative metric that simultaneously evaluates OOD detection capability while interpreting in-distribution performance. Experimental validation in real-world welding scenarios demonstrates that our framework effectively maintains robust quality prediction capabilities across significant distribution shifts, addressing critical challenges in dynamic manufacturing environments where process parameters frequently change. This research makes a substantial contribution to applied artificial intelligence by providing an explainable and at the same time adaptive solution for quality assurance in dynamic manufacturing processes-a crucial step towards robust, practical AI systems in the industrial environment. Yannik Hahn, Jan Voets, Antonin Königsfeld, Hasan Tercan, Tobias Meisen |
CIKM | 5 |
| 2024 | Quality Prediction in Arc Welding: Leveraging Transformer Models and Discrete Representations from Vector Quantised-VAEabstractModern manufacturing relies heavily on fusion welding processes, including gas metal arc welding (GMAW), which efficiently converts electrical energy into thermal energy to join metals. Despite decades of research and extensive application in the automotive and aerospace sectors, weld quality assessment in the GMAW process remains a major challenge. This paper presents a novel learning-based approach relying on a vector quantised variational autoencoder (VQ-VAE) for data representation. In addition, we are the first to provide a time series dataset to the research community that combines labeled and unlabeled time series data from the GMAW domain, thereby enabling further research. The core idea of our approach consists of two stages: In the first stage, we use a learned automatic extraction of local features of the input signal using a VQ-VAE architecture. Based on this, in the second stage, we use a transformer model that processes the discretized features and performs weld quality prediction and classification. Our approach addresses real-world scenarios and improves the prediction of quality and fill existing data gaps by providing a reliable approach for quality assessment during manufacturing based on sensor data. Yannik Hahn, Robert F. Maack, Hasan Tercan, Tobias Meisen, Marion Purrio, Guido Buchholz, Matthias Angerhausen |
CIKM | 4 |
| 2023 | Online Quality Prediction in Windshield Manufacturing using Data-Efficient Machine LearningabstractThe digitization of manufacturing processes opens up the possibility of using machine learning methods on process data to predict future product quality. Based on the model predictions, quality improvement actions can be taken at an early stage. However, significant challenges must be overcome to successfully implement the predictions. Production lines are subject to hardware and memory limitations and are characterized by constant changes in quality influencing factors. In this paper, we address these challenges and present an online prediction approach for real-world manufacturing processes. On the one hand, it includes methods for feature extraction and selection from multimodal process and sensor data. On the other hand, a continual learning method based on memory-aware synapses is developed to efficiently train an artificial neural network over process changes. We deploy and evaluate the approach in a windshield production process. Our experimental evaluation shows that the model can accurately predict windshield quality and achieve significant process improvement. By comparing with other learning strategies such as transfer learning, we also show that the continual learning method both prevents catastrophic forgetting of the model and maintains its data efficiency. Hasan Tercan, Tobias Meisen |
KDD | 2 |
| 2022 | DocSemMap 2.0: Semantic Labeling based on Textual Data Documentations Using Seq2Seq Context LearnerabstractMethods for automated semantic labeling of data are an indispensable basis for increasing the usability of data. On the one hand, they contribute to the homogenization of the annotations and thus to the increase in quality; on the other hand, they reduce the modeling effort, provided that the quality of the used methodology is sufficient. In the past, research has focused primarily on data- and label-based methods. Another approach that has received recent attention is the incorporation of textual data documentations to support the automatic mapping of datasets to a knowledge graph. However, upon deeper analysis, our recent approach called DocSemMap gives away potential in a number of places. In this paper, we extend the current state of the art approach by uncovering existing shortcomings and presenting our own improvements. Using a sequence-to-sequence model (Seq2Seq), we exploit the context of datasets. An additional introduced classifier provides the linkage of documentation and labels for prediction. Our extended approach achieves a sustainable improvement in comparison to the reference approach. Andreas Burgdorf, Alexander Paulus, André Pomp, Tobias Meisen |
CIKM | 4 |
| 2022 | Will This Online Shopping Session Succeed? Predicting Customer's Purchase Intention Using EmbeddingsabstractCustomers are increasingly using online channels to buy products. For e-commerce companies, this offers new opportunities to tailor the shopping experience to customers' needs. Therefore, it is of great importance for a company to know their customers' intentions while browsing their webpage. A major challenge is the real-time analysis of a customer's intention during browsing sessions. To this end, a representation of the customer's browsing behavior must be retrieved from their live interactions on the webpage. Typically, characteristic behavioral features are extracted manually based on the knowledge of marketing experts. In this paper, we propose a customer embedding representation that is based on the customer's click-events recorded during browsing sessions. Thus, our approach does not use manually extracted features and is not based on marketing expert domain knowledge, which makes it transferable to different webpages and different online markets. We demonstrate our approach using three different e-commerce datasets to successfully predict whether a customer is going to purchase a specific product. For the prediction, we utilize the customer embedding representations as input for different machine learning models. We compare our approach with existing state-of-the-art approaches for real-time purchase prediction and show that our proposed customer representation with an LSTM predictor outperforms the state-of-the-art approach on all three datasets. Additionally, the creation process of our customers' representation is on average 235 times faster than the creation process of the baseline. Miguel Alves Gomes, Richard Meyes, Philipp Meisen, Tobias Meisen |
CIKM | 4 |
| 2022 | PLASMA: A Semantic Modeling Tool for Domain ExpertsabstractIn recent years, Knowledge Graphs and Ontology-based Data Management have proven to be particularly effective in the efficient management and consolidation of heterogeneous data sources. In this context, semantic modeling has proven to be a useful approach for creating semantic data annotations. However, automatically generated semantic models usually need to be revised by a domain expert, who is often not familiar with semantic technologies. For addressing this issue, we propose the PLASMA semantic modeling tool, which aims at enabling domain experts to build semantic models from scratch or refine models created by automatic algorithms. We demonstrate the use of the tool and its user interface in two different semantic data management use cases for integrating smart city data in a public funded project, called City Dataspace, and for creating semantic models in an industrial use case at Siemens AG. Alexander Paulus, Andreas Burgdorf, Tristan Langer, André Pomp, Tobias Meisen, Sebastian Pol |
CIKM | 5 |
| 2022 | Domain-independent Data-to-Text Generation for Open Data
Andreas Burgdorf, Micaela Barkmann, André Pomp, Tobias Meisen |
DATA | 4 |
| 2021 | A Semantic Data Marketplace for Easy Data Sharing within a Smart CityabstractToday, smart city applications are largely based on data collected from different stakeholders. This presupposes that the required data sources are publicly available. While open data platforms already provide a number of urban data sources, enterprises and citizens have few opportunities to make their data available. To complicate things further, if the data is published, the processing of this data is already extremely time-consuming today, as the data sources are heterogeneous and the corresponding homogenization has to be carried out by the data consumers themselves. In this paper, we present a data marketplace that enables different stakeholders (public institutions, enterprises, citizens) to easily provide data that can especially contribute to the further realization of smart cities. This marketplace is based on the principles of semantic data management, i.e., data providers annotate their added data with semantic models. With the help of these models, the data sources can be found and understood by data consumers and finally homogenized in a way that is suitable for their application. André Pomp, Alexander Paulus, Andreas Burgdorf, Tobias Meisen |
CIKM | 4 |
| 2021 | Applied Feature-oriented Project Life Cycle Classification
Oliver Böhme, Tobias Meisen |
DATA | 2 |
| 2008 | Efficient similarity search using the Earth Mover's Distance for large multimedia databasesabstractMultimedia similarity search in large databases requires efficient query processing. The Earth mover's distance, introduced in computer vision, is successfully used as a similarity model in a number of small-scale applications. Its computational complexity hindered its adoption in large multimedia databases. We enable directly indexing the Earth mover's distance in structures such as the R-tree and the VA-file by providing the accurate 'MinDist' function to any bounding rectangle in the index. We exploit the computational structure of the new MinDist to derive a new lower bound for the EMD MinDist which is assembled from quantized partial solutions yielding very fast query processing times. We prove completeness of our approach in a multistep scheme. Extensive experiments on real world data demonstrate the high efficiency. Ira Assent, Marc Wichterich, Tobias Meisen, Thomas Seidl 0001 |
ICDE | 3 |