VLDB 2026 Research / reviewers in the wild / expert
Robert Baumgartner
dblp:b/RobertBaumgartner
· DBLP profile ↗
27ranked-venue papers
12as first author
2since 2021 · last 2024
0000-0003-0899-4903ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 16 · 6 first-authorArtificial intelligence and machine learning · 7 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-authorTheory of computation · 4 · 4 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PDFA Distillation with Error Bound Guarantees
Robert Baumgartner, Sicco Verwer |
CIAA | 1 |
| 2023 | SoK: Explainable Machine Learning for Computer Security ApplicationsabstractExplainable Artificial Intelligence (XAI) aims to improve the transparency of machine learning (ML) pipelines. We systematize the increasingly growing (but fragmented) microcosm of studies that develop and utilize XAI methods for defensive and offensive cybersecurity tasks. We identify 3 cybersecurity stakeholders, i.e., model users, designers, and adversaries, who utilize XAI for 4 distinct objectives within an ML pipeline, namely 1) XAI-enabled user assistance, 2) XAI-enabled model verification, 3) explanation verification & robustness, and 4) offensive use of explanations. Our analysis of the literature indicates that many of the XAI applications are designed with little understanding of how they might be integrated into analyst workflows – user studies for explanation evaluation are conducted in only 14% of the cases. The security literature sometimes also fails to disentangle the role of the various stakeholders, e.g., by providing explanations to model users and designers while also exposing them to adversaries. Additionally, the role of model designers is particularly minimized in the security literature. To this end, we present an illustrative tutorial for model designers, demonstrating how XAI can help with model verification. We also discuss scenarios where interpretability by design may be a better alternative. The systematization and the tutorial enable us to challenge several assumptions, and present open problems that can help shape the future of XAI research within cybersecurity. Azqa Nadeem, Daniël Vos, Clinton Cao, Luca Pajola, Simon Dieck, Robert Baumgartner, Sicco Verwer |
EuroS&P | 6 |
| 2015 | Efficient Approximation of Head-Related Transfer Functions in Subbands for Accurate Sound LocalizationabstractHead-related transfer functions (HRTFs) describe the acoustic filtering of incoming sounds by the human morphology and are essential for listeners to localize sound sources in virtual auditory displays. Since rendering complex virtual scenes is computationally demanding, we propose four algorithms for efficiently representing HRTFs in subbands, i.e., as an analysis filterbank (FB) followed by a transfer matrix and a synthesis FB. All four algorithms use sparse approximation procedures to minimize the computational complexity while maintaining perceptually relevant HRTF properties. The first two algorithms separately optimize the complexity of the transfer matrix associated to each HRTF for fixed FBs. The other two algorithms jointly optimize the FBs and transfer matrices for complete HRTF sets by two variants. The first variant aims at minimizing the complexity of the transfer matrices, while the second one does it for the FBs. Numerical experiments investigate the latency-complexity trade-off and show that the proposed methods offer significant computational savings when compared with other available approaches. Psychoacoustic localization experiments were modeled and conducted to find a reasonable approximation tolerance so that no significant localization performance degradation was introduced by the subband representation. Damián Marelli, Robert Baumgartner, Piotr Majdak |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Web data extraction, applications and techniques: A survey
Emilio Ferrara, Pasquale De Meo, Giacomo Fiumara, Robert Baumgartner |
Knowl. Based Syst. | 4 |
| 2011 | A versatile model for web page representation, information extraction and content re-packagingabstractOn today's Web, designers take huge efforts to create visually rich websites that boast a magnitude of interactive elements. Contrarily, most web information extraction (WIE) algorithms are still based on attributed tree methods which struggle to deal with this complexity. In this paper, we introduce a versatile model to represent web documents. The model is based on gestalt theory principles---trying to capture the most important aspects in a formally exact way. It (i) represents and unifies access to visual layout, content and functional aspects; (ii) is implemented with semantic web techniques that can be leveraged for i.e. automatic reasoning. Considering the visual appearance of a web page, we view it as a collection of gestalt figures---based on gestalt primitives---each representing a specific design pattern, be it navigation menus or news articles. Based on this model, we introduce our WIE methodology, a re-engineering process involving design patterns, statistical distributions and text content properties. The complete framework consists of the UOM model, which formalizes the mentioned components, and the MANM layer that hints on structure and serialization, providing document re-packaging foundations. Finally, we discuss how we have applied and evaluated our model in the area of web accessibility. Bernhard Krüpl, Ruslan R. Fayzrakhmanov, Wolfgang Holzinger, Mathias Panzenböck, Robert Baumgartner |
ACM Symposium on Document Engineering | 5 |
| 2011 | Design of Automatically Adaptable Web Wrappers
Emilio Ferrara, Robert Baumgartner |
ICAART (1) | 2 |
| 2010 | Semantic Online Tourism Market Monitoring
Norbert Walchhofer, Milan Hronsky, Michael Pöttler, Robert Baumgartner, Karl Anton Froeschl |
ENTER | 4 |
| 2010 | A unified ontology-based web page model for improving accessibilityabstractFast technological advancements and little compliance with accessibility standards by Web page authors pose serious obstacles to the Web experience of the blind user. We propose a unified Web document model that enables us to create a richer browsing experience and improved navigability for blind users. The model provides an integrated view on all aspects of a Web page and is leveraged to create a multi-axial user interface. Ruslan R. Fayzrakhmanov, Max C. Göbel, Wolfgang Holzinger, Bernhard Krüpl, Robert Baumgartner |
WWW | 5 |
| 2009 | Automated Ontology-Driven Metasearch Generation with Metamorph
Wolfgang Holzinger, Bernhard Krüpl, Robert Baumgartner |
WISE | 3 |
| 2009 | A flight meta-search engine with metamorphabstractWe demonstrate a flight meta-search engine that is based on the Metamorph framework. Metamorph provides mechanisms to model web forms together with the interactions which are needed to fulfil a request, and can generate interaction sequences that pose queries using these web forms and collect the results. In this paper, we discuss an interesting new feature that makes use of the forms themselves as an information source. We show how data can be extracted from web forms (rather than the data behind web forms) to generate a graph of flight connections between cities. Bernhard Krüpl, Wolfgang Holzinger, Yansen Darmaputra, Robert Baumgartner |
WWW | 4 |
| 2009 | Scalable Web Data Extraction for Online Market IntelligenceabstractOnline market intelligence (OMI), in particular competitive intelligence for product pricing, is a very important application area for Web data extraction. However, OMI presents non-trivial challenges to data extraction technology. Sophisticated and highly parameterized navigation and extraction tasks are required. On-the-fly data cleansing is necessary in order two identify identical products from different suppliers. It must be possible to smoothly define data flow scenarios that merge and filter streams of extracted data stemming from several Web sites and store the resulting data into a data warehouse, where the data is subjected to market intelligence analytics. Finally, the system must be highly scalable, in order to be able to extract and process massive amounts of data in a short time. Lixto (www.lixto.com), a company offering data extraction tools and services, has been providing OMI solutions for several customers. In this paper we show how Lixto has tackled each of the above challenges by improving and extending its original data extraction software. Most importantly, we show how high scalability is achieved through cloud computing. This paper also features a case study from the computers and electronics market. Robert Baumgartner, Georg Gottlob, Marcus Herzog |
Proc. VLDB Endow. | 1 |
| 2008 | Exploiting semantic web technologies to model web form interactionsabstractForm mapping is the key problem that needs to be solved in order to get access to the hidden web. Currently available solutions for fully automatic mapping are not ready for commercial meta-search engines, which still have to rely on hand crafted code and are hard to maintain. Wolfgang Holzinger, Bernhard Krüpl, Robert Baumgartner |
WWW | 3 |
| 2007 | The Lixto Systems Applications in Business Intelligence and Semantic WebabstractThis paper shows how technologies for Web data extraction, syndication and integration allow for new applications and services in the Business Intelligence and the Semantic Web domain. First, we demonstrate how knowledge about market developments and competitor activities on the market can be extracted dynamically and automatically from semi-structured information sources on the Web. Then, we show how the data can be integrated in Business Intelligence Systems and how data can be classified, re-assigned and transformed with the aid of Semantic Web ontological domain knowledge. Existing Semantic Web and Business Intelligence applications and scenarios using our technology illustrate the whole process. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Robert Baumgartner, Oliver Frölich, Georg Gottlob |
ESWC | 1 |
| 2007 | Table Recognition and Understanding from PDF FilesabstractWe propose a flexible method for detecting and understanding tables in PDF files, which is not reliant upon one particular feature being present, for example ruling lines or indentations, and is therefore applicable to a wide variety of visual presentations. We describe the steps required in transforming the low-level PDF instructions into text segments, lines and boxes on a page. We propose three different classifications for published tables, and develop methods to detect these tables and correctly identify their respective rows and columns. We also explain how to recognize spanning rows and columns, and multi-line rows. Experimental results show that our algorithm is effective in converting a wide variety of tabular presentations into HTML for information extraction purposes. Tamir Hassan, Robert Baumgartner |
ICDAR | 2 |
| 2006 | Semantically Integrating Portlets in Portals Through Annotation
Iñaki Paz, Oscar Díaz 0001, Robert Baumgartner, Sergio Fernández Anzuola |
WISE | 3 |
| 2006 | Using graph matching techniques to wrap data from PDF documentsabstractWrapping is the process of navigating a data source, semi-automatically extracting data and transforming it into a form suitable for data processing applications. There are currently a number of established products on the market for wrapping data from web pages. One such approach is Lixto [1], a product of research performed at our institute.Our work is concerned with extending the wrapping functionality of Lixto to PDF documents. As the PDF format is relatively unstructured, this is a challenging task. We have developed a method to segment the page into blocks, which are represented as nodes in a relational graph. This paper describes our current research in the use of relational matching techniques on this graph to locate wrapping instances. Tamir Hassan, Robert Baumgartner |
WWW | 2 |
| 2005 | The Personal Publication Reader: Illustrating Web Data Extraction, Personalization and Reasoning for the Semantic Web
Robert Baumgartner, Nicola Henze, Marcus Herzog |
ESWC | 1 |
| 2005 | Semantic Web Enabled Information Systems: Personalized Views on Web Data
Robert Baumgartner, Christian Enzi, Nicola Henze, Marc Herrlich, Marcus Herzog, Matthias Kriesell, Kai Tomaschewski |
ICCSA (2) | 1 |
| 2005 | The Personal Publication Reader
Fabian Abel, Robert Baumgartner, Adrian Brooks, Christian Enzi, Georg Gottlob, Nicola Henze, Marcus Herzog, Matthias Kriesell, Wolfgang Nejdl, Kai Tomaschewski |
ISWC | 2 |
| 2004 | The Lixto Data Extraction Project - Back and Forth between Theory and PracticeabstractDATA Georg Gottlob, Christoph Koch 0001, Robert Baumgartner, Marcus Herzog, Sergio Flesca |
PODS | 3 |
| 2003 | Web Information Acquisition with Lixto SuiteabstractWe demonstrate the Lixto Suite, a Web data extraction and transformation software kit for retrieving and converting information from various sources to various customer devices. With the Lixto Suite, nontechnical content managers can rapidly develop applications in the areas of m-commerce, e-commerce, content integration and corporate portals. Robert Baumgartner, Michal Ceresna, Georg Gottlob, Marcus Herzog, Viktor Zigo |
ICDE | 1 |
| 2002 | Propositional default logics made easier: computational complexity of model checking
Robert Baumgartner, Georg Gottlob |
Theor. Comput. Sci. | 1 |
| 2001 | The Elog Web Extraction Language
Robert Baumgartner, Sergio Flesca, Georg Gottlob |
LPAR | 1 |
| 2001 | Declarative Information Extraction, Web Crawling, and Recursive Wrapping with Lixto
Robert Baumgartner, Sergio Flesca, Georg Gottlob |
LPNMR | 1 |
| 2001 | Visual Web Information Extraction with Lixto
Robert Baumgartner, Sergio Flesca, Georg Gottlob |
VLDB | 1 |
| 2001 | Supervised Wrapper Generation with Lixto
Robert Baumgartner, Sergio Flesca, Georg Gottlob |
VLDB | 1 |
| 1999 | On the Complexity of Model Checking for Propositional Default Logics: New Results and Tractable Cases
Robert Baumgartner, Georg Gottlob |
IJCAI | 1 |