Marcus Herzog

dblp:91/4750 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8Applied, interdisciplinary, general and emerging computing · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Data integration and cleaning · 83% Information retrieval · 12% Indexing and storage engines · 4%
Artificial intelligence
2 papers
Information extraction and text analysis · 90% Knowledge representation and reasoning · 10%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning › data extraction
web data extraction
0.232009
Scalable Web Data Extraction for Online Market Intelligence · Proc. VLDB Endow. 2009
The Lixto Data Extraction Project - Back and Forth between Theory and Practice · PODS 2004
Web Information Acquisition with Lixto Suite · ICDE 2003
Data integration and cleaning › data preprocessing
data cleaning
0.112009
Scalable Web Data Extraction for Online Market Intelligence · Proc. VLDB Endow. 2009
Natural language and speech › Information extraction and text analysis › document understanding › table recognition
web table extraction
0.112007
Towards domain-independent information extraction from web tables · WWW 2007
Natural language and speech › Information extraction and text analysis › document understanding › table recognition
table detection
0.112006
Visually guided bottom-up table detection and segmentation in web documents · WWW 2006
Cloud and datacenter computing › cloud data management
cloud data processing
0.012009
Scalable Web Data Extraction for Online Market Intelligence · Proc. VLDB Endow. 2009
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge acquisition
0.012007
Towards domain-independent information extraction from web tables · WWW 2007
Indexing and storage engines
tree structures
0.012004
The Lixto Data Extraction Project - Back and Forth between Theory and Practice · PODS 2004

Methods — techniques the papers use, named apart from their topics

stream merging and filtering · 0.2data flow scenarios · 0.2visual box model · 0.1topological analysis · 0.1visual rendering analysis · 0.1heuristic grouping · 0.1web data transformation · 0.0
YearPublicationVenuePosition
2009 Scalable Web Data Extraction for Online Market Intelligence
abstract
Online market intelligence (OMI), in particular competitive intelligence for product pricing, is a very important application area for Web data extraction. However, OMI presents non-trivial challenges to data extraction technology. Sophisticated and highly parameterized navigation and extraction tasks are required. On-the-fly data cleansing is necessary in order two identify identical products from different suppliers. It must be possible to smoothly define data flow scenarios that merge and filter streams of extracted data stemming from several Web sites and store the resulting data into a data warehouse, where the data is subjected to market intelligence analytics. Finally, the system must be highly scalable, in order to be able to extract and process massive amounts of data in a short time. Lixto (www.lixto.com), a company offering data extraction tools and services, has been providing OMI solutions for several customers. In this paper we show how Lixto has tackled each of the above challenges by improving and extending its original data extraction software. Most importantly, we show how high scalability is achieved through cloud computing. This paper also features a case study from the computers and electronics market.
Robert Baumgartner, Georg Gottlob, Marcus Herzog
Proc. VLDB Endow.3
2007 Towards domain-independent information extraction from web tables
abstract
Traditionally, information extraction from web tables has focused on small, more or less homogeneous corpora, often based on assumptions about the use of tags. A multitude of different HTML implementations of web tables make these approaches difficult to scale. In this paper, we approach the problem of domain-independent information extraction from web tables by shifting our attention from the tree-based representation of web pages to a variation of the two-dimensional visual box model used by web browsers to display the information on the screen. The thereby obtained topological and style information allows us to fill the gap created by missing domain-specific knowledge about content and table templates. We believe that, in a future step, this approach can become the basis for a new way of large-scale knowledge acquisition from the current “Visual Web.”
Wolfgang Gatterbauer, Paul Bohunsky, Marcus Herzog, Bernhard Krüpl, Bernhard Pollak
WWW3
2006 Using Ontologies for Extracting Product Features from Web Pages
Wolfgang Holzinger, Bernhard Krüpl, Marcus Herzog
ISWC3
2006 Visually guided bottom-up table detection and segmentation in web documents
abstract
In the AllRight project, we are developing an algorithm for unsupervised table detection and segmentation that uses the visual rendition of a Web page rather than the HTML code. Our algorithm works bottom-up by grouping word bounding boxes into larger groups and uses a set of heuristics. It has already been implemented and a preliminary evaluation on about 6000 Web documents has been carried out.
Bernhard Krüpl, Marcus Herzog
WWW2
2005 The Personal Publication Reader: Illustrating Web Data Extraction, Personalization and Reasoning for the Semantic Web
Robert Baumgartner, Nicola Henze, Marcus Herzog
ESWC3
2005 Semantic Web Enabled Information Systems: Personalized Views on Web Data
Robert Baumgartner, Christian Enzi, Nicola Henze, Marc Herrlich, Marcus Herzog, Matthias Kriesell, Kai Tomaschewski
ICCSA (2)5
2005 The Personal Publication Reader
Fabian Abel, Robert Baumgartner, Adrian Brooks, Christian Enzi, Georg Gottlob, Nicola Henze, Marcus Herzog, Matthias Kriesell, Wolfgang Nejdl, Kai Tomaschewski
ISWC7
2004 The Lixto Data Extraction Project - Back and Forth between Theory and Practice
abstract
DATA
Georg Gottlob, Christoph Koch 0001, Robert Baumgartner, Marcus Herzog, Sergio Flesca
PODS4
2003 Web Information Acquisition with Lixto Suite
abstract
We demonstrate the Lixto Suite, a Web data extraction and transformation software kit for retrieving and converting information from various sources to various customer devices. With the Lixto Suite, nontechnical content managers can rapidly develop applications in the areas of m-commerce, e-commerce, content integration and corporate portals.
Robert Baumgartner, Michal Ceresna, Georg Gottlob, Marcus Herzog, Viktor Zigo
ICDE4