EDBT 2026 Demo / reviewers in the wild / expert
Marc Spaniol
dblp:s/MarcSpaniol
· DBLP profile ↗
15ranked-venue papers in the field
0as first author
4since 2021 · last 2025
0000-0002-5094-4523ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 9Database Systems & Data Management · 6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Introduction to the Special Issue on Temporal Web: Studying Time and the Temporal Dimension
Omar Alonso, Marc Spaniol, Ricardo Baeza-Yates |
ACM Trans. Web | 2 |
| 2021 | FETD2: A Framework for Enabling Textual Data Denoising via Robust Contextual Embeddings
Govind, Céline Alec, Jean-Luc Manguin, Marc Spaniol |
TPDL | 4 |
| 2021 | Semantic Tagging via Entity-Level Analytics: Assessment of Concise Content Tagging
Amit Kumar 0044, Marc Spaniol |
TPDL | 2 |
| 2021 | AnnoTag: Concise Content Annotation via LOD Tags derived from Entity-Level Analytics
Amit Kumar 0044, Marc Spaniol |
TPDL | 2 |
| 2018 | Semantic Fingerprinting: A Novel Method for Entity-Level Content Classification
Govind, Céline Alec, Marc Spaniol |
ICWE | 3 |
| 2018 | ELEVATE-Live: Assessment and Visualization of Online News Virality via Entity-Level Analytics
Govind, Céline Alec, Marc Spaniol |
ICWE | 3 |
| 2012 | PRAVDA-live: interactive knowledge harvestingabstractAcquiring high-quality (temporal) facts for knowledge bases is a labor-intensive process. Although there has been recent progress in the area of semi-supervised fact extraction, these approaches still have limitations, including a restricted corpus, a fixed set of relations to be extracted or a lack of assessment capabilities. In this paper we introduce PRAVDA-live, a framework that overcomes these limitations and supports the entire pipeline of interactive knowledge harvesting. To this end, our demo exhibits fact extraction from ad-hoc corpus creation, via relation specification, labeling and assessment all the way to ready-to-use RDF exports. Yafang Wang, Maximilian Dylla, Zhaochun Ren, Marc Spaniol, Gerhard Weikum |
CIKM | 4 |
| 2012 | Predicting the Evolution of Taxonomy Restructuring in Collective Web Catalogues
Natalia Boldyrev, Marc Spaniol, Gerhard Weikum |
WebDB | 2 |
| 2011 | Longitudinal Analytics on Web Archive Data: It's About Time!
Gerhard Weikum, Nikos Ntarmos, Marc Spaniol, Peter Triantafillou, András A. Benczúr, Scott Kirkpatrick, Philippe Rigaux, Mark Williamson |
CIDR | 3 |
| 2011 | Harvesting facts from textual web sources by constrained label propagationabstractThere have been major advances on automatically constructing large knowledge bases by extracting relational facts from Web and text sources. However, the world is dynamic: periodic events like sports competitions need to be interpreted with their respective timepoints, and facts such as coaching a sports team, holding political or business positions, and even marriages do not hold forever and should be augmented by their respective timespans. This paper addresses the problem of automatically harvesting temporal facts with such extended time-awareness. We employ pattern-based gathering techniques for fact candidates and construct a weighted pattern-candidate graph. Our key contribution is a system called PRAVDA based on a new kind of label propagation algorithm with a judiciously designed loss function, which iteratively processes the graph to label good temporal facts for a given set of target relations. Our experiments with online news and Wikipedia articles demonstrate the accuracy of this method. Yafang Wang, Bin Yang 0002, Lizhen Qu, Marc Spaniol, Gerhard Weikum |
CIKM | 4 |
| 2011 | AIDA: An Online Tool for Accurate Disambiguation of Named Entities in Text and Tables
Mohamed Amir Yosef, Johannes Hoffart, Ilaria Bordino, Marc Spaniol, Gerhard Weikum |
Proc. VLDB Endow. | 4 |
| 2011 | The SHARC framework for data quality in Web archiving
Dimitar Denev, Arturas Mazeika, Marc Spaniol, Gerhard Weikum |
VLDB J. | 3 |
| 2010 | Timely YAGO: harvesting, querying, and visualizing temporal knowledge from WikipediaabstractRecent progress in information extraction has shown how to automatically build large ontologies from high-quality sources like Wikipedia. But knowledge evolves over time; facts have associated validity intervals. Therefore, ontologies should include time as a first-class dimension. In this paper, we introduce Timely YAGO, which extends our previously built knowledge base YAGO with temporal aspects. This prototype system extracts temporal facts from Wikipedia infoboxes, categories, and lists in articles, and integrates these into the Timely YAGO knowledge base. We also support querying temporal facts, by temporal predicates in a SPARQL-style language. Visualization of query results is provided in order to better understand of the dynamic nature of knowledge. Yafang Wang, Mingjie Zhu, Lizhen Qu, Marc Spaniol, Gerhard Weikum |
EDBT | 4 |
| 2009 | SHARC: Framework for Quality-Conscious Web ArchivingabstractWeb archives preserve the history of born-digital content and offer great potential for sociologists, business analysts, and legal experts on intellectual property and compliance issues. Data quality is crucial for these purposes. Ideally, crawlers should gather sharp captures of entire Web sites, but the politeness etiquette and completeness requirement mandate very slow, long-duration crawling while Web sites undergo changes. This paper presents the SHARC framework for assessing the data quality in Web archives and for tuning capturing strategies towards better quality with given resources. We define quality measures, characterize their properties, and derive a suite of quality-conscious scheduling strategies for archive crawling. It is assumed that change rates of Web pages can be statistically predicted based on page types, directory depths, and URL names. We develop a stochastically optimal crawl algorithm for the offline case where all change rates are known. We generalize the approach into an online algorithm that detect information on a Web site while it is crawled. For dating a site capture and for assessing its quality, we propose several strategies that revisit pages after their initial downloads in a judiciously chosen order. All strategies are fully implemented in a testbed, and shown to be effective by experiments with both synthetically generated sites and a daily crawl series for a medium-sized site. Dimitar Denev, Arturas Mazeika, Marc Spaniol, Gerhard Weikum |
Proc. VLDB Endow. | 3 |
| 2007 | Watching the Blogosphere: Knowledge Sharing in the Web 2.0
Ralf Klamma, Yiwei Cao, Marc Spaniol |
ICWSM | 3 |