EDBT 2026 Demo / reviewers in the wild / expert
Valerio Cetorelli
dblp:290/0355
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0002-5406-702XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 25% Web and social media mining · 25% Recommender systems · 25% | |
| Theoretical computer science
1 paper |
Automata and formal languages · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems › beyond-accuracy recommendation
coverage maximization |
1.0 | 1 | 2026 | URLBank: Data-Driven URL Discovery via Temporal Link Graphs · WWW 2026 |
Web and social media mining › social network analysis › influence maximization
seed selection |
1.0 | 1 | 2026 | URLBank: Data-Driven URL Discovery via Temporal Link Graphs · WWW 2026 |
Information retrieval › search engines
web crawling |
1.0 | 1 | 2026 | URLBank: Data-Driven URL Discovery via Temporal Link Graphs · WWW 2026 |
Data integration and cleaning › data extraction › web data extraction
template-based extraction |
0.5 | 1 | 2021 | The Smallest Extraction Problem · Proc. VLDB Endow. 2021 |
Data integration and cleaning › data extraction
web data extraction |
0.5 | 1 | 2021 | The Smallest Extraction Problem · Proc. VLDB Endow. 2021 |
Automata and formal languages › formal grammars
context-free grammar |
0.5 | 1 | 2021 | The Smallest Extraction Problem · Proc. VLDB Endow. 2021 |
Methods — techniques the papers use, named apart from their topics
unsupervised grammar induction · 1.0temporal link graph analysis · 1.0optimization · 1.0greedy marginal gain · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | URLBank: Data-Driven URL Discovery via Temporal Link GraphsabstractWeb-scale editorial crawling must balance coverage and freshness within tight politeness and request budgets. Nonetheless, in production systems, manual seed management remains common despite being inefficient—oversampling redundant seeds while missing high-yield ones. URLBank replaces manual curation with a label-free controller that infers optimal seed selection directly from temporal crawl telemetry. It identifies candidate entry points, estimates stability—the persistence of links across crawler revolutions—and productivity—the rate of first-seen publications—and ranks them through greedy marginal gain on a shared-credit coverage objective. In a shadow A/B evaluation spanning 5,238 sites, URLBank consistently achieves higher coverage, greater efficiency, and earlier discovery under identical conditions. Gains remain stable across Top-K budgets, approaching near-complete coverage with far fewer seeds. Deployed alongside Meltwater's production crawler Pulitzer, URLBank operates with versioned policies, ranked prefixes for crawl budgets, and integrated health diagnostics, making allocation transparent, auditable, and reversible. Together, these results demonstrate that temporal signals, through an interpretable greedy objective, yield large, measurable improvements in industrial-scale coverage, resource efficiency, and freshness. Felipe Marineli, Valerio Cetorelli, Valter Crescenzi, Tim Furche, Xiaonan Guo 0001 |
WWW | 2 |
| 2021 | The Smallest Extraction ProblemabstractWe introduce landmark grammars , a new family of context-free grammars aimed at describing the HTML source code of pages published by large and templated websites and therefore at effectively tackling Web data extraction problems. Indeed, they address the inherent ambiguity of HTML, one of the main challenges of Web data extraction, which, despite over twenty years of research, has been largely neglected by the approaches presented in literature. We then formalize the Smallest Extraction Problem (SEP), an optimization problem for finding the grammar of a family that best describes a set of pages and contextually extract their data. Finally, we present an unsupervised learning algorithm to induce a landmark grammar from a set of pages sharing a common HTML template, and we present an automatic Web data extraction system. The experiments on consolidated benchmarks show that the approach can substantially contribute to improve the state-of-the-art. Valerio Cetorelli, Paolo Atzeni, Valter Crescenzi, Franco Milicchio |
Proc. VLDB Endow. | 1 |