Valerio Cetorelli

dblp:290/0355 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0002-5406-702XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 25% Web and social media mining · 25% Recommender systems · 25%
Theoretical computer science
1 paper
Automata and formal languages · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems › beyond-accuracy recommendation
coverage maximization
1.012026
URLBank: Data-Driven URL Discovery via Temporal Link Graphs · WWW 2026
Web and social media mining › social network analysis › influence maximization
seed selection
1.012026
URLBank: Data-Driven URL Discovery via Temporal Link Graphs · WWW 2026
Information retrieval › search engines
web crawling
1.012026
URLBank: Data-Driven URL Discovery via Temporal Link Graphs · WWW 2026
Data integration and cleaning › data extraction › web data extraction
template-based extraction
0.512021
The Smallest Extraction Problem · Proc. VLDB Endow. 2021
Data integration and cleaning › data extraction
web data extraction
0.512021
The Smallest Extraction Problem · Proc. VLDB Endow. 2021
Automata and formal languages › formal grammars
context-free grammar
0.512021
The Smallest Extraction Problem · Proc. VLDB Endow. 2021

Methods — techniques the papers use, named apart from their topics

unsupervised grammar induction · 1.0temporal link graph analysis · 1.0optimization · 1.0greedy marginal gain · 1.0
YearPublicationVenuePosition
2026 URLBank: Data-Driven URL Discovery via Temporal Link Graphs
abstract
Web-scale editorial crawling must balance coverage and freshness within tight politeness and request budgets. Nonetheless, in production systems, manual seed management remains common despite being inefficient—oversampling redundant seeds while missing high-yield ones. URLBank replaces manual curation with a label-free controller that infers optimal seed selection directly from temporal crawl telemetry. It identifies candidate entry points, estimates stability—the persistence of links across crawler revolutions—and productivity—the rate of first-seen publications—and ranks them through greedy marginal gain on a shared-credit coverage objective. In a shadow A/B evaluation spanning 5,238 sites, URLBank consistently achieves higher coverage, greater efficiency, and earlier discovery under identical conditions. Gains remain stable across Top-K budgets, approaching near-complete coverage with far fewer seeds. Deployed alongside Meltwater's production crawler Pulitzer, URLBank operates with versioned policies, ranked prefixes for crawl budgets, and integrated health diagnostics, making allocation transparent, auditable, and reversible. Together, these results demonstrate that temporal signals, through an interpretable greedy objective, yield large, measurable improvements in industrial-scale coverage, resource efficiency, and freshness.
Felipe Marineli, Valerio Cetorelli, Valter Crescenzi, Tim Furche, Xiaonan Guo 0001
WWW2
2021 The Smallest Extraction Problem
abstract
We introduce landmark grammars , a new family of context-free grammars aimed at describing the HTML source code of pages published by large and templated websites and therefore at effectively tackling Web data extraction problems. Indeed, they address the inherent ambiguity of HTML, one of the main challenges of Web data extraction, which, despite over twenty years of research, has been largely neglected by the approaches presented in literature. We then formalize the Smallest Extraction Problem (SEP), an optimization problem for finding the grammar of a family that best describes a set of pages and contextually extract their data. Finally, we present an unsupervised learning algorithm to induce a landmark grammar from a set of pages sharing a common HTML template, and we present an automatic Web data extraction system. The experiments on consolidated benchmarks show that the approach can substantially contribute to improve the state-of-the-art.
Valerio Cetorelli, Paolo Atzeni, Valter Crescenzi, Franco Milicchio
Proc. VLDB Endow.1