Antonio Gulli

dblp:65/5164 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
0since 2021 · last 2016
0000-0002-1024-2324ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5Artificial intelligence and machine learning · 4Theory of computation · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Information retrieval · 81% Knowledge graphs · 10% Data mining · 7%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › search engines › semantic search › entity retrieval
entity ranking
0.212016
WSDM Cup 2016: Entity Ranking Challenge · WSDM 2016
Knowledge graphs › domain-specific knowledge graph
academic knowledge graph
0.112016
WSDM Cup 2016: Entity Ranking Challenge · WSDM 2016
Information retrieval › ranking › text ranking
news ranking
0.112005
Ranking a stream of news · WWW 2005
Information retrieval › document retrieval › domain-specific retrieval
news retrieval
0.112005
Ranking a stream of news · WWW 2005
Information retrieval › ranking
ranking model
0.112005
Ranking a stream of news · WWW 2005
Information retrieval
retrieval models
0.112005
Ranking a stream of news · WWW 2005
Information retrieval › ranking › context-aware ranking
temporal ranking
0.112005
Ranking a stream of news · WWW 2005
Data mining › clustering
hierarchical clustering
0.012004
The Anatomy of a Hierarchical Clustering Engine for Web-page, News and Book Snippets · ICDM 2004
Information retrieval › search engines
search result clustering
0.012004
The Anatomy of a Hierarchical Clustering Engine for Web-page, News and Book Snippets · ICDM 2004
Information retrieval › query reformulation
query refinement
0.012004
The Anatomy of a Hierarchical Clustering Engine for Web-page, News and Book Snippets · ICDM 2004

Methods — techniques the papers use, named apart from their topics

temporal modeling · 0.1ranking framework · 0.1clustering · 0.1ranking function · 0.0ephemeral clustering · 0.0
YearPublicationVenuePosition
2016 WSDM Cup 2016: Entity Ranking Challenge
abstract
In this paper, we describe the WSDM Cup entity ranking challenge held in conjunction with the 2016 Web Search and Data Mining conference (WSDM 2016). Participants in the challenge were provided access to the Microsoft Academic Graph (MAG), a large heterogeneous graph of academic entities, and were invited to calculate the query-independent importance of each publication in the graph. Submissions for the challenge were open from August through November 2015, and a public leaderboard displayed teams? progress against a set of training judgements. Final evaluations were performed against a separate, withheld portion of the evaluation judgements. The top eight performing teams were then invited to submit papers to the WSDM Cup workshop, held at the WSDM 2016 conference.
Alex D. Wade, Kuansan Wang, Yizhou Sun, Antonio Gulli
WSDM4
2009 TC-SocialRank: Ranking the Social Web
Antonio Gulli, Stefano Cataudella 0003, Luca Foschini 0002
WAW1
2008 A personalized search engine based on Web-snippet hierarchical clustering
abstract
Abstract We propose a (meta‐)search engine, called SnakeT (SNippet Aggregation for Knowledge ExtracTion), which queries more than 18 commodity search engines and offers two complementary views on their returned results. One is the classical flat‐ranked list, the other consists of a hierarchical organization of these results into folders created on‐the‐fly at query time and labeled with intelligible sentences that capture the themes of the results contained in them. Users can browse this hierarchy with various goals: knowledge extraction, query refinement and personalization of search results. In this novel form of personalization, the user is requested to interact with the hierarchy by selecting the folders whose labels (themes) best fit her query needs. SnakeT then personalizes on‐the‐fly the original ranked list by filtering out those results that do not belong to the selected folders. Consequently, this form of personalization is carried out by the users themselves and thus results fully adaptive, privacy preserving, scalable and non‐intrusive for the underlying search engines. We have extensively tested SnakeT and compared it against the best available Web‐snippet clustering engines. SnakeT is efficient and effective, and shows that a mutual reinforcement relationship between ranking and Web‐snippet clustering does exist. In fact, the better the ranking of the underlying search engines, the more relevant the results from which SnakeT distills the hierarchy of labeled folders, and hence the more useful this hierarchy is to the user. Vice versa, the more intelligible the folder hierarchy, the more effective the personalization offered by SnakeT on the ranking of the query results. Copyright © 2007 John Wiley & Sons, Ltd.
Paolo Ferragina, Antonio Gulli
Softw. Pract. Exp.2
2005 Ranking a stream of news
abstract
According to a recent survey made by Nielsen NetRatings, searching on news articles is one of the most important activity online. Indeed, Google, Yahoo, MSN and many others have proposed commercial search engines for indexing news feeds. Despite this commercial interest, no academic research has focused on ranking a stream of news articles and a set of news sources. In this paper, we introduce this problem by proposing a ranking framework which models: (1) the process of generation of a stream of news articles, (2) the news articles clustering by topics, and (3) the evolution of news story over the time. The ranking algorithm proposed ranks news information, finding the most authoritative news sources and identifying the most interesting events in the di#erent categories to which news article belongs. All these ranking measures take in account the time and can be obtained without a predefined sliding window of observation over the stream. The complexity of our algorithm is linear in the number of pieces of news still under consideration at the time of a new posting. This allow a continuous on-line process of ranking. Our ranking framework is validated on a collection of more than 300,000 pieces of news, produced in two months by more then 2000 news sources belonging to 13 di#erent categories (World, U.S, Europe, Sports, Business, etc). This collection is extracted from the index of comeToMyHead, an academic news search engine available online. Categories and Subject Descriptors H.3.1 [Information Storage And Retrieval]: Content Analysis and IndexingRetrieval models; Search process; H.3.3 [Information Storage And Retrieval]: Information Search and Retrieval; H.3.5 [Information Storage And Retrieval]: OnlineInformation Services General Terms Algorithms, Experimentat...
Gianna M. Del Corso, Antonio Gulli, Francesco Romani
WWW2
2004 The Anatomy of a Hierarchical Clustering Engine for Web-page, News and Book Snippets
abstract
In this paper, we investigate the Web snippet hierarchical clustering problem in its full extent by devising an algorithmic solution, and a software prototype called SnakeT (accessible at http://roquefort.di.unipi.it/), that: (1) draws the snippets from 16 Web search engines, the Amazon collection of books a9.com, the news of Google News and the blogs of Blogline; (2) builds the clusters on-the-fly (ephemeral clustering (Maarek et al., 2000)) in response to a user query without adopting any predefined organization in categories; (3) labels the clusters with sentences of variable length, drawn from the snippets and possibly missing some terms, provided they are not too many; (4) uses some ranking functions which exploit two knowledge bases properly built by our engine at preprocessing time for the sentences selection and cluster-assignment process; (5) organizes the clusters into a hierarchy, and assigns to the nodes intelligible sentences in order to allow post-navigation for query refinement. Our clustering algorithm possibly let the clusters overlap at different levels of the hierarchy.
Paolo Ferragina, Antonio Gulli
ICDM2
2004 The Anatomy of SnakeT: A Hierarchical Clustering Engine for Web-Page Snippets
Paolo Ferragina, Antonio Gulli
PKDD2
2004 Experimenting SnakeT: A Hierarchical Clustering Engine for Web-Page Snippets
Paolo Ferragina, Antonio Gulli
PKDD2
2004 Fast PageRank Computation Via a Sparse Linear System (Extended Abstract)
Gianna M. Del Corso, Antonio Gulli, Francesco Romani
WAW2