Eric Crestan

dblp:39/2237 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 6 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Knowledge graphs · 41% Web and social media mining · 22% Data integration and cleaning · 19%
Artificial intelligence
1 paper
Information extraction and text analysis · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge graphs
knowledge graph construction
0.222011
Web-scale table census and classification · WSDM 2011
Web-scale knowledge extraction from semi-structured tables · WWW 2010
Web and social media mining
web mining
0.112011
Web-scale table census and classification · WSDM 2011
Data integration and cleaning
entity disambiguation
0.112010
Web-scale knowledge extraction from semi-structured tables · WWW 2010
Natural language and speech › Information extraction and text analysis › distributional semantics
distributional similarity
0.112009
Web-Scale Distributional Similarity and Entity Set Expansion · EMNLP 2009
Natural language and speech › Information extraction and text analysis › named entity processing
entity set expansion
0.112009
Web-Scale Distributional Similarity and Entity Set Expansion · EMNLP 2009
Information retrieval › interactive information retrieval
browsing
0.012004
Natural language processing for browse help · SIGIR 2004
Information retrieval
interactive information retrieval
0.012004
Natural language processing for browse help · SIGIR 2004
Query processing and optimization
search space reduction
0.012004
Natural language processing for browse help · SIGIR 2004

Methods — techniques the papers use, named apart from their topics

supervised classification · 0.1feature analysis · 0.1feature engineering · 0.1classification · 0.1named entity recognition · 0.0
YearPublicationVenuePosition
2014 Modelling and Detecting Changes in User Satisfaction
abstract
Informational needs behind queries, that people issue to search engines, are inherently sensitive to external factors such as breaking news, new models of devices, or seasonal changes as 'black Friday'. Mostly these changes happen suddenly and it is natural to suppose that they may cause a shift in user satisfaction with presented old search results and push users to reformulate their queries. For instance, if users issued the query 'CIKM conference' in 2013 they were satisfied with results referring to the page cikm2013.org and this page gets a majority of clicks. However, the confernce site has been changed and the same query issued in 2014 should be linked to the different page cikm2014.fudan.edu.cn. If the link to the fresh page is not among the retrieved results then users will reformulate the query to find desired information.
Julia Kiseleva, Eric Crestan, Riccardo Brigo, Roland Ditte
CIKM2
2011 Web-scale table census and classification
abstract
We report on a census of the types of HTML tables on the Web according to a fine-grained classification taxonomy describing the semantics that they express. For each relational table type, we describe open challenges for extracting from them semantic triples, i.e., knowledge. We also present TabEx, a supervised framework for web-scale HTML table classification and apply it to the task of classifying HTML tables into our taxonomy. We show empirical evidence, through a large-scale experimental analysis over a crawl of the Web, that classification accuracy significantly outperforms several baselines. We present a detailed feature analysis and outline the most salient features for each table type.
Eric Crestan, Patrick Pantel
WSDM1
2010 A fine-grained taxonomy of tables on the web
abstract
We propose a classification taxonomy over a large crawl of HTML tables on the Web, focusing primarily on those tables that express structured knowledge. The taxonomy separates tables into two top-level classes: a) those used for layout purposes, including navigational and formatting; and b) those containing relational knowledge, including listings, attribute/value, matrix, enumeration, and form. We then propose a classification algorithm for automatically detecting a subset of the classes in our taxonomy, namely layout tables and attribute/value tables. We report on the performance of our system over a large sample of manually annotated HTML tables on the Web.
Eric Crestan, Patrick Pantel
CIKM1
2010 Web-scale knowledge extraction from semi-structured tables
abstract
A wealth of knowledge is encoded in the form of tables on the World Wide Web. We propose a classification algorithm and a rich feature set for automatically recognizing layout tables and attribute/value tables. We report the frequencies of these table types over a large analysis of the Web and propose open challenges for extracting from attribute/value tables semantic triples (knowledge). We then describe a solution to a key problem in extracting semantic triples: protagonist detection, i.e., finding the subject of the table that often is not present in the table itself. In 79% of our Web tables, our method finds the correct protagonist in its top three returned candidates.
Eric Crestan, Patrick Pantel
WWW1
2009 Helping editors choose better seed sets for entity set expansion
abstract
Sets of named entities are used heavily at commercial search engines such as Google, Yahoo and Bing. Acquiring sets of entities typically consists of combining semi-supervised expansion algorithms with manual cleaning of the resulting expanded sets. In this paper, we study the effects of different seed sets in a state-of-the-art semi-supervised expansion system and show a tremendous variation in expansion performance depending on the choice of seeds. We further show that human editors, in general, provide very bad seed sets, which perform well-below the average random seed set. We identify three factors of seed set composition, namely prototypicality, ambiguity and coverage, and we investigate their effects on expansion performance. Finally, we propose various automatic systems for improving editor-generated seed sets, which seek to remove ambiguous and other error-prone seed instances. An extensive experimental analysis shows that expansion quality, measured in R-precision, can be improved on average by a maximum of 46% by removing the right seeds from a seed set. Our automatic methods outperform the human editors seed sets and on average improve expansion performance by up to 34% over the original seed sets.
Vishnu Vyas, Patrick Pantel, Eric Crestan
CIKM3
2009 Web-Scale Distributional Similarity and Entity Set Expansion
Patrick Pantel, Eric Crestan, Arkady Borkovsky, Ana-Maria Popescu, Vishnu Vyas
EMNLP2
2004 Browsing Help for a Faster Retrieval
Eric Crestan, Claude de Loupy
COLING1
2004 Natural language processing for browse help
abstract
In this paper, we will present three "browsing" systems that should save user's time. The first uses named entities and gives a way to reduce search space. By using a information visualization system, the user can comprehend more easily the content of a corpus or a document. Named entities are highlighted for quick reading, temporal and geographic representation gives a global view of the result of a query. All these browse and search helps seem to be very useful. Nevertheless, an evaluation would give more practical results.
Eric Crestan, Claude de Loupy
SIGIR1