Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Olena Medelyan

dblp:74/4895 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 2 first-authorArtificial intelligence and machine learning · 5 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 61% Data mining · 30% Knowledge graphs · 9%
Artificial intelligence
3 papers
Information extraction and text analysis · 100%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
keyphrase extraction
0.112009
Human-competitive tagging using automatic keyphrase extraction · EMNLP 2009
Data mining › structured data mining › graph mining
community detection
0.112009
Analysis of community structure in Wikipedia · WWW 2009
Natural language and speech › Information extraction and text analysis › discourse analysis
lexical chains
0.112007
Computing Lexical Chains with Graph Clustering · ACL 2007
Information retrieval
cross-language information retrieval
0.112005
Bootstrapping dictionaries for cross-language information retrieval · SIGIR 2005
Information retrieval
multilingual dictionary construction
0.112005
Bootstrapping dictionaries for cross-language information retrieval · SIGIR 2005
Information retrieval › document processing › document analysis
document annotation
0.012009
Human-competitive tagging using automatic keyphrase extraction · EMNLP 2009
Information retrieval › ranking › graph-based ranking
pagerank
0.012009
Analysis of community structure in Wikipedia · WWW 2009
Information retrieval
ranking
0.012009
Analysis of community structure in Wikipedia · WWW 2009
Graph algorithms and graph theory
graph clustering
0.012007
Computing Lexical Chains with Graph Clustering · ACL 2007

Methods — techniques the papers use, named apart from their topics

automatic keyphrase extraction · 0.2graph clustering · 0.1pagerank · 0.1community detection · 0.1parallel corpora · 0.1cognate mapping · 0.1co-occurrence analysis · 0.1
YearPublicationVenuePosition
2013 Constructing a Focused Taxonomy from a Document Collection
Olena Medelyan, Steve Manion, Jeen Broekstra, Anna Divoli, Anna-Lan Huang, Ian H. Witten
ESWC1
2009 Human-competitive tagging using automatic keyphrase extraction
Olena Medelyan, Eibe Frank, Ian H. Witten
EMNLP1
2009 "All You Can Eat" Ontology-Building: Feeding Wikipedia to Cyc
abstract
In order to achieve genuine web intelligence, building some kind of large general machine-readable conceptual scheme (i.e. ontology) seems inescapable. Yet the past 20 years have shown that manual ontology-building is not practicable. The recent explosion of free user-supplied knowledge on the Web has led to great strides in automatic ontology-building, but quality-control is still a major issue. Ideally one should automatically build onto an already intelligent base. We suggest that the long-running Cyc project is able to assist here. We describe methods used to add 35K new concepts mined from Wikipedia to collections in ResearchCyc entirely automatically. Evaluation with 22 human subjects shows high precision both for the new concepts’ categorization, and their assignment as individuals or collections. Most importantly we show how Cyc itself can be leveraged for ontological quality control by ‘feeding’ it assertions one by one, enabling it to reject those that contradict its other knowledge.
Samuel Sarjant, Catherine Legg, Olena Medelyan
Web Intelligence4
2009 Analysis of community structure in Wikipedia
abstract
We present the results of a community detection analysis of the Wikipedia graph. Distinct communities in Wikipedia contain semantically closely related articles. The central topic of a community can be identified using PageRank. Extracted communities can be organized hierarchically similar to manually created Wikipedia category structure.
Dmitry Lizorkin, Olena Medelyan, Maria P. Grineva
WWW2
2009 Mining meaning from Wikipedia
Olena Medelyan, David N. Milne, Catherine Legg, Ian H. Witten
Int. J. Hum. Comput. Stud.1
2008 Domain-independent automatic keyphrase indexing with small training sets
abstract
Abstract Keyphrases are widely used in both physical and digital libraries as a brief, but precise, summary of documents. They help organize material based on content, provide thematic access, represent search results, and assist with navigation. Manual assignment is expensive because trained human indexers must reach an understanding of the document and select appropriate descriptors according to defined cataloging rules. We propose a new method that enhances automatic keyphrase extraction by using semantic information about terms and phrases gleaned from a domain‐specific thesaurus. The key advantage of the new approach is that it performs well with very little training data. We evaluate it on a large set of manually indexed documents in the domain of agriculture, compare its consistency with a group of six professional indexers, and explore its performance on smaller collections of documents in other domains and of French and Spanish documents.
Olena Medelyan, Ian H. Witten
J. Assoc. Inf. Sci. Technol.1
2007 Computing Lexical Chains with Graph Clustering
Olena Medelyan
ACL1
2006 Language Specific and Topic Focused Web Crawling
Olena Medelyan, Stefan Schulz 0001, Jan Paetzold, Michael Poprat, Kornél G. Markó
LREC1
2006 Mining Domain-Specific Thesauri from Wikipedia: A Case Study
abstract
Domain-specific thesauri are high-cost, high-maintenance, high-value knowledge structures. We show how the classic thesaurus structure of terms and links can be mined automatically from Wikipedia. In a comparison with a professional thesaurus for agriculture we find that Wikipedia contains a substantial proportion of its concepts and semantic relations; furthermore it has impressive coverage of contemporary documents in the domain. Thesauri derived using our techniques capitalize on existing public efforts and tend to reflect contemporary language usage better than their costly, painstakingly-constructed manual counterparts
David N. Milne, Olena Medelyan, Ian H. Witten
Web Intelligence2
2005 Bootstrapping dictionaries for cross-language information retrieval
abstract
The bottleneck for dictionary-based cross-language information retrieval is the lack of comprehensive dictionaries, in particular for many different languages. We here introduce a methodology by which multilingual dictionaries (for Spanish and Swedish) emerge automatically from simple seed lexicons. These seed lexicons are automatically generated, by cognate mapping, from (previously manually constructed) Portuguese and German as well as English sources. Lexical and semantic hypotheses are then validated and new ones iteratively generated by making use of co-occurrence patterns of hypothesized translation synonyms in parallel corpora. We evaluate these newly derived dictionaries on a large medical document collection within a cross-language retrieval setting.
Kornél G. Markó, Stefan Schulz 0001, Olena Medelyan, Udo Hahn
SIGIR3