Günes Erkan

dblp:66/3087 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 58% Information extraction and text analysis · 42%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › statistical genetics › genotype-phenotype association
gene-disease association prediction
0.112008
Identifying gene-disease associations using centrality on a literature mined gene-interaction network · ISMB 2008
Bioinformatics and computational biology › network bioinformatics › biological network analysis
network analysis
0.112008
Identifying gene-disease associations using centrality on a literature mined gene-interaction network · ISMB 2008
Natural language and speech › Information extraction and text analysis
relation extraction
0.112007
Semi-Supervised Classification for Extracting Protein Interaction Sentences using Dependency Parsing · EMNLP-CoNLL 2007
Natural language and speech › Language models and text generation › text summarization
graph-based summarization
0.112006
LexNet: A Graphical Environment for Graph-Based NLP · ACL 2006
Natural language and speech › Language models and text generation › text generation
paraphrase generation
0.112006
Adding Syntax to Dynamic Programming for Aligning Comparable Texts for the Generation of Paraphrases · ACL 2006
Natural language and speech › Information extraction and text analysis › text matching
text alignment
0.112006
Adding Syntax to Dynamic Programming for Aligning Comparable Texts for the Generation of Paraphrases · ACL 2006
Natural language and speech › Language models and text generation
text summarization
0.112006
LexNet: A Graphical Environment for Graph-Based NLP · ACL 2006
Information retrieval › retrieval models
graph-based retrieval
0.112006
LexNet: A Graphical Environment for Graph-Based NLP · ACL 2006
Information retrieval › retrieval models
random walk models
0.112006
LexNet: A Graphical Environment for Graph-Based NLP · ACL 2006
Information retrieval › text summarization
multi-document summarization
0.012004
LexPageRank: Prestige in Multi-Document Text Summarization · EMNLP 2004
Information retrieval
text summarization
0.012004
LexPageRank: Prestige in Multi-Document Text Summarization · EMNLP 2004
Graph algorithms and graph theory › network analysis
graph ranking
0.012004
LexPageRank: Prestige in Multi-Document Text Summarization · EMNLP 2004
Bioinformatics and computational biology
biomedical text mining
0.012007
Semi-Supervised Classification for Extracting Protein Interaction Sentences using Dependency Parsing · EMNLP-CoNLL 2007

Methods — techniques the papers use, named apart from their topics

dependency parsing · 0.3semi-supervised classification · 0.2random walk · 0.1lexrank · 0.1lexpagerank · 0.1graph ranking · 0.1text mining · 0.1support vector machine · 0.1centrality metrics · 0.1syntactic parsing · 0.1dynamic programming · 0.1
YearPublicationVenuePosition
2009 Biased LexRank: Passage retrieval using random walks with question-based priors
Jahna Otterbacher, Günes Erkan, Dragomir R. Radev
Inf. Process. Manag.2
2008 Identifying gene-disease associations using centrality on a literature mined gene-interaction network
abstract
MOTIVATION: Understanding the role of genetics in diseases is one of the most important aims of the biological sciences. The completion of the Human Genome Project has led to a rapid increase in the number of publications in this area. However, the coverage of curated databases that provide information manually extracted from the literature is limited. Another challenge is that determining disease-related genes requires laborious experiments. Therefore, predicting good candidate genes before experimental analysis will save time and effort. We introduce an automatic approach based on text mining and network analysis to predict gene-disease associations. We collected an initial set of known disease-related genes and built an interaction network by automatic literature mining based on dependency parsing and support vector machines. Our hypothesis is that the central genes in this disease-specific network are likely to be related to the disease. We used the degree, eigenvector, betweenness and closeness centrality metrics to rank the genes in the network. RESULTS: The proposed approach can be used to extract known and to infer unknown gene-disease associations. We evaluated the approach for prostate cancer. Eigenvector and degree centrality achieved high accuracy. A total of 95% of the top 20 genes ranked by these methods are confirmed to be related to prostate cancer. On the other hand, betweenness and closeness centrality predicted more genes whose relation to the disease is currently unknown and are candidates for experimental study. AVAILABILITY: A web-based system for browsing the disease-specific gene-interaction networks is available at: http://gin.ncibi.org.
Arzucan Özgür, Thuy Vu, Günes Erkan, Dragomir R. Radev
ISMB3
2008 Blind men and elephants: What do citation summaries tell us about a research article?
abstract
Abstract The old Asian legend about the blind men and the elephant comes to mind when looking at how different authors of scientific papers describe a piece of related prior work. It turns out that different citations to the same paper often focus on different aspects of that paper and that neither provides a full description of its full set of contributions. In this article, we will describe our investigation of this phenomenon. We studied citation summaries in the context of research papers in the biomedical domain. A citation summary is the set of citing sentences for a given article and can be used as a surrogate for the actual article in a variety of scenarios. It contains information that was deemed by peers to be important. Our study shows that citation summaries overlap to some extent with the abstracts of the papers and that they also differ from them in that they focus on different aspects of these papers than do the abstracts. In addition to this, co‐cited articles (which are pairs of articles cited by another article) tend to be similar. We show results based on a lexical similarity metric called cohesion to justify our claims.
Aaron Elkiss, Siwei Shen, Anthony Fader, Günes Erkan, David J. States, Dragomir R. Radev
J. Assoc. Inf. Sci. Technol.4
2007 Semi-Supervised Classification for Extracting Protein Interaction Sentences using Dependency Parsing
Günes Erkan, Arzucan Özgür, Dragomir R. Radev
EMNLP-CoNLL1
2006 LexNet: A Graphical Environment for Graph-Based NLP
abstract
This interactive presentation describes LexNet, a graphical environment for graph-based NLP developed at the University of Michigan. LexNet includes LexRank (for text summarization), biased LexRank (for passage retrieval), and TUMBL (for binary classification). All tools in the collection are based on random walks on lexical graphs, that is graphs where different NLP objects (e.g., sentences or phrases) are represented as nodes linked by edges proportional to the lexical similarity between the two nodes. We will demonstrate these tools on a variety of NLP tasks including summarization, question answering, and prepositional phrase attachment.
Dragomir R. Radev, Günes Erkan, Anthony Fader, Patrick Jordan, Siwei Shen, James P. Sweeney
ACL2
2006 Adding Syntax to Dynamic Programming for Aligning Comparable Texts for the Generation of Paraphrases
Siwei Shen, Dragomir R. Radev, Agam Patel, Günes Erkan
ACL4
2006 Language Model-Based Document Clustering Using Random Walks
Günes Erkan
HLT-NAACL1
2004 LexPageRank: Prestige in Multi-Document Text Summarization
Günes Erkan, Dragomir R. Radev
EMNLP1
2004 LexRank: Graph-based Lexical Centrality as Salience in Text Summarization
abstract
We introduce a stochastic graph-based method for computing relative importance of textual units for Natural Language Processing. We test the technique on the problem of Text Summarization (TS). Extractive TS relies on the concept of sentence salience to identify the most important sentences in a document or set of documents. Salience is typically defined in terms of the presence of particular important words or in terms of similarity to a centroid pseudo-sentence. We consider a new approach, LexRank, for computing sentence importance based on the concept of eigenvector centrality in a graph representation of sentences. In this model, a connectivity matrix based on intra-sentence cosine similarity is used as the adjacency matrix of the graph representation of sentences. Our system, based on LexRank ranked in first place in more than one task in the recent DUC 2004 evaluation. In this paper we present a detailed analysis of our approach and apply it to a larger data set including data from earlier DUC evaluations. We discuss several methods to compute centrality using the similarity graph. The results show that degree-based methods (including LexRank) outperform both centroid-based methods and other systems participating in DUC in most of the cases. Furthermore, the LexRank with threshold method outperforms the other degree-based techniques including continuous LexRank. We also show that our approach is quite insensitive to the noise in the data that may result from an imperfect topical clustering of documents.
Günes Erkan, Dragomir R. Radev
J. Artif. Intell. Res.1