VLDB 2026 Research / reviewers in the wild / expert
Günes Erkan
dblp:66/3087
· DBLP profile ↗
9ranked-venue papers
4as first author
0since 2021 · last 2009
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 58% Information extraction and text analysis · 42% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › statistical genetics › genotype-phenotype association
gene-disease association prediction |
0.1 | 1 | 2008 | Identifying gene-disease associations using centrality on a literature mined gene-interaction network · ISMB 2008 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
network analysis |
0.1 | 1 | 2008 | Identifying gene-disease associations using centrality on a literature mined gene-interaction network · ISMB 2008 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.1 | 1 | 2007 | Semi-Supervised Classification for Extracting Protein Interaction Sentences using Dependency Parsing · EMNLP-CoNLL 2007 |
Natural language and speech › Language models and text generation › text summarization
graph-based summarization |
0.1 | 1 | 2006 | LexNet: A Graphical Environment for Graph-Based NLP · ACL 2006 |
Natural language and speech › Language models and text generation › text generation
paraphrase generation |
0.1 | 1 | 2006 | Adding Syntax to Dynamic Programming for Aligning Comparable Texts for the Generation of Paraphrases · ACL 2006 |
Natural language and speech › Information extraction and text analysis › text matching
text alignment |
0.1 | 1 | 2006 | Adding Syntax to Dynamic Programming for Aligning Comparable Texts for the Generation of Paraphrases · ACL 2006 |
Natural language and speech › Language models and text generation
text summarization |
0.1 | 1 | 2006 | LexNet: A Graphical Environment for Graph-Based NLP · ACL 2006 |
Information retrieval › retrieval models
graph-based retrieval |
0.1 | 1 | 2006 | LexNet: A Graphical Environment for Graph-Based NLP · ACL 2006 |
Information retrieval › retrieval models
random walk models |
0.1 | 1 | 2006 | LexNet: A Graphical Environment for Graph-Based NLP · ACL 2006 |
Information retrieval › text summarization
multi-document summarization |
0.0 | 1 | 2004 | LexPageRank: Prestige in Multi-Document Text Summarization · EMNLP 2004 |
Information retrieval
text summarization |
0.0 | 1 | 2004 | LexPageRank: Prestige in Multi-Document Text Summarization · EMNLP 2004 |
Graph algorithms and graph theory › network analysis
graph ranking |
0.0 | 1 | 2004 | LexPageRank: Prestige in Multi-Document Text Summarization · EMNLP 2004 |
Bioinformatics and computational biology
biomedical text mining |
0.0 | 1 | 2007 | Semi-Supervised Classification for Extracting Protein Interaction Sentences using Dependency Parsing · EMNLP-CoNLL 2007 |
Methods — techniques the papers use, named apart from their topics
dependency parsing · 0.3semi-supervised classification · 0.2random walk · 0.1lexrank · 0.1lexpagerank · 0.1graph ranking · 0.1text mining · 0.1support vector machine · 0.1centrality metrics · 0.1syntactic parsing · 0.1dynamic programming · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | Biased LexRank: Passage retrieval using random walks with question-based priors
Jahna Otterbacher, Günes Erkan, Dragomir R. Radev |
Inf. Process. Manag. | 2 |
| 2008 | Identifying gene-disease associations using centrality on a literature mined gene-interaction networkabstractMOTIVATION: Understanding the role of genetics in diseases is one of the most important aims of the biological sciences. The completion of the Human Genome Project has led to a rapid increase in the number of publications in this area. However, the coverage of curated databases that provide information manually extracted from the literature is limited. Another challenge is that determining disease-related genes requires laborious experiments. Therefore, predicting good candidate genes before experimental analysis will save time and effort. We introduce an automatic approach based on text mining and network analysis to predict gene-disease associations. We collected an initial set of known disease-related genes and built an interaction network by automatic literature mining based on dependency parsing and support vector machines. Our hypothesis is that the central genes in this disease-specific network are likely to be related to the disease. We used the degree, eigenvector, betweenness and closeness centrality metrics to rank the genes in the network. RESULTS: The proposed approach can be used to extract known and to infer unknown gene-disease associations. We evaluated the approach for prostate cancer. Eigenvector and degree centrality achieved high accuracy. A total of 95% of the top 20 genes ranked by these methods are confirmed to be related to prostate cancer. On the other hand, betweenness and closeness centrality predicted more genes whose relation to the disease is currently unknown and are candidates for experimental study. AVAILABILITY: A web-based system for browsing the disease-specific gene-interaction networks is available at: http://gin.ncibi.org. Arzucan Özgür, Thuy Vu, Günes Erkan, Dragomir R. Radev |
ISMB | 3 |
| 2008 | Blind men and elephants: What do citation summaries tell us about a research article?abstractAbstract The old Asian legend about the blind men and the elephant comes to mind when looking at how different authors of scientific papers describe a piece of related prior work. It turns out that different citations to the same paper often focus on different aspects of that paper and that neither provides a full description of its full set of contributions. In this article, we will describe our investigation of this phenomenon. We studied citation summaries in the context of research papers in the biomedical domain. A citation summary is the set of citing sentences for a given article and can be used as a surrogate for the actual article in a variety of scenarios. It contains information that was deemed by peers to be important. Our study shows that citation summaries overlap to some extent with the abstracts of the papers and that they also differ from them in that they focus on different aspects of these papers than do the abstracts. In addition to this, co‐cited articles (which are pairs of articles cited by another article) tend to be similar. We show results based on a lexical similarity metric called cohesion to justify our claims. Aaron Elkiss, Siwei Shen, Anthony Fader, Günes Erkan, David J. States, Dragomir R. Radev |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2007 | Semi-Supervised Classification for Extracting Protein Interaction Sentences using Dependency Parsing
Günes Erkan, Arzucan Özgür, Dragomir R. Radev |
EMNLP-CoNLL | 1 |
| 2006 | LexNet: A Graphical Environment for Graph-Based NLPabstractThis interactive presentation describes LexNet, a graphical environment for graph-based NLP developed at the University of Michigan. LexNet includes LexRank (for text summarization), biased LexRank (for passage retrieval), and TUMBL (for binary classification). All tools in the collection are based on random walks on lexical graphs, that is graphs where different NLP objects (e.g., sentences or phrases) are represented as nodes linked by edges proportional to the lexical similarity between the two nodes. We will demonstrate these tools on a variety of NLP tasks including summarization, question answering, and prepositional phrase attachment. Dragomir R. Radev, Günes Erkan, Anthony Fader, Patrick Jordan, Siwei Shen, James P. Sweeney |
ACL | 2 |
| 2006 | Adding Syntax to Dynamic Programming for Aligning Comparable Texts for the Generation of Paraphrases
Siwei Shen, Dragomir R. Radev, Agam Patel, Günes Erkan |
ACL | 4 |
| 2006 | Language Model-Based Document Clustering Using Random Walks
Günes Erkan |
HLT-NAACL | 1 |
| 2004 | LexPageRank: Prestige in Multi-Document Text Summarization
Günes Erkan, Dragomir R. Radev |
EMNLP | 1 |
| 2004 | LexRank: Graph-based Lexical Centrality as Salience in Text SummarizationabstractWe introduce a stochastic graph-based method for computing relative importance of textual units for Natural Language Processing. We test the technique on the problem of Text Summarization (TS). Extractive TS relies on the concept of sentence salience to identify the most important sentences in a document or set of documents. Salience is typically defined in terms of the presence of particular important words or in terms of similarity to a centroid pseudo-sentence. We consider a new approach, LexRank, for computing sentence importance based on the concept of eigenvector centrality in a graph representation of sentences. In this model, a connectivity matrix based on intra-sentence cosine similarity is used as the adjacency matrix of the graph representation of sentences. Our system, based on LexRank ranked in first place in more than one task in the recent DUC 2004 evaluation. In this paper we present a detailed analysis of our approach and apply it to a larger data set including data from earlier DUC evaluations. We discuss several methods to compute centrality using the similarity graph. The results show that degree-based methods (including LexRank) outperform both centroid-based methods and other systems participating in DUC in most of the cases. Furthermore, the LexRank with threshold method outperforms the other degree-based techniques including continuous LexRank. We also show that our approach is quite insensitive to the noise in the data that may result from an imperfect topical clustering of documents. Günes Erkan, Dragomir R. Radev |
J. Artif. Intell. Res. | 1 |