EDBT 2026 Demo / reviewers in the wild / expert
Olcay Taner Yildiz
dblp:24/1166
· DBLP profile ↗
14ranked-venue papers in the field
3as first author
6since 2021 · last 2023
0000-0001-5838-4615ORCID · reported
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 8Knowledge Engineering, Semantic Web & Information Systems · 3 (1 first)Database Systems & Data Management · 1 (1 first)Data Mining & Knowledge Discovery · 1 (1 first)Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | StarNet: A WordNet Editor InterfaceabstractIn this paper, we introduce StarNet WordNet Editor, an open-source annotation tool designed for natural language processing.It's mainly used for creating and maintaining machinereadable dictionaries like WordNet (Miller, 1995) or domain-specific dictionaries.Word-Net editor provides a user friendly interface and since it is open-source, it is easy to use and develop.Besides English and Turkish WordNet (KeNet) (Bakay et al., 2020), it is also applicable to several languages and their domain specific dictionaries. Oguzhan Kuyrukçu, Ezgi Saniyar, Olcay Taner Yildiz |
GWC | 3 |
| 2023 | A CCGbank for Turkish: From Dependency to CCGabstractIn this paper, we present the building of a CCGbank for Turkish by using standardised dependency corpora.We automatically induce Combinatory Categorial Grammar (CCG) categories for each word token in the Turkish dependency corpora.The CCG induction algorithm we present here is based on the dependency relations that are defined in the latest release of the Universal Dependencies (UD) framework.We aim for an algorithm that can easily be used in all the Turkish treebanks that are annotated in this framework.Therefore, we employ a lexicalist approach in order to make full use of the dependency relations while creating a semantically transparent corpus.We present the treebanks we employed in this study as well as their annotation framework.We introduce the structure of the algorithm we used along with the specific issues that are different from previous studies.Lastly, we show how the results change with this lexical approach in CCGbank for Turkish compared to the previous CCGbank studies in Turkish. Asli Kuzgun, Oguz Kerem Yildiz, Olcay Taner Yildiz |
GWC | 3 |
| 2021 | Creating Domain Dependent Turkish WordNet and SentiNetabstractA WordNet is a thesaurus that has a structured list of words organized depending on their meanings. WordNet represents word senses, all meanings a single lemma may have, the relations between these senses, and their definitions. Another study within the domain of Natural Language Processing is sentiment analysis. With sentiment analysis, data sets can be scored according to the emotion they contain. In the sentiment analysis we did with the data we received on the Tourism WordNet, we performed a domain-specific sentiment analysis study by annotating the data. In this paper, we propose a method to facilitate Natural Language Processing tasks such as sentiment analysis performed in specific domains via creating a specific-domain subset of an original Turkish dictionary. As the preliminary study, we have created a WordNet for the tourism domain with 14,000 words and validated it on simple tasks. Bilge Nas Arican, Merve Özçelik, Deniz Baran Aslan, Elif Sarmis, Selen Parlar, Olcay Taner Yildiz |
GWC | 6 |
| 2021 | Turkish WordNet KeNetabstractÖzge Bakay, Özlem Ergelen, Elif Sarmış, Selin Yıldırım, Bilge Nas Arıcan, Atilla Kocabalcıoğlu, Merve Özçelik, Ezgi Sanıyar, Oğuzhan Kuyrukçu, Begüm Avar, Olcay Taner Yıldız. Proceedings of the 11th Global Wordnet Conference. 2021. Özge Bakay, Özlem Ergelen, Elif Sarmis, Selin Yildirim, Bilge Nas Arican, Atilla Kocabalcioglu, Merve Özçelik, Ezgi Saniyar, Oguzhan Kuyrukçu, Begüm Avar, Olcay Taner Yildiz |
GWC | 11 |
| 2021 | Building the Turkish FrameNetabstractFrameNet (Lowe, 1997;Baker et al., 1998;Fillmore and Atkins, 1998;Johnson et al., 2001) is a computational lexicography project that aims to offer insight into the semantic relationships between predicate and arguments.Having uses in many NLP applications, FrameNet has proven itself as a valuable resource.The main goal of this study is laying the foundation for building a comprehensive and cohesive Turkish FrameNet that is compatible with other resources like PropBank (Kara et al., 2020) or WordNet (Bakay et al., 2019; Büsra Marsan, Neslihan Kara, Merve Özçelik, Bilge Nas Arican, Neslihan Cesur, Asli Kuzgun, Ezgi Saniyar, Oguzhan Kuyrukçu, Olcay Taner Yildiz |
GWC | 9 |
| 2021 | HisNet: A Polarity Lexicon based on WordNet for Emotion AnalysisabstractDictionary-based methods in sentiment analysis have received scholarly attention recently, the most comprehensive examples of which can be found in English.However, many other languages lack polarity dictionaries, or the existing ones are small in size as in the case of Senti-TurkNet, the first and only polarity dictionary in Turkish.Thus, this study aims to extend the content of SentiTurkNet by comparing the two available WordNets in Turkish, namely KeNet and TR-wordnet of BalkaNet.To this end, a current Turkish polarity dictionary has been created relying on 76,825 synsets matching KeNet, where each synset has been annotated with three polarity labels, which are positive, negative and neutral.Meanwhile, the comparison of KeNet and TR-wordnet of BalkaNet has revealed their weaknesses such as the repetition of the same senses, lack of necessary merges of the items belonging to the same synset and the presence of redundant narrower versions of synsets, which are discussed in light of their potential to the improvement of the current lexical databases of Turkish. Merve Özçelik, Bilge Nas Arican, Özge Bakay, Elif Sarmis, Özlem Ergelen, Nilgün Güler Bayezit, Olcay Taner Yildiz |
GWC | 7 |
| 2019 | A Hybrid Approach to Dynamic Enterprise Data PlatformabstractToday, corporations aim to make maximum use of the data produced in business applications. One of the most important goals is to convert the data to the commercial benefit in the fastest way. For this purpose, it is critical to receive the data from source systems, process this data and use it as a support for business decisions. There are many approaches to the proceeding of acquiring, processing and making the data useful. In this study, we took advantage of most of the existing approaches and produced a hybrid solution. This solution can be integrated with new data sources very quickly and reduces the amount of time for data integration, preprocessing, deduplication and entity mapping by using open source software components. Mehmet Selman Sezgin, Ahmet Tugrul Bayrak, Olcay Taner Yildiz |
IEEE BigData | 3 |
| 2019 | English-Turkish Parallel Semantic Annotation of Penn-TreebankabstractThis paper reports our efforts in constructing a sense-labeled English-Turkish parallel corpus using the traditional method of manual tagging.We tagged a pre-built parallel treebank which was translated from the Penn Treebank corpus.This approach allowed us to generate a resource combining syntactic and semantic information.We provide statistics about the corpus itself as well as information regarding its development process. Bilge Nas Arican, Özge Bakay, Begüm Avar, Olcay Taner Yildiz, Özlem Ergelen |
GWC | 4 |
| 2019 | Comparing Sense Categorization Between English PropBank and English WordNetabstractGiven the fact that verbs play a crucial role in language comprehension, this paper presents a study which compares the verb senses in English PropBank with the ones in English WordNet through manual tagging.After analyzing 1554 senses in 1453 distinct verbs, we have found out that while the majority of the senses in Prop-Bank have their one-to-one correspondents in WordNet, a substantial amount of them are differentiated.Furthermore, by analysing the differences between our manually-tagged and an automaticallytagged resource, we claim that manual tagging can help provide better results in sense annotation. Özge Bakay, Begüm Avar, Olcay Taner Yildiz |
GWC | 3 |
| 2013 | Omnivariate Rule Induction Using a Novel Pairwise Statistical TestabstractRule learning algorithms, for example, Ripper, induces univariate rules, that is, a propositional condition in a rule uses only one feature. In this paper, we propose an omnivariate induction of rules where under each condition, both a univariate and a multivariate condition are trained, and the best is chosen according to a novel statistical test. This paper has three main contributions: First, we propose a novel statistical test, the combined 5 × 2 cv t test, to compare two classifiers, which is a variant of the 5 × 2 cv t test and give the connections to other tests as 5 × 2 cv F test and k-fold paired t test. Second, we propose a multivariate version of Ripper, where support vector machine with linear kernel is used to find multivariate linear conditions. Third, we propose an omnivariate version of Ripper, where the model selection is done via the combined 5 × 2 cv t test. Our results indicate that 1) the combined 5 × 2 cv t test has higher power (lower type II error), lower type I error, and higher replicability compared to the 5 × 2 cv t test, 2) omnivariate rules are better in that they choose whichever condition is more accurate, selecting the right model automatically and separately for each condition in a rule. Olcay Taner Yildiz |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Eigenclassifiers for combining correlated classifiers
Aydin Ulas, Olcay Taner Yildiz, Ethem Alpaydin |
Inf. Sci. | 2 |
| 2011 | Model selection in omnivariate decision trees using Structural Risk Minimization
Olcay Taner Yildiz |
Inf. Sci. | 1 |
| 2009 | Incremental construction of classifier and discriminant ensembles
Aydin Ulas, Murat Semerci, Olcay Taner Yildiz, Ethem Alpaydin |
Inf. Sci. | 3 |
| 2005 | Model Selection in Omnivariate Decision Trees
Olcay Taner Yildiz, Ethem Alpaydin |
ECML | 1 |