Olcay Taner Yildiz

dblp:24/1166 · DBLP profile ↗
← Back
14ranked-venue papers in the field
3as first author
6since 2021 · last 2023
0000-0001-5838-4615ORCID · reported

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 8Knowledge Engineering, Semantic Web & Information Systems · 3 (1 first)Database Systems & Data Management · 1 (1 first)Data Mining & Knowledge Discovery · 1 (1 first)Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2023 StarNet: A WordNet Editor Interface
abstract
In this paper, we introduce StarNet WordNet Editor, an open-source annotation tool designed for natural language processing.It's mainly used for creating and maintaining machinereadable dictionaries like WordNet (Miller, 1995) or domain-specific dictionaries.Word-Net editor provides a user friendly interface and since it is open-source, it is easy to use and develop.Besides English and Turkish WordNet (KeNet) (Bakay et al., 2020), it is also applicable to several languages and their domain specific dictionaries.
Oguzhan Kuyrukçu, Ezgi Saniyar, Olcay Taner Yildiz
GWC3
2023 A CCGbank for Turkish: From Dependency to CCG
abstract
In this paper, we present the building of a CCGbank for Turkish by using standardised dependency corpora.We automatically induce Combinatory Categorial Grammar (CCG) categories for each word token in the Turkish dependency corpora.The CCG induction algorithm we present here is based on the dependency relations that are defined in the latest release of the Universal Dependencies (UD) framework.We aim for an algorithm that can easily be used in all the Turkish treebanks that are annotated in this framework.Therefore, we employ a lexicalist approach in order to make full use of the dependency relations while creating a semantically transparent corpus.We present the treebanks we employed in this study as well as their annotation framework.We introduce the structure of the algorithm we used along with the specific issues that are different from previous studies.Lastly, we show how the results change with this lexical approach in CCGbank for Turkish compared to the previous CCGbank studies in Turkish.
Asli Kuzgun, Oguz Kerem Yildiz, Olcay Taner Yildiz
GWC3
2021 Creating Domain Dependent Turkish WordNet and SentiNet
abstract
A WordNet is a thesaurus that has a structured list of words organized depending on their meanings. WordNet represents word senses, all meanings a single lemma may have, the relations between these senses, and their definitions. Another study within the domain of Natural Language Processing is sentiment analysis. With sentiment analysis, data sets can be scored according to the emotion they contain. In the sentiment analysis we did with the data we received on the Tourism WordNet, we performed a domain-specific sentiment analysis study by annotating the data. In this paper, we propose a method to facilitate Natural Language Processing tasks such as sentiment analysis performed in specific domains via creating a specific-domain subset of an original Turkish dictionary. As the preliminary study, we have created a WordNet for the tourism domain with 14,000 words and validated it on simple tasks.
Bilge Nas Arican, Merve Özçelik, Deniz Baran Aslan, Elif Sarmis, Selen Parlar, Olcay Taner Yildiz
GWC6
2021 Turkish WordNet KeNet
abstract
Özge Bakay, Özlem Ergelen, Elif Sarmış, Selin Yıldırım, Bilge Nas Arıcan, Atilla Kocabalcıoğlu, Merve Özçelik, Ezgi Sanıyar, Oğuzhan Kuyrukçu, Begüm Avar, Olcay Taner Yıldız. Proceedings of the 11th Global Wordnet Conference. 2021.
Özge Bakay, Özlem Ergelen, Elif Sarmis, Selin Yildirim, Bilge Nas Arican, Atilla Kocabalcioglu, Merve Özçelik, Ezgi Saniyar, Oguzhan Kuyrukçu, Begüm Avar, Olcay Taner Yildiz
GWC11
2021 Building the Turkish FrameNet
abstract
FrameNet (Lowe, 1997;Baker et al., 1998;Fillmore and Atkins, 1998;Johnson et al., 2001) is a computational lexicography project that aims to offer insight into the semantic relationships between predicate and arguments.Having uses in many NLP applications, FrameNet has proven itself as a valuable resource.The main goal of this study is laying the foundation for building a comprehensive and cohesive Turkish FrameNet that is compatible with other resources like PropBank (Kara et al., 2020) or WordNet (Bakay et al., 2019;
Büsra Marsan, Neslihan Kara, Merve Özçelik, Bilge Nas Arican, Neslihan Cesur, Asli Kuzgun, Ezgi Saniyar, Oguzhan Kuyrukçu, Olcay Taner Yildiz
GWC9
2021 HisNet: A Polarity Lexicon based on WordNet for Emotion Analysis
abstract
Dictionary-based methods in sentiment analysis have received scholarly attention recently, the most comprehensive examples of which can be found in English.However, many other languages lack polarity dictionaries, or the existing ones are small in size as in the case of Senti-TurkNet, the first and only polarity dictionary in Turkish.Thus, this study aims to extend the content of SentiTurkNet by comparing the two available WordNets in Turkish, namely KeNet and TR-wordnet of BalkaNet.To this end, a current Turkish polarity dictionary has been created relying on 76,825 synsets matching KeNet, where each synset has been annotated with three polarity labels, which are positive, negative and neutral.Meanwhile, the comparison of KeNet and TR-wordnet of BalkaNet has revealed their weaknesses such as the repetition of the same senses, lack of necessary merges of the items belonging to the same synset and the presence of redundant narrower versions of synsets, which are discussed in light of their potential to the improvement of the current lexical databases of Turkish.
Merve Özçelik, Bilge Nas Arican, Özge Bakay, Elif Sarmis, Özlem Ergelen, Nilgün Güler Bayezit, Olcay Taner Yildiz
GWC7
2019 A Hybrid Approach to Dynamic Enterprise Data Platform
abstract
Today, corporations aim to make maximum use of the data produced in business applications. One of the most important goals is to convert the data to the commercial benefit in the fastest way. For this purpose, it is critical to receive the data from source systems, process this data and use it as a support for business decisions. There are many approaches to the proceeding of acquiring, processing and making the data useful. In this study, we took advantage of most of the existing approaches and produced a hybrid solution. This solution can be integrated with new data sources very quickly and reduces the amount of time for data integration, preprocessing, deduplication and entity mapping by using open source software components.
Mehmet Selman Sezgin, Ahmet Tugrul Bayrak, Olcay Taner Yildiz
IEEE BigData3
2019 English-Turkish Parallel Semantic Annotation of Penn-Treebank
abstract
This paper reports our efforts in constructing a sense-labeled English-Turkish parallel corpus using the traditional method of manual tagging.We tagged a pre-built parallel treebank which was translated from the Penn Treebank corpus.This approach allowed us to generate a resource combining syntactic and semantic information.We provide statistics about the corpus itself as well as information regarding its development process.
Bilge Nas Arican, Özge Bakay, Begüm Avar, Olcay Taner Yildiz, Özlem Ergelen
GWC4
2019 Comparing Sense Categorization Between English PropBank and English WordNet
abstract
Given the fact that verbs play a crucial role in language comprehension, this paper presents a study which compares the verb senses in English PropBank with the ones in English WordNet through manual tagging.After analyzing 1554 senses in 1453 distinct verbs, we have found out that while the majority of the senses in Prop-Bank have their one-to-one correspondents in WordNet, a substantial amount of them are differentiated.Furthermore, by analysing the differences between our manually-tagged and an automaticallytagged resource, we claim that manual tagging can help provide better results in sense annotation.
Özge Bakay, Begüm Avar, Olcay Taner Yildiz
GWC3
2013 Omnivariate Rule Induction Using a Novel Pairwise Statistical Test
abstract
Rule learning algorithms, for example, Ripper, induces univariate rules, that is, a propositional condition in a rule uses only one feature. In this paper, we propose an omnivariate induction of rules where under each condition, both a univariate and a multivariate condition are trained, and the best is chosen according to a novel statistical test. This paper has three main contributions: First, we propose a novel statistical test, the combined 5 × 2 cv t test, to compare two classifiers, which is a variant of the 5 × 2 cv t test and give the connections to other tests as 5 × 2 cv F test and k-fold paired t test. Second, we propose a multivariate version of Ripper, where support vector machine with linear kernel is used to find multivariate linear conditions. Third, we propose an omnivariate version of Ripper, where the model selection is done via the combined 5 × 2 cv t test. Our results indicate that 1) the combined 5 × 2 cv t test has higher power (lower type II error), lower type I error, and higher replicability compared to the 5 × 2 cv t test, 2) omnivariate rules are better in that they choose whichever condition is more accurate, selecting the right model automatically and separately for each condition in a rule.
Olcay Taner Yildiz
IEEE Trans. Knowl. Data Eng.1
2012 Eigenclassifiers for combining correlated classifiers
Aydin Ulas, Olcay Taner Yildiz, Ethem Alpaydin
Inf. Sci.2
2011 Model selection in omnivariate decision trees using Structural Risk Minimization
Olcay Taner Yildiz
Inf. Sci.1
2009 Incremental construction of classifier and discriminant ensembles
Aydin Ulas, Murat Semerci, Olcay Taner Yildiz, Ethem Alpaydin
Inf. Sci.3
2005 Model Selection in Omnivariate Decision Trees
Olcay Taner Yildiz, Ethem Alpaydin
ECML1