VLDB 2026 Research / reviewers in the wild / expert
Olcay Taner Yildiz
dblp:24/1166
· DBLP profile ↗
48ranked-venue papers
20as first author
9since 2021 · last 2026
0000-0001-5838-4615ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 14 first-author · 9 since 2021Databases, data management, data science and information retrieval · 14 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Dependency to CCG to Incremental CCG: Approaches to Flexible Word Order in TurkishabstractCombinatory Categorial Grammar (CCG), a lexicalized formalism known for its flexible constituency, is well-suited for modeling headfinal languages with flexible word order like Turkish.Building on Kuzgun et al. (2023), we first develop a Turkish CCG lexicon by automatically inducing categories from a dependency treebank.By leveraging standard and extended operations tailored to Turkish syntax, our parser achieves a robust coverage of 92.5%.Furthermore, we introduce the first (partially) incremental, left-to-right CCG parser for Turkish, designed to facilitate the immediate integration of words into the evolving representation.Finally, we present an example experiment showing that CCG parsers can model psycholinguistic evidence for extra processing costs associated with arguments in noncanonical positions, via the frequency of orderreversing operations.These findings provide evidence that CCG offers a cognitively plausible framework for modeling real-time processing in languages like Turkish. RulesCategory Özge Bakay, Oguz Kerem Yildiz, Rajesh Bhatt, Brian Dillon, Olcay Taner Yildiz |
CoNLL | 5 |
| 2023 | Point of Sale Fraud Detection Methods via Machine LearningabstractRestaurant cash registers frequently experience fraudulent transactions, leading to substantial financial losses for operators. Despite several methods aimed at preventing fraud at the cash register, addressing this issue remains an ongoing concern. In this study, machine learning methods are used to detect fraudulent transactions at the cash register in fast-food restaurants. By using POS logs, transactions in restaurants are recorded and these logs are analyzed to detect fraudulent transactions on an unbalanced dataset. Random forest, XGBoost and LGBM algorithms are used in the study and different resampling techniques (ADASYN etc.) are applied to improve the performance of these algorithms. In addition, it is aimed to find the best parameters with the randomized search method. In conclusion, this study offers a solution for detecting fraudulent transactions at the cash register in fast-food restaurants. The results of the study are promising in its current state. Emre Begen, Ismail Utku Sayan, Ahmet Tugrul Bayrak, Olcay Taner Yildiz |
INISTA | 4 |
| 2023 | StarNet: A WordNet Editor InterfaceabstractIn this paper, we introduce StarNet WordNet Editor, an open-source annotation tool designed for natural language processing.It's mainly used for creating and maintaining machinereadable dictionaries like WordNet (Miller, 1995) or domain-specific dictionaries.Word-Net editor provides a user friendly interface and since it is open-source, it is easy to use and develop.Besides English and Turkish WordNet (KeNet) (Bakay et al., 2020), it is also applicable to several languages and their domain specific dictionaries. Oguzhan Kuyrukçu, Ezgi Saniyar, Olcay Taner Yildiz |
GWC | 3 |
| 2023 | A CCGbank for Turkish: From Dependency to CCGabstractIn this paper, we present the building of a CCGbank for Turkish by using standardised dependency corpora.We automatically induce Combinatory Categorial Grammar (CCG) categories for each word token in the Turkish dependency corpora.The CCG induction algorithm we present here is based on the dependency relations that are defined in the latest release of the Universal Dependencies (UD) framework.We aim for an algorithm that can easily be used in all the Turkish treebanks that are annotated in this framework.Therefore, we employ a lexicalist approach in order to make full use of the dependency relations while creating a semantically transparent corpus.We present the treebanks we employed in this study as well as their annotation framework.We introduce the structure of the algorithm we used along with the specific issues that are different from previous studies.Lastly, we show how the results change with this lexical approach in CCGbank for Turkish compared to the previous CCGbank studies in Turkish. Asli Kuzgun, Oguz Kerem Yildiz, Olcay Taner Yildiz |
GWC | 3 |
| 2022 | A Learning-Based Dependency to Constituency Conversion Algorithm for the Turkish LanguageabstractThis study aims to create the very first dependency-to-constituency conversion algorithm optimised for Turkish language. For this purpose, a state-of-the-art morphologic analyser and a feature-based machine learning model was used. In order to enhance the performance of the conversion algorithm, bootstrap aggregating meta-algorithm was integrated. While creating the conversation algorithm, typological properties of Turkish were carefully considered. A comprehensive and manually annotated UD-style dependency treebank was the input, and constituency trees were the output of the conversion algorithm. A team of linguists manually annotated a set of constituency trees. These manually annotated trees were used as the gold standard to assess the performance of the algorithm. The conversion process yielded more than 8000 constituency trees whose UD-style dependency trees are also available on GitHub. In addition to its contribution to Turkish treebank resources, this study also offers a viable and easy-to-implement conversion algorithm that can be used to generate new constituency treebanks and training data for NLP resources like constituency parsers. Büsra Marsan, Oguz Kerem Yildiz, Asli Kuzgun, Neslihan Cesur, Arife Betül Yenice, Ezgi Saniyar, Oguzhan Kuyrukçu, Bilge Nas Arican, Olcay Taner Yildiz |
LREC | 9 |
| 2021 | Creating Domain Dependent Turkish WordNet and SentiNetabstractA WordNet is a thesaurus that has a structured list of words organized depending on their meanings. WordNet represents word senses, all meanings a single lemma may have, the relations between these senses, and their definitions. Another study within the domain of Natural Language Processing is sentiment analysis. With sentiment analysis, data sets can be scored according to the emotion they contain. In the sentiment analysis we did with the data we received on the Tourism WordNet, we performed a domain-specific sentiment analysis study by annotating the data. In this paper, we propose a method to facilitate Natural Language Processing tasks such as sentiment analysis performed in specific domains via creating a specific-domain subset of an original Turkish dictionary. As the preliminary study, we have created a WordNet for the tourism domain with 14,000 words and validated it on simple tasks. Bilge Nas Arican, Merve Özçelik, Deniz Baran Aslan, Elif Sarmis, Selen Parlar, Olcay Taner Yildiz |
GWC | 6 |
| 2021 | Turkish WordNet KeNetabstractÖzge Bakay, Özlem Ergelen, Elif Sarmış, Selin Yıldırım, Bilge Nas Arıcan, Atilla Kocabalcıoğlu, Merve Özçelik, Ezgi Sanıyar, Oğuzhan Kuyrukçu, Begüm Avar, Olcay Taner Yıldız. Proceedings of the 11th Global Wordnet Conference. 2021. Özge Bakay, Özlem Ergelen, Elif Sarmis, Selin Yildirim, Bilge Nas Arican, Atilla Kocabalcioglu, Merve Özçelik, Ezgi Saniyar, Oguzhan Kuyrukçu, Begüm Avar, Olcay Taner Yildiz |
GWC | 11 |
| 2021 | Building the Turkish FrameNetabstractFrameNet (Lowe, 1997;Baker et al., 1998;Fillmore and Atkins, 1998;Johnson et al., 2001) is a computational lexicography project that aims to offer insight into the semantic relationships between predicate and arguments.Having uses in many NLP applications, FrameNet has proven itself as a valuable resource.The main goal of this study is laying the foundation for building a comprehensive and cohesive Turkish FrameNet that is compatible with other resources like PropBank (Kara et al., 2020) or WordNet (Bakay et al., 2019; Büsra Marsan, Neslihan Kara, Merve Özçelik, Bilge Nas Arican, Neslihan Cesur, Asli Kuzgun, Ezgi Saniyar, Oguzhan Kuyrukçu, Olcay Taner Yildiz |
GWC | 9 |
| 2021 | HisNet: A Polarity Lexicon based on WordNet for Emotion AnalysisabstractDictionary-based methods in sentiment analysis have received scholarly attention recently, the most comprehensive examples of which can be found in English.However, many other languages lack polarity dictionaries, or the existing ones are small in size as in the case of Senti-TurkNet, the first and only polarity dictionary in Turkish.Thus, this study aims to extend the content of SentiTurkNet by comparing the two available WordNets in Turkish, namely KeNet and TR-wordnet of BalkaNet.To this end, a current Turkish polarity dictionary has been created relying on 76,825 synsets matching KeNet, where each synset has been annotated with three polarity labels, which are positive, negative and neutral.Meanwhile, the comparison of KeNet and TR-wordnet of BalkaNet has revealed their weaknesses such as the repetition of the same senses, lack of necessary merges of the items belonging to the same synset and the presence of redundant narrower versions of synsets, which are discussed in light of their potential to the improvement of the current lexical databases of Turkish. Merve Özçelik, Bilge Nas Arican, Özge Bakay, Elif Sarmis, Özlem Ergelen, Nilgün Güler Bayezit, Olcay Taner Yildiz |
GWC | 7 |
| 2020 | TRopBank: Turkish PropBank V2.0abstractIn this paper, we present and explain TRopBank “Turkish PropBank v2.0”. PropBank is a hand-annotated corpus of propositions which is used to obtain the predicate-argument information of a language. Predicate-argument information of a language can help understand semantic roles of arguments. “Turkish PropBank v2.0”, unlike PropBank v1.0, has a much more extensive list of Turkish verbs, with 17.673 verbs in total. Neslihan Kara, Deniz Baran Aslan, Büsra Marsan, Özge Bakay, Koray Ak, Olcay Taner Yildiz |
LREC | 6 |
| 2019 | A Hybrid Approach to Dynamic Enterprise Data PlatformabstractToday, corporations aim to make maximum use of the data produced in business applications. One of the most important goals is to convert the data to the commercial benefit in the fastest way. For this purpose, it is critical to receive the data from source systems, process this data and use it as a support for business decisions. There are many approaches to the proceeding of acquiring, processing and making the data useful. In this study, we took advantage of most of the existing approaches and produced a hybrid solution. This solution can be integrated with new data sources very quickly and reduces the amount of time for data integration, preprocessing, deduplication and entity mapping by using open source software components. Mehmet Selman Sezgin, Ahmet Tugrul Bayrak, Olcay Taner Yildiz |
IEEE BigData | 3 |
| 2019 | English-Turkish Parallel Semantic Annotation of Penn-TreebankabstractThis paper reports our efforts in constructing a sense-labeled English-Turkish parallel corpus using the traditional method of manual tagging.We tagged a pre-built parallel treebank which was translated from the Penn Treebank corpus.This approach allowed us to generate a resource combining syntactic and semantic information.We provide statistics about the corpus itself as well as information regarding its development process. Bilge Nas Arican, Özge Bakay, Begüm Avar, Olcay Taner Yildiz, Özlem Ergelen |
GWC | 4 |
| 2019 | Comparing Sense Categorization Between English PropBank and English WordNetabstractGiven the fact that verbs play a crucial role in language comprehension, this paper presents a study which compares the verb senses in English PropBank with the ones in English WordNet through manual tagging.After analyzing 1554 senses in 1453 distinct verbs, we have found out that while the majority of the senses in Prop-Bank have their one-to-one correspondents in WordNet, a substantial amount of them are differentiated.Furthermore, by analysing the differences between our manually-tagged and an automaticallytagged resource, we claim that manual tagging can help provide better results in sense annotation. Özge Bakay, Begüm Avar, Olcay Taner Yildiz |
GWC | 3 |
| 2018 | AnlamVer: Semantic Model Evaluation Dataset for Turkish - Word Similarity and RelatednessabstractIn this paper, we present AnlamVer, which is a semantic model evaluation dataset for Turkish designed to evaluate word similarity and word relatedness tasks while discriminating those two relations from each other. Our dataset consists of 500 word-pairs annotated by 12 human subjects, and each pair has two distinct scores for similarity and relatedness. Word-pairs are selected to enable the evaluation of distributional semantic models by multiple attributes of words and word-pair relations such as frequency, morphology, concreteness and relation types (e.g., synonymy, antonymy). Our aim is to provide insights to semantic model researchers by evaluating models in multiple attributes. We balance dataset word-pairs by their frequencies to evaluate the robustness of semantic models concerning out-of-vocabulary and rare words problems, which are caused by the rich derivational and inflectional morphology of the Turkish language. Gökhan Ercan, Olcay Taner Yildiz |
COLING | 2 |
| 2018 | Constructing a WordNet for Turkish Using Manual and Automatic AnnotationabstractIn this article, we summarize the methodology and the results of our 2-year-long efforts to construct a comprehensive WordNet for Turkish. In our approach, we mine a dictionary for synonym candidate pairs and manually mark the senses in which the candidates are synonymous. We marked every pair twice by different human annotators. We derive the synsets by finding the connected components of the graph whose edges are synonym senses. We also mined Turkish Wikipedia for hypernym relations among the senses. We analyzed the resulting WordNet to highlight the difficulties brought about by the dictionary construction methods of lexicographers. After splitting the unusually large synsets, we used random walk–based clustering that resulted in a Zipfian distribution of synset sizes. We compared our results to BalkaNet and automatic thesaurus construction methods using variation of information metric. Our Turkish WordNet is available online. Razieh Ehsani, Ercan Solak, Olcay Taner Yildiz |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2016 | Incremental construction of rule ensembles using classifiers produced by different class orderingsabstractIn this paper, we discuss a novel approach to incrementally construct a rule ensemble. The approach constructs an ensemble from a dynamically generated set of rule classifiers. Each classifier in this set is trained by using a different class ordering. We investigate criteria including accuracy, ensemble size, and the role of starting point in the search. Fusion is done by averaging. Using 22 data sets, floating search finds small, accurate ensembles in polynomial time. Olcay Taner Yildiz, Aydin Ulas |
ICPR | 1 |
| 2016 | English-Turkish Parallel Treebank with Morphological Annotations and its Use in Tree-based SMTabstractIn this paper, we report our tree based statistical translation study from English to Turkish. We describe our data generation process and report the initial results of tree-based translation under a simple model. For corpus construction, we used the Penn Treebank in the English side. We manually translated about 5K trees from English to Turkish under grammar constraints with adaptations to accommodate the agglutinative nature of Turkish morphology. We used a permutation model for subtrees together with a word to word mapping. We report BLEU scores under simple choices of inference algorithms. Onur Görgün, Olcay Taner Yildiz, Ercan Solak, Razieh Ehsani |
ICPRAM | 2 |
| 2016 | A novel kernel to predict software defectiveness
Ahmet Okutan, Olcay Taner Yildiz |
J. Syst. Softw. | 2 |
| 2016 | Tree Ensembles on the Induced Discrete SpaceabstractDecision trees are widely used predictive models in machine learning. Recently, K -tree is proposed, where the original discrete feature space is expanded by generating all orderings of values of k discrete attributes and these orderings are used as the new attributes in decision tree induction. Although K -tree performs significantly better than the proper one, their exponential time complexity can prohibit their use. In this brief, we propose K -forest, an extension of random forest, where a subset of features is selected randomly from the induced discrete space. Simulation results on 17 data sets show that the novel ensemble classifier has significantly lower error rate compared with the random forest based on the original feature space. Olcay Taner Yildiz |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Chunking in Turkish with Conditional Random Fields
Olcay Taner Yildiz, Ercan Solak, Razieh Ehsani, Onur Görgün |
CICLing (1) | 1 |
| 2015 | Quadratic programming for class ordering in rule induction
Olcay Taner Yildiz |
Pattern Recognit. Lett. | 1 |
| 2015 | VC-Dimension of Univariate Decision TreesabstractIn this paper, we give and prove the lower bounds of the Vapnik-Chervonenkis (VC)-dimension of the univariate decision tree hypothesis class. The VC-dimension of the univariate decision tree depends on the VC-dimension values of its subtrees and the number of inputs. Via a search algorithm that calculates the VC-dimension of univariate decision trees exhaustively, we show that our VC-dimension bounds are tight for simple trees. To verify that the VC-dimension bounds are useful, we also use them to get VC-generalization bounds for complexity control using structural risk minimization in decision trees, i.e., pruning. Our simulation results show that structural risk minimization pruning using the VC-dimension bounds finds trees that are more accurate as those pruned using cross validation. Olcay Taner Yildiz |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Budding TreesabstractWe propose a new decision tree model, named the budding tree, where a node can be both a leaf and an internal decision node. Each bud node starts as a leaf node, can then grow children, but then later on, if necessary, its children can be pruned. This contrasts with traditional tree construction algorithms that only grows the tree during the training phase, and prunes it in a separate pruning phase. We use a soft tree architecture and show that the tree and its parameters can be trained using gradient-descent. Our experimental results on regression, binary classification, and multi-class classification data sets indicate that our newly proposed model has better performance than traditional trees in terms of accuracy while inducing trees of comparable size. Ozan Irsoy, Olcay Taner Yildiz, Ethem Alpaydin |
ICPR | 2 |
| 2014 | VC-Dimension of Rule SetsabstractIn this paper, we give and prove lower bounds of the VC-dimension of the rule set hypothesis class where the input features are binary or continuous. The VC-dimension of the rule set depends on the VC-dimension values of its rules and the number of inputs. Olcay Taner Yildiz |
ICPR | 1 |
| 2014 | Bilingual Software Requirements Tracing using Vector Space Model
Olcay Taner Yildiz, Ahmet Okutan, Ercan Solak |
ICPRAM | 1 |
| 2014 | Software defect prediction using Bayesian networks
Ahmet Okutan, Olcay Taner Yildiz |
Empir. Softw. Eng. | 2 |
| 2014 | On the feature extraction in discrete space
Olcay Taner Yildiz |
Pattern Recognit. | 1 |
| 2013 | A Novel Regression Method for Software Defect Prediction with Kernel Methods
Ahmet Okutan, Olcay Taner Yildiz |
ICPRAM | 2 |
| 2013 | Omnivariate Rule Induction Using a Novel Pairwise Statistical TestabstractRule learning algorithms, for example, Ripper, induces univariate rules, that is, a propositional condition in a rule uses only one feature. In this paper, we propose an omnivariate induction of rules where under each condition, both a univariate and a multivariate condition are trained, and the best is chosen according to a novel statistical test. This paper has three main contributions: First, we propose a novel statistical test, the combined 5 × 2 cv t test, to compare two classifiers, which is a variant of the 5 × 2 cv t test and give the connections to other tests as 5 × 2 cv F test and k-fold paired t test. Second, we propose a multivariate version of Ripper, where support vector machine with linear kernel is used to find multivariate linear conditions. Third, we propose an omnivariate version of Ripper, where the model selection is done via the combined 5 × 2 cv t test. Our results indicate that 1) the combined 5 × 2 cv t test has higher power (lower type II error), lower type I error, and higher replicability compared to the 5 × 2 cv t test, 2) omnivariate rules are better in that they choose whichever condition is more accurate, selecting the right model automatically and separately for each condition in a rule. Olcay Taner Yildiz |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Searching for the optimal ordering of classes in rule induction
Sezin Ata, Olcay Taner Yildiz |
ICPR | 2 |
| 2012 | Soft decision trees
Ozan Irsoy, Olcay Taner Yildiz, Ethem Alpaydin |
ICPR | 2 |
| 2012 | On the VC-Dimension of Univariate Decision Trees
Olcay Taner Yildiz |
ICPRAM (1) | 1 |
| 2012 | Univariate Decision Tree Induction using Maximum Margin ClassificationabstractIn many pattern recognition applications, first decision trees are used due to their simplicity and easily interpretable nature. In this paper, we propose a new decision tree learning algorithm called univariate margin tree where, for each continuous attribute, the best split is found using convex optimization. Our simulation results on 47 data sets show that the novel margin tree classifier performs at least as good as C4.5 and linear discriminant tree (LDT) with a similar time complexity. For two-class data sets, it generates significantly smaller trees than C4.5 and LDT without sacrificing from accuracy, and generates significantly more accurate trees than C4.5 and LDT for multiclass data sets with one-vs-rest methodology. Olcay Taner Yildiz |
Comput. J. | 1 |
| 2012 | Eigenclassifiers for combining correlated classifiers
Aydin Ulas, Olcay Taner Yildiz, Ethem Alpaydin |
Inf. Sci. | 2 |
| 2012 | Cost-conscious comparison of supervised learning algorithms over multiple data sets
Aydin Ulas, Olcay Taner Yildiz, Ethem Alpaydin |
Pattern Recognit. | 2 |
| 2012 | Design and Analysis of Classifier Learning Experiments in Bioinformatics: Survey and Case StudiesabstractIn many bioinformatics applications, it is important to assess and compare the performances of algorithms trained from data, to be able to draw conclusions unaffected by chance and are therefore significant. Both the design of such experiments and the analysis of the resulting data using statistical tests should be done carefully for the results to carry significance. In this paper, we first review the performance measures used in classification, the basics of experiment design and statistical tests. We then give the results of our survey over 1,500 papers published in the last two years in three bioinformatics journals (including this one). Although the basics of experiment design are well understood, such as resampling instead of using a single training set and the use of different performance metrics instead of error, only 21 percent of the papers use any statistical test for comparison. In the third part, we analyze four different scenarios which we encounter frequently in the bioinformatics literature, discussing the proper statistical methodology as well as showing an example case study for each. With the supplementary software, we hope that the guidelines we discuss will play an important role in future studies. Ozan Irsoy, Olcay Taner Yildiz, Ethem Alpaydin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2011 | Mapping classifiers and datasets
Olcay Taner Yildiz |
Expert Syst. Appl. | 1 |
| 2011 | Model selection in omnivariate decision trees using Structural Risk Minimization
Olcay Taner Yildiz |
Inf. Sci. | 1 |
| 2010 | Feature Extraction from Discrete AttributesabstractIn many pattern recognition applications, first decision trees are used due to their simplicity and easily interpretable nature. In this paper, we extract new features by combining k discrete attributes, where for each subset of size k of the attributes, we generate all orderings of values of those attributes exhaustively. We then apply the usual univariate decision tree classifier using these orderings as the new attributes. Our simulation results on 16 datasets from UCI repository show that the novel decision tree classifier performs better than the proper in terms of error rate and tree complexity. The same idea can also be applied to other univariate rule learning algorithms such as C4.5 Rules and Ripper. Olcay Taner Yildiz |
ICPR | 1 |
| 2009 | An Incremental Model Selection Algorithm Based on Cross-Validation for Finding the Architecture of a Hidden Markov Model on Hand Gesture Data SetsabstractIn a multi-parameter learning problem, besides choosing the architecture of the learner, there is the problem of finding the optimal parameters to get maximum performance. When the number of parameters to be tuned increases, it becomes infeasible to try all the parameter sets, hence we need an automatic mechanism to find the optimum parameter setting using computationally feasible algorithms. In this paper, we define the problem of optimizing the architecture of a Hidden Markov Model (HMM) as a state space search and propose the MSUMO (Model Selection Using Multiple Operators) framework that incrementally modifies the structure and checks for improvement using cross-validation. There are five variants that use forward/backward search, single/multiple operators, and depth-first/breadth-first search. On four hand gesture data sets, we compare the performance of MSUMO with the optimal parameter set found by exhaustive search in terms of expected error and computational complexity. Aydin Ulas, Olcay Taner Yildiz |
ICMLA | 2 |
| 2009 | An Incremental Framework Based on Cross-Validation for Estimating the Architecture of a Multilayer PerceptronabstractWe define the problem of optimizing the architecture of a multilayer perceptron (MLP) as a state space search and propose the MOST (Multiple Operators using Statistical Tests) framework that incrementally modifies the structure and checks for improvement using cross-validation. We consider five variants that implement forward/backward search, using single/multiple operators and searching depth-first/breadth-first. On 44 classification and 30 regression datasets, we exhaustively search for the optimal and evaluate the goodness based on: (1) Order, the accuracy with respect to the optimal and (2) Rank, the computational complexity. We check for the effect of two resampling methods (5 × 2, ten-fold cv), four statistical tests (5 × 2 cv t, ten-fold cv t, Wilcoxon, sign) and two corrections for multiple comparisons (Bonferroni, Holm). We also compare with Dynamic Node Creation (DNC) and Cascade Correlation (CC). Our results show that: (1) On most datasets, networks with few hidden units are optimal, (2) forward searching finds simpler architectures, (3) variants using single node additions (deletions) generally stop early and get stuck in simple (complex) networks, (4) choosing the best of multiple operators finds networks closer to the optimal, (5) MOST variants generally find simpler networks having lower or comparable error rates than DNC and CC. Oya Aran, Olcay Taner Yildiz, Ethem Alpaydin |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2009 | Incremental construction of classifier and discriminant ensembles
Aydin Ulas, Murat Semerci, Olcay Taner Yildiz, Ethem Alpaydin |
Inf. Sci. | 3 |
| 2007 | Parallel univariate decision trees
Olcay Taner Yildiz, Onur Dikmen |
Pattern Recognit. Lett. | 1 |
| 2006 | Ordering and Finding the Best of K > 2 Supervised Learning AlgorithmsabstractGiven a data set and a number of supervised learning algorithms, we would like to find the algorithm with the smallest expected error. Existing pairwise tests allow a comparison of two algorithms only; range tests and ANOVA check whether multiple algorithms have the same expected error and cannot be used for finding the smallest. We propose a methodology, the MultiTest algorithm, whereby we order supervised learning algorithms taking into account 1) the result of pairwise statistical tests on expected error (what the data tells us), and 2) our prior preferences, e.g., due to complexity. We define the problem in graph-theoretic terms and propose an algorithm to find the "best" learning algorithm in terms of these two criteria, or in the more general case, order learning algorithms in terms of their "goodness." Simulation results using five classification algorithms on 30 data sets indicate the utility of the method. Our proposed method can be generalized to regression and other loss functions by using a suitable pairwise test. Olcay Taner Yildiz, Ethem Alpaydin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Model Selection in Omnivariate Decision Trees
Olcay Taner Yildiz, Ethem Alpaydin |
ECML | 1 |
| 2005 | Linear discriminant treesabstractWe discuss and test empirically the effects of six dimensions along which existing decision tree induction algorithms differ. These are: Node type (univariate versus multivariate), branching factor (two or more), grouping of classes into two if the tree is binary, error (impurity) measure, and the methods for minimization to find the best split vector and threshold. We then propose a new decision tree induction method that we name linear discriminant trees (LDT) which uses the best combination of these criteria in terms of accuracy, simplicity and learning time. This tree induction method can be univariate or multivariate. The method has a supervised outer optimization layer for converting a K > 2-class problem into a sequence of two-class problems and each two-class problem is solved analytically using Fisher's Linear Discriminant Analysis (LDA). On twenty datasets from the UCI repository, we compare the linear discriminant trees with the univariate decision tree methods C4.5 and C5.0, multivariate decision tree methods CART, OC1, QUEST, neural trees and LMDT. Our proposed linear discriminant trees learn fast, are accurate, and the trees generated are small. Olcay Taner Yildiz, Ethem Alpaydin |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2001 | Omnivariate decision treesabstractUnivariate decision trees at each decision node consider the value of only one feature leading to axis-aligned splits. In a linear multivariate decision tree, each decision node divides the input space into two with a hyperplane. In a nonlinear multivariate tree, a multilayer perceptron at each node divides the input space arbitrarily, at the expense of increased complexity and higher risk of overfitting. We propose omnivariate trees where the decision node may be univariate, linear, or nonlinear depending on the outcome of comparative statistical tests on accuracy thus matching automatically the complexity of the node with the subproblem defined by the data reaching that node. Such an architecture frees the designer from choosing the appropriate node type, doing model selection automatically at each node. Our simulation results indicate that such a decision tree induction method generalizes better than trees with the same types of nodes everywhere and induces small trees. Olcay Taner Yildiz, Ethem Alpaydin |
IEEE Trans. Neural Networks | 1 |
| 2000 | Linear Discriminant Trees
Olcay Taner Yildiz, Ethem Alpaydin |
ICML | 1 |