VLDB 2026 Research / reviewers in the wild / expert
Murat Can Ganiz
dblp:85/1397
· DBLP profile ↗
26ranked-venue papers
5as first author
5since 2021 · last 2022
0000-0001-8338-991XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-authorSecurity and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Traditional Machine Learning and Deep Learning-based Text Classification for Turkish Law Documents using Transformers and Domain AdaptationabstractNatural Language Processing (NLP) is an interdisciplinary field between linguistics and computer science. Its main aim is to process natural (human) language using computer programs. Text classification is one of the main tasks of this field, and they are widely used in many different applications such as spam filtering, sentiment analysis, and document categorization. Nonetheless, there is only very little text classification work in the law domain and even less for the Turkish language. This may be attributed to the complexity within the domain. The length, complexity of documents, and use of extensive technical jargon are some of the reasons that distinguish this domain from others. Similar to the medical domain, understanding these documents requires extensive specialization. Another reason can be the scarcity of publicly available datasets. In this study, we compile sizeable unsupervised and supervised datasets from publicly available sources and experiment with several classification algorithms ranging from traditional classifiers to much more complicated deep learning and transformer-based models along with different text representations. We focus on classifying Court of Cassation decisions for their crime labels. Interestingly, the majority of the models we experiment with could be able to obtain good results. This suggests that although understanding the documents in the legal domain is complicated and requires expertise from humans, it may be relatively easier for machine learning models despite the extensive presence of the technical terms. This seems to be especially the case for transformer-based pre-trained neural language models which can be adapted to the law domain, showing high potential for future real-world applications. Onur Akça, Giyaseddin Bayrak, Abdul Majeed Issifu, Murat Can Ganiz |
INISTA | 4 |
| 2022 | Biomedical Named Entity Recognition Using Transformers with biLSTM + CRF and Graph Convolutional Neural NetworksabstractOne of the applications of Natural Language Processing (NLP) is to process free text data for extracting information. Information extraction has various forms like Named Entity Recognition (NER) for detecting the named entities in the free text. Biomedical named-entity extraction task is about extracting named entities like drugs, diseases, organs, etc. from texts in medical domain. In our study, we improve commonly used models in this domain, such as biLSTM+CRF model, using transformer based language models like BERT and its domain-specific variant BioBERT in the embedding layer. We conduct several experiments on several different benchmark biomedical datasets using a variety of combination of models and embeddings such as BioBERT+biLSTM+CRF, BERT+biLSTM+CRF, Fasttext+biLSTM+CRF, and Graph Convolutional Networks. Our results show a quite visible, 4% to 13%, improvements when baseline biLSTM+CRF model is initialized with pretrained language models such as BERT and especially with domain specific one like BioBERT on several datasets. Gökberk Çelkmasat, Muhammed Enes Aktürk, Yunus Emre Ertunç, Abdul Majeed Issifu, Murat Can Ganiz |
INISTA | 5 |
| 2022 | Individual Stock Price Prediction by Using KAP and Twitter Sentiments with Machine Learning for BIST30abstractIn this study we used machine learning models for predicting individual stock price and volume changes using sentiments from public disclosures and tweets. Public Disclosure Platform (KAP) is the mandated regulatory platform for disclosing news about companies listed in Borsa Istanbul. Investors in Borsa Istanbul use Twitter to express their sentiments about stocks. By combining people’s sentiment on Twitter and companies’ disclosures, our prediction model predicts the volume and price changes of individual company stocks listed in BIST30. Financial data regarding market conditions consisting of daily price changes of BIST30, DJI, USD, and Gold per Ounce are also added to enhance the prediction accuracy of the model. Our model achieves an maximum of 80% individual stock price prediction accuracy for companies with high social media presence and public disclosure count. We also achieve 74.7% mean volume prediction accuracy across all BIST30 companies. Muhlis Sariyer, Ahmet Akil, Feyza Nur Bulgurcu, Fatih Emin Öge, Murat Can Ganiz |
INISTA | 5 |
| 2021 | Log and Execution Trace Analytics SystemabstractLog files are available on every computer system. They automatically record important run time events of operating systems or software applications. They are mainly used to find the root cause of failures. Analyzing such log files allows us to detect anomalies, problems and improve the system. Since the log files are usually unstructured or semi-structured, the important task of log analysis is to parse usually immense amount of log message strings into the human readable and actionable reports. In this paper, we propose an implementation of a machine learning based log parsing system using Named Entity Recognition which is the process of identifying and categorizing entities in the text. Our approach makes use of Bidirectional Encoder Representations from Transformers (BERT) to extract entities. The paper reports the results of experiments on two benchmark; macOS and Linux OS datasets. Nazrin Abbasli, Murat Can Ganiz |
INISTA | 2 |
| 2021 | Descriptive and Prescriptive Analysis of Construction Site Incidents Using Decision Tree Classification and Association Rule MiningabstractLearning from previous incidents is essential for preventing future incidents and taking the necessary precautions. We analyze construction site incidents by employing data science process and machine learning algorithms such as decision trees and Apriori. Patterns that are extracted using machine learning algorithms provides interesting insights on the causes of incidents and their relations with other factors. The dataset we use is a novel dataset containing hundreds of construction site incidents between 2014 and 2020 from an international construction company. The data consist of a wide range of features such as activity during incident, incident condition, hazard source, incident severity, location, and time. The decision tree is used in a descriptive analytics setting to extract patterns in our dataset. Additionally, the Apriori algorithm is employed to extract patterns in the form of frequent itemsets and association rules. The patterns we extract using machine learning algorithms shed light on associations between different factors and different types of incidents. One of the interesting results of our study is that the patterns extracted from a supervised classifier in a descriptive analytics setting collides with the patterns extracted using the unsupervised machine learning algorithm of Apriori. The generated rules can be used for informing the health and safety experts by developing a decision support mechanism for taking necessary precautions for minimizing the risk of different types of incidents. Özgür Ugur, Ali Atilla Arisoy, Murat Can Ganiz, Berkay Bolac |
INISTA | 3 |
| 2019 | Effects of Positivization on the Paragraph Vector ModelabstractNatural language processing (NLP) is an important field of Artificial Intelligence. One of the fundamental problems in NLP is to create vector (distributed) representations of words so that vectors of words that have similar meaning lie closer in space. One of the most popular algorithms for creating these representations are word embedding models such as word2vec and fastText. Similarly the paragraph vector model (doc2vec) is used to create distributed representations of documents while simultaneously creating distributed representations for the words in these documents. These models create a dense, and low dimensional (usually in the low hundreds) vector representations which may include negative values. In this study we focus on these negative values and introduce a family of regularization methods in which document, word and/or context vectors of the paragraph vector model are forced to have only positive components. We measure its effects on several tasks; text classification, semantic similarity, and analogy tasks. Although positivization greatly increases the sparsity of the word embeddings, and should be expected to result in a loss of information, our results show that there is almost no reduction in the performance of the regularized embeddings in these tasks. We also observe an increase in the classification accuracy in one case. We foresee that these approaches can be beneficial in machine learning systems which require non-negative vectors. Aydin Gerek, Mehmet Can Yüney, Erencan Erkaya, Murat Can Ganiz |
INISTA | 4 |
| 2019 | Diffused Label Propagation based Transductive Classification Algorithm for Word Sense DisambiguationabstractA major natural language processing problem, word sense disambiguation is the task of identifying the correct sense of a polysemous word based on its context. In terms of machine learning, this can be considered as a supervised classification problem. A better alternative can be the use of semi-supervised classifiers since labeled data is usually scarce yet we can access large quantities of unlabeled textual data. We propose an improvement to Label Propagation which is a well-known transductive classification algorithm for word sense disambiguation. Our approach make use of a semantic diffusion kernel. We name this new algorithm as diffused label propagation algorithm (DILP). We evaluate our proposed algorithm with experiments utilizing various sizes of training sets of disambiguated corpora. With these experiments we try to answer the following questions: 1. Does our algorithm with semantic kernel formulation yield higher classification performance than the popular kernels? 2. Under which conditions does a kernel design perform better than others? 3. What kind of regularization methods result with better performance? Our experiments demonstrate that our approach can outperform baseline in terms of accuracy in several conditions. Gökhan Kocaman, Bilge Sipal, Aydin Gerek, Berna Altinel, Murat Can Ganiz |
INISTA | 5 |
| 2019 | Waste Not: Meta-Embedding of Word and Context Vectors
Selin Degirmenci, Aydin Gerek, Murat Can Ganiz |
NLDB | 3 |
| 2018 | Semantic text classification: A survey of past and recent advances
Berna Altinel, Murat Can Ganiz |
Inf. Process. Manag. | 2 |
| 2017 | A Feature Based Simple Machine Learning Approach with Word Embeddings to Named Entity Recognition on Tweets
Mete Taspinar, Murat Can Ganiz, Tankut Acarman |
NLDB | 2 |
| 2017 | Instance labeling in semi-supervised learning with meaning values of words
Berna Altinel, Murat Can Ganiz, Banu Diri |
Eng. Appl. Artif. Intell. | 2 |
| 2016 | Semi-supervised learning using higher-order co-occurrence paths to overcome the complexity of data representationabstractWe present a novel approach to semi-supervised learning for text classification based on the higher-order co-occurrence paths of words. We name the proposed method as Semi-Supervised Semantic Higher-Order Smoothing (S3HOS). The S3HOS is built on a tri-partite graph based data representation of labeled and unlabeled documents that allows semantics in higher-order co-occurrence paths between terms (words) to be exploited. There are several graph-based techniques proposed in the literature to diffuse class labels from labeled documents to the unlabeled documents. In this study we propose a different and natural way of estimating class conditional probabilities for the terms in unlabeled documents without need to label the documents first. The proposed approach allows estimating class conditional probabilities for the terms in unlabeled documents and improve the estimation of terms in the labeled documents at the same time. We experimentally show that S3HOS can highly improve the parameter estimation and hence increase the classification accuracy particularly when the amount of the labeled data is scarce but unlabeled data is plentiful. Murat Can Ganiz |
SMC | 1 |
| 2016 | Helmholtz principle based supervised and unsupervised feature selection methods for text miningabstractOne of the important problems in text classification is the high dimensionality of the feature space. Feature selection methods are used to reduce the dimensionality of the feature space by selecting the most valuable features for classification. Apart from reducing the dimensionality, feature selection methods have potential to improve text classifiers’ performance both in terms of accuracy and time. Furthermore, it helps to build simpler and as a result more comprehensible models. In this study we propose new methods for feature selection from textual data, called Meaning Based Feature Selection (MBFS) which is based on the Helmholtz principle from the Gestalt theory of human perception which is used in image processing. The proposed approaches are extensively evaluated by their effect on the classification performance of two well-known classifiers on several datasets and compared with several feature selection algorithms commonly used in text mining. Our results demonstrate the value of the MBFS methods in terms of classification accuracy and execution time. Melike Tutkan, Murat Can Ganiz, Selim Akyokus |
Inf. Process. Manag. | 2 |
| 2016 | A new hybrid semi-supervised algorithm for text classification with class-based semantics
Berna Altinel, Murat Can Ganiz |
Knowl. Based Syst. | 2 |
| 2015 | A novel classifier based on meaning for text classificationabstractText classification is one of the key methods used in text mining. Generally, traditional classification algorithms from machine learning field are used in text classification. These algorithms are primarily designed for structured data. In this paper, we propose a new classifier for textual data, called Supervised Meaning Classifier (SMC). The new SMC classifier uses meaning measure, which is based on Helmholtz principle from Gestalt Theory. In SMC, meaningfulness of terms in the context of classes are calculated and used for classification of a document. Experiment results show that new SMC classifier outperforms traditional classifiers of Multinomial Naïve Bayes (MNB) and Support Vector Machine (SVM) especially when the training data limited. Murat Can Ganiz, Melike Tutkan, Selim Akyokus |
INISTA | 1 |
| 2015 | Evaluation of classification models for language processingabstractNaïve Bayes is a commonly used algorithm in text categorization because of its easy implementation and low complexity. Naïve Bayes has mainly two event models used for text categorization which are multivariate Bernoulli and multinomial models. A very large number of studies choose multinomial model and Laplace smoothing just based on the assumption that it performs better than multivariate model under almost any conditions. This study aims to shed some light into this widely adopted assumption by analyzing Naïve Bayes event models and smoothing methods from a different perspective. To clarify the difference between events models of Naïve Bayes, their classification performance are compared on different languages - English and Turkish - datasets. Results of our extensive experiments demonstrate that superior performance of multinomial model does not observed all the time. On the other hand, multivariate Bernoulli model can perform well when combined with an appropriate smoothing method under different training data size conditions. Zeynep Hilal Kilimci, Murat Can Ganiz |
INISTA | 2 |
| 2015 | A corpus-based semantic kernel for text classification by using meaning values of terms
Berna Altinel, Murat Can Ganiz, Banu Diri |
Eng. Appl. Artif. Intell. | 2 |
| 2015 | A novel semantic smoothing kernel for text classification with class-based weighting
Berna Altinel, Banu Diri, Murat Can Ganiz |
Knowl. Based Syst. | 3 |
| 2014 | A simple semantic kernel approach for SVM using higher-order pathsabstractThe bag of words (BOW) representation of documents is very common in text classification systems. However, the BOW approach ignores the position of the words in the document and more importantly, the semantic relations between the words. In this study, we present a simple semantic kernel for Support Vector Machines (SVM) algorithm. This kernel uses higher-order relations between terms in order to incorporate semantic information into the SVM. This is an easy to implement algorithm which forms a basis for future improvements. We perform a serious of experiments on different well known textual datasets. Experiment results show that classification performance improves over the traditional kernels used in SVM such as linear kernel which is commonly used in text classification. Berna Altinel, Murat Can Ganiz, Banu Diri |
INISTA | 2 |
| 2014 | Higher-Order Smoothing: A Novel Semantic Smoothing Method for Text Classification
Mitat Poyraz, Zeynep Hilal Kilimci, Murat Can Ganiz |
J. Comput. Sci. Technol. | 3 |
| 2012 | A Novel Semantic Smoothing Method Based on Higher Order Paths for Text ClassificationabstractIt has been shown that Latent Semantic Indexing (LSI) takes advantage of implicit higher-order (or latent) structure in the association of terms and documents. Higher order relations in LSI capture "latent semantics". Inspired by this, a novel Bayesian framework for classification named Higher Order Naïve Bayes (HONB), which can explicitly make use of these higher-order relations, has been introduced previously. We present a novel semantic smoothing method named Higher Order Smoothing (HOS) for the Naive Bayes algorithm. HOS is built on a similar graph based data representation of HONB which allows semantics in higher order paths to be exploited. Additionally, we take the concept one step further in HOS and exploited the relationships between instances of different classes in order to improve the parameter estimation when dealing with insufficient labeled data. As a result, we have not only been able to move beyond instance boundaries, but also class boundaries to exploit the latent information in higher-order paths. The results of our extensive experiments demonstrate the value of HOS on several benchmark datasets. Mitat Poyraz, Zeynep Hilal Kilimci, Murat Can Ganiz |
ICDM | 3 |
| 2012 | Discrete-Time Hopfield Neural Network Based Text Clustering Algorithm
Zekeriya Uykan, Murat Can Ganiz, Çagla Sahinli |
ICONIP (1) | 2 |
| 2011 | Higher Order Naïve Bayes: A Novel Non-IID Approach to Text ClassificationabstractThe underlying assumption in traditional machine learning algorithms is that instances are Independent and Identically Distributed (IID). These critical independence assumptions made in traditional machine learning algorithms prevent them from going beyond instance boundaries to exploit latent relations between features. In this paper, we develop a general approach to supervised learning by leveraging higher order dependencies between features. We introduce a novel Bayesian framework for classification termed Higher Order Naïve Bayes (HONB). Unlike approaches that assume data instances are independent, HONB leverages higher order relations between features across different instances. The approach is validated in the classification domain on widely used benchmark data sets. Results obtained on several benchmark text corpora demonstrate that higher order approaches achieve significant improvements in classification accuracy over the baseline methods, especially when training data is scarce. A complexity analysis also reveals that the space and time complexity of HONB compare favorably with existing approaches. Murat Can Ganiz, Cibin George, William M. Pottenger |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2009 | Leveraging Higher Order Dependencies between Features for Text Classification
Murat Can Ganiz, Nikita I. Lytkin, William M. Pottenger |
ECML/PKDD (1) | 1 |
| 2007 | Mining Higher-Order Association Rules from Distributed Named Entity DatabasesabstractThe burgeoning amount of textual data in distributed sources combined with the obstacles involved in creating and maintaining central repositories motivates the need for effective distributed information extraction and mining techniques. Recently, as the need to mine patterns across distributed databases has grown, Distributed Association Rule Mining (D-ARM) algorithms have been developed. These algorithms, however, assume that the databases are either horizontally or vertically distributed. In the special case of databases populated from information extracted from textual data, existing D-ARM algorithms cannot discover rules based on higher-order associations between items in distributed textual documents that are neither vertically nor horizontally distributed, but rather a hybrid of the two. In this article we present D-HOTM, a framework for Distributed Higher Order Text Mining. Unlike existing algorithms, D-HOTM requires neither full knowledge of the global schema nor that the distribution of data be horizontal or vertical. D-HOTM discovers rules based on higher-order associations between distributed database records containing the extracted entities. In this paper, two approaches to the definition and discovery of higher order itemsets are presented. The implementation of D-HOTM is based on the TMI [20] and tested on a cluster at the National Center for Supercomputing Applications (NCSA). Results on a real-world dataset from the Richmond, VA police department demonstrate the performance and relevance of D-HOTM in law enforcement and homeland defense. Shenzhi Li, Christopher D. Janneck, Aditya P. Belapurkar, Murat Can Ganiz, Xiaoning Yang, Mark Dilsizian, Tianhao Wu 0004, John M. Bright, William M. Pottenger |
ISI | 4 |
| 2006 | Detection of Interdomain Routing Anomalies Based on Higher-Order Path AnalysisabstractAnomalous interdomain border gateway protocol (BGP) events including misconfigurations, attacks and large-scale power failures often affect the global routing infrastructure. Thus, the ability to detect and categorize such events is extremely useful. In this article we present a novel anomaly detection technique for BGP that distinguishes between different anomalies in BGP traffic. This technique is termed higher order path analysis (HOPA) and focuses on the discovery of patterns in higher order paths in supervised learning datasets. Our results demonstrate that not only worm events but also different types of worms as well as blackout events are cleanly separable and can be classified in real time based on our incremental approach. This novel approach to supervised learning has potential applications in cybersecurity/forensics and text/data mining in general. Murat Can Ganiz, Sudhan Kanitkar, Mooi Choo Chuah, William M. Pottenger |
ICDM | 1 |