Burcu Can

dblp:82/8420 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-1700-0395ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 6 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A comprehensive analysis of adversarial attacks against spam filters
Esra Hotoglu, Sevil Sen, Burcu Can
Comput. Secur.3
2025 Semantically-Informed Graph Neural Networks for Irony Detection in Turkish
abstract
Social media plays an important role in expressing the thoughts and sentiments of users. Irony is a way of stating a sentiment about something by expressing the opposite of the intended literal meaning. Irony detection is a recent emerging task in low-resource languages, although other tasks related to sentiment, such as sentiment analysis and emotion detection, have been widely tackled. In this study, we investigate Graph Neural Networks (GNNs) for irony detection in Turkish, a low-resource language in sentiment-related tasks. We incorporate semantic information into the GNNs using the Universal Conceptual Cognitive Annotation (UCCA) framework. Extensive experimental results and in-depth analysis show that our models outperform state-of-the-art irony detection models in Turkish. Our UCCA-GAT (UCCA-Graph Attention Network) model achieves an F 1 -score of 94.85% (7.362% gain over the state-of-the-art) on the Turkish-Irony-Dataset and an accuracy of 72.82% (4.39% gain over the state-of-the-art) on the IronyTR Dataset. We also provide a comprehensive analysis of the proposed models to understand their limitations. 1
Necva Bölücü, Burcu Can
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2024 Error Analysis of NLP Models and Non-Native Speakers of English Identifying Sarcasm in Reddit Comments
abstract
This paper summarises the differences and similarities found between humans and three natural language processing models when attempting to identify whether English online comments are sarcastic or not. Three models were used to analyse 300 comments from the FigLang 2020 Reddit Dataset, with and without context. The same 300 comments were also given to 39 non-native speakers of English and the results were compared. The aim was to find whether there were any results that could be applied to English as a Foreign Language (EFL) teaching. The results showed that there were similarities between the models and non-native speakers, in particular the logistic regression model. They also highlighted weaknesses with both non-native speakers and the models in detecting sarcasm when the comments included political topics or were phrased as questions. This has potential implications for how the EFL teaching industry could implement the results of error analysis of NLP models in teaching practices.
Oliver Cakebread-Andrews, Le An Ha, Ingo Frommholz, Burcu Can
LREC/COLING4
2023 Multi-label emotion classification in texts using transfer learning
Iqra Ameer, Necva Bölücü, Muhammad Hammad Fahim Siddiqui, Burcu Can, Grigori Sidorov, Alexander F. Gelbukh
Expert Syst. Appl.4
2023 A Siamese Neural Network for Learning Semantically-Informed Sentence Embeddings
Necva Bölücü, Burcu Can, Harun Artuner
Expert Syst. Appl.2
2022 Turkish Universal Conceptual Cognitive Annotation
abstract
Universal Conceptual Cognitive Annotation (UCCA) (Abend and Rappoport, 2013a) is a cross-lingual semantic annotation framework that provides an easy annotation without any requirement for linguistic background. UCCA-annotated datasets have been already released in English, French, and German. In this paper, we introduce the first UCCA-annotated Turkish dataset that currently involves 50 sentences obtained from the METU-Sabanci Turkish Treebank (Atalay et al., 2003; Oflazeret al., 2003). We followed a semi-automatic annotation approach, where an external semantic parser is utilised for an initial annotation of the dataset, which is partially accurate and requires refinement. We manually revised the annotations obtained from the semantic parser that are not in line with the UCCA rules that we defined for Turkish. We used the same external semantic parser for evaluation purposes and conducted experiments with both zero-shot and few-shot learning. While the parser cannot predict remote edges in zero-shot setting, using even a small subset of training data in few-shot setting increased the overall F-1 score including the remote edges. This is the initial version of the annotated dataset and we are currently extending the dataset. We will release the current Turkish UCCA annotation guideline along with the annotated dataset.
Necva Bölücü, Burcu Can
LREC2
2022 Joint learning of morphology and syntax with cross-level contextual information flow
abstract
Abstract We propose an integrated deep learning model for morphological segmentation, morpheme tagging, part-of-speech (POS) tagging, and syntactic parsing onto dependencies, using cross-level contextual information flow for every word, from segments to dependencies, with an attention mechanism at horizontal flow. Our model extends the work of Nguyen and Verspoor ((2018). Proceedings of the CoNLL Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies. The Association for Computational Linguistics, pp. 81–91.) on joint POS tagging and dependency parsing to also include morphological segmentation and morphological tagging. We report our results on several languages. Primary focus is agglutination in morphology, in particular Turkish morphology, for which we demonstrate improved performance compared to models trained for individual tasks. Being one of the earlier efforts in joint modeling of syntax and morphology along with dependencies, we discuss prospective guidelines for future comparison.
Burcu Can, Huseyin Alecakir, Suresh Manandhar, Cem Bozsahin
Nat. Lang. Eng.1
2021 Transfer learning for Turkish named entity recognition on noisy text
abstract
Abstract In this article, we investigate using deep neural networks with different word representation techniques for named entity recognition (NER) on Turkish noisy text. We argue that valuable latent features for NER can, in fact, be learned without using any hand-crafted features and/or domain-specific resources such as gazetteers and lexicons. In this regard, we utilize character-level, character n-gram-level, morpheme-level, and orthographic character-level word representations. Since noisy data with NER annotation are scarce for Turkish, we introduce a transfer learning model in order to learn infrequent entity types as an extension to the Bi-LSTM-CRF architecture by incorporating an additional conditional random field (CRF) layer that is trained on a larger (but formal) text and a noisy text simultaneously. This allows us to learn from both formal and informal/noisy text, thus improving the performance of our model further for rarely seen entity types. We experimented on Turkish as a morphologically rich language and English as a relatively morphologically poor language. We obtained an entity-level F1 score of 67.39% on Turkish noisy data and 45.30% on English noisy data, which outperforms the current state-of-art models on noisy text. The English scores are lower compared to Turkish scores because of the intense sparsity in the data introduced by the user writing styles. The results prove that using subword information significantly contributes to learning latent features for morphologically rich languages.
Emre Kagan Akkaya, Burcu Can
Nat. Lang. Eng.2
2021 Incorporating word embeddings in unsupervised morphological segmentation
abstract
Abstract We investigate the usage of semantic information for morphological segmentation since words that are derived from each other will remain semantically related. We use mathematical models such as maximum likelihood estimate (MLE) and maximum a posteriori estimate (MAP) by incorporating semantic information obtained from dense word vector representations. Our approach does not require any annotated data which make it fully unsupervised and require only a small amount of raw data together with pretrained word embeddings for training purposes. The results show that using dense vector representations helps in morphological segmentation especially for low-resource languages. We present results for Turkish, English, and German. Our semantic MLE model outperforms other unsupervised models for Turkish language. Our proposed models could be also used for any other low-resource language with concatenative morphology.
Ahmet Üstün, Burcu Can
Nat. Lang. Eng.2
2021 A Cascaded Unsupervised Model for PoS Tagging
abstract
Part of speech (PoS) tagging is one of the fundamental syntactic tasks in Natural Language Processing, as it assigns a syntactic category to each word within a given sentence or context (such as noun, verb, adjective, etc.). Those syntactic categories could be used to further analyze the sentence-level syntax (e.g., dependency parsing) and thereby extract the meaning of the sentence (e.g., semantic parsing). Various methods have been proposed for learning PoS tags in an unsupervised setting without using any annotated corpora. One of the widely used methods for the tagging problem is log-linear models. Initialization of the parameters in a log-linear model is very crucial for the inference. Different initialization techniques have been used so far. In this work, we present a log-linear model for PoS tagging that uses another fully unsupervised Bayesian model to initialize the parameters of the model in a cascaded framework. Therefore, we transfer some knowledge between two different unsupervised models to leverage the PoS tagging results, where a log-linear model benefits from a Bayesian model’s expertise. We present results for Turkish as a morphologically rich language and for English as a comparably morphologically poor language in a fully unsupervised framework. The results show that our framework outperforms other unsupervised models proposed for PoS tagging.
Necva Bölücü, Burcu Can
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2019 Unsupervised Joint PoS Tagging and Stemming for Agglutinative Languages
abstract
The number of possible word forms is theoretically infinite in agglutinative languages. This brings up the out-of-vocabulary (OOV) issue for part-of-speech (PoS) tagging in agglutinative languages. Since inflectional morphology does not change the PoS tag of a word, we propose to learn stems along with PoS tags simultaneously. Therefore, we aim to overcome the sparsity problem by reducing word forms into their stems. We adopt a Bayesian model that is fully unsupervised. We build a Hidden Markov Model for PoS tagging where the stems are emitted through hidden states. Several versions of the model are introduced in order to observe the effects of different dependencies throughout the corpus, such as the dependency between stems and PoS tags or between PoS tags and affixes. Additionally, we use neural word embeddings to estimate the semantic similarity between the word form and stem. We use the semantic similarity as prior information to discover the actual stem of a word since inflection does not change the meaning of a word. We compare our models with other unsupervised stemming and PoS tagging models on Turkish, Hungarian, Finnish, Basque, and English. The results show that a joint model for PoS tagging and stemming improves on an independent PoS tagger and stemmer in agglutinative languages.
Necva Bölücü, Burcu Can
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2018 Tree Structured Dirichlet Processes for Hierarchical Morphological Segmentation
abstract
This article presents a probabilistic hierarchical clustering model for morphological segmentation. In contrast to existing approaches to morphology learning, our method allows learning hierarchical organization of word morphology as a collection of tree structured paradigms. The model is fully unsupervised and based on the hierarchical Dirichlet process. Tree hierarchies are learned along with the corresponding morphological paradigms simultaneously. Our model is evaluated on Morpho Challenge and shows competitive performance when compared to state-of-the-art unsupervised morphological segmentation systems. Although we apply this model for morphological segmentation, the model itself can also be used for hierarchical clustering of other types of data.
Burcu Can, Suresh Manandhar
Comput. Linguistics1
2017 Joint PoS Tagging and Stemming for Agglutinative Languages
Necva Bölücü, Burcu Can
CICLing (1)2
2017 A Trie-structured Bayesian Model for Unsupervised Morphological Segmentation
Murathan Kurfali, Ahmet Üstün, Burcu Can
CICLing (1)3
2017 Building Morphological Chains for Agglutinative Languages
Serkan Özen, Burcu Can
CICLing (1)2
2017 The Role of Letter Frequency on Eye Movements in Sentential Pseudoword Reading
Cengiz Acartürk, Özkan Kiliç, Bilal Kirkici, Burcu Can, Aysegül Özkan
CogSci4
2016 Turkish PoS Tagging by Reducing Sparsity with Morpheme Tags in Small Datasets
Burcu Can, Ahmet Üstün, Murathan Kurfali
CICLing (1)1
2014 Methods and Algorithms for Unsupervised Learning of Morphology
Burcu Can, Suresh Manandhar
CICLing (1)1
2013 Dirichlet Processes for Joint Learning of Morphology and PoS Tags
Burcu Can, Suresh Manandhar
IJCNLP1
2012 Probabilistic Hierarchical Clustering of Morphological Paradigms
Burcu Can, Suresh Manandhar
EACL1