VLDB 2026 Research / reviewers in the wild / expert
Burcu Can
dblp:82/8420
· DBLP profile ↗
20ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-1700-0395ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 6 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A comprehensive analysis of adversarial attacks against spam filters
Esra Hotoglu, Sevil Sen, Burcu Can |
Comput. Secur. | 3 |
| 2025 | Semantically-Informed Graph Neural Networks for Irony Detection in TurkishabstractSocial media plays an important role in expressing the thoughts and sentiments of users. Irony is a way of stating a sentiment about something by expressing the opposite of the intended literal meaning. Irony detection is a recent emerging task in low-resource languages, although other tasks related to sentiment, such as sentiment analysis and emotion detection, have been widely tackled. In this study, we investigate Graph Neural Networks (GNNs) for irony detection in Turkish, a low-resource language in sentiment-related tasks. We incorporate semantic information into the GNNs using the Universal Conceptual Cognitive Annotation (UCCA) framework. Extensive experimental results and in-depth analysis show that our models outperform state-of-the-art irony detection models in Turkish. Our UCCA-GAT (UCCA-Graph Attention Network) model achieves an F 1 -score of 94.85% (7.362% gain over the state-of-the-art) on the Turkish-Irony-Dataset and an accuracy of 72.82% (4.39% gain over the state-of-the-art) on the IronyTR Dataset. We also provide a comprehensive analysis of the proposed models to understand their limitations. 1 Necva Bölücü, Burcu Can |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | Error Analysis of NLP Models and Non-Native Speakers of English Identifying Sarcasm in Reddit CommentsabstractThis paper summarises the differences and similarities found between humans and three natural language processing models when attempting to identify whether English online comments are sarcastic or not. Three models were used to analyse 300 comments from the FigLang 2020 Reddit Dataset, with and without context. The same 300 comments were also given to 39 non-native speakers of English and the results were compared. The aim was to find whether there were any results that could be applied to English as a Foreign Language (EFL) teaching. The results showed that there were similarities between the models and non-native speakers, in particular the logistic regression model. They also highlighted weaknesses with both non-native speakers and the models in detecting sarcasm when the comments included political topics or were phrased as questions. This has potential implications for how the EFL teaching industry could implement the results of error analysis of NLP models in teaching practices. Oliver Cakebread-Andrews, Le An Ha, Ingo Frommholz, Burcu Can |
LREC/COLING | 4 |
| 2023 | Multi-label emotion classification in texts using transfer learning
Iqra Ameer, Necva Bölücü, Muhammad Hammad Fahim Siddiqui, Burcu Can, Grigori Sidorov, Alexander F. Gelbukh |
Expert Syst. Appl. | 4 |
| 2023 | A Siamese Neural Network for Learning Semantically-Informed Sentence Embeddings
Necva Bölücü, Burcu Can, Harun Artuner |
Expert Syst. Appl. | 2 |
| 2022 | Turkish Universal Conceptual Cognitive AnnotationabstractUniversal Conceptual Cognitive Annotation (UCCA) (Abend and Rappoport, 2013a) is a cross-lingual semantic annotation framework that provides an easy annotation without any requirement for linguistic background. UCCA-annotated datasets have been already released in English, French, and German. In this paper, we introduce the first UCCA-annotated Turkish dataset that currently involves 50 sentences obtained from the METU-Sabanci Turkish Treebank (Atalay et al., 2003; Oflazeret al., 2003). We followed a semi-automatic annotation approach, where an external semantic parser is utilised for an initial annotation of the dataset, which is partially accurate and requires refinement. We manually revised the annotations obtained from the semantic parser that are not in line with the UCCA rules that we defined for Turkish. We used the same external semantic parser for evaluation purposes and conducted experiments with both zero-shot and few-shot learning. While the parser cannot predict remote edges in zero-shot setting, using even a small subset of training data in few-shot setting increased the overall F-1 score including the remote edges. This is the initial version of the annotated dataset and we are currently extending the dataset. We will release the current Turkish UCCA annotation guideline along with the annotated dataset. Necva Bölücü, Burcu Can |
LREC | 2 |
| 2022 | Joint learning of morphology and syntax with cross-level contextual information flowabstractAbstract We propose an integrated deep learning model for morphological segmentation, morpheme tagging, part-of-speech (POS) tagging, and syntactic parsing onto dependencies, using cross-level contextual information flow for every word, from segments to dependencies, with an attention mechanism at horizontal flow. Our model extends the work of Nguyen and Verspoor ((2018). Proceedings of the CoNLL Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies. The Association for Computational Linguistics, pp. 81–91.) on joint POS tagging and dependency parsing to also include morphological segmentation and morphological tagging. We report our results on several languages. Primary focus is agglutination in morphology, in particular Turkish morphology, for which we demonstrate improved performance compared to models trained for individual tasks. Being one of the earlier efforts in joint modeling of syntax and morphology along with dependencies, we discuss prospective guidelines for future comparison. Burcu Can, Huseyin Alecakir, Suresh Manandhar, Cem Bozsahin |
Nat. Lang. Eng. | 1 |
| 2021 | Transfer learning for Turkish named entity recognition on noisy textabstractAbstract In this article, we investigate using deep neural networks with different word representation techniques for named entity recognition (NER) on Turkish noisy text. We argue that valuable latent features for NER can, in fact, be learned without using any hand-crafted features and/or domain-specific resources such as gazetteers and lexicons. In this regard, we utilize character-level, character n-gram-level, morpheme-level, and orthographic character-level word representations. Since noisy data with NER annotation are scarce for Turkish, we introduce a transfer learning model in order to learn infrequent entity types as an extension to the Bi-LSTM-CRF architecture by incorporating an additional conditional random field (CRF) layer that is trained on a larger (but formal) text and a noisy text simultaneously. This allows us to learn from both formal and informal/noisy text, thus improving the performance of our model further for rarely seen entity types. We experimented on Turkish as a morphologically rich language and English as a relatively morphologically poor language. We obtained an entity-level F1 score of 67.39% on Turkish noisy data and 45.30% on English noisy data, which outperforms the current state-of-art models on noisy text. The English scores are lower compared to Turkish scores because of the intense sparsity in the data introduced by the user writing styles. The results prove that using subword information significantly contributes to learning latent features for morphologically rich languages. Emre Kagan Akkaya, Burcu Can |
Nat. Lang. Eng. | 2 |
| 2021 | Incorporating word embeddings in unsupervised morphological segmentationabstractAbstract We investigate the usage of semantic information for morphological segmentation since words that are derived from each other will remain semantically related. We use mathematical models such as maximum likelihood estimate (MLE) and maximum a posteriori estimate (MAP) by incorporating semantic information obtained from dense word vector representations. Our approach does not require any annotated data which make it fully unsupervised and require only a small amount of raw data together with pretrained word embeddings for training purposes. The results show that using dense vector representations helps in morphological segmentation especially for low-resource languages. We present results for Turkish, English, and German. Our semantic MLE model outperforms other unsupervised models for Turkish language. Our proposed models could be also used for any other low-resource language with concatenative morphology. Ahmet Üstün, Burcu Can |
Nat. Lang. Eng. | 2 |
| 2021 | A Cascaded Unsupervised Model for PoS TaggingabstractPart of speech (PoS) tagging is one of the fundamental syntactic tasks in Natural Language Processing, as it assigns a syntactic category to each word within a given sentence or context (such as noun, verb, adjective, etc.). Those syntactic categories could be used to further analyze the sentence-level syntax (e.g., dependency parsing) and thereby extract the meaning of the sentence (e.g., semantic parsing). Various methods have been proposed for learning PoS tags in an unsupervised setting without using any annotated corpora. One of the widely used methods for the tagging problem is log-linear models. Initialization of the parameters in a log-linear model is very crucial for the inference. Different initialization techniques have been used so far. In this work, we present a log-linear model for PoS tagging that uses another fully unsupervised Bayesian model to initialize the parameters of the model in a cascaded framework. Therefore, we transfer some knowledge between two different unsupervised models to leverage the PoS tagging results, where a log-linear model benefits from a Bayesian model’s expertise. We present results for Turkish as a morphologically rich language and for English as a comparably morphologically poor language in a fully unsupervised framework. The results show that our framework outperforms other unsupervised models proposed for PoS tagging. Necva Bölücü, Burcu Can |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2019 | Unsupervised Joint PoS Tagging and Stemming for Agglutinative LanguagesabstractThe number of possible word forms is theoretically infinite in agglutinative languages. This brings up the out-of-vocabulary (OOV) issue for part-of-speech (PoS) tagging in agglutinative languages. Since inflectional morphology does not change the PoS tag of a word, we propose to learn stems along with PoS tags simultaneously. Therefore, we aim to overcome the sparsity problem by reducing word forms into their stems. We adopt a Bayesian model that is fully unsupervised. We build a Hidden Markov Model for PoS tagging where the stems are emitted through hidden states. Several versions of the model are introduced in order to observe the effects of different dependencies throughout the corpus, such as the dependency between stems and PoS tags or between PoS tags and affixes. Additionally, we use neural word embeddings to estimate the semantic similarity between the word form and stem. We use the semantic similarity as prior information to discover the actual stem of a word since inflection does not change the meaning of a word. We compare our models with other unsupervised stemming and PoS tagging models on Turkish, Hungarian, Finnish, Basque, and English. The results show that a joint model for PoS tagging and stemming improves on an independent PoS tagger and stemmer in agglutinative languages. Necva Bölücü, Burcu Can |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2018 | Tree Structured Dirichlet Processes for Hierarchical Morphological SegmentationabstractThis article presents a probabilistic hierarchical clustering model for morphological segmentation. In contrast to existing approaches to morphology learning, our method allows learning hierarchical organization of word morphology as a collection of tree structured paradigms. The model is fully unsupervised and based on the hierarchical Dirichlet process. Tree hierarchies are learned along with the corresponding morphological paradigms simultaneously. Our model is evaluated on Morpho Challenge and shows competitive performance when compared to state-of-the-art unsupervised morphological segmentation systems. Although we apply this model for morphological segmentation, the model itself can also be used for hierarchical clustering of other types of data. Burcu Can, Suresh Manandhar |
Comput. Linguistics | 1 |
| 2017 | Joint PoS Tagging and Stemming for Agglutinative Languages
Necva Bölücü, Burcu Can |
CICLing (1) | 2 |
| 2017 | A Trie-structured Bayesian Model for Unsupervised Morphological Segmentation
Murathan Kurfali, Ahmet Üstün, Burcu Can |
CICLing (1) | 3 |
| 2017 | Building Morphological Chains for Agglutinative Languages
Serkan Özen, Burcu Can |
CICLing (1) | 2 |
| 2017 | The Role of Letter Frequency on Eye Movements in Sentential Pseudoword Reading
Cengiz Acartürk, Özkan Kiliç, Bilal Kirkici, Burcu Can, Aysegül Özkan |
CogSci | 4 |
| 2016 | Turkish PoS Tagging by Reducing Sparsity with Morpheme Tags in Small Datasets
Burcu Can, Ahmet Üstün, Murathan Kurfali |
CICLing (1) | 1 |
| 2014 | Methods and Algorithms for Unsupervised Learning of Morphology
Burcu Can, Suresh Manandhar |
CICLing (1) | 1 |
| 2013 | Dirichlet Processes for Joint Learning of Morphology and PoS Tags
Burcu Can, Suresh Manandhar |
IJCNLP | 1 |
| 2012 | Probabilistic Hierarchical Clustering of Morphological Paradigms
Burcu Can, Suresh Manandhar |
EACL | 1 |