Arzucan Özgür

dblp:26/6952 · DBLP profile ↗
← Back
3ranked-venue papers in the field
0as first author
2since 2021 · last 2025
0000-0001-8376-1056ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 Automatic Labeling of Bank Transfer Categories Using a Hybrid Retrieval Augmented Generation Architecture
Hasan Ersan Yagci, Aygül Dikmen, Cansu Gürel, Ömer Burak Akgün, Ilgin Safak, Nailcan Kara, Özge Özcan Metinkaya, Arzucan Özgür
IEEE Big Data8
2022 A SHAP-based Active Learning Approach for Creating High-Quality Training Data
abstract
Machine learning-based text classification models require labeled data for training. However, manual labeling is a costly and time-consuming process. This task is particularly difficult in domains such as banking, where outsourcing data labeling is generally not allowed due to privacy laws. We propose a novel active learning-based approach in which the most difficult instances in the pool of unlabeled data are selected based on the Shapley Additive Explanations (SHAP) values of the words in the texts to be classified and passed to human annotators for labeling. At each iteration of this human-in-the-loop strategy, newly labeled instances are added to the training set. We demonstrate the effectiveness of this approach in classifying customer comments in the banking domain surveys. Our experiments indicate that better results are achieved when the proposed approach is used to expand the training set, compared to a baseline strategy of expanding the training set with randomly selected instances. Further analysis shows that the difference in performance between the two approaches becomes more pronounced as class imbalance increases. This study suggests that human-in-the-loop based active learning is a powerful strategy for creating high-quality training datasets by effectively leveraging human annotation effort.
Nailcan Kara, Yagiz Levent Gume, Umit Tigrak, Gokce Ezeroglu, Serdar Mola, Ömer Burak Akgün, Arzucan Özgür
IEEE Big Data7
2018 Segmenting hashtags and analyzing their grammatical structure
abstract
Originated as a label to mark specific tweets, hashtags are increasingly used to convey messages that people like to see in the trending hashtags list. Complex noun phrases and even sentences can be turned into a hashtag. Breaking hashtags into their words is a challenging task due to the irregular and compact nature of the language used in Twitter. In this study, we investigate feature‐based machine learning and language model (LM)‐based approaches for hashtag segmentation. Our results show that LM alone is not successful at segmenting nontrivial hashtags. However, when the N‐best LM‐based segmentations are incorporated as features into the feature‐based approach, along with context‐based features proposed in this study, state‐of‐the‐art results in hashtag segmentation are achieved. In addition, we provide an analysis of over two million distinct hashtags, autosegmented by using our best configuration. The analysis reveals that half of all 60 million hashtag occurrences contain multiple words and 80% of sentiment is trapped inside multiword hashtags, justifying the need for hashtag segmentation. Furthermore, we analyze the grammatical structure of hashtags by parsing them and observe that 77% of the hashtags are noun‐based, whereas 11.9% are verb‐based.
Arda Çelebi, Arzucan Özgür
J. Assoc. Inf. Sci. Technol.2