Furkan Sahinuç

dblp:283/8873 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0001-9104-2860ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Reward Modeling for Scientific Writing Evaluation
abstract
Scientific writing is an expert-domain task that demands deep domain knowledge, task-specific requirements and reasoning capabilities that leverage the domain knowledge to satisfy the task specifications.While scientific text generation has been widely studied, its evaluation remains a challenging and open problem.It is critical to develop models that can be reliably deployed for evaluating diverse openended scientific writing tasks while adhering to their distinct requirements.However, existing LLM-based judges and reward models are primarily optimized for general-purpose benchmarks with fixed scoring rubrics and evaluation criteria.Consequently, they often fail to reason over sparse knowledge of scientific domains when interpreting task-dependent and multi-faceted criteria.Moreover, fine-tuning for each individual task is costly and impractical for low-resource settings.To bridge these gaps, we propose cost-efficient, open-source reward models tailored for scientific writing evaluation.We introduce a two-stage training framework that initially optimizes scientific evaluation preferences and then refines reasoning capabilities.Our multi-aspect evaluation design and joint training across diverse tasks enable fine-grained assessment and robustness to dynamic criteria and scoring rubrics.Experimental analysis shows that our training regime strongly improves LLM-based scientific writing evaluation.Our models generalize effectively across tasks and to previously unseen scientific writing evaluation settings, allowing a single trained evaluator to be reused without task-specific retraining.We make our code 1 and data 2 publicly available.
Furkan Sahinuç, Subhabrata Dutta, Iryna Gurevych
ACL (1)1
2024 Systematic Task Exploration with LLMs: A Study in Citation Text Generation
abstract
Large language models (LLMs) bring unprecedented flexibility in defining and executing complex, creative natural language generation (NLG) tasks.Yet, this flexibility brings new challenges, as it introduces new degrees of freedom in formulating the task inputs and instructions and in evaluating model performance.To facilitate the exploration of creative NLG tasks, we propose a three-component research framework that consists of systematic input manipulation, reference data, and output measurement.We use this framework to explore citation text generation -a popular scholarly NLP task that lacks consensus on the task definition and evaluation metric and has not yet been tackled within the LLM paradigm.Our results highlight the importance of systematically investigating both task instruction and input configuration when prompting LLMs, and reveal non-trivial relationships between different evaluation metrics used for citation text generation.Additional human generation and human evaluation experiments provide new qualitative insights into the task to guide future research in citation text generation.We make our code 1 and data 2 publicly available.
Furkan Sahinuç, Ilia Kuznetsov, Yufang Hou 0001, Iryna Gurevych
ACL (1)1
2024 MiDe22: An Annotated Multi-Event Tweet Dataset for Misinformation Detection
abstract
The rapid dissemination of misinformation through online social networks poses a pressing issue with harmful consequences jeopardizing human health, public safety, democracy, and the economy; therefore, urgent action is required to address this problem. In this study, we construct a new human-annotated dataset, called MiDe22, having 5,284 English and 5,064 Turkish tweets with their misinformation labels for several recent events between 2020 and 2022, including the Russia-Ukraine war, COVID-19 pandemic, and Refugees. The dataset includes user engagements with the tweets in terms of likes, replies, retweets, and quotes. We also provide a detailed data analysis with descriptive statistics and the experimental results of a benchmark evaluation for misinformation detection.
Cagri Toraman, Oguzhan Ozcelik, Furkan Sahinuç, Fazli Can
LREC/COLING3
2024 Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards
abstract
Scientific leaderboards are standardized ranking systems that facilitate evaluating and comparing competitive methods.Typically, a leaderboard is defined by a task, dataset, and evaluation metric (TDM) triple, allowing objective performance assessment and fostering innovation through benchmarking.However, the exponential increase in publications has made it infeasible to construct and maintain these leaderboards manually.Automatic leaderboard construction has emerged as a solution to reduce manual labor.Existing datasets for this task are based on the community-contributed leaderboards without additional curation.Our analysis shows that a large portion of these leaderboards are incomplete, and some of them contain incorrect information.In this work, we present SCILEAD, a manually-curated Scientific Leaderboard dataset that overcomes the aforementioned problems.Building on this dataset, we propose three experimental settings that simulate real-world scenarios where TDM triples are fully defined, partially defined, or undefined during leaderboard construction.While previous research has only explored the first setting, the latter two are more representative of real-world applications.To address these diverse settings, we develop a comprehensive LLM-based framework for constructing leaderboards.Our experiments and analysis reveal that various LLMs often correctly identify TDM triples while struggling to extract result values from publications.We make our code 1 and data 2 publicly available.
Furkan Sahinuç, Thy Thy Tran, Yulia Grishina, Yufang Hou 0001, Iryna Gurevych
EMNLP1
2023 Gender bias in legal corpora and debiasing it
abstract
Abstract Word embeddings have become important building blocks that are used profoundly in natural language processing (NLP). Despite their several advantages, word embeddings can unintentionally accommodate some gender- and ethnicity-based biases that are present within the corpora they are trained on. Therefore, ethical concerns have been raised since word embeddings are extensively used in several high-level algorithms. Studying such biases and debiasing them have recently become an important research endeavor. Various studies have been conducted to measure the extent of bias that word embeddings capture and to eradicate them. Concurrently, as another subfield that has started to gain traction recently, the applications of NLP in the field of law have started to increase and develop rapidly. As law has a direct and utmost effect on people’s lives, the issues of bias for NLP applications in legal domain are certainly important. However, to the best of our knowledge, bias issues have not yet been studied in the context of legal corpora. In this article, we approach the gender bias problem from the scope of legal text processing domain. Word embedding models that are trained on corpora composed by legal documents and legislation from different countries have been utilized to measure and eliminate gender bias in legal documents. Several methods have been employed to reveal the degree of gender bias and observe its variations over countries. Moreover, a debiasing method has been used to neutralize unwanted bias. The preservation of semantic coherence of the debiased vector space has also been demonstrated by using high-level tasks. Finally, overall results and their implications have been discussed in the scope of NLP in legal domain.
Nurullah Sevim, Furkan Sahinuç, Aykut Koç
Nat. Lang. Eng.2
2023 Impact of Tokenization on Language Models: An Analysis for Turkish
abstract
Tokenization is an important text preprocessing step to prepare input tokens for deep language models. WordPiece and BPE are de facto methods employed by important models, such as BERT and GPT. However, the impact of tokenization can be different for morphologically rich languages, such as Turkic languages, in which many words can be generated by adding prefixes and suffixes. We compare five tokenizers at different granularity levels, that is, their outputs vary from the smallest pieces of characters to the surface form of words, including a Morphological-level tokenizer. We train these tokenizers and pretrain medium-sized language models using the RoBERTa pretraining procedure on the Turkish split of the OSCAR corpus. We then fine-tune our models on six downstream tasks. Our experiments, supported by statistical tests, reveal that the morphological-level tokenizer delivers a challenging performance with de facto tokenizers. Furthermore, we find that increasing the vocabulary size improves the performance of Morphological- and Word-level tokenizers more than that of de facto tokenizers. The ratio of the number of vocabulary parameters to the total number of model parameters can be empirically chosen as 20% for de facto tokenizers and 40% for other tokenizers to obtain a reasonable trade-off between model size and performance.
Cagri Toraman, Eyup Halit Yilmaz, Furkan Sahinuç, Oguzhan Ozcelik
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2022 Large-Scale Hate Speech Detection with Cross-Domain Transfer
abstract
The performance of hate speech detection models relies on the datasets on which the models are trained. Existing datasets are mostly prepared with a limited number of instances or hate domains that define hate topics. This hinders large-scale analysis and transfer learning with respect to hate domains. In this study, we construct large-scale tweet datasets for hate speech detection in English and a low-resource language, Turkish, consisting of human-labeled 100k tweets per each. Our datasets are designed to have equal number of tweets distributed over five domains. The experimental results supported by statistical tests show that Transformer-based language models outperform conventional bag-of-words and neural models by at least 5% in English and 10% in Turkish for large-scale hate speech detection. The performance is also scalable to different training sizes, such that 98% of performance in English, and 97% in Turkish, are recovered when 20% of training instances are used. We further examine the generalization ability of cross-domain transfer among hate domains. We show that 96% of the performance of a target domain in average is recovered by other domains for English, and 92% for Turkish. Gender and religion are more successful to generalize to other domains, while sports fail most.
Cagri Toraman, Furkan Sahinuç, Eyup Halit Yilmaz
LREC2
2022 Learning interpretable word embeddings via bidirectional alignment of dimensions with semantic concepts
Lutfi Kerem Senel, Furkan Sahinuç, Veysel Yücesoy, Hinrich Schütze, Tolga Çukur, Aykut Koç
Inf. Process. Manag.2
2022 Fractional Fourier Transform Meets Transformer Encoder
abstract
Utilizing signal processing tools in deep learning models has been drawing increasing attention. Fourier transform (FT), one of the most popular signal processing tools, is employed in many deep learning models. Transformer-based sequential input processing models have also started to make use of FT. In the existing FNet model, it is shown that replacing the attention layer, which is computationally expensive, with FT accelerates model training without sacrificing task performances significantly. We further improve this idea by introducing the fractional Fourier transform (FrFT) into the transformer architecture. As a parameterized transform with a fraction order, FrFT provides an opportunity to access any intermediate domain between time and frequency and find better-performing transformation domains. According to the needs of downstream tasks, a suitable fractional order can be used in our proposed model FrFNet. Our experiments on downstream tasks show that FrFNet leads to performance improvements over the ordinary FNet1.
Furkan Sahinuç, Aykut Koç
IEEE Signal Process. Lett.1
2021 Tweet Length Matters: A Comparative Analysis on Topic Detection in Microblogs
Furkan Sahinuç, Cagri Toraman
ECIR (2)1
2021 Zipfian regularities in "non-point" word representations
Furkan Sahinuç, Aykut Koç
Inf. Process. Manag.1
2021 Imparting interpretability to word embeddings while preserving semantic structure
abstract
Abstract As a ubiquitous method in natural language processing, word embeddings are extensively employed to map semantic properties of words into a dense vector representation. They capture semantic and syntactic relations among words, but the vectors corresponding to the words are only meaningful relative to each other. Neither the vector nor its dimensions have any absolute, interpretable meaning. We introduce an additive modification to the objective function of the embedding learning algorithm that encourages the embedding vectors of words that are semantically related to a predefined concept to take larger values along a specified dimension, while leaving the original semantic learning mechanism mostly unaffected. In other words, we align words that are already determined to be related, along predefined concepts. Therefore, we impart interpretability to the word embedding by assigning meaning to its vector dimensions. The predefined concepts are derived from an external lexical resource, which in this paper is chosen as Roget’s Thesaurus. We observe that alignment along the chosen concepts is not limited to words in the thesaurus and extends to other related words as well. We quantify the extent of interpretability and assignment of meaning from our experimental results. Manual human evaluation results have also been presented to further verify that the proposed method increases interpretability. We also demonstrate the preservation of semantic coherence of the resulting vector space using word-analogy/word-similarity tests and a downstream task. These tests show that the interpretability-imparted word embeddings that are obtained by the proposed framework do not sacrifice performances in common benchmark tests.
Lutfi Kerem Senel, Ihsan Utlu, Furkan Sahinuç, Haldun M. Özaktas, Aykut Koç
Nat. Lang. Eng.3