Aykut Koç

dblp:191/2590 · DBLP profile ↗
← Back
30ranked-venue papers
2as first author
21since 2021 · last 2026
0000-0002-6348-2663ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 scHyperLink: Revealing Cell-Type-Specific Gene Regulation With Hypergraph Neural Networks
abstract
Single-cell RNA sequencing (scRNA-seq) allows gene expression to be measured at single-cell resolution, offering new opportunities to investigate Gene Regulatory Networks (GRNs), which represent the regulatory interactions between transcription factors (TFs) and their target genes. Given their relational structure, GRNs are formulated as graphs, enabling gene interaction inference to be framed as a link prediction task among graph nodes (i.e., genes). Prior work adopts Graph Neural Networks (GNNs) to this end, employing their unique ability to model inter-node relationships. However, since GNNs are inherently limited to pair-wise node interactions, they struggle to capture the higher-order dependencies characteristic of GRNs. Gene expression is regulated through multi-way feedback loops involving multiple TFs and targets, and disregarding these higher-order dependencies can lower accuracy in gene interaction inference. To overcome this limitation, we introduce scHyperLink, a hypergraph-based framework for GRN reconstruction. scHyperLink models gene interactions using Hypergraph Neural Networks (HGNNs), where hyperedges allow the simultaneous representation of multi-gene regulatory relationships. scHyperLink integrates experimentally derived interaction graphs with dynamically learned hyperedges to better reflect the underlying regulatory structure. We demonstrate that scHyperLink achieves higher accuracy than state-of-the-art on cell-type-specific benchmark datasets, particularly in sparse regimes with few known interactions. Moreover, we validate the biological relevance of scHyperLink via interpretability analyses on inferred hypergraphs and showcase its scalability to tissue-level analyses. We share the analyzed datasets and source codes for reproducibility.
Emre Kulkul, Tolga Çukur, Aykut Koç
IEEE J. Biomed. Health Informatics3
2025 Joint time-vertex fractional Fourier transform
Tuna Alikasifoglu, Bünyamin Kartal, Eray Özgünay, Aykut Koç
Signal Process.4
2025 Emotion Classification With Visibility Graphs
abstract
Transformers have gained prominence in natural language processing due to their representational capabilities and performances. Transformers process natural language as a sequence on finite context windows; however, global relationships among words beyond these windows cannot be completely modeled via sequence processing only. Graph neural network (GNN) based models have been proposed to alleviate this problem, as they provide geometric extensions to neural networks, enabling models to learn associations within a text. However, regular graph-based methods ignore the sequential nature of underlying texts. In this paper, we propose EmoVis, the first generic graph-based neural network that utilizes visibility graphs, which converts classical time-series information to graph representations. We cast the problem as an emotion classification task, enabling the proposed model to learn associations between the labels and words in a sentence. Moreover, EmoVis can be used as a highly modular graph-based extension to any transformer-based model, significantly improving their performance and learning capabilities in various languages. We experimentally show that EmoVis enables transformer-based models to outperform the state-of-the-art baselines across three diverse datasets in different languages in the SemEval2018 competition datasets and the GoEmotions dataset.
Ecem Simsek, Atakan Topcu, Emirhan Koç, Emine Ulku Saritas, Aykut Koç
IEEE Signal Process. Lett.5
2024 Wiener Filtering in Joint Time-Vertex Fractional Fourier Domains
abstract
Graph signal processing (GSP) uses network structures to analyze and manipulate interconnected signals. These graph signals can also be time-varying. The established joint time-vertex processing framework and corresponding joint timevertex Fourier transform provide a basis to endeavor such timevarying graph signals. The optimal Wiener filtering problem has been deliberated within the joint time-vertex framework. However, the ordinary Fourier domain is only sometimes optimal for separating the signal and noise; one can achieve lower error rates in a fractional Fourier domain. In this paper, we solve the optimal Wiener filtering problem in the joint time-vertex fractional Fourier domains. We provide a theoretical analysis and numerical experiments with comprehensive comparisons to existing filtering approaches for time-varying graph signals to demonstrate the advantages of our approach.
Tuna Alikasifoglu, Bünyamin Kartal, Aykut Koç
IEEE Signal Process. Lett.3
2024 Text-RGNNs: Relational Modeling for Heterogeneous Text Graphs
abstract
Text-graph convolutional Network (TextGCN) is the fundamental work representing corpus with heterogeneous text graphs. Its innovative application of GCNs for text classification has garnered widespread recognition. However, GCNs are inherently designed to operate within homogeneous graphs, potentially limiting their performance. To address this limitation, we present Text Relational Graph Neural Networks (Text-RGNNs), which offer a novel methodology by assigning dedicated weight matrices to each relation within the graph by using heterogeneous GNNs. This approach leverages RGNNs, enabling nuanced and compelling modeling of relationships inherent in the heterogeneous text graphs, ultimately resulting in performance enhancements. We present a theoretical framework for the relational modeling of GNNs for text classification within the context of document classification and demonstrate its effectiveness through extensive experimentation on benchmark datasets. Conducted experiments reveal that Text-RGNNs outperform the existing state-of-the-art in scenarios with complete labeled nodes and minimal labeled training data proportions by incorporating relational modeling into heterogeneous text graphs. Text-RGNNs outperform the second-best models by up to 10.61% for the corresponding evaluation metric.
Arda Can Aras, Tuna Alikasifoglu, Aykut Koç
IEEE Signal Process. Lett.3
2024 Trainable Fractional Fourier Transform
abstract
Recently, the fractional Fourier transform (FrFT) has been integrated into distinct deep neural network (DNN) models such as transformers, sequence models, and convolutional neural networks (CNNs). However, in previous works, the fraction order a is merely considered a hyperparameter and selected heuristically or tuned manually to find the suitable values, which hinders the applicability of FrFT in deep neural networks. We extend the scope of FrFT and introduce it as a trainable layer in neural network architectures, where a is learned in the training stage along with the network weights. We mathematically show that a can be updated in any neural network architecture through backpropagation in the network training phase. We also conduct extensive experiments on benchmark datasets encompassing image classification and time series prediction tasks. Our results show that the trainable FrFT layers alleviate the need to search for suitable a and improve performance over time and Fourier domain approaches. We share our publicly available source codes for reproducibility.
Emirhan Koç, Tuna Alikasifoglu, Arda Can Aras, Aykut Koç
IEEE Signal Process. Lett.4
2024 Automatic Construction of Sememe Knowledge Bases From Machine Readable Dictionaries
abstract
Sememes are the minimum semantic units of natural languages. Words annotated with sememes are organized into Sememe Knowledge Bases (SKBs). SKBs are successfully applied to various high-level language processing tasks as external knowledge bases. However, existing SKBs are manually or semi-manually constructed by linguistic experts over long periods, inhibiting their widespread utilization, updating, and expansion. To automatically construct an SKB from Machine-Readable Dictionaries (MRDs), which are readily available, we propose MRD2SKB as an automatic SKB generation approach. Well-established MRDs exist, and their construction is much simpler than SKBs. Therefore, the proposed MRD2SKB allows for fast, flexible, and extendable generation of SKBs. Building upon matrix factorization and topic modeling, we proposed several variants of MRD2SKB and constructed SKBs fully automatically. Both quantitative and qualitative results of extensive experiments are presented to demonstrate that the performances of the proposed automatically created SKBs are on par with manually and semi-manually prepared SKBs.
Omer Musa Battal, Aykut Koç
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Measuring and Mitigating Gender Bias in Legal Contextualized Language Models
abstract
Transformer-based contextualized language models constitute the state-of-the-art in several natural language processing (NLP) tasks and applications. Despite their utility, contextualized models can contain human-like social biases, as their training corpora generally consist of human-generated text. Evaluating and removing social biases in NLP models has been a major research endeavor. In parallel, NLP approaches in the legal domain, namely, legal NLP or computational law, have also been increasing. Eliminating unwanted bias in legal NLP is crucial, since the law has the utmost importance and effect on people. In this work, we focus on the gender bias encoded in BERT-based models. We propose a new template-based bias measurement method with a new bias evaluation corpus using crime words from the FBI database. This method quantifies the gender bias present in BERT-based models for legal applications. Furthermore, we propose a new fine-tuning-based debiasing method using the European Court of Human Rights (ECtHR) corpus to debias legal pre-trained models. We test the debiased models’ language understanding performance on the LexGLUE benchmark to confirm that the underlying semantic vector space is not perturbed during the debiasing process. Finally, we propose a bias penalty for the performance scores to emphasize the effect of gender bias on model performance.
Mustafa Bozdag, Nurullah Sevim, Aykut Koç
ACM Trans. Knowl. Discov. Data3
2023 Named-entity recognition in Turkish legal texts
abstract
Abstract Natural language processing (NLP) technologies and applications in legal text processing are gaining momentum. Being one of the most prominent tasks in NLP, named-entity recognition (NER) can substantiate a great convenience for NLP in law due to the variety of named entities in the legal domain and their accentuated importance in legal documents. However, domain-specific NER models in the legal domain are not well studied. We present a NER model for Turkish legal texts with a custom-made corpus as well as several NER architectures based on conditional random fields and bidirectional long-short-term memories (BiLSTMs) to address the task. We also study several combinations of different word embeddings consisting of GloVe, Morph2Vec, and neural network-based character feature extraction techniques either with BiLSTM or convolutional neural networks. We report 92.27% F1 score with a hybrid word representation of GloVe and Morph2Vec with character-level features extracted with BiLSTM. Being an agglutinative language, the morphological structure of Turkish is also considered. To the best of our knowledge, our work is the first legal domain-specific NER study in Turkish and also the first study for an agglutinative language in the legal domain. Thus, our work can also have implications beyond the Turkish language.
Can Çetindag, Berkay Yazicioglu, Aykut Koç
Nat. Lang. Eng.3
2023 Gender bias in legal corpora and debiasing it
abstract
Abstract Word embeddings have become important building blocks that are used profoundly in natural language processing (NLP). Despite their several advantages, word embeddings can unintentionally accommodate some gender- and ethnicity-based biases that are present within the corpora they are trained on. Therefore, ethical concerns have been raised since word embeddings are extensively used in several high-level algorithms. Studying such biases and debiasing them have recently become an important research endeavor. Various studies have been conducted to measure the extent of bias that word embeddings capture and to eradicate them. Concurrently, as another subfield that has started to gain traction recently, the applications of NLP in the field of law have started to increase and develop rapidly. As law has a direct and utmost effect on people’s lives, the issues of bias for NLP applications in legal domain are certainly important. However, to the best of our knowledge, bias issues have not yet been studied in the context of legal corpora. In this article, we approach the gender bias problem from the scope of legal text processing domain. Word embedding models that are trained on corpora composed by legal documents and legislation from different countries have been utilized to measure and eliminate gender bias in legal documents. Several methods have been employed to reveal the degree of gender bias and observe its variations over countries. Moreover, a debiasing method has been used to neutralize unwanted bias. The preservation of semantic coherence of the debiased vector space has also been demonstrated by using high-level tasks. Finally, overall results and their implications have been discussed in the scope of NLP in legal domain.
Nurullah Sevim, Furkan Sahinuç, Aykut Koç
Nat. Lang. Eng.3
2023 Multi-Label Sentiment Analysis on 100 Languages With Dynamic Weighting for Label Imbalance
abstract
We investigate cross-lingual sentiment analysis, which has attracted significant attention due to its applications in various areas including market research, politics, and social sciences. In particular, we introduce a sentiment analysis framework in multi-label setting as it obeys Plutchik's wheel of emotions. We introduce a novel dynamic weighting method that balances the contribution from each class during training, unlike previous static weighting methods that assign non-changing weights based on their class frequency. Moreover, we adapt the focal loss that favors harder instances from single-label object recognition literature to our multi-label setting. Furthermore, we derive a method to choose optimal class-specific thresholds that maximize the macro-f1 score in linear time complexity. Through an extensive set of experiments, we show that our method obtains the state-of-the-art performance in seven of nine metrics in three different languages using a single model compared with the common baselines and the best performing methods in the SemEval competition. We publicly share our code for our model, which can perform sentiment analysis in 100 languages, to facilitate further research.
Selim F. Yilmaz, Ergün Batuhan Kaynak, Aykut Koç, Hamdi Dibeklioglu, Suleyman Serdar Kozat
IEEE Trans. Neural Networks Learn. Syst.3
2022 Learning interpretable word embeddings via bidirectional alignment of dimensions with semantic concepts
Lutfi Kerem Senel, Furkan Sahinuç, Veysel Yücesoy, Hinrich Schütze, Tolga Çukur, Aykut Koç
Inf. Process. Manag.6
2022 Fractional Fourier Transform in Time Series Prediction
abstract
Several signal processing tools are integrated into machine learning models for performance and computational cost improvements. Fourier transform (FT) and its variants, which are powerful tools for spectral analysis, are employed in the prediction of univariate time series by converting them to sequences in the spectral domain to be processed further by recurrent neural networks (RNNs). This approach increases the prediction performance and reduces training time compared to conventional methods. In this letter, we introduce fractional Fourier transform (FrFT) to time series prediction by RNNs. As a parametric transformation, FrFT allows us to seek and select better-performing transformation domains by providing access to a continuum of domains between time and frequency. This flexibility yields significant improvements in the prediction power of the underlying models without sacrificing computational efficiency. We evaluated our FrFT-based time series prediction approach on synthetic and real-world datasets. Our results show that FrFT gives rise to performance improvements over ordinary FT.11Source codes are available athttps://github.com/koc-lab/FrFTimeSeries
Emirhan Koç, Aykut Koç
IEEE Signal Process. Lett.2
2022 Fractional Fourier Transform Meets Transformer Encoder
abstract
Utilizing signal processing tools in deep learning models has been drawing increasing attention. Fourier transform (FT), one of the most popular signal processing tools, is employed in many deep learning models. Transformer-based sequential input processing models have also started to make use of FT. In the existing FNet model, it is shown that replacing the attention layer, which is computationally expensive, with FT accelerates model training without sacrificing task performances significantly. We further improve this idea by introducing the fractional Fourier transform (FrFT) into the transformer architecture. As a parameterized transform with a fraction order, FrFT provides an opportunity to access any intermediate domain between time and frequency and find better-performing transformation domains. According to the needs of downstream tasks, a suitable fractional order can be used in our proposed model FrFNet. Our experiments on downstream tasks show that FrFNet leads to performance improvements over the ordinary FNet1.
Furkan Sahinuç, Aykut Koç
IEEE Signal Process. Lett.2
2022 Multivariate Time Series Imputation With Transformers
abstract
Processing time series with missing segments is a fundamental challenge that puts obstacles to advanced analysis in various disciplines such as engineering, medicine, and economics. One of the remedies is imputation to fill the missing values based on observed values properly without undermining performance. We propose the Multivariate Time-Series Imputation with Transformers (MTSIT), a novel method that uses transformer architecture in an unsupervised manner for missing value imputation. Unlike the existing transformer architectures, this model only uses the encoder part of the transformer due to computational benefits. Crucially, MTSIT trains the autoencoder by jointly reconstructing and imputing stochastically-masked inputs via an objective designed for multivariate time-series data. The trained autoencoder is then evaluated for imputing both simulated and real missing values. Experiments show that MTSIT outperforms state-of-the-art imputation methods over benchmark datasets.
Ayberk Yarkin Yildiz, Emirhan Koç, Aykut Koç
IEEE Signal Process. Lett.3
2021 Natural language processing in law: Prediction of outcomes in the higher courts of Turkey
Emre Mumcuoglu, Ceyhun E. Öztürk, Haldun M. Özaktas, Aykut Koç
Inf. Process. Manag.4
2021 Zipfian regularities in "non-point" word representations
Furkan Sahinuç, Aykut Koç
Inf. Process. Manag.2
2021 Imparting interpretability to word embeddings while preserving semantic structure
abstract
Abstract As a ubiquitous method in natural language processing, word embeddings are extensively employed to map semantic properties of words into a dense vector representation. They capture semantic and syntactic relations among words, but the vectors corresponding to the words are only meaningful relative to each other. Neither the vector nor its dimensions have any absolute, interpretable meaning. We introduce an additive modification to the objective function of the embedding learning algorithm that encourages the embedding vectors of words that are semantically related to a predefined concept to take larger values along a specified dimension, while leaving the original semantic learning mechanism mostly unaffected. In other words, we align words that are already determined to be related, along predefined concepts. Therefore, we impart interpretability to the word embedding by assigning meaning to its vector dimensions. The predefined concepts are derived from an external lexical resource, which in this paper is chosen as Roget’s Thesaurus. We observe that alignment along the chosen concepts is not limited to words in the thesaurus and extends to other related words as well. We quantify the extent of interpretability and assignment of meaning from our experimental results. Manual human evaluation results have also been presented to further verify that the proposed method increases interpretability. We also demonstrate the preservation of semantic coherence of the resulting vector space using word-analogy/word-similarity tests and a downstream task. These tests show that the interpretability-imparted word embeddings that are obtained by the proposed framework do not sacrifice performances in common benchmark tests.
Lutfi Kerem Senel, Ihsan Utlu, Furkan Sahinuç, Haldun M. Özaktas, Aykut Koç
Nat. Lang. Eng.5
2021 Operator theory-based computation of linear canonical transforms
Aykut Koç, Haldun M. Özaktas
Signal Process.1
2021 Graph Signal Processing: Vertex Multiplication
abstract
On the Euclidean domains of classical signal processing, linking of signal samples to underlying coordinate structures is straightforward. While graph adjacency matrices totally define the quantitative associations among the underlying graph vertices, a major problem in graph signal processing is the lack of explicit association of vertices with an underlying coordinate structure. To make this link, we propose an operation, called the vertex multiplication (VM), which is defined for graphs and can operate on graph signals. VM, which generalizes the coordinate multiplication (CM) operation in time series signals, can be interpreted as an operator that assigns a coordinate structure to a graph. By using the graph domain extension of differentiation and graph Fourier transform (GFT), VM is defined such that it shows Fourier duality that differentiation and CM operations are duals of each other under Fourier transformation (FT). Numerical examples and applications are also presented.
Bünyamin Kartal, Yigit E. Bayiz, Aykut Koç
IEEE Signal Process. Lett.3
2021 Semantic Change Detection With Gaussian Word Embeddings
abstract
Diachronic study of the evolution of languages is of importance in natural language processing (NLP). Recent years have witnessed a surge of computational approaches for the detection and characterization of lexical semantic change (LSC) by using advancing word representation techniques and the availability of diachronic corpora. We propose a Gaussian word embedding (w2g) based methodology and present a comprehensive study for the LSC detection. W2g is a probabilistic distribution-based word embedding model and represents words as Gaussian mixture models using covariance information along with the existing mean (word vector). We also extensively study several aspects of w2g-based LSC detection under the SemEval-2020 Task 1 evaluation framework as well as using Google N-gram corpus. In the Sub-task 1 (LSC binary classification) of the SemEval-2020 Task 1, we report the highest overall ranking as well as the highest ranks for the two (German and Swedish) of the four languages (English, Swedish, German and Latin). We also report the highest Spearman correlation in the Sub-task 2 (LSC ranking) for Swedish. Our overall rankings in the LSC classification and ranking sub-tasks are 1st and 7th, respectively. Qualitative analysis has also been presented.
Arda Yüksel, Berke Ugurlu, Aykut Koç
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Quadruplet Selection Methods for Deep Embedding Learning
abstract
Recognition of objects with subtle differences has been used in many practical applications, such as car model recognition and maritime vessel identification. For discrimination of the objects in fine-grained detail, we focus on deep embedding learning by using a multi-task learning framework, in which the hierarchical labels (coarse and fine labels) of the samples are utilized both for classification and a quadruplet-based loss function. In order to improve the recognition strength of the learned features, we present a novel feature selection method specifically designed for four training samples of a quadruplet. By experiments, it is observed that the selection of very hard negative samples with relatively easy positive ones from the same coarse and fine classes significantly increases some performance metrics in a fine-grained dataset when compared to selecting the quadruplet samples randomly. The feature embedding learned by the proposed method achieves favorable performance against its state-of-the-art counterparts.
Kaan Karaman, Erhan Gundogdu, Aykut Koç, A. Aydin Alatan
ICIP3
2019 Co-occurrence Weight Selection in Generation of Word Embeddings for Low Resource Languages
abstract
This study aims to increase the performance of word embeddings by proposing a new weighting scheme for co-occurrence counting. The idea behind this new family of weights is to overcome the disadvantage of distant appearing word pairs, which are indeed semantically close, while representing them in the co-occurrence counting. For high-resource languages, this disadvantage might not be effective due to the high frequency of co-occurrence. However, when there are not enough available resources, such pairs suffer from being distant. To favour such pairs, a weighting scheme based on a polynomial fitting procedure is proposed to shift the weights up for distant words while the weights of nearby words are left almost unchanged. The parameter optimization for new weights and the effects of the weighting scheme are analysed for the English, Italian, and Turkish languages. A small portion of English resources and a quarter of Italian resources are utilized for demonstration purposes, as if these languages are low-resource languages. Performance increase is observed in analogy tests when the proposed weighting scheme is applied to relatively small corpora (i.e., mimicking low-resource languages) of both English and Italian. To show the effectiveness of the proposed scheme in small corpora, it is also shown for a large English corpus that the performance of the proposed weighting scheme cannot outperform the original weights. Since Turkish is relatively a low-resource language, it is demonstrated that the proposed weighting scheme can increase the performance of both analogy and similarity tests when all Turkish Wikipedia pages are utilized as a corpus. The positive effect of the proposed scheme has also been demonstrated in a standard sentiment analysis task for the Turkish language.
Veysel Yücesoy, Aykut Koç
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2019 Statistically Segregated k-Space Sampling for Accelerating Multiple-Acquisition MRI
abstract
A central limitation of multiple-acquisition magnetic resonance imaging (MRI) is the degradation in scan efficiency as the number of distinct datasets grows. Sparse recovery techniques can alleviate this limitation via randomly undersampled acquisitions. A frequent sampling strategy is to prescribe for each acquisition a different random pattern drawn from a common sampling density. However, naive random patterns often contain gaps or clusters across the acquisition dimension that, in turn, can degrade reconstruction quality or reduce scan efficiency. To address this problem, a statistically segregated sampling method is proposed for multiple-acquisition MRI. This method generates multiple patterns sequentially while adaptively modifying the sampling density to minimize k-space overlap across patterns. As a result, it improves incoherence across acquisitions while still maintaining similar sampling density across the radial dimension of k-space. Comprehensive simulations and in vivo results are presented for phase-cycled balanced steady-state free precession and multi-echo [Formula: see text]-weighted imaging. Segregated sampling achieves significantly improved quality in both Fourier and compressed-sensing reconstructions of multiple-acquisition datasets.
Lutfi Kerem Senel, Toygan Kilic, Alper Güngör, Emre Kopanoglu, H. Emre Guven, Emine Ulku Saritas, Aykut Koç, Tolga Çukur
IEEE Trans. Medical Imaging7
2018 Generating Semantic Similarity Atlas for Natural Languages
abstract
Cross-lingual studies attract a growing interest in natural language processing (NLP) research, and several studies showed that similar languages are more advantageous to work with than fundamentally different languages in transferring knowledge. Different similarity measures for the languages are proposed by researchers from different domains. However, a similarity measure focusing on semantic structures of languages can be useful for selecting pairs or groups of languages to work with, especially for the tasks requiring semantic knowledge such as sentiment analysis or word sense disambiguation. For this purpose, in this work, we leverage a recently proposed word embedding based method to generate a language similarity atlas for 76 different languages around the world. This atlas can help researchers select similar language pairs or groups in cross-lingual applications. Our findings suggest that semantic similarity between two languages is strongly correlated with the geographic proximity of the countries in which they are used.
Lutfi Kerem Senel, Ihsan Utlu, Veysel Yücesoy, Aykut Koç, Tolga Çukur
SLT4
2018 Fine-grained recognition of maritime vessels and land vehicles by deep feature embedding
abstract
Recent advances in large‐scale image and video analysis have empowered the potential capabilities of visual surveillance systems. In particular, deep learning‐based approaches bring in substantial benefits in solving certain computer vision problems such as fine‐grained object recognition. Here, the authors mainly concentrate on classification and identification of maritime vessels and land vehicles, which are the key constituents of visual surveillance systems. Employing publicly available data sets for maritime vessels and land vehicles, the authors aim to improve visual recognition. Specifically, the authors focus on five tasks regarding visual recognition; coarse‐grained classification, fine‐grained classification, coarse‐grained retrieval, fine‐grained retrieval, and verification. To increase the performance in these tasks, the authors utilise a multi‐task learning framework and present a novel loss function which simultaneously considers deep feature learning and classification by exploiting the available hierarchical labels of individual samples and the global statistics of distances between the data pairs. The authors observe that the proposed multi‐task learning model improves the fine‐grained recognition performance on MARVEL and Stanford Cars data sets, compared to training of a model targeting a single recognition task.
Berkan Solmaz, Erhan Gundogdu, Veysel Yücesoy, Aykut Koç, A. Aydin Alatan
IET Comput. Vis.4
2018 Semantic Structure and Interpretability of Word Embeddings
abstract
Dense word embeddings, which encode meanings of words to low-dimensional vector spaces, have become very popular in natural language processing (NLP) research due to their state-of-the-art performances in many NLP tasks. Word embeddings are substantially successful in capturing semantic relations among words, so a meaningful semantic structure must be present in the respective vector spaces. However, in many cases, this semantic structure is broadly and heterogeneously distributed across the embedding dimensions making interpretation of dimensions a big challenge. In this study, we propose a statistical method to uncover the underlying latent semantic structure in the dense word embeddings. To perform our analysis, we introduce a new dataset (SEMCAT) that contains more than 6500 words semantically grouped under 110 categories. We further propose a method to quantify the interpretability of the word embeddings. The proposed method is a practical alternative to the classical word intrusion test that requires human intervention.
Lutfi Kerem Senel, Ihsan Utlu, Veysel Yücesoy, Aykut Koç, Tolga Çukur
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Sparse representation of two- and three-dimensional images with fractional Fourier, Hartley, linear canonical, and Haar wavelet transforms
Aykut Koç, Burak Bartan, Erhan Gundogdu, Tolga Çukur, Haldun M. Özaktas
Expert Syst. Appl.1
2016 MARVEL: A Large-Scale Image Dataset for Maritime Vessels
Erhan Gundogdu, Berkan Solmaz, Veysel Yücesoy, Aykut Koç
ACCV (5)4
2016 Object classification in infrared images using deep representations
abstract
In this study, we address the problem of infrared (IR) object classification that divides the object appearance space hierarchically with a binary decision tree structure. Binary decisions are made by using the special features of the object appearances. These features are extracted using a fully connected deep neural network learnt by training samples. At each node of the tree, we train individual deep CNNs such that each node specializes in its corresponding subspace. The proposed classification algorithm is evaluated in our generated dataset, which consists of IR targets collected from different video records obtained from different IR sensors (both midwave and longwave) and taken from real world field. The generated dataset consists of four different class labels as ship/boat, tank, plane and helicopter containing a total of 16K samples. Using the proposed tree-based classifier, we observe a favourable performance increase in our dataset against a single deep CNN classifier.
Erhan Gundogdu, Aykut Koç, A. Aydin Alatan
ICIP2