Ramazan Gokberk Cinbis

dblp:54/2808 · also Ramazan Gökberk Cinbis · DBLP profile ↗
← Back
39ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0003-0962-7101ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021
YearPublicationVenuePosition
2026 Representation recycling for streaming video analysis
Can Ufuk Ertenli, Ramazan Gokberk Cinbis, Emre Akbas
Neurocomputing2
2025 Interchangeable Token Embeddings for Extendable Vocabulary and Alpha-Equivalence
abstract
Language models lack the notion of interchangeable tokens: symbols that are semantically equivalent yet distinct, such as bound variables in formal logic. This limitation prevents generalization to larger vocabularies and hinders the model's ability to recognize alpha-equivalence, where renaming bound variables preserves meaning. We formalize this machine learning problem and introduce alpha-covariance, a metric for evaluating robustness to such transformations. To tackle this task, we propose a dual-part token embedding strategy: a shared component ensures semantic consistency, while a randomized component maintains token distinguishability. Compared to a baseline that relies on alpha-renaming for data augmentation, our approach demonstrates improved generalization to unseen tokens in linear temporal logic solving, propositional logic assignment prediction, and copying with an extendable vocabulary, while introducing a favorable inductive bias for alpha-equivalence. Our findings establish a foundation for designing language models that can learn interchangeable token representations, a crucial step toward more flexible and systematic reasoning in formal domains. Our code and project page are available at https://necrashter.github.io/interchangeable-token-embeddings
Ilker Isik, Ramazan Gokberk Cinbis, Ebru Aydin Gol
ICML2
2025 DARKIN: a zero-shot benchmark for phosphosite-dark kinase association using protein language models
abstract
MOTIVATION: Protein language models (pLMs) have emerged as powerful tools for capturing the intricate information encoded in protein sequences, facilitating various downstream protein prediction tasks. With numerous pLMs available, there is a critical need for diverse benchmarks to systematically evaluate their performance across biologically relevant tasks. Here, we introduce DARKIN, a zero-shot classification benchmark designed to assign phosphosites to understudied kinases, termed dark kinases. Kinases, which catalyze phosphorylation, are central to cellular signaling pathways. While phosphoproteomics enables the large-scale identification of phosphosites, determining the cognate kinase responsible for the phosphorylation event remains an experimental challenge. RESULTS: In DARKIN, we prepared training, validation, and test folds that respect the zero-shot nature of this classification problem, incorporating stratification based on kinase groups and sequence similarity. We evaluated multiple pLMs using two zero-shot classifiers: a novel, training-free k-NN-based method, and a bilinear classifier. Our findings indicate that ESM, ProtT5-XL, and SaProt exhibit superior performance on this task. DARKIN provides a challenging benchmark for assessing pLM efficacy and fosters deeper exploration of under-characterized (dark) kinases by offering a biologically relevant test bed. AVAILABILITY AND IMPLEMENTATION: The DARKIN benchmark data and the scripts for generating additional splits are publicly available at: https://github.com/tastanlab/darkin.
Emine Ayse Sunar, Zeynep Isik, Mert Pekey, Ramazan Gokberk Cinbis, Öznur Tastan
Bioinform.4
2024 Cross-lingual few-shot sign language recognition
Yunus Can Bilge, Nazli Ikizler-Cinbis, Ramazan Gokberk Cinbis
Pattern Recognit.3
2023 Meta-Tuning Loss Functions and Data Augmentation for Few-Shot Object Detection
abstract
Few-shot object detection, the problem of modelling novel object detection categories with few training instances, is an emerging topic in the area of few-shot learning and object detection. Contemporary techniques can be divided into two groups: fine-tuning based and meta-learning based approaches. While meta-learning approaches aim to learn dedicated meta-models for mapping samples to novel class models, fine-tuning approaches tackle few-shot detection in a simpler manner, by adapting the detection model to novel classes through gradient based optimization. Despite their simplicity, fine-tuning based approaches typically yield competitive detection results. Based on this observation, we focus on the role of loss functions and augmentations as the force driving the fine-tuning process, and propose to tune their dynamics through meta-learning principles. The proposed training scheme, therefore, allows learning inductive biases that can boost few-shot detection, while keeping the advantages of fine-tuning based approaches. In addition, the proposed approach yields interpretable loss functions, as opposed to highly parametric and complex few-shot meta-models. The experimental results highlight the merits of the proposed scheme, with significant improvements over the strong fine-tuning based few-shot detection baselines on benchmark Pascal VOC and MS-COCO datasets, in terms of both standard and generalized few-shot performance metrics.
Berkan Demirel, Orhun Bugra Baran, Ramazan Gokberk Cinbis
CVPR3
2023 HybridAugment++: Unified Frequency Spectra Perturbations for Model Robustness
abstract
Convolutional Neural Networks (CNN) are known to exhibit poor generalization performance under distribution shifts. Their generalization have been studied extensively, and one line of work approaches the problem from a frequency-centric perspective. These studies highlight the fact that humans and CNNs might focus on different frequency components of an image. First, inspired by these observations, we propose a simple yet effective data augmentation method HybridAugment that reduces the reliance of CNNs on high-frequency components, and thus improves their robustness while keeping their clean accuracy high. Second, we propose HybridAugment++, which is a hierarchical augmentation method that attempts to unify various frequency-spectrum augmentations. HybridAugment++ builds on HybridAugment, and also reduces the reliance of CNNs on the amplitude component of images, and promotes phase information instead. This unification results in competitive to or better than state-of-the-art results on clean accuracy (CIFAR-10/100 and ImageNet), corruption benchmarks (ImageNet-C, CIFAR-10-C and CIFAR-100-C), adversarial robustness on CIFAR-10 and out-of-distribution detection on various datasets. HybridAugment and HybridAugment++ are implemented in a few lines of code, does not require extra data, ensemble models or additional networks1.
Mehmet Kerim Yucel, Ramazan Gokberk Cinbis, Pinar Duygulu
ICCV2
2023 Towards Zero-Shot Sign Language Recognition
abstract
This paper tackles the problem of zero-shot sign language recognition (ZSSLR), where the goal is to leverage models learned over the seen sign classes to recognize the instances of unseen sign classes. In this context, readily available textual sign descriptions and attributes collected from sign language dictionaries are utilized as semantic class representations for knowledge transfer. For this novel problem setup, we introduce three benchmark datasets with their accompanying textual and attribute descriptions to analyze the problem in detail. Our proposed approach builds spatiotemporal models of body and hand regions. By leveraging the descriptive text and attribute embeddings along with these visual representations within a zero-shot learning framework, we show that textual and attribute based class definitions can provide effective knowledge for the recognition of previously unseen sign classes. We additionally introduce techniques to analyze the influence of binary attributes in correct and incorrect zero-shot predictions. We anticipate that the introduced approaches and the accompanying datasets will provide a basis for further exploration of zero-shot learning in sign language recognition.
Yunus Can Bilge, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 DeepSide: A Deep Learning Approach for Drug Side Effect Prediction
abstract
Drug failures due to unforeseen adverse effects at clinical trials pose health risks for the participants and lead to substantial financial losses. Side effect prediction algorithms have the potential to guide the drug design process. LINCS L1000 dataset provides a vast resource of cell line gene expression data perturbed by different drugs and creates a knowledge base for context specific features. The state-of-the-art approach that aims at using context specific information relies on only the high-quality experiments in LINCS L1000 and discards a large portion of the experiments. In this study, our goal is to boost the prediction performance by utilizing this data to its full extent. We experiment with 5 deep learning architectures. We find that a multi-modal architecture produces the best predictive performance among multi-layer perceptron-based architectures when drug chemical structure (CS), and the full set of drug perturbed gene expression profiles (GEX) are used as modalities. Overall, we observe that the CS is more informative than the GEX. A convolutional neural network-based model that uses only SMILES string representation of the drugs achieves the best results and provides 13.0% macro-AUC and 3.1% micro-AUC improvements over the state-of-the-art. We also show that the model is able to predict side effect-drug pairs that are reported in the literature but was missing in the ground truth side effect dataset. DeepSide is available at http://github.com/OnurUner/DeepSide.
Onur Can Uner, Halil Ibrahim Kuru, Ramazan Gokberk Cinbis, Öznur Tastan, A. Ercüment Çiçek
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 Streaming Multiscale Deep Equilibrium Models
Can Ufuk Ertenli, Emre Akbas, Ramazan Gokberk Cinbis
ECCV (11)3
2022 Closed-form Sample Probing for Learning Generative Models in Zero-shot Learning
Samet Çetin, Orhun Bugra Baran, Ramazan Gokberk Cinbis
ICLR3
2022 MaskSplit: Self-supervised Meta-learning for Few-shot Semantic Segmentation
abstract
Just like other few-shot learning problems, few-shot segmentation aims to minimize the need for manual annotation, which is particularly costly in segmentation tasks. Even though the few-shot setting reduces this cost for novel test classes, there is still a need to annotate the training data. To alleviate this need, we propose a self-supervised training approach for learning few-shot segmentation models. We first use unsupervised saliency estimation to obtain pseudo-masks on images. We then train a simple prototype based model over different splits of pseudo masks and augmentations of images. Our extensive experiments show that the proposed approach achieves promising results, highlighting the potential of self-supervised training. To the best of our knowledge this is the first work that addresses unsupervised few-shot segmentation problem on natural images.
Mustafa Sercan Amac, Ahmet Sencan, Orhun Bugra Baran, Nazli Ikizler-Cinbis, Ramazan Gokberk Cinbis
WACV5
2022 Semantics-driven attentive few-shot learning over clean and noisy samples
Orhun Bugra Baran, Ramazan Gokberk Cinbis
Neurocomputing2
2022 Caption generation on scenes with seen and unseen object categories
Berkan Demirel, Ramazan Gokberk Cinbis
Image Vis. Comput.2
2022 How robust are discriminatively trained zero-shot learning models?
Mehmet Kerim Yucel, Ramazan Gokberk Cinbis, Pinar Duygulu
Image Vis. Comput.2
2021 Red Carpet to Fight Club: Partially-supervised Domain Transfer for Face Recognition in Violent Videos
abstract
In many real-world problems, there is typically a large discrepancy between the characteristics of data used in training versus deployment. A prime example is the analysis of aggression videos: in a criminal incidence, typically suspects need to be identified based on their clean portraitlike photos, instead of their prior video recordings. This results in three major challenges; large domain discrepancy between violence videos and ID-photos, the lack of video examples for most individuals and limited training data availability. To mimic such scenarios, we formulate a realistic domain-transfer problem, where the goal is to transfer the recognition model trained on clean posed images to the target domain of violent videos, where training videos are available only for a subset of subjects. To this end, we introduce the "WildestFaces" dataset, tailored to study cross-domain recognition under a variety of adverse conditions. We divide the task of transferring a recognition model from the domain of clean images to the violent videos into two sub-problems and tackle them using (i) stacked affine-transforms for classifier-transfer, (ii) attention-driven pooling for temporal-adaptation. We additionally formulate a self-attention based model for domain-transfer. We establish a rigorous evaluation protocol for this "clean-to-violent" recognition task, and present a detailed analysis of the proposed dataset and the methods. Our experiments highlight the unique challenges introduced by the WildestFaces dataset and the advantages of the proposed approach.
Yunus Can Bilge, Mehmet Kerim Yucel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis, Pinar Duygulu
WACV3
2020 Key protected classification for collaborative learning
Mert Bülent Sariyildiz, Ramazan Gokberk Cinbis, Erman Ayday
Pattern Recognit.2
2019 Zero-Shot Sign Language Recognition: Can Textual Data Uncover Sign Languages?
Yunus Can Bilge, Nazli Ikizler-Cinbis, Ramazan Gokberk Cinbis
BMVC3
2019 Image Captioning with Unseen Objects
Berkan Demirel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis
BMVC2
2019 Gradient Matching Generative Networks for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) is one of the most promising problems where substantial progress can potentially be achieved through unsupervised learning, due to distributional differences between supervised and zero-shot classes. For this reason, several works investigate the incorporation of discriminative domain adaptation techniques into ZSL, which, however, lead to modest improvements in ZSL accuracy. In contrast, we propose a generative model that can naturally learn from unsupervised examples, and synthesize training examples for unseen classes purely based on their class embeddings, and therefore, reduce the zero-shot learning problem into a supervised classification task. The proposed approach consists of two important components: (i) a conditional Generative Adversarial Network that learns to produce samples that mimic the characteristics of unsupervised data examples, and (ii) the Gradient Matching (GM) loss that measures the quality of the gradient signal obtained from the synthesized examples. Using our GM loss formulation, we enforce the generator to produce examples from which accurate classifiers can be trained. Experimental results on several ZSL benchmark datasets show that our approach leads to significant improvements over the state of the art in generalized zero-shot classification.
Mert Bülent Sariyildiz, Ramazan Gokberk Cinbis
CVPR2
2019 Cross-Task Weakly Supervised Learning From Instructional Videos
abstract
In this paper we investigate learning visual models for the steps of ordinary tasks using weak supervision via instructional narrations and an ordered list of steps instead of strong supervision via temporal annotations. At the heart of our approach is the observation that weakly supervised learning may be easier if a model shares components while learning different steps: ``pour egg'' should be trained jointly with other tasks involving ``pour'' and ``egg''. We formalize this in a component model for recognizing steps and a weakly supervised learning framework that can learn this model under temporal constraints from narration and the list of steps. Past data does not permit systematic studying of sharing and so we also gather a new dataset aimed at assessing cross-task sharing. Our experiments demonstrate that sharing across tasks improves performance, especially when done at the component level and that our component model can parse previously unseen tasks by virtue of its compositionality.
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David F. Fouhey, Ivan Laptev, Josef Sivic
CVPR3
2019 Learning Visually Consistent Label Embeddings for Zero-Shot Learning
abstract
In this work, we propose a zero-shot learning method to effectively model knowledge transfer between classes via jointly learning visually consistent word vectors and label embedding model in an end-to-end manner. The main idea is to project the vector space word vectors of attributes and classes into the visual space such that word representations of semantically related classes become more closer, and use the projected vectors in the proposed embedding model to identify unseen classes. We evaluate the proposed approach on two benchmark datasets and the experimental results show that our method yields significant improvements in recognition accuracy.
Berkan Demirel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis
ICIP2
2019 Weakly Supervised Deep Convolutional Networks for Fine-Grained Object Recognition in Multispectral Images
abstract
The challenging task of training object detectors for fine-grained classification faces additional difficulties when there are registration errors between the image data and the ground truth. We propose a weakly supervised learning methodology for the classification of 40 types of trees by using fixed-sized multispectral images with a class label but with no exact knowledge of the object location. Our approach consists of an end-to-end trainable convolutional neural network with separate branches for learning class-specific and location-specific scoring of image regions. Comparative experiments show that the proposed method simultaneously learns to detect and classify the objects of interest with high accuracy.
Bulut Aygünes, Selim Aksoy, Ramazan Gokberk Cinbis
IGARSS3
2019 Multisource Region Attention Network for Fine-Grained Object Recognition in Remote Sensing Imagery
abstract
Fine-grained object recognition concerns the identification of the type of an object among a large number of closely related subcategories. Multisource data analysis that aims to leverage the complementary spectral, spatial, and structural information embedded in different sources is a promising direction toward solving the fine-grained recognition problem that involves low between-class variance, small training set sizes for rare classes, and class imbalance. However, the common assumption of coregistered sources may not hold at the pixel level for small objects of interest. We present a novel methodology that aims to simultaneously learn the alignment of multisource data and the classification model in a unified framework. The proposed method involves a multisource region attention network that computes per-source feature representations, assigns attention scores to candidate regions sampled around the expected object locations by using these representations, and classifies the objects by using an attention-driven multisource representation that combines the feature representations and the attention scores from all sources. All components of the model are realized using deep neural networks and are learned in an end-to-end fashion. Experiments using RGB, multispectral, and LiDAR elevation data for classification of street trees showed that our approach achieved 64.2% and 47.3% accuracies for the 18-class and 40-class settings, respectively, which correspond to 13% and 14.3% improvement relative to the commonly used feature concatenation approach from multiple sources.
Gencer Sumbul, Ramazan Gokberk Cinbis, Selim Aksoy
IEEE Trans. Geosci. Remote. Sens.2
2018 Zero-Shot Object Detection by Hybrid Region Embedding
Berkan Demirel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis
BMVC2
2018 Fine-Grained Object Recognition and Zero-Shot Learning in Remote Sensing Imagery
abstract
Fine-grained object recognition that aims to identify the type of an object among a large number of subcategories is an emerging application with the increasing resolution that exposes new details in image data. Traditional fully supervised algorithms fail to handle this problem where there is low between-class variance and high within-class variance for the classes of interest with small sample sizes. We study an even more extreme scenario named zero-shot learning (ZSL) in which no training example exists for some of the classes. ZSL aims to build a recognition model for new unseen categories by relating them to seen classes that were previously learned. We establish this relation by learning a compatibility function between image features extracted via a convolutional neural network and auxiliary information that describes the semantics of the classes of interest by using training samples from the seen classes. Then, we show how knowledge transfer can be performed for the unseen classes by maximizing this function during inference. We introduce a new data set that contains 40 different types of street trees in 1-ft spatial resolution aerial data, and evaluate the performance of this model with manually annotated attributes, a natural language model, and a scientific taxonomy as auxiliary information. The experiments show that the proposed model achieves 14.3% recognition accuracy for the classes with no training examples, which is significantly better than a random guess accuracy of 6.3% for 16 test classes, and three other ZSL algorithms.
Gencer Sumbul, Ramazan Gokberk Cinbis, Selim Aksoy
IEEE Trans. Geosci. Remote. Sens.2
2017 Attributes2Classname: A Discriminative Model for Attribute-Based Unsupervised Zero-Shot Learning
abstract
We propose a novel approach for unsupervised zero-shot learning (ZSL) of classes based on their names. Most existing unsupervised ZSL methods aim to learn a model for directly comparing image features and class names. However, this proves to be a difficult task due to dominance of non-visual semantics in underlying vector-space embeddings of class names. To address this issue, we discriminatively learn a word representation such that the similarities between class and combination of attribute names fall in line with the visual similarity. Contrary to the traditional zero-shot learning approaches that are built upon attribute presence, our approach bypasses the laborious attribute-class relation annotations for unseen classes. In addition, our proposed approach renders text-only training possible, hence, the training can be augmented without the need to collect additional image data. The experimental results show that our method yields state-of-the-art results for unsupervised ZSL in three benchmark datasets.
Berkan Demirel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis
ICCV2
2017 Weakly Supervised Object Localization with Multi-Fold Multiple Instance Learning
abstract
Object category localization is a challenging problem in computer vision. Standard supervised training requires bounding box annotations of object instances. This time-consuming annotation process is sidestepped in weakly supervised learning. In this case, the supervised information is restricted to binary labels that indicate the absence/presence of object instances in the image, without their locations. We follow a multiple-instance learning approach that iteratively trains the detector and infers the object locations in the positive training images. Our main contribution is a multi-fold multiple instance learning procedure, which prevents training from prematurely locking onto erroneous object locations. This procedure is particularly important when using high-dimensional representations, such as Fisher vectors and convolutional neural network features. We also propose a window refinement method, which improves the localization accuracy by incorporating an objectness prior. We present a detailed experimental evaluation using the PASCAL VOC 2007 dataset, which verifies the effectiveness of our approach.
Ramazan Gokberk Cinbis, Jakob Verbeek, Cordelia Schmid
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Approximate Fisher Kernels of Non-iid Image Models for Image Categorization
abstract
The bag-of-words (BoW) model treats images as sets of local descriptors and represents them by visual word histograms. The Fisher vector (FV) representation extends BoW, by considering the first and second order statistics of local descriptors. In both representations local descriptors are assumed to be identically and independently distributed (iid), which is a poor assumption from a modeling perspective. It has been experimentally observed that the performance of BoW and FV representations can be improved by employing discounting transformations such as power normalization. In this paper, we introduce non-iid models by treating the model parameters as latent variables which are integrated out, rendering all local regions dependent. Using the Fisher kernel principle we encode an image by the gradient of the data log-likelihood w.r.t. the model hyper-parameters. Our models naturally generate discounting effects in the representations; suggesting that such transformations have proven successful because they closely correspond to the representations obtained for non-iid models. To enable tractable computation, we rely on variational free-energy bounds to learn the hyper-parameters and to compute approximate Fisher kernels. Our experimental evaluation results validate that our models lead to performance improvements comparable to using power normalization, as employed in state-of-the-art feature aggregation methods.
Ramazan Gokberk Cinbis, Jakob Verbeek, Cordelia Schmid
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Multi-fold MIL Training for Weakly Supervised Object Localization
abstract
Object category localization is a challenging problem in computer vision. Standard supervised training requires bounding box annotations of object instances. This time-consuming annotation process is sidestepped in weakly supervised learning. In this case, the supervised information is restricted to binary labels that indicate the absence/presence of object instances in the image, without their locations. We follow a multiple-instance learning approach that iteratively trains the detector and infers the object locations in the positive training images. Our main contribution is a multi-fold multiple instance learning procedure, which prevents training from prematurely locking onto erroneous object locations. This procedure is particularly important when high-dimensional representations, such as the Fisher vectors, are used. We present a detailed experimental evaluation using the PASCAL VOC 2007 dataset. Compared to state-of-the-art weakly supervised detectors, our approach better localizes objects in the training images, which translates into improved detection performance.
Ramazan Gokberk Cinbis, Jakob Verbeek, Cordelia Schmid
CVPR1
2013 Segmentation Driven Object Detection with Fisher Vectors
abstract
We present an object detection system based on the Fisher vector (FV) image representation computed over SIFT and color descriptors. For computational and storage efficiency, we use a recent segmentation-based method to generate class-independent object detection hypotheses, in combination with data compression techniques. Our main contribution is a method to produce tentative object segmentation masks to suppress background clutter in the features. Re-weighting the local image features based on these masks is shown to improve object detection significantly. We also exploit contextual features in the form of a full-image FV descriptor, and an inter-category rescoring mechanism. Our experiments on the VOC 2007 and 2010 datasets show that our detector improves over the current state-of-the-art detection results.
Ramazan Gokberk Cinbis, Jakob Verbeek, Cordelia Schmid
ICCV1
2012 Image categorization using Fisher kernels of non-iid image models
abstract
The bag-of-words (BoW) model treats images as an unordered set of local regions and represents them by visual word histograms. Implicitly, regions are assumed to be identically and independently distributed (iid), which is a poor assumption from a modeling perspective. We introduce non-iid models by treating the parameters of BoW models as latent variables which are integrated out, rendering all local regions dependent. Using the Fisher kernel we encode an image by the gradient of the data log-likelihood w.r.t. hyper-parameters that control priors on the model parameters. Our representation naturally involves discounting transformations similar to taking square-roots, providing an explanation of why such transformations have proven successful. Using variational inference we extend the basic model to include Gaussian mixtures over local descriptors, and latent topic models to capture the co-occurrence structure of visual words, both improving performance. Our models yield state-of-the-art categorization performance using linear classifiers; without using non-linear transformations such as taking square-roots of features, or using (approximate) explicit embeddings of non-linear kernels.
Ramazan Gokberk Cinbis, Jakob Verbeek, Cordelia Schmid
CVPR1
2012 Contextual Object Detection Using Set-Based Classification
Ramazan Gokberk Cinbis, Stan Sclaroff
ECCV (6)1
2011 Unsupervised metric learning for face identification in TV video
abstract
The goal of face identification is to decide whether two faces depict the same person or not. This paper addresses the identification problem for face-tracks that are automatically collected from uncontrolled TV video data. Face-track identification is an important component in systems that automatically label characters in TV series or movies based on subtitles and/or scripts: it enables effective transfer of the sparse text-based supervision to other faces. We show that, without manually labeling any examples, metric learning can be effectively used to address this problem. This is possible by using pairs of faces within a track as positive examples, while negative training examples can be generated from pairs of face tracks of different people that appear together in a video frame. In this manner we can learn a cast-specific metric, adapted to the people appearing in a particular video, without using any supervision. Identification performance can be further improved using semi-supervised learning where we also include labels for some of the face tracks. We show that our cast-specific metrics not only improve identification, but also recognition and clustering.
Ramazan Gokberk Cinbis, Jakob Verbeek, Cordelia Schmid
ICCV1
2010 Image Mining Using Directional Spatial Constraints
abstract
Spatial information plays a fundamental role in building high-level content models for supporting analysts' interpretations and automating geospatial intelligence. We describe a framework for modeling directional spatial relationships among objects and using this information for contextual classification and retrieval. The proposed model first identifies image areas that have a high degree of satisfaction of a spatial relation with respect to several reference objects. Then, this information is incorporated into the Bayesian decision rule as spatial priors for contextual classification. The model also supports dynamic queries by using directional relationships as spatial constraints to enable object detection based on the properties of individual objects as well as their spatial relationships to other objects. Comparative experiments using high-resolution satellite imagery illustrate the flexibility and effectiveness of the proposed framework in image mining with significant improvements in both classification and retrieval performance.
Selim Aksoy, Ramazan Gokberk Cinbis
IEEE Geosci. Remote. Sens. Lett.2
2009 Learning actions from the Web
abstract
This paper proposes a generic method for action recognition in uncontrolled videos. The idea is to use images collected from the Web to learn representations of actions and use this knowledge to automatically annotate actions in videos. Our approach is unsupervised in the sense that it requires no human intervention other than the text querying. Its benefits are two-fold: 1) we can improve retrieval of action images, and 2) we can collect a large generic database of action poses, which can then be used in tagging videos. We present experimental evidence that using action images collected from the Web, annotating actions is possible.
Nazli Ikizler-Cinbis, Ramazan Gokberk Cinbis, Stan Sclaroff
ICCV2
2008 Human action recognition with line and flow histograms
abstract
We present a compact representation for human action recognition in videos using line and optical flow histograms. We introduce a new shape descriptor based on the distribution of lines which are fitted to boundaries of human figures. By using an entropy-based approach, we apply feature selection to densify our feature representation, thus, minimizing classification time without degrading accuracy. We also use a compact representation of optical flow for motion information. Using line and flow histograms together with global velocity information, we show that high-accuracy action recognition is possible, even in challenging recording conditions.
Nazli Ikizler, Ramazan Gokberk Cinbis, Pinar Duygulu
ICPR2
2008 Recognizing actions from still images
abstract
In this paper, we approach the problem of understanding human actions from still images. Our method involves representing the pose with a spatial and orientational histogramming of rectangular regions on a parse probability map. We use LDA to obtain a more compact and discriminative feature representation and binary SVMs for classification. Our results over a new dataset collected for this problem show that by using a rectangle histogramming approach, we can discriminate actions to a great extent. We also show how we can use this approach in an unsupervised setting. To our best knowledge, this is one of the first studies that try to recognize actions within still images.
Nazli Ikizler, Ramazan Gokberk Cinbis, Selen Pehlivan, Pinar Duygulu
ICPR2
2008 Automatic Mapping of Linearwoody Vegetation Features in Agricultural Landscapes
abstract
Development of automatic methods for agricultural mapping and monitoring using remotely sensed imagery has been an important research problem. We describe algorithms that exploit the spectral, textural and object shape information using hierarchical feature extraction and decision making steps for automatic mapping of linear strips of woody vegetation in very high-resolution imagery. First, combinations of multispectral values and multi-scale Gabor and entropy texture features are used for training pixel level statistical classifiers for characterizing individual trees and tree groups with respect to their surroundings. Then, decisions based on object level texture features and morphological shape analysis provide the final detection of woody vegetation having a linear structure. Experiments on QuickBird imagery from different sites show that the proposed algorithms provide good localization of linear strips of woody vegetation in different landscapes.
Selim Aksoy, Huseyin Gokhan Akcay, Ramazan Gokberk Cinbis, Tom Wassenaar
IGARSS (4)3
2007 Relative Position-Based Spatial Relationships using Mathematical Morphology
abstract
Spatial information is a crucial aspect of image understanding for modeling context as well as resolving the uncertainties caused by the ambiguities in low-level features. We describe intuitive, flexible and efficient methods for modeling pairwise directional spatial relationships and the ternary between relation using fuzzy mathematical morphology. First, a fuzzy landscape is constructed where each point is assigned a value that quantifies its relative position according to the reference object(s) and the type of the relationship. Then, the degree of satisfaction of this relation by a target object is computed by integrating the corresponding landscape over the support of the target region. Our models support sensitivity to visibility to handle areas that are partially enclosed by objects and are not visible from image points along the direction of interest. They can also cope with the cases where one object is significantly spatially extended relative to others. Experiments using synthetic and real images show that our models produce more intuitive results than other techniques.
Ramazan Gokberk Cinbis, Selim Aksoy
ICIP (2)1